Speech recording integrated control system, method, and program

JP7858121B1Active Publication Date: 2026-05-13竹内祐树 +4
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
竹内祐树
Filing Date
2025-10-22
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing voice recording systems lack versatility, as they are limited to mobile communication terminals and do not effectively integrate and manage speech recordings to enhance evidentiary value, particularly in contexts like hospital consent forms and insurance contracts, and are prone to disputes over what was said and what was not said.

Method used

A speech recording integrated control system that includes a terminal device for voice recording, voice input, speaker authentication, score calculation, recording activation, speech recording, integration, and control units to manage and authenticate speech records, adding a signature to generate evidence, and detect tampering, with features for social trust scoring and submission control.

Benefits of technology

The system provides a highly versatile platform for enhancing the evidentiary value of speech records by integrating voice data with contracts and consent forms, managing and authenticating speech evidence, and handling fraudulent claims, while ensuring secure and reliable submission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007858121000001_ABST
    Figure 0007858121000001_ABST
Patent Text Reader

Abstract

We provide an integrated control system for recording speeches that enables a highly versatile platform and improves the evidentiary value of speech records when creating meeting minutes, various contracts, etc. [Solution] The system is characterized by comprising: an audio input unit that registers the speaker's first voice information in advance in a terminal device equipped with an audio recording function and inputs the speaker's second voice information and terminal information before recording; an integrated authentication unit that performs user authentication; a user authentication score calculation unit; a recording start activation unit that activates the recording start button based on the calculated score value; a speech recording unit that records audio data, transcribed text information, and speaker metadata after the recording start button is activated; an integration unit that integrates session data of the recorded audio information, text information, and metadata; and an integrated control unit that performs integrated control of the processes up to speech preparation, speaker authentication, speaker recording, and score calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a speech recording integrated control system. For example, it integrates voice information spoken by a speaker and text obtained by digitizing the voice information, further generates a speech recording evidence with the speaker's signature added, and relates to a speech recording integrated control system suitable for utilizing the generated speech recording evidence.

Background Art

[0002] Conventionally, even when documents such as an operation consent form in a hospital, an insurance contract form for property insurance in automobile insurance, and a product contract form in product sales are exchanged between a doctor and a patient, an insurance company and an insurance policyholder, and a product salesperson and a customer, respectively, problems such as what was said and what was not said may occur later.

[0003] Also, when creating the above consent forms and contract forms, even if the speech recordings of both parties are saved in a terminal device having a voice recording function, since the information of the speaker himself / herself is not recorded, the evidentiary value as evidence material was not high. Even in this case, disputes may occur between the speakers, and problems such as what was said and what was not said may occur.

[0004] As an example of solving such problems, a voice recording system and a voice recording service method that ensure the evidentiary value of conversation recording data have been proposed (for example, Patent Document 1).

[0005] Patent Document 1 discloses a voice recording system including a mobile communication terminal, a mobile communication base station, a server system connected to a wide area network such as the Internet, an information system of a mobile communication service provider, etc. The server system includes a conversation recording means for recording a user's conversation in digital data via the mobile communication terminal to a storage medium in the server system, a conversation information holding means for holding conversation information such as the time and place where the conversation is made, an authenticity ensuring means for ensuring the authenticity of the digital data of the conversation, and a conversation recording transfer means for transferring the digital data and conversation information, etc. in response to a request from the user.

Prior Art Documents

[0006] [Patent Document 1] Japanese Patent Publication No. 2002-230203 [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] Patent Document 1 describes a method for digitally recording conversation data on a conversation recording management server using a mobile communication terminal, thereby ensuring the authenticity of the digital conversation data and the evidentiary value of the conversation recording data. However, the target audio recording devices are limited to mobile communication terminals such as mobile phones and PHS phones, and dedicated recording devices (e.g., IC recorders and voice recorders), security cameras, and dashcams are not included, which presents a problem in that the system lacks versatility.

[0008] Furthermore, the system is configured such that recording of conversations between sales representatives and customers begins when a mobile communication terminal connects to the conversation recording management server of the conversation recording service center via a mobile communication base station. This presents a challenge in that it is difficult to provide services using various types of voice recording devices and diverse usage patterns.

[0009] To solve the above problems, the object of the present invention is to provide a speech recording integrated control system, method, and program that can realize a highly versatile platform and improve the evidentiary value of speech records by adding the voice data of the person to meeting minutes, various contracts, surgical consent forms at hospitals, automobile damage insurance contracts, etc.

[0010] Another object of the present invention is to provide a speech record integrated control system, method, and program that can appropriately integrate and manage generated speech record evidence and handle fraudulent contracts and "he said, she said" claims. [Means for solving the problem]

[0011] To solve the above problems, the first invention of the present invention provides a speech recording integrated control system that includes: a terminal device equipped with a recording function for recording the voice spoken by the speaker, which in advance registers the speaker's first voice information; a voice input unit that inputs the speaker's second voice information and terminal information to obtain the speaker's permission before recording as preparation for speaking; an integrated authentication unit that compares the second voice information input by the voice input unit with the first voice information to authenticate the speaker; a score calculation unit that calculates a score value from the degree of comparison by the integrated authentication unit; a recording start activation unit that activates a recording start button when the score value calculated by the score calculation unit is above a predetermined threshold; a speech recording unit that, after confirming the activation state of the recording start button, starts recording the speaker and records the voice data, text information obtained by real-time transcription, and the speaker's metadata; an integration unit that integrates the session data of the recorded voice information, text information, and metadata into a file; and an integrated control unit that performs integrated control of the processes up to speaker permission, speaker authentication, speaker recording, and score calculation.

[0012] The speech recording integrated control system according to the second invention of the present invention further comprises, in the first invention, a score calculation unit that calculates a confidence score, a responsibility score, and an emotion score of the speech based on the results of voice feature extraction, language analysis, and emotion estimation of the speech, and a score visualization unit that visualizes the confidence score value, responsibility score value, and emotion score value of the speech calculated by the score calculation unit, wherein the integrated control unit further performs integrated control up to the score calculation and visualization of the score values.

[0013] The third invention of the present invention is a speech recording integrated control system characterized in that, in the first or second invention, it includes a tampering warning unit that adds the speaker's signature to the file to generate speech recording evidence, and also detects and warns of tampering operations.

[0014] The fourth invention of the present invention is a speech recording integrated control system, characterized in that, in the third invention, it includes a speech authenticity evaluation responsibility determination unit that determines whether the content of the speech is a normal speech, a false speech, or an ambiguous speech, and sorts the speech based on the urgency of the speech, reliability, and past history.

[0015] The fifth version of the speech record integrated control system according to the present invention is characterized in that, in the fourth version, it comprises: an evidence output unit that outputs the speech record evidence file in PDF / CSV / ZIP format; a submission destination control unit that processes the speech record evidence file and controls submission depending on whether the submission destination is a court, an administrative agency, or a customer; and a submission destination history recording unit that records the history of the processed speech record evidence file submitted to the submission destination.

[0016] The sixth invention of the present invention provides a speech recording integrated control system, characterized in that, in the fifth invention, it comprises a social trust scoring unit that accumulates and weights all scores, converts them into social trust scores for individuals and / or organizational units, and obtains social trust score values, and a social trust scoring display unit that displays the social trust score values ​​obtained by the social trust scoring unit using a social trust score dashboard.

[0017] The seventh invention of the present invention is a speech recording integrated control system. 6 In the invention A scoring unit that calculates a confidence score, responsibility score, and sentiment score for the aforementioned statement based on the results of voice feature extraction, language analysis, and sentiment estimation of the aforementioned statement, The aforementioned The score calculated by the score calculation unit The system is characterized by comprising a submission priority score generation unit that calculates a submission priority score based on the responsibility score RS, the social trust score STS, and the social impact score SIS using the following formula: PPS = α·RS + β·STS + γ·SIS (where α, β, and γ are coefficients).

[0018] The speech recording integrated control system according to the eighth invention of the present invention is characterized by comprising a submission control unit that, in the seventh invention, automatically submits when the PPS is 85 or more, holds when it is 60 or more but less than 85, and does not allow submission when it is less than 60.

[0019] The speech recording integrated control system according to the ninth invention of the present invention is characterized in that, in the eighth invention, it comprises an update unit that sequentially updates the social trust score STS by the following equation STS_t = STS_{t-1} + Δs ​​- λ·Δt (t: time, Δs: score change value, λ: coefficient, Δt: time change value).

[0020] The speech recording integrated control system according to the tenth invention of the present invention is characterized by comprising a submission destination level database that manages the submission destinations according to confidence levels A to C and E in the ninth invention, and a submission destination control unit that performs stepwise control of the output format, anonymization, and log granularity according to the confidence level.

[0021] The speech recording integrated control system according to the 11th invention of the present invention adds a chain hash to the submission log showing the log information submitted to the submission destination in the 10th invention, and includes ISO8601 time, submitter ID, terminal signature, and submission destination ID. D It is characterized by having an audit unit that holds audit logs, including audit logs.

[0022] The speech recording integrated control system according to the twelfth invention of the present invention is characterized in that, in the eleventh invention, SHA-256 is used for the chain hash, ECDSA for the electronic signature, and RFC3161 compliant TSA for time proof, and includes a sealing part that womens the package.

[0023] The speech recording integration control method according to the 13th invention of the present invention registers the first voice information of the speaker in advance in a terminal device having a recording function for recording the voice spoken by the speaker. A voice input step in which the voice input unit inputs the second voice information and terminal information of the speaker in order to obtain the permission of the speaker before the recording as a speech preparation; An integration authentication step in which the integration authentication unit collates the second voice information input by the voice input step with the first voice information to perform authentication of the speaker; A score calculation step in which the score calculation unit calculates a score value from the degree of collation by the integration authentication step; A recording start activation step in which the recording start activation unit activates a recording start button when the score value calculated by the score calculation step is equal to or higher than a predetermined threshold; A speech recording step in which the speech recording unit starts recording the speaker after confirming the activation state of the recording start button and records voice data, text information by real-time speech-to-text conversion, and meta information of the speaker; An integration step in which the integration unit integrates the session data of the voice information, text information, and meta information recorded as a file; An integration control step in which the integration control unit performs integrated control of the processes up to the speech preparation, authentication of the speaker, recording of the speaker, and calculation of the score value. It is characterized by comprising the above steps.

[0024] The speech recording integrated control program according to the 14th invention of the present invention registers the first voice information of the speaker in advance in a computer having a recording function for recording the voice spoken by the speaker. The computer has a voice input unit for inputting the second voice information and terminal information of the speaker in order to obtain the permission of the speaker before the recording as a speech preparation, an integrated authentication unit for collating the second voice information input by the voice input unit with the first voice information to perform speaker authentication, a score calculation unit for calculating a score value from the degree of collation by the integrated authentication unit, a recording start activation unit for activating a recording start button when the score value calculated by the score calculation unit is equal to or greater than a predetermined threshold value, a speech recording unit for starting the recording of the speaker after confirming the activation state of the recording start button and recording voice data, text information by real-time speech-to-text conversion, and meta information of the speaker, an integration unit for integrating the recorded voice information, text information, and meta information as session data into a file, and is executed as an integrated control unit for performing integrated control of the processes up to the speech preparation, the speaker authentication, the speaker recording, and the calculation of the score value.

Effect of the Invention

[0025] According to the present invention, it is possible to realize a highly versatile platform, and it is possible to add personal voice data to meeting minutes, various contracts, surgical consent forms in hospitals, automobile damage insurance contracts, etc., and improve the evidentiary value of speech records.

[0026] Also, according to the present invention, it is possible to realize a speech recording integrated control system that can appropriately integrate and manage the generated speech recording evidence materials and can handle claims of unauthorized contracts and false statements.

Brief Description of the Drawings

[0027] The drawings show specific embodiments of the present invention according to the present disclosure, including not only essential configurations of the invention but also optional and preferred embodiments. [Figure 1]Figure 1 is a schematic block diagram of a communication network including a speech recording integrated control system to which the present invention is applied. [Figure 2] Figure 2 is a schematic functional block diagram of a speech recording integrated control system illustrating an embodiment of the present invention. [Figure 3] Figures 3(a) to 3(l) are detailed functional block diagrams of the speech recording integrated control system shown in Figure 2. [Figure 4] Figures 4(a) to 4(d) show an example of a terminal having a voice recording function to which the present invention applies. [Figure 5] Figure 5 shows an example of the hardware configuration of the mobile communication terminal shown in Figure 1. [Figure 6] Figure 6(a) shows an example screen of a mobile communication terminal with the speech recording integrated control application installed, and Figure 6(b) shows an example of the rear camera of the mobile communication terminal. [Figure 7] Figure 7 shows an example of the hardware configuration of the management server shown in Figure 1. [Figure 8] Figure 8 shows an example of the hardware configuration of an LLM server. [Figure 9] Figure 9(a) shows an example of a speech preparation screen, and Figure 9(b) shows an example of a speech recording screen. [Figure 10] Figure 10 shows an example of the score analysis results. [Figure 11] Figure 11 shows the overall integrated control flowchart for the speech recording integrated control application. [Figure 12] Figure 12 is a flowchart showing the process for the speaking preparation step. [Figure 13] Figure 13 is a processing flowchart for the speech recording step. [Figure 14] Figure 14 shows the processing flowchart for the score calculation and visualization step. [Figure 15] Figure 15 shows the processing flowchart for the tamper detection and warning step. [Figure 16] Figure 16 shows the processing flowchart for the step of evaluating the veracity of a report and determining responsibility. [Figure 17] Figure 17 is a flowchart of the process for the evidence output submission step. [Figure 18]Figure 18 shows the processing flowchart for the social trust score visualization step. [Figure 19] Figure 19 shows the processing flowchart for the role control access restriction step. [Figure 20] Figure 20 is a flowchart illustrating a speech recording integrated control method that shows another embodiment of the present invention. [Figure 21] Figure 21 is a flowchart showing a series of processes from evidence generation to submission of evidence in an embodiment of the present invention. [Figure 22] Figure 22 is a verification flowchart of the evidence package in an embodiment of the present invention. [Modes for carrying out the invention]

[0028] The embodiments illustrated below, applying the present invention, will be described with reference to the drawings.

[0029] <Embodiment 1> <Overview of Embodiment 1> The first embodiment of the speech recording integrated control system is characterized by comprising: a terminal device equipped with a recording function for recording the speaker's voice, which has a first voice information of the speaker registered in advance; a voice input unit which inputs the speaker's second voice information and terminal information in order to obtain the speaker's permission before recording as preparation for speaking; an integrated authentication unit (integrated authentication engine) which compares the second voice information input by the voice input unit with the first voice information to authenticate the speaker's identity; a score calculation unit which calculates a score value from the degree of comparison by the integrated authentication unit; a recording start activation unit which activates a recording start button when the score value calculated by the score calculation unit is above a predetermined threshold; a speech recording unit which, after confirming the activation state of the recording start button, starts recording the speaker and records the voice data, text information by real-time transcription, and the speaker's attribute information; an integration unit which integrates the session data of the recorded voice information, text information, and metadata into a file; and an integrated control unit which performs integrated control of the processes up to the preparation for speaking, the authentication of the speaker's identity, the recording of the speaker, and the calculation of the score value. This enables the creation of a highly versatile platform and a speech recording integrated control system that enhances the evidentiary value of speech records by adding the individual's voice data to meeting minutes, various contracts, surgical consent forms at hospitals, automobile insurance contracts, etc.

[0030] Figure 1 is a schematic block diagram of a communication network system including a speech recording integrated control system to which the present invention is applied. Figure 2 is a schematic functional block diagram of a speech recording integrated control system showing an embodiment of the present invention, and Figure 3 is a detailed functional block diagram of the speech recording integrated control system of Figure 2. Figure 3(a) is a functional block diagram of the speech preparation unit, Figure 3(b) is a functional block diagram of the speech recording unit, Figure 3(c) is a functional block diagram of the score calculation visualization unit, Figure 3(d) is a functional block diagram of the report authenticity evaluation unit, Figure 3(e) is a functional block diagram of the AI ​​responsibility degree branching unit, Figure 3(f) is a functional block diagram of the evidence output unit, Figure 3(g) is a functional block diagram of the submission destination control unit, Figure 3(h) is a submission destination history recording unit, Figure 3(i) is a functional block diagram of the tampering warning unit, Figure 3(j) is a functional block diagram of the access control unit, Figure 3(k) is a functional block diagram of the social trust score display unit, and Figure 3(l) is a functional block diagram of the dashboard unit. Figure 4 shows an example of a terminal having an audio recording function to which the present invention applies. Figure 4(a) shows an example of a mobile communication terminal, Figure 4(b) shows an example of an IC recorder, Figure 4(c) shows an example of a security camera, and Figure 4(d) shows an example of a drive recorder. The speech recording integrated control system will be described below using Figures 1 to 4.

[0031] In Figure 1, the communication network system is connected to a mobile communication terminal 1, an IC recorder 2, a security camera 3, a drive recorder 4, a speech recording integrated control system 5, and a cloud 6. The management server 7, LLM server 8, and submission terminal 9 are also connected to the cloud 6. The communication network connection method is not limited to this; it may also be connected via a mobile communication line (4G LTE, 5G). It may also be connected via WiFi® or Bluetooth®.

[0032] The mobile communication terminal 1 can be a smartphone or other mobile communication terminal, a tablet PC, or a touch-input computer. As shown in Figures 4(a) to 4(d), the terminals with voice recording functionality to which the present invention applies include the mobile communication terminal 1, IC recorder 2, security camera 3, and drive recorder 4, enabling a highly versatile platform. Specifically, since the terminals with voice recording functionality need to connect to the management server 7, they have internet connectivity, and their communication capabilities include WiFi®, 4G LTE, 5G, or a combination thereof. The mobile communication terminal 1, like a computer, offers a high degree of freedom in the applications that can be installed. However, the IC recorder 2, security camera 3, and drive recorder 4 are dedicated devices, so they upload speech to the management server 7 by partially adding functionality to generate speech recording evidence. In the case of terminals that cannot communicate bidirectionally, data verification of the management server 7 is performed by the mobile communication terminal 1. In the case of IC recorder 2, recording can be started when the speaker gives the instruction to start recording. However, with security camera 3, a security schedule is set in advance, and surveillance is performed continuously or during specific times such as at night. Therefore, recordings are made according to the security schedule, and evidence of speech (sound generation) is generated.

[0033] This section will explain how to set up time-based schedules for security cameras, using specific examples. <Example of a typical time-slot schedule setting> <Residential settings> Weekday schedule: 6:00-8:00 (morning) - Standard recording mode, motion detection ON Family members' commuting / school schedules 8:00-17:00 (Daytime) - High sensitivity mode, instant notification ON This is a time when people are away, so the highest level of vigilance is necessary. 17:00-23:00 (Evening to Night) - Standard recording mode, motion detection ON Commuting time 23:00-6:00 (late night) - High sensitivity mode, infrared shooting, instant notification ON The risk of intrusion is high during sleeping hours. Weekend schedule: Because I'm spending more time at home, I've reduced the notification frequency and am continuing to record in standard mode.

[0034] <Settings for stores and offices> Business days: 7:00-9:00 (Opening preparations) - Standard recording, entry / exit records 9:00-18:00 (Business Hours) - Continuous recording, high-resolution recording around the cash register. 18:00-20:00 (Closing procedures) - Standard recording, entry / exit records 20:00-7:00 (after closing) - High-sensitivity mode, instant notification upon motion detection, security alarm linked. Closed days: High sensitivity mode all day, records and notifies of all movements.

[0035] <Settings for parking lots> 6:00-22:00 (Daytime) - Standard recording mode Peak hours 22:00-6:00 (Nighttime) - High-sensitivity mode, specialized for vehicle detection Time of day when the risk of theft and vandalism is higher.

[0036] <Key points for settings> Adjusting recording quality: The critical time period is high resolution and high frame rate. Standard definition is used during normal hours to save storage space. Notification settings: Immediate notification when absent When you're at home, suppress notifications to prevent false alarms. Storage period: Long-term storage during important events Normally, rotations occur every 1-2 weeks. By flexibly scheduling operations to match lifestyle patterns and business hours, effective security and efficient operation become possible.

[0037] <Remote start function for dashcam> This section explains the remote start function of a dashcam. Remote start of a dashcam refers to a function that automatically starts recording based on specific conditions or triggers, without the driver having to manually operate it. Main types 1. Automatic recording triggered by impact detection. The G-sensor automatically starts recording when it detects an impact above a certain level. Ensure that important evidence is recorded in the event of an accident or hit-and-run. 2. Parking surveillance mode Monitoring continues even after the engine is shut off. Automatic recording when impact or motion is detected on the vehicle. Battery drain prevention features are common. 3. Smartphone connectivity Start recording remotely using the dedicated app. Operated via Wi-Fi / Bluetooth Some models also allow for real-time video monitoring. 4. Voice commands Operated by voice commands such as "Start recording". Safe to operate even while driving.

[0038] In the case of terminals that cannot communicate bidirectionally, for example, in the case of IC recorder 2, the management server 7 sends the recorded data and transcript data from IC recorder 2 to the mobile communication terminal 1 for verification. Alternatively, the management server 7 could send the data by email after recording is finished for verification, or the recorder could connect to the management server 7 via cloud connection for verification.

[0039] In the case of drive recorder 4, the power to drive recorder 4 can be turned on when the engine is started, or recording can be started by the drive recorder monitor operator via remote monitoring. Drive recorder 4 is configured to send still images and videos of the camera footage to management server 7.

[0040] When offline, encrypted data is temporarily stored on the device. After the connection is restored, the TSA is obtained, recalculated, and synchronized.

[0041] In Figure 2, the speech recording integrated control system 5 includes a speech preparation unit 21, a speech recording unit 22, a score calculation visualization unit 23, a report authenticity evaluation unit 24, an AI responsibility degree branching unit 25, an evidence output unit 26, a submission destination control unit 27, a submission destination history recording unit 28, a tampering warning unit 29, an access control unit 30, a social trust score display unit 31, and a dashboard unit 32. When these functions are applied to a mobile communication terminal (smartphone), they can be realized simply by installing the speech recording evidence generation and integrated control application on the standard hardware of the mobile communication terminal (iPhone® or Android® terminal), without the need for any special machinery or equipment. However, for processing with a high processing load, the management server 7 also has similar functions and executes the processing on the management server 7. In this case, the speech recording integrated control application has an interface to display the processing results, and on the mobile communication terminal side, recording can be started and stopped, and simple text corrections can be made. The speech recording integrated control system 1 may be implemented not only in software but also in hardware configuration.

[0042] The speech recording integrated control system 5 will be described below with reference to Figures 3(a) to 3(l).

[0043] The speech preparation unit 21 prepares the speaker's speech and verifies their identity before recording permission is granted.

[0044] The speech recording unit 22 performs audio data recording, real-time transcription, and metadata recording.

[0045] The score calculation and visualization unit 23 calculates and visualizes the credibility, responsibility, and emotional scores of statements.

[0046] The report authenticity evaluation unit 24 evaluates the authenticity of the report based on the confidence score and responsibility score.

[0047] The AI ​​responsibility branching unit 25 processes information such as warnings, blocking, and evidence creation according to the responsibility score. For legitimate reports, evidence is saved. For false reports, the report is blocked and a warning is recorded. For ambiguous reports, the information is put on hold and its contents are recorded.

[0048] The evidence output unit 26 outputs the generated evidence files in PDF / CSV / ZIP output format. It also selects the recipient and records the log.

[0049] The submission destination control unit 27 performs control according to the submission destination, such as courts, administrative agencies, and customers. If the submission destination is domestic, it complies with the Personal Information Protection Act, and if the submission destination is an international organization, it provides protection equivalent to, for example, GDPR and complies with the "anonymization and retention period" standards. GDPR is an abbreviation for General Data Protection Regulation, which is the EU's General Data Protection Regulation. Submission destinations are managed by trust levels A to C and E, with A being a signed + TSA-protected ZIP file, B being read-only, C being automatically masked for personal identifiers and location, and E being non-submission. The granularity of the submission log is controlled in stages corresponding to each level. The submission log includes at least {ts, submitter_id, device_sig, dest_id, ip, user_agent, view_count, reexport_count, pkg_hash}, and chain hashing is used for tamper detection.

[0050] Security measures will be implemented to detect tampering. Key management will be performed using signing keys, verification keys, and role keys. Hardware Security Module (HSM) / secret sharing, key rotation, and certificate revocation processing will be performed. To address NTP / TS dependency and drift, clock synchronization and time authenticity (timestamp) will be added, and verification procedures will follow server time source, handling of offline terminals, and TS priority. Anonymization and masking will be performed by setting rule-based / ML masking for proper nouns, locations, and IDs, and differential privacy thresholds. Evidence files will be processed and submitted in a format that satisfies the requirements of the recipient. In addition, the recipient control unit 27 will perform limited access and history saving, etc.

[0051] The submission history recording unit 28 records evidence output, evidence submission, and evidence manipulation in its history. The "schema definition" of the submission log includes ISO8601 date and time, submitter ID, terminal signature / evidence ID, submission destination, IP / UA, number of views / re-outputs, and hash.

[0052] The tampering warning unit 29 detects and reports any attempts to tamper with evidence.

[0053] The access control unit 30 identifies the speaker, the administrator of the management server 7, and external recipients (e.g., courts, administrative agencies), and restricts display and output based on their roles.

[0054] The social trust score display unit 31 displays an accumulated graph of trust score and responsibility score. This display is provided by the social trust dashboard.

[0055] The dashboard section 32 displays all records, scores, submissions, history, and warning logs in one place.

[0056] Figure 5 shows an example of the hardware configuration of mobile communication terminal 1.

[0057] In Figure 4, the computer configuration of the mobile communication terminal 1 includes a CPU 11, non-volatile memory (ROM) 12 such as an HDD and ROM, main memory (RAM) 13 such as D-RAM, a display 17, a keyboard (software keyboard) 18, a communication interface 16, a camera 14, a microphone 19 as an audio input unit, and a speaker 15 as an audio output unit, all connected to a system bus. Access by the mobile communication terminal is selectively connected to a network as appropriate depending on the communication environment, such as 3G, 4G, 5G mobile wireless network services or WiFi® wireless connection. The above configuration is typical for the hardware configuration of a mobile communication terminal, so a detailed explanation will be omitted here.

[0058] The mobile communication terminal 1 can log in to the management server 7 via a communication network 10 such as the internet through the communication interface 16. Login authentication is performed using the email address and password used during registration. If the customer registered with an SNS email address, login authentication is performed using the email address and password of that SNS.

[0059] Figure 6(a) shows an example screen of a mobile communication terminal with the speech recording integrated control application installed, and Figure 6(b) shows an example of the rear camera of a mobile communication terminal.

[0060] In Figures 6(a) and 6(b), the mobile communication terminal 50 includes a touchscreen display screen 51, an icon for a speech recording integrated control application 52, a front camera 53a, rear cameras 53b and 53c, a microphone 54, a speaker 55, and a flash 56.

[0061] Figure 7 shows an example of the hardware configuration of the management server 7 shown in Figure 1. In Figure 7, the management server 7 includes a communication control unit 71, a reception unit 72, a program processing unit 73, a storage unit 74, a display unit 75, an integrated authentication engine 76, a speech record database 77, an evaluation interpretation logic / evaluation criteria database 78, a submission history / management database 79, a prompt database 80, and an output unit 81. The evaluation interpretation logic / evaluation criteria database 78 stores the verification criteria used when updating the model. The prompt database 80 stores settings such as prompt templates, prohibited outputs, and rate control. The prompt database 80 works in conjunction with the LLM server 8 and acts as a safety valve. Table 1 shows an example of the criteria for the Social Trust Score (STS).

[0062] [Table 1]

[0063] Table 2 shows an example of the criteria for the Responsibility Score (RS). The evaluation system consists of points, ratings, symbols, and meanings, and assesses the logic, consistency, and composure of statements and reports. The evaluation criteria are shown on a scale of 0 to 100 points, with higher scores indicating a more "responsible statement."

[0064] [Table 2]

[0065] An example of a Social Impact Score (SIS) is shown in Table 3. The Social Impact Score (SIS) indicates the risk, ripple effect, and urgency of a reported incident on society. The evaluation criteria are expressed on a scale of 0 to 100 points and are automatically calculated by AI.

[0066] [Table 3]

[0067] Table 4 shows an example of a Priority Processing Score (PPS). The priority score PPS is calculated using the following formula: PPS = α×RS + β×STS + γ×SIS

[0068] [Table 4]

[0069] Management Server 7 has a database table "matching" and a UI / flow mapping table, and performs mapping between primary keys and main columns such as reports / scores / submissions / logs / feedback and flows (Figures 11-20). Table 5 shows an example of the main table names and main columns of the database tables.

[0070] [Table 5]

[0071] The voice data from the speaker is stored in various databases on the management server 7, and real-time AI learning can be performed. The AI ​​learning results can also be used for authenticating past speakers. In addition, AI requests can be made from the management server 7 to the LLM server 8 using the prompt database 80 and the evaluation interpretation logic / evaluation criteria database 78, and analysis can be performed by the generating AI.

[0072] Figure 8 shows an example of the hardware configuration of the LLM server 8. In Figure 8, the server includes a processor 83, an input unit 84, a large-scale language model 85, a storage unit 86, a general-purpose generative AI module 87, an output unit 88, and a communication control unit 89.

[0073] Figure 9(a) shows an example of the user interface (UI) screen for preparing to speak, and Figure 9(b) shows an example of the user interface screen for recording a speech. In addition to Figures 9(a) and (b), symbols and color displays as shown in Tables 2 and 4 may also be used. As for authentication methods, the mobile communication terminal supports multi-factor authentication such as facial recognition, voice recognition, and fingerprint recognition. In this embodiment, voice (voiceprint) authentication is particularly used to block requests from fraudulent requesters and prevent unnecessary requests. In addition, there are input buttons for terminal information and a user authentication button. A user authentication score is calculated, and if the score is 85 points or higher, the recording start button is activated. If the score is less than 85 points, the recording start button is not activated, and recording is stopped at the initial stage. Table 6 shows an example of a threshold corresponding to the overall score of the system.

[0074] [Table 6]

[0075] Figure 9(b) shows an example of the user interface screen during speech recording. The speaker's recording is saved to the management server 7 via the cloud 6, the audio data is recorded, and transcription is performed in real time. As a result, the text is displayed in the text display area of ​​the mobile communication terminal. The speaker ID, time, and GPS location information are also displayed. When the audio playback button is pressed, the audio linked to the text is played. The text is not 100% accurate, so the speaker may make corrections. Meeting minutes can be created with real-time display and immediate corrections, or they can be done after the meeting has ended. As evidence of speech recording, it is necessary to identify the speaker at the time the meeting is held. If there are new meeting participants, their voices should be registered first.

[0076] Figure 10 shows an example of a user interface screen for score analysis of trust score, responsibility score, and sentiment score. In addition to Figure 10, symbols and color displays as shown in Tables 2 and 4 may also be used. Trust score, responsibility score, and sentiment score are displayed as indicators from 0 to 100, but different colors may be used for each score.

[0077] Figure 11 is an overall control flowchart of the speech recording integrated control system illustrating this embodiment. Here, Figure 11 shows an overall flow as an example, and each step does not necessarily have to be performed in this order, and can be performed in any order, as long as there are no problems with the processing. Figures 12 to 19 are detailed flowcharts of each step in Figure 11. The processing method of the speech recording integrated control system will be described below using Figures 11 to 19.

[0078] The speech recording integrated control system 5 includes a speech preparation step 101, a speech recording step 102, a score calculation and visualization step 103, a tampering detection and warning step 104, a report authenticity evaluation and responsibility determination step 105, evidence output and submission step 106, social trust scoring and visualization step 107, role control and access restriction step 108, and integrated control and termination processing step 109.

[0079] <Preparation Step 101 for Speaking (Figure 12)> Step 101 of preparing to speak is performed by activating the speaking preparation UI on the user authentication input screen before recording permission is granted. Specifically, as shown in Figure 12, the speaking preparation UI is activated to perform speaking preparation and start user authentication (step 201). In the speaking preparation UI as shown in Figure 9(a), the speaker's face / voice / fingerprint / terminal information is entered (step 202). In this embodiment, biometric authentication of the user is performed by authenticating the speaker's voice (voiceprint). When the mobile communication terminal 1 and the management server 7 are connected via the cloud 6, selecting "Voice" in the speaking preparation UI and executing the "Authenticate" button will display a message such as "Please enter the speaker's voice," and when the speaker "speaks," verification of the speaker's voiceprint, which has been registered in advance on the management server 7, is performed. Verification is performed by the integrated authentication engine 76 shown in Figure 7. Here, as a countermeasure against GPS and voice spoofing, GPS deception detection (speed, signal strength, wireless fingerprint matching, threshold setting for voiceprint + speaker conversion detection, Wi-Fi / BLE matching) and voice deepfake indicators are employed. The integrated authentication engine 76 performs voiceprint matching and calculates a score (step 202). If a score of 85 points or higher out of a possible 100 points is output (step 203), authentication is successful. Subsequently, the recording start button shown in Figure 9(a) is activated (step 204), and recording of the meeting or other conversation begins. This state is shown in Figure 9(b).

[0080] <Speech Recording Step 102 (Figure 13)> Step 102 of the speech recording process involves recording, real-time transcription, and metadata recording. Specifically, in the speech recording UI shown in Figure 9(b), recording begins, and the speaker's recording (recording of audio data) is performed (Step 301). Real-time transcription (Step 302) and metadata recording are then performed. Next, the speaker ID, time, location (GPS), and terminal log are recorded (Step 303). Finally, session data integration of audio information, text information, and metadata is performed (Step 304). From the perspective of protecting personal information, location information and personal information are minimized, and consent management is implemented. Minimization of metadata acquisition items, consent logs, and a deletion request (DSR) workflow are incorporated. When offline, encrypted data is temporarily stored on the device. After the connection is restored, the TSA is obtained, recalculated, and synchronized.

[0081] <Score calculation visualization step 103 (Figure 14)> Step 103 of the score calculation and visualization process launches the score analysis UI to calculate the trust score, responsibility score, and sentiment score of the statement, and visualizes the trust score, responsibility score, and sentiment score of the statement. In this embodiment, the submission priority score (PPS) is defined as PPS = α·RS + β·STS + γ·SIS, using the responsibility score RS, social trust score STS, and social impact score SIS. For example, if α=0.3, β=0.3, and γ=0.4, then automatic submission is performed when PPS ≥ 85, it is put on hold when 60 ≤ PPS < 85, and submission is not permitted when PPS < 60.

[0082] Specifically, the following processing is performed in score calculation visualization step 103. First, the speaker's speech (audio data) input via the microphone of the mobile communication terminal 1 is sent to the integrated authentication engine 76 of the management server 7 via the cloud 6. The integrated authentication engine 76 of the management server 7 performs real-time transcription from the received audio data and obtains text information and speaker metadata (meta information) (step 401).

[0083] Next, the integrated authentication engine 76 extracts speech features such as tone, intonation, and speaking speed from the speaker's voice data (step 402). Furthermore, the integrated authentication engine 76 performs linguistic analysis on the speaker's voice data, identifying imperatives, ambiguous words, assertive words, etc. (step 403).

[0084] Next, the integrated authentication engine 76 uses the voice feature extraction results and language analysis results to estimate the speaker's emotions, such as anger, tension, and sincerity (step 404). Based on these analysis results and emotion estimation results, it calculates a trust score, responsibility score, and emotion score (step 405). The trust score, responsibility score, and emotion score values ​​from these calculations are determined in the range of 0 to 100 using a predetermined score calculation formula. These scores are visualized numerically and in color on the user interface (Figure 10 shows an example of indicator display for the trust score, responsibility score, and emotion score) and displayed on the screen of the mobile communication terminal 1 (step 406). Here, the "feature list and model example" for RS / trust, responsibility, and emotion scores is used to create a learning model from the feature list of voice (MFCC (Mel-frequency cepstrum coefficient), F0 (fundamental frequency), speech rate, silence rate, etc.) / language (imperative words, negation, ambiguity) / estimator (SVM / GBDT / NN) / evaluation index (AUC, etc.) and perform voice authentication.

[0085] <Tampering detection warning step 104 (Figure 15)> Tamper detection warning step 104 activates the tamper warning UI to detect and report tampering operations. Specifically, the following processing is performed in the tamper detection warning step 104. The management server 7 performs hash encryption, adds a signature, and attaches a TSA (Time-Stamping Authority) to the speech record evidence file generated by the above processing (step 501). The hash is SHA-256, the digital signature (ECDSA), and the package to which the RFC3161 compliant TSA (Time-Stamping Authority) time certificate is attached is sealed with the WORM attribute, the clock is synchronized with NTP, and the TS time is prioritized for verification.

[0086] The branching process for exceptions in this embodiment will now be described. In the event of a TSA failure, signature failure, or inability to verify, the system will respond according to the rules for retrying, isolation, user notification, and non-submission of evidence. For example, the number of retries may be set to 3, and if it exceeds 3, the data will be isolated and saved, the user will be notified, and a decision will be made that submission is not possible.

[0087] Next, the monitoring module detects any tampering with the speech record evidence file (step 502). Security measures are taken to detect tampering. Key management is performed using signing keys, verification keys, and role keys. Hardware Security Module (HSM) / secret sharing, key rotation, and certificate revocation are performed. As a countermeasure against NTP / TS dependency and drift, clock synchronization, time authenticity (timestamp) is added, and verification procedures follow, including server time source, handling of terminal offline situations, and TS priority. The tampering flag is turned ON, the notification judgment UI is displayed (step 503), automatic containment processing is performed, and read-only mode is implemented (step 504).

[0088] <Evaluating the veracity of a report and determining responsibility: Step 105 (Figure 16)> Step 105, which involves evaluating the authenticity of a report and determining responsibility, activates the report judgment UI, the AI ​​responsibility level branching UI, and the report x responsibility matrix UI. It then performs a report judgment based on the authenticity, reliability, and urgency of the statement, processes such as warnings, blocking, and evidence creation based on the responsibility score, and displays a four-quadrant judgment of the responsibility score x report appropriateness as shown in Table 7 below.

[0089] [Table 7]

[0090] Table 8 shows an example of a threshold corresponding to the overall score of the system. [Table 8]

[0091] Specifically, in the report authenticity evaluation and responsibility determination step 105, the following processing is performed. The management server 7 performs score calculation on the report content using the report UI × responsibility matrix UI (step 601). Next, it determines whether the report is a normal report, a false report, or an ambiguous report, and performs a matrix evaluation of the report validity and responsibility score (step 602). The social impact score (SIS) is normalized to 0 to 100 using the threat word intensity s_threat, target scale s_scale, secondary damage s_secondary, and domain weight w_domain as SIS = norm( a·s_threat + b·s_scale + c·s_secondary ) · w_domain (a, b, c are calibration parameters). In this embodiment, the submission priority score (PPS) is defined as PPS = α·RS + β·STS + γ·SIS, using the responsibility score RS, social trust score STS, and social impact score SIS. For example, let α=0.3, β=0.3, γ=0.4. If PPS ≥ 85, it will be automatically submitted; if 60 ≤ PPS < 85, it will be held; and if PPS < 60, it will not be submitted. If the above determination is "ambiguous report," the conditions for holding the report and the reconfirmation SLA will be shown, and the score will be recalculated and reconfirmed according to the holding period, reevaluation trigger, and PPS / STS recalculation conditions.

[0092] Based on the matrix evaluation results, a classification is performed based on urgency, reliability, and past history (Step 603). The Social Trust Score (STS) is calculated using the following formula, where STS is the score change value, Δs is the score change value, and t is time. STS(t) = STS(t-1) + Δs Here, the upper limit of the score is set to 100, and the lower limit to 0 (the cutoff line). Depending on whether the report was legitimate, valid, or false, the amount will be adjusted as shown in Table 9 below.

[0093] [Table 9]

[0094] If the Priority Processing Score (PPS) is defined as PPS, and the responsibility score RS, trust score STS, social impact score SIS, α, β, and γ are used as coefficients, the priority processing score PPS can be calculated using the following formula. PPS = α * RS + β * STS + γ * SIS For example, α = 0.3, β = 0.3, γ = 0.4 (the most important factor is social impact). A higher priority score (PPS) means "immediate reporting, immediate submission, and priority processing." Based on the results of the report, various scores are calculated by adding or subtracting score fluctuation values. The social impact score (SIS) is linked to the priority score (PPS), and the cases are categorized based on urgency, reliability, and past history. In step 603, if the AI ​​responsibility level branching UI determines that the report is highly responsible and legitimate, the evidence is saved, processed according to the recipient, and automatically submitted (step 604). If the AI ​​responsibility level branching UI determines that the report is false, the report is rejected and recorded (step 605). In other words, if the report is false, the report is blocked and a warning is recorded. If the AI ​​responsibility level branching UI determines that the report is ambiguous, it is put on hold and a record of reconfirmation is made (step 606).

[0095] <Evidence Output / Submission Step 106 (Figure 17)> Step 106, which involves outputting and submitting evidence, activates the evidence output UI and submission control UI to output and submit evidence. It also implements controls (such as limited access and history saving) according to the recipient, based on the trust level.

[0096] Specifically, the following processes are performed in the evidence output / submission step 106. On management server 7, if the score visualization shows that the score is above the threshold, the score is deemed normal, and the evidence output UI outputs the evidence materials in PDF / CSV / ZIP format (step 701). The evidence package includes audio, RTT, metadata, hash list, signing certificate, TS token, and verification README within the ZIP file. The receiving party performs signature verification, TS token verification, and hash consistency using the public key according to the attached README. Anonymization and masking are performed by setting thresholds for rule-based / ML masking of proper nouns, locations, and IDs, and differential privacy.

[0097] When an output operation occurs via the evidence output UI, if the evidence output is of the highest level, for example, for submission to a court or auditing body, a signature + TSA + log information is added to the evidence output (step 702), and the recipient (court, auditing body, etc.) is selected (step 703). Here, the recipient is managed by trust levels A to C and E, A is a signed + TSA-attached ZIP file, B is read-only, C is automatically masked for personal identifiers and location, and E is not submittable. The granularity of the submission log is controlled in stages corresponding to each level. The submission log includes at least {ts, submitter_id, device_sig, dest_id, ip, user_agent, view_count, reexport_count, pkg_hash}, and chain hashing is used for tamper detection. Security measures are taken for tamper detection. Key management is operated using signing keys, verification keys, and role keys. HSM (Hardware Security Module) / secret sharing, key rotation, and certificate revocation processing are performed. As a countermeasure against NTP / TS dependency and drift, clock synchronization, time authenticity (timestamp) is added, and verification procedures follow, including server time source, handling of terminal offline situations, and TS priority.

[0098] The submission history recording UI records the history of evidence output and submission operations in the submission destination history / management database 79 (step 704). The "schema definition" of the submission log includes ISO8601 date and time, submitter ID, terminal signature / evidence ID, submission destination, IP / UA, number of views / re-outputs, and hash. The input / output, signature, and retry API definitions include request / response examples for the submission API / audit log API / verification API, signature method, and timeout. A chain log is created from evidence acquisition → evidence sealing → evidence submission → evidence viewing → evidence re-output, and a signature chain is added to maintain the continuity of evidence preservation.

[0099] <Social Trust Score and Visualization Step 107 (Figure 18)> Step 107 of the Social Trust Score (STS) calculation and visualization process displays the Social Trust Score (STS) dashboard and accumulates and visualizes the score based on behavioral history. If a re-evaluation flow or objection flow is applicable, a recalculation is performed based on the log evidence. Specifically, a recalculation request for PPS and STS is sent, approved by the auditor, and the historical difference of the results is saved.

[0100] Specifically, in step 107 of the social trust scoring and visualization process, the following processing is performed: On the management server 7, upon completion of processing, all scores are accumulated and weighted (step 801). The social trust score (STS) is updated sequentially. The STS is updated based on the time series as STS_t = STS_{t-1} + Δs ​​- λ·Δt. Δs is assigned according to the truthfulness and severity of the event, for example, truthfulness = True: +5 to +10, False: -10 to -30. The social trust score dashboard is used to score STS at the individual / organizational level (step 802). Here, the "sequential update rule" and the addition / deducting ranges for the social trust score STS are explained. STS(t) = STS(t-1) + Δs ​​(True: +5 to +10 / False: -10 to -30, etc.) and time decay are performed. The display screen of mobile communication terminal 1 shows a timeline, history graph, and notifications (step 803).

[0101] <Roll control / access restriction step 108 (Figure 19)> Role control and access restriction step 108 activates the access control UI (RBAC) and restricts viewing, outputting, and operations according to permissions. Access restrictions are performed using a "role x operation matrix." For example, it is based on a matrix evaluation of the permissions of the roles of user / administrator / auditor / external recipient and operations such as viewing scope / output permission / re-output permission. Restrictions are applied to user / administrator / auditor / external recipient × viewing scope / output permission / re-output permission. Table 100 shows the judgment axes, explanations, and functional significance of role control and access control.

[0102] [Table 10]

[0103] Specifically, in the role control / access restriction step 108, the following processing is performed. The management server 7 uses the access control UI to identify users such as the user, administrator, and external submission destination (step 901), and controls the displayed items and output permissions according to the operation privileges (step 902). For example, as shown in Table 11, if the user is the user, output control is performed such as viewing their own cases = OK / output = up to C / re-output = requires approval. Similarly, if the user is an auditor, output control is performed such as viewing all cases = OK / output = NOT OK / viewing logs = OK.

[0104] [Table 11]

[0105] <Matrix Evaluation> The following provides specific examples of matrix evaluation. Administrator × High Score (RBAC: Admin, STS: 95) • Access to all data → Full access privileges are available for investigation, submission, and log viewing. General Staff × Medium Score (RBAC: Staff, STS: 72) • Viewable: Yes (partially masked) Output: Requires approval (after passing through administrator workflow) →The minimum necessary information is provided when handling reports. Reporter (Self) × Low Score (RBAC: Self, STS: 44) • Access: Blocked / Sealed • Re-output: Not possible / Notification with reason for sealing →Having a prior record of making false reports, this is restricted due to the high social risk.

[0106] Then, all operation logs are recorded in the submission history / management database 79 (step 903). If a re-evaluation flow or objection flow is applicable, a recalculation is performed based on the log trail. That is, a recalculation request for PPS and STS is sent, auditor approval is obtained, and the historical difference of the results is saved. The "retention period," "tamper resistance," and "inquiry UI" items of the audit log are also saved, and the retention period, hash chain, and query / export API are recorded.

[0107] <Integrated control / termination processing step 109> The integrated control and termination process step 109 activates the final integrated control flow UI and performs integrated control and termination processing for the steps of speech → scoring → evidence creation → submission → social trust scoring. Specifically, the integrated flow controls the steps from speech → authentication → recording → scoring → evidence → submission → trust visualization.

[0108] This section provides a brief explanation of UI (User Interface) screen transitions. First, the user prepares to speak (user authentication). If the score is 85 points or higher, authentication is successful, and the user starts the speech recording UI (recording + transcription). Recording and real-time transcription are performed, and once recording is complete, the score visualization UI (trust / responsibility / emotion) is activated, and the score is visualized. If the score falls below the threshold, it is considered an abnormal score.

[0109] Next, the report authenticity evaluation UI is launched, and the AI ​​responsibility level branching UI makes a report based on the trust and responsibility score. A normal report is saved as evidence and stored on management server 7. A false report is blocked and a warning is recorded. An ambiguous report is put on hold and recorded. The process proceeds to the next step depending on the report result.

[0110] If the score is normal, the evidence output UI will initiate an output operation to export the evidence files in PDF / CSV / ZIP format.

[0111] The submission destination-specific control UI allows for the following: If the submission destination is a court, the file will be ZIP-bound, TSA-registered, and signed. If the submission destination is a government agency, the output will be partially restricted. If the submission destination is a customer service center, a limited link and a lightweight PDF output will be provided.

[0112] If the submission history record UI (output log) detects an attempt at tampering or leakage, the tampering warning UI will automatically perform containment processing (read-only mode). If unauthorized access or external viewing operations are detected, the access control UI (RBAC) will restrict display and output based on roles.

[0113] Upon completion, the Social Trust Score UI (STS) will display an accumulated graph of the trust score and responsibility score.

[0114] When accessed by administrators / auditors, the administrator dashboard UI displays all records, scores, submission history, and warning logs in one place. Here, the UI design completes the process from "trust generation → tamper prevention → responsibility visualization → automatic judgment → social integration" in a single series of screen transitions. All evidence output is standardized as TSA + signature + hashed ZIP. In the reporting judgment, blocking / warning / evidence preservation is automatically branched based on the score, so the user interface also has a rule-based linked structure.

[0115] <Other Embodiments> Figure 20 is a flowchart illustrating a speech recording integrated control method that shows another embodiment of the present invention.

[0116] The speech recording integrated control method shown in Figure 20 includes a voice input step 1001, an integrated authentication step 1002, a score calculation step 1003, a recording start activation step 1004, a speech recording step 1005, an integration step 1006, and an integrated control step 1007. Here, the speaker's first voice information is registered in advance in a terminal device equipped with a recording function for recording the speaker's voice.

[0117] In the voice input step 1001, the voice input unit inputs the speaker's second voice information and terminal information to obtain permission from the speaker before recording as preparation for speaking. Next, in the integrated authentication step 1002, the integrated authentication unit compares the second voice information input in the voice input step 1001 with the first voice information to authenticate the speaker's identity.

[0118] Next, in the score calculation step 1003, the score calculation unit calculates a score value from the degree of matching in the integrated authentication step 1002. Next, in the recording start activation step 1004, the recording start activation unit activates the recording start button if the score value calculated in the score calculation step 1003 is above a predetermined threshold (for example, 85 points or more).

[0119] Next, in the speech recording step 1005, the speech recording unit starts recording the speaker after confirming that the recording start button is activated, recording the audio data, the text information obtained through real-time transcription, and the speaker's metadata. Next, in the integration step 1006, the integration unit integrates the session data of the recorded audio information, text information, and metadata into a file.

[0120] Next, in the integrated control step 1007, the integrated control unit performs integrated control of the processes up to the preparation of the statement, the authentication of the speaker, the recording of the speaker's statement, and the calculation of the score value.

[0121] The aforementioned voice input unit, integrated authentication unit, score calculation unit, recording start activation unit, speech recording unit, integration unit, and integrated control unit may be implemented as a speech recording integrated control program on a computer. The specific processing details have been described above, so a detailed explanation is omitted here.

[0122] Figure 21 is a flowchart showing the entire process from evidence generation to evidence submission. The following explanation of the process flow will use the flowchart in Figure 21. First, to prevent disputes over what was said or not said, a pre-recording gate process is used to verify the speaker's identity (authentication) and set a confidence score threshold as a gate function. Until these verifications are completed, the recording start button will not be activated, and the speaker will not be able to start recording (step 1101). When the recording start button is activated, the speaker's voice is recorded (step 1102), and evidence is generated. Next, once the recording is complete, the evidence materials are sealed (Step 1103). This sealing process involves generating the evidence materials according to the requirements of the recipient (PDF, ZIP, encryption, etc.). Details regarding the evidence material sealing process have been described above, so we will omit further explanation here. Next, an evidence package is created according to the recipient and submitted to the recipient specified by the user (Step 1104). Furthermore, the evidence is submitted to the recipient only after measures have been taken to ensure the security of the evidence. The security of the evidence (security measures, etc.) has been described in detail in the description of the above embodiment, so it will not be explained here. The above processing can improve the evidentiary value of the evidence.

[0123] Figure 22 is a flowchart for verifying the evidence package. The verification flow will be explained below using the processing flowchart in Figure 22. First, when a user triggers the output of evidence (step 1201), the management server 7 starts recording the output action (step 1202). Next, add the viewing / re-output / submission log (step 1203). Encrypt and save the submission log (step 1204). By linking to the administrator / auditor dashboard (step 1205), the submission destination and submission log information are saved, allowing for verification of the evidence package.

[0124] The speech recording integrated control system of this embodiment integrates "visualization of accountability for speech → prevention of tampering → reliability evaluation → reporting decision → social trust score," enabling a highly versatile platform. Furthermore, it can realize a speech recording integrated control system that enhances the evidentiary value of speech records by adding the individual's voice data to meeting minutes, various contracts, surgical consent forms at hospitals, automobile insurance contracts, etc.

[0125] Furthermore, in this embodiment, it is possible to realize a speech record integrated control system that can appropriately integrate and manage the generated speech record evidence materials, and can handle fraudulent contracts and claims about what was said or not said.

[0126] <Problem solving case study collection 1> This section presents 15 case studies illustrating how the integrated speech recording control system of the above embodiment can improve current "problems and challenges" and how it will be effective in the future. Each item is explained in the following order: (1) current problem → (2) additional elements → (3) improvement / future vision → (4) measurement indicator (example).

[0127] <1. "Unclear / Feasible" points raised during the review> (1) The formulas and thresholds are ambiguous and open to a wide range of interpretations. (2) Added explicit formulas for PPS / STS / SIS and threshold tables such as 85 / 60. (3) The invention is clearly defined, and the number of grounds for rejection is drastically reduced. This lays the foundation for initial approval. (4) Number of OA cases / number of responses, number of claim corrections, first-time approval rate.

[0128] <2. The claim of inventive step is weak.> (1) Often confused with the "post-recording authenticity" method. (2) The inevitable chain reaction of pre-recording gate control × notification priority (PPS) × submission destination control is definitively stated using formulas and branching tables. (3) The difference from prior art is demonstrated through "combination of elements," and the room for design avoidance is also reduced. (4) Number of cited references, citation accuracy rate for rejection reasons, and acceptance rate of counterarguments during examiner interviews.

[0129] <3. Concerns about the "strength of evidence" before courts and regulatory authorities> (1) The signature / timestamp / sealing procedures are abstract. (2) Explicitly define SHA-256 / ECDSA / RFC3161 TSA+WORM prevention+chain hashing. (3) The admissibility and verifiability of evidence will dramatically improve, making it more persuasive in disputes and administrative responses. (4) Verification recall rate, verification time, and evidence rejection rate at the receiving end.

[0130] <4. Operational burden of false and malicious reporting> (1) Handling standards are human-dependent, increasing the risk of damage escalation. (2) Automatic blocking / holding / submission based on SIS (Social Impact) and PPS branching. (3) Achieve both the suppression of false information and the expedited handling of urgent cases, thereby reducing response costs. (4) False reporting rate, FPR / FNR in the initial assessment, and compliance rate with response SLA.

[0131] <5. The "buried" status of legitimate whistleblowing> (1) Important cases are being held up due to manual review. (2) PPS >= 85 means immediate submission, pending means timer / re-evaluation. (3) Prioritizing important matters minimizes damage and increases user satisfaction. (4) Processing delays, escalation times, and estimated damages for high-PPS cases.

[0132] <6. Concerns about the leakage of personal and confidential information> (1) Fixed output particle size, risk of providing data for purposes other than intended. (2) Submission level table (A / B / C / E) + anonymization / masking requirements. (3) Minimum disclosure ensures compliance with laws and internal regulations, and makes it easy to explain any differences in disclosure. (4) Masking rate, re-identification risk assessment, number of complaints / incidents.

[0133] <7. Ambiguity of authority and responsibility> (1) It is unclear who can see what and produce the output. (2) RBAC × Operation Matrix (View / Output / Re-output / Log Reference). (3) The principle of "least authority" is institutionalized to deter internal fraud and operational errors. (4) Excess privilege rate, number of privilege violations detected, and lead time to corrective action.

[0134] <8. Weakness in auditing and accountability> (1) Insufficient log entries make it difficult to trace the problem. (2) Schema of the submission / audit log (ISO time, terminal signature, submission destination, IP / UA, number of re-outputs, etc.). (3) Complete traceability enables immediate response to audits and objections. (4) Audit time, log loss rate, and investigation costs.

[0135] <9. Inconsistencies when using multiple devices or offline> (1) Device differences and offline conditions can cause inconsistencies in time / signature. (2) Offline temporary encryption → post-synchronization, TS time priority, unification of start conditions. (3) Eliminates inconsistencies in consistency and reduces the implementation burden of field deployment. (4) Post-synchronization failure rate, time difference detection rate, and terminal-specific failure rate.

[0136] <10. Implementation rework and maintenance costs> (1) The specifications are text-heavy and require a lot of reinterpretation. (2) Fix the API definition / DB name matching (reports / scores / submissions / logs / ...) and the breakdown of the evidence ZIP file. (3) Improved implementation reproducibility / interoperability, and reduced MTTR for maintenance and modifications. (4) Number of implementation bugs, number of specification interpretation differences pointed out, and lead time for modifications.

[0137] <11. Variations in Model Quality and Safety> (1) The behavior changes when the LLM / classifier is updated. (2) Specify the prompt DB, evaluation criteria DB, and regression tests / rollbacks. (3) Early detection and suppression of quality degradation and harmful emissions, and safe continuous learning. (4) Post-release quality metrics (AUC / recall) fluctuations, number of deviations detected, and number of rollbacks triggered.

[0138] <12. Ad hoc responses to exceptions and failures> (1) TSA / signature failure leads to divided opinions on the ground. (2) Define the abnormal flow (retry / isolation / notification / submission not allowed). (3) Standardized safety shutdown and recovery procedures ensure stable operations. (4) Number of incidents, average recovery time, and number of users affected.

[0139] <13. International Expansion and Growth Potential for PPH> (1) Terminology and standards differ slightly across different agencies. (2) Specify that the glossary / translations are fixed and that the standards are compliant. (3) Increased consistency when using PPH, leading to shorter review times and higher approval rates. (4) PPH adoption rate, median time from commencement of review to approval.

[0140] <14. Negotiating power in litigation and licensing negotiations> (1) Often considered an invention of “business rules.” (2) Specify the mapping of formula, threshold, RBAC / log to technical means. (3) The interpretation of rights becomes clearer, which is advantageous in proving infringement and preventing design evasion. (4) Settlement / license unit price, negotiation period, and number of advantages listed in the technology comparison table.

[0141] <15. User Trust and Expanding Recruitment> (1) Distrust of the "black box" approach. (2) UI state transitions (color / activation / description label) + re-evaluation / objection mechanism. (3) Increased transparency and the Social Trust Score (STS) becoming the basis of trust in the ecosystem. (4) NPS / CSAT, satisfaction with the handling of re-evaluation requests, retention rate / churn rate.

[0142] <Overall effect> Short term: The focus of the review process shifted from "abstractness" to "narrow comparison of prior art," resulting in fewer open access requests and an upward revision of the first-time approval rate. Operations improved the SLA by suppressing false submissions and prioritizing important cases. Medium-term: Improved responsiveness to audits and legal compliance, along with international harmonization, will shorten lead times for PPHs and overseas applications. Implementation and maintenance costs will also decrease. Long-term: Expanding STS (Security Consulting System) as a core trust foundation to other areas (insurance, crime prevention, reporting hotlines, BPO, etc.).

[0143] <Problem solving case study collection 2> This book presents 15 examples of service / business models that illustrate how current societal challenges can be improved. Each section is explained in the following order: (1) Current challenges → (2) Proposed service (BM) → (3) How this patent is effective → (4) Expected effects / KPI examples (definitions) → (5) Revenue model (example).

[0144] <1. Administrative 119 / 110 Reporting Triage SaaS (B2G)> (1) The number of calls is overwhelmed by varying levels of urgency and false reports. (2) A cloud service that automatically prioritizes requests using PPS and switches between immediate submission / hold / blocking to 119 / 110. (3) If PPS ≥ 85, immediate coordination is required, crowd risk is emphasized in SIS, and verification can be immediately performed upon arrival with evidence ZIP (signed + TSA). (4) KPIs: Arrival → Initial decision SLA, false reporting rate, median time to site for high PPS cases. (5) Price: Municipal subscription + high PPS / unit variable rate.

[0145] <2. DV / Stalker Hotline 110 (B2C / B2G) Support App> (1) The victims cited "the difficulty of reporting the incident" and "the weakness of the evidence." (2) One-tap recording → integrated authentication → evidence sealing, secure UI that cannot be seen by the other party with RBAC, immediate submission to 110. (3) "Impersonation" is deterred at the pre-recording gate, and personal information is protected by masked output (submission level C). (4) KPIs: Adoption rate of protection order applications, time from victim report to intervention, and rate of reduction of re-victimization (follow-up). (5) Price: Free to use + subsidies from local government / funds, with long-term storage options for evidence preservation.

[0146] <3. School Bullying and Harassment "Safe Lines" (B2B / Board of Education)> (1) The school finds it difficult to take action due to false / controversial anonymous reports. (2) Prioritize student and parent reports with PPS, and deter "habitual false reporting" with STS (Continued Trust). (3) Submissions are made in stages to the school / board of education / police according to the submission level table, and transparency is ensured through audit logs. (4) KPIs: false reporting rate, time to intervention, and relapse prevention rate. (5) Pricing: School-based subscription (linked to the number of students) + audit dashboard.

[0147] <4. Corporate Whistleblowing (Compliance / Fraud) SaaS (B2B)> (1) Investigations can be prolonged due to disputes over what was said or not said, increasing the risk of retaliation. (2) Evidence pack (audio + RTT + metadata + signature / TSA), least privilege with RBAC, re-evaluation workflow. (3) Chain of custody is guaranteed by logs and confidentiality is maintained through anonymization (C level). (4) KPIs: Investigation cycle time, evidence acquisition rate, number of retaliatory complaints. (5) Pricing: Per-seat charge + per-item charge; audit export API is a higher-tier plan.

[0148] <5. Support for medical incident reporting and accountability (B2B / hospitals)> (1) Discrepancies in oral explanations and insufficient evidence during litigation. (2) Important scenes are unified and authenticated → immediately sealed, access is restricted to within the hospital at recipient B, and in the event of a dispute, they are fully submitted to A. (3) WORM sealing and time authenticity increase evidentiary value, and role matrix controls reuse. (4) KPIs: Explanation and consent dispute rate, dispute settlement rate, and time required for internal audits. (5) Pricing: Hospital-scale pricing + dispute resolution module.

[0149] <6. Insurance FNOL (First Report of Accident) Triage (B2B / B2C)> (1) False statements, exaggerations, and delays in initial response led to an increase in fraudulent payments. (2) PPS x SIS allows for immediate / holding, dashcam / smartphone multi-device linkage, and the assessment side verifies at submission level A. (3) Terminal signature + location deter spoofing, and abnormal flow (isolation in case of signature / TSA failure) ensures quality. (4) KPIs: Fraud detection accuracy, average payment cycle, initial on-site response time. (5) Price: Monthly fee per contract + FNOL charges.

[0150] <7. Traffic Accident / Road Rage Evidence Escrow (B2C / B2B)> (1) The video was tampered with / deleted, leaving the victim with no recourse. (2) Automatically upload the seal from the dashcam / smartphone and submit it to the police / insurance company at A level. (3) Authenticity is guaranteed by signature / TSA / chain hash, and even offline, data is preserved through post-synchronization. (4) KPIs: Percentage of evidence accepted, number of cases involving enforcement cooperation, number of days shortened in settlement negotiations. (5) Price: Monthly storage + fee per submission.

[0151] <8. Construction Site Safety Violations / Corrective Action Records (B2B / General Contractors)> (1) Suppressing or falsifying warnings about dangers. (2) Workers report violations via audio / photo → seal → report to supervisor / general contractor via RBAC, and the corrective action history is reflected in STS. (3) Submission level C makes it difficult to mask personal information and conceal it in logs. (4) KPIs: Correction lead time, recurrence frequency, number of incidents. (5) Price: Monthly fee per site ID + performance-based pricing (discounts for reducing disasters, etc.).

[0152] <9. Reporting / Auditing Abuse at Nursing Care and Welfare Facilities (B2B / B2G)> (1) The person receiving care had difficulty reporting the incident, and the evidence was missing in chronological order. (2) Family members / staff members record information via the safety UI → automatically routed to the relevant administrative body (A / B). (3) Visualize the trend of facility reliability with STS and support guidance with audit logs. (4) KPIs: Correction rate, recurrence rate, and family satisfaction. (5) Price: Facility subscription + government audit license.

[0153] <10. E-commerce Platform Dispute / Fake Review Countermeasures (B2B2C)> (1) False statements / comments in reviews or customer service. (2) Important communications are unified and authenticated → sealed, and in the event of a dispute, they are disclosed in stages at the submission level. (3) Prioritize serious complaints with PPS, and minimize personal information with RBAC. (4) KPIs: Chargeback rate, refund dispute duration, fake review detection rate. (5) Price: Tenant-based rate + dispute resolution package.

[0154] <11. Environmental Pollution Hotline (B2G / Citizen Science)> (1) On-site eyewitness accounts are not being used as evidence, and administrative responses are slow. (2) Prioritize based on SIS (Severity of Damage / Secondary Damage) and submit evidence to the supervisory authority at level A. (3) Location and time can be traced authentically, and responses are made transparent through audit logs. (4) KPIs: Time from reporting to water sampling / inspection, corrective order rate, and recurrence reduction rate. (5) Price: Municipal contract + corporate CSR support.

[0155] <12. Eliminating "He Said / She Said" Disputes in Contact Centers (B2B)> (1) Logs were lost, and excerpts were falsified, leading to legal disputes. (2) Conversation + meta is sealed on a session basis, shared internally at the C / B level, and disputes are handled by A. (3) Reproducible verification with evidence packs, and reduced operator errors through UI state transitions. (4) KPIs: Reinvestigation rate, escalation rate, resolution lead time. (5) Price: Charge per seat + storage.

[0156] <13. Journalism and Whistleblowing: "Evidence Vault" (B2B / Media)> (1) It is difficult to maintain both confidentiality of the source and authenticity of the evidence. (2) Anonymized C-level data will be sent to the editorial department, and an A-level set of evidence will be submitted for auditing purposes upon publication. (3) Access is restricted to journalists and legal staff via RBAC, and chain hashes are used to counter suspicions of tampering. (4) KPIs: Adoption rate, number of corrections / lawsuits, zero incidents of information source protection. (5) Price: Media subscription + per project fee.

[0157] <14. Crowd Safety / Event / Traffic Congestion Risk Detection (B2G / Business Operators)> (1) Accidents happen before the voices from the scene can be heard. (2) Audio reports from participants will be given priority in SIS (crowd density / suggestion of congestion) and immediately submitted to the operators / railway. (3) Multiple devices (smartphone / security camera sound) × offline synchronization reduces the chance of missing captures. (4) KPI: Hazard detection → Time to start guidance, time to clear congestion, number of accidents prevented. (5) Price: Per-event license + usage-based charges.

[0158] <15. Disaster Evidence Pack → Automated Support Application (B2G / B2C)> (1) Insufficient evidence / procedure delays in disaster certificates and benefits. (2) Record on site → seal ZIP → submit to local government at A level, application form automatically created via API. (3) Time-series verification using TSA timestamps, no data missing from submitted logs, and missing information supplemented by re-evaluation. (4) KPIs: Number of days from application to payment, rejection rate, and fraud prevention rate. (5) Price: Free for municipalities and citizens, with charges based on storage limits.

[0159] <Cross-cutting "evidence (why it works)"> Technical authenticity: Pre-recording gate (authenticity verification) → Signature / TSA / WORM → Chain hash to systematize resistance to tampering. Privacy & Minimal Disclosure: "Only those who need to see it, and only the minimum necessary," based on the recipient level and RBAC. Prioritization and fairness: PPS x SIS prioritizes urgent / critical cases, and the STS ledger deters fraud and abuse. Auditability: Third-party verification is possible with submitted / audit logs and evidence packs, making it strong in litigation and administrative matters. Field applicability: With multi-device / offline synchronization and abnormal flow handling, operations continue uninterrupted even in outdoor and high-load environments. [Explanation of Symbols]

[0160] 1. Mobile communication terminal 2 IC recorder 3. Security cameras 4. Dashcam 5. Integrated Control System for Recording Speech 6 Cloud 7. Management Server 8 LLM Servers 9. Submission terminal

Claims

1. A terminal device equipped with a recording function to record the voice spoken by the speaker is pre-registered with the speaker's first voice information. As preparation for speaking, the audio input unit inputs the speaker's second voice information and terminal information in order to obtain the speaker's permission before recording, An integrated authentication unit compares the second voice information input by the voice input unit with the first voice information and performs authentication of the speaker's identity. The integrated authentication unit calculates a score value from the degree of matching, A recording start activation unit activates the recording start button when the score value calculated by the score calculation unit is equal to or greater than a predetermined threshold, After confirming the activation status of the recording start button, the recording unit starts recording the speaker's voice data, records text information through real-time transcription, and records the speaker's metadata. An integration unit that combines recorded audio information, text information, and metadata into a file as session data, An integrated control unit that performs integrated control of the processes up to the preparation of the statement, the authentication of the speaker, the recording of the speaker, and the calculation of the score value, A speech recording integrated control system characterized by comprising the following features.

2. A scoring unit that calculates a confidence score, responsibility score, and sentiment score for the aforementioned statement based on the results of voice feature extraction, language analysis, and sentiment estimation of the aforementioned statement, A score visualization unit visualizes the confidence score, responsibility score, and emotion score of the statement calculated by the score calculation unit, Equipped with, The integrated control unit further performs integrated control of the processes up to the calculation of the score and the visualization of the score value, as described in claim 1.

3. The integrated speech recording control system according to claim 1 or 2, characterized in that it includes a tampering warning unit that adds the speaker's signature to the file to generate speech recording evidence, and also detects and warns of tampering operations.

4. The integrated speech recording control system according to claim 3, further comprising a speech authenticity evaluation responsibility determination unit that determines whether the content of the aforementioned speech is a normal speech, a false speech, or an ambiguous speech, and sorts the speech based on the urgency of the speech, reliability, and past history.

5. An evidence output unit that outputs the aforementioned speech record evidence file in one of the following formats: PDF, CSV, or ZIP. A submission destination control unit processes the file of the statement record evidence and controls its submission depending on whether the recipient is a court, an administrative agency, or a customer, and A submission history record unit that records the history of the file of the statement record evidence submitted to the aforementioned recipient, The speech recording integrated control system according to claim 4, characterized by comprising the above.

6. A social trust scoring unit accumulates and weights all scores, converts them into social trust scores for individuals and / or organizational units, and calculates social trust score values. A social trust score display unit that displays the social trust score value obtained by the social trust score calculation unit on a social trust score dashboard, The speech recording integrated control system according to claim 5, characterized by comprising the above.

7. Based on the results of voice feature extraction, language analysis, and sentiment estimation of the utterance, the confidence score of the utterance A. A score calculation unit that performs score calculations for responsibility score and emotional score, If the responsibility score value of the statement calculated by the score calculation unit is RS, the social trust score value is STS, the social impact score value is SIS, and the priority score value indicating the priority of the recipient is PPS, then the submission priority score generation unit calculates PPS based on the responsibility score value RS, the social trust score value STS, and the social impact score value SIS using the following formula: PPS = α・RS + β・STS + γ・SIS (α, β, γ are coefficients), The speech recording integrated control system according to claim 6, characterized by comprising the above.

8. The submission control unit automatically submits the data when the PPS is 85 or higher, holds it on hold when it is 60 or higher but less than 85, and refuses to submit it when it is less than 60. The speech recording integrated control system according to claim 7, characterized by comprising the above.

9. The update unit sequentially updates the aforementioned social trust score value STS using the following formula: STS_t = STS_{t-1} + Δs ​​- λ・Δt (t: time, Δs: score change value, λ: coefficient, Δt: time change value) The speech recording integrated control system according to claim 8, characterized by comprising the above.

10. The system includes a submission destination level database that manages the aforementioned submission destinations with trust levels A to C and E. The submission destination control unit has steps to control the output format, anonymization, and log granularity according to the confidence level. The speech recording integrated control system according to feature 9.

11. The audit department adds a chain hash to the submission log showing the log information submitted to the aforementioned recipient and maintains an audit log that includes the ISO 8601 time, submitter ID, terminal signature, and recipient ID. The speech recording integrated control system according to claim 10, characterized by comprising the above.

12. The aforementioned chain hash uses SHA-256, the electronic signature uses ECDSA, and the timestamp uses RFC3161 compliant TSA, and the package is sealed to create a WORM package. The speech recording integrated control system according to claim 11, characterized by comprising the above.

13. A terminal device equipped with a recording function to record the voice spoken by the speaker has the speaker's first voice information registered in advance. A voice input step in which the voice input unit inputs the speaker's second voice information and terminal information to obtain the speaker's permission before recording as preparation for speaking. An integrated authentication step in which the integrated authentication unit compares the second voice information input in the voice input step with the first voice information to authenticate the speaker's identity. A score calculation step in which the score calculation unit calculates a score value from the degree of comparison in the integrated authentication step. A recording start activation unit activates when the score value calculated in the score calculation step exceeds a predetermined threshold. A speech recording integrated control method characterized by comprising: a recording start activation step which activates the recording start button when the above condition is met; a speech recording step in which the speech recording unit, after confirming the activation status of the recording start button, starts recording the speaker and records audio data, text information obtained by real-time transcription, and metadata of the speaker; an integration step in which the integration unit integrates the session data of the recorded audio information, text information, and metadata into a file; and an integrated control step in which the integrated control unit performs integrated control of the processes up to the preparation of the speech, the authentication of the speaker, the recording of the speaker, and the calculation of the score value.

14. A speech recording integrated control program characterized in that a computer equipped with a recording function for recording the voice spoken by a speaker has a first voice information of the speaker registered in advance, the computer is configured to be executed as an integrated control unit that performs integrated control of the processes up to the preparation for speaking, the authentication of the speaker, the recording of the speaker, and the calculation of the score value by comparing the second voice information input by the voice input unit with the first voice information to obtain permission from the speaker before recording, the integrated authentication unit calculates a score value from the degree of matching, the recording start activation unit activates the recording start button when the score value calculated by the score calculation unit is above a predetermined threshold, the speech recording unit starts recording the speaker after confirming the activation status of the recording start button and records the voice data, text information from real-time transcription, and the speaker's metadata, the integration unit integrates the session data of the recorded voice information, text information, and metadata into a file, and the computer is configured to perform integrated control of the processes up to the preparation for speaking, the authentication of the speaker, the recording of the speaker, and the calculation of the score value.