Meeting support system and meeting support method

The conference support system uses personal voice models on user terminals to identify and integrate speaker information for offline meetings, addressing speaker identification challenges and ensuring privacy, thus generating accurate meeting minutes for both online and offline settings.

JP7836140B1Active Publication Date: 2026-03-26CLOUDBASE INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing technologies struggle to identify speakers and generate meeting minutes for offline conferences where multiple participants speak in the same physical space, making it difficult to apply conventional minute-making technologies.

Method used

A conference support system utilizing user terminals with personal voice models to identify user speech and transmit feature vectors to a management server, which integrates the information chronologically to associate speakers with their statements, enabling the creation of meeting minutes for both online and offline meetings while maintaining privacy by processing voice identification locally.

Benefits of technology

Enables the generation of meeting minutes for offline conferences by accurately identifying speakers and reducing privacy risks through local voice identification, applicable to both online and offline meetings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007836140000001_ABST
    Figure 0007836140000001_ABST
Patent Text Reader

Abstract

To provide a meeting support system and a meeting support method applicable to creating meeting minutes for offline meetings. [Solution] A conference support system comprising multiple user terminals and a management server, wherein each user terminal includes a model learning unit that generates a personal voice model to identify the user's voice, a data acquisition unit that acquires conference audio data, and a feature extraction unit that identifies the user's utterances from the audio data using the personal voice model and extracts the voice feature vectors when the user speaks. The management server includes a data receiving unit that acquires feature information indicating the voice feature vectors from each user terminal, and an integration processing unit that identifies the speaker of each statement in the conference by arranging and integrating the feature information acquired from multiple user terminals in chronological order.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a conference support system and a conference support method that execute processes related to creating minutes of a conference.

Background Art

[0002] In recent years, technologies have been proposed that convert voice during a conference into character information and automatically generate minutes and the like using the character information. For example, Patent Document 1 discloses extracting voice subtitle data from cloud recording data of an online conference and generating minutes by natural language processing from the voice subtitle data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the online conferences targeted in Patent Document 1, it is common for participants to participate in the conference using their respective individual terminals, and since the remarks by each participant are input from their respective individual terminals, it is easy to identify the speaker. On the other hand, in an offline conference held in a face-to-face format, since multiple participants speak within the same space such as a conference room, it is difficult to identify the speaker, and it has been difficult to apply conventional minute-making technologies such as those in Patent Document 1 to offline conferences.

[0005] One of the objectives of exemplary embodiments of the present disclosure is to provide a conference support system and a conference support method applicable to creating minutes of an offline conference.

Means for Solving the Problems

[0006] A conference support system relating to one aspect of this disclosure comprises multiple user terminals and a management server. Each of the user terminals is: A model learning unit that learns the speech of users using the user terminal and generates a personal voice model that identifies the user's voice, A data acquisition unit that acquires audio data from meetings, A feature extraction unit that identifies the user's speech from the audio data using the personal voice model and extracts the speech feature vector when the user speaks, The system includes a transmission processing unit that transmits feature information generated by adding the speech time to the aforementioned speech feature vector to the management server. The aforementioned management server A data receiving unit that acquires the feature information from each of the user terminals, The system includes an integration processing unit that identifies the speaker of each statement in the meeting by arranging and integrating the feature information acquired from multiple user terminals in chronological order.

[0007] According to the above-described meeting support system, meeting minutes can be created not only for online meetings but also for face-to-face offline meetings, while associating speakers with their statements. Other issues and solutions disclosed in this application will be made clearer by the embodiments and drawings of this disclosure. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is a diagram illustrating the configuration of a conference support system according to one embodiment of the present disclosure. [Figure 2] Figure 2 is a block diagram illustrating the hardware configuration of the user terminal shown in Figure 1. [Figure 3] Figure 3 is a block diagram illustrating the software configuration of the conference support system shown in Figure 1. [Figure 4] Figure 4 is a flowchart showing an example of information processing performed by the conference support system shown in Figure 1. [Modes for carrying out the invention]

[0009] Hereinafter, an information processing system relating to one embodiment of this disclosure will be described with reference to the drawings. In the attached drawings, identical or similar elements are given identical or similar reference numerals and names, and redundant descriptions of identical or similar elements may be omitted in the description of the embodiment. The contents shown in each drawing are merely examples for explaining this embodiment and are only schematic examples to facilitate explanation of this embodiment. The contents of each drawing may be modified or changed to the extent that no technical problems arise.

[0010] <System Overview> The meeting support system according to this embodiment (hereinafter referred to as "this system") is a system that performs information processing related to the creation of meeting minutes. This system is applicable to the creation of minutes for offline meetings. Here, "offline meeting" means a meeting in which at least some participants participate face-to-face in the same physical space (e.g., a conference room). In contrast, an online meeting (web conference) means a meeting in which each participant participates from their respective terminal via a network.

[0011] In this embodiment, the process of generating meeting minutes using this system will be described for offline meetings in which each participant participates in the meeting within the same physical space. However, this system is also applicable to generating meeting minutes for online meetings, and is applicable to meetings that include both online and offline participants, such as when some participants participate online from remote locations. The type of meeting targeted by this system is not necessarily limited.

[0012] When attempting to automatically generate meeting minutes, it is common practice to convert the audio data from the meeting into text and then generate the minutes based on that text data. The audio data from the meeting is collected, for example, by devices owned by each participant (e.g., smartphones, portable devices, laptops, and other general-purpose computers). In the case of online meetings where each participant accesses the meeting from a different location, each device collects only the voice of its owner and shares the collected audio with the devices of other participants. Therefore, in the case of online meetings, it is possible to record the speaker and the content of their remarks by associating them with each participant's account information or device identification information. In contrast, in the case of offline meetings held in person, when audio is collected by each participant's device, the collected data on each device includes not only the owner's voice but also the voices of other participants in the surrounding area. Therefore, in the case of offline meetings, it is not easy for the system to identify whose remarks are being made.

[0013] In this system, each user's (participant's) terminal (hereinafter referred to as "user terminal") acquires audio data during the meeting, and then uses each user terminal's personal voice model to identify the speech of the user using the terminal from the audio data during the meeting. Here, the "personal voice model" is a user-specific AI module built to identify the user's voice by machine learning the voice of the user who owns the user terminal. The personal voice model may also be an inference model that operates to complete processing locally within the user terminal (e.g., edge AI, local AI). The personal voice model identifies only the voice of the user using the terminal from the audio data during the meeting, which may contain the voices of multiple people, and extracts the speech feature vector when the user speaks. Each user terminal sends the generated feature information, which is the speech feature vector with the time of utterance attached, to the management server. The management server identifies the speaker of each statement in the meeting by arranging and integrating the feature information received from each user terminal in chronological order.

[0014] A "voice feature vector" refers to information obtained by converting voice features calculated from an audio signal into a vector format represented by an array of numbers. The individual voice model converts the user's speech segments, identified through analysis of the meeting's audio data, into voice feature vectors and sends them to the management server. Here, the user's raw voice data and voiceprint information, such as its frequency characteristics, are biometric information that can be used to identify a specific individual. If the system is configured to store voice data and biometric information such as voiceprint information on a management server (including a cloud server), it will face privacy risks such as the leakage of personal information. Therefore, in this system, the voice identification process is completed locally on each user terminal, and the management server aggregates voice feature vectors converted into a format that does not include biometric information, rather than the voice data itself. With this configuration, this system can provide a meeting minutes generation service applicable not only to online meetings but also to offline meetings, while reducing privacy risks. The details of this system will be explained below using the examples shown in the diagram.

[0015] <System Configuration> As shown in Figure 1, the information processing system of this embodiment comprises a plurality of user terminals 1 and a management server 2. The user terminals 1 and the management server 2 are connected to each other via a network NW so that they can communicate with one another. In this embodiment, the network NW is mainly assumed to be the internet, but the network NW is not limited to the internet and may be constructed using, for example, a public telephone network, a mobile phone network, a wireless communication network, Ethernet (registered trademark), etc. Note that the illustrated configuration is just an example and is not limited thereto.

[0016] <User Terminal 1> The user terminal 1 is an information processing device used by users participating in a meeting. The user terminal 1 may be, for example, a mobile terminal such as a smartphone or a tablet terminal, or a general-purpose computer such as a workstation and a personal computer. Also, each user terminal 1 may include a plurality of devices. For example, it may include wearable terminals such as smartwatches and smart rings.

[0017] FIG. 2 is a block diagram illustrating the hardware configuration of the user terminal 1. Note that the illustrated configuration is an example, and the user terminal 1 may have other configurations. The user terminal 1 includes, for example, a processor 10, a memory 11, a storage 12, a transmission / reception unit 13, an input / output unit 14, etc., which are electrically connected to each other through a bus 16.

[0018] The processor 10 is an arithmetic unit that controls the operation of the entire user terminal 1, controls the transmission and reception of data between each element, and performs information processing necessary for the execution and authentication processing of applications. For example, the processor 10 is a CPU (Central Processing Unit) and / or a GPU (Graphics Processing Unit). Each function provided in the user terminal 1 is realized by the processor 10 executing a program or the like stored in the storage 12 and expanded in the memory 11.

[0019] The memory 11 includes a main memory composed of a volatile storage device such as a DRAM (Dynamic Random Access Memory), and an auxiliary memory composed of a non-volatile storage device such as a flash memory and an HDD (Hard Disc Drive). The memory 11 is used as a work area of the processor 10, and stores a BIOS (Basic Input / Output System) executed when the user terminal 1 is started up, and various setting information.

[0020] Storage 12 stores various programs, such as application programs, and in particular, programs for executing each function of the user terminal 1 in this system. A database containing data used for each process may also be built in storage 12. For example, the memory unit 120 described later is implemented as part of the memory area of ​​memory 11 and / or storage 12.

[0021] The transmitting / receiving unit 13 is a communication interface for the user terminal 1 to communicate with various information processing terminals such as the management server 2 via a communication network. The transmitting / receiving unit 13 may further include a short-range communication interface such as Bluetooth® and BLE (Bluetooth Low Energy) and / or a USB (Universal Serial Bus) terminal.

[0022] The input / output section 14 includes information input devices such as keyboards, mice, and microphones for collecting sound, and output devices such as displays. The input / output section 14 may also include a touch panel that combines both information input and output functions, and may include a printer, speaker, etc., as output devices.

[0023] Bus 16 is connected in common to all of the above elements and transmits, for example, address signals, data signals, and various control signals.

[0024] <Management Server 2> The management server 2 is an information processing device that performs various information processing tasks to create meeting minutes based on information sent from each user terminal 1. The management server 2 may be configured on-premises using a general-purpose computer such as a workstation or personal computer, or it may be logically implemented through cloud computing. Like the user terminal 1, the management server 2 is equipped with a processor 20, memory 21, storage 22, a transceiver 23, an input / output unit 24, etc., which are electrically connected to each other via a bus 26. Each element of the hardware configuration of the management server 2 can be configured in the same way as the user terminal 1 shown in Figure 2, and a detailed explanation of each hardware element of the management server 2 is omitted.

[0025] <System Functions (Software Configuration)> Figure 3 is a block diagram illustrating the functions (software configuration) implemented in the conference support system of this embodiment. Each user terminal 1 may include, for example, a model learning unit 101, a data acquisition unit 102, a feature extraction unit 103, a text conversion processing unit 104, and a transmission processing unit 105, as functions realized by the processor 10 executing a program. These functional units are exemplified as functions executed by the processor 10 of the user terminal 1, but some of these functional units may be executed by the processor of another information processing device, such as the management server 2, instead of the user terminal 1. In addition, the storage unit 120 of each user terminal 1 may include various databases such as a model information storage unit 121, an acquired data storage unit 122, and a reference information storage unit 123.

[0026] The model information storage unit 121 stores information about the individual voice model. For example, the model information storage unit 121 may store model data that constitutes the individual voice model of a user, associated with identification information (user ID) that uniquely identifies the user. The model data of the individual voice model may include trained model parameters, information indicating the model structure, a model identifier, version information, creation date and time, update date and time, and other management information.

[0027] Furthermore, the model information storage unit 121 may store training audio data used to generate or update the personal voice model, linked to the user's identification information. The training audio data may include the user's voice data itself for machine learning, as well as management information such as the number of training sessions, the total amount of audio data used for training, and condition information at the time of training execution. The model information storage unit 121 may be configured to temporarily store the training audio data only during the training of the personal voice model, or it may be configured to continuously retain the training audio data not only during training but also during the operation phase of the personal voice model. In other words, the training audio data does not necessarily need to be continuously stored and may be deleted after the completion of the training process.

[0028] In addition to the data described above, the model information storage unit 121 may also store indicator information showing the effectiveness of the individual voice model, status information for controlling the availability of the model, and various setting information for voice recognition. The setting information for voice recognition may include setting information that defines rules for adjusting parameters according to the reference information described later during voice recognition.

[0029] The acquired data storage unit 122 stores information about the meeting acquired by the user terminal 1. For example, the acquired data storage unit 122 may store audio data collected during the meeting, linked to identification information (meeting ID) for uniquely identifying the meeting. The audio data may include primary audio data in its original collected state and secondary audio data obtained by applying predetermined processing to the primary audio data. The primary audio data includes not only the speech of the user using the user terminal 1, but also the speech of other users during the meeting and ambient sounds. The secondary audio data may include, for example, audio data obtained by applying predetermined preprocessing such as noise reduction to the primary audio data, audio frames obtained by dividing the primary audio data into predetermined sections such as speech segments, and audio data corresponding to the user's own speech segments extracted from the primary audio data.

[0030] The acquired data storage unit 122 may be configured to temporarily store the above-mentioned audio data only for the period until the processing related to speech identification is completed, or it may be configured to continuously retain the data even after the processing related to speech identification is completed. In other words, the audio data may be deleted after the processing related to speech identification is completed (for example, after the speech feature vector is sent to the management server 2), or it may be retained in the acquired data storage unit 122 for any period until a deletion instruction is entered by the user.

[0031] The acquired data storage unit 122 may also store supplementary information associated with the audio data. This supplementary information may include, for example, information to identify the speaking time of each utterance in a meeting, such as the start time of audio data collection, the end time of audio data collection, and the length of the audio data (playback time); information indicating the data format of the audio data; and conditions at the time of recording (information identifying the microphone or recording device, microphone parameters at the time of recording, etc.). The acquired data storage unit 122 may also store meeting-related information such as the meeting name, location, date and time, and participants (or prospective participants) linked to the meeting identification information. Such meeting-related information may be obtained from a schedule management application such as a calendar, or it may be obtained through user input.

[0032] The reference information storage unit 123 stores reference information that can be referenced in speech recognition processing using a personal voice model. The reference information may include information indicating the possibility of changes in the user's voice quality. For example, the reference information storage unit 123 may store the user's schedule information, user status information such as the user's health status, and similar information. The schedule information may include, for example, the user's attendance history, the user's work schedule, work history, and the user's private schedule. The data format, acquisition method, and update frequency of the reference information stored in the reference information storage unit 123 are not particularly limited. Furthermore, the reference information may be information generated within the user terminal 1, or information acquired from an external information processing device or application.

[0033] In addition to the model information storage unit 121, the acquired data storage unit 122, and the reference information storage unit 123, the storage unit 120 may also store information generated or used in various information processing performed by the user terminal 1. For example, the storage unit 120 may store speech analysis results such as speech feature quantities and speech feature vectors extracted from speech data, or it may store text data after the speech data has been transcribed. In the above description, the model information storage unit 121, the acquired data storage unit 122, and the reference information storage unit 123 have been described as logically distinct, but these databases may be physically implemented in the same or different storage areas of the memory 11 and / or storage 12. Furthermore, the data items, data configurations, and management units stored in each of these storage units can be appropriately changed according to the manner of information processing in the user terminal 1, and are not limited to the examples of this embodiment.

[0034] The model learning unit 101 executes the process of building a personal voice model that identifies the user's voice. For example, the model learning unit 101 may generate a personal voice model specifically for the user using the user terminal 1 by machine learning a reference model for voice identification using the user's speech for training.

[0035] The method for acquiring audio data used to train a personal voice model is not particularly limited. For example, the model training unit 101 may output a training screen that prompts the user to read a predetermined text aloud or to speak freely for a predetermined amount of time, and use the audio data acquired through the training screen as training audio data. The model training unit 101 may use only the user's normal voice as training audio data, or it may use, in addition to such normal voice data, audio data from when the user's voice has been altered due to a predetermined reason such as a cold as training data. The machine learning algorithm used to construct the personal voice model is not particularly limited. For example, known training algorithms such as neural networks, convolutional neural networks, and self-supervised learning may be employed.

[0036] The model learning unit 101 may generate a personal voice model based on speech features extracted from user speech for training. In this case, the model learning unit 101 may analyze the user's speech data for training as described above and extract speech features contained in the user's speech, such as intonation, spectral characteristics, and formants. The model learning unit 101 may then construct a personal voice model that selectively identifies only the user's speech by having a reference model for speech recognition perform machine learning on the speech features extracted from the training data. The types of speech features extracted during machine learning of the personal voice model are not particularly limited. Furthermore, the method of constructing the personal voice model is not limited to this method; the user's speech data itself may be directly provided to the reference model to perform machine learning. The model learning unit 101 may perform additional learning or retraining on an already generated personal voice model to update the personal voice model. The personal voice model generated or updated by the model learning unit 101 is stored in the model information storage unit 121.

[0037] The data acquisition unit 102 acquires audio data of the meeting. The data acquisition unit 102 may acquire audio data in real time from the start to the end of the meeting, or it may acquire all the audio data collected during the meeting after the meeting has ended. The method of acquiring audio data is not particularly limited. For example, the data acquisition unit 102 may acquire audio data by collecting audio during the meeting using an input device such as a microphone installed in the user terminal 1. In this case, the data acquisition unit 102 may present a UI for audio collection on the display unit (display) of the user terminal 1. Then, through this UI, it receives instructions from the user to start and stop recording at the start and end of the offline meeting, and collects audio data from the offline meeting.

[0038] Audio data collection may be performed on a device other than the user terminal 1. In this case, the data acquisition unit 102 of the user terminal 1 may receive the audio data acquired by the audio collection device in real time or after the meeting has ended. When acquiring audio data after the meeting has ended, the audio data may be acquired by reading a predetermined recording medium such as a USB drive on which the audio data is recorded.

[0039] The data acquisition unit 102 records the acquired audio data in the acquired data storage unit 122. When acquiring audio data, the data acquisition unit 102 may accept input from the user for information to identify the meeting (e.g., meeting ID, meeting name, etc.) and record the audio data linked to this input information. In this case, the data acquisition unit 102 may also accept input of meeting-related information such as the meeting location and the meeting agenda (purpose). Furthermore, the data acquisition unit 102 may record time information (time stamps), such as the start and end times of audio collection, linked to the audio data. Note that the information to identify the meeting linked to the audio data is not limited to input by the user; it may also be obtained from the user's schedule management tool, such as a calendar application. In this case, the data acquisition unit 102 may identify the meeting of the audio being collected by obtaining the user's schedule information registered at the time corresponding to the audio collection start time from the schedule management tool using a predetermined method such as API integration.

[0040] The feature extraction unit 103 uses a personal voice model to identify the user's speech from the audio data recorded from the meeting. Specifically, the feature extraction unit 103 applies the personal voice model to the meeting audio data and identifies the speech segment corresponding to the speech of the user using the user terminal 1 from among the multiple utterances contained in the audio data. In this case, the personal voice model may analyze each utterance contained in the audio data and extract speech features from each utterance. Then, by comparing the speech features extracted from the audio data with the user's speech features defined as parameters in the personal voice model, the speech segment corresponding to the user's speech may be identified.

[0041] The feature extraction unit 103 may perform predetermined preprocessing on the meeting audio data before executing the speech recognition processing described above. For example, the feature extraction unit 103 may perform preprocessing to remove noise such as ambient sounds from the audio data. The method of noise reduction processing is not particularly limited, and known single or multiple noise reduction processing methods such as spectral subtraction processing, filtering processing, or noise reduction processing using a trained noise control model can be applied. The feature extraction unit 103 may also perform preprocessing to divide the audio data into utterance units or predetermined time units. For example, the feature extraction unit 103 may divide the audio data into multiple audio sections using silent sections or low-volume sections with a volume below a predetermined threshold as boundaries, and determine whether each divided audio section corresponds to the user's speech. The execution order of preprocessing and the types of processing performed as preprocessing may be appropriately changed depending on the audio data acquisition environment, the format of the meeting, or the configuration of the individual voice model.

[0042] The feature extraction unit 103 may perform a process to adjust the threshold parameter for determining whether or not the voice is the user's speech during the voice recognition process. Specifically, the feature extraction unit 103 may control the threshold parameter for identifying the user's voice by estimating the possibility of a change in the user's voice quality based on reference information including at least one of the user's schedule information and the user's status information. For example, by referring to the user's schedule information, if attendance history such as absence due to poor health is registered before the date of the meeting, the threshold parameter may be changed by estimating the possibility of a change in voice quality due to poor health. Alternatively, before (or after) the start of recording the meeting's audio, the user may be presented with a questionnaire regarding their current physical condition, and the answer to the questionnaire may be obtained as the user's status information. If an answer such as hoarseness is entered in the questionnaire, the threshold parameter may be changed according to that answer.

[0043] The method for controlling threshold parameters is not particularly limited. For example, threshold parameters may be modified to widen the range between the upper and lower limits in order to relax the voice recognition judgment, or threshold parameters may be modified to shift the interval of parameters that are judged as user voice by increasing or decreasing the upper and lower limits by a certain amount. Threshold parameters may also be modified in accordance with predetermined control rules. If thresholds are set for multiple types of parameters for user voice recognition, the thresholds may be changed for each type of parameter.

[0044] The feature extraction unit 103 identifies speech segments corresponding to user utterances and then converts the identified user utterances into speech feature vectors to extract speech feature vectors corresponding to user utterances from the conference audio data. For example, speech features extracted from user utterances (e.g., features representing the frequency distribution shape such as power spectra and Mel-frequency spectra, features showing the periodic structure of spectra such as Mel-frequency cepstrum coefficients, features corresponding to the physical characteristics of the vocal organs such as linear prediction coefficients and formant bandwidths, pitch variation, speech velocity, etc.) may be converted into 128-dimensional speech feature vectors. The types of speech features to be extracted, the calculation methods, the number of dimensions of the speech feature vectors, and the methods for converting to vectors are not particularly limited and may be set as appropriate depending on the configuration of the personal speech model and the manner of speech identification processing.

[0045] The feature extraction unit 103 may detect user utterances in real time from audio data acquired during a meeting and convert the content of the utterances into audio feature vectors each time user utterances are detected. Alternatively, after the meeting ends, the system may analyze the audio data from the start (start of recording) to the end (end of recording) of the meeting in chronological order to identify audio segments corresponding to user utterances and extract audio feature vectors for each identified audio segment.

[0046] The feature extraction unit 103 may discard the audio data that generated the audio feature vectors after extracting the audio feature vectors related to the user's speech. In this case, the user's speech is not sent to the management server 2, and only the audio feature vectors are sent to the management server 2 as data originating from the conference audio. In other words, the speech identification process is completed locally within the user terminal 1. The timing of deleting the original audio data that has been analyzed is not necessarily limited; the deletion process may be performed immediately after conversion to audio feature vectors, after the process of sending the audio feature vectors to the management server 2 is performed, or after a predetermined time has elapsed after sending the audio feature vectors.

[0047] The text conversion processing unit 104 performs the process of converting the meeting audio data into text data. Specifically, the text conversion processing unit 104 takes the meeting audio data as input, recognizes the content of human speech, and generates text data in which the recognized speech content is represented as a string of characters. The method of speech recognition processing for transcription is not necessarily limited. For example, a personal voice model may be used to convert the audio data into text data, or a speech recognition model specifically for transcription may be used, or the conversion process to text data may be performed by referring to a speech recognition library. Other known methods may also be adopted.

[0048] The text conversion processing unit 104 may generate text data that records each statement in the meeting in chronological order by identifying the speaking time of each statement recorded in the meeting's audio data based on the recording start time of the audio data and the elapsed time from the start time. The text conversion processing unit 104 may generate text data in real time processing of the audio data collected during the meeting, or it may generate text data after the meeting has ended by analyzing the audio data from the start to the end of the meeting all at once. The text conversion processing unit 104 sends the generated text data to the management server 2. In this case, the audio data that was the source of the text data may not be sent to the management server 2, and only the text information may be sent.

[0049] The text conversion processing unit 104 may generate text data from the entire audio data during the meeting, or it may generate text data from only the audio sections corresponding to the user's utterances identified by the feature extraction unit 103. In the latter case, the text conversion processing unit 104 generates text data that transcribes only the utterances of the user using the user terminal 1. Furthermore, the text data generation process by the text conversion processing unit 104 may be performed on any one user terminal 1, or on each of the multiple user terminals 1. For example, the configuration may involve generating text data from the entire audio data during the meeting on any one user terminal 1, or each user terminal 1 may involve generating text data that transcribes only the utterances of the user who owns that terminal.

[0050] Note that the text data generation process may be performed by an information processing device other than the user terminal 1. For example, a transcription device dedicated to transcription processing may be used to convert the conference audio into text data. In this case, the user terminal 1 does not need to be equipped with the text conversion processing unit 104.

[0051] The transmission processing unit 105 executes the process of sending the speech feature vector extracted by the feature extraction unit 103 to the management server 2. Specifically, the transmission processing unit 105 adds information indicating the speech time to the speech feature vector converted from the user's speech, generates feature quantity information representing the speech feature vector, and sends this feature quantity information to the management server 2. The information indicating the speech time may be, for example, a timestamp expressed as an absolute time, or information indicating a relative time based on the recording start time, and the method of representing the speech time is not necessarily limited.

[0052] The transmission processing unit 105 may add information for identifying the user to the speech feature vector and send the speech feature vector, which includes information indicating the user as the speaker, to the management server 2. The format of the information for identifying the user is not particularly limited; it may be an identifier assigned to the user terminal 1, identification information corresponding to the user account (user ID), or similar information. In addition, information for identifying the meeting (such as a meeting ID or other meeting identification information) is added to the feature information indicating the speech feature vector.

[0053] The transmission processing unit 105 may transmit the audio feature vector to the management server 2 in its original data format, or it may perform a predetermined encryption process on the audio feature vector before transmission. The encryption method is not particularly limited, and any known encryption technology, such as symmetric-key cryptography, public-key cryptography, or a combination thereof, may be applied. The feature information transmitted from the transmission processing unit 105 to the management server 2 may include the audio feature vector (or the encrypted audio feature vector) but not the audio data itself.

[0054] As shown in Figure 3, the management server 2 may include, for example, a data receiving unit 201, a text data acquisition unit 202, an integration processing unit 203, a meeting minutes generation unit 204, and an output control unit 205, as functions realized by the processor 20 executing a program. These functional units are exemplified as functions executed by the processor 20 of the management server 2, but some of these functional units may be executed by the processor of another information processing device, such as the user terminal 1, instead of the management server 2. In addition, the storage unit 220 of the management server 2 may include various databases such as a received information storage unit 221 and a meeting information storage unit 222.

[0055] The received information storage unit 221 stores received information, such as feature information transmitted from the user terminal 1. For example, the received information storage unit 221 may store feature information, in which information indicating the time of utterance is attached to the speech feature vector, in association with identification information for identifying the meeting. The received information storage unit 221 may also store text data of the meeting received from the user terminal 1 or an information processing terminal such as a transcription device. The feature information stored in the received information storage unit 221 may be temporarily stored only for the period until it is used for subsequent processing such as meeting minutes generation, or it may be retained for a predetermined period as a processing history related to the meeting.

[0056] The meeting information storage unit 222 stores various information related to the meeting. For example, the meeting information storage unit 222 may store meeting-related information such as the meeting name, date and time, location, and participant information, associated with identification information (meeting ID) to uniquely identify the meeting. The meeting information storage unit 222 may also store meeting minutes data generated in response to the meeting, or similar information indicating the content of the meeting. This information is generated based on feature information stored in the receiving information storage unit 221 and / or text data acquired from the user terminal 1, etc., and stored in the meeting information storage unit 222.

[0057] In addition to the received information storage unit 221 and the conference information storage unit 222 mentioned above, the storage unit 220 of the management server 2 also stores various types of information used for information processing, such as configuration information for each process executed by the management server 2.

[0058] For example, the memory unit 220 may store security standard information defined by the organization to which the user participating in the meeting belongs. This security standard information may include information security policies set by the user's organization, operational rules for AI assets used by the user's organization, such as personal voice models, configuration information defining items of confidential information such as confidential company information and confidential keywords, and configuration information regarding the handling of personal information. Such security standard information may be configured in the AI ​​Security Posture Management System (AI-SPM system). In this case, the memory unit 220 may also store configuration information for coordinating with the AI-SPM system, information indicating the location where the security standard information is stored in the AI-SPM system, and so on.

[0059] The data receiving unit 201 executes a process to receive feature information transmitted from each user terminal 1. For example, the data receiving unit 201 may receive feature information corresponding to the content spoken by each user in the meeting from each user terminal 1, associate meeting identification information (meeting ID) with this feature information, and register it in the received information storage unit 221. The data receiving unit 201 may also identify the corresponding meeting identification information based on meeting identification information (meeting name, date and time, location, etc.) sent from the user terminal 1 along with the feature information, and attach it to the feature information. The meeting identification information may be attached to the feature information by the transmission processing unit 105 at the stage when the user terminal 1 transmits the feature information. The data receiving unit 201 may store the feature information received from each user terminal 1 in the received information storage unit 221 in a distinguishable state. As information for distinguishing user terminals 1, terminal identification information, user identification information, or similar information may be used, and its format is not particularly limited.

[0060] The text data acquisition unit 202 acquires text data that records each statement made during the meeting in chronological order. For example, the text data acquisition unit 202 may acquire text data generated on the entire audio data during the meeting from any one of the user terminals 1. Alternatively, the system may be configured to acquire text data generated on each user terminal 1, targeting only the utterances of the user who owns that terminal, individually from multiple user terminals 1. In this case, the text data acquisition unit 202 registers the text data acquired from each user terminal 1 as text data belonging to the same meeting based on the meeting identification information in the received information storage unit 221. If the meeting text data is generated using an information processing device other than a user terminal 1, the text data acquisition unit 202 acquires the meeting text data from that information processing device.

[0061] The integration processing unit 203 performs a process to integrate the feature information received from each user terminal 1. Specifically, the integration processing unit 203 arranges and integrates the feature information obtained from multiple user terminals 1 in chronological order based on the information indicating the speaking time attached to each speech feature vector. Based on the integrated feature information, the integration processing unit 203 identifies the speaker of each statement in the meeting. The integration processing unit 203 may also perform a process to decrypt each piece of feature information before integrating the feature information.

[0062] For example, the integration processing unit 203 identifies the speaker of each statement in the meeting by comparing the time-series integrated feature information with the text data acquired from the text data acquisition unit 202, and adds speaker information to each statement in the text data. This makes it possible to generate integrated data that associates the content of the statement with the speaker for each statement in the meeting. When identifying the speaker, the statement and the feature information can be associated by checking whether the utterance time (start time and / or end time of the utterance) attached to each speech feature vector matches the time of each statement in the text data within a predetermined error range.

[0063] The method for associating speakers with text data is not necessarily limited to the method of referring to separately acquired text data as described above. For example, the integrated processing unit 203 may perform a process to generate meeting text information from each speech feature vector. Specifically, the integrated processing unit 203 may generate text data indicating the utterance content of each speech feature vector by reconstructing a phoneme sequence (string) from each speech feature vector arranged in chronological order. In this case, the information of the user terminal 1 that transmitted each speech feature vector, or the user identification information attached to the speech feature vector, is associated with the text generated from the speech feature vector to generate text data for the entire meeting that associates the content of the speech with the speaker. The method for converting speech feature vectors to phoneme sequences (strings) is not particularly limited, and for example, an acoustic model (such as an ASR model) that has been trained to enable conversion from vector information to phoneme sequences may be used, or other known methods may be adopted.

[0064] The integrated processing unit 203 stores text data (text data for the entire meeting) with speaker identification information linked to meeting identification information in the meeting information storage unit 222. In the process of identifying speakers and generating text data by the integrated processing unit 203, the audio data itself and voiceprint information are not referenced (not used). The management server 2 identifies speakers using audio feature vectors without handling audio data or voiceprint information, thereby reducing privacy risks such as leakage of personal information, while allowing the management server 2 to consolidate the information necessary for creating meeting minutes.

[0065] The minutes generation unit 204 performs a process to create meeting minutes by performing natural language processing on text data to which speaker information has been added. The natural language processing method is not particularly limited, and known natural language processing techniques such as morphological analysis, dependency parsing, sentence segmentation, semantic analysis, and summarization processing may be applied individually or in combination. For example, the minutes generation unit 204 may divide the text data into sentence units or utterance units, analyze the semantic content of each sentence, and generate the main body of the minutes organized according to the flow of the meeting. In generating the minutes, the minutes generation unit 204 may also generate summary information that summarizes the content of the meeting. For example, it may extract the main statements made in the meeting, the key points of each agenda item, or the conclusion of the meeting as a whole, and generate a summary sentence that concisely expresses the content of the meeting. The method of summary generation is not particularly limited, and known summarization methods such as calculating sentence importance, extracting key phrases, or summarization processing using a trained natural language processing model may be applied. The generated summary information may be included in the meeting minutes data either at the beginning of the minutes or in a separate summary section.

[0066] The format of the meeting minutes data generated by the minutes generation unit 204 is not particularly limited. For example, it may be a verbatim record of what was said in chronological order, or it may be a summary of the meeting's main points. The minutes generation unit 204 may also generate meeting minutes data based on a predetermined meeting minutes format. The meeting minutes format may define items to be included, such as the meeting name, date and time, participants, agenda, summary, statements, decisions, and task list, as well as their order of arrangement. The minutes generation unit 204 may extract information corresponding to the data items defined in the format from the meeting's text data, and arrange the extracted information in correspondence with each item in the meeting minutes format to generate minutes in a predetermined format. Multiple meeting minutes formats may be available depending on the organization, type of meeting, or purpose of use, and the minutes generation unit 204 may select a format specified by any of the users (e.g., the meeting organizer or administrator) to generate the minutes.

[0067] The minutes generation unit 204 may perform a process to extract decisions, agreements, or important conclusions from the text data of the meeting. For example, it may extract matters that were finalized during the meeting as decisions by detecting statements containing predetermined keywords such as "decide," "approve," or "agree," or sentence structures that suggest a decision. The minutes generation unit 204 may generate minutes containing the extracted decisions, or it may output the extracted decisions as a list of decisions separate from the main body of the minutes. The minutes generation unit 204 may also perform a process to extract tasks or action items that should be performed after the meeting from the text data of the meeting. For example, it may detect statements containing expressions indicating actions such as "address," "consider," "confirm," or "create," and identify the person in charge of the task, its content, deadline, etc., based on speaker information or contextual information attached to the statement. The minutes generation unit 204 may generate minutes containing the extracted task information, or it may output a task list showing the extracted task information.

[0068] The minutes generation unit 204 stores the generated minutes data in the meeting information storage unit 222, linking it to the meeting identification information.

[0069] The output control unit 205 executes output control processing to present the meeting minutes data generated by the meeting minutes generation unit 204 to the user. For example, the output control unit 205 may read the meeting minutes data stored in the meeting information storage unit 222 and output the meeting minutes data in a predetermined display format to an information processing device used by meeting participants, such as the user terminal 1. In this case, the output control unit 205 may generate a user interface for displaying the meeting minutes data on a display screen.

[0070] The output control unit 205 may convert the meeting minutes data into a predetermined file format and output it. For example, it may generate meeting minutes data for presentation in text format, PDF format, spreadsheet format, or a similar electronic file format and provide it to the user terminal 1 in a state where it can be transmitted or downloaded. The output control unit 205 may also control the output content or output manner according to output conditions specified by the user. For example, it may accept specifications to output only the summary, to extract and output only the decisions or task information, or to extract and display only the statements of specific participants. Based on these specified conditions, the output control unit 205 may extract or reconstruct a portion of the meeting minutes data and output it.

[0071] The output control unit 205 may determine whether the generated meeting minutes contain confidential information, personal information, or other information that should not be disclosed, and may perform processing to make the meeting minutes presentable to the user and / or related parties. Specifically, the output control unit 205 may refer to security standard information defined by the organization to which the user belongs, identify specific information from the meeting minutes that falls under at least one of the confidential information and personal information defined in the security standard information, and apply masking processing to such specific information. The security standard information may be set in the AI-SPM filter. In this case, the output control unit 205 may identify specific information from the meeting minutes by filtering with AI-SPM.

[0072] The output control unit 205 may accept a request from the user to correct the meeting minutes data. In this case, the output control unit 205 may correct the meeting minutes data based on the user's correction request, manage the correction history of the meeting minutes data, and store the latest version of the meeting minutes data in the meeting information storage unit 222 in a state that allows for the distinction between previous versions.

[0073] <An example of an information processing method> Next, with reference to the flowchart illustrated in Figure 4, an example of an information processing method (meeting support method) performed by the meeting support system of this embodiment will be described.

[0074] First, in each user terminal 1, the model learning unit 101 executes a process to generate a personal voice model for identifying the user's speech (step SQ101). The model learning unit 101 constructs a personal voice model specifically for that user by performing machine learning on a reference model for speech identification using the user's speech for training. The generated personal voice model is stored in the model information storage unit 121 of the user terminal 1. Next, the data acquisition unit 102 of each user terminal 1 executes a process to acquire audio data during the meeting (step SQ102). For example, the data acquisition unit 102 acquires audio collected from the start to the end of the meeting via the user terminal 1 or an external sound collection device, and records the acquired audio data in the acquired data storage unit 122. Alternatively, the data acquisition unit 102 may acquire audio data during the meeting in real time.

[0075] Next, the feature extraction unit 103 of each user terminal 1 identifies the user's speech from the conference audio data and performs a process to extract the speech feature vector at the time of the user's speech (step SQ103). Specifically, the feature extraction unit 103 uses an individual speech model to identify the speech section corresponding to the user's speech and generates a speech feature vector based on that speech section. The extraction of the speech feature vector corresponding to the user's speech by the feature extraction unit 103 may be performed in real time during the conference. The transmission processing unit 105 of the user terminal 1 generates feature information by adding information indicating the time of speech, etc., to the speech feature vector extracted by the feature extraction unit 103, and performs a process to send this feature information to the management server 2 (step SQ104).

[0076] At least one user terminal 1 may generate text data from audio data, recording each statement made during the meeting in chronological order. In this case, the text conversion processing unit 104 of user terminal 1 converts the meeting audio data into text data (step SQ105). The text conversion processing unit 104 may transcribe the audio data acquired during the meeting in real time, or it may transcribe the entire meeting audio after the meeting has ended. The text conversion processing unit 104 or the transmission processing unit 105 sends the generated text data to the management server 2 (step SQ106).

[0077] The management server 2 receives feature information for each user from each user terminal. The integration processing unit 203 of the management server 2 then performs a process to integrate the feature information obtained from each user terminal 1 (step SQ107). Specifically, the integration processing unit 203 arranges and integrates multiple feature information in chronological order based on the information indicating the time of utterance attached to each feature information. After that, the integration processing unit 203 performs a process to identify the speaker of each statement in the meeting based on the integrated feature information (step SQ108). For example, the integration processing unit 203 may identify the speaker corresponding to each statement by comparing the feature information with the text data obtained by the management server 2. Alternatively, the integration processing unit 203 may generate text data corresponding to the content of the user's statement by converting the speech feature vector indicated by each feature information into a phoneme sequence. After identifying the speaker, the integration processing unit 203 attaches information indicating the identified speaker to each statement in the text data (step SQ109).

[0078] Subsequently, the minutes generation unit 204 of the management server 2 executes a process to generate meeting minutes based on text data that includes information indicating the speakers (step SQ110). The minutes generation unit 204 uses natural language processing to organize and summarize the meeting content, extract decisions and tasks, and generate meeting minutes data in a predetermined format.

[0079] Note that the flowchart in Figure 4 is merely an example, and the meeting support method of this embodiment is not limited to the example in Figure 4. Steps may be added, deleted, modified, or rearranged as appropriate.

[0080] In this embodiment of the conference support system, a personal voice model is implemented in each user terminal 1 used by each participant in the conference. The feature extraction unit 103 of each user terminal 1 uses the personal voice model to identify the speech of the user using the terminal from the conference audio data and extracts the speech feature vector when the user speaks. The transmission processing unit 105 of each user terminal 1 adds information indicating the time of speech to the speech feature vector converted by the speech identification process and transmits the feature quantity information indicating the speech feature vector to the management server 2. The integration processing unit 203 of the management server 2 aggregates the feature quantity information sent from each user terminal 1, arranges it in chronological order and integrates it to identify the speaker of each statement in the conference.

[0081] With this configuration, the system can accurately identify the speaker of each statement made during a meeting and create meeting minutes, whether it is a face-to-face offline meeting where all participants speak in the same physical space, or a hybrid online-offline meeting where some participants join online from remote locations.

[0082] In this system, instead of sending the audio data itself during the meeting to the management server 2, the voice identification process is completed locally at each user terminal 1, and the voice feature vectors extracted by this process are sent to the management server 2 in a format that does not allow voiceprint information to be obtained. As a result, it is not necessary to aggregate the user's raw audio data and biometric information that can identify specific individuals, such as voiceprint information, at the management server 2, and while the management server 2 can aggregate and process information related to the meeting, it is possible to reduce privacy risks such as the leakage of personal information.

[0083] This system estimates the possibility of changes in the user's voice quality based on reference information, including the user's schedule and status, and controls the threshold parameters for voice recognition according to the estimation result. For example, if there is a possibility of changes in voice quality due to poor health, the system can appropriately adjust the judgment conditions for voice recognition. This configuration improves the flexibility in identifying the user's spoken voice and contributes to improving the accuracy of voice recognition.

[0084] The embodiments described above are illustrative to facilitate understanding of this disclosure and are not intended to limit it. This disclosure may be modified and improved without departing from its intent, and its equivalents are included.

[0085] For example, in the above embodiment, a configuration was illustrated in which a personal voice model is generated at each user terminal 1 and voice recognition processing is performed using the personal voice model. However, the manner in which the personal voice model is generated and updated is not limited to this. For example, an initial model of the personal voice model may be generated at the management server 2 or an external information processing device, and after the initial model is distributed to each user terminal 1, additional learning or retraining may be performed at each user terminal 1. With such a configuration, it is possible to reduce the load of initial learning at the user terminal 1 while achieving individual optimization of the personal voice model.

[0086] In the above embodiment, a configuration for transmitting speech feature vectors to the management server 2 was illustrated, but the transmission unit and timing of the speech feature vectors are not necessarily limited. For example, the user terminal 1 may transmit a predetermined number of speech feature vectors corresponding to multiple utterances in batch processing, or it may be configured to transmit accumulated speech feature vectors in batches at regular intervals. Such a configuration can reduce the number of communications and lower the network load. In addition, the feature extraction unit 103 may convert the speech segments corresponding to the user's utterances in the meeting into text data and transmit feature information including the text data to the management server 2.

[0087] In the above embodiment, a configuration in which feature information is integrated in time series on the management server 2 was illustrated, but a configuration in which part or all of the integration process is performed on the user terminal 1 side is also possible. For example, one of the user terminals 1 of the multiple user terminals 1 may receive feature information from the other user terminals 1, integrate the feature information in time series, and then send it to the management server 2. Such a configuration can distribute the processing load on the management server 2.

[0088] In the above embodiment, a configuration for identifying the speaker using speech feature vectors was illustrated, but the information used for speaker identification is not limited to speech feature vectors. For example, in addition to speech feature vectors, a configuration may be used to identify the speaker by combining additional information such as the terminal's position information, orientation of the terminal, or microphone directivity information at the time of speaking. Such a configuration can improve the accuracy of speaker identification, especially in environments where multiple participants are speaking in close proximity.

[0089] In the above embodiment, a configuration was illustrated in which the minutes generation unit 204 generates minutes after the meeting has ended, but the timing of minutes generation is not limited to this. For example, during the meeting, minutes or summary information may be generated as an interim report at regular intervals or at the end of each agenda item, and presented sequentially to the user terminal 1. Such a configuration makes it possible for meeting participants to proceed with the discussion while reviewing the content during the meeting.

[0090] Furthermore, the series of processes performed by the conference support system described herein may be implemented using software, hardware, or a combination of software and hardware. It is also possible to create computer programs to implement each function of the user terminal 1 or management server 2 according to this embodiment and implement them on a PC or the like. A computer-readable recording medium containing such computer programs can also be provided. Examples of recording media include magnetic disks, optical disks, magneto-optical disks, and flash memory. Alternatively, the computer programs may be distributed without using a recording medium, for example, via a network.

[0091] Furthermore, the effects described herein are merely descriptive or illustrative and not limiting. In other words, the technology relating to this disclosure may produce other effects that are obvious to those skilled in the art from the description herein, in addition to or in lieu of the effects described herein.

[0092] The conference support system and conference support method described herein may have the following configuration. [Item 1] A conference support system comprising multiple user terminals and a management server, Each of the user terminals is: A model learning unit that learns the speech of users using the user terminal and generates a personal voice model that identifies the user's voice, A data acquisition unit that acquires audio data from meetings, A feature extraction unit that identifies the user's speech from the audio data using the personal voice model and extracts the speech feature vector when the user speaks, The system includes a transmission processing unit that transmits feature information generated by adding the speech time to the aforementioned speech feature vector to the management server. The aforementioned management server A data receiving unit that acquires the feature information from each of the user terminals, A conference support system comprising: an integration processing unit that identifies the speaker of each statement in the conference by arranging and integrating the feature information acquired from multiple user terminals in chronological order. [Item 2] The aforementioned management server The system further includes a text data acquisition unit that acquires text data recording each statement made at the aforementioned meeting in chronological order. The conference support system described in item 1, wherein the integrated processing unit identifies the speaker of each statement in the conference by comparing the feature information integrated in time series with the text data, and assigns information indicating the speaker to each statement in the text data. [Item 3] At least one of the user terminals is The system further comprises a text conversion processing unit that converts the audio data of the meeting into text data, The meeting support system according to item 2, wherein the text data acquisition unit acquires the text data from at least one user terminal. [Item 4] The aforementioned management server The meeting support system according to item 2, further comprising a minutes generation unit that generates minutes of the meeting by performing natural language processing on the text data to which information indicating the speaker has been added. [Item 5] The feature extraction unit, after extracting the speech feature vectors related to the user's speech, discards the speech data. The feature information transmitted from the transmission processing unit to the management server does not include the voice data, and is a conference support system as described in any of items 1 to 4. [Item 6] The aforementioned management server The meeting support system according to item 4, further comprising an output control unit that, by referring to security standard information stipulated by the organization to which the user belongs, identifies specific information from the meeting minutes that falls under at least one of confidential information and personal information as defined in the security standard information, and applies masking processing to said specific information. [Item 7] The aforementioned security criteria information is set in the AI-SPM filter, The output control unit identifies the specific information from the meeting minutes by filtering the AI-SPM, as described in item 6. [Item 8] The conference support system according to item 1, wherein the feature extraction unit controls a threshold parameter for identifying the user's voice by estimating the possibility of a change in the user's voice quality based on reference information including at least one of the user's schedule information and status information. [Item 9] A conference support method in which multiple user terminals and a management server perform information processing for conference support, Each of the user terminals is: This involves learning the speech of users using the user terminal and generating a personal voice model that identifies the user's voice, To acquire audio data of the meeting, Using the aforementioned personal voice model, the user's speech is identified from the voice data, and the voice feature vectors of the speech spoken by the user are extracted. The process involves sending the generated feature information, which is created by adding the utterance time to the aforementioned speech feature vector, to the management server. The aforementioned management server Obtaining the feature information from each of the user terminals, A conference support method that performs the following: identifying the speaker of each statement in the conference by arranging and integrating the feature information obtained from multiple user terminals in chronological order. [Explanation of Symbols]

[0093] 1 User terminal 101 Model Learning Department 102 Data Acquisition Unit 103 Feature Extraction Unit 105 Transmission Processing Unit 2 Management Server 2 201 Data Receiving Unit 203 Integrated Processing Unit

Claims

1. A conference support system comprising multiple user terminals and a management server, Each of the user terminals is: A model learning unit that learns the speech of users using the user terminal and generates a personal voice model that identifies the user's voice, A data acquisition unit that acquires audio data from meetings, A feature extraction unit that identifies the user's speech from the audio data using the personal voice model and extracts the speech feature vector when the user speaks, The system includes a transmission processing unit that transmits feature information generated by adding the speech time to the aforementioned speech feature vector to the management server. The aforementioned management server A data receiving unit that acquires the feature information from each of the user terminals, A text data acquisition unit acquires text data that records each statement made at the aforementioned meeting in chronological order. A conference support system comprising: an integration processing unit that arranges and integrates the feature information acquired from multiple user terminals in chronological order, compares the chronologically integrated feature information with the text data to identify the speaker of each statement in the conference, and assigns speaker-indicating information to each statement in the text data.

2. At least one of the user terminals is The system further comprises a text conversion processing unit that converts the audio data of the meeting into text data, The meeting support system according to claim 1, wherein the text data acquisition unit acquires the text data from at least one user terminal.

3. The aforementioned management server The meeting support system according to claim 1, further comprising a minutes generation unit that generates minutes of the meeting by performing natural language processing on the text data to which information indicating the speaker has been added.

4. The feature extraction unit, after extracting the speech feature vectors related to the user's speech, discards the speech data. The conference support system according to any one of claims 1 to 3, wherein the feature information transmitted from the transmission processing unit to the management server does not include the voice data.

5. The aforementioned management server The meeting support system according to claim 3, further comprising an output control unit that, by referring to security standard information stipulated by the organization to which the user belongs, identifies specific information from the meeting minutes that falls under at least one of confidential information and personal information as defined in the security standard information, and applies masking processing to said specific information.

6. The aforementioned security criteria information is set in the AI-SPM filter. The meeting support system according to claim 5, wherein the output control unit identifies the specific information from the meeting minutes by filtering the AI-SPM.

7. The conference support system according to claim 1, wherein the feature extraction unit controls a threshold parameter for identifying the user's voice by estimating the possibility of a change in the user's voice quality based on reference information including at least one of the user's schedule information and status information.

8. A conference support method in which multiple user terminals and a management server perform information processing for conference support, Each of the user terminals is: This involves learning the speech of users using the user terminal and generating a personal voice model that identifies the user's voice, To acquire audio data of the meeting, Using the aforementioned personal voice model, the user's speech is identified from the voice data, and the voice feature vectors of the speech spoken by the user are extracted. The process involves sending the generated feature information, which is created by adding the utterance time to the aforementioned speech feature vector, to the management server. The aforementioned management server Obtaining the feature information from each of the user terminals, To obtain text data that records each statement made at the aforementioned meeting in chronological order, A conference support method that involves arranging and integrating the feature information obtained from multiple user terminals in chronological order, comparing the chronologically integrated feature information with the text data to identify the speaker of each statement in the conference, and adding information indicating the speaker to each statement in the text data.

Citation Information

Patent Citations

  • Speech information control method and terminal apparatus

    JP2016029468A

  • Voice processing device, method and program

    JP2024119245A

  • SYSTEM AND METHOD FOR IMPLEMENTING AN ARTIFICIAL INTELLIGENCE SECURITY PLATFORM - Patent application

    JP2025512674A

  • Program, information processing device, and information processing method

    WO2022215140A1

  • System

    JP2025044888A