Electronic device and control method thereof
The electronic device addresses the lack of participant information in video conferencing by performing voice recognition, speaker identification, and data processing to acquire and manage speaker information, including translations, improving participant management.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-24
- Publication Date
- 2026-05-21
AI Technical Summary
Existing video conferencing systems lack the ability to identify and store participant information, necessitating a solution for acquiring and managing speaker identification and summary data during non-face-to-face interactions.
An electronic device equipped with a display, memory, and processor that performs voice recognition, speaker identification, and data processing to acquire, summarize, and store speaker information, using neural networks when necessary, and translate data for display.
Effectively identifies and stores speaker information, including names, job titles, and contact details, and translates data for display, enhancing participant management in video conferencing.
Smart Images

Figure KR2025017129_21052026_PF_FP_ABST
Abstract
Description
Electronic device and method of controlling the same
[0001] The present disclosure relates to an invention concerning an electronic device and a method for controlling the same, and more specifically, includes an electronic device and a method for controlling the same for identifying a speaker and obtaining text data regarding a voice uttered by the identified speaker.
[0002] Recently, driven by advancements in electronic technology, various types of electronic products are being developed and distributed. In particular, various display devices such as TVs, mobile phones, PCs, laptop PCs, and PDAs are widely used. Most of these display devices include a camera, a microphone, and a display, and utilize this configuration to provide video conferencing services.
[0003] In particular, with the recent increase in non-face-to-face work, the need for video call services (or video conferencing services) is steadily growing.
[0004] However, since information on video conference participants is not stored in the past, it is necessary to explore ways to provide participant information.
[0005] According to one embodiment of the present disclosure, an electronic device comprises: a display; a memory; and at least one processor; wherein, when the instructions are executed collectively or individually by the at least one processor, the electronic device may acquire audio data containing voices spoken by a plurality of speakers, perform voice recognition on the audio data to acquire text data corresponding to the voices spoken by the plurality of speakers, acquire summary data for the voices spoken by the plurality of speakers based on the text data, acquire identification information of the plurality of speakers based on the text data, and store the summary data for the voices spoken by the plurality of speakers and the identification information of the plurality of speakers by matching them.
[0006] It is determined whether the above text data contains identification information of a speaker, and if it is determined that the above text data contains identification information of a speaker, identification information of the plurality of speakers is obtained based on the identification information contained in the above text data, and the identification information may include at least one of a name, job title, telephone number, and email information.
[0007] When the above instructions are executed collectively or individually by the at least one processor, if the electronic device identifies that the text data does not contain speaker identification information, it analyzes the text data to obtain the speaker's context data, and the context data may include the content of the text data and technical terms included in the text data.
[0008] When the above instructions are executed collectively or individually by the at least one processor, if the electronic device identifies that the text data does not contain speaker identification information, it may input the text data into a learned neural network model to obtain the speaker identification information.
[0009] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may identify whether the acquired identification information exists in a database stored in the memory, and if it is identified that the acquired identification information exists, update the acquired identification information in a database containing identification information of multiple speakers.
[0010] When the above instructions are executed collectively or individually by the at least one processor, if the electronic device identifies that the acquired identification information does not exist, it may update the acquired identification information in a database containing the identification information of multiple speakers.
[0011] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may share identification information included in the updated database with an external device.
[0012] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may translate the summary data into the language of the voice signal spoken by the plurality of speakers and store the translated summary data when the language of the text data and the language of the voice signal spoken by the plurality of speakers are different.
[0013] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may control the display to display the translated summary data and identification information for the plurality of speakers.
[0014] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may control the display to display a UI requesting a user's response to the acquired identification information.
[0015] According to one embodiment of the present disclosure, a method for controlling an electronic device may include: acquiring audio data containing voices spoken by a plurality of speakers and performing voice recognition on the audio data to acquire text data corresponding to the voices spoken by the plurality of speakers; acquiring summary data for the voices spoken by the plurality of speakers based on the text data; acquiring identification information of the plurality of speakers based on the text data; and matching and storing the summary data for the voices spoken by the plurality of speakers with the identification information of the plurality of speakers.
[0016] The method includes: a step of identifying whether the text data contains identification information of a speaker; and a step of obtaining identification information of a plurality of speakers based on the identification information contained in the text data when it is identified that the text data contains identification information of a speaker, wherein the identification information may include at least one of a name, a job title, a telephone number, and email information.
[0017] If it is determined that the above text data does not contain identification information of the speaker, the method includes the step of analyzing the above text data to obtain context data of the speaker; wherein the context data may include the content of the above text data and technical terms included in the above text data.
[0018] If it is determined that the above text data does not contain speaker identification information, the method may include the step of inputting the above text data into a trained neural network model to obtain the speaker identification information.
[0019] The method may include: a step of determining whether the acquired identification information exists in a database stored in the memory; and a step of updating the acquired identification information in a database containing identification information of multiple speakers if it is determined that the acquired identification information exists.
[0020] If it is determined that the above-mentioned acquired identification information does not exist, the above-mentioned acquired identification information can be updated in a database containing identification information of multiple speakers.
[0021] It may include a step of sharing identification information included in the above-mentioned updated database with an external device.
[0022] If the language of the text data and the language of the voice signals spoken by the plurality of speakers are different, the method may include the step of translating the summary data into the language of the voice signals spoken by the plurality of speakers; and the step of storing the translated summary data.
[0023] It may include the step of controlling the display to display the translated summary data and identification information for the plurality of speakers.
[0024] It may include a step of controlling the display to display a UI requesting a user's response to the above-mentioned identification information.
[0025] FIG. 1 is a drawing for explaining an embodiment of an electronic device according to one embodiment of the present disclosure, and
[0026] FIG. 2 is a block diagram for explaining the configuration of an electronic device according to one embodiment of the present disclosure, and
[0027] FIGS. 3 to 5 are drawings for explaining a method for obtaining speaker identification information and summary data for a voice uttered by a speaker according to an embodiment of the present disclosure.
[0028] FIG. 6 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.
[0029] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.
[0030] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.
[0031] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.
[0032] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. The singular expression includes the plural expression unless the context clearly indicates otherwise.
[0033] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.
[0034] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.
[0035] Expressions such as “first,” “second,” “first,” or “second” used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.
[0036] Where it is stated that a certain component (e.g., a first component) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., a second component), it should be understood that the said certain component may be directly connected to the said other component or connected through another component (e.g., a third component).
[0037] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between said certain component and said other component.
[0038] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.
[0039] Instead, in some situations, the expression “device configured to do something” may mean that the device is “capable of doing something” together with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) capable of performing those operations by executing one or more software programs stored in a memory device.
[0040] In the embodiments, a 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for the 'module' or 'part' that needs to be implemented in specific hardware.
[0041] Meanwhile, the various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.
[0042] Hereinafter, various embodiments of the present invention will be described in detail using the attached drawings.
[0043] FIG. 1 is a drawing for illustrating an embodiment of an electronic device according to one embodiment of the present disclosure. As shown in FIG. 1, the electronic device may be implemented as a notebook, but this is merely one embodiment and may be implemented as various electronic devices such as a TV, a user terminal device or a desktop.
[0044] As illustrated in FIG. 1, when multiple speakers hold a meeting, the electronic device (100) can acquire audio data containing the voices of multiple speakers.
[0045] At this time, the electronic device (100) can convert audio data into text. Based on the data converted into text, the electronic device (100) can obtain identification information about the speaker who spoke at the meeting.
[0046] Meanwhile, 'identification information' may refer to information about the speaker used to identify the speaker. For example, 'identification information' may include one of the following: the speaker's name, job title, telephone number, email, department, company affiliation, or whether they possess expertise.
[0047] In one embodiment, the electronic device (100) can 'directly identify' the speaker's identification information included in the data converted into text. At this time, 'direct identification' may mean a method of identifying text corresponding to the speaker's identification information, such as the name, phone number, or job title included in the text.
[0048] For example, the electronic device (100) can identify information about the name 'Cheolsu' and the job title 'Manager' when the data converted into text is "I am Cheolsu, the manager of this project." At this time, the electronic device (100) can identify that the name of Participant 1 is 'Cheolsu' and the job title is 'Manager'.
[0049] In another embodiment, the electronic device (100) may ‘indirectly identify’ the speaker’s identification information when the speaker’s identification information is not directly identified in the data converted into text. In this case, ‘indirect identification’ may mean obtaining identification information by inputting the data converted into text into a neural network model. Alternatively, it may mean obtaining identification information based on whether the speaker’s speech patterns or technical terms exist by analyzing the data converted into text.
[0050] For example, when the electronic device (100) inputs data converted into text of the voice spoken by Participant 3 into a neural network model, the neural network model can output information that Participant 3 is a 'customer' based on the speech pattern of Participant 3.
[0051] As another example, the electronic device (100) can identify technical terms included in the data that converts the voice spoken by Participant 2 into text. At this time, the electronic device (100) may identify that the technical terms are terms related to the development field and then obtain identification information that Participant 2 is a 'developer'.
[0052] The electronic device (100) can identify whether the identification information included in a database containing the identification information of a plurality of users stored in memory (120) matches the identification information of a speaker obtained by the method described above.
[0053] For example, the electronic device (100) can identify a person whose name is 'Cheolsu' among the identification information of multiple users included in the database. At this time, the electronic device (100) can obtain information from the database stored in the memory (120) that 'Cheolsu's' affiliated company is 'Company A' and his email is '1@a.com'.
[0054] The electronic device (100) can update the information that the acquired Cheolsu is a ‘manager’ in the database.
[0055] However, if the electronic device (100) identifies that information about a speaker named Cheol-su is not already stored in the database, it may add information about the name and rank or occupation of the speaker named Cheol-su to the database.
[0056] The electronic device (100) can input data in which the voice spoken by Cheol-su is converted into text into a neural network model to obtain summarized data such as “scheduled to be delivered in accordance with the project result schedule.”
[0057] As illustrated in FIG. 1, the electronic device (100) can display summarized data by speaker and identification information about the speaker.
[0058] However, this is merely one example, and the electronic device (100) may also transmit the speaker's identification information to other attendees to share it with other attendees attending the meeting.
[0059] FIG. 2 is a block diagram for explaining the configuration of an electronic device (100).
[0060] The configuration illustrated in FIG. 2 is merely an example of various embodiments, and some components may be omitted or new components may be added. As illustrated in FIG. 2, the electronic device (100) may include a communication interface (110), a memory (120), a display (130), a microphone (140), and a processor (150). The configuration illustrated in FIG. 2 is merely an example, and it goes without saying that some components may be deleted or added depending on the configuration of the electronic device (100).
[0061] First, the communication interface (110) is configured to communicate with various types of external devices according to various types of communication methods. In particular, the communication interface (110) can receive audio data containing voice from an external device or data regarding a method for obtaining identification information of multiple speakers. Additionally, the communication interface (110) may receive summary data regarding voice spoken by multiple speakers or identification information of multiple speakers from an external device.
[0062] The communication interface (110) may transmit or receive to an external device a database containing identification information of multiple speakers or summary data regarding voices spoken by multiple speakers.
[0063] A wireless communication module may be a module that communicates wirelessly with an external device. For example, the wireless communication module may include at least one module among a Wi-Fi module, a Bluetooth module, an infrared communication module, an Ultra Wide-Band (UWB) module, or other communication modules.
[0064] A wired communication module may be a module that communicates with an external device via a wire. For example, a wired communication module may include at least one of a Local Area Network (LAN) module, an Ethernet module, a pair cable, a coaxial cable, or a fiber optic cable.
[0065] The memory (120) can store an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and instructions or data related to the components of the electronic device (100). In particular, the memory (120) can store a database containing identification information of multiple speakers. Additionally, the memory (120) can store a neural network model that outputs speaker identification information when text data is input. Additionally, the memory (120) can store a neural network model that outputs summarized data when text data is input. The memory (120) can store an updated database when the speaker identification information is updated.
[0066] However, for the sake of convenience of explanation, the following description assumes that the neural network model is stored in memory (120), but this is merely one embodiment and the neural network model can be stored in an external device.
[0067] Memory (120) can be implemented in various forms such as volatile memory (e.g., DRAM (dynamic RAM), SRAM (static RAM), or SDRAM (synchronous dynamic RAM), non-volatile memory (e.g., OTPROM (one time programmable ROM), PROM (programmable ROM), EPROM (erasable and programmable ROM), EEPROM (electrically erasable and programmable ROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD).
[0068] The display (130) can display various information. In particular, the display (130) can display speaker identification information and summary data regarding the voice spoken by the speaker. Additionally, the display (130) can display a UI requesting a user's response regarding the acquired speaker identification information.
[0069] The display (130) can be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, and a PDP (Plasma Display Panel). The display may also include a driving circuit, a backlight unit, etc., which can be implemented in forms such as an a-si TFT (amorphous silicon thin film transistor), an LTPS (low temperature poly silicon) TFT, and an OTFT (organic TFT). The display can be implemented as a touch screen combined with a touch sensor, a flexible display, a 3D display, a three-dimensional display, etc. According to various embodiments of the present disclosure, the display (130) may include not only a display panel that outputs an image, but also a bezel that houses the display panel.
[0070] The microphone (140) may refer to a module that acquires a voice signal and converts it into an electrical signal, and may be a condenser microphone, ribbon microphone, moving coil microphone, piezoelectric element microphone, carbon microphone, or MEMS (Micro Electro Mechanical System) microphone. Additionally, it may be implemented in omnidirectional, bidirectional, unidirectional, subcardioid, supercardioid, or hypercardioid modes. In particular, the microphone (130) can acquire audio data containing the voice spoken by the speaker.
[0071] The processor (150) may include one or more processors. Specifically, the one or more processors may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. The processor (150) may control one or any combination of other components of an electronic device and may perform operations or data processing related to communication. The one or more processors may execute one or more programs or instructions stored in memory. For example, the one or more processors may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory.
[0072] One or more processors may be implemented as a single-core processor comprising one core, or as one or more multicore processors comprising multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as cache memory or on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor. Additionally, each of the multiple cores included in the multicore processor (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one embodiment of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one embodiment of the present disclosure.
[0073] In particular, the processor (150) can acquire audio data containing voices spoken by multiple speakers and perform voice recognition on the audio data to acquire text data corresponding to the voices spoken by multiple speakers. At this time, the processor (150) can acquire summary data for the voices spoken by multiple speakers based on the text data and acquire identification information of multiple speakers based on the text data. Additionally, the processor (150) can match and store the summary data for the voices spoken by multiple speakers with the identification information of multiple speakers.
[0074] When the processor (150) identifies that the text data contains the identification information of a speaker, it can obtain the identification information of multiple speakers based on the identification information contained in the text data.
[0075] On the other hand, if the processor (150) determines that the text data does not contain the speaker's identification information, it can analyze the text data to obtain the speaker's context data. At this time, the context data may include the content of the text data and technical terms included in the text data.
[0076] Additionally, if the processor (150) determines that the text data does not contain the speaker's identification information, it may input the text data into a trained neural network model to obtain the speaker's identification information. At this time, when the neural network model inputs the text data, it may output identification information regarding the speaker's occupation or name.
[0077] The processor (150) identifies whether the acquired identification information exists in a database stored in memory (120), and if it is identified that the acquired identification information exists, it can update the acquired identification information in a database containing identification information of multiple speakers.
[0078] In contrast, if the processor (150) determines that there is no acquired identification information, it can update the acquired identification information in a database containing identification information of multiple speakers.
[0079] The processor (150) can transmit a database containing identification information for multiple speakers obtained by the method described above to an external device.
[0080] Additionally, the processor (150) can translate summary data into the language of the voice signal spoken by multiple speakers and store the translated summary data when the language of the text data and the language of the voice signal spoken by multiple speakers are different.
[0081] The processor (150) can control the display (130) to display identification information and translated summary data for multiple speakers acquired.
[0082] FIG. 3 is a flowchart illustrating the operation of an electronic device (100) according to one embodiment of the present disclosure matching and storing identification information for a plurality of speakers and summary data for voices spoken by a plurality of speakers, and then sharing it with an external device or displaying it on a display.
[0083] First, the electronic device (100) can acquire audio data containing voice signals spoken by multiple speakers (S305). For convenience of explanation, an embodiment in which the electronic device (100) acquires data containing voice signals using a microphone (140) will be described later, but this is merely one embodiment, and it goes without saying that the electronic device (100) can load data containing voice signals stored in memory (120).
[0084] The electronic device (100) can identify voice segments included in the acquired data.
[0085] In one embodiment, the electronic device (100) can detect a voice segment by using the energy of a voice signal to identify a voice segment.
[0086] In another embodiment, the electronic device (100) may detect a voice segment by using a method for detecting a voice segment using a zero crossing rate for each voice signal, extracting a feature vector from a voice signal, and determining the presence or absence of a voice signal from the extracted feature vector using a Support Vector Machine (SVM) to detect a voice segment.
[0087] However, this is only one example, and it goes without saying that the electronic device (100) can identify voice segments using the MFCC (Mel-Frequency Cepstral Coefficients) method.
[0088] The electronic device (100) can distinguish multiple speakers by identifying a voice segment by the method described above and extracting characteristic information of the speaker who output the voice signal (S310). For convenience of explanation, the process of distinguishing multiple speakers who uttered the voice is defined as 'speaker diarization'.
[0089] When distinguishing a speaker, the electronic device (100) can extract speaker feature information. In this case, the speaker feature information can be a speaker embedding vector. Specifically, the electronic device (100) may use acoustic signal analysis methods such as Mel-Frequency Cepstral Coefficients (MFCC), spectrograms, and f0 pitch information to extract speaker feature information from a voice signal. In this case, information including not only the physical characteristics of the voice signal but also unique features that distinguish a speaker, such as speech habits, timbre, and pronunciation characteristics, can be extracted.
[0090] The electronic device (100) can convert the extracted feature information into a speaker embedding vector by mapping it into a low-dimensional eigenvector space. The speaker embedding vector can be represented by compressing the speaker's voice features into a single vector. This vector is configured to be distinguishable from other speakers and can reflect unique characteristics such as the speaker's timbre or pronunciation style.
[0091] The electronic device (100) can identify speech segments for each speaker using the acquired speaker embedding vector. For example, if three speakers speak alternately in the voice data, three clusters are generated through clustering, and each cluster can represent a specific speaker.
[0092] Additionally, the electronic device (100) can separate the voice signals by speaker when multiple speakers simultaneously output voice signals (S315). For convenience of explanation, the process of separating the voice signals is defined as 'speaker separation'. The electronic device (100) may use a deep learning-based voice separation model for speaker separation.
[0093] The electronic device (100) can identify the speaker after separating the voice signal by speaker using the method described above and obtaining feature information of the separated voice signal. Specifically, the electronic device (100) can match the speaker by extracting a feature vector from the separated voice signal. Meanwhile, since the method for extracting an embedding vector from the voice signal has been described above, a detailed explanation will be omitted.
[0094] At this time, the electronic device (100) can compare the extracted embedding vector with the already stored speaker embedding vector and match with a speaker with a high similarity.
[0095] In one embodiment, the electronic device (100) can match with a speaker with high similarity using a cosine similarity method. In another embodiment, the electronic device (100) can match with a speaker with high similarity using a Euclidean distance method.
[0096] By the method described above, the electronic device (100) can match a speaker corresponding to a separated voice signal.
[0097] The electronic device (100) can perform voice recognition on audio data to obtain text data corresponding to voices spoken by multiple speakers (S320).
[0098] Specifically, the electronic device (100) can convert multiple separated voice signals into text using an Automatic Speech Recognition (ASR) method. ASR is a voice recognition system and is a technology that converts voice signals into text using voice signals.
[0099] The electronic device (100) can obtain summary data for voices spoken by multiple speakers based on text data (S325).
[0100] Specifically, the electronic device (100) can obtain summary data by inputting text data into a neural network model. At this time, the neural network model can be a Large Language Model (LLM).
[0101] Specifically, when text data is input into the neural network model, it can be a model trained to analyze the input text data, identify the structure and meaning of the sentences, summarize the input text, and output summarized data.
[0102] The electronic device (100) can identify whether the text data contains the speaker's identification information. At this time, if the electronic device (100) identifies that the text data contains the speaker's identification information (S330-Y), it can obtain the identification information of multiple speakers based on the identification information contained in the text data (S335).
[0103] For example, text data may include the text "Hello, my name is Younghee, and I am a developer. If you have any further questions, I would appreciate it if you could contact me at 010-2222-2222."
[0104] At this time, the electronic device (100) can identify the speaker's identification information as 'Young-hee', occupation or department affiliation as 'developer', and phone number as '010-2222-2222'.
[0105] However, if the electronic device (100) identifies that the text data does not contain the speaker's identification information (S330-N), it can analyze the text data to obtain the speaker's context data (S340). At this time, the context data may be technical terms or speech patterns, etc.
[0106] For example, if the text data contains design-related technical terms, the electronic device (100) can identify the speaker as a speaker belonging to a design team.
[0107] As another example, Speaker 1 may repeatedly use the text “frankly speaking.” In this case, the electronic device (100) may identify the text “frankly speaking” as the voice spoken by Speaker 1 when it is identified.
[0108] Alternatively, if the electronic device (100) identifies that the text data does not contain the speaker's identification information, it can input the text data into a learned neural network model to obtain the speaker's identification information.
[0109] When the electronic device (100) inputs text data into a neural network model, it can obtain information about the speaker's occupation or rank.
[0110] For example, if the electronic device (100) inputs text data containing the content “Database optimization work is needed” or “I finished the code review for this project” into a neural network model, it can obtain information that the speaker’s occupation is ‘developer’ and the department to which they belong is ‘development team’.
[0111] When the electronic device (100) inputs text data into a learned neural network model or obtains speaker identification information using context data, it can control the display (130) to display a UI requesting a user's response to the obtained identification information.
[0112] This will be described later with reference to Fig. 4.
[0113] As illustrated in FIG. 4, when the electronic device (100) obtains speaker identification information using context information or a neural network model, it can control the display (130) to display a UI requesting a user's response.
[0114] For example, the electronic device (100) may display a UI containing the text “Is the participant currently speaking “Cheolsu”?” when the speaker is identified as “Cheolsu” using context information included in text data or a neural network model.
[0115] However, this is merely one example, and it goes without saying that a separate step of obtaining a user's response regarding the speaker's identification information may not be required.
[0116] After the electronic device (100) obtains identification information of multiple speakers, it can update the obtained identification information in a database (S345).
[0117] Specifically, the electronic device (100) identifies whether the acquired identification information exists in a database stored in memory (120), and if it is identified that the acquired identification information exists, it can update the acquired identification information in a database containing identification information of multiple speakers.
[0118] For example, if the electronic device (100) identifies among the acquired speaker identification information that the name is "JANE" and the affiliated company is Company B, it can identify whether there is a speaker whose name is "JANE" and whose affiliated company is B in a database containing identification information.
[0119] If the electronic device (100) identifies that a speaker named "JANE" is stored in the database but that the information that the affiliated company is "Company B" is not stored, it can update the information that "Company B" is stored in the database.
[0120] On the other hand, if the electronic device (100) determines that there is no acquired identification information, it can update the acquired identification information in a database containing identification information of multiple speakers.
[0121] For example, the electronic device (100) can identify among the acquired speaker's identification information that the name is "JANE" and the affiliated company is Company B. At this time, the electronic device (100) can identify that there is no speaker whose name is "JANE" and whose affiliated company is Company B in the database containing the identification information. The electronic device (100) can update the information regarding the speaker's name and company in the database.
[0122] After acquiring identification information of multiple speakers, the electronic device (100) can store summary data of voices spoken by multiple speakers and identification information of multiple speakers by matching them (S350).
[0123] For example, among the speaker's identification information, information such as 'Cheolsu', 'Manager', and '010-1111-1111' can be matched and stored with summary data about the voice spoken by Cheolsu, such as 'Project results to be delivered according to schedule'.
[0124] The electronic device (100) can display identification information and summary data of a plurality of matched speakers through a display (120) or transmit them to an external device.
[0125] FIG. 5 is a drawing for explaining an embodiment of providing summary data in different languages for each speaker according to one embodiment of the present disclosure.
[0126] Specifically, the electronic device (100) can translate summary data into the language of the voice signal spoken by multiple speakers when the language of the text data and the language of the voice signal spoken by multiple speakers are different. Additionally, the electronic device (100) can store the translated summary data.
[0127] For example, the electronic device (100) can identify that the voice spoken by the speaker is in English. At this time, the electronic device (100) can identify that the text data obtained by performing speech recognition on the audio data spoken by the speaker is in Korean. The electronic device (100) can translate the summary data into English and store the translated summary data.
[0128] At this time, as illustrated in FIG. 5, the electronic device (100) can display summary data translated into English through the display (130).
[0129] FIG. 6 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.
[0130] First, the electronic device (100) can acquire data containing a voice signal (S610). At this time, the electronic device (100) can acquire the data through a microphone, but this is only one embodiment, and it is obvious that it can load data stored in memory (120).
[0131] The electronic device (100) can perform voice recognition on audio data to obtain text data corresponding to voices spoken by multiple speakers (S620).
[0132] As the method for acquiring text data corresponding to speech uttered by multiple speakers has been described in detail above, a detailed explanation will be omitted.
[0133] Additionally, the electronic device (100) can obtain summary data for voices spoken by multiple speakers based on text data (S630). At this time, the electronic device (100) can obtain summary data by inputting the text data into an artificial intelligence model.
[0134] The electronic device (100) can obtain identification information of multiple speakers based on text data (S640).
[0135] The electronic device (100) can store summary data for voices spoken by multiple speakers and identification information of multiple speakers by matching them (S650).
[0136] Additionally, methods according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user (20) devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0137] A method according to various embodiments of the present disclosure may be implemented as software comprising instructions stored on a machine-readable storage medium (e.g., a computer). The machine may include a server device or an electronic device according to the disclosed embodiments, which is a device capable of calling instructions stored from the storage medium and operating according to the called instructions.
[0138] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory readable recording medium. Here, 'non-transitory readable recording medium' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.
[0139] When the above instruction is executed by a processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or an interpreter.
[0140] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.
Claims
1. In an electronic device, display; Memory; and Includes at least one processor; and When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Acquire audio data containing voices spoken by multiple speakers, and Speech recognition is performed on the above audio data to obtain text data corresponding to the speech uttered by the plurality of speakers, and Based on the above text data, summary data for the voices spoken by the plurality of speakers is obtained, and Based on the above text data, identification information of the plurality of speakers is obtained, and An electronic device that stores summary data of voices spoken by the plurality of speakers and identification information of the plurality of speakers by matching them.
2. In Paragraph 1, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Identify whether the above text data includes speaker identification information, and If it is determined that the above text data contains speaker identification information, the identification information of the plurality of speakers is obtained based on the identification information contained in the above text data, and The above identification information is an electronic device comprising at least one of a name, job title, telephone number, and email information.
3. In Paragraph 2, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, If it is determined that the above text data does not contain speaker identification information, the above text data is analyzed to obtain the speaker's context data, and Based on the above context data, identification information of the above speaker is obtained, and The above context data is an electronic device containing speech patterns or technical terms.
4. In Paragraph 3, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that inputs the text data into a trained neural network model to obtain the speaker's identification information when it is determined that the text data does not contain the speaker's identification information.
5. In Paragraph 4, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Identify whether the acquired identification information exists in the database stored in the memory, and An electronic device that updates the acquired identification information in a database containing identification information of multiple speakers when it is determined that the acquired identification information exists.
6. In Paragraph 5, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that, if it is determined that the above-mentioned acquired identification information does not exist, updates the above-mentioned acquired identification information in a database containing identification information of multiple speakers.
7. In Paragraph 6, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that transmits identification information included in the above-mentioned updated database to an external device.
8. In Paragraph 1, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, If the language of the text data and the language of the voice signal uttered by the plurality of speakers are different, the summary data is translated into the language of the voice signal uttered by the plurality of speakers, and An electronic device that stores the above-mentioned translated summary data and the above-mentioned identification information of multiple speakers by matching them.
9. In Paragraph 8, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device for controlling the display to display the above-mentioned translated summary data and identification information for the above-mentioned plurality of speakers.
10. In Paragraph 4, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that controls the display to display a UI requesting a user's response to the above-mentioned acquired identification information.
11. A method for controlling an electronic device, Acquire audio data containing voices spoken by multiple speakers, and A step of performing speech recognition on the above audio data to obtain text data corresponding to the speech uttered by the plurality of speakers; A step of obtaining summary data for voices spoken by the plurality of speakers based on the above text data; A step of obtaining identification information of the plurality of speakers based on the text data above; and A control method comprising the step of matching and storing summary data for voices spoken by the plurality of speakers with identification information of the plurality of speakers.
12. In Paragraph 11, A step of identifying whether the above text data includes speaker identification information; and If it is determined that the above text data contains identification information of a speaker, the method includes the step of obtaining identification information of the plurality of speakers based on the identification information contained in the above text data; The above identification information is a control method comprising at least one of a name, job title, telephone number, and email information.
13. In Paragraph 12, If it is determined that the above text data does not contain speaker identification information, a step of analyzing the above text data to obtain the speaker's context data; and The method includes the step of obtaining identification information of the speaker based on the context data obtained above; The above context data is a control method comprising the content of the text data and specialized terms included in the text data.
14. In Paragraph 13, A control method comprising the step of, if it is determined that the above text data does not contain speaker identification information, inputting the above text data into a learned neural network model to obtain the speaker identification information.
15. In Paragraph 14, A step of identifying whether the acquired identification information exists in the database stored in the memory; and A control method comprising the step of updating the acquired identification information in a database containing identification information of multiple speakers when it is determined that the acquired identification information exists.