Method for providing group call service and electronic device supporting the same

KR103012793B1Active Publication Date: 2026-09-02SAMSUNG ELECTRONICS CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
KR1020210027314
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-02
Publication Date
2026-09-02
Estimated Expiration
2041-03-02

Smart Images

  • Figure 112021024266759-PAT00003_ABST
    Figure 112021024266759-PAT00003_ABST
Patent Text Reader

Abstract

An electronic device according to various embodiments includes a communication module and a processor operatively connected to the communication module, wherein the processor receives and stores a first spoken voice associated with at least a first external device and a second spoken voice associated with a second external device, and when detecting a single speech based on the first spoken voice and the second spoken voice, transmits the first spoken voice or the second spoken voice having a first playback speed to the at least first electronic device and the second electronic device, and when detecting a simultaneous speech based on the first spoken voice and the second spoken voice, the processor may be configured to convert at least a portion of a synthesized voice in which at least a first superimposed speech of the first spoken voice and at least a second superimposed speech of the second spoken voice are connected in succession to a second playback speed different from the first playback speed and transmit it to the at least first electronic device and the second electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The various embodiments disclosed in this document relate to electronic devices, specifically to a method for providing group call services and an electronic device supporting the same. Background Technology

[0002] With the development of digital technology, a wide variety of electronic devices capable of communication and personal information processing while on the move are being released, such as mobile communication terminals, electronic notebooks, smartphones, tablet PCs, laptop PCs (or notebook PCs), and wearable devices. Through rapid technological advancements, these electronic devices have evolved from simple voice calls and short message transmission functions to include various functions such as video calls, electronic notebook functions, document functions, email functions, and internet functions.

[0003] Meanwhile, recent electronic devices provide group calling services that allow at least two people to speak simultaneously. Group calling services are used by people in different locations to socialize personally through voice or video calls, or for business purposes such as remote video conferencing. The problem to be solved

[0004] In the operation of the aforementioned conventional group call service, the electronic device can acquire the spoken voice of a speaker and transmit it to the electronic device of another speaker participating in the group call, or receive the spoken voice of another speaker.

[0005] However, in cases where simultaneous speech occurs in which two or more speakers speak practically at the same time, the utterances of the simultaneous speakers may be transmitted in an overlapping manner, and in this process, some of the utterances of the simultaneous speakers may be lost.

[0006] Accordingly, at least one of the various embodiments provides a method for providing a group call service that generates and transmits a synthetic voice in which the spoken voices of simultaneous speakers are connected in a continuous manner when simultaneous speech occurs, and an electronic device that supports the same. means of solving the problem

[0007] An electronic device according to various embodiments includes a communication module and a processor operatively connected to the communication module, wherein the processor receives and stores a first spoken voice associated with at least a first external device and a second spoken voice associated with a second external device, and when detecting a single speech based on the first spoken voice and the second spoken voice, transmits the first spoken voice or the second spoken voice having a first playback speed to the at least first electronic device and the second electronic device, and when detecting a simultaneous speech based on the first spoken voice and the second spoken voice, the processor may be configured to convert at least a portion of a synthesized voice in which at least a first superimposed speech of the first spoken voice and at least a second superimposed speech of the second spoken voice are connected in succession to a second playback speed different from the first playback speed and transmit it to the at least first electronic device and the second electronic device.

[0008] A method of operation of an electronic device according to various embodiments may include: receiving and storing a first spoken voice associated with at least a first external device and a second spoken voice associated with a second external device; detecting a single spoken voice or a simultaneous spoken voice based on the first spoken voice and the second spoken voice; when the single spoken voice is detected, transmitting the first spoken voice or the second spoken voice having a first playback speed to the at least first electronic device and the second electronic device; and when the simultaneous spoken voice is detected, converting at least a portion of a synthetic voice in which at least a first superimposed utterance of the first spoken voice and at least a second superimposed utterance of the second spoken voice are connected in succession to a second playback speed different from the first playback speed and transmitting it to the at least first electronic device and the second electronic device.

[0009] An electronic device according to various embodiments may include a communication module, a microphone, an output module, and a processor operatively connected to the communication module, the microphone, and the output module, wherein the processor transmits a spoken voice obtained through the microphone to at least a first communication device and a second communication device, receives a first spoken voice obtained by the first communication device and a second spoken voice obtained by the second communication device, detects a single spoken voice or a simultaneous spoken voice based on the received first spoken voice and the second spoken voice, and when the single spoken voice is detected, outputs the first spoken voice or the second spoken voice having a first playback speed through the output module, and when the simultaneous spoken voice is detected, generates a synthetic voice in which at least a first superimposed speech of the first spoken voice and at least a second superimposed speech of the second spoken voice are connected in succession and outputs it at a second playback speed different from the first playback speed. Effects of the invention

[0010] The electronic device according to the various embodiments disclosed in this document can support the clear transmission of each speaker's voice without overlap by generating and transmitting a synthesized voice formed by sequentially connecting the voices of the simultaneous speakers when simultaneous speech occurs while providing a group call service.

[0011] The effects obtainable from this document are not limited to those mentioned above. Brief explanation of the drawing

[0012] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments. FIG. 2a is a schematic diagram illustrating the configuration of a group call system according to various embodiments. FIG. 2b is a schematic diagram illustrating the configuration of an external device according to various embodiments. FIG. 3a is a diagram illustrating the operation of acquiring (or extracting) overlapping utterances from an external device according to various embodiments. FIG. 3b is a diagram illustrating the operation of generating synthesized speech from an external device according to various embodiments. FIG. 3c is a diagram illustrating the operation of playing synthesized speech from an external device according to various embodiments. FIGS. 3D and FIGS. 3E are drawings for illustrating different operations of generating synthesized speech from an external device according to various embodiments. FIG. 4 is a flowchart illustrating the operation of providing a group call service in an electronic device according to various embodiments. FIG. 5 is a flowchart illustrating the operation of acquiring superimposed ignition in an electronic device according to various embodiments. FIG. 6 is a flowchart illustrating different operations for obtaining superimposed ignition in an electronic device according to various embodiments. FIG. 7 is a flowchart illustrating the operation of determining the speech rate of a synthesized voice in an electronic device according to various embodiments. FIG. 8 is a diagram illustrating the operation of a group call system according to various embodiments. FIGS. 9a and 9b are drawings illustrating different operations of a group call system according to various embodiments. FIG. 10 is a diagram illustrating another operation of a group call system according to various embodiments. FIG. 11 is a diagram illustrating the operation of setting parameters of a synthesized speech according to various embodiments. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Specific details for implementing the invention

[0013] Hereinafter, various embodiments of this document are described with reference to the accompanying drawings. However, this is not intended to limit the technology described in this document to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives to the embodiments of this document. In relation to the description of the drawings, similar reference numerals may be used for similar components.

[0015] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or may communicate with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0016] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), for example, and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0017] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0018] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0019] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0020] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0021] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0022] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0023] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102), speaker or headphones, etc.) connected directly or wirelessly to the electronic device (101).

[0024] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0025] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0026] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0027] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0028] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0029] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0030] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0031] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0032] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.

[0033] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0034] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0035] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0036] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0037] The electronic device (101) according to the various embodiments disclosed in this document may be of various forms. The electronic device (101) may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. The electronic device (101) according to the embodiments of this document is not limited to the aforementioned devices.

[0038] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, each of phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as “first,” “second,” or “first” or “second” may be used simply to distinguish a component from another component and do not limit the components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationally,” it means that said component may be connected to said other component directly (e.g., wired), wirelessly, or through a third component.

[0039] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0040] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0041] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0042] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations among the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0044] FIG. 2a is a schematic diagram illustrating the configuration of a group call system (200) according to various embodiments.

[0045] Referring to FIG. 2a, a group call system (200) according to various embodiments may be composed of a plurality of electronic devices (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) and an external device (240). Each electronic device (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) may communicate with the external device (240) through a network (e.g., a second network (199)).

[0046] According to various embodiments, each electronic device (e.g., first electronic device (210), second electronic device (220), and third electronic device (230)) may acquire the spoken voice of a speaker and transmit it to another speaker's electronic device participating in the group call, or receive the spoken voice of another speaker. According to one embodiment, each electronic device (e.g., first electronic device (210), second electronic device (220), and third electronic device (230)) may transmit the user's spoken voice received through a microphone to another electronic device via an external device (240). Additionally, each electronic device (e.g., first electronic device (210), second electronic device (220), and third electronic device (230)) may receive the spoken voice acquired by another electronic device via the external device (240).

[0047] According to various embodiments, the external device (240) may include at least one server device that provides a group call service enabling a plurality of electronic devices (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) to make a call simultaneously, or a portable electronic device capable of providing the group call service. According to one embodiment, an external device (240) may assign (or form) a channel (e.g., audio channel and / or video channel) to each electronic device (e.g., first electronic device (210), second electronic device (220), and third electronic device (230)) participating in a group call, and may receive spoken voice from each electronic device (e.g., first electronic device (210), second electronic device (220), and third electronic device (230)) or transmit spoken voice to each electronic device (e.g., first electronic device (210), second electronic device (220), and third electronic device (230)) through the assigned channel. For example, the external device (240) may assign a first channel to the first electronic device (210), assign a second channel to the second electronic device (220), and assign a third channel to the third electronic device (230). Additionally, the external device (240) can transmit the spoken voice of the first electronic device (210) received through the first channel to the second electronic device (220) and the third electronic device (230) through the second and third channels. Similarly, the external device (240) can transmit the spoken voice of the second electronic device (220) received through the second channel to the first electronic device (210) and the third electronic device (230) through the first and third channels.

[0048] According to various embodiments, the external device (240) may detect simultaneous speech while providing a group call service. Simultaneous speech may be a situation in which at least two speakers speak simultaneously or at close proximity, causing the speech voices of two or more simultaneous speakers to overlap. According to one embodiment, when simultaneous speech is detected, the external device (240) may generate a synthetic voice processed to sequentially play the speeches overlapped by the simultaneous speech and provide it to at least one electronic device participating in the group call (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)). The synthetic voice is formed by connecting overlapping speeches separated from each speech voice, thereby preventing the speech voice of a specific speaker from being overlapped and lost due to simultaneous speech.

[0050] FIG. 2b is a schematic diagram illustrating the configuration of an external device (240) according to various embodiments. FIG. 3a is a diagram illustrating the operation of acquiring (or extracting) a superimposed utterance in an external device (240) according to various embodiments, FIG. 3b, FIG. 3d, and FIG. 3e are diagrams illustrating the operation of generating a synthesized voice in an external device (240) according to various embodiments, and FIG. 3c is a diagram illustrating the operation of playing a synthesized voice in an external device (240) according to various embodiments.

[0051] Referring to FIG. 2b, an external device (240) according to various embodiments may correspond to at least one of the electronic device (101) described above through FIG. 1, an external electronic device (e.g., electronic device (102), electronic device (104)), and a server (108).

[0052] According to various embodiments, the external device (240) may include a communication module (2410) (e.g., communication module (190)), a processor (2420) (e.g., processor (120)), and a memory (2430) (e.g., memory (130)). However, this is merely exemplary, and the aforementioned configurations are not limited to being essential configurations of the external device (200). For example, the external device (240) may be implemented with more configurations than those shown in FIG. 2b, or with fewer configurations. For example, the external device (240) may be configured to include configurations such as at least one input module (e.g., input module (150)), at least one display module (e.g., display module (160)), at least one sensor module (e.g., sensor module (176)), or a power management module (e.g., power management module (188)).

[0053] The communication module (2410) may support communication with at least one electronic device (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)). According to various embodiments, the communication module (2410) may be a device comprising hardware and software for transmitting and receiving signals (e.g., commands or data) between at least one electronic device (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) and an external device (240).

[0054] The processor (2420) can be operatively connected to the communication module (2410) and memory (2430) and can control various components (e.g., hardware or software components) of the external device (240).

[0055] According to various embodiments, the processor (2420) may provide a group call service so that a plurality of electronic devices (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) can make a call simultaneously. According to one embodiment, the processor (2420) may transmit a spoken voice received from at least one electronic device (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) to at least one other electronic device participating in the group call while the group call is in progress.

[0056] According to various embodiments, the processor (2420) can detect a simultaneous speech situation in which at least two speakers speak substantially simultaneously during a group call, and the speech sounds of two or more simultaneous speakers overlap. According to one embodiment, the processor (2420) can detect simultaneous speech based on time information in which speech sounds are received through each channel. For example, the processor (2420) can detect that simultaneous speech occurs when speech sounds are received through at least two channels at simultaneous or close times.

[0057] According to various embodiments, the processor (2420) may generate a synthetic voice based on the spoken voices of the simultaneous speakers when simultaneous speech occurs, and transmit the generated synthetic voice to the group call participants (e.g., the first electronic device (210), the second electronic device (220), and the third electronic device (230)). The synthetic voice may be data processed so that each spoken voice uttered by the simultaneous speakers is played sequentially.

[0058] According to one embodiment, when a first spoken voice and a second spoken voice are substantially received simultaneously, the processor (2420) can obtain a first overlapping voice from the first spoken voice and obtain a second overlapping voice from the second spoken voice. For example, as shown in FIG. 3a, assuming a situation where a first spoken voice (310) (e.g., "Yes. Now I understand") spoken from time t1 to time t4 is received from a first electronic device (210) and a second spoken voice (320) (e.g., "I see your point") spoken from time t3 to time t5 is received from a second electronic device (220), the processor (2420) can determine time t3 to time t4 as a simultaneous spoken voice interval. Accordingly, the processor (2420) can acquire the first speech voice (310) from time t3 to time t4 as the first overlapping speech (312) (e.g., I understand) and the second speech voice (320) from time t3 to time t4 as the second overlapping speech (322) (e.g., I see).

[0059] According to one embodiment, the processor (2420) can generate a synthesized speech (330) (e.g., I understand I see) by sequentially connecting a first superimposed utterance (312) obtained from a first speech voice (310) and a second superimposed utterance (322) obtained from a second speech voice (320), as illustrated in FIG. 3B. In this regard, the processor (2420) may optionally or additionally add a silent interval (e.g., a short pause interval or a silence interval) (332) of a specified length between the first superimposed utterance (312) and the second superimposed utterance (322).

[0060] As another example, the processor (2420) may obtain additional utterances based on at least one overlapping utterance and use them for generating synthetic speech. For example, as illustrated in FIG. 3b, the processor (2420) may obtain a portion (314) (e.g., Now) corresponding to a certain time before (and / or after) the first overlapping utterance (312) (e.g., I understand) among the first speech voices (e.g., Yes. Now I understand) as a first additional utterance and connect it to the first overlapping utterance (312), and obtain a portion (324) (e.g., your point) corresponding to a certain time before (and / or after) the second overlapping utterance (322) (e.g., I see) among the second speech voices (e.g., I see your point) as a second additional utterance and connect it to the second overlapping utterance (322) to generate synthetic speech (340) (e.g., Now I understand I see your point). Accordingly, the synthetic voice (340) including additional utterances (314, 324) can provide the effect of clearly conveying the context before and after the overlapping utterances compared to the synthetic voice (330) composed only of overlapping utterances (312, 322).

[0061] According to various embodiments, the processor (2420) may determine the playback order of overlapping utterances based on the utterance order of simultaneous speakers. For example, if the utterance time of the first utterance voice (310) (e.g., t1) is earlier than the utterance time of the second utterance voice (320) (e.g., t3), the processor (2420) may generate a synthesized voice (e.g., I understand I see) such that the first overlapping utterance (e.g., I understand) (312) is played before the second overlapping utterance (e.g., I see) (322). Conversely, if the utterance time of the second speech voice (320) is earlier than the utterance time of the first speech voice (310), for example, as illustrated in FIG. 3d, if the utterance time of the second speech voice (230) is substantially earlier than the utterance time of the first speech voice (310) due to an interjection (e.g., hmm) (327), the processor (2420) may generate a synthesized voice (e.g., I see I understand) so that the second overlapping utterance (e.g., I see) (322) is played before the first overlapping utterance (e.g., I understand) (312). However, this is merely illustrative and the present document is not limited thereto. For example, the playback order of the overlapping utterances may be determined by considering the utterance speed, utterance size, etc. of the simultaneous speakers.

[0062] Additionally, as illustrated in FIG. 3e, if the utterance time of the second utterance voice (230) itself is earlier than the utterance time of the first utterance voice (310) itself, the processor (2420) may obtain the overlapping portion (e.g., your point) (328) of the second utterance voice (e.g., I see your point) (320) that overlaps with the first utterance voice (e.g., Yes. Now I understand) (310) as the second overlapping utterance, and obtain the overlapping portion (e.g., Yes. Now) (318) of the first utterance voice (310) that overlaps with the second utterance voice (320) as the first overlapping utterance to generate a synthesized voice (e.g., your point Yes, Now).

[0063] According to various embodiments, the synthesized speech may be played back through at least one electronic device participating in the group call (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) after the simultaneous speech has ceased. For example, as illustrated in 350 of FIG. 3c, a first overlapping speech (e.g., I understand) (352) in the t'1 to t'2 interval of the synthesized speech (350) and a second overlapping speech (e.g., I see) (354) in the t'3 to t'4 interval of the synthesized speech (350) may be played back at a first rate (or speech rate). The first rate may be substantially the same rate as the speaker's speech rate (normal rate or standard rate) (e.g., 1x rate). Such playback of the synthesized speech (350) may delay the speaker's subsequent speech. Accordingly, the processor (2420) can reduce the delay of subsequent speech and allow smooth group calls to proceed by processing at least one superimposed speech included in the synthesized speech (350), as shown in 360 to 390 of FIG. 3c, to be played back at a second speed (e.g., a speed greater than 1.2 times and less than 2 times, e.g., a speed of 1.5 times) faster than the speech rate of the speaker (e.g., a first speed).

[0064] For example, as illustrated through 360 in FIG. 3c, the processor (2420) can process the first overlapping utterance (352) (e.g., I understand I see) of the synthesized speech (350) (e.g., I understand I see) to be played at a first speed, and the second overlapping utterance (354) (e.g., I see) to be played at a second speed (362). As previously mentioned, the period between the first overlapping utterance (e.g., I understand) and the second overlapping utterance (e.g., I see), that is, during the t'2 to t'3 interval of the synthesized speech (350), may be provided as a silent period.

[0065] As another example, as illustrated through 370 in FIG. 3c, the processor (2420) may process so that part (372) (e.g., I or understand) of the first overlapping utterance (352) (e.g., I understand) is played at a first speed, and only another part (374) (e.g., understand or I) of the first overlapping utterance (352) is played at a second speed. At this time, the processor (2420) may process so that at least part of the second overlapping utterance (354) (e.g., I see) is played at a first speed or at a second speed.

[0066] As another example, as illustrated through 380 in FIG. 3c, the processor (2420) can process the second overlapping utterance (354) (e.g., I see) in addition to the first overlapping utterance (352) (e.g., I understand) to be played back at a second speed (382, 384).

[0067] As another example, as illustrated by 390 in FIG. 3c, the processor (2420) may process the second overlapping utterance (354) (e.g., I see) in addition to the first overlapping utterance (352) (e.g., I understand) to be played back at a second speed (392, 394), and may remove the silent interval between the first overlapping utterance (392) and the second overlapping utterance (394). However, this is merely illustrative and is not limited to this. For example, the processor (2420) may adjust the speech rate of the overlapping speech by removing the silent interval (e.g., silent interval between words) present in the first overlapping utterance (352) (e.g., I understand) or the second overlapping utterance (354) (e.g., I see).

[0068] In the above-described embodiment, an embodiment was described in which overlapping utterances are separated from each utterance of simultaneous speakers to generate synthesized speech. However, this is merely illustrative and the present document is not limited thereto. For example, the processor (2420) may generate synthesized speech by connecting overlapping utterances obtained from the speech of one speaker in succession with the speech of another speaker.

[0069] For example, when a first speech voice (310) (e.g., Yes. Now I understand) is received from a first electronic device (210) and a second speech voice (320) (e.g., I see your point) is received from a second electronic device (220), the processor (2420) may generate a synthesized voice (e.g., I see your point I understand or I understand I see your point) by connecting a first superimposed speech (312) (e.g., I understand) of the first speech voice (310) to the second speech voice (320). In this regard, the processor (2420) may process at least a portion of the synthesized voice (e.g., at least a portion of the second speech voice (320) or at least a portion of the first superimposed speech (312)) to be played back at a second speed faster than the first speed. Conversely, the processor (2420) can generate a synthesized speech (e.g., I see) by connecting a second superimposed utterance (322) (e.g., I see) of the second speech (320) to the first speech (310) (e.g., I see Yes. Now I understand or Yes. Now I understand I see).

[0070] As another example, when a first spoken voice (310) (e.g., Yes. Now I understand) is received from a first electronic device (210) and a second spoken voice (320) (e.g., I see your point) is received from a second electronic device (220), the processor (2420) may generate a synthesized voice (e.g., Yes. Now I understand I see your point) composed of a first segment of the first spoken voice (310) that is not overlapping (e.g., Yes. Now), a second segment of the first spoken voice and the second spoken voice that is overlapping (e.g., I understand I see), and a third segment of the second spoken voice (320) that is not overlapping (e.g., your point). In this regard, the playback speed of at least some segments of the synthesized voice may be adjusted. For example, the first section (e.g., Yes. Now) and the third section (e.g., your point) may be played back at a first speed that is substantially the same as the speed of the original speech, and the overlapping speech of the second section (e.g., I understand I see) may be played back at a second speed that is faster than the first speed. As another example, the processor (2420) may increase the playback speed of the first section (e.g., Yes. Now) and the third section (e.g., your point) to secure a playback interval (e.g., a silent interval) between the first section (e.g., Yes. Now) and the second section (e.g., I understand I see) or a playback interval between the second section (e.g., I understand I see) and the third section (e.g., your point), or secure a playback interval between the first overlapping speech (e.g., I understand) and the second overlapping speech (e.g., I see) of the second section (e.g., I understand I see).

[0071] The memory (330) may store commands or data related to at least one other component of the electronic device (300). According to one embodiment, the memory (330) may store at least a portion of the spoken voice generated during a group call.

[0072] According to various embodiments, the memory (2430) may include at least one program module. The program module may include the program (140) of FIG. 1. The at least one program module may include a service providing module (2432), an extraction module (2434), and a generation module (2436). However, this is merely illustrative and the present document is not limited thereto. For example, at least one of the aforementioned modules may be excluded from the configuration of the memory (2430), and conversely, other modules other than the aforementioned modules may be added to the configuration of the memory (2430). Additionally, some of the aforementioned modules may be integrated into other modules.

[0073] According to one embodiment, the service providing module (2432) may include a command to provide a group call service that enables a plurality of electronic devices (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)) to make a call simultaneously, and to detect simultaneous utterances in which at least two speakers speak substantially simultaneously while the group call is in progress. According to one embodiment, the extraction module (2434) may include a command to obtain overlapping utterances from the spoken voice. According to one embodiment, the generation module (2436) may include a command to generate synthetic voice based on overlapping utterances. In this regard, the generation module (2436) may include a command to control the playback speed of the synthetic voice.

[0074] In the above-described embodiment, a configuration in which the synthesized voice is generated by an external device (240) has been described, but this document is not limited thereto. For example, the synthesized voice may be generated by at least one electronic device (e.g., a first electronic device (210), a second electronic device (220), and a third electronic device (230)), as described later through FIGS. 9a and 9b.

[0076] An electronic device (e.g., electronic device (240)) according to various embodiments includes a communication module (e.g., communication module (2410)) and a processor (e.g., processor (2420)) operatively connected to the communication module, and the processor receives and stores a first spoken voice associated with at least a first external device and a second spoken voice associated with a second external device, and when detecting a solitary speech based on the first spoken voice and the second spoken voice, transmits the first spoken voice or the second spoken voice having a first playback speed to the at least first electronic device and the second electronic device, and when detecting a simultaneous speech based on the first spoken voice and the second spoken voice, converts at least a portion of a synthesized speech in which at least a first superimposed speech of the first spoken voice and at least a second superimposed speech of the second spoken voice are connected in succession to a second playback speed different from the first playback speed, and transmits to the at least first electronic device and the second It can be configured to be transmitted to an electronic device.

[0077] According to various embodiments, the first speed is substantially the same as the speaker's speech speed, and the second speed may include a speed faster than the first speed.

[0078] According to various embodiments, the processor may be configured to identify a first utterance time associated with the first overlapping utterance and a second utterance time associated with the second overlapping utterance, and to determine the second playback speed such that the synthesized voice is played within a time shorter than the sum of the first utterance time and the second utterance time.

[0079] According to various embodiments, the processor may be configured to convert at least one of the first overlapping utterance and the second overlapping utterance into the second playback speed.

[0080] According to various embodiments, the processor may be configured to convert at least one of the first overlapped utterance and the second overlapped utterance into the second playback speed.

[0081] According to various embodiments, the processor may be configured to generate the synthesized speech by adding a silent interval between the first superimposed utterance and the second superimposed utterance.

[0082] According to various embodiments, the processor may be configured to acquire a portion corresponding to a certain range based on the first superimposed utterance among the first spoken voices as a first additional utterance, acquire a portion corresponding to a certain range based on the second superimposed utterance among the second spoken voices as a second additional utterance, and use the first additional utterance and the second additional utterance for the generation of the synthesized voice.

[0083] According to various embodiments, the processor may be configured to receive information related to the second playback speed from the first external device or the second external device, and to convert the synthesized speech based on the received information.

[0084] According to various embodiments, the processor may be configured to convert the synthesized speech such that a certain level of pitch is maintained for the first superimposed utterance and the second superimposed utterance.

[0085] An electronic device (e.g., electronic device (240)) according to various embodiments comprises a communication module (e.g., communication module (2410)), a microphone (e.g., input module (150)), an output module (e.g., sound output module (155)), and a processor operatively connected to the communication module, the microphone, and the output module, wherein the processor transmits a spoken voice obtained through the microphone to at least a first communication device and a second communication device, receives a first spoken voice obtained by the first communication device and a second spoken voice obtained by the second communication device, detects a single spoken voice or a simultaneous spoken voice based on the received first spoken voice and the second spoken voice, outputs the first spoken voice or the second spoken voice having a first playback speed when the single spoken voice is detected, and when the simultaneous spoken voice is detected, at least a first superimposed utterance of the first spoken voice and at least a second superimposed utterance of the second spoken voice It can be set to generate a series of connected synthetic voices and output them at a second playback speed different from the first playback speed.

[0086] According to various embodiments, the first speed is substantially the same as the speaker's speech speed, and the second speed may include a speed faster than the first speed.

[0088] FIG. 4 is a flowchart illustrating the operation of providing a group call service in an electronic device according to various embodiments. In the following description, the electronic device may be the external device described above through FIG. 2a, and the first call device and the second call device may be at least one electronic device described above through FIG. 2a.

[0089] Referring to FIG. 4, an electronic device (240) (or processor (2420)) according to various embodiments may receive spoken voices from at least a first call device and a second call device participating in a group call in operation 410. According to one embodiment, the electronic device (240) may receive a first spoken voice through a first channel assigned to the first call device and receive a second spoken voice through a second channel assigned to the second call device. However, this is merely illustrative and is not limited thereto. For example, if n call devices are participating in a group call, the electronic device (240) may receive n spoken voices.

[0090] According to various embodiments, the electronic device (240) can determine whether simultaneous speech is detected based on the first spoken voice and the second spoken voice in operation 420. Simultaneous speech may be a situation in which at least the speaker of the first call device and the speaker of the second call device speak substantially simultaneously. According to one embodiment, the electronic device (240) can detect simultaneous speech based on the time when spoken voices are received through the first channel and the second channel.

[0091] According to various embodiments, when simultaneous speech is not detected, that is, when a single speech occurs by the first or second call device, the electronic device (240) may, in operation 460, transmit a speech voice of a first speech rate to the call device. The first speech rate may be substantially the same as the speech rate of the speaker. For example, the electronic device (240) may transmit a first speech voice corresponding to the speech rate of a first speaker using the first call device to the second call device. Additionally, the electronic device (240) may transmit a second speech voice corresponding to the speech rate of a second speaker using the second call device to the first call device.

[0092] According to various embodiments, when simultaneous utterances are detected, the electronic device (240) can obtain a superimposed utterance from the received utterance voice in operation 430. According to one embodiment, the electronic device (240) can obtain a first superimposed utterance that overlaps with a second utterance voice in a first utterance voice. Additionally, the electronic device (240) can obtain a second superimposed utterance that overlaps with the first utterance voice in a second utterance voice. For example, the first superimposed utterance may include at least a portion belonging to the first utterance voice among the superimposed portions where the first utterance voice and the second utterance voice overlap. Additionally, the second superimposed utterance may include at least a portion belonging to the second utterance voice among the superimposed portions where the first utterance voice and the second utterance voice overlap.

[0093] According to various embodiments, the electronic device (240) can generate a synthetic voice in which a first superimposed utterance and a second superimposed utterance are connected in operation 440. According to one embodiment, the electronic device (240) can generate a synthetic voice by connecting superimposed utterances extracted from a speech voice. For example, the electronic device (240) can generate a synthetic voice by connecting a second superimposed utterance after a first superimposed utterance. As another example, the electronic device (240) can generate a synthetic voice by connecting a first superimposed utterance after a second superimposed utterance. As yet another example, the electronic device (240) can generate a synthetic voice in which a silent interval of a certain length is formed between a first superimposed utterance and a second superimposed utterance. According to another embodiment, the electronic device (240) can generate a synthetic voice by connecting a superimposed utterance extracted from a speech voice to the speech voice. For example, the electronic device (240) can generate a synthesized voice by connecting a second superimposed utterance after a first utterance. As another example, the electronic device (240) can generate a synthesized voice by connecting a second utterance after a first superimposed utterance.

[0094] According to various embodiments, the electronic device (240) can transmit synthesized speech at a second speech rate in operation 450. The second speech rate may be faster than the speaker's speech rate. According to one embodiment, the electronic device (240) can process at least one of the first overlapping speech and the second overlapping speech included in the synthesized speech to be played back at a second speech rate faster than the first speech rate. Additionally, the electronic device (240) can adjust the speech rate for at least a portion of the first overlapping speech and at least a portion of the second overlapping speech.

[0096] FIG. 5 is a flowchart illustrating the operation of acquiring superimposed ignition in an electronic device according to various embodiments. The operations of FIG. 5 described below may represent various embodiments of at least one of operations 410 to 430 of FIG. 4.

[0097] Referring to FIG. 5, an electronic device (240) (or processor (2420)) according to various embodiments may, in operation 510, store speech voice received from at least a first call device and a second call device based on a slot (or window) of a first size. The slot may be a range in which overlapping speech can be extracted from speech voice. According to one embodiment, the electronic device (240) may store a first speech voice received from a first call device based on a first slot of a first size, and store a second speech voice received from a second call device based on a second slot of a first size. For example, the slot of a first size may be a minimum range in which overlapping speech can be extracted, and as the size of the slot increases, the range in which overlapping speech can be extracted from speech voice may increase.

[0098] According to various embodiments, the electronic device (240) can identify a silent period in the spoken voice while simultaneous speech is detected in the 520 operation. The silent period may be a period during which the speaker's speech is interrupted for a specified time (e.g., 3 seconds). For example, the electronic device (240) can identify a silent period per channel by checking the point in time when the spoken voice is not received for a specified time after being received through each channel.

[0099] According to various embodiments, the electronic device (240) can adjust the size of the slot to a second size larger than the first size based on a silent period in operation 530. The second size may correspond to the period from when speech begins to when a silent period occurs. According to one embodiment, while simultaneous speech is detected, the electronic device (240) can expand the size of the first slot to the second size based on a silent period of the first speech sound, and expand the size of the second slot to the second size based on a silent period of the second speech sound.

[0100] According to various embodiments, the electronic device (240) can obtain a superimposed utterance based on a second-sized slot in operation 540. According to one embodiment, the electronic device (240) can obtain a speech voice corresponding to a second-sized slot at the time when simultaneous speech is interrupted as a superimposed utterance. For example, the electronic device (240) can obtain a first superimposed utterance corresponding to a first-sized slot of the second size in a first speech voice, and obtain a second superimposed utterance corresponding to a second-sized slot of the second size in a second speech voice.

[0102] FIG. 6 is a flowchart illustrating different operations for obtaining superimposed ignition in an electronic device according to various embodiments. The operations of FIG. 6 described below may represent various embodiments of at least one of operations 410 to 430 of FIG. 4.

[0103] Referring to FIG. 6, an electronic device (240) (or processor (2420)) according to various embodiments may, in operation 610, store speech voice received from at least a first call device and a second call device based on a slot (or window) of a first size. According to one embodiment, as described above through FIG. 5, the electronic device (240) may store a first speech voice received from a first call device based on a first slot of a first size and store a second speech voice received from a second call device based on a second slot of a first size.

[0104] According to various embodiments, the electronic device (240) can obtain voice information having a certain level of similarity by comparing a stored spoken voice with a voice information database while simultaneous speech is detected in operation 620. The voice information database may include at least one voice information defined by a combination of at least one word (e.g., a short answer word) or two or more words that may occur simultaneously depending on the characteristics of a group call (e.g., a meeting, a class, etc.). According to one embodiment, the electronic device (240) can obtain at least one voice information having a certain level of similarity with at least one voice information included in the voice information database from a first spoken voice and a second spoken voice.

[0105] According to various embodiments, the electronic device (240) can adjust the size of the slot to a second size corresponding to the acquired voice information in operation 630. According to one embodiment, the electronic device (240) can expand the size of the first slot to a corresponding second size based on voice information acquired from the first spoken voice, and expand the size of the second slot to a corresponding second size based on voice information acquired from the second spoken voice.

[0106] According to various embodiments, the electronic device (240) can obtain a superimposed utterance based on a slot of a second size in a 640 operation. According to one embodiment, the electronic device (240) can obtain a first superimposed utterance corresponding to a first slot of a second size in a first utterance voice, and obtain a second superimposed utterance corresponding to a second slot of a second size in a second utterance voice.

[0108] FIG. 7 is a flowchart illustrating the operation of determining the speech rate of a synthesized voice in an electronic device according to various embodiments. The operations of FIG. 7 described below may represent various embodiments of at least one of operations 440 to 450 of FIG. 4.

[0109] Referring to FIG. 7, an electronic device (240) (or processor (2420)) according to various embodiments can determine a first utterance time associated with a first overlapping utterance in operation 710. According to one embodiment, the electronic device (240) can determine a time interval defined as the start time and end time of the first overlapping utterance.

[0110] According to various embodiments, the electronic device (240) can determine the second utterance time associated with the second overlapping utterance in operation 720. According to one embodiment, the electronic device (240) can determine a time interval defined as the start time and end time of the second overlapping utterance.

[0111] According to various embodiments, the electronic device (240) may determine a second speech rate associated with the synthesized speech based on a second speech time and a second speech time in operation 730. According to one embodiment, the electronic device (240) may process the synthesized speech to be played back in a time shorter than the time (e.g., 2 seconds) of the sum of the first speech time (e.g., 2 seconds) and the second speech time (e.g., 1 second) (e.g., 3 seconds). For example, the electronic device (240) may process the first overlapping speech and at least some of the second overlapping speech to be played back faster than the first rate (normal rate or standard rate).

[0113] FIG. 8 is a diagram illustrating the operation of a group call system according to various embodiments.

[0114] Referring to FIG. 8, a group call system according to various embodiments may be composed of a plurality of electronic devices (e.g., a first electronic device (802), a second electronic device (806), and a third electronic device (808)) and an external device (804).

[0115] According to various embodiments, such as operations 810 through 814, each electronic device (e.g., first electronic device (802), second electronic device (806), and third electronic device (808)) can transmit a user's spoken voice received through a microphone to an external device (804). According to one embodiment, the first electronic device (802) can transmit a first spoken voice through a first channel, the second electronic device (806) can transmit a second spoken voice through a second channel, and the third electronic device (808) can transmit a third spoken voice through a third channel.

[0116] According to various embodiments, the external device (804) can detect simultaneous utterances, such as in the operation of 816. According to one embodiment, the external device (804) can detect that simultaneous utterances have occurred when utterance sounds are received through at least two channels at the same or close time.

[0117] According to various embodiments, such as in operation 818, the external device (804) may generate a synthesized voice based on the utterances of the simultaneous speakers in response to detecting the occurrence of simultaneous speech. According to one embodiment, the external device (802) may perform at least some of operations 430 to 450 of FIG. 4 described above in order to generate a synthesized voice.

[0118] According to various embodiments, such as in operation 820, the external device (804) may transmit the synthesized voice to at least one electronic device (e.g., a first electronic device (802), a second electronic device (806), and a third electronic device (808)). According to one embodiment, the external device (804) may transmit the synthesized voice only to at least one electronic device (e.g., a first electronic device (802)) where no simultaneous speech has occurred. As another example, the external device (804) may transmit the synthesized voice to all electronic devices participating in a group call.

[0119] According to various embodiments, such as in operation 822, at least one electronic device (e.g., first electronic device (802), second electronic device (806) and third electronic device (808)) can play synthesized voice received from an external device (802).

[0121] FIG. 9a is a diagram illustrating different operations of a group call system according to various embodiments. The group call system described below differs from the group call system described through FIG. 8 in that it detects the occurrence of simultaneous speech and generates synthesized speech from the perspective of an electronic device rather than an external device.

[0122] Referring to FIG. 9a, a group call system according to various embodiments may be composed of a plurality of electronic devices (e.g., a first electronic device (902), a second electronic device (906), and a third electronic device (908)) and an external device (904).

[0123] According to various embodiments, at least one electronic device (e.g., a first electronic device (902), a second electronic device (906), and a third electronic device (908)) may receive a spoken voice obtained by another electronic device. According to one embodiment, the first electronic device (902) may receive a user's spoken voice received through the microphones of the second electronic device (906) and the third electronic device (908). In this regard, such as in operations 910 and 912, the second electronic device (906) may transmit the user's spoken voice to an external device (904), and the external device (904) may transmit the received spoken voice to the first electronic device (902) through a first channel. Additionally, as with operations 914 and 916, the third electronic device (908) transmits the user's spoken voice to an external device (904), and the external device (904) can transmit the received spoken voice to the first electronic device (902) through a second channel. For example, the first channel may be a channel set in the first electronic device (902) to receive the spoken voice of the second electronic device (906), and the second channel may be a channel set in the first electronic device (902) to receive the spoken voice of the third electronic device (908).

[0124] According to various embodiments, such as in operation 918, at least one electronic device (e.g., first electronic device (902)) can detect simultaneous utterances based on received speech sounds. According to one embodiment, at least one electronic device (e.g., first electronic device (902)) can detect that simultaneous utterances have occurred when speech sounds are received through the first channel and the second channel at simultaneous or close together.

[0125] According to various embodiments, such as in operation 920, at least one electronic device (e.g., first electronic device (902)) may generate synthetic speech based on the speech of simultaneous speakers in response to detecting the occurrence of simultaneous speech. According to one embodiment, at least one electronic device (e.g., first electronic device (902)) may perform at least some of operations 430 to 450 of FIG. 4 described above in order to generate synthetic speech.

[0126] According to various embodiments, such as in operation 922, at least one electronic device (e.g., first electronic device (902)) can play the synthesized voice generated.

[0127] In the above-described embodiment, a group call system composed of a plurality of electronic devices (e.g., a first electronic device (902), a second electronic device (906), and a third electronic device (908)) and an external device (904) was described. However, this is merely illustrative and the present document is not limited thereto. For example, as illustrated in FIG. 9b, a group call system according to various embodiments may be composed only of a plurality of electronic devices (e.g., a first electronic device (902), a second electronic device (906), and a third electronic device (908)).

[0128] According to one embodiment, at least one electronic device (e.g., a first electronic device (902), a second electronic device (906), and a third electronic device (908)) can receive a spoken voice obtained by another electronic device. In this regard, such as in operations 911 and 913, the second electronic device (906) can transmit the user's spoken voice to the first electronic device (902) through a first channel, and the third electronic device (908) can transmit the user's spoken voice to the first electronic device (902) through a second channel.

[0129] According to one embodiment, as with operations 918 to 922 described above through FIG. 9a, at least one electronic device (e.g., first electronic device (902)) may detect simultaneous speech based on received speech sounds, generate synthetic speech based on the speech sounds of simultaneous speakers, and play the generated synthetic speech.

[0131] FIG. 10 is a diagram illustrating another operation of a group call system according to various embodiments. FIG. 11 is a diagram for explaining the operation of setting parameters of a synthesized voice according to various embodiments. The group call system described below is similar to the group call system described through FIG. 8 in that an external device detects the occurrence of simultaneous speech and generates synthesized voice, but differs in that parameters for the synthesized voice are set from the perspective of an electronic device.

[0132] Referring to FIG. 10, a group call system according to various embodiments may be composed of a plurality of electronic devices (e.g., a first electronic device (1002) and a second electronic device (1006)) and an external device (1004).

[0133] According to various embodiments, such as in operation 1010, at least one electronic device (e.g., first electronic device (1002) and second electronic device (1006)) may set parameters related to the synthesized speech. The parameters may include a method for separating (or extracting) overlapping utterances from the spoken speech, a method for reproducing overlapping utterances, and the number of simultaneous speakers allowed. According to one embodiment, the parameters may be set by an electronic device that has opened a group call service. However, this is merely illustrative and the present document is not limited thereto.

[0134] According to one embodiment, at least one electronic device (e.g., a first electronic device (1002) and a second electronic device (1006)) may output a user interface including at least one menu for setting parameters, as illustrated in FIG. 11 (a), before or during the execution of a group call service. For example, the user interface may include an object (1102) for setting a method for separating (or extracting) overlapping utterances from spoken voice, an object (1104) for setting a method for playing overlapping utterances, and an object (1106) for setting the number of allowed simultaneous speakers.

[0135] According to one embodiment, at least one electronic device (e.g., a first electronic device (1002) and a second electronic device (1006)) may select at least one of a method (1112) for obtaining a superimposed utterance in a spoken voice based on a silent interval described through FIG. 5 based on user input, as shown in FIG. 11 (b), or a method (1114) for obtaining a superimposed utterance in a spoken voice based on a database described through FIG. 6.

[0136] According to one embodiment, at least one electronic device (e.g., a first electronic device (1002) and a second electronic device (1006)) may select one of a method for playing overlapping utterances based on utterance quality (1122) or a method for playing overlapping utterances based on utterance speed (1124) based on user input, as illustrated in FIG. 11 (c). The method based on utterance quality may be a method in which the playback speed is controlled within a range where the pitch of the overlapping utterances is maintained at a certain level. Additionally, the method based on utterance speed may be a method in which the playback speed is controlled at a faster playback speed than the method based on utterance quality, while the pitch of the overlapping utterances is maintained below a certain level.

[0137] According to one embodiment, at least one electronic device (e.g., a first electronic device (1002) and a second electronic device (1006)) can set (1132) the number of allowed simultaneous speakers based on user input, as shown in (d) of FIG. 11. The set number of simultaneous speakers may be the maximum number of overlapping speeches that can be included in the synthesized speech.

[0138] According to various embodiments, such as in operation 1012, at least one electronic device (e.g., first electronic device (1002) and second electronic device (1006)) can transmit parameter setting information to an external device (1004).

[0139] According to various embodiments, each electronic device (e.g., first electronic device (1002) and second electronic device (1006)) can transmit a user's spoken voice received through a microphone to an external device (240). According to one embodiment, as in operation 1014-1 and operation 1016-1, the first electronic device (1002) can acquire the spoken voice and transmit it to an external device (1004). Additionally, as in operation 1014-2 and operation 1016-2, the second electronic device (1006) can also acquire the spoken voice and transmit it to an external device (1004).

[0140] According to various embodiments, such as operations 1018 and 1020, the external device (1004) can detect simultaneous speech and generate synthetic speech based on parameter setting information. For example, the external device (1004) can generate synthetic speech based on the method of acquiring overlapping speech, the method of reproducing overlapping speech, and the number of simultaneous speakers set by at least one electronic device (e.g., the first electronic device (1002) and the second electronic device (1006)). Regarding the method of acquiring overlapping speech, if both a separation method based on silent intervals and a separation method based on a database are selected, the external device (1004) can acquire overlapping speech using both methods simultaneously. In this case, if the overlapping speech is acquired by either the acquisition method based on silent intervals or the acquisition method based on a database (e.g., based on silent intervals), the external device (1004) may stop the operation based on the other one (e.g., based on a database).

[0141] According to various embodiments, as in operation 1022, the external device (1004) can transmit the synthesized voice to at least one electronic device (e.g., the first electronic device (1002) and the second electronic device (1006)). Accordingly, at least one electronic device (e.g., the first electronic device (1002) and the second electronic device (1006)) can play the received synthesized voice as in operation 1024.

[0143] A method of operation of an electronic device (e.g., electronic device (240)) according to various embodiments may include: receiving and storing a first spoken voice associated with at least a first external device and a second spoken voice associated with a second external device; detecting a single spoken voice or a simultaneous spoken voice based on the first spoken voice and the second spoken voice; when the single spoken voice is detected, transmitting the first spoken voice or the second spoken voice having a first playback speed to the at least first electronic device and the second electronic device; and when the simultaneous spoken voice is detected, converting at least a portion of a synthetic voice in which at least a first superimposed speech of the first spoken voice and at least a second superimposed speech of the second spoken voice are connected in succession to a second playback speed different from the first playback speed and transmitting it to the at least first electronic device and the second electronic device.

[0144] According to various embodiments, the first speed is substantially the same as the speaker's speech speed, and the second speed may include a speed faster than the first speed.

[0145] According to various embodiments, the method may include an operation to confirm a first utterance time associated with the first overlapping utterance and a second utterance time associated with the second overlapping utterance, and an operation to determine a second playback speed such that the synthesized voice is played within a time shorter than the sum of the first utterance time and the second utterance time.

[0146] According to various embodiments, the operation may include converting at least one of the first superimposed utterance and the second superimposed utterance into the second playback speed.

[0147] According to various embodiments, the operation may include converting at least one of the first superimposed utterance and the second superimposed utterance into the second playback speed.

[0148] According to various embodiments, the operation of generating the synthesized speech by adding a silent interval between the first superimposed utterance and the second superimposed utterance may be included.

[0149] According to various embodiments, the method may include the operation of obtaining a portion corresponding to a certain range based on the first superimposed utterance among the first spoken voices as a first additional utterance, the operation of obtaining a portion corresponding to a certain range based on the second superimposed utterance among the second spoken voices as a second additional utterance, and the operation of using the first additional utterance and the second additional utterance in the synthesis of the voice.

[0150] According to various embodiments, the method may include receiving information related to the second playback speed from the first external device or the second external device, and converting the synthesized speech based on the received information.

[0151] According to various embodiments, the operation of converting the synthesized speech so that a certain level of pitch is maintained for the first superimposed utterance and the second superimposed utterance may be included.

Claims

Claim 1 In an electronic device, a communication module; and includes a processor operatively connected to the communication module, wherein the processor receives a first spoken voice associated with at least a first external device and a second spoken voice associated with a second external device, transmits at least one of the first spoken voice and the second spoken voice having a first playback speed to at least one of the first external device and the second external device, and when simultaneous speech is identified based on a first speech time defined by the start and end times of the first spoken voice and a second speech time defined by the start and end times of the second spoken voice: a portion of the first spoken voice corresponding to the second speech time is acquired as a first superimposed speech, a portion of the second spoken voice corresponding to the first speech time is acquired as a second superimposed speech, and at least a portion of a synthesized speech in which the first superimposed speech and the second superimposed speech are consecutively connected is converted to a second playback speed different from the first playback speed and the first external device and the second external device An electronic device configured to transmit to at least one of the following. Claim 2 An electronic device according to claim 1, wherein the first playback speed is substantially the same speed as the speaker's speech speed, and the second playback speed includes a speed faster than the first playback speed. Claim 3 An electronic device configured such that, in claim 1, the processor identifies a first utterance time associated with the first overlapping utterance and a second utterance time associated with the second overlapping utterance, and determines the second playback speed such that the synthesized voice is played within a time shorter than the sum of the first utterance time and the second utterance time. Claim 4 In claim 1, the processor is an electronic device configured to convert at least one of the first superimposed utterance and the second superimposed utterance into the second playback speed. Claim 5 In claim 1, the processor is an electronic device configured to convert at least one of the first overlapping utterance and the second overlapping utterance into the second playback speed. Claim 6 In claim 1, the processor is an electronic device configured to generate the synthesized speech by adding a silent interval between the first superimposed utterance and the second superimposed utterance. Claim 7 An electronic device according to claim 1, wherein the processor acquires a portion corresponding to a certain range based on the first superimposed utterance among the first spoken voices as a first additional utterance, acquires a portion corresponding to a certain range based on the second superimposed utterance among the second spoken voices as a second additional utterance, and is configured to use the first additional utterance and the second additional utterance for the generation of the synthesized voice. Claim 8 In claim 1, the processor is an electronic device configured to receive information related to the second playback speed from the first external device or the second external device and to convert the synthesized speech based on the received information. Claim 9 In claim 1, the processor is an electronic device configured to convert the synthesized speech such that a certain level of pitch is maintained for the first superimposed utterance and the second superimposed utterance. Claim 10 A method of operating an electronic device comprises: receiving a first spoken voice associated with at least a first external device and a second spoken voice associated with a second external device; and transmitting at least one of the first spoken voice and the second spoken voice, having a first playback speed, to at least one of the first external device and the second external device. and, when simultaneous utterances are identified based on a first utterance time defined by the start and end times of the first utterance voice and a second utterance time defined by the start and end times of the second utterance voice: an operation of acquiring a portion of the first utterance voice corresponding to the second utterance time as a first superimposed utterance and acquiring a portion of the second utterance voice corresponding to the first utterance time as a second superimposed utterance; and an operation of converting at least a portion of the synthesized voice in which the first superimposed utterance and the second superimposed utterance are consecutively connected to a second playback speed different from the first playback speed and transmitting it to at least one of the first external device and the second external device. Claim 11 A method according to claim 10, wherein the first playback speed is substantially the same speed as the speaker's speech speed, and the second playback speed includes a speed faster than the first playback speed. Claim 12 A method according to claim 10, comprising: an operation of confirming a first utterance time associated with the first overlapping utterance and a second utterance time associated with the second overlapping utterance; and an operation of determining a second playback speed such that the synthesized voice is played within a time shorter than the sum of the first utterance time and the second utterance time. Claim 13 A method according to claim 10, comprising the operation of converting at least one of the first superimposed utterance and the second superimposed utterance into the second playback speed. Claim 14 A method according to claim 10, comprising the operation of converting at least one of the part of the first superimposed utterance and the part of the second superimposed utterance into the second playback speed. Claim 15 A method according to claim 10, comprising the operation of generating the synthesized speech by adding a silent interval between the first superimposed utterance and the second superimposed utterance. Claim 16 A method according to claim 10, comprising: an operation of obtaining a portion corresponding to a certain range based on the first superimposed utterance among the first spoken voices as a first additional utterance; an operation of obtaining a portion corresponding to a certain range based on the second superimposed utterance among the second spoken voices as a second additional utterance; and an operation of using the first additional utterance and the second additional utterance in the generation of the synthesized voice. Claim 17 A method according to claim 10, comprising: receiving information related to the second playback speed from the first external device or the second external device; and converting the synthesized speech based on the received information. Claim 18 A method according to claim 10, comprising the operation of converting the synthesized speech so that a certain level of pitch is maintained for the first superimposed utterance and the second superimposed utterance. Claim 19 In an electronic device, a communication module; a microphone; an output module; and includes a processor operatively connected to the communication module, the microphone, and the output module, wherein the processor transmits a spoken voice obtained through the microphone to at least a first counterparty communication device and a second counterparty communication device, receives a first spoken voice obtained by the first counterparty communication device and a second spoken voice obtained by the second counterparty communication device, outputs the first spoken voice and the second spoken voice having a first playback speed, and when simultaneous speech is identified based on a first speech time defined by the start and end times of the first spoken voice and a second speech time defined by the start and end times of the second spoken voice: a portion of the first spoken voice corresponding to the second speech time is obtained as a first superimposed speech and a portion of the second spoken voice corresponding to the first speech time is obtained as a second superimposed speech, and generates a synthesized voice in which the first superimposed speech and the second superimposed speech are connected in succession, and has a first playback speed different from An electronic device set to output at a second playback speed. Claim 20 An electronic device according to claim 19, wherein the first playback speed is substantially the same speed as the speaker's speech speed, and the second playback speed includes a speed faster than the first playback speed.

Citation Information

Patent Citations

  • Communication system and communication terminal

    JP2009033298A

  • Method and apparatus for recognizing a voice

    KR1020190065201A

  • Method for operating speech recognition service and electronic device supporting the same

    KR1020190097483A

  • Method and device for providing information

    KR1020190106887A

  • Speech processing method and apparatus therefor

    KR1020190118996A