Electronic device and control method thereof

The electronic device addresses the issue of improper image display and speaker identification in multi-user video calls by adjusting image transmission and display based on device characteristics and speech analysis, ensuring clear participant visibility and speaker recognition.

WO2025170186A1PCT designated stage Publication Date: 2025-08-14SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/021347
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2024-12-27
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Video calls involving multiple users with varying hardware specifications can result in improper display of participant images, making it difficult to identify the speaker, especially when the number of users exceeds the hardware capabilities of the terminal devices.

Method used

An electronic device that transmits characteristic information of each participant's device to a server, receives images based on speech situations, and adjusts the display output to ensure proper image transmission and identification of the primary speaker.

Benefits of technology

The solution ensures that all participants' images are displayed correctly, even when hardware limitations are exceeded, by selectively transmitting and replacing images with graphic objects, and accurately identifying the primary speaker based on speech analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024021347_14082025_PF_FP_ABST
    Figure KR2024021347_14082025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device is disclosed. When a video call is performed with a plurality of users, a processor transmits characteristic information of the electronic device to a server through a communication unit, receives, from the server through the communication unit, the number of images corresponding to respective external device characteristic information from among images of the plurality of users performing the video call, and controls a display to change and output the received images on the basis of an utterance situation of each of the plurality of users.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method of controlling the same

[0001] The invention relates to an electronic device that provides a video call involving multiple users and a method for controlling the same.

[0002] As demand for non-face-to-face video conferencing services increases during the COVID-19 pandemic, related technologies and devices for video calls are rapidly advancing.

[0003] In particular, demand for video calls and video chats via large-screen displays such as TVs is increasing.

[0004] To make a video call, multiple users must use their own terminal devices. In this case, hardware specifications may vary depending on the terminal type. If a terminal with low hardware specifications is used, the video of some participants in the video call may not be displayed properly.

[0005] In this case, it is pointed out as a limitation of video calls that it is not possible to properly confirm who is speaking.

[0006] An electronic device according to at least one embodiment of the present disclosure includes a communication unit, a display, a memory, and a processor.

[0007] The processor, when performing a video call with multiple users, transmits characteristic information of the electronic device to a server through the communication unit, receives a number of images corresponding to the characteristic information of each external device among the images of the multiple users performing the video call from the server through the communication unit, and controls the display so that the received images are changed and output based on the speech situations of each of the multiple users.

[0008] A method for controlling an electronic device according to at least one embodiment of the present disclosure includes the steps of transmitting characteristic information of the electronic device to a server when performing a video call with a plurality of users, receiving a number of images corresponding to the characteristic information of each external device among images of the plurality of users performing the video call from the server, and changing the received images and outputting them through a display based on the speech situations of each of the plurality of users.

[0009] According to at least one embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, the operation includes: when performing a video call with a plurality of users, transmitting characteristic information of the electronic device to a server; receiving, from the server, a number of images corresponding to the characteristic information of each of the external devices among images of the plurality of users performing the video call; and changing the received images and outputting them through a display based on the speech situations of each of the plurality of users.

[0010] FIG. 1 is a perspective view illustrating the operation of an electronic device according to at least one embodiment of the present disclosure.

[0011] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to at least one embodiment of the present disclosure.

[0012] FIG. 3 is a diagram for explaining an image display process based on characteristic information of an electronic device according to at least one embodiment of the present disclosure.

[0013] FIG. 4 is a diagram illustrating a process of transmitting a changed image of an electronic device according to at least one embodiment of the present disclosure.

[0014] FIG. 5 is a diagram illustrating a method for selecting a primary speaker of an electronic device according to at least one embodiment of the present disclosure.

[0015] FIG. 6 is a diagram illustrating a speaker prediction model of an electronic device according to at least one embodiment of the present disclosure.

[0016] FIG. 7 is a diagram illustrating a speaker estimation process based on identification information of an electronic device according to at least one embodiment of the present disclosure.

[0017] FIG. 8 is a drawing illustrating an electronic device including a display according to at least one embodiment of the present disclosure.

[0018] FIG. 9 is a flowchart illustrating a method for controlling an electronic device according to at least one embodiment of the present disclosure.

[0019] FIG. 10 is a diagram illustrating a process for identifying a primary speaker based on priority information of an electronic device according to at least one embodiment of the present disclosure.

[0020] FIG. 11 is a diagram illustrating a process for identifying a primary speaker based on a trigger word of an electronic device according to at least one embodiment of the present disclosure.

[0021] FIG. 12 is a flowchart illustrating the overall operation process of an electronic device according to at least one embodiment of the present disclosure.

[0022] The terms used in the various embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should be defined based on the meaning of the terms and the overall content of this disclosure, rather than simply their names.

[0023] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.

[0024] The expression "at least one of A and / or B" should be understood to mean either "A" or "B" or "A and B".

[0025] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.

[0026] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).

[0027] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this disclosure, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0028] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor (not shown), excluding any "modules" or "parts" that need to be implemented as specific hardware.

[0029] In this disclosure, the term user may refer to a person using an electronic device or a device used by the person.

[0030] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.

[0031] FIG. 1 is a perspective view illustrating the operation of an electronic device according to at least one embodiment of the present disclosure.

[0032] An electronic device (100) can provide video calls between a plurality of external devices (200 to 203). Specifically, the electronic device (100) can be implemented as various devices such as a server device, a PC, a laptop PC, a mobile phone, a tablet PC, a cloud server, etc. The external devices (200 to 203) can be electronic devices used by a user to make video calls. Each of the external devices (200 to 203) can be implemented as various devices such as a TV, a PC, a laptop PC, a smart monitor, a mobile phone, a tablet PC, an electronic whiteboard, a kiosk, a home appliance, etc. Each of the external devices (200 to 203) can be implemented as an electronic device that directly includes a display or is connected to an external display (e.g., a monitor).

[0033] In FIG. 1, the external device (200 to 203) is referred to as an external device in the sense that it is a separate device from the electronic device (100), but it may also be referred to as an electronic device, a user terminal device, or, if it directly includes a display, a display device.

[0034] A video call can be a service that supports online conversations between multiple users. When a video call involving multiple users is initiated, the electronic device (100) can transmit and receive video, audio, and other data between the multiple external devices (200-203) used by each user. While referred to as a video call in this disclosure, it can also be referred to in various other forms, such as a video conferencing service or an online conversation service. Furthermore, the electronic device (100) will be described herein as a server device.

[0035] Referring to FIG. 1, multiple users can conduct a video conference using their own electronic devices. When multiple users are conducting a video conference, each external device (200-203) used by the multiple users can receive other users' videos or transmit its own videos via the electronic device (100).

[0036] Video calls can be implemented in a variety of ways. For example, the electronic device (100) can transmit video from external devices (200-203) used by each user to external devices of other users. In this case, each external device (200-203) can directly configure a video conference screen by combining the user video it has with other received video.

[0037] In this case, depending on the characteristics of the external devices (200 to 203), there may be limitations on the number of images that can be received or displayed from the electronic device (100). That is, the external devices (200 to 203) require hardware components such as a decoder, a scaler, etc. to process and display the received images. In the present disclosure, the characteristics of the external devices (200 to 203) may include the type or number of hardware components required for image processing. In addition, the characteristic information may include information indicating the type or number of such hardware components.

[0038] When the number of users participating in a video conference does not exceed the number of hardware configurations included in the external devices (200-203), the external devices (200-203) can display the videos of all users. However, when the number of users participating in a video conference exceeds the number of hardware configurations included in the external devices (200-203), the external device (200) may experience problems displaying the videos.

[0039] In the case of the latest TVs, multiple hardware components may be implemented to display multiple contents on a single screen. However, even in this case, there may be cases where the number of users exceeds the number of hardware components. Furthermore, at least one of the external devices (200-203) may have only one decoder and one scaler. In such cases, at least one of the external devices may not be able to display multiple received images simultaneously on a single video conference screen, or may only process and display them in low-quality format. Consequently, it may be difficult to conduct a video conference normally.

[0040] The electronic device (100) can determine the number of images to be transmitted to each of the external devices (200 to 203) based on the characteristic information of each of the external devices (200 to 203) and transmit that number of images. For example, based on the characteristic information of external device 1 (200) and external device 2 (201), the electronic device (100) can determine that the number of images that can be simultaneously played in external device 1 (200) is 2 and that the number of images that can be simultaneously played in external device 2 (201) is 3.

[0041] The electronic device (100) can select two images from the images received from all external devices (200 to 203), excluding the image of external device 1 (200), and transmit them to external device 1 (200). The electronic device (100) can transmit the three images selected in the same manner to external device 2 (201).

[0042] Meanwhile, during a video conference, the speaking situations of multiple users can change at any time. These situations can include which user is the primary speaker, which user is currently speaking, which user speaks most frequently or for the most time, who will be the next speaker in the conference, which user has been designated as the next speaker, and which user is predicted to be the next speaker.

[0043] The electronic device (100) can change the image transmitted to each external device (200-203) based on this firing situation.

[0044] Accordingly, in the present disclosure, the electronic device (100) can overcome limitations due to differences in specifications of each external device (200) by considering the characteristic information of the external device (200) and provide a better video call. Hereinafter, various embodiments of the present disclosure will be described in detail.

[0045] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to at least one embodiment of the present disclosure.

[0046] According to FIG. 2, the electronic device (100) includes a communication unit (110), a memory (120), a processor (130), and a display (140). However, the present invention is not limited thereto, and the electronic device (100) may be implemented in a form in which some components are excluded, or may be implemented in a form in which other components are further included.

[0047] The communication unit (110) is a configuration for performing communication with at least one external device (200). The communication unit (110) may include at least one wireless communication module, at least one wired communication module, etc. Each communication module may be implemented in the form of at least one hardware chip. For example, the wireless communication module may include at least one module among a Wi-Fi module, a Bluetooth module, an infrared communication module, or other communication modules. In addition, the communication unit (110) may include at least one communication chip for performing communication according to various wireless communication standards such as Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc.

[0048] The wired communication module may include, for example, at least one of a Local Area Network (LAN) module, an Ethernet module, a paired cable, a coaxial cable, a fiber optic cable, or an Ultra Wide-Band (UWB) module.

[0049] The memory (120) can store at least one command, data, program, etc. required for the operation of the electronic device (100). For example, the memory (120) can store characteristic information of an external device (200), identification information of each of a plurality of users conducting a video conference, characteristic information of an external device used by each user, etc.

[0050] The memory (120) may be implemented in the form of memory embedded in the electronic device (100) or in the form of memory detachable from the electronic device (100) depending on the purpose of data storage. For example, data for driving the electronic device (100) may be stored in a memory embedded in the electronic device (100), and data for expanding the functions of the electronic device (100) may be stored in a memory detachable from the electronic device (100).

[0051] In the case of memory embedded in the electronic device (100), it may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).

[0052] The memory (120) may be implemented as a single memory that stores data generated in various operations according to the present disclosure, but is not limited thereto, and the memory (120) may be implemented to include multiple memories that each store different types of data or each store data generated in different stages.

[0053] The display (140) is a configuration for displaying various screens. The display (140) can be implemented as various types of displays such as an LCD (liquid crystal display), an OLED (organic light-emitting diode), an LCoS (Liquid Crystal on Silicon), a DLP (Digital Light Processing), a QD (quantum dot) display panel, a QLED (quantum dot light-emitting diodes), a μLED (Micro light-emitting diodes), a Mini LED, etc. Meanwhile, the display (140) can also be implemented as a touch screen combined with a touch sensor, a flexible display, a rollable display, a 3D display, a display in which a plurality of display modules are physically connected, etc.

[0054] The processor (130) is a component for controlling the operation of the electronic device (100). The processor (130) may be implemented as a digital signal processor (DSP) for processing digital signals, a microprocessor, but is not limited thereto, and may include one or more of a central processing unit (CPU), a microcontroller unit (MCU), a microprocessing unit (MPU), a controller, an application processor (AP), a communication processor (CP), an ARM processor, and an artificial intelligence (AI) processor, or may be defined by the relevant terminology. In addition, the processor (130) may be implemented as a system on chip (SoC) having a built-in processing algorithm, a large scale integration (LSI), or may be implemented in the form of a field programmable gate array (FPGA). The processor (130) may perform various functions by executing computer executable instructions stored in the memory (120).

[0055] The processor (130) may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores). When one or more processors (130) are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. In addition, each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute a program instruction for implementing a method according to an embodiment of the present disclosure, or all (or some) of the multiple cores may be linked to read and execute a program instruction for implementing a method according to an embodiment of the present disclosure.

[0056] When performing a video call with multiple users, the processor (130) can transmit characteristic information of the electronic device (100) to the server through the communication unit (110). The processor (130) can receive characteristic information of multiple external devices used by multiple users from the server through the communication unit (110) and store the information in the memory (120). The processor (130) can receive a number of images corresponding to the characteristic information of each external device among the images of multiple users performing the video call from the server through the communication unit (110). The processor (130) controls the display (140) so that the received images are changed and output based on the speech situations of each of the multiple users.

[0057] When a user input requesting a video call from at least one user is received through the communication unit (110), the processor (130) may read and execute an application corresponding to the video call from the memory (120) according to the user input. Upon execution of the application, the processor (130) may transmit a message requesting various information necessary for the video call to an external device to which the requester is logged on through the communication unit (110). For example, the processor (130) may transmit a message requesting various information such as the phone number, name, email address, messenger address, ID, etc. of other users who will participate in the video call, as well as the video conference execution time, video conference name, etc. When a predetermined time has arrived, the processor (130) may execute the video call. In the present disclosure, execution of a video call is not limited to a case where a conference is in progress after multiple users have actually completed participating, but may also include a case where an application is executed upon a video call request and various operations are performed according to the control of the application.

[0058] When a video call involving multiple users is executed, the processor (130) can receive characteristic information of multiple external devices used by multiple users through the communication unit (110) and store the information in the memory (120).

[0059] The processor (130) can transmit a number of images corresponding to the characteristic information of each external device (200 to 203) among the images for multiple users to the multiple external devices (200) through the communication unit (110). The processor (130) can change the images to be transmitted to the multiple external devices (200 to 203) based on the speech situations of each of the multiple users.

[0060] The characteristic information of the external device (200 to 203) may also be referred to as capability information, resource information, etc. of the electronic device. The characteristic information may specifically include decoder information, scaler information, memory capacity, graphics card performance, battery life, security features, etc. of the electronic device.

[0061] The processor (130) may receive characteristic information of an external device (200 to 203) at the time a video call is executed. Alternatively, the processor (130) may receive characteristic information of an external device (200 to 203) according to a user input before or after the video call is executed.

[0062] The processor (130) can receive multiple user images from each of the external devices (200 to 203). Thereafter, the processor (130) can transmit user images that match the number of images that can be displayed according to the specifications of each external device (200 to 203) based on the characteristic information of the external devices (200 to 203) received from the external devices (200 to 203). The specific details thereof will be described later in FIG. 3.

[0063] In order to transmit a number of images corresponding to the characteristic information of the external device (200 to 203) among the received plurality of user images to the external device, the processor (130) must select the images to be transmitted from among the plurality of images. Thereafter, the processor (130) can transmit the selected images to the external device (200 to 203).

[0064] The processor (130) can transmit an image of a user speaking to an external device (200 to 203). In addition, the processor (130) can select a primary speaker based on specific criteria and transmit the selected speaker to the external device (200 to 203). In addition, the processor (130) can estimate a speaker based on context information obtained from the spoken voices of multiple users and transmit an image of the estimated speaker to the external device (200 to 203).

[0065] In this way, the processor (130) can select and transmit an image to be transmitted to an external device (200 to 203) from among a plurality of images based on the speech situations of each of the plurality of users. When the speech situation changes, the processor (130) can transmit the changed image based on the changed situation to the external device (200 to 203).

[0066] When performing a video call with multiple users, the processor (130) can transmit characteristic information of the electronic device (100) to the server via the communication unit (110). The processor (130) can receive, from the server via the communication unit (110), a number of images corresponding to the characteristic information of each external device among the images of the multiple users performing the video call. The processor (130) can control the display (140) to change and output the received images based on the speech situations of each of the multiple users.

[0067] The processor (130) according to the present disclosure can selectively receive from the server, from among the images for a plurality of users, a number of images corresponding to the number of decoders or scalers included in the characteristic information of a plurality of external devices. The processor (130) can control the display (140) so that, from among the images for a plurality of users, the remaining images other than the images selectively received are displayed by replacing them with graphic objects.

[0068] The processor (130) according to the present disclosure can receive audio data of at least one of a plurality of users through the communication unit (110) when at least one of the users speaks. The processor (130) can receive images of at least one user corresponding to the audio data from a server and output a number of images corresponding to the characteristic information of the electronic device (100) through the display (140). When audio data of another user among the plurality of users is received through the communication unit (110), the processor (130) can replace the image being output through the display (140) with an image of the user corresponding to the audio data of the other user.

[0069] The processor (130) according to the present disclosure can identify a primary speaker among multiple users based on audio data received through the communication unit (110). The processor (130) can output an image of the identified primary speaker through the display (140). If the primary speaker of the audio data received through the communication unit (110) changes, the processor (130) can replace the image being output through the display (140) with an image of the primary speaker changed based on the received audio data.

[0070] The processor (130) according to the present disclosure can identify a primary speaker based on at least one of the cumulative speech time, speech frequency, and speech topic of each of a plurality of users.

[0071] The processor (130) according to the present disclosure can identify the next speaker corresponding to the context information among the plurality of users based on the context information obtained from the speech of each of the plurality of users. The processor (130) can transmit an image of the identified next speaker to the server via the communication unit (110).

[0072] The processor (130) according to the present disclosure can identify the other user as the primary speaker if the speech of at least one of the multiple users includes identification information of another user. The processor (130) can transmit an image of the identified primary speaker to the server via the communication unit (110).

[0073] The processor (130) according to the present disclosure can control the display (140) to output an image of a primary speaker identified based on priority information of multiple users.

[0074] The processor (130) according to the present disclosure can control the display (140) to output an image of a user corresponding to the trigger text as an image of the next speaker when the context information obtained from the speech of each of the plurality of users includes a trigger text corresponding to the identification of each of the plurality of users.

[0075] The processor (130) according to the present disclosure can control the display (140) to output an image of the next speaker identified through the speaker prediction model when input data including the speech voices of each of a plurality of users is received through the communication unit (110).

[0076] Below, the process of selecting an image to be transmitted to an external device (200 to 203) by the processor (130) will be described in detail.

[0077] FIG. 3 is a drawing for explaining the operation of an electronic device according to at least one embodiment of the present disclosure.

[0078] The processor (130) can receive images (31 to 34) of users from a plurality of external devices (200 to 203). The processor (130) can selectively transmit, from among the received images (31 to 34), a number of images corresponding to characteristic information of each of the external devices (200 to 203) to each of the plurality of external devices (200 to 203).

[0079] For example, the number of images can be selected and transmitted as many as the number of decoders or scalers. A decoder is a hardware configuration that can decode a digital data stream to restore data. The processor (130) can transmit a number of images, in particular, the number of video decoders for processing a video signal, among various decoders such as a video decoder, an audio decoder, and a data decoder. A scaler is a configuration for performing a scaling operation that converts the resolution of a video signal to a different resolution. The video decoder and the scaler may be implemented as a single hardware chip or as separate hardware chips.

[0080] The processor (130) can transmit the same number of images as the number of scalers. The processor (130) can select and transmit images from among the images transmitted from the remaining devices excluding the images transmitted from each external device (200-203), but is not necessarily limited thereto, and can transmit all images from all external devices (200-203).

[0081] Each external device (200-203) can process images corresponding to the number of decoders and scalers and display them on the video conference screen. Each external device (200-203) can display the remaining images that it cannot display by replacing them with graphic objects (35, 36). The graphic objects (35, 35-1, 35-2, 36) are visually expressed objects and can be expressed as avatars, images, texts, shapes, icons, thumbnails, photos, etc.

[0082] Graphic objects can be created in a variety of ways. For example, the processor (130) can detect facial features of multiple users participating in a video conference from captured images and create graphic objects, such as avatars, based on these features. The processor (130) can create graphic objects in the form of vector images using 3D modeling software. Multiple users can also create their own avatars using an avatar creation tool linked to the video conference platform and use these avatars on the video conference platform.

[0083] The processor (130) can replace the remaining images, other than the images selectively transmitted among the images (31 to 34) for multiple users, with graphic objects (35, 36) and transmit them to each of the multiple external devices (200 to 203). However, this is not limited thereto, and each external device (200 to 203) can directly generate graphic objects (35, 35-1, 35-2, 36) and display them instead of the images of the remaining users.

[0084] Referring to FIG. 3, the electronic device (100) transmits a number of images that match the characteristic information of each external device (200 to 203) among the images of multiple users to each external device (200 to 203), and indicates a state in which each external device (200-203) displays the images.

[0085] According to FIG. 3, the processor (130) can transmit one or two images excluding the user image of external device 1 (200) to external device 1 (200) having two images that can be played simultaneously. The processor (130) can transmit three or four images excluding the user image of external device 2 (202) to external device 2 (201) having four images that can be played simultaneously. The processor (130) can transmit one or two images excluding the user image of external device 3 (202) to external device 3 (202) having two images that can be played simultaneously. The processor (130) can transmit two or three images excluding the user image of external device 4 (203) to external device 4 (203) having three images that can be played simultaneously.

[0086] That is, the processor (130) can provide a number of images equal to or less than the number of images that can be displayed on each external device (200 to 203). Each external device (200 to 203) can display its own images and received images, or display only the received images excluding its own images. The external devices (200 to 203) can determine whether to display their own images based on user settings.

[0087] Each external device (200 to 203) can display images that exceed the number of simultaneously playable images by replacing them with graphic objects (35, 35-1, 35-2, 36). In Fig. 3, as an example of a graphic object, an icon in the shape of an avatar (35, 35-1, 35-2) or a microphone (36) can be seen.

[0088] In the case of external device 1 (200), the processor (130) can transmit to external device 1 (200) the remaining images, other than one or two images selectively transmitted among multiple user images, by replacing them with an avatar (35-1) and a microphone-shaped icon (36). In addition, in the case of external device 3 (202), the processor (130) can transmit to external device 3 (202) the remaining images, other than one or two images selectively transmitted among multiple user images, by replacing them with an avatar (35) and a microphone-shaped icon (36).

[0089] Alternatively, as described above, a graphic object may be directly created to replace the remaining images other than those displayable on each external device (200-203). In this case, each external device (200-203) may replace the image captured by the user of the device with a graphic object based on user input.

[0090] When some images are replaced with graphic objects, the processor (130) may also provide information indicating which user the graphic object belongs to. For example, when a microphone-shaped icon is displayed, the processor (130) may display the user's name, ID, nickname, etc. corresponding to the icon on one side of the icon.

[0091] According to the above-described embodiment, the images to be replaced with graphic objects in each of the external devices (200 to 203) may be different. However, this is not necessarily limited to this, and multiple users may directly select a user to be replaced with a graphic object. For example, if all or a certain percentage of users decide to replace the image of the host of a video conference with a graphic object, the processor (130) may generate a graphic object to replace the image of the selected user (i.e., the host's image) and transmit the graphic object to all of the external devices (200 to 203).

[0092] In this way, the processor (130) can selectively transmit at least some of the images of the plurality of users to each external device (200 to 203) based on the characteristic information of each external device (200 to 203). When the images to be transmitted to the external devices (200 to 203) are changed, the processor (130) can select the changed images and transmit them to the external devices (200 to 203). In addition, the processor (130) can also change the images based on the speech situations of each of the plurality of users. A detailed description thereof will be provided later.

[0093] FIG. 4 is a diagram illustrating a process of transmitting a changed image of an electronic device according to at least one embodiment of the present disclosure.

[0094] When at least one of the plurality of users speaks, the processor (130) can transmit the video of at least one of the users speaking to the plurality of external devices (200 to 203) via the communication unit (110). When the speaker among the plurality of users changes, the processor (130) can replace the video being transmitted to the plurality of external devices (200 to 203) with the video of the changed speaker.

[0095] FIG. 4 illustrates an example of a screen change state of external device 1 (200), which is one of a plurality of external devices. Referring to FIG. 4, when one of the plurality of users (40) is speaking, the processor (130) provides the image (32) of the speaker (40) to external device 1 (200). In the above example, if there are four total conference users and two images can be displayed simultaneously on external device 1 (200), external device 1 (200) displays a video conference screen including the received image (32), the device's own image (31), and one graphic object (35-1, 36).

[0096] In this state, if the speaking user has changed, the processor (130) can transmit the image (33) of the changed user to external device 1 (200). For example, if a user (40) who was speaking previously among the meeting users ends speaking and another user (41) speaks, the processor (130) transmits the image (33) of the other user (41) to external device 1 (200).

[0097] External device 1 (200) replaces the previously displayed image (32) with a newly received image (33), and changes the graphic object (35-1) displayed to correspond to another user (41) to a graphic object (35-2) corresponding to the previously speaking user (40) and displays it.

[0098] In Fig. 4, the display screen of the external device 1 (200) is illustrated with a structure in which the user's own image (31) is displayed in the upper right corner, the image (32) of a user selected from among four users is displayed in the center left corner, and graphic objects (35-1, 35-2, 36), etc. are displayed below it. However, the display positions of the user images and graphic objects are not limited thereto. In addition, the processor (130) can display the images of the remaining users, excluding the user's image, as avatars (35-1, 35-2) or microphone-shaped icons (36) according to user input, and all of the user's images can be replaced with avatars (35-1, 35-2) or microphone-shaped icons (36) or other graphic objects.

[0099] FIGS. 3 and 4 illustrate a video conference screen in which at least one video and graphic object are distributed and displayed on the desktop. At this time, the desktop may display an arbitrary background image set by each external device or electronic device. However, the configuration of the video conference screen is not necessarily limited to this and may be modified in various ways. For example, the processor (130) may configure the video conference screen in a manner in which the video of the main speaker is displayed on the entire screen and other videos and graphic objects are superimposed on a certain area of ​​the entire screen. When the main speaker is changed, the processor (130) may replace the video of the main speaker shown on the entire screen with the video of another user. Alternatively, the processor (130) may configure the video conference screen in a manner in which the entire screen is divided into multiple areas and videos and graphic objects are arranged in each area.

[0100] Meanwhile, the processor (130) can determine the speech situation in various ways. For example, the processor (130) can determine the speech situation based on audio signals transmitted from each external device (200 to 203).

[0101] Each external device (200 to 203) can receive voices spoken by each user through a built-in microphone or a connected external microphone. When a voice having an audio level above a certain level is input, each external device (200 to 203) can transmit an audio signal including the input voice to the server device (100). The server device (100) can share the received audio signal with all external devices (200 to 203). Accordingly, users of each external device (200 to 203) can conduct a video conference while listening to the audio signal along with the screen.

[0102] When an audio signal is received from one of the external devices (200 to 203) during a video conference, the processor (130) can determine that the user of the external device that transmitted the audio signal is speaking. Accordingly, the processor (130) can select an image of the determined speaker and transmit it to each external device (200 to 203). If there is an external device among the external devices (200 to 203) that is displaying a graphic object for the currently speaking speaker, the external device replaces the graphic object with the received image and displays the image of the currently speaking speaker as soon as it is transmitted from the server device (100). Accordingly, all users can see the image of the currently speaking user and easily determine who is speaking.

[0103] As another example, the processor (130) may estimate a primary speaker based on the reception frequency of audio signals received from each external device (200 to 203). The primary speaker may be a moderator who speaks most frequently or a lot in a video conference with multiple users, or who conducts a conversation or conference, a user who is scheduled to speak when it is their turn to speak, or a presenter who is giving a presentation during a seminar or lecture.

[0104] For example, if the frequency of reception of audio signals transmitted from external device 1 (200) among a plurality of external devices (200 to 203) is the highest, the processor (130) may estimate the user of external device 1 as the primary speaker. The processor (130) transmits the image of external device 1 estimated as the primary speaker to other external devices, and the other external devices receiving the image display the image of external device 1. If the primary speaker is identified, the processor (130) may maintain the image of the primary speaker as it is even if another speaker speaks later. The processor (130) may also periodically check the frequency of reception of audio signals to update the primary speaker.

[0105] In the above, it has been described that the primary speaker is estimated based on the reception frequency, but the processor (130) may also estimate the primary speaker based on the speech time. When an audio signal is received, the processor (130) may calculate the speech time based on the length of the user's voice included in the audio signal, and accumulate the same to count the total speech time. For example, if it is determined that the user of external device 2 does not speak frequently but has spoken for the longest time, the processor (130) may estimate the user of external device 2 as the primary speaker. The processor (130) transmits an image of external device 2 estimated as the primary speaker to other external devices, and the other external devices that receive the image display the image of external device 2.

[0106] As another example, the processor (130) can analyze a plurality of image frames included in an image transmitted from each external device (200 to 203) to identify whether the user is speaking. Specifically, the processor (130) divides all pixels included in each of the plurality of consecutive image frames into a plurality of block units each consisting of n*m pixels. The processor (130) can detect a representative value representing the characteristics of the pixels in each block. The representative value may be, but is not limited to, an average pixel value of the pixels in each block, and may also be a maximum pixel value, a minimum pixel value, or an RMS (Root Means Square) value.

[0107] The processor (130) can connect blocks that form a closed loop and are arranged in a continuous position while having a similar range of representative values ​​among a plurality of blocks, and can identify the closed loop as an edge of an object included in a captured image. The processor (130) can identify the size of the object based on the number of blocks included in the edge. In addition, the processor (130) can identify the shape of the object based on the shape of the edge. Since the user's face is mostly circular or oval, the processor (130) can identify an edge arranged in a circular or oval shape and identify pixel values ​​within the edge to determine whether it is a part of the user's face. If it is determined to be a part of the face, the processor (130) can specify the user's mouth in a similar manner. The processor (130) can track changes in the positions of pixel blocks of the mouth in a plurality of consecutive image frames, and if a change in the position is detected, it can determine that the user is speaking.

[0108] Once the speaker is determined based on the image, the processor (130) can transmit the image of the speaker to another external device via the communication unit (110). In this case, similar to the example described above, the processor (130) can also detect the frequency of speaking or the cumulative time of speaking. The processor (130) can identify a speaker with a high frequency of speaking or a long cumulative time of speaking as the primary speaker. Once the primary speaker is identified, the processor (130) can transmit the image of the primary speaker to another external device (200 to 203) via the communication unit (110). The processor (130) can periodically re-identify the primary speaker to determine whether there has been a change.

[0109] When the primary speaker changes, the processor (130) can replace one of the images being transmitted to multiple external devices (200 to 203) with the image of the changed primary speaker. Accordingly, each external device can display the existing primary speaker image as a graphic object or image and display the newly received image of the primary speaker.

[0110] Additionally, the processor (130) may use an AI model or speaker prediction model to determine the primary speaker. For example, in a video conference for a seminar or lecture, the processor (130) may determine the presenter as the primary speaker. In cases where another participant asks a question, the processor (130) may determine the questioner as the primary speaker. For example, the processor (130) may determine the host of the video conference as the primary speaker.

[0111] According to another example, the processor (130) may estimate the next speaker based on the content of the video conference and transmit a video of the estimated speaker to each external device (200 to 203). When the speaker is estimated, the processor (130) may transmit the video to each external device (200 to 203) in advance before the estimated speaker speaks.

[0112] The processor (130) can recognize the speech content of each speaker using voice recognition technology and estimate the next speaker based on the recognition results. This will be further described in detail in the following section.

[0113] The above describes a case where the electronic device (100) determines the primary speaker, but it is not necessarily limited thereto, and external devices (200 to 203) can also determine the primary speaker among multiple users. The method for determining the primary speaker in the external devices (200 to 203) may be the same as in the electronic device (100) described above, so a duplicate description is omitted. When the external devices (200 to 203) determine the primary speaker, they can transmit information about the primary speaker to the electronic device (100).

[0114] The processor (130) can receive information about the primary speaker determined from each external device (200 to 203). The processor (130) can compare the information about the primary speaker received from each external device (200 to 203) with the information about the primary speaker determined by the processor (130) itself to ultimately determine the primary speaker. Alternatively, the processor (130) can independently determine the primary speaker regardless of the information about the primary speaker determined by the external devices (200 to 203).

[0115] The processor (130) can determine the primary speaker among multiple users in real time and transmit the image of the primary speaker to each external device (200 to 203). The processor (130) can determine the primary speaker according to a preset time cycle and transmit the image of the determined primary speaker to each external device (200 to 203).

[0116] The processor (130) can determine the primary speaker based on the voice signal received from each external device (200-203) or the speech situation for a preset period of time. Even if the primary speaker changes due to another user's speech for a preset period of time, the processor (130) can maintain the existing speaker as the primary speaker.

[0117] If multiple users are conversing with each other, the processor (130) may determine all users in the conversation as primary speakers. The processor (130) may split and display the images of the users in the conversation on a single screen and transmit them to external devices (200 to 203). If at least one of the external devices (200 to 203) cannot display multiple user images due to hardware characteristics, the processor (130) may configure a predetermined number of images that can be processed according to hardware characteristics as high-definition images and process the remaining images through software to provide them as standard or low-definition images.

[0118] FIG. 5 is a diagram illustrating a method for selecting a primary speaker of an electronic device according to at least one embodiment of the present disclosure.

[0119] The processor (130) can select a primary speaker based on at least one of the cumulative speech time, speech frequency, and speech topic of each of the plurality of users.

[0120] Referring to Fig. 5, a table is provided showing the cumulative speaking time (510), speaking frequency (520), and speaking topic suitability (530) of a total of four users. The processor (130) can compare the cumulative speaking time, speaking frequency, and speaking topic suitability of participants 1 to 4 to determine the main speaker. According to Fig. 5, the cumulative speaking time is expressed in minutes, the speaking frequency is expressed in units of the number of times spoken during a video conference, and the speaking topic suitability can be expressed in units of % indicating the degree of suitability to the conference topic. As described above, the speaking frequency, the cumulative speaking time, etc. can be determined based on the frequency of reception of the audio signal or the length of the audio signal, or the frequency of appearance of an image frame in which the mouth part changes within an image frame or the number of such image frames.

[0121] Videoconference topic relevance is information indicating the extent to which each speaker's utterances are appropriate for the videoconference topic. Speakers with high relevance to the videoconference topic may be selected as the primary speaker, even if their frequency or cumulative speaking time is lower than that of other speakers. Videoconference topic relevance can be determined based on contextual information, as described below.

[0122] When data such as that in FIG. 5 is secured, the processor (130) can determine (500) that participant 2, among the plurality of participants, has a cumulative time of 30 minutes, a frequency of 6 times, and a topic of utterance suitability of 70%, is the main speaker. The processor (130) can select a video of participant 2, who is determined to be the main speaker, and transmit the video of participant 2 to the external devices (200 to 203). In addition, when the speech of other participants is finished and no participant among the plurality of participants speaks, the processor (130) can automatically select participant 2, who is the main speaker, and transmit the video of participant 2 to the external devices (200 to 203).

[0123] When there is an external device capable of playing two videos simultaneously, the processor (130) can transmit the videos of participants 2 and 3 by considering the cumulative utterance time, utterance frequency, and utterance topic suitability, and can replace the videos of participant 1 and participant 4 with graphic objects and transmit them to the external device.

[0124] When a user finishes speaking during a video conference with multiple users, the processor (130) can automatically estimate the next speaker. The processor (130) can select the video of the estimated user and transmit it to each external device (200 to 203). Details regarding this will be described below.

[0125] FIG. 6 is a diagram illustrating a speaker prediction model of an electronic device according to at least one embodiment of the present disclosure.

[0126] The speaker prediction model of FIG. 6 may be stored in the memory (120) of FIG. 2, but is not necessarily limited thereto, and may also be stored in an external device connected to the electronic device (100).

[0127] When the processor (130) receives input data (610) including the speech voices of each of a plurality of users, it executes the speaker prediction model (600). The speaker prediction model (600) may be a software module for predicting the next speaker by analyzing the context of the speech contained in the input data (610).

[0128] The processor (130) can obtain context information using the speaker prediction model (600) and estimate the next speaker among multiple users based on the context information.

[0129] Specifically, the processor (130) can recognize the user's voice using the speaker prediction model (600). The processor (130) can receive voice signals from multiple users through microphones (not shown) located in each external device (200 to 203), or can receive voice signals from multiple users through a remote control (not shown) including a microphone. In addition, the processor (130) can receive voice signals from multiple users using user terminal devices through a remote control application.

[0130] For example, the processor (130) can extract a feature vector after removing noise from an input speech signal. The processor (130) can generate a phoneme sequence from the feature vector using an acoustic model. A phoneme sequence can be a continuous arrangement of phonemes, which are the smallest units of phonology that can distinguish meaning.

[0131] The processor (130) can generate a word sequence by decoding a phoneme sequence in word units, or can generate a syllable sequence by decoding a phoneme sequence in syllable units. The word sequence can be a continuous arrangement of words separated by spaces in a text that is a result of speech recognition. The syllable sequence can be a continuous arrangement of syllables, which are units of speech sounds having one phonetic value. The processor (130) can generate a plurality of words or a plurality of syllables from a phoneme sequence based on dictionary data stored in the memory (120).

[0132] The processor (130) can determine text as a result of recognition of a speech signal based on at least one of a word sequence and a syllable sequence. The processor (130) can recognize the meaning of the text based on previously stored dictionary data.

[0133] The processor (130) can estimate the next speaker among multiple users based on the meaning of the recognized text.

[0134] For example, if three users take turns giving presentations one by one, and the processor (130) determines that users 1 and 2 have finished presenting, it can estimate user 3 as the next speaker and automatically select user 3's video to transmit to each external device (200 to 203). In this case, the processor (130) may estimate the next speaker based on the reception pattern of the audio signal, without estimating the meaning of the text. For example, if audio signals are continuously input in a pattern such as user 1, user 2, user 1, user 3, user 1, and user 4, the processor (130) can estimate that the next speaker is user 4 when the audio signals of user 3 and user 1 are sequentially received.

[0135] For example, when one presenter and multiple users are having a video conference, and at least one user says “I will finish my presentation,” the processor (130) can assume that the presenter is the next speaker and automatically select the presenter’s video to transmit to each external device (200 to 203).

[0136] Additionally, the processor (130) may perform voice recognition to determine the speech content of multiple users when determining the appropriateness of the speech topic, as shown in FIG. 5, and may determine the primary speaker based on this. For example, if the topic of the video conference is "global warming," the processor (130) may identify how many texts related to global warming are included in the user speech content, and determine the appropriateness of the speech topic based on the results of this identification.

[0137] Alternatively, if the processor (130) includes a word (e.g., a name or nickname) that identifies another user in the speech content spoken by one user, the processor (130) may estimate the other user as the next speaker as soon as the user's speech ends.

[0138] As described above, the processor (130) can transmit output data (620) including the image of the next estimated speaker to multiple external devices via the communication unit (110).

[0139] The speaker prediction model may be referred to in various ways, such as a voice recognition estimation model, an AI model, a context estimation model, a speaker judgment model, etc., but in the present disclosure, it is referred to as a speaker prediction model.

[0140] The processor (130) can estimate the next speaker based on identification information for multiple users. Details on this will be described below.

[0141] FIG. 7 is a diagram for explaining a speaker estimation process of an electronic device according to at least one embodiment of the present disclosure.

[0142] If the speech of one of the multiple users includes identification information of another user, the processor (130) can transmit the other user's video to multiple external devices via the communication unit (110). The identification information can be information that can identify the user. For example, various information such as an ID, name, number, or IP address can be used.

[0143] Referring to Figure 7, multiple users can set their own identification information before or during a video conference. The set identification information may be displayed (72) on one side of the screen of each external device (200-203), but is not limited thereto, and may or may not be displayed at the bottom.

[0144] In Fig. 7, it can be seen that user KIM (10) has set his / her identification information to KIM. Except for user KIM (10), the remaining users can set various identification information (73) such as LEE, PARK, KANG, etc. For example, the processor (130) can select user PARK's video (70) excluding user KIM's own video from an external device (200 to 203) that can play two videos simultaneously, and transmit it to the external device (200 to 203).

[0145] At this time, if the user KIM speaks “LEE, please speak” through a microphone (not shown) built into the external device (200 to 203), the processor (130) can perform voice recognition to recognize the meaning of each word by dividing it into text words such as “LEE / speak / please” and combine the meanings to recognize the meaning of the entire text. Based on the recognition result, the processor (130) can select an image (71) of the user LEE whose identification information is set to LEE and transmit it to the external device (200 to 203).

[0146] In this way, the processor (130) can perform voice recognition to select an image of another user based on the identification information set by the user, if the user's spoken voice contains identification information of another user. Thereafter, the processor (130) can transmit the selected image of the other user to each external device (200 to 203).

[0147] Although the above description is limited to the case where the electronic device (100) is a server device, the description in the various embodiments above can also be performed on devices of other types than server devices. Specifically, various electronic devices such as PCs, laptop PCs, smartphones, tablet PCs, TVs, and kiosks can directly provide video calls, and in this case, the operations described in the various embodiments above can be performed.

[0148] FIG. 8 is a drawing for explaining an electronic device according to at least one embodiment of the present disclosure.

[0149] The electronic device (800) may refer to an electronic device that directly includes a display or is connected to an external display (e.g., a monitor). Specifically, the electronic device (800) may be implemented as various devices such as a set-top box, a PC, a laptop PC, a smartphone, a tablet PC, a TV, a kiosk, etc.

[0150] When a user of an electronic device (800) wishes to hold a video conference with other users, the user can run a video conference program. In this case, a video conference can be established between the electronic device (800) and other electronic devices (801 to 803) without going through a server device. Each of the electronic devices (800, 801 to 803) may have a built-in camera or be connected to a camera or an external device including the camera. Accordingly, each of the electronic devices (800, 801 to 803) can generate its own video and transmit it to other external devices (801, 802, 803). Each video may be a video taken of the user, but is not necessarily limited thereto, and may also be an image of a conference data screen, etc.

[0151] The electronic device (800) may include a communication unit (810), a memory (820), and a processor (830). However, the present invention is not limited thereto, and the electronic device (100) may be implemented in a form in which some components are excluded, or may be implemented in a form in which other components are additionally included. In the description of the communication unit (810), the memory (820), and the processor (830), any content that overlaps with that described in other embodiments described above will be omitted.

[0152] The communication unit (810) can receive images from external devices (801 to 803) participating in a video call. In addition, the communication unit (810) can also transmit images of the electronic device (800) itself to external devices (801 to 803).

[0153] When receiving images of all users participating in a video call, the processor (830) can provide a video conference screen including a number of images corresponding to the characteristic information of the electronic device (800) itself. The processor (830) can change the images displayed on the video conference screen based on the speech status of all users. Information on the speech status of all users can be determined based on audio signals or image signals received from external devices (801 to 803). Since the details thereof are the same as those of the processor (130) of the electronic device (100), a detailed description thereof will be omitted.

[0154] Additionally, the processor (830) can perform speech recognition using a speaker prediction model, estimate the next speaker based on speech signals from multiple users, and automatically select an image of the estimated speaker and transmit it to each external device (801 to 803). A detailed description thereof is omitted as it has been described above.

[0155] The electronic device (800) may also operate as one of the external devices (200 to 203) of FIG. 1. For example, the electronic device (800) may transmit a user's image to each external device (801 to 803) via a server device (not shown) and may also receive a user's image from the external devices (801 to 803).

[0156] In addition, when a video call involving the user of the electronic device (800) and at least one other user is executed, the processor (830) can transmit characteristic information of an external device stored in the memory (820) to the server device through the communication unit (810). The processor (830) can receive a number of images corresponding to the characteristic information from the server device through the communication unit (810). The processor (830) can provide a video conference image including the received image to the external device.

[0157] FIG. 9 is a flowchart illustrating a method for controlling an electronic device according to at least one embodiment of the present disclosure.

[0158] Referring to FIG. 9, when an electronic device performs a video call with multiple users, the electronic device transmits characteristic information of the electronic device to the server (S910). The electronic device receives and stores characteristic information of multiple external devices used by multiple users from the server (S920). The electronic device receives a number of images corresponding to the characteristic information of each external device among the images of multiple users performing the video call from the server (S930). The electronic device modifies the received images based on the speech situations of each of the multiple users, and outputs the modified images through the display (S940). Since the specific method for this has been specifically described in the various embodiments described above, a redundant description will be omitted.

[0159] FIG. 10 is a diagram illustrating a process for identifying a primary speaker based on priority information of an electronic device according to at least one embodiment of the present disclosure.

[0160] The processor (130) can identify a primary speaker based on priority information of multiple users. The processor (130) can control the display (140) to output an image of the identified primary speaker. Here, priority information refers to information set in advance to determine which user among multiple users speaking simultaneously can become the primary speaker.

[0161] Referring to FIG. 10, if a user sets priority information (1010) among multiple users to KIM as the first priority, LEE as the second priority, and PARK as the third priority, and if both KIM and LEE speak, the processor (130) can identify KIM as the primary speaker (1020) based on the preset priority information.

[0162] A user can set priority information collectively for multiple users, or can set it solely for the electronic device, unlike external devices used by multiple users. The processor (130) can display the image of the user identified as the primary speaker through the display (140) based on the user-defined priority information. The processor (130) can transmit the image of the user identified as the primary speaker to the server through the communication unit (110) based on the priority information.

[0163] FIG. 11 is a diagram illustrating a process for identifying a primary speaker based on a trigger word of an electronic device according to at least one embodiment of the present disclosure.

[0164] If the context information obtained from the speech of each of multiple users includes a trigger text, the processor (130) can identify the user corresponding to the trigger text as the next speaker. The processor (130) can control the display (140) to output an image of the next identified speaker.

[0165] Trigger text is text that is set to trigger a special process when a specific keyword is entered, either to refer to a specific speaker or to move on to the next speaker.

[0166] Referring to FIG. 11, the processor (130) can identify the next speaker based on a trigger text set by a user. The user can set the trigger text by himself or herself, or can set it together with multiple users. The processor (130) can set the trigger text "I will finish my presentation" from the user. When the processor (130) acquires the voice "I will finish my presentation" from the voice of a specific user (1120) among multiple users (1110), the processor (130) can identify a user who has not spoken among the multiple users as the next speaker (1130), or can identify the next speaker (1130) based on the order in which they previously spoke.

[0167] For example, when discussing a particular topic, the processor (130) may identify the opposing user as the next speaker based on the trigger text "Please state the opposing side's position."

[0168] FIG. 12 is a flowchart illustrating the overall operation process of an electronic device according to at least one embodiment of the present disclosure.

[0169] Referring to FIG. 12, when an electronic device performs a video call with multiple users (S1210), the electronic device transmits characteristic information of the electronic device to a server (S1220). The electronic device receives and stores characteristic information of multiple external devices used by multiple users from the server (S1230). The electronic device receives from the server a number of images corresponding to characteristic information of each external device among the images of multiple users performing the video call (S1240). The electronic device changes the received images based on the speech situations of each of the multiple users and outputs them through a display (S1250).

[0170] The explanation for this has been provided in the above section, so the specific details will be omitted.

[0171] The various embodiments described above may be implemented as a single embodiment, or at least one embodiment may be combined with each other in whole or in part and implemented together in one device.

[0172] According to the various embodiments described above, high-quality video conferencing is possible in the same environment when video conferencing with each electronic device having different specifications.

[0173] Meanwhile, the various embodiments described above may be applied to a product as an embodiment alone, but at least some of the contents may be implemented in combination with other embodiments of the present disclosure.

[0174] The various embodiments described above can be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored in the storage medium and operate according to the called instructions, and may include an electronic device (100) according to the disclosed embodiments. When an instruction is executed by a processor, the processor can perform a function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter. The machine-readable storage medium can be provided in the form of a non-transitory computer-readable storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.

[0175] Additionally, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product.

[0176] Specifically, a non-transitory readable storage medium or a computer program product storing computer instructions that cause the computer to perform an operation including a step of determining the number of images to be transmitted to each electronic device based on characteristic information of a plurality of external devices used by a plurality of users participating in a video call, a step of transmitting the determined number of images for each of the plurality of external devices among the images for the plurality of users to each external device, and a step of changing the images to be transmitted to the plurality of external devices based on the speech situations of each of the plurality of users may be provided.

[0177] The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0178] In addition, computer instructions or programs for performing the control methods of electronic devices according to the various embodiments described above may be stored in a non-transitory computer-readable medium. The computer instructions stored in such a non-transitory computer-readable medium, when executed by a processor of a specific device, cause the specific device to perform processing operations in the device according to the various embodiments described above. A non-transitory computer-readable medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of non-transitory computer-readable media may include a CD, DVD, hard disk, Blu-ray disk, USB, memory card, or ROM.

[0179] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person skilled in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, Department of Communications; display; memory; and Processor; including; The above processor, When making a video call with multiple users, the characteristic information of the electronic device is transmitted to the server through the communication unit, Receives a number of images corresponding to the external device characteristic information of each of the images of multiple users performing the video call from the server through the communication unit, An electronic device that controls the display so that the received image is changed and output based on the speech situations of each of the plurality of users.

2. In paragraph 1, The above processor, Among the videos for the plurality of users, a number of videos corresponding to the number of decoders or scalers included in the characteristic information of the plurality of external devices is selectively received from the server, An electronic device that controls the display so that, among the images for the plurality of users, the remaining images other than the selectively received images are displayed by replacing them with graphic objects.

3. In paragraph 1, The above processor, When at least one of the above plurality of users speaks, audio data of at least one of the speaking users is received through the communication unit, At least one user's video corresponding to the audio data is received from the server and a number of videos corresponding to the characteristic information of the electronic device are output through the display, An electronic device that, when audio data of another user among the plurality of users is received through the communication unit, replaces the image output through the display with a user image corresponding to the audio data of the other user.

4. In paragraph 1, The above processor, Identifying a primary speaker among the plurality of users based on audio data received through the communication unit, The image of the identified primary speaker is output through the display, An electronic device that, when the primary speaker of audio data received through the above communication unit changes, replaces the image output through the display with the image of the primary speaker changed based on the received audio data.

5. In paragraph 4, The above processor, An electronic device that identifies the primary speaker based on at least one of the cumulative speaking time, speaking frequency, and speaking topic of each of the plurality of users.

6. In paragraph 1, The above processor, Identifying the next speaker corresponding to the context information among the plurality of users based on the context information obtained from the speech voices of each of the plurality of users, An electronic device that transmits an image of the identified next speaker to the server through the communication unit.

7. In paragraph 1, The above memory contains identification information of each of the plurality of users, The above processor, If the speech voice of at least one of the above multiple users includes identification information of another user, the other user corresponding to the identification information is identified as the primary speaker, An electronic device that transmits an image of an identified primary speaker to the server through the communication unit.

8. In paragraph 1, The above processor, An electronic device that controls the display so that an image of a primary speaker identified based on priority information of the plurality of users is output.

9. In paragraph 6, The above processor, An electronic device that controls the display so that, when the context information obtained from the speech of each of the plurality of users includes the trigger text corresponding to the identification of each of the plurality of users, the image of the user corresponding to the trigger text is output as the image of the next speaker.

10. In paragraph 1, The above processor, An electronic device that controls the display so that, when input data including the speech voices of each of the plurality of users is received through the communication unit, an image of the speaker identified through a speaker prediction model is output.

11. In a method for controlling an electronic device, A step of transmitting characteristic information of the electronic device to a server when performing a video call with multiple users; A step of receiving, from the server, a number of images corresponding to each external device characteristic information among images of multiple users performing the video call; and A control method, comprising: a step of changing the received image and outputting it through a display based on the speech situation of each of the plurality of users.

12. In paragraph 11, A step of selectively receiving from the server a number of images corresponding to the number of decoders or the number of scalers, based on characteristic information of the plurality of external devices, among images for the plurality of users, including information on the number of decoders or the number of scalers included in each external device; and A control method, comprising: a step of displaying, through a display, the remaining images, excluding the selectively received images, among the images for the plurality of users, by replacing them with graphic objects.

13. In paragraph 11, When at least one of the plurality of users speaks, a step of receiving audio data of at least one of the speaking users; A step of receiving images of at least one user corresponding to the audio data from the server and outputting a number of images corresponding to the characteristic information of the electronic device; and A control method comprising: a step of replacing the output image with a user image corresponding to the audio data of another user among the plurality of users when audio data of another user is received; 14. In paragraph 11, A step of identifying a primary speaker among the plurality of users based on the received audio data; A step of outputting an image of the identified primary speaker; and A control method, comprising: a step of replacing the output image with an image of the changed main speaker based on the received audio data when the main speaker of the received audio data is changed.

15. A non-transitory computer-readable storage medium storing computer instructions that, when executed by a processor of an electronic device, cause the electronic device to perform an operation, the operation comprising: A step of transmitting characteristic information of the electronic device to a server when performing a video call with multiple users; A step of receiving, from the server, a number of images corresponding to the respective external device characteristic information among the images of a plurality of users performing the video call; and A non-transitory computer-readable storage medium, comprising: a step of changing the received image and outputting it through a display based on the speech situations of each of the plurality of users.

Citation Information

Patent Citations

  • A method for performing a video conference in a portable terminal and an apparatus thereof

    KR1020090063608A

  • Video conferencing with unlimited dynamic active participants

    KR1020140138609A

  • Composition for prevention of pulmonary hypertension, prevention of liver disease caused by pulmonary hypertension, and improvement of liver function and Functional food comprising the same

    KR1020230054303A

  • Apparatus and method for controlling air pump

    KR1020250118910A

  • Server providing business establishment supprot service and method thereof

    KR102214687B1