Electronic device, method, and non-transitory computer-readable storage medium for translating utterance in call
By capturing and analyzing user interface images before and after touch inputs, the electronic device accurately identifies and responds to functions like mute during VoIP calls, addressing inconsistencies across different service providers and enhancing user experience.
Patent Information
- Application Number
- PCT/KR2025/006767
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-03
- Filing Date
- 2025-05-19
- Publication Date
- 2026-01-08
AI Technical Summary
Existing digital assistants struggle to accurately identify and respond to user inputs for functions like mute during voice over internet protocol (VoIP) calls due to varying object arrangements in user interfaces across different service providers, leading to inconsistencies and user inconvenience.
An electronic device and method that utilize a processor to capture images of the user interface before and after a touch input, identifying the function activated by the input and adjusting the execution state of a second application accordingly, ensuring consistent functionality across different service providers.
Ensures accurate identification and response to user inputs for functions like mute during VoIP calls, maintaining user expectations and reducing inconsistencies across different service providers.
Smart Images

Figure KR2025006767_08012026_PF_FP_ABST
Abstract
Description
Electronic device, method, and non-transitory computer-readable storage medium for translating speech within a call
[0001] The following descriptions relate to electronic devices, methods, and non-transitory computer-readable storage media for translating speech within a call.
[0002] Digital assistants (or virtual assistants or intelligent automated assistants) can provide a beneficial human-machine interface. For example, the digital assistant can be used to translate from one language to another. For example, the digital assistant can be used to translate a user's speech during a phone call.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] An electronic device is provided. The electronic device may include a display. The electronic device may include communication circuitry. The electronic device may include at least one processor including a processing circuit. The electronic device may include a memory storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, using the display, a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, using the display, a second UI of a second application for performing a translation service for displaying text related to a translation of a first utterance of a user of the electronic device and a translation of a second utterance of a user of the other electronic device, the translation service being performed through a second application during the VoIP call. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a second image for the first UI based on a release of the touch input.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify that a mute function of the first application is activated based on the first image and the second image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to stop displaying text related to a translation of the first utterance within the second UI while the mute function is activated, based on the identification that the mute function is activated.
[0005] A method is provided. The method can be executed in an electronic device including a communication circuit and a display. The method can include an operation of displaying a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuit, using the display. The method can include an operation of displaying a second UI of a second application for performing a translation service for displaying text related to a translation of a first utterance of a user of the electronic device and a translation of a second utterance of the user of the other electronic device, performed through a second application during the VoIP call, using the display together with the first UI. The method can include an operation of obtaining a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI. The method can include an operation of obtaining a second image for the first UI based on a release of the touch input. The method may include an operation for identifying that a mute function of the first application is activated based on the first image and the second image. The method may include an operation for stopping displaying text related to the translation of the first utterance within the second UI while the mute function is activated, based on the identification that the mute function is activated.
[0006] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device having a communication circuit and a display, cause the electronic device to display, using the display, a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuit. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, using the display, a second UI of a second application for performing a translation service for displaying text related to a translation of a first utterance of a user of the electronic device and a translation of a second utterance of a user of the other electronic device, the translation service being performed through a second application during the VoIP call. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a second image for the first UI based on a release of the touch input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify, based on the first image and the second image, that a mute function of the first application is activated in response to the touch input.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to stop displaying text related to the translation of the first utterance within the second UI while the mute function is activated, based on the identification that the mute function is activated.
[0007] An electronic device is provided. The electronic device may include a display. The electronic device may include communication circuitry. The electronic device may include at least one processor including a processing circuit. The electronic device may include a memory storing instructions and including one or more storage media. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, using the display, a first user interface (UI) of a first application for a voice over internet protocol (VoIP) call with another electronic device using the communication circuitry. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, using the display, a second UI of a second application for performing a translation service for displaying text related to a translation of a first utterance of a user of the electronic device and a translation of a second utterance of a user of the other electronic device, the translation service being performed through a second application during the VoIP call. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a second image for the first UI based on a release of the touch input.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify that a video call function of the first application is activated based on the first image and the second image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to stop displaying the second UI while the video call function is activated based on the identification that the video call function is activated.
[0008] A method is provided. The method can be executed in an electronic device including a display and a communication circuit. The method can include an operation of displaying a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuit, using the display. The method can include an operation of displaying a second UI of a second application for performing a translation service for displaying text related to a translation of a first utterance of a user of the electronic device and a translation of a second utterance of the user of the other electronic device, performed through a second application during the VoIP call, using the display together with the first UI. The method can include an operation of obtaining a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI. The method can include an operation of obtaining a second image for the first UI based on a release of the touch input. The method may include an operation of identifying that the video call function of the first application is activated based on the first image and the second image. The method may include an operation of stopping displaying the second UI while the video call function is activated based on the identification that the video call function is activated.
[0009] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device having a communication circuit and a display, cause the electronic device to display, using the display, a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuit. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display, using the display, a second UI of a second application for performing a translation service for displaying text related to a translation of a first utterance of a user of the electronic device and a translation of a second utterance of a user of the other electronic device, the translation service being performed through a second application during the VoIP call. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a second image for the first UI based on a release of the touch input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify that a video call function of the first application is activated based on the first image and the second image.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to stop displaying the second UI while the video call function is activated, based on the identification that the video call function is activated.
[0010] Figure 1 illustrates examples of user interfaces (UIs) of an application for VoIP (voice over internet protocol) calls.
[0011] Figure 2 is a simplified block diagram of an exemplary electronic device.
[0012] FIG. 3a is a flowchart illustrating an exemplary method for identifying execution of a function of a first application for a VoIP call.
[0013] FIG. 3b is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application for a VoIP call.
[0014] Figure 4 illustrates an example of the first UI.
[0015] Figure 5 illustrates an example of a second UI displayed together with a first UI.
[0016] Figure 6 shows examples of the first image and the second image.
[0017] Figure 7 illustrates an example of stopping displaying the results of translation of a user's speech of an electronic device.
[0018] FIG. 8 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application using an area on a display that includes a location where a touch input is received.
[0019] FIG. 9 illustrates an example of a portion of a first UI and a portion of a second UI positioned within an area on a display including a location where a touch input is received.
[0020] Figure 10 shows an example of comparing a first image and a second image.
[0021] FIG. 11 is a flowchart illustrating an exemplary method for identifying whether an area on a display including a location where a touch input is received is included within a reference area on the display.
[0022] Figure 12 illustrates an exemplary method of using a reference area on a display.
[0023] FIG. 13 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application using a trained model in connection with identifying a location where a touch input was received.
[0024] Figure 14 illustrates a method of using a training model.
[0025] FIG. 15 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application through comparison between a first image and a second image.
[0026] Figure 16 illustrates an example of a portion of a first image that is different from a second image and a portion of a second image that is different from the first image.
[0027] FIG. 17 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application using a trained model in relation to a comparison between a first image and a second image.
[0028] FIG. 18 is a flowchart illustrating an exemplary method for identifying the deactivation of the mute function of a first application for a VoIP call.
[0029] FIG. 19a is a flowchart illustrating an exemplary method for identifying activation of a video call function of a first application for a VoIP call.
[0030] Figure 19b illustrates an example of a first UI having a changed state according to the release of a touch input.
[0031] FIG. 20 is a block diagram of an electronic device within a network environment according to various embodiments.
[0032] Figure 21 is a schematic diagram of an exemplary AI system.
[0033] VoIP calls can be a method of voice communication for transmitting voice communication sessions over an IP network, such as the Internet. VoIP calls can be described as Internet telephony, broadband telephony, and broadband telephony services. VoIP calls can also include providing other communication services (e.g., short message service (SMS) or voice messaging) over the Internet.
[0034] The user interface (UI) of the first application for the VoIP call may include objects (or executable objects) for performing functions available during the VoIP call. The arrangement of the objects within the UI may vary. Examples of various arrangements of the objects within the UI are described with reference to FIG. 1.
[0035] Figure 1 illustrates examples of user interfaces (UIs) of an application for VoIP (voice over internet protocol) calls.
[0036] Referring to FIG. 1, the UI (110) of an application for a VoIP call of a first service provider may include a first object (111) for a voice filter applied to a user's speech, a second object (112) for a mute function, a third object (113) for terminating a VoIP call, a fourth object (114) for a speaker phone function, and a fifth object (115) for a video call function.
[0037] The UI (140) of the application for a VoIP call of a second service provider may include a first object (141) for a speaker phone function, a second object (142) for a video call function, a third object (143) for a mute function, and a fourth object (144) for terminating a VoIP call. The arrangement of the first object (141), the second object (142), the third object (143), and the fourth object (144) within the UI (140) may be different from the arrangement of the first object (111), the second object (112), the third object (113), and the fourth object (114) within the UI (110). For example, the position (or coordinates) of the second object (112) for the mute function within the UI (110) may be different from the position (or coordinates) of the third object (143) for the mute function. For example, the shape of the second object (112) within the UI (110) for the mute function may be different from the shape of the third object (143) for the mute function. For example, the visual characteristics of the second object (112) within the UI (110) for the mute function may be different from the visual characteristics of the third object (143) for the mute function.
[0038] The UI (170) of the application for VoIP calls of a third-party service provider may include a first object (171) for a mute function, a second object (172) for a video call function, a third object (173) for changing a device that outputs the other party's speech, a fourth object (174) for pausing a VoIP call, a fifth object (175) for displaying objects for performing additional functions, and a sixth object (176) for terminating a VoIP call. The arrangement of the first object (171), the second object (172), the third object (173), the fourth object (174), the fifth object (175), and the sixth object (176) within the UI (170) may be different from the arrangement of the first object (111), the second object (112), the third object (113), and the fourth object (114) within the UI (110). The arrangement of the first object (171), the second object (172), the third object (173), the fourth object (174), the fifth object (175), and the sixth object (176) within the UI (170) may be different from the arrangement of the first object (141), the second object (142), the third object (143), and the fourth object (144) within the UI (140). For example, the position (or coordinates) of the first object (171) within the UI (170) for the mute function may be different from the position (or coordinates) of the second object (112) within the UI (110) for the mute function. For example, the position (or coordinates) of the first object (171) within the UI (170) for the mute function may be different from the position (or coordinates) of the third object (143) for the mute function.
[0039] As illustrated in FIG. 1, since the arrangement of objects within the UI of a first application for VoIP calls varies depending on the service provider, the processor may not be able to identify which object among the objects a touch input received while the UI of the first application is displayed is for. Since the object associated with the touch input is not identified, the processor may not be able to change the execution status of a second application that provides a service in conjunction with the first application in relation to an object within the UI of the first application associated with the touch input. The inability to change the execution status of the second application in relation to an object within the UI of the first application may cause inconvenience.
[0040] As illustrated in FIG. 1, since the arrangement of objects within the UI of a first application for VoIP calls varies depending on the service provider, a function executed in response to a touch input to one of the objects within the UI may not be recognized by a second application that is distinct from the first application. For example, a function executed in response to a touch input to one of the objects within the UI may not be identified by the second application that is distinct from the first application. Since the function of the first application is not identified by the second application, a method for identifying the execution status of the first application for the second application may be required.
[0041] For example, the function of the other application may be performed in conjunction with the function of the first application, which is performed in response to the touch input. As a non-limiting example, the second application may be used to translate utterances made during a VoIP call conducted via the first application.
[0042] An application for VoIP calls may support a mute function during a VoIP call. The application for VoIP calls may be an application that deactivates a microphone while the mute function is activated, or an application that activates a microphone while the mute function is activated. The first application may be an application for VoIP calls that activates a microphone while the mute function is activated. For example, the first application may maintain the state of a microphone activated for the VoIP call while the mute function of the first application is activated, and may refrain from transmitting data regarding speech received through the microphone.
[0043] A touch input for an object for a mute function within the UI of the first application may not be identified by the second application. Since the second application cannot identify which object within the UI of the first application the touch input received while the UI of the first application is displayed is for, the second application may not be able to identify that the mute function of the first application is activated based on the touch input.
[0044] For example, since the location of an object for a mute function within the UI of the first application is not known to the second application, the second application may not recognize that the mute function of the first application is activated based on a touch input to the object within the UI of the first application. For example, contrary to the user's perception that the mute function of the first application is activated, the second application may display the result of a translation of a speech received through the microphone while the mute function is activated. For example, the result of the translation that is different from the user's perception may cause inconvenience.
[0045] To reduce this inconvenience, the operation of identifying the mute function of the first application for the second application may be performed within the electronic device described below. For example, since the mute function of the first application keeps the microphone activated, the operation may be distinguished from identifying the state of the microphone. Components of the electronic device that execute (or perform) the operation are described with reference to FIG. 2.
[0046] Figure 2 is a simplified block diagram of an exemplary electronic device.
[0047] Referring to FIG. 2, the electronic device (200) may include at least one processor (210), memory (220), communication circuit (230), display (240), microphone (250), and speaker (260).
[0048] At least one processor (210) may include any processing circuitry operative to control the performance and operations of one or more components (e.g., communication circuitry (230), display (240), microphone (250), and / or speaker (260)) of the electronic device (200). The at least one processor (210) may execute instructions stored in the memory (220) to perform the operations described with reference to FIGS. 3A to 19B. The at least one processor (210) may be implemented as a single chip (or a single chip set), such as a system on chip (SoC). For example, the at least one processor (210) may also be implemented as multiple chips (or multiple chip sets).
[0049] For example, at least one processor (210) may include a central processing unit (CPU) (e.g., including a processing circuit). For example, at least one processor (210) may include a neural processing unit (NPU). As a non-limiting example, the CPU and the NPU may be configured to interact with the AI system illustrated in FIG. 21 or control the AI system illustrated in FIG. 21. For example, at least one processor (210) may include at least a portion of the processor (2020) of FIG. 20 or may correspond to at least a portion of the processor (2020) of FIG. 20.
[0050] For example, at least one processor (210) may be used to execute or run one or more software applications, such as an operating system software application, a firmware software application, a media playback software application, a media editing software application, a VoIP calling software application, a translation software application, a digital assistant software application, and / or any other suitable software applications.
[0051] The memory (220) may include one or more storage media. For example, the one or more storage media may include a hard drive, flash memory, permanent memory such as read-only memory (ROM), semi-permanent memory such as random access memory (RAM), any other suitable type of storage assembly, or any combination thereof. The memory (220) may include a cache memory, which is one or more different types of memory used to temporarily store data for the function or feature of the electronic device (200). The memory (220) may be fixedly embedded within the electronic device (200) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and / or a secure digital (SD) memory card) that can be repeatedly inserted into and removed from the electronic device (200). For example, the memory (220) may include at least a portion of the memory (2030) of FIG. 20 or correspond to at least a portion of the memory (2030) of FIG. 20.
[0052] The memory (220) may store one or more software applications, such as an operating system software application, a firmware software application, a media playback software application, a media editing software application, a VoIP calling software application, a translation software application, a digital assistant software application, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by at least a portion of at least one processor (210).
[0053] The communication circuit (230) may be used for communication with another electronic device (or an external electronic device). For example, the communication circuit (230) may be used to transmit packets for a VoIP call to the other electronic device and / or to receive packets for a VoIP call from the other electronic device. For example, the communication circuit (230) may be used to transmit the results of a translation obtained from a translation software application to the other electronic device. For example, the communication circuit (230) may include at least a portion of the communication module (2090) of FIG. 20 or may correspond to at least a portion of the communication module (2090) of FIG. 20.
[0054] The display (240) may be used to display visual information, visual data, screens, and / or user interfaces. For example, the display (240) may be used to display a UI of a first software application for VoIP calls. For example, the display (240) may be used to display a UI of a second software application for translation. For example, the display (240) may include at least a portion of the display module (2060) of FIG. 20 or correspond to at least a portion of the display module (2060) of FIG. 20.
[0055] The microphone (250) may be used to capture audio generated in connection with the electronic device (200). For example, the microphone (250) may receive speech from a user of the electronic device (200) while making a VoIP call.
[0056] The speaker (260) can be used to output audio. For example, the speaker (260) can output speech from a user of another electronic device acquired through a VoIP call conducted via the first application. For example, the speaker (260) can output the results of a translation conducted via the second application as audio.
[0057] FIG. 3a is a flowchart illustrating an exemplary method for identifying execution of a function of a first application for a VoIP call.
[0058] Referring to FIG. 3A, in operation 301, at least one processor (210) may display a first UI (e.g., UI (110, 140, 170)) for a VoIP call with another electronic device performed through a first application on a display (240). As a non-limiting example, at least some of the functions provided through the first application may be linked to a second application.
[0059] In operation 302, at least one processor (210) may display, on the display (240), a second UI for translating a first utterance of the user of the electronic device (200) and a second utterance of a user of the other electronic device (e.g., a call partner of the user of the electronic device (200)), performed through a second application (e.g., the second application described with reference to FIG. 2) during the VoIP call, together with the first UI. The second application may be described as an application that translates the first utterance in a first language into a second language and translates the second utterance in the second language into the first language. As a non-limiting example, the second application may be described as an application that cannot recognize what function the first application provides (or is providing).
[0060] A second UI displayed together with a first UI may be overlapped (or overlaid) on the first UI. The second UI overlapped on the first UI may be translucent to allow the first UI to be visible. The second UI displayed together with the first UI may be overlapped on a portion of the first UI. The second UI overlapped on the first UI may be a pop-up window positioned over the first UI. The second UI displayed together with the first UI may be displayed simultaneously with the first UI in a split view.
[0061] The first UI and the second UI can be displayed as PIP (picture in picture) or PBP (picture by picture).
[0062] In operation 303, at least one processor (210) may obtain a first image for the first UI based on a touch input received on the display (240) while displaying the second UI together with the first UI. The first image may be an image displayed at a location corresponding to a location where the touch input was received.
[0063] For example, acquiring the first image may be performed before recognizing for which function the touch input is received. For example, acquiring the first image may be performed before the touch input is released.
[0064] In operation 304, at least one processor (210) may obtain a second image for the first UI based on the release of the touch input received in operation 303. The second image may be an image displayed at a location corresponding to a location where the touch input is released.
[0065] In operation 305, at least one processor (210) can at least partially change the execution state of the second application using the first image and the second image. At least one processor (210) can identify a function of the first application executed by the touch input using the first image and the second image, and change the execution state of the second application according to the identified function.
[0066] For example, at least one processor (210) can identify a difference between the first image and the second image by comparing the first image and the second image. Based on the difference, at least one processor (210) can identify a function of the first application that is activated (or executed) (or performed) according to the touch input. As a non-limiting example, at least one processor (210) can identify a change in a display state (or view) of an object (or an executable object) (or a UI (user interface) object) within the first application through the difference and recognize the object whose display state has changed, thereby identifying the function.
[0067] For example, at least one processor (210) may determine whether to change the state of the second application running in conjunction with the first application according to the function. For example, at least one processor (210) may maintain the execution state of the second application based on a determination that the function is not related to the second application. For example, at least one processor (210) may change the execution state of the second application based on a determination that the function is related to the second application. For example, at least one processor (210) may change the execution state of the second application by activating (or executing) a function of the second application related to the function of the first application. For example, at least one processor (210) may change the execution state of the second application by deactivating the function of the second application according to the function of the first application. The change of the function of the first application and the execution state of the second application is described in more detail with reference to FIG. 3B.
[0068] FIG. 3b is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application for a VoIP call.
[0069] Referring to FIG. 3B, in operation 310, at least one processor (210) may display, on the display (240), a first UI for a VoIP call with another electronic device performed through a first application (e.g., the first application described with reference to FIG. 2). The first UI of the first application may be available for performing the VoIP call. The first application may be configured to maintain the state of a microphone (e.g., a microphone (250) or a microphone wirelessly or wiredly connected to the electronic device (200)) activated for the VoIP call while activating a mute function of the first application, and to refrain from transmitting data (or packets) regarding a first utterance of a user of the electronic device (200) received through the microphone to the other electronic device. The first UI is described with reference to FIG. 4.
[0070] Figure 4 illustrates an example of the first UI.
[0071] Referring to FIG. 4, a first UI (400) may be displayed on a display (240) for the VoIP call performed through the first application. The first UI (400) may include objects for performing each of the functions of the first application provided in relation to the VoIP call. For example, the first UI (400) may include a first object (401) for a speaker phone function of the first application, a second object (402) for a video call function of the first application, a third object (403) for a mute function of the first application, and a fourth object (404) for an end function of the VoIP call. The first object (401) of FIG. 4 may indicate that the speaker phone function of the first application is disabled. For example, the speaker phone function may be activated through the first application based on a touch input to the first object (401) of FIG. 4. The second object (402) of FIG. 4 may indicate that the video call function of the first application is disabled. For example, the video call function may be activated through the first application based on a touch input to the second object (402) of FIG. 4. The third object (403) of FIG. 4 may indicate that the mute function of the first application is disabled. For example, the mute function may be activated through the first application based on a touch input to the third object (403) of FIG. 4. The fourth object (404) of FIG. 4 may indicate that the end function of the VoIP call of the first application is disabled. For example, the end function of the VoIP call may be activated through the first application based on a touch input to the fourth object (404) of FIG. 4.
[0072] Referring back to FIG. 3B, at operation 320, at least one processor (210) may display, on the display (240), a second UI for translating a first utterance of a user of the electronic device (200) and a second utterance of a user of the other electronic device (e.g., a call partner of the user of the electronic device (200)), performed via a second application (e.g., the second application described with reference to FIG. 2) during the VoIP call, together with the first UI. The second UI of the second application may be available for performing a translation service for displaying text related to the translation of the first utterance and the translation of the second utterance performed during the VoIP call. For example, the second application may be described as an application that cannot recognize some of the information related to the execution of the first application. For example, the arrangement of objects within the first UI may be unnoticeable to the second application. For example, the arrangement of the objects within the first UI may be transparent to the second application. For example, the second application may be described as an application that is unaware that a touch input received in relation to the first UI is intended to execute a function of the first application.
[0073] For example, the second application can be used to support real-time translation (or live translation).
[0074] For example, the second application may be used to generate or obtain a first text in a first language by performing a speech-to-text (STT) of the first utterance in the first language. For example, the second application may be used to generate or obtain a second text in a second language by translating the first text in the first language. The second text may include content corresponding to content in the first text. For example, the second application may display the first text and the second text within the second UI. For example, the second application may transmit a voice signal for the second text (or a voice signal for the first text and the second text) to the other electronic device. For example, the voice signal may be output through a speaker associated with the other electronic device.
[0075] For example, the second application may be used to generate or obtain a third text in the second language by performing STT of the second utterance in the second language. For example, the second application may be used to generate or obtain a fourth text in the first language by translating the third text in the second language. The fourth text may include content corresponding to content in the third text. For example, the second application may display the third text and the fourth text within the second UI. For example, the second application may output a voice signal for the fourth text (or a voice signal for the third text and the fourth text) through a speaker associated with the electronic device.
[0076] For example, since the second application is executed in conjunction with the first application, the second UI may be displayed in relation to the first UI. The second UI displayed in relation to the first UI is described with reference to FIG. 5.
[0077] Figure 5 illustrates an example of a second UI displayed together with a first UI.
[0078] Referring to FIG. 5, a second UI (500) may be displayed together with the first UI for translation of the first utterance of the user of the electronic device (200) and translation of the second utterance of the user of the other electronic device, performed through the second application during the VoIP call performed through the first application. For example, the second UI (500) may include an object (501) including an executable element (501-1) for setting the first language of the first utterance of the user of the electronic device (200) and an executable element (501-2) for setting the second language of the second utterance of the user of the other electronic device. For example, the second UI (500) may include an object (502) for maintaining the VoIP call and terminating the translation of the first utterance of the user of the electronic device (200) and the translation of the second utterance of the user of the other electronic device performed through the second application. For example, the second UI (500) may include a translation window (503) for displaying the first text, the second text, the third text, and the fourth text.
[0079] For example, the second UI (500) may overlap or float on at least a portion of the first UI (400). For example, the second UI (500) may include a translucent area, such as a window (503). For example, the second UI (500) may include a transparent area. For example, the second UI (500) may include an area (511) that is transparent and does not provide objects and / or content. For example, the first object (401), the second object (402), the third object (403), and the fourth object (404) may be visible (or recognizable) through the area (511) of the second UI (500). However, the present invention is not limited thereto. For example, the second UI (500) may also be displayed on the area (510) of the display (240). For example, the size of the second UI (500) may be smaller than the size of the first UI (400).
[0080] As a non-limiting example, the first UI (400) and the second UI (500) may be displayed simultaneously. For example, at least one processor (210) may display the first UI (400) and the second UI (500) simultaneously on the display (240) according to a split view.
[0081] Referring back to FIG. 3B, at operation 330, at least one processor (210) may obtain a first image for the first UI based on a touch input received on the display (240) while displaying the second UI together with the first UI. For example, the first image for the first UI may be obtained while the touch input is maintained on the display (240). For example, the first image for the first UI may be obtained in response to the touch input. For example, the first image for the first UI may be obtained immediately (or after) the touch input is received. For example, the first image for the first UI may be obtained before the touch input is released. For example, the first image for the first UI may be obtained using a composition list used to display the first UI and the second UI. The above configuration list may include information indicating a stacking order of a layer including the first UI and a stacking order of a layer including the second UI, and information indicating a position of the layer including the first UI and a position of the layer including the second UI.
[0082] For example, the first image for the first UI may be described as an image obtained by capturing the entire first UI. As another example, the first image for the first UI may be described as an image obtained by capturing a portion of the first UI. The first image is described with reference to FIG. 6.
[0083] Figure 6 shows examples of the first image and the second image.
[0084] Referring to FIG. 6, the first image (600) can be obtained by capturing at least a portion of the first UI (400). For example, the first image (600) can be obtained in a state where the second application does not recognize which object among the objects (e.g., the first object (401), the second object (402), the third object (403), and the fourth object (404)) within the first UI (400) for which the touch input was received. For example, the second application can obtain the first image (600) in a state where it does not identify (or recognize) that the touch input in operation 330 was received in relation to the third object (403). The first image (600) may include a first visual object (601) corresponding to the first object (401), a second visual object (602) corresponding to the second object (402), a third visual object (603) corresponding to the third object (403), and a fourth visual object (604) corresponding to the fourth object (404). For example, since the first image (600) is acquired before the touch input received in operation 330 is released, the third visual object (603) in the first image (600) may be in a first state (610) indicating that the mute function of the first application is deactivated. For example, the fact that the third visual object (603) is in the first state (610) may not be recognized by the second application. For example, the fact that the third visual object (603) is in the first state (610) may not be known to the second application.
[0085] Although FIG. 6 illustrates a first image (600) related to the first UI (400) among the first UI (400) and the second UI (500), this is merely exemplary. The first image obtained in operation 330 may be related to both the first UI (400) and the second UI (500). For example, the first image may include the first UI (400) illustrated in FIG. 5 and the second UI (500) overlapping at least a portion of the first UI (400).
[0086] Referring back to FIG. 3B , at operation 340, at least one processor (210) may obtain a second image for the first UI based on the release of the touch input in operation 330. For example, the second image for the first UI may be described as an image obtained after obtaining the first image for the first UI. For example, the second image for the first UI may be described as an image obtained based on identifying a single tap input on the display (240) based on the reception of the touch input and the release of the touch input. The single tap input is merely exemplary. For example, various types of touch inputs on the display (240) may be applied in connection with operations 330 and 340. For example, the touch input and the release of the touch input may be a double tap input. For example, the touch input and the release of the touch input may be a long press input caused by maintaining the touch input for a reference time. For example, the touch input and the release of the touch input may be force tap inputs. In other words, the touch input and the release of the touch input may indicate the identification and release of inputs received while displaying the first UI and the second UI.
[0087] For example, the second image for the first UI may be described as an image obtained by capturing the entire first UI after (or immediately after) the touch input is released. As another example, the second image for the first UI may be described as an image obtained by capturing a portion of the first UI after (or immediately after) the touch input is released. The second image is described with reference to FIG. 6.
[0088] Referring to FIG. 6, the second image (650) can be obtained by capturing at least a portion of the first UI (400) after the touch input is released. For example, the second image (650) can be obtained in a state where the second application does not recognize which object among the objects in the first UI (400) (e.g., the first object (401), the second object (402), the third object (403), and the fourth object (404)) the touch input was received for. For example, the second application can obtain the second image (650) in a state where it does not identify (or recognize) that the release of the touch input in operation 340 is related to the third object (403). The second image (650) may include a first visual object (601) corresponding to the first object (401), a second visual object (602) corresponding to the second object (402), a third visual object (603) corresponding to the third object (403), and a fourth visual object (604) corresponding to the fourth object (404). For example, since the second image (650) is acquired after the touch input received in operation 330 is released, the third visual object (603) in the second image (650) may be in a second state (620) indicating that the mute function of the first application is activated. For example, the fact that the third visual object (603) is in the second state (620) may not be recognized by the second application. For example, the fact that the third visual object (603) is in the second state (620) may not be known to the second application.
[0089] Although Fig. 6 illustrates a second image (650) related to the first UI (400) among the first UI (400) and the second UI (500), this is merely exemplary. The second image obtained in operation 340 may be related to both the first UI (400) and the second UI (500). For example, the second image may include the first UI (400) illustrated in Fig. 5, and the second UI (500) overlapping at least a portion of the first UI (400). A visual object in the second image corresponding to a third object (403) (e.g., identical or similar to the third visual object (603)) may be in a second state (620).
[0090] Referring back to FIG. 3B, in operation 350, at least one processor (210) may identify (or determine) the mute function of the first application activated according to the touch input in operation 330 using the first image and the second image. The identification that the mute function of the first application is activated may be performed based on the first image and the second image. For example, at least one processor (210) may determine the activation of the mute function of the first application based on comparing the first image and the second image and identifying the second image as being at least partially different from the first image based on the comparison. For example, at least one processor (210) may determine the activation of the mute function of the first application based on a difference between the first image and the second image. The comparison between the first image and the second image is described with reference to FIG. 6.
[0091] Referring to FIG. 6, at least one processor (210) may compare the first image (600) and the second image (650) through the second application. For example, at least one processor (210) may identify, through the second application, that the state of the third visual object (603) in the first image (600) (e.g., the first state (610)) is different from the state of the third visual object (603) in the second image (650) (e.g., the second state (620)). For example, the second application may not recognize that the third visual object (603) in the first state (610) indicates that the mute function of the first application is disabled and that the third visual object (603) in the second state (620) indicates that the mute function of the first application is activated, but may identify the activation of the mute function of the first application based on the state of the third visual object (603) changing from the first state (610) to the second state (620).
[0092] Referring back to FIG. 3B , at operation 360, at least one processor (210) may stop displaying a result of the translation of the first utterance (e.g., the second text) within the second UI based on the identification at operation 350. The second text may be related to a translation of the first utterance. Stopping displaying the second text may be performed while the mute function is activated. Stopping displaying the second text may be implemented by controlling (the second application) not to perform a translation of the first utterance while the mute function is activated, based on the identification. Stopping displaying the second text may include not performing a translation of the first utterance, or performing a translation of the first utterance and not displaying a text related to the translation of the first utterance.
[0093] As a non-limiting example, at least one processor (210) may, based on the identification, stop displaying the result of the translation of the first utterance within the second UI by stopping the translation of the first utterance. As a non-limiting example, at least one processor (210) may, based on the identification, stop displaying the result of the translation of the first utterance within the second UI by stopping the STT of the first utterance. As a non-limiting example, at least one processor (210) may, based on the identification, stop obtaining the first text by performing the STT of the first utterance, maintaining obtaining the second text by translating the first text, and displaying the second text within the second UI (e.g., displaying the result of the translation of the first utterance within the second UI). Stopping displaying the result of the translation of the first utterance within the second UI is described with reference to FIG. 7.
[0094] Figure 7 illustrates an example of stopping displaying the results of translation of a user's speech of an electronic device.
[0095] Referring to FIG. 7, at least one processor (210) may display, on the display (240), a first UI (400) including a third object (403) in a state (710) through the first application, according to the touch input in operation 330 and the release of the touch input in operation 340. For example, at least one processor (210) may stop displaying the result of the translation of the first utterance in the second UI (500) (or the translation window (503) in the second UI (500)) according to the identification in operation 350 performed through the second application. For example, after the identification in operation 350 performed through the second application, utterances (or audio) represented by the state (760) may be caused around the electronic device (200). For example, at least one processor (210) may, independently of receiving the utterances (or the audio) via the microphone (250), refrain from displaying the results of the translation of the utterances (or the audio) within a translation window (503) within the second UI (500), as indicated by the arrow (770).
[0096] Referring again to FIG. 3B , in accordance with operation 360, as a non-limiting example, at least one processor (210) may, based on the identification in operation 350, instead of ceasing to display the result of the translation of the first utterance (e.g., the second text) within the second UI, continue to display the result of the translation of the first utterance (e.g., the second text) within the second UI, and refrain from or block transmitting information about the result of the translation of the first utterance to the other electronic device. In a non-limiting example, at least one processor (210) may additionally display, on the display (240), an indication indicating that the information is not to be transmitted to the other electronic device when continuing to display the result of the translation of the first utterance within the second UI.
[0097] As described above, the electronic device (200) may be unable to identify at least a part of the execution state of the first application, but may be able to inform the second application of at least a part of the execution state of the first application by acquiring images (e.g., the first image and the second image) before and after a touch input for the second application that is executed in conjunction with the first application. For example, the electronic device (200) may enhance the service quality of the second application related to the first application by acquiring images (e.g., the first image and the second image) before and after a touch input.
[0098] FIG. 8 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application using an area on a display that includes a location where a touch input is received.
[0099] Referring to FIG. 8, in operation 801, at least one processor (210) may receive a touch input on the display (240) while displaying the second UI together with the first UI. The touch input may correspond to the touch input in operation 330.
[0100] In operation 802, at least one processor (210) may capture the first UI (or a screen on a display (240) including the first UI and the second UI) based on the touch input in operation 801. The captured first UI is described with reference to FIG. 9.
[0101] FIG. 9 illustrates an example of a portion of a first image and a portion of a second image positioned within an area on the display including a location where a touch input was received.
[0102] Referring to FIG. 9, at least one processor (210) may obtain an image (900) by capturing the first UI (or the screen including the first UI and the second UI). For example, at least a portion of the information expressed by the image (900) (e.g., the first UI or the first UI and the second UI) may be unknown or unnoticeable to the second application. For example, the image (900) may be obtained while the touch input is maintained on the display (240). For example, the image (900) may be obtained before feedback (e.g., a change in the state of the third object (403)) according to the touch input is provided. For example, the image (900) may be obtained immediately (or immediately before) the touch input.
[0103] Referring again to FIG. 8, at operation 803, at least one processor (210) may identify a release of the touch input received (or identified) at operation 801. The release of the touch input at operation 803 may correspond to the release of the touch input at operation 340.
[0104] In operation 804, at least one processor (210) may capture the first UI (or the screen including the first UI and the second UI) based on the release of the touch input in operation 803. The first UI captured based on the release of the touch input is described with reference to FIG. 9.
[0105] Referring to FIG. 9, at least one processor (210) may obtain an image (950) by capturing the first UI (or the screen including the first UI and the second UI). For example, at least a portion of the information expressed by the image (950) (e.g., the first UI or the first UI and the second UI) may not be known to or noticed by the second application. For example, the image (950) may be obtained after feedback (e.g., a change in the state of the third object (403)) according to the touch input is provided.
[0106] Referring again to FIG. 8, at operation 805, at least one processor (210) may identify an area on the display (240) that includes a location where the touch input was received. Although FIG. 8 illustrates an example in which operation 805 is performed after operation 803, this is merely exemplary. Operation 805 may also be performed before operation 803 is performed, based on operation 801. The area on the display (240) identified at operation 805 is described with reference to FIG. 9.
[0107] Referring to FIG. 9, at least one processor (210) can identify a location (910) (or area (910)) of the touch input received in operation 801. For example, at least one processor (210) can identify an area (920) on the display (240) that includes the location (910). The area (920) can include at least a portion of the location (910). The area (920) can be identified based on the location (910), which is location data of the touch input obtained from a touch IC (integrated circuitry) included in the display (240).
[0108] Referring back to FIG. 8, in operation 806, at least one processor (210) may obtain the first image by cropping the first UI (or the screen including the first UI and the second UI) captured based on the touch input using the area identified in operation 805. The first image obtained according to operation 806 is described with reference to FIG. 10.
[0109] Figure 10 shows an example of comparing a first image and a second image.
[0110] Referring to FIG. 10, at least one processor (210) may crop an image (900) using an area (920), as indicated by arrow (1010), according to operation 806. For example, at least one processor (210) may obtain a first image (1000) based on the cropping of the image (900) performed using the area (920). For example, information represented by the first image (1000) (or information within the first image (1000)) may be unknown or unnoticed by the second application.
[0111] Referring back to FIG. 8, in operation 807, at least one processor (210) may obtain the second image by cropping the first UI (or the screen including the first UI and the second UI) captured based on the release of the touch input using the area identified in operation 805. The second image obtained according to operation 807 is described with reference to FIG. 10.
[0112] Referring to FIG. 10, at least one processor (210) may crop an image (950) using an area (920), as indicated by arrow (1060), according to operation 807. For example, at least one processor (210) may obtain a second image (1050) based on the cropping of the image (950) performed using the area (920). For example, information represented by the second image (1050) (or information within the second image (1050)) may be unknown or unnoticed by the second application.
[0113] Referring back to FIG. 8, at operation 808, at least one processor (210) may compare the first image acquired at operation 806 with the second image acquired at operation 807. As a non-limiting example, at least one processor (210) may compare a color of the first image with a color of the second image. As a non-limiting example, at least one processor (210) may compare a shape of a visual object included in the first image with a shape of a visual object included in the second image. As a non-limiting example, at least one processor (210) may compare feature points of the first image with feature points of the second image. The comparison between the first image and the second image is described with reference to FIG. 10.
[0114] Referring to FIG. 10, at least one processor (210) may, through the second application, compare a first image (1000) including information not recognized by the second application with a second image (1050) including information not recognized by the second application, as indicated by an arrow (1090). For example, the at least one processor (210) may, based on the comparison between the first image (1000) and the second image (1050), identify the second image (1050) as being at least partially different from the first image (1000), such as in a state (1095).
[0115] Referring again to FIG. 8, at operation 809, at least one processor (210) may identify whether the second image is at least partially different from the first image based on the comparison at operation 808. For example, at least one processor (210) may perform operation 810 based on the second image being at least partially different from the first image, and may perform operation 811 based on the second image being identical to the first image.
[0116] In operation 810, at least one processor (210) may identify the mute function of the first application activated in response to the touch input under conditions in which the second image is at least partially different from the first image. Operation 810, performed under conditions in which the second image is at least partially different from the first image, may further include operations performed in relation to a trained model (e.g., an artificial intelligence model) within the electronic device (200). These operations will be described with reference to FIGS. 13 and 14.
[0117] In operation 811, at least one processor (210) can identify that the mute function of the first application is not activated according to the touch input under the condition that the second image corresponds to the first image. For example, at least one processor (210) can identify that the mute function of the first application is maintained independently of the touch input based on the second image that is identical to the first image. For example, at least one processor (210) can identify that the touch input is not related to the mute function of the first application based on the second image that is identical to the first image.
[0118] As described above, the electronic device (200) can determine whether the mute function of the first application is activated for the second application by using screen captures performed before and after a touch input and the position data of the touch input. Through this operation, the electronic device (200) can enhance the service quality of the second application.
[0119] As a non-limiting example, the arrangement of objects of the first UI may vary depending on the service provider, but the objects of the first UI may be primarily adjacent to a bottom area on the display (240). For example, at least one processor (210) may compare the area on the display (240) identified in operation 805 with a reference area on the display (240) corresponding to the bottom area on the display (240) where the objects of the first UI are primarily (or typically) arranged. This comparison is described with reference to FIG. 11.
[0120] FIG. 11 is a flowchart illustrating an exemplary method for identifying whether an area on a display including a location where a touch input is received is included within a reference area on the display.
[0121] Referring to FIG. 11, in operation 1101, at least one processor (210) may identify an area on the display (240) that includes a location where the touch input was received in operation 801. Operation 1101 may correspond to operation 805.
[0122] In operation 1102, at least one processor (210) may identify whether the area identified in operation 1101 is included within a reference area on the display (240). The reference area may be an area estimated to include objects within the first UI. For example, the reference area may correspond to a lower area on the display (240). However, the present invention is not limited thereto. For example, at least one processor (210) may perform operation 1103 based on identifying the area included within the reference area, and may perform operation 1104 based on identifying the area including a portion not included within the reference area. Identifying whether the area identified in operation 1101 is included within the reference area on the display (240) is described with reference to FIG. 12.
[0123] Figure 12 illustrates an exemplary method of using a reference area on a display.
[0124] Referring to FIG. 12, the touch input in operation 801 may be received at a location (910). For example, at least one processor (210) may identify an area (920) of the display (240) including the location (910) according to operation 1101. For example, when the area (920) is included within the reference area (1200), such as in state (1230), at least one processor (210) may perform operation 1103.
[0125] The touch input in operation 801 may be received at a location (1210). For example, at least one processor (210) may identify an area (1215) of the display (240) including the location (1210) according to operation 1101. For example, when a part of the area (1215) is not included in the reference area (1200), such as in state (1260), at least one processor (210) may perform operation 1104. However, the present invention is not limited thereto. For example, when a part of the area (1215) is not included in the reference area (1200), such as in state (1260), at least one processor (210) may also perform operation 1103.
[0126] The touch input in operation 801 may be received at a location (1220). For example, at least one processor (210) may identify an area (1225) of the display (240) including the location (1220) according to operation 1101. For example, when the area (1225) is not (completely) included within the reference area (1200), such as in state (1260), at least one processor (210) may perform operation 1104.
[0127] Referring back to FIG. 11, in operation 1103, at least one processor (210) may acquire a first image and a second image based on the area included in the reference area. For example, operation 1103 may correspond to operations 806 and 807 of FIG. 8. For example, at least one processor (210) may identify, by comparing the first image and the second image, that the mute function of the first application is activated according to the touch input in operation 801 or that the mute function of the first application is not activated according to the touch input in operation 801.
[0128] At operation 1104, at least one processor (210) can identify that the mute function of the first application is not activated according to the touch input at operation 801, based on the area not included in the reference area.
[0129] For example, the electronic device (200) can strengthen the conditions for performing operations 806 and 807 by utilizing the reference area. For example, the electronic device (200) can reduce the amount of computation (or power consumption) required to identify whether the mute function of the first application is activated through the second application by strengthening these conditions.
[0130] For example, even if the second application identifies the second image as being partially different from the first image based on the comparison between the first image and the second image in operation 809 of FIG. 8, it may not recognize that a visual object corresponding to an object for the mute function (e.g., the third object (403)) is included in the first image and that a visual object corresponding to the object is included in the second image. For example, a trained model within the electronic device (200) may be utilized to recognize that the first image includes a visual object related to the mute function and that the second image includes a visual object related to the mute function. Operations related to the trained model are described with reference to FIG. 13.
[0131] FIG. 13 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application using a trained model in connection with identifying a location where a touch input was received.
[0132] Referring to FIG. 13, in operation 1301, at least one processor (210) may identify a second image that is at least partially different from the first image. For example, operation 1301 may be performed based on the comparison in operation 809.
[0133] In operation 1302, at least one processor (210) may provide first data related to the first image and second data related to the second image to a trained model within the electronic device (200). The trained model may be an artificial intelligence model usable for image analysis. The trained model may be usable for object recognition, for example.
[0134] As a non-limiting example, at least one processor (210) may generate or obtain the first data, which is input data to the trained model, by processing the first image. For example, the at least one processor (210) may generate the first data by resizing the first image. For example, the size of the resized first image may correspond to the size of the learning data used for training (or learning) the model.
[0135] As a non-limiting example, at least one processor (210) may generate or obtain the second data, which is another input data to the trained model, by processing the second image. For example, the at least one processor (210) may generate the second data by resizing the second image. For example, the size of the resized second image may correspond to the size of the training data.
[0136] In operation 1303, at least one processor (210) may obtain information related to the first data and the second data from the model. The information may indicate whether each of the first data (or the first image) and the second data (or the second image) includes an object (e.g., a microphone-shaped object) indicating a mute function. The information is described with reference to FIG. 14.
[0137] Figure 14 illustrates a method of using a training model.
[0138] Referring to FIG. 14, at least one processor (210) may input first data (1410) and second data (1420) to the model (1400). For example, content within each of the first data (1410) and the second data (1420) may not be recognized by the first application.
[0139] The model (1400) can output information (1430) in response to the first data (1410) and the second data (1420). For example, the information (1430) output from the model (1400) can indicate that the first data (1410) includes an object (1415) representing a mute function (or an object (1415) representing that the mute function is disabled), such as a state (1440), and that the second data (1420) includes an object (1425) representing a mute function (or an object (1425) representing that the mute function is enabled). For example, the information (1430) output from the model (1400) can indicate that the first data (1410) and the second data (1420) are not related to the mute function, such as a state (1450). For example, information (1430) may indicate that the first data (1410) includes an object (1417) that does not express a mute function (or an object (1417) that expresses that the speaker phone function is disabled), such as a state (1450), and that the second data (1420) includes an object (1427) that does not express a mute function (or an object (1427) that expresses that the speaker phone function is enabled). For example, information (1430) output from the model (1400) may include information indicating activation or deactivation of the mute function using the first data (1410) and the second data (1420).
[0140] Referring again to FIG. 13, at operation 1304, at least one processor (210) may identify, based on the information, whether each of the first data and the second data includes an object for a mute function. Alternatively, at least one processor (210) may identify, based on the information, whether each of the first data and the second data is associated with an object for a mute function.
[0141] For example, at least one processor (210) may perform operation 1305 based on the first data and the second data, each of which includes an object for a mute function, and otherwise perform operation 1306.
[0142] In operation 1305, at least one processor (210) may identify the mute function of the first application activated according to the touch input, based on a condition that an object for the mute function is included in each of the first data and the second data. For example, operation 1305 may correspond to operation 810.
[0143] In operation 1306, at least one processor (210) may identify that the mute function of the first application is not activated according to the touch input, provided that an object for the mute function is not included in each of the first data and / or the second data. For example, operation 1306 may correspond to operation 811.
[0144] FIG. 15 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application through comparison between a first image and a second image.
[0145] Referring to FIG. 15, in operation 1501, at least one processor (210) may obtain a first image by capturing the first UI (or a screen including the first UI and the second UI) based on the touch input in operation 330.
[0146] In operation 1502, at least one processor (210) may obtain a second image by capturing the first UI (or the screen including the first UI and the second UI) based on the release of the touch input in operation 340.
[0147] The above first image and the above second image are described with reference to FIG. 16.
[0148] Figure 16 illustrates an example of a portion of a first image that is different from a second image and a portion of a second image that is different from the first image.
[0149] Referring to FIG. 16, at least one processor (210) may obtain a first image (1600) by capturing the first UI (or the screen including the first UI and the second UI) based on the touch input. For example, at least a portion of the information expressed by the first image (1600) (e.g., the first UI or the first UI and the second UI) may not be known or noticed by the second application. For example, the first image (1600) may be obtained while the touch input is maintained on the display (240). For example, the first image (1600) may be obtained before feedback (e.g., a change in the state of the third object (403)) according to the touch input is provided. For example, the first image (1600) may be obtained immediately (or immediately before) the touch input.
[0150] At least one processor (210) may capture the first UI (or the screen including the first UI and the second UI) based on the release of the touch input to obtain a second image (1650). For example, at least a portion of the information expressed by the second image (1650) (e.g., the first UI or the first UI and the second UI) may not be known or noticed by the second application. For example, the second image (1650) may be obtained after feedback (e.g., a change in the state of the third object (403)) according to the touch input is provided.
[0151] Referring again to FIG. 15, at operation 1503, at least one processor (210) may compare the first image acquired in operation 1501 with the second image acquired in operation 1502.
[0152] In operation 1504, at least one processor (210) may identify, based on the comparison, whether the second image is at least partially different from the first image. For example, the at least one processor (210) may perform operation 1505 based on the second image being at least partially different from the first image, and may perform operation 1508 based on the second image being identical to the first image.
[0153] In operation 1505, at least one processor (210) can identify a portion of the first image that is different from the second image and a portion of the second image that is different from the first image, provided that the second image is at least partially different from the first image. This operation is described with reference to FIG. 16.
[0154] Referring to FIG. 16, at least one processor (210) may identify a first region (1605) of the first image (1600) and a second region (1615) of the first image (1600) that is a portion of the first image (1600) that is not common to the second image (1650) based on the comparison (e.g., operation 1503) between the first image (1600) and the second image (1650). For example, at least one processor (210) may identify a first region (1655) of the second image (1650) and a second region (1665) of the second image (1650) that is a portion of the second image (1650) that is not common to the first image (1610) based on the comparison (e.g., operation 1503) between the first image (1600) and the second image (1650). For example, although the content in the first image (1600) and the content in the second image (1650) are not recognized by the second application, at least one processor (210) can identify, through the second application, a first area (1605) of the first image (1600) and a second area (1615) of the first image (1600) corresponding to the difference between the first image (1600) and the second image (1650). For example, although the content in the first image (1600) and the content in the second image (1650) are not recognized by the second application, at least one processor (210) can identify, through the second application, a first area (1655) of the second image (1650) and a second area (1665) of the second image (1650) corresponding to the difference between the first image (1600) and the second image (1650).
[0155] Referring again to FIG. 15, at operation 1506, at least one processor (210) may identify whether a difference between the portion of the first image identified in operation 1505 and the portion of the second image satisfies a condition.
[0156] For example, the condition may be whether the color data of the part of the second image is outside the reference range with respect to the color data of the part of the first image. For example, at least one processor (210) may determine that the condition is satisfied based on the color data of the part of the second image that is outside the reference range with respect to the color data of the part of the first image. For example, at least one processor (210) may determine that the condition is not satisfied based on the color data of the part of the second image that is within the reference range with respect to the color data of the part of the first image. For example, as represented by the visual object (603) in the first state (610) of the first image (600) of FIG. 6 and the visual object (603) in the second state (620) of the second image (650) of FIG. 6, the color of the object indicating that the mute function is disabled may be substantially opposite to the color of the object indicating that the mute function is enabled. As a non-limiting example, the reference data may be about 90 (%)(percent).
[0157] For example, the condition may be whether the shape data (or structural data) of the part of the second image is within a threshold range with respect to the shape data (or structural data) of the part of the first image. For example, at least one processor (210) may determine that the condition is satisfied based on the shape data of the part of the second image that is outside the reference range with respect to the shape data of the part of the first image. For example, at least one processor (210) may determine that the condition is not satisfied based on the shape data of the part of the second image that is within the reference range with respect to the shape data of the part of the first image. For example, as represented by the visual object (603) in the first state (610) of the first image (600) of FIG. 6 and the visual object (603) in the second state (620) of the second image (650) of FIG. 6, the shape of the object indicating that the mute function is disabled may substantially correspond to the shape of the object indicating that the mute function is enabled. As a non-limiting example, the threshold data may be about 10 (%).
[0158] For example, at least one processor (210) can identify that a difference between a first region (1605) of a first image (1600) (e.g., including an object for a mute function) and a first region (1655) of a second image (1650) (e.g., including an object for a mute function) satisfies the above condition. For example, at least one processor (210) can identify that a difference between a second region (1615) of a first image (1600) (e.g., including a translation result in a translation window) and a second region (1665) of a second image (1650) (e.g., including a translation result in a translation window) do not satisfy the above condition.
[0159] Referring again to FIG. 15, at operation 1507, at least one processor (210) may identify the mute function of the first application activated according to the touch input when the above condition is satisfied.
[0160] In operation 1508, at least one processor (210) can identify that the mute function of the first application is not activated according to the touch input when the second image is identical to the first image or does not satisfy the condition.
[0161] For example, the second application may not recognize that a visual object corresponding to an object for the mute function (e.g., a third object (403)) is included in the portion of the first image and that a visual object corresponding to the object is included in the portion of the second image, even if the difference between the portion of the first image and the portion of the second image in operation 1506 satisfies the condition. For example, a trained model within the electronic device (200) may be used to recognize that the portion of the first image includes a visual object associated with the mute function and that the portion of the second image includes a visual object associated with the mute function. Operations related to the trained model are described with reference to FIG. 17.
[0162] FIG. 17 is a flowchart illustrating an exemplary method for identifying activation of a mute function of a first application using a trained model in relation to a comparison between a first image and a second image.
[0163] Referring to FIG. 17, in operation 1701, at least one processor (210) may identify that a difference between the portion of the first image and the portion of the second image satisfies a condition. For example, operation 1701 may correspond to operation 1506.
[0164] In operation 1702, at least one processor (210) may provide first data related to the portion of the first image and second data related to the portion of the second image to a trained model within the electronic device (200) based on the identification that the condition is satisfied. The trained model may be an artificial intelligence model usable for image analysis. The trained model may be usable for object recognition, for example.
[0165] As a non-limiting example, at least one processor (210) may generate or obtain the first data, which is input data to the trained model, by processing the portion of the first image. For example, the at least one processor (210) may generate the first data by resizing the portion of the first image. For example, the size of the portion of the resized first image may correspond to the size of learning data used for training (or learning) the model.
[0166] As a non-limiting example, at least one processor (210) may generate or obtain the second data, which is another input data to the trained model, by processing the portion of the second image. For example, the at least one processor (210) may generate the second data by resizing the portion of the second image. For example, the size of the portion of the resized second image may correspond to the size of the training data.
[0167] In operation 1703, at least one processor (210) may obtain information related to the first data and the second data from the model. The information may indicate whether each of the first data (or the portion of the first image) and the second data (or the portion of the second image) includes an object (e.g., a microphone-shaped object) indicating a mute function. For example, the information may include information indicating whether the mute function is activated or deactivated using the first data and the second data.
[0168] In operation 1704, at least one processor (210) may identify, based on the information, whether each of the first data and the second data includes an object for a mute function. Alternatively, at least one processor (210) may identify, based on the information, whether each of the first data and the second data is related to an object for a mute function.
[0169] For example, at least one processor (210) may perform operation 1705 based on the first data and the second data, each including an object for a mute function, and otherwise perform operation 1706.
[0170] In operation 1705, at least one processor (210) may identify the mute function of the first application activated according to the touch input, based on a condition that an object for the mute function is included in each of the first data and the second data. For example, operation 1705 may correspond to operation 1507.
[0171] In operation 1706, at least one processor (210) may identify that the mute function of the first application is not activated according to the touch input, provided that an object for the mute function is not included in each of the first data and / or the second data. For example, operation 1706 may correspond to operation 1508.
[0172] FIG. 18 is a flowchart illustrating an exemplary method for identifying the deactivation of the mute function of a first application for a VoIP call.
[0173] Referring to FIG. 18, in operation 1801, at least one processor (210) may stop displaying the result of the translation of the first utterance within the second UI displayed together with the first UI. For example, operation 1801 may correspond to operation 360.
[0174] In operation 1802, at least one processor (210) may obtain a third image for the first UI based on another touch input received while the result of the translation of the first utterance is stopped from being displayed within the second UI displayed together with the first UI.
[0175] In operation 1803, at least one processor (210) may obtain a fourth image for the first UI based on the release of the other touch input in operation 1802.
[0176] At operation 1804, at least one processor (210) may identify the mute function of the first application that is disabled based on the other touch input using the third image and the fourth image. For example, operation 1804 may be performed based on a comparison between the third image and the fourth image.
[0177] At operation 1805, at least one processor (210) may resume displaying the result of the translation of the first utterance within the second UI based on the identification that the mute function of the first application is disabled in response to the other touch input.
[0178] For example, at least some of the operations described with reference to FIG. 18 may include operations described with reference to FIGS. 8 to 17.
[0179] For example, the first application for the VoIP call may provide a video call function performed through a camera (not shown) of the electronic device (200) or a camera (e.g., a camera of an external electronic device) connected to the electronic device (200) by wire or wirelessly. As a non-limiting example, providing services (e.g., STT and translation services) of the second application while providing the video call function may cause an overload of the electronic device (200). As a non-limiting example, simultaneously displaying an image acquired while providing the video call function and the translation window of the second UI on the display (240) may reduce the visual quality of the display (240). For example, the electronic device (200) may stop displaying the result of the translation of the first utterance based on the activation of the video call function of the first application. Based on the activation of the video call function of the first application, the displaying of the result of the translation of the first utterance is stopped as described with reference to FIGS. 19A and 19B . In one example, at least one processor (210) may stop displaying the result of the translation of the first utterance based on the activation of the camera function. For example, when the camera function is activated, at least one processor (210) may stop the translation function, and based on the termination of the camera function, may re-enable the translation function. For example, at least one processor (210) may display an object capable of performing the translation function on the UI based on the termination of the camera function.
[0180] FIG. 19a is a flowchart illustrating an exemplary method for identifying activation of a video call function of a first application for a VoIP call.
[0181] Referring to FIG. 19A, in operation 1901, at least one processor (210) may display, on a display (240), a first UI (user interface) for a VoIP (voice over internet protocol) call with another electronic device performed through the first application using a communication circuit (230). The first UI may include an object requesting activation (or driving) of a camera of the electronic device (200). For example, referring to FIG. 4, the first UI (400) may include a second object (402) for a video call function of the first application.
[0182] Referring back to FIG. 19A, at operation 1902, at least one processor (210) may display, using the display (240), a second UI for the translation of the first utterance of the user of the electronic device (200) and the translation of the second utterance of the user of the other electronic device, performed via the second application during the VoIP call, together with the first UI. For example, referring to FIG. 5, the second UI (500) may be superimposed on the first UI (400). For example, the second UI (500) may be superimposed on the first UI (400) that includes a second object (402) that triggers operation of a camera of the electronic device (200).
[0183] Referring back to FIG. 19A, at operation 1903, at least one processor (210) may obtain a first image for the first UI based on a received touch input on the display (240) while displaying the second UI together with the first UI. For example, referring to FIG. 6, at least one processor (210) may obtain a first image (600). For example, the first image (600) may be obtained before a function of the first application (e.g., the video call function) is executed in response to the touch input. For example, the shape of a visual object (602) in the first image (600) may correspond to the shape of a second object (402) in the first UI (400) provided while the video call function is not executed.
[0184] Referring back to FIG. 19A, at operation 1904, at least one processor (210) may obtain a second image for the first UI based on the release of the touch input. For example, at least one processor (210) may at least partially change the state of the first UI based on the touch input received in operation 1903. The change in the state of the first UI is described with reference to FIG. 19B.
[0185] Figure 19b illustrates an example of a first UI having a changed state according to the release of a touch input.
[0186] Referring to FIG. 19B, as in a state (1990), at least one processor (210) may, based on the release of the touch input, display a second object (402) within the first UI (400) that has a different state than the state of the second object (402) illustrated in FIG. 5 . For example, the shape of the second object (402) within the first UI (400) illustrated in FIG. 19B may be at least partially different from the shape of the second object (402) illustrated in FIG. 5 . For example, the second object (402) within the first UI (400) illustrated in FIG. 19B may have a state indicating that the video call function can be deactivated (or a state indicating that the video call function is activated), unlike the second object (402) illustrated in FIG. 5 , which has a state indicating that the video call function can be activated (or a state indicating that the video call function is deactivated). For example, at least one processor (210) can obtain the second image including the first UI (400) (or the first UI (400) and the second UI (500)) within the state (1990).
[0187] Referring back to FIG. 19A, at operation 1905, at least one processor (210) may identify a video call function of the first application activated in response to the touch input using the first image and the second image. For example, the at least one processor (210) may identify a difference between the first image and the second image (e.g., the second object (402) in FIG. 19B having a different state from the state of the second object (402) in FIG. 5) by performing a comparison between the first image and the second image, and may identify the video call function of the first application based on the identification.
[0188] At operation 1906, at least one processor (210) may stop displaying the second UI based on the identification.
[0189] For example, referring to FIG. 19B, at least one processor (210) may change the state (1990) to the state (1991) based on the identification of the video call function. Within the state (1991), at least one processor (210) may stop displaying the second UI (500) and display the UI (1995) of the first application according to the video call function. For example, the UI (1995) may include a preview image (1992) acquired through a camera of the electronic device (200). For example, the UI (1995) may include an object (1994) for deactivating the video call function and activating the voice call function.
[0190] As a non-limiting example, at least one processor (210) may display a message (1993) indicating that the translation service of the second application has been suspended by ceasing to display the second UI (500) along with the UI (1995). For example, the message (1993) may be superimposed on the UI (1995). For example, the message (1993) may disappear from the display (240) over the elapse of a reference time.
[0191] The operations of the electronic device (200) described above may also be performed by the electronic device (2001) of FIG. 20.
[0192] FIG. 20 is a block diagram of an electronic device (2001) within a network environment (2000) according to various embodiments. Referring to FIG. 20, in the network environment (2000), the electronic device (2001) may communicate with the electronic device (2002) via a first network (2098) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (2004) or the server (2008) via a second network (2099) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (2001) may communicate with the electronic device (2004) via the server (2008). According to one embodiment, the electronic device (2001) may include a processor (2020), a memory (2030), an input module (2050), an audio output module (2055), a display module (2060), an audio module (2070), a sensor module (2076), an interface (2077), a connection terminal (2078), a haptic module (2079), a camera module (2080), a power management module (2088), a battery (2089), a communication module (2090), a subscriber identification module (2096), or an antenna module (2097). In some embodiments, the electronic device (2001) may omit at least one of these components (e.g., the connection terminal (2078)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (2076), camera module (2080), or antenna module (2097)) may be integrated into a single component (e.g., display module (2060)).
[0193] The processor (2020) may, for example, execute software (e.g., a program (2040)) to control at least one other component (e.g., a hardware or software component) of the electronic device (2001) connected to the processor (2020) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (2020) may store commands or data received from other components (e.g., a sensor module (2076) or a communication module (2090)) in the volatile memory (2032), process the commands or data stored in the volatile memory (2032), and store the resulting data in the non-volatile memory (2034). According to one embodiment, the processor (2020) may include a main processor (2021) (e.g., a central processing unit or an application processor) or a secondary processor (2023) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (2021). For example, when the electronic device (2001) includes the main processor (2021) and the secondary processor (2023), the secondary processor (2023) may be configured to use less power than the main processor (2021) or to be specialized for a given function. The secondary processor (2023) may be implemented separately from the main processor (2021) or as a part thereof.
[0194] The auxiliary processor (2023) may control at least a portion of functions or states associated with at least one component (e.g., the display module (2060), the sensor module (2076), or the communication module (2090)) of the electronic device (2001), for example, on behalf of the main processor (2021) while the main processor (2021) is in an inactive (e.g., sleep) state, or together with the main processor (2021) while the main processor (2021) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (2023) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (2080) or a communication module (2090)). In one embodiment, the auxiliary processor (2023) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (2001) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (2008)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0195] The memory (2030) can store various data used by at least one component (e.g., the processor (2020) or the sensor module (2076)) of the electronic device (2001). The data can include, for example, software (e.g., the program (2040)) and input data or output data for commands related thereto. The memory (2030) can include volatile memory (2032) or non-volatile memory (2034).
[0196] The program (2040) may be stored as software in memory (2030) and may include, for example, an operating system (2042), middleware (2044), or an application (2046).
[0197] The input module (2050) can receive commands or data to be used in a component of the electronic device (2001) (e.g., a processor (2020)) from an external source (e.g., a user) of the electronic device (2001). The input module (2050) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0198] The audio output module (2055) can output audio signals to the outside of the electronic device (2001). The audio output module (2055) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0199] The display module (2060) can visually provide information to an external party (e.g., a user) of the electronic device (2001). The display module (2060) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (2060) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0200] The audio module (2070) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (2070) can acquire sound through the input module (2050), output sound through the sound output module (2055), or an external electronic device (e.g., electronic device (2002)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (2001).
[0201] The sensor module (2076) can detect the operating status (e.g., power or temperature) of the electronic device (2001) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (2076) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0202] The interface (2077) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (2001) to an external electronic device (e.g., the electronic device (2002)). In one embodiment, the interface (2077) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0203] The connection terminal (2078) may include a connector through which the electronic device (2001) may be physically connected to an external electronic device (e.g., the electronic device (2002)). In one embodiment, the connection terminal (2078) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0204] The haptic module (2079) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (2079) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0205] The camera module (2080) can capture still images and videos. In one embodiment, the camera module (2080) may include one or more lenses, image sensors, image signal processors, or flashes.
[0206] The power management module (2088) can manage power supplied to the electronic device (2001). According to one embodiment, the power management module (2088) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0207] A battery (2089) may power at least one component of the electronic device (2001). In one embodiment, the battery (2089) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0208] The communication module (2090) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (2001) and an external electronic device (e.g., electronic device (2002), electronic device (2004), or server (2008)), and the performance of communication through the established communication channel. The communication module (2090) may operate independently from the processor (2020) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (2090) may include a wireless communication module (2092) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (2094) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (2004) via a first network (2098) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (2099) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (2092) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (2096) to identify or authenticate the electronic device (2001) within a communication network such as the first network (2098) or the second network (2099).
[0209] The wireless communication module (2092) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (2092) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (2092) can support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (2092) can support various requirements specified in the electronic device (2001), an external electronic device (e.g., the electronic device (2004)), or a network system (e.g., the second network (2099)). According to one embodiment, the wireless communication module (2092) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0210] The antenna module (2097) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (2097) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (2097) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (2098) or the second network (2099), may be selected from the plurality of antennas, for example, by the communication module (2090). A signal or power may be transmitted or received between the communication module (2090) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (2097).
[0211] According to various embodiments, the antenna module (2097) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0212] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0213] According to one embodiment, commands or data may be transmitted or received between the electronic device (2001) and an external electronic device (2004) via a server (2008) connected to a second network (2099). Each of the external electronic devices (2002 or 2004) may be the same or a different type of device as the electronic device (2001). According to one embodiment, all or part of the operations executed in the electronic device (2001) may be executed in one or more of the external electronic devices (2002, 2004, or 2008). For example, when the electronic device (2001) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (2001) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (2001). The electronic device (2001) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (2001) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (2004) may include an Internet of Things (IoT) device. The server (2008) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (2004) or server (2008) may be included within the second network (2099). The electronic device (2001) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0214] Some of the operations described above may be executed (or performed) by an AI (artificial intelligence) system as described with reference to FIG. 21.
[0215] Figure 21 is a schematic diagram of an exemplary AI system.
[0216] Referring to FIG. 21, the AI system (2100) may include an input / output interface (2110), an AI (artificial intelligence) framework (2120), a generative AI model (2130), an application / service component (2180), and / or a knowledge repository (2190).
[0217] The input / output interface (2110) can receive input. The input can include user input and / or data acquired or generated by an electronic device (e.g., the electronic device (200) or the electronic device (2001) described above). The data can include images, videos, and / or sensor data generated by at least one processor (e.g., at least one processor (210) or processor (2020)) of the electronic device (e.g., illuminance data around the electronic device acquired from a sensor or sensor hub (e.g., a coprocessor (2023), posture data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., temperature of the display (240) or temperature of the at least one processor (210)), size information of a display area of the display (240), and / or images acquired through an image sensor (e.g., included in a camera module (2080)) of the electronic device). The user input may include natural language, touch data obtained via touch circuitry included within the display (240) (e.g., used to identify input from a finger and / or a stylus), images displayed (and / or to be displayed) on the display (240), and / or video. As a non-limiting example, the user input may be received by the input / output interface (2110) together with context information. The context information may be described as additional information obtained in connection with the user input. The context information may relate to a state when the user input is received (e.g., including a state of the electronic device and / or a state surrounding the electronic device (e.g., a user state)). For example, the context information may include information about one or more software applications running within the electronic device when the user input is received.For example, the contextual information may include information about the location of the electronic device (or the location of the user of the electronic device) at the time the user input is received. For example, the user input may be integrated with the contextual information. For example, the user input integrated with the contextual information may be received by the input / output interface (2110).
[0218] The input / output interface (2110) can transmit (or provide) output. The output may include a result (or result information) generated or acquired by the AI system (2100) based at least in part on the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., including media content and / or multimedia content). For example, the output may include an action related to a user of the electronic device. For example, the output may have a format according to a user setting of the electronic device.
[0219] The input / output interface (2110) can be described as a user query / response interface (2110).
[0220] The AI framework (2120) can be used to obtain information (or data) about the input from the input / output interface (2110) and control one or more components related to the AI system (2100) using the obtained information.
[0221] For example, the prompt design component (2121) within the AI framework (2120) can use the acquired information to generate or obtain a prompt for a generative AI model (2130) (e.g., including a large language model (LLM) or a large multimodal model (LMM)). For example, the prompt design component (2121) can be described as an AI component that uses a learning algorithm and / or a neural network to provide enhanced prompts over time. For example, the prompt design component (2121) can use the acquired information to access a knowledge component (e.g., a knowledge repository (2190)) that includes user preference data, a prompt library, and / or prompt examples to generate or obtain a prompt. The generated prompt can be provided to the generative AI model (2130) (e.g., including an LLM or LMM).
[0222] For example, the API / plugin management component (2122) within the AI framework (2120) may be utilized to support communication for additional information requested (or induced) in connection with the prompt provided (or to be provided) to the generative AI model (2130). For example, the API / plugin management component (2122) may be utilized to create or establish channels for communication with various data sources (e.g., knowledge repositories (2190)). For example, the API / plugin management component (2122) may support access to at least some of the data sources. For example, the API / plugin management component (2122) may be utilized to request another component (e.g., an application / service component (2180)) to perform feedback (or response) according to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (2122) may be provided to the prompt design component (2121) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (2122) may be provided to the generative AI model (2130).
[0223] For example, the improvement component (2123) within the AI framework (2120) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (2130). For example, the improvement component (2123) can determine or verify whether the content obtained from the generative AI model (2130) is related to the input. For example, the improvement component (2123) can determine or verify whether the content obtained from the generative AI model (2130) contains biased content. For example, the improvement component (2123) can determine or verify whether the content obtained from the generative AI model (2130) contains harmful content. For example, the improvement component (2123) can support or assist in performing additional processing to improve the content obtained from the generative AI model (2130). For example, the improvement component (2123) may support providing hints to the user to improve the content.
[0224] A generative AI model (2130) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information related to the prompt, but relative to the prompt. For example, the feedback may include new content related to the prompt. For example, the generative AI model (2130) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, a generative AI model (2130) may include an LMM that generates the feedback by recognizing text, images, and / or speech.
[0225] As a non-limiting example, the AI framework (2120) and / or the generative AI model (2130) may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (210) or processor (2120)). For example, the AI module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.
[0226] As described above, an electronic device (e.g., electronic device (200)) may include a display (e.g., display (240)), a communication circuit (e.g., communication circuit (230)), at least one processor (e.g., at least one processor (210)) including a processing circuit, and a memory (e.g., memory (220)) that stores instructions and includes one or more storage media. The instructions, when individually or collectively executed by the at least one processor, display a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuit, using the display, and display a second UI of a second application for performing a translation service for displaying text related to translation of a first utterance of a user of the electronic device and translation of a second utterance of a user of the other electronic device performed through a second application during the VoIP call, using the display together with the first UI, and obtaining a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI, and obtaining a second image for the first UI based on a release of the touch input, and identifying that a mute function of the first application is activated based on the first image and the second image, and based on the identification that the mute function is activated, while the mute function is activated, The electronic device may be caused to stop displaying text related to the translation of the first utterance within the second UI.
[0227] For example, the first application may be configured to maintain the state of the microphone activated for the VoIP call while activating the mute function of the first application, and to refrain from transmitting data about the first utterance received through the microphone to the other electronic device.
[0228] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the second UI together with the first UI by displaying the first UI and the second UI overlapping at least a portion of the first UI. For example, the touch input may be received through an area of the display where the at least a portion of the first UI and the second UI are displayed.
[0229] For example, the second UI overlapping at least a portion of the first UI may be at least partially translucent.
[0230] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, using the display, text related to the translation of the first utterance within the second UI prior to the identification.
[0231] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: capture a screen including the first UI and the second UI based on a touch input received on the display while displaying the second UI together with the first UI; capture a screen including the first UI and the second UI based on the release of the touch input; identify an area on the display including a location where the touch input was received; obtain the first image by cropping the screen captured based on the touch input using the area; obtain the second image by cropping the screen captured based on the release of the touch input using the area; compare the first image and the second image; and identify the second image as being at least partially different from the first image based on the comparison, thereby identifying that the mute function of the first application is activated based on the touch input.
[0232] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify that the mute function of the first application is not activated in response to the touch input based on identifying the second image corresponding to the first image based on the comparison.
[0233] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide a trained model within the electronic device with first data related to the first image and second data related to the second image based on identifying the second image as at least partially different from the first image based on the comparison, and to obtain information from the model indicating that each of the first data and the second data includes an object for a mute function, thereby identifying that the mute function of the first application is activated in response to the touch input.
[0234] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the first data by resizing the first image, and to obtain the second data by resizing the second image.
[0235] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the area on the display that includes the location where the touch input was received, identify whether the area is included within a reference area on the display, obtain the first image by cropping the screen captured based on the touch input using the area based on the area included within the reference area, and obtain the second image by cropping the screen captured based on the release of the touch input using the area, and obtain the first image based on the area not included within the reference area, and refrain from obtaining the second image.
[0236] For example, the instructions, when individually or collectively executed by the at least one processor, capture a screen including the first UI and the second UI based on the touch input received on the display while displaying the second UI together with the first UI, thereby obtaining the first image; capture the screen based on the release of the touch input, thereby obtaining the second image; compare the first image with the second image, and identify, based on the comparison, a portion of the first image that is different from the second image and a portion of the second image that is different from the first image; identify, based on color data of the portion of the second image that is within a reference range with respect to color data of the portion of the first image, that the mute function of the first application is not activated in response to the touch input; and identify, based on color data of the portion of the second image that is outside the reference range with respect to the color data of the portion of the first image, that the mute function of the first application is activated in response to the touch input. May cause electronic devices to malfunction.
[0237] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify that the mute function of the first application is not activated in response to the touch input based on identifying the second image corresponding to the first image based on the comparison between the first image and the second image.
[0238] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify that the mute function of the first application is activated in response to the touch input, based on the color data of the portion of the second image that is outside the reference range for the color data of the portion of the first image, by providing first data related to the portion of the first image and second data related to the portion of the second image to a trained model within the electronic device, and obtaining information from the model indicating that each of the first data and the second data includes an object for the mute function.
[0239] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the first data by resizing the portion of the first image based on the color data of the portion of the second image that is outside the reference range with respect to the color data of the portion of the first image, and to obtain the second data by resizing the portion of the second image.
[0240] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control not to perform translation of the first utterance based on the identification that the mute function is activated.
[0241] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: receive another touch input while ceasing to display text related to the translation of the first utterance within the second UI displayed together with the first UI; obtain a third image for the first UI based on the other touch input; obtain a fourth image for the first UI based on release of the other touch input; identify, using the third image and the fourth image, that the mute function of the first application is disabled based on the other touch input; and resume displaying, within the second UI, a result of the translation of the first utterance based on the identification that the mute function of the first application is disabled based on the other touch input.
[0242] As described above, an electronic device (e.g., electronic device (200)) may include a display (e.g., display (240)), a communication circuit (e.g., communication circuit (230)), at least one processor (e.g., at least one processor (210)) including a processing circuit, and a memory (e.g., memory (220)) that stores instructions and includes one or more storage media. The instructions, when individually or collectively executed by the at least one processor, display a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuit, using the display, and display a second UI of a second application for performing a translation service for displaying text related to translation of a first utterance of a user of the electronic device and translation of a second utterance of a user of the other electronic device performed through a second application during the VoIP call, using the display together with the first UI, and obtaining a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI, obtaining a second image for the first UI based on a release of the touch input, and identifying that a video call function of the first application is activated based on the first image and the second image, and based on the identification that the video call function is activated, displaying the second UI while the video call function is activated. This may cause the electronic device to stop displaying.
[0243] As described above, a non-transitory computer-readable storage medium can store one or more programs. The one or more programs, when executed by an electronic device having a communication circuit and a display, display a first user interface (UI) of a first application for performing a voice over internet protocol (VoIP) call with another electronic device using the communication circuit, using the display, and display a second UI of a second application for performing a translation service for displaying text related to translation of a first utterance of a user of the electronic device and translation of a second utterance of a user of the other electronic device performed through a second application during the VoIP call, using the display together with the first UI, and obtaining a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI, and obtaining a second image for the first UI based on release of the touch input, and identifying that a mute function of the first application is activated based on the first image and the second image, and based on the identification that the mute function is activated, deactivating the first UI while the mute function is activated. It may include instructions that cause the electronic device to stop displaying text related to the translation of the utterance within the second UI.
[0244] For example, the one or more programs may include instructions that cause the electronic device, when executed by the electronic device, to: capture a screen including the first UI and the second UI based on a touch input received on the display while displaying the second UI together with the first UI; capture the screen based on the release of the touch input; identify an area on the display including a location where the touch input was received; crop the screen captured based on the touch input using the area to obtain the first image; crop the screen captured based on the release of the touch input using the area to obtain the second image; compare the first image with the second image; and identify the second image as being at least partially different from the first image based on the comparison, thereby identifying that the mute function of the first application is activated based on the touch input.
[0245] For example, the one or more programs may include instructions that cause the electronic device, when executed by the electronic device, to: obtain the first image by capturing the screen based on the touch input received on the display while displaying the second UI together with the first UI; obtain the second image by capturing the screen based on the release of the touch input; compare the first image with the second image; and, based on the comparison, identify a portion of the first image that is different from the second image and a portion of the second image that is different from the first image; identify that the mute function of the first application is not activated based on color data of the portion of the second image that is within a reference range with respect to color data of the portion of the first image; and identify that the mute function of the first application is activated based on color data of the portion of the second image that is outside the reference range with respect to color data of the portion of the first image.
[0246] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0247] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0248] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0249] Various embodiments of the present document may be implemented as software (e.g., a program (2040)) including one or more instructions stored in a storage medium (e.g., an internal memory (2036) or an external memory (2038)) readable by a machine (e.g., an electronic device (2001)). For example, a processor (e.g., a processor (2020)) of the machine (e.g., an electronic device (2001)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0250] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0251] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, display; communication circuit; At least one processor comprising a processing circuit; and A memory storing instructions, comprising one or more storage media, wherein the instructions, when individually or collectively executed by the at least one processor, Using the above communication circuit, displaying a first UI (user interface) of a first application for performing a VoIP (voice over internet protocol) call with another electronic device using the display; A second UI of the second application for performing a translation service for displaying text related to the translation of a first utterance of a user of the electronic device and the translation of a second utterance of a user of the other electronic device performed through a second application during the VoIP call, is displayed using the display together with the first UI; Acquire a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI; Based on the release of the above touch input, a second image for the first UI is acquired; Based on the first image and the second image, identifying that the mute function of the first application is activated; and Based on the above identification that the mute function is activated, to stop displaying text related to the translation of the first utterance within the second UI while the mute function is activated; causing the above electronic device, Electronic devices.
2. In claim 1, the first application, While activating the mute function of the first application, the state of the microphone activated for the VoIP call is maintained, and data about the first utterance received through the microphone is refrained from being transmitted to the other electronic device. Electronic devices.
3. In claim 1, the instructions, when individually or collectively executed by the at least one processor, By displaying the second UI overlapping the first UI and at least a portion of the first UI, the second UI is displayed together with the first UI. causing the above electronic device, The above touch input is, Received through an area of the display where at least a portion of the first UI and the second UI are displayed, Electronic devices.
4. In claim 3, the second UI overlapping at least a portion of the first UI is at least partially translucent, Electronic devices.
5. In claim 1, the instructions, when individually or collectively executed by the at least one processor, Before the above identification, display the text related to the translation of the first utterance within the second UI using the display. causing the above electronic device, Electronic devices.
6. In claim 1, the instructions, when individually or collectively executed by the at least one processor, Capturing a screen including the first UI and the second UI based on a touch input received on the display while displaying the second UI together with the first UI; Based on the release of the touch input, a screen including the first UI and the second UI is captured, Identifying an area on the display that includes a location where the touch input was received; By cropping the screen captured based on the touch input using the area, the first image is obtained, By cropping the captured screen based on the release of the touch input using the area, the second image is obtained, Compare the first image and the second image above, Based on the comparison, identifying the second image as being at least partially different from the first image, to identify that the mute function of the first application is activated in response to the touch input. causing the above electronic device, Electronic devices.
7. In claim 6, the instructions, when individually or collectively executed by the at least one processor, Based on the identification of the second image corresponding to the first image according to the comparison, to identify that the mute function of the first application is not activated according to the touch input, causing the above electronic device, Electronic devices.
8. In claim 6, the instructions, when individually or collectively executed by the at least one processor, Providing first data related to the first image and second data related to the second image to a trained model within the electronic device based on identifying the second image as being at least partially different from the first image based on the comparison, Based on obtaining information from the model indicating that each of the first data and the second data includes an object for a mute function, identifying that the mute function of the first application is activated according to the touch input, causing the above electronic device, Electronic devices.
9. In claim 8, the instructions, when individually or collectively executed by the at least one processor, Obtaining the first data by resizing the first image, and To obtain the second data by resizing the second image, causing the above electronic device, Electronic devices.
10. In claim 6, the instructions, when individually or collectively executed by the at least one processor, Identifying the area on the display that includes the location where the touch input is received, Identify whether the above area is included within the reference area on the display, Based on the area included in the reference area, the first image is obtained by cropping the screen captured based on the touch input using the area, and the second image is obtained by cropping the screen captured based on the release of the touch input using the area. Based on the area not included in the reference area, acquire the first image and refrain from acquiring the second image. causing the above electronic device, Electronic devices.
11. In claim 1, the instructions, when individually or collectively executed by the at least one processor, Acquire the first image by capturing the screen including the first UI and the second UI based on the touch input received on the display while displaying the second UI together with the first UI, Based on the release of the touch input, the second image is obtained by capturing the screen, Compare the first image and the second image above, Based on the above comparison, a part of the first image that is different from the second image and a part of the second image that is different from the first image are identified, Based on the color data of the part of the second image that is within the reference range for the color data of the part of the first image, it is identified that the mute function of the first application is not activated according to the touch input, To identify that the mute function of the first application is activated according to the touch input based on the color data of the part of the second image that is outside the reference range for the color data of the part of the first image; causing the above electronic device, Electronic devices.
12. In claim 11, the instructions, when individually or collectively executed by the at least one processor, Identifying that the mute function of the first application is not activated according to the touch input based on identifying the second image corresponding to the first image based on the comparison between the first image and the second image; causing the above electronic device, Electronic devices.
13. In claim 11, the instructions, when individually or collectively executed by the at least one processor, Based on the color data of the part of the second image that is outside the reference range for the color data of the part of the first image, providing the trained model in the electronic device with first data related to the part of the first image and second data related to the part of the second image, Based on obtaining information from the model indicating that each of the first data and the second data includes an object for the mute function, identifying that the mute function of the first application is activated according to the touch input, causing the above electronic device, Electronic devices.
14. In a non-transitory computer-readable storage medium storing one or more programs, when the one or more programs are executed by an electronic device having a communication circuit and a display, Using the above communication circuit, displaying a first UI (user interface) of a first application for performing a VoIP (voice over internet protocol) call with another electronic device using the display; A second UI of the second application for performing a translation service for displaying text related to the translation of a first utterance of a user of the electronic device and the translation of a second utterance of a user of the other electronic device performed through a second application during the VoIP call, is displayed using the display together with the first UI; Acquire a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI; Based on the release of the above touch input, a second image for the first UI is acquired; Based on the first image and the second image, identifying that the mute function of the first application is activated; and instructions that cause the electronic device to stop displaying text related to the translation of the first utterance within the second UI while the mute function is activated, based on the identification that the mute function is activated; Computer readable storage medium.
15. A method for an electronic device having a display and communication circuit, An operation of displaying a first UI (user interface) of a first application for performing a VoIP (voice over internet protocol) call with another electronic device using the above communication circuit, using the display; An action of displaying, using the display, together with the first UI, a second UI of the second application for performing a translation service for displaying text related to the translation of the first utterance of the user of the electronic device and the translation of the second utterance of the user of the other electronic device performed through the second application during the VoIP call; An operation of obtaining a first image for the first UI based on a touch input received on the display while displaying the second UI together with the first UI; An operation of obtaining a second image for the first UI based on the release of the above touch input, An operation for identifying that the mute function of the first application is activated based on the first image and the second image; Based on the identification that the mute function is activated, an action is included to stop displaying text related to the translation of the first utterance within the second UI while the mute function is activated. method.
Citation Information
Patent Citations
Mobile terminal and control method for mobile terminal
KR101867514B1
BeamData Auto Scan System
KR1020220102474A
Manufacturing method of light emitting element
KR102039090B1
Translation Method and Electronic Device
US20210385328A1
Classification device, classification method, and classification program
US20240153241A1