Electronic device and control method thereof

The electronic device synchronizes audio and text output using threshold values and similarity analysis to address timing discrepancies, enhancing accuracy and coherence in audio and text presentation.

WO2026151148A1PCT designated stage Publication Date: 2026-07-16SAMSUNG ELECTRONICS CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-12-30
Publication Date
2026-07-16

AI Technical Summary

Technical Problem

Existing systems face challenges in accurately matching audio and text information due to time discrepancies in recognizing and acquiring audio and text data, leading to inefficiencies in synchronization.

Method used

An electronic device with a processor that synchronizes audio and text output by adjusting timing based on threshold values, length comparisons, and similarity analysis to ensure accurate matching.

Benefits of technology

Enhances the accuracy and synchronization of audio and text information output, improving user experience by ensuring timely and coherent presentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025023220_16072026_PF_FP_ABST
    Figure KR2025023220_16072026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device comprises: a display; a communication interface; a memory storing instructions; and at least one processor, wherein the instructions, when collectively or individually executed by the at least one processor, instruct the electronic device to: when an output time of first text information and first audio information included in video data is greater than or equal to a threshold value, acquire second text information corresponding to the first audio information or acquire second audio information corresponding to the first text information, on the basis of stored information; and, on the basis of the acquired second text information or second audio information, control the display to output the first audio information together with the second text information or control the display to output the second audio information together with the first text information.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and control method thereof

[0001] The present disclosure relates to an invention concerning an electronic device and a method for controlling the same, and more specifically, includes an electronic device for matching text information and a speaker and a method for controlling the same.

[0002] Recently, audio information and text information matching the audio information are provided to the user simultaneously.

[0003] However, conventionally, there was a problem where audio information and text information were not matched due to the time required to recognize audio information and text information or the time required to acquire information.

[0004] Accordingly, there is a recent demand to explore ways to improve the accuracy of matching audio and text information.

[0005] An electronic device according to one embodiment of the present disclosure comprises: a display; a communication interface; a memory for storing instructions; and at least one processor; wherein, when the instructions are executed collectively or individually by the at least one processor, the electronic device may, if the output time of first audio information and first text information included in image data is greater than or equal to a threshold value, acquire second text information corresponding to the first audio information or acquire second audio information corresponding to the first text information based on the stored information, and control the display to output the first audio information and the second text information together or control the display to output the second audio information and the first text information together based on the acquired second text information or second audio information.

[0006] The output time of the first audio information and the first text information included in the above video data is, if the first audio information is output before the first text information, the time corresponding to the time from when the first audio information is output until when the output of the first text information is completed, and if the first text information is output before the first audio information, the time corresponding to the time from when the first text information is output until when the output of the first audio information is completed, and the stored information may include at least one non-corresponding audio information and text information.

[0007] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may obtain the length of audio information corresponding to the first text information based on the length of the first text information, obtain second audio information corresponding to the first text information among a plurality of audio information included in the stored information based on the length of the obtained audio information, and control the display to output the obtained second audio information and the first text information.

[0008] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may obtain the length of text information corresponding to the first audio information based on the length of the first audio information, obtain second text information including the length of the obtained text information among the information corresponding to at least one text information included in the stored information, and control the display to output the first audio information and the second text information simultaneously.

[0009] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may control the display to output the second audio information and the first text information, if the second audio information including the length of the acquired audio information is not acquired, acquire the corresponding third text information, and if the similarity value between the third text information and the first text information is greater than or equal to a preset value, acquire the second audio information corresponding to the first text information based on information regarding the third text information and the audio information corresponding to the third text information.

[0010] The above similarity value may be obtained based on connecting phrases, connecting symbols, or repeated words included in the above first text information and the above third text information.

[0011] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may control the display to output the first audio information and the second text information, if the second text information including the length of the acquired text information is not acquired, acquire a video frame containing the corresponding third text information, and if the similarity value between the object included in the video frame containing the second text information and the video frame corresponding to the first audio information is greater than or equal to a preset value, acquire the second text information corresponding to the first audio information based on the video frame containing the third text information.

[0012] The above similarity value may be obtained based on the face of a person, the gender of a person, or a background image included in the video frame containing the third text information and the video frame containing the first audio information.

[0013] When the above instructions are executed collectively or individually by the at least one processor, the electronic device may make the threshold value proportional to the length of the first audio information or the length of the second text information.

[0014] When the above instructions are executed collectively or individually by the at least one processor, the electronic device can control the display to output the first audio information and the first text information simultaneously if the output time of the first audio information and the first text information included in the image data is less than a threshold value.

[0015] A method for controlling an electronic device according to one embodiment of the present disclosure may include: a step of, if the output time of first audio information and first text information included in image data is greater than or equal to a threshold value, acquiring second text information corresponding to the first audio information or acquiring second audio information corresponding to the first text information based on stored information; and a step of, based on the acquired second text information or second audio information, controlling the display to output the first audio information and the second text information together or to output the second audio information and the first text information together.

[0016] The output time of the first audio information and the first text information included in the above video data is, if the first audio information is output before the first text information, the time corresponding to the time from when the first audio information is output until when the output of the first text information is completed, and if the first text information is output before the first audio information, the time corresponding to the time from when the first text information is output until when the output of the first audio information is completed, and the stored information may include at least one non-corresponding audio information and text information.

[0017] The method may include: a step of obtaining the length of audio information corresponding to the first text information based on the length of the first text information; a step of obtaining second audio information corresponding to the first text information among a plurality of audio information included in the stored information based on the length of the obtained audio information; and a step of outputting the obtained second audio information and the first text information.

[0018] The method may include: a step of obtaining the length of text information corresponding to the first audio information based on the length of the first audio information; a step of obtaining second text information including the length of the obtained text information among information corresponding to at least one text information included in the stored information; and a step of simultaneously outputting the first audio information and the second text information.

[0019] If a second audio information including the length of the audio information obtained above is not obtained, the method may include the step of obtaining a previously corresponding third text information; if the similarity value between the third text information and the first text information is greater than or equal to a previously set value, the method may include the step of obtaining a second audio information corresponding to the first text information based on information regarding the third text information and the audio information corresponding to the third text information; and the step of outputting the second audio information and the first text information.

[0020] The above similarity value may include a step of obtaining it based on connecting phrases, connecting symbols, or repeated words included in the first text information and the third text information.

[0021] If a second text information including the length of the text information obtained above is not obtained, the method may include: a step of obtaining a video frame including a previously corresponding third text information; if the similarity value between an object included in the video frame including the second text information and a video frame corresponding to the first audio information is greater than or equal to a preset value, a step of obtaining a second text information corresponding to the first audio information based on the video frame including the third text information; and a step of outputting the first audio information and the second text information.

[0022] The above similarity value may be obtained based on the face of a person, the gender of a person, or a background image included in the video frame containing the third text information and the video frame containing the first audio information.

[0023] The above threshold value may be proportional to the length of the first audio information or the length of the second text information.

[0024] A non-transient computer-readable recording medium comprising a program for executing a method for controlling an electronic device according to one embodiment of the present disclosure, wherein the method for controlling the electronic device may include: a step of, if the output time of first audio information and first text information included in image data is greater than or equal to a threshold value, acquiring second text information corresponding to the first audio information or acquiring second audio information corresponding to the first text information based on stored information; and a step of, based on the acquired second text information or second audio information, controlling the display to output the first audio information and the second text information together or to output the second audio information and the first text information together.

[0025] FIG. 1 is a drawing for explaining an embodiment of an electronic device according to one embodiment of the present disclosure, and

[0026] FIG. 2 is a block diagram for explaining an electronic device according to one embodiment of the present disclosure, and

[0027] FIGS. 3 to 11 are drawings for explaining a method for identifying audio information and text information corresponding to text information and audio information, respectively, according to an embodiment of the present disclosure.

[0028] FIG. 12 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.

[0029] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.

[0030] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.

[0031] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.

[0032] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. The singular expression includes the plural expression unless the context clearly indicates otherwise.

[0033] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.

[0034] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.

[0035] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.

[0036] Where it is stated that a certain component (e.g., a first component) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., a second component), it should be understood that the said certain component may be directly connected to the said other component or connected through another component (e.g., a third component).

[0037] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between said certain component and said other component.

[0038] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.

[0039] Instead, in some situations, the expression “device configured to do something” may mean that the device is “capable of doing something” together with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) capable of performing those operations by executing one or more software programs stored in a memory device.

[0040] In the embodiments, a 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for the 'module' or 'part' that needs to be implemented in specific hardware.

[0041] Meanwhile, the various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.

[0042] Hereinafter, various embodiments of the present invention will be described in detail using the attached drawings.

[0043] FIG. 1 is a drawing for explaining an embodiment of an electronic device according to one embodiment of the present disclosure. The electronic device (100) may be implemented as a TV as shown in FIG. 1, but this is only one embodiment, and it is obvious that it may be implemented as various types of electronic devices such as a tablet, laptop computer, smartphone, camera, PC, etc.

[0044] Meanwhile, for the sake of convenience of explanation, the following description is based on the premise that the electronic device (100) includes a display (130), but this is merely one embodiment, and it is obvious that the electronic device (100) can output video data or audio signals using HDMI or DP (Display Port) without including a display (130).

[0045] If the electronic device (100) does not include a display (130), the electronic device (100) may transmit data to an external device that includes a display and output video data or audio signals from the external device.

[0046] First, the electronic device (100) can acquire video data including audio information and text information (or subtitle information).

[0047] As illustrated in FIG. 1, the electronic device (100) can identify that the time from the time (t1) when speaker 1 outputs audio information (audio 1) to the time (t2) when the output of text information (subtitle 1) is completed is less than a threshold value. At this time, the electronic device (100) can output subtitle 1 and audio 1 together after identifying that subtitle 1 and audio 1 match.

[0048] Additionally, the electronic device (100) can identify that subtitle 2 and audio 2 do not match if the time from the point in time (t3) when the text information (subtitle 2) of speaker 2 is output to the point in time (t4) when the output of audio information (audio 2) is completed is greater than or equal to a threshold value.

[0049] At this time, the electronic device (100) can store audio information (audio 2) and text information (subtitle 2) in a waiting queue.

[0050] The waiting queue can store multiple text information that does not match the audio information. In this case, the multiple text information can store scene information at the time when the text information is output together with the text information.

[0051] Additionally, the waiting queue can store multiple audio information that does not match the text information. In this case, the multiple audio information may also store scene information at the time when the audio information is output together with the audio information.

[0052] Meanwhile, for the sake of convenience of explanation, the scene information at the time when text information is output is defined as 'scene information corresponding to the text information,' and the scene information at the time when audio information is output is defined as 'scene information corresponding to the audio information,' and will be described below.

[0053] The electronic device (100) can identify the length of audio 2 and, based on the identified length, identify the length of text 2 that matches audio 2. Additionally, the electronic device (100) can identify text information containing the length of the identified text 2 among a plurality of text information stored in a waiting queue.

[0054] For example, the electronic device (100) can identify that the length of audio 2, “Someday a painting might be hung in a gallery,” is 5 seconds. The electronic device (100) can identify that the text information corresponding to the length of the 5-second audio information is 15 to 20 characters. At this time, the electronic device (100) can identify that among the multiple text information stored in the waiting queue, “Someday a painting might be hung in a gallery” is text information having a length of 15 to 20 characters.

[0055] After identifying that audio 2 and text information of 'Maybe someday a painting will be hung in a gallery' match, the electronic device (100) can output audio 2 and the text information identified as matching together.

[0056] Meanwhile, for the convenience of explanation, an embodiment in which text information matching audio information is obtained using the length of audio information has been described above, but this is only one embodiment, and it is obvious that the electronic device (100) can obtain audio information matching text information using the length of text information.

[0057] However, if the electronic device (100) fails to obtain matching text information using the length of the audio information, it can obtain subtitle information that matches the audio information using scene information stored in the waiting queue.

[0058] Specifically, the electronic device (100) can acquire audio information and corresponding scene information. At this time, the electronic device (100) can acquire information about the speaker by using information about the face and gender of the person included in the audio information and corresponding scene information.

[0059] The electronic device (100) can identify the speaker's information included in the scene information corresponding to the multiple text information stored in the waiting queue.

[0060] At this time, the electronic device (100) can identify whether the similarity value between the speaker's information is greater than or equal to a preset value. If the similarity value is greater than or equal to a preset value, the electronic device (100) can identify that it contains information about the same speaker. At this time, the electronic device (100) can obtain information about text information stored in a waiting queue using the speaker's information.

[0061] For example, as illustrated in FIG. 1, the electronic device (100) can acquire audio information of a second speaker and corresponding scene information. At this time, the electronic device (100) can acquire information that the speaker included in the scene information is a "man," a "Western person," or that "the color of the hair is brown." However, this is merely an example, and the electronic device (100) can, of course, acquire information about "the shape of the man's eyes," "the shape of the man's nose," or "the size of the speaker's mouth."

[0062] The electronic device (100) can identify information regarding text information containing information about the same speaker among a plurality of scene information stored in a waiting queue using information about the appearance of the acquired speaker. That is, the electronic device (100) can identify scene information corresponding to subtitles output by the same speaker.

[0063] For example, the electronic device (100) can acquire a scene identified by the same male speaker among a plurality of scenes stored in a waiting queue. At this time, the electronic device (100) can acquire text information corresponding to the acquired scene, such as "Someday a painting might be hung in a gallery."

[0064] Meanwhile, as an embodiment for identifying text information using audio information has been described above, an embodiment for identifying audio information using text information will be described below.

[0065] Specifically, the electronic device (100) can identify whether there is text information and audio information matched prior to a preset time. If the electronic device (100) identifies that there is no text information and audio information matched prior to a preset time, it can identify whether text information prior to a preset time and current text information are contextually related.

[0066] If the electronic device (100) identifies that the speaker of the previously matched audio information and the speaker of the current text information are the same when identified as being contextually related, the electronic device (100) can identify the audio information in the waiting queue. Based on this, the electronic device (100) can identify the audio information in the waiting queue.

[0067] For example, as illustrated in FIG. 1, if audio information corresponding to the subtitle “The place where the makeup-up sunlight kisses me every morning” cannot be identified, it is possible to identify whether there is a subtitle among past subtitles with a similarity value greater than or equal to a threshold. In this case, if a subtitle greater than or equal to the threshold is identified, it is possible to identify that the speaker is the same. The electronic device (100) can match the voice information of the speaker associated with the past subtitle to the current subtitle.

[0068] In the step of matching text information and audio information, if the method described above is used, the matching accuracy of text information and audio information can be increased.

[0069] Meanwhile, for the sake of convenience of explanation, an embodiment in which an electronic device (100) identifies whether audio information and text information match has been described in detail; however, this is merely one embodiment, and an external device may receive identification result data in which the electronic device (100) identifies whether audio information and text information match. At this time, it goes without saying that the electronic device (100) can output text information and audio information based on the data received from the external device.

[0070] Meanwhile, although the present disclosure has been described above on the premise that an audio signal is already stored in the video data, it is obvious that the electronic device (100) can acquire the audio signal separately. At this time, the electronic device (100) can acquire the audio signal using a microphone (not shown) included in the electronic device (100), and can also acquire the audio signal using a microphone included in an external remote control.

[0071] FIG. 2 is a block diagram for explaining the configuration of an electronic device (100). The configuration shown in FIG. 2 is merely an example of various embodiments, and some components may be omitted, and new components may be added. As shown in FIG. 2, the electronic device (100) may include a communication interface (110), a memory (120), a display (130), a speaker (140), and a processor (150). The configuration shown in FIG. 2 is merely an example, and it goes without saying that some components may be deleted or added depending on the configuration of the electronic device (100).

[0072] First, the communication interface (110) is configured to communicate with various types of external devices according to various types of communication methods. In particular, the communication interface (110) can receive video data including audio information and text information from an external device. Additionally, the communication interface (110) can receive information about a waiting queue from an external device.

[0073] A wireless communication module may be a module that communicates wirelessly with an external device. For example, the wireless communication module may include at least one module among a Wi-Fi module, a Bluetooth module, an infrared communication module, an Ultra Wide-Band (UWB) module, or other communication modules.

[0074] Specifically, the wireless communication module may include at least one communication chip that performs communication according to various wireless communication standards, such as Zigbee, USB (Universal Serial Bus), MIPI CSI (Mobile Industry Processor Interface Camera Serial Interface), 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), and 5G (5th Generation), in addition to the communication method described above. However, this is merely one embodiment, and the communication interface (120) may use at least one communication module among the various communication modules.

[0075] A wired communication module may be a module that communicates with an external device via a wire. For example, a wired communication module may include at least one of a Local Area Network (LAN) module, an Ethernet module, a pair cable, a coaxial cable, or a fiber optic cable.

[0076] The memory (120) may store instructions or data related to an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and the components of the electronic device (100). In particular, the memory (120) may store information about a waiting queue or image data including text information and audio information. Additionally, the memory (120) may store information about a method for identifying audio information that matches text information.

[0077] Memory (120) can be implemented in various forms such as volatile memory (e.g., DRAM (dynamic RAM), SRAM (static RAM), or SDRAM (synchronous dynamic RAM), non-volatile memory (e.g., OTPROM (one time programmable ROM), PROM (programmable ROM), EPROM (erasable and programmable ROM), EEPROM (electrically erasable and programmable ROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD).

[0078] The display (130) can display various information. In particular, the display (130) can display scene information and text information.

[0079] The display (130) can be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, and a PDP (Plasma Display Panel). The display may also include a driving circuit, a backlight unit, etc., which can be implemented in forms such as an a-si TFT (amorphous silicon thin film transistor), an LTPS (low temperature poly silicon) TFT, and an OTFT (organic TFT). The display can be implemented as a touch screen combined with a touch sensor, a flexible display, a 3D display, a three-dimensional display, etc. According to various embodiments of the present disclosure, the display (130) may include not only a display panel that outputs an image, but also a bezel that houses the display panel.

[0080] The speaker (140) can output audio information. In particular, the speaker (140) can output audio information received through the communication interface (120) or audio information stored in the memory (120). At this time, the output of the speaker (140) can be controlled by the processor (150). For example, the timing of the speaker (140) outputting audio information can be controlled by the processor (150).

[0081] Meanwhile, the speaker (140) may include various types of speaker modules. For example, the speaker (140) may include a woofer speaker, a mid-range speaker, and a tweeter.

[0082] Additionally, the speaker (140) can be placed at various locations on the electronic device (100). For example, if the electronic device (100) is a TV, the speaker (140) can be placed on the top of the electronic device (100). However, this is merely an example, and the speaker (140) can be placed on the bottom of the electronic device (100). Additionally, the speaker (140) can output audio information data in various directions. For example, the speaker (140) can output audio information data toward the top of the electronic device (100). However, this is merely an example, and the speaker (140) can output audio information data toward the bottom, side, or front of the electronic device (100). When multi-channel audio information data is received, the speaker (140) can output audio information according to the channel information.

[0083] The processor (150) may include one or more processors. Specifically, the one or more processors may include one or more of a CPU (Central Processing Unit), GPU (Graphics Processing Unit), APU (Accelerated Processing Unit), MIC (Many Integrated Core), DSP (Digital Signal Processor), NPU (Neural Processing Unit), hardware accelerator, or machine learning accelerator. The processor (150) may control one or any combination of other components of an electronic device and may perform operations or data processing related to communication. The one or more processors may execute one or more programs or instructions stored in memory. For example, the one or more processors may perform a method according to one embodiment of the present disclosure by executing one or more instructions stored in memory.

[0084] One or more processors may be implemented as a single-core processor comprising one core, or as one or more multicore processors comprising multiple cores (e.g., homogeneous multicore or heterogeneous multicore). When one or more processors are implemented as multicore processors, each of the multiple cores included in the multicore processor may include internal processor memory such as cache memory or on-chip memory, and a common cache shared by multiple cores may be included in the multicore processor. Additionally, each of the multiple cores included in the multicore processor (or some of the multiple cores) may independently read and execute program instructions for implementing a method according to one embodiment of the present disclosure, or all (or some) of the multiple cores may be linked together to read and execute program instructions for implementing a method according to one embodiment of the present disclosure.

[0085] In particular, if the output time of the first audio information and the first text information included in the video data is greater than or equal to a threshold value, the processor (150) may obtain second text information corresponding to the first audio information or obtain second audio information corresponding to the first text information based on the stored information. Additionally, based on the obtained second text information or second audio information, the processor (150) may control the display to output the first audio information and the second text information together or control the display to output the second audio information and the first text information together.

[0086] At this time, the output time of the first audio information and the first text information included in the video data may be a time corresponding to the time from when the first audio information is output to when the output of the first text information is completed, if the first audio information is output before the first text information. Additionally, if the first text information is output before the first audio information, it is a time corresponding to the time from when the first text information is output to when the output of the first audio information is completed, and the stored information may include at least one audio information and text information that does not correspond.

[0087] The processor (150) can obtain the length of audio information corresponding to the first text information based on the length of the first text information, and obtain second audio information corresponding to the first text information among a plurality of audio information included in the stored information based on the length of the obtained audio information. Additionally, the processor (150) can control the display to output the obtained second audio information and the first text information.

[0088] Based on the length of the first audio information, the length of text information corresponding to the first audio information can be obtained, and a second text information including the length of the obtained text information among the information corresponding to at least one text information included in the stored information can be obtained. Additionally, the processor (150) can control the display to output the first audio information and the second text information simultaneously.

[0089] If the processor (150) does not acquire second audio information including the length of the acquired audio information, it may acquire the corresponding third text information. Additionally, if the similarity value between the third text information and the first text information is greater than or equal to a preset value, the processor (150) may acquire second audio information corresponding to the first text information based on information regarding the third text information and audio information corresponding to the third text information, and control the display to output the second audio information and the first text information.

[0090] Additionally, the processor (150) may obtain a similarity value based on a connecting phrase, connecting code, or repeated word included in the first text information and the third text information.

[0091] If the second text information including the length of the acquired text information is not acquired, the processor (150) acquires a video frame containing the corresponding third text information, and if the similarity value between the object included in the video frame containing the second text information and the video frame corresponding to the first audio information is greater than or equal to a preset value, the processor (150) can acquire the second text information corresponding to the first audio information based on the video frame containing the third text information. The processor (150) can control the display to output the first audio information and the second text information.

[0092] At this time, the similarity value can be obtained based on the face of a person, the gender of a person, or a background image included in the video frame containing the third text information and the video frame containing the first audio information.

[0093] Additionally, the threshold value may be proportional to the length of the first audio information or the length of the second text information. If the output time of the first audio information and the first text information included in the video data is less than the threshold value, the display may be controlled to output the first audio information and the first text information simultaneously.

[0094] Hereinafter, each step described above will be described in detail with reference to FIGS. 3 to 7.

[0095] First, FIGS. 3 to 5 are flowcharts and drawings for explaining the case where the time at which audio information and text information are output according to one embodiment of the present disclosure is less than a threshold value.

[0096] First, the electronic device (100) can acquire image data (S310).

[0097] The electronic device (100) can recognize text information contained in the image (S320).

[0098] For example, the electronic device (100) can recognize a 'text area' displayed in the bottom area of ​​the scene included in the image as shown in FIG. 5.

[0099] Additionally, the electronic device (100) can obtain text information corresponding to a text area (S330).

[0100] For example, as shown in FIG. 5, the electronic device (100) can obtain text information such as ‘worked at a supermarket for 2 years’ by using an OCR recognition method in the ‘text information’ area of ​​the scene included in the image.

[0101] Meanwhile, the electronic device (100) can recognize audio information included in the video (S330). At this time, the electronic device (100) can obtain information about the length of the audio information included in the video.

[0102] The electronic device (100) can store text information and audio information corresponding to text information in a waiting queue. At this time, the waiting queue will be described later with reference to FIG. 4.

[0103] FIG. 4 is a drawing for explaining a waiting queue according to one embodiment of the present disclosure.

[0104] As illustrated in FIG. 4, the electronic device (100) can store voice 1 at 0 seconds and scene 1 displayed at the time voice 1 is output in the waiting queue. Additionally, the electronic device (100) can store subtitle 1 and scene 1 displayed at the time subtitle 1 is output in the waiting queue.

[0105] The electronic device (100) can identify whether the time during which voice 1 and subtitle 1 are output is less than a threshold value. If the threshold value is set to 2 seconds, the electronic device (100) can identify that voice 1 and subtitle 1 are matched. At this time, the electronic device (100) can delete voice 1 and subtitle 1 stored in memory (120).

[0106] Additionally, at 2 seconds, the electronic device (100) can acquire voice 2. At this time, the electronic device (100) can store voice 2 and scene 2 corresponding to voice 2 in a waiting queue.

[0107] The electronic device (100) can store unmatched voice information or text information (S350). Specifically, the electronic device (100) can store unmatched voice information or text information in a waiting queue or memory.

[0108] Additionally, the electronic device (100) can identify that there is no text information or audio information that is not matched in the waiting queue (or, memory) (S360).

[0109] At this time, the electronic device (100) can match subtitle information and audio information when the output time of audio information and text information included in the video data is less than a threshold (S370).

[0110] Specifically, the electronic device (100) can obtain the output time of audio information and text information included in the image data. At this time, the time obtained by the electronic device (100) may mean the time from the time when the first audio information is output to the time when the output of the first text information is completed, in the case where the first audio information is output before the first text information, and the time from the time when the first text information is output to the time when the output of the first audio information is completed, in the case where the first text information is output before the first audio information.

[0111] For example, as illustrated in FIG. 5, after audio information (audio information 1) is output at time t1, text information (subtitle 1) may be output at time t2. Additionally, audio information (audio information 1) may end at time t3, and output of text information (subtitle 1) may end at time t4.

[0112] Meanwhile, the threshold value may be proportional to the length of the first audio information or the second text information data.

[0113] The electronic device (100) can control the display (130) and the speaker (140) to output text information and audio information together, as shown in FIG. 5, when the time for recognizing text information and audio information is less than a threshold value. Specifically, the electronic device (100) can output the first audio information and the first text information simultaneously when the time for outputting the first audio information and the first text information included in the image data is less than a threshold value.

[0114] FIG. 6 is a diagram illustrating a method for identifying audio information and text information that match each of the text information and audio information when the text information and audio information recognition time according to one embodiment of the present disclosure is greater than or equal to a threshold value.

[0115] As illustrated in FIG. 6, the electronic device (100) can identify that the text information and audio information recognition time is from t1 to t4, and thus is greater than or equal to a threshold value. At this time, the electronic device (100) can identify that the text information and audio information do not match.

[0116] The electronic device (100) can identify text information that matches the audio information based on the length of the audio information. Specifically, the electronic device (100) can identify the length of the first audio information, identify the length of text information that matches the first audio information based on the length of the first audio information, and identify a second text information that includes the length of the identified text information among a plurality of text information stored in a waiting queue. Additionally, the electronic device (100) can output the first audio information and the second text information simultaneously.

[0117] This will be described later with reference to Fig. 7.

[0118] The electronic device (100) can acquire audio 1 of 500ms length as shown in FIG. 7. At this time, since the length of audio 1 is 500ms, the electronic device (100) can identify that the length of the corresponding text information is 2 to 5 characters.

[0119] The electronic device (100) can identify text information between 2 and 5 characters among a plurality of text information stored in a waiting queue. When the electronic device (100) identifies text information between 2 and 5 characters, it can identify that the identified text information and audio information 1 are matched.

[0120] Similar to the method described above, the electronic device (100) can acquire audio information 2 having a length of 1500ms. At this time, the electronic device (100) can identify that the length of the text information corresponding to the 1500ms audio information is 4 to 15 characters.

[0121] The electronic device (100) can identify text information between 4 and 15 characters among a plurality of text information stored in a waiting queue. When the electronic device (100) identifies text information between 4 and 15 characters, it can identify that the identified text information and audio information 2 are matched.

[0122] An embodiment for identifying text information that matches audio information based on the length of audio information has been described in detail. Below, an embodiment for identifying audio information that matches text information based on the length of text information will be described.

[0123] Specifically, the electronic device (100) can identify the length of the first text information, identify the length of the audio information corresponding to the first text information based on the length of the first text information, and identify the second audio information that matches the first text information among a plurality of audio information stored in a queue based on the identified audio information length. Additionally, the electronic device (100) can output the identified second audio information and the first text information.

[0124] FIG. 8 is a drawing for explaining an embodiment of identifying audio information that matches text information based on the length of text information, according to one embodiment of the present disclosure.

[0125] Specifically, the electronic device (100) can identify the length of a first text, identify the length of audio information corresponding to the first text based on the length of the first text, and identify a second audio information that matches the first text among a plurality of audio information stored in a queue based on the identified length of the audio information. Additionally, the electronic device (100) can output the identified second audio information and the first text.

[0126] As illustrated in FIG. 8, the electronic device (100) can acquire text information including 1 to 15 characters. At this time, the electronic device (100) can identify that the length of audio information corresponding to the length of 15 characters is 500ms.

[0127] The electronic device (100) can identify audio information containing a length of 500ms among a plurality of audio information stored in a waiting queue. When the electronic device (100) identifies audio information containing a length of 500ms, it can identify that the identified audio information and text information match.

[0128] At this time, when the electronic device (100) identifies matching audio information, it can identify matching text information by also considering information about whether the time of recognition of text information is close.

[0129] Similar to the method described above, the electronic device (100) can obtain text information including 2 to 5 characters. At this time, the electronic device (100) can identify that the length of audio information corresponding to the length of 5 characters is 200ms.

[0130] The electronic device (100) can identify audio information containing a length of 200ms among a plurality of audio information stored in a waiting queue. When the electronic device (100) identifies audio information containing a length of 200ms, it can identify that the identified audio information and text information match.

[0131] Meanwhile, although an embodiment capable of identifying matching text information or audio information based on the length of audio information or text information has been described above, it is of course possible that matching text information or audio information cannot be identified based on the length of audio information or text information.

[0132] In the following cases, if the corresponding text information cannot be identified based on the length of the audio information, the text information can be identified using previous scene data included in the video.

[0133] Specifically, if the second text information including the length of the identified text information is not identified, the electronic device (100) may identify a video frame containing previously matched third text information, obtain a video frame containing the second text information and a video frame corresponding to the first audio information, and obtain a similarity value between an object included in the video frame containing the second text information and a video frame corresponding to the first audio information. Additionally, if the similarity value is greater than or equal to a preset value, the electronic device (100) may identify second text information corresponding to the first audio information based on the video frame containing the third text information and output the first audio information and the second text information.

[0134] This will be described later with reference to FIGS. 9 and FIGS. 10.

[0135] FIG. 9 is a flowchart illustrating a method for identifying text information corresponding to audio information according to one embodiment of the present disclosure.

[0136] FIG. 10 is a drawing for explaining an embodiment of obtaining text information using similarity between a scene prior to a preset time and a current scene, according to one embodiment of the present disclosure.

[0137] First, the electronic device (100) can load multiple subtitle (or text information) information stored in a waiting queue (S910).

[0138] Afterwards, if there is no subtitle from a pre-set time prior to the electronic device (100), the electronic device (100) can obtain a similarity value between a scene from a pre-set time prior to the current scene and a scene from a pre-set time prior to the current scene (S920).

[0139] A similarity value can be obtained based on the face of a person, the gender of a person, or a background image included in a video frame containing third text information and a video frame containing first audio information.

[0140] As illustrated in FIG. 10, the electronic device (100) may include information about the speaker in the scene included on the left screen, such as ‘woman’s face’ and ‘the speaker is a woman’.

[0141] The electronic device (100) may equally include information about the speaker in the scene included in the right screen, such as ‘a woman’s face’ and ‘the speaker is a woman.’ Therefore, the electronic device (100) can identify that the same person appears in both the left screen and the right screen, and that facial features (e.g., a woman’s face and hairstyle) are similar. The electronic device (100) can obtain a similarity value between the left screen and the right screen.

[0142] When the electronic device (100) identifies that the similarity value is greater than or equal to a threshold value (S930-Y), it can obtain information about the speaker included in the scene where the similarity value is greater than or equal to the threshold value (S940).

[0143] As illustrated in FIG. 10, if the speaker included in the left screen and the speaker included in the right screen are both identified as the same speaker, the electronic device (100) can identify that the similarity value of the left screen and the right screen is greater than or equal to a threshold value.

[0144] In contrast, if it is identified that the similarity value is not greater than or equal to the threshold value (S930-N), multiple subtitle data stored in the waiting queue can be loaded (S910).

[0145] Meanwhile, although a method for identifying text information that matches audio information has been described in detail, this is merely one embodiment, and it is obvious that the electronic device (100) can identify audio information that matches text information.

[0146] Specifically, if the second audio information, including the length of the identified audio information, is not identified, the electronic device (100) may identify the previously matched third text information and obtain a correlation value between the third text information and the first text information. Additionally, if the correlation value is greater than or equal to a preset value, the electronic device (100) may identify the second audio information corresponding to the first text information based on information regarding the third text information and the audio information corresponding to the third text information. The electronic device (100) may output the second audio information and the first text information. Additionally, the correlation value may be obtained based on connecting phrases, connecting symbols, or repeated words included in the first text information and the third text information.

[0147] This will be described in detail later with reference to Fig. 11.

[0148] FIG. 11 is a drawing for explaining an embodiment that uses similarity between text information from a pre-set past time and text information when identifying audio information corresponding to text information according to one embodiment of the present disclosure.

[0149] First, the electronic device (100) can load multiple text data stored in a waiting queue (S1110).

[0150] At this time, if the electronic device (100) identifies that there is text information and audio information matched within a preset time (S1120-Y), it can identify audio information matched to text information matched before the preset time (S1140).

[0151] Specifically, if the electronic device (100) identifies text information and audio information matched prior to a preset time, it can identify information about the speaker included in the matched audio information (e.g., the speaker's voice tone, the speaker's speech rate, etc.). Additionally, the electronic device (100) can identify text information and the context of the text information prior to a preset time (e.g., the content of the text information, location information, or timestamp), etc.

[0152] At this time, the electronic device (100) can identify whether the content of the text information and the previously matched text information are similar. If the electronic device (100) identifies that the content of the text information is similar, it can identify that the speaker of the previously matched text information and the speaker of the current text information are the same speaker.

[0153] The electronic device (100) can identify audio information output by a speaker identified in a waiting queue and identify audio information that matches text information.

[0154] The electronic device (100) can compare speaker information extracted from a voice signal with the current subtitle if it is identified that the current text information and the text information matched prior to a preset time are similar.

[0155] In contrast, if the electronic device (100) fails to identify a matched subtitle within a preset time (S1120-N), it can obtain a similarity value using the relationship level between past text information and current text information prior to the preset time (S1130).

[0156] Specifically, if the electronic device (100) fails to identify a pair of audio information and text information matched prior to a preset time, it can identify whether past text information and current text information are contextually connected. Specifically, whether past text information and current text information are contextually connected can be identified using a similarity value.

[0157] At this time, the electronic device (100) can identify connecting phrases, connecting codes, or repeated words, or analyze text information, in order to identify text information that has a high similarity to the current text information among a plurality of text information.

[0158] For example, past text information could be “I feel good when the weather is nice in the morning,” and current text information could be “I mean a place where the bright sunshine kisses me every morning.” At this time, the electronic device (100) can identify that the similarity value between the past text information and the current text information is large.

[0159] When the similarity value is greater than or equal to a threshold value (S1150-Y), the electronic device (100) can obtain audio corresponding to text information from a waiting queue (S1160).

[0160] Specifically, the electronic device (100) can identify that the time stamp interval between the past text information and the current text information is short in terms of timestamps when it is identified that the past text information and the information and the current text information are contextually connected.

[0161] The electronic device (100) can extract audio information that is temporally close to the current text information. At this time, when the extracted audio information is converted into text information, it can identify that the converted text information and the current text information are similar. At this time, the electronic device (100) can match the extracted audio information and the current text information.

[0162] In contrast, when the similarity value of the electronic device (100) is less than the threshold value (S1150-N), the electronic device (100) can load multiple text information data stored in a waiting queue.

[0163] FIG. 12 is a flowchart illustrating a method for controlling an electronic device according to one embodiment of the present disclosure.

[0164] First, if the output time of the first audio information and the first text information included in the image data is greater than or equal to a threshold value, the electronic device (100) can obtain second text information corresponding to the first audio information or obtain second audio information corresponding to the first text information based on the stored information (S1210).

[0165] Meanwhile, regarding the operation of the electronic device (100) identifying matching audio information or text information using a waiting queue, as described above, a detailed explanation will be omitted.

[0166] Additionally, the electronic device (100) can output the first audio information and the second text information together or output the second audio information and the first text information together based on the acquired second text information or second audio information (S1220).

[0167] Specifically, if the electronic device (100) identifies that there is audio information among the multiple audio information stored in the waiting queue that matches the text information, it can output the audio information identified as matching and the text information together.

[0168] However, this is merely one embodiment, and if the electronic device (100) identifies that there is text information among the plurality of text information stored in the waiting queue that matches the audio information, it may output the text information identified as matching and the audio information together.

[0169] According to the above-described disclosure, if the audio information and text information are not matched, the electronic device (100) may match using the length of the audio information or the length of the text information. If matching cannot be done using the length of the audio information or the length of the text information, the electronic device (100) may match using a scene output together with the audio information or text information, or match and output the audio information or text information using information about the previous text information.

[0170] Additionally, methods according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user (20) devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0171] A method according to various embodiments of the present disclosure may be implemented as software comprising instructions stored on a machine-readable storage medium (e.g., a computer). The machine may include a server device or an electronic device according to the disclosed embodiments, which is a device capable of calling instructions stored from the storage medium and operating according to the called instructions.

[0172] Meanwhile, a device-readable storage medium may be provided in the form of a non-transitory readable recording medium. Here, 'non-transitory readable recording medium' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.

[0173] When the above instruction is executed by a processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or an interpreter.

[0174] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.

Claims

1. In an electronic device, display; Communication interface; Memory for storing instructions; and Includes at least one processor; and When the above instructions are executed collectively or individually by the at least one processor, the electronic device, If the output time of the first audio information and the first text information included in the video data is greater than or equal to a threshold value, the second text information corresponding to the first audio information is obtained or the second audio information corresponding to the first text information is obtained based on the stored information, and An electronic device that controls the display to output the first audio information and the second text information together, or controls the display to output the second audio information and the first text information together, based on the second text information or second audio information obtained above.

2. In Paragraph 1, The output time of the first audio information and the first text information included in the above video data is, If the first audio information is output before the first text information, the time corresponds to the time from when the first audio information is output until when the output of the first text information is completed. If the first text information is output before the first audio information, the time corresponds to the time from when the first text information is output until when the output of the first audio information is completed. The above-mentioned stored information is an electronic device comprising at least one non-corresponding audio information and text information.

3. In Paragraph 1, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Based on the length of the first text information, the length of audio information corresponding to the first text information is obtained, and Based on the length of the audio information obtained above, among the plurality of audio information included in the stored information, a second audio information corresponding to the first text information is obtained, and An electronic device that controls the display to output the second audio information and the first text information obtained above.

4. In Paragraph 1, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, Based on the length of the first audio information, the length of text information corresponding to the first audio information is obtained, and Acquiring a second text information including the length of the acquired text information among the information corresponding to at least one text information included in the stored information, An electronic device that controls the display to simultaneously output the first audio information and the second text information.

5. In Paragraph 3, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, If the second audio information including the length of the above-mentioned acquired audio information is not acquired, the previously corresponding third text information is acquired, and If the similarity value between the third text information and the first text information is greater than or equal to a preset value, second audio information corresponding to the first text information is obtained based on information regarding the third text information and audio information corresponding to the third text information, and An electronic device that controls the display to output the second audio information and the first text information.

6. In Paragraph 5, An electronic device that obtains the above similarity value based on connecting phrases, connecting symbols, or repeating words included in the above first text information and the above third text information.

7. In Paragraph 4, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, If the second text information including the length of the above-mentioned acquired text information is not acquired, an image frame including the previously corresponding third text information is acquired, and If the similarity value between an object included in a video frame containing the second text information and a video frame corresponding to the first audio information is greater than or equal to a preset value, the second text information corresponding to the first audio information is obtained based on a video frame containing the third text information, and An electronic device that controls the display to output the first audio information and the second text information.

8. In Paragraph 7, An electronic device that obtains the above similarity value based on the face of a person, the gender of a person, or a background image included in the video frame containing the above third text information and the video frame containing the above first audio information.

9. In Paragraph 1, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device in which the above threshold is proportional to the length of the first audio information or the length of the second text information.

10. In Paragraph 1, When the above instructions are executed collectively or individually by the at least one processor, the electronic device, An electronic device that controls the display to output the first audio information and the first text information simultaneously when the output time of the first audio information and the first text information included in the above video data is less than a threshold value.

11. In a method for controlling an electronic device, If the output time of the first audio information and the first text information included in the video data is greater than or equal to a threshold value, a step of obtaining second text information corresponding to the first audio information or obtaining second audio information corresponding to the first text information based on the stored information; and A control method comprising the step of controlling the display to output the first audio information and the second text information together, or to output the second audio information and the first text information together, based on the second text information or the second audio information obtained above.

12. In Paragraph 11, The output time of the first audio information and the first text information included in the above video data is, If the first audio information is output before the first text information, the time corresponds to the time from when the first audio information is output until when the output of the first text information is completed. If the first text information is output before the first audio information, the time corresponds to the time from when the first text information is output until when the output of the first audio information is completed. A control method wherein the above-mentioned stored information includes at least one non-corresponding audio information and text information.

13. In Paragraph 11, A step of obtaining the length of audio information corresponding to the first text information based on the length of the first text information; A step of obtaining a second audio information corresponding to the first text information among a plurality of audio information included in the stored information based on the length of the audio information obtained above; and A control method comprising the step of outputting the second audio information obtained above and the first text information.

14. In Paragraph 11, A step of obtaining the length of text information corresponding to the first audio information based on the length of the first audio information; A step of obtaining second text information including the length of the acquired text information among information corresponding to at least one text information included in the stored information; and A control method comprising the step of simultaneously outputting the first audio information and the second text information.

15. A non-transient computer-readable recording medium comprising a program for executing a method of controlling an electronic device, The control method of the above electronic device is, If the output time of the first audio information and the first text information included in the video data is greater than or equal to a threshold value, a step of obtaining second text information corresponding to the first audio information or obtaining second audio information corresponding to the first text information based on the stored information; and A computer-readable recording medium comprising the step of controlling the display to output the first audio information and the second text information together or to output the second audio information and the first text information together, based on the second text information or second audio information obtained above.