Display device and control method thereof
The display device addresses the challenge of visually impaired users accessing subtitles by using a neural network to detect and convert subtitles to audio, improving multi-view functionality for foreign language videos.
Patent Information
- Application Number
- PCT/KR2025/001874
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-06
- Filing Date
- 2025-02-07
- Publication Date
- 2025-09-11
AI Technical Summary
Visually impaired individuals face difficulty in watching videos with foreign language audio or content due to the inability to see translated subtitles when using multi-view functionality on display devices.
A display device equipped with a display, memory, and processor that utilizes a neural network model to detect and analyze text areas, identify subtitle regions, and perform character recognition to convert subtitles into audible form.
Enables visually impaired users to access subtitles in foreign language videos by converting them into audio, enhancing the usability of multi-view features for this demographic.
Smart Images

Figure KR2025001874_12092025_PF_FP_ABST
Abstract
Description
Display device and control method thereof
[0001] The present disclosure relates to a display device and a control method thereof, and more particularly, to a display device capable of providing audible subtitles and a control method thereof.
[0002] Typically, display devices can display a single screen on a single display. However, recently, multi-view functionality has been introduced, allowing multiple partial screens to be displayed on a single display. This multi-view feature allows multiple users to view different content on the display, or a single user to view multiple content simultaneously.
[0003] Meanwhile, when visually impaired people use the multi-view function to watch videos containing foreign language audio or foreign language content, they have difficulty watching the videos because they cannot see the translated subtitles.
[0004] According to the present disclosure, a display device according to at least one embodiment includes a display, a memory storing a neural network model, and a processor. The processor detects at least one text area of an image displayed on the display, analyzes the at least one text area, and obtains information about the at least one text area. The processor inputs information about the at least one text area into the neural network model, identifies a subtitle area among the at least one text area, and performs character recognition on text in the subtitle area to obtain subtitle data.
[0005] Meanwhile, a method for controlling a display device according to one or more embodiments of the present disclosure includes a step of detecting at least one text area of an image displayed on the display device, a step of analyzing the at least one text area to obtain information about the at least one text area, a step of inputting the information about the at least one text area obtained into a neural network model to identify a subtitle area among the at least one text area, and a step of performing character recognition on text in the subtitle area to obtain subtitle data.
[0006] Meanwhile, in a computer-readable recording medium including a program for executing a control method of a display device according to one or more embodiments of the present disclosure, the control method includes the steps of: detecting at least one text area of an image displayed on the display device; analyzing the at least one text area to obtain information about the at least one text area; inputting the information about the at least one text area obtained into a neural network model to identify a subtitle area among the at least one text area; performing character recognition on text in the subtitle area to obtain subtitle data; and, when a screen on which the subtitle area is located is selected, converting the subtitle data into voice and outputting it.
[0007] FIG. 1 is a block diagram showing the configuration of a display device according to various embodiments of the present disclosure.
[0008] FIG. 2 is a drawing for explaining a detailed configuration of a display device according to various embodiments of the present disclosure.
[0009] FIGS. 3 and 4 are drawings for explaining the operation of a processor according to various embodiments of the present disclosure.
[0010] FIG. 5 and FIG. 6 are drawings for explaining the continuity of images according to various embodiments of the present disclosure.
[0011] FIG. 7 is a diagram illustrating an advertisement identification factor according to various embodiments of the present disclosure.
[0012] FIGS. 8 to 10 are flowcharts for explaining a method of controlling a display device according to various embodiments of the present disclosure.
[0013] The terms used in this specification will be briefly explained, and the present disclosure will be described in detail.
[0014] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of this disclosure.
[0015] In this specification, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a feature (e.g., a number, function, operation, or component such as a part), and do not exclude the presence of additional features.
[0016] In this disclosure, expressions such as “A or B,” “at least one of A and / or B,” or “one or more of A or / and B” can include all possible combinations of the listed items. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.
[0017] As used herein, the expressions “first,” “second,” “first,” or “second,” etc., may describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0018] When it is said that a component (e.g., a first component) is “operatively or communicatively coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0019] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "consist of" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0020] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor (not shown), excluding any "modules" or "parts" that need to be implemented as specific hardware.
[0021] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.
[0022] FIG. 1 is a block diagram showing the configuration of a display device according to various embodiments of the present disclosure.
[0023] The display device (100) is a device capable of providing a single-view screen or a multi-view screen. A multi-view screen is a screen displayed on the display of the display device (100) composed of multiple sub-screens, each of which can output different content. When a user selects a function to provide a multi-view screen, the display device (100) outputs multiple contents on the display, allowing the user to view multiple contents simultaneously. For example, multiple users can use the function to provide a multi-view screen to simultaneously output and view multiple dramas selected according to their individual preferences. Alternatively, when a user selects the function to provide a multi-view screen, a single user can simultaneously output and view a sports broadcast and a drama.
[0024] Meanwhile, when a visually impaired person uses a display device (100) to watch a foreign film or foreign language video, the visually impaired person must select and watch the video dubbed in their native language or be able to understand the voice dialogue implemented in the foreign language, as they cannot see the translated subtitles. Here, the term "foreign film" refers to a foreign film produced in a foreign country based on the user watching the video, and the term "foreign language video" refers to video content that provides voice dialogue in a foreign language based on the user.
[0025] Referring to FIG. 1, the display device (100) includes a display (110), a memory (120), and a processor (130). At this time, the display device (100) is a device that displays image content and may be a TV, a desktop PC, a laptop, a video wall, a large format display (LFD), a digital signage, a digital information display (DID), a projector display, a digital video disk (DVD) player, a smartphone, a tablet PC, a monitor, smart glasses, a smart watch, etc. Any device that can display an input image may be used. This is merely an example and is not limited thereto, and the display device (100) may be implemented as various types of electronic devices. For example, the display device (100) may be implemented as a set-top box. If the display device (100) is implemented as a set-top box, the display (110) may not be included.
[0026] The display (110) is configured to display a video signal. The display (110) can display a single-view screen or a multi-view screen. For example, if the screen displayed on the display (110) is selected as a multi-view screen composed of multiple partial screens, the processor (130) can divide the screen displayed on the display (110) into multiple partial screens and display them, and display multiple contents on each partial screen.
[0027] The display (110) may be implemented as a display including a self-luminous element or a display including a non-luminous element and a backlight. For example, the display (110) may be implemented as various types of displays such as an LCD (Liquid Crystal Display), an OLED (Organic Light Emitting Diodes) display, an LED (Light Emitting Diodes), a micro LED, a Mini LED, a PDP (Plasma Display Panel), a QD (Quantum dot) display, a QLED (Quantum dot light-emitting diodes), etc.
[0028] The memory (120) can store at least one command, data, program, etc. required for the operation of the display device (100) or the processor (130). The memory (120) can store a neural network model. Specifically, the memory (120) can store information on a text area detected from an image displayed on the display (110). In this case, the information on the text area can include at least one of the size of the text area, the number of text areas, the position of the text area in the image, and the width / height ratio of the text area. The size of the text area can be a relative size with respect to the image displayed on the screen. The position of the text area can be a relative position with respect to the image displayed on the screen.
[0029] For example, the size of a text area can be set to the size of the minimum bounding rectangle (MBR) that encloses the text area in the image. In computer geometry, the minimum bounding rectangle (MBR) represents a bounding box (BBOX) used to represent the maximum extent of a two-dimensional object or set of objects within an xy coordinate system. The minimum rectangle can be specified using two coordinates for the minimum and maximum x and y coordinates.
[0030] The memory (120) may be implemented in the form of memory embedded in the display device (100) or in the form of memory detachable from the display device (100) depending on the purpose of data storage. For example, data for driving the display device (100) may be stored in a memory embedded in the display device (100), and data for the expansion function of the display device (100) may be stored in a memory detachable from the display device (100).
[0031] Meanwhile, the memory embedded in the display device (100) may be implemented as at least one of volatile memory (e.g., dynamic RAM (DRAM), static RAM (SRAM), or synchronous dynamic RAM (SDRAM)), non-volatile memory (e.g., one time programmable ROM (OTPROM), programmable ROM (PROM), erasable and programmable ROM (EPROM), electrically erasable and programmable ROM (EEPROM), mask ROM, flash ROM, flash memory (e.g., NAND flash or NOR flash), hard drive, or solid state drive (SSD)).
[0032] The memory (120) may be implemented as a single memory that stores data generated from various operations according to the present disclosure. However, according to another embodiment, the memory (130) may be implemented to include multiple memories that each store different types of data or each store data generated at different stages.
[0033] The processor (130) is a component that is connected to each component of the display device (100) and controls the overall operation of the display device (100). The processor (130) may be implemented as a digital signal processor (DSP), a microprocessor, a GPU (Graphics Processing Unit), an AI (Artificial Intelligence) processor, an NPU (Neural Processing Unit), etc. However, the present invention is not limited thereto, and may include one or more of a central processing unit (CPU), a MCU (Micro Controller Unit), an MPU (micro processing unit), a controller, an application processor (AP), a communication processor (CP), an ARM processor, or may be defined by the relevant terminology. In addition, the processor (130) may be implemented as a SoC (System on Chip), an LSI (Large Scale Integration) having a built-in processing algorithm, or may be implemented in the form of an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array).
[0034] In addition, a processor (130) for executing a neural network model according to one embodiment may be implemented through a combination of software and a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU, a VPU (Vision Processing Unit), or a neural network-only processor such as an NPU.
[0035] The processor (130) can be controlled to process input data according to predefined operating rules or neural network models stored in memory. Alternatively, if the processor (130) is a dedicated processor (or a neural network-dedicated processor), it can be designed with a hardware structure specialized for processing a specific neural network model. For example, hardware specialized for processing a specific neural network model can be designed as a hardware chip such as an ASIC or FPGA.
[0036] When the processor (130) is implemented as a dedicated processor, it may be implemented to include a memory for implementing the embodiments of the present disclosure, or it may be implemented to include a memory processing function for utilizing external memory. The processor (130) may be implemented as one or more processors.
[0037] Meanwhile, functions related to artificial intelligence according to the present disclosure may be operated through a processor (130) and a memory (120). One or more processors are controlled to process input data according to predefined operation rules or neural network models stored in the memory (120). Alternatively, if one or more processors are artificial intelligence-specific processors, the artificial intelligence-specific processors may be designed with a hardware structure specialized for processing a specific neural network model. The predefined operation rules or neural network models are characterized by being created through learning.
[0038] Here, "created through learning" means that a basic neural network model is trained using a learning algorithm using a large amount of learning data, thereby creating a predefined set of operating rules or a neural network model configured to perform a desired characteristic (or purpose). This learning may be performed on the device itself where the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0039] A neural network model can be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the computational results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers can be optimized based on the learning results of the neural network model. For example, during the learning process, the multiple weights can be updated so that the loss or cost values obtained from the neural network model are reduced or minimized.
[0040] The artificial neural network may include a Conditional Neural Process (CNP) or a Deep Neural Network (DNN), and examples thereof include, but are not limited to, a Convolutional Neural Network (CNN), a Deep Neural Network (DNN), a Recurrent Neural Network (RNN), a Restricted Boltzmann Machine (RBM), a Deep Belief Network (DBN), a Bidirectional Recurrent Deep Neural Network (BRDNN), a Generative Adversarial Network (GAN), a Latent Neural Process (LNP), Attentive Neural Processes (ANP), Residual Neural Processes (RNP), Doubly Stochastic Neural Processes (DSNP), Transformer Neural Processes (TNP), Memory-Augmented Neural Networks (MANN), Simple Neural Attentive Meta-Learner (SNAIL), or Deep Q-Networks.
[0041] The processor (130) detects at least one text area of an image displayed on the display (110), analyzes the at least one text area, and obtains information about the at least one text area. The processor (130) inputs information about the at least one text area into a neural network model to identify a subtitle area among the at least one text area. In this case, the neural network model is an artificial intelligence model trained to obtain information about a text area included in a captured image, and may include at least one weight value for identifying a subtitle area among the text areas.
[0042] The processor (130) can input at least one of the size of the text area, the number of text areas, the position of the text area, and the width / height ratio of the text area into the neural network model to obtain information on the probability that the text area corresponds to the subtitle area. The processor (130) can identify the subtitle area among at least one text area based on the information on the probability that each of the at least one text areas corresponds to the subtitle area.
[0043] For example, a neural network model can be trained such that a text region located at the bottom of an image is more likely to correspond to a subtitle region than a text region located at the top or center of the image. Furthermore, the neural network model can be trained such that, if the text region is located at the bottom of an image, a rectangle whose width is relatively longer than its height is more likely to correspond to a subtitle region than a rectangle or square whose height is relatively longer than its width.
[0044] The processor (130) can identify a text area having a probability greater than a threshold value among at least one text area as a subtitle area.
[0045] When a subtitle area is identified among text areas, the processor (130) performs character recognition on the text in the identified subtitle area to obtain subtitle data. For example, the processor (130) may perform character recognition on the text in the subtitle area using OCR (Optical Character Recognition).
[0046] Traditionally, OCR (optical character recognition) refers to a device that automatically recognizes and reads characters printed on receipts and other forms by shining light on them and detecting differences in the amount of reflected light, or the intensity of the light. OCR input methods used include direct input, which directly recognizes and reads handwritten postal codes or order receipts, and turnaround input, which re-enters the contents of receipts printed by a calculator, such as tax bills or electricity and gas bills, into a calculator.
[0047] However, recent OCR technologies incorporate AI technologies that automatically recognize text within images. Examples include recognizing license plate numbers through a camera, recognizing personal information text on ID cards, and detecting text in images and videos.
[0048] FIG. 2 is a drawing for explaining a detailed configuration of a display device according to various embodiments of the present disclosure.
[0049] According to FIG. 2, the display device (100) may include a display (110), a memory (120), a processor (130), a speaker (140), and an interface (150). Among the configurations illustrated in FIG. 2, a detailed description of configurations that overlap with those illustrated in FIG. 1 will be omitted.
[0050] The speaker (140) is a component for outputting an audio signal. In FIG. 2, only the speaker (140) is illustrated, but various components such as an audio decoder, an audio output mixer, an audio signal processor, and an amplifier circuit for processing an audio signal may be further included in the display device (100). The speaker (140) may be configured as one or more, and when implemented as multiple speakers, they may be arranged symmetrically on the exterior of the main body to output audio information in all directions, that is, in all 360 degrees. The processor (130) may output a voice signal through the speaker (140). Specifically, when a screen where a subtitle area is located is selected, the processor (130) may convert the acquired subtitle data into voice and output it through the speaker (140).
[0051] For example, the processor (130) can convert subtitle data into speech using TTS (Text to Speech). TTS refers to implementing a human voice through a computer program. The processor (130) can implement TTS by recording a human voice in advance, dividing it into certain sound units and storing it, and searching for and combining voices corresponding to sentences when text is input. Using TTS, sentences can be converted into speech to implement a human voice without a voice actor. TTS has the disadvantage of unnatural pronunciation and intonation, but has the advantage of outputting speech as soon as letters are input.
[0052] In this case, the processor (130) can output the original sound of a foreign language or foreign language video and the voice-converted subtitle data (hereinafter, voice subtitle signal) through a plurality of speakers (140). Alternatively, if the user selects an option to output only the voice subtitle signal from the display device (100), the processor (130) can mute the original sound output in the foreign language from the foreign language or foreign language video and output only the voice subtitle signal.
[0053] The interface (150) is configured to receive various data from a user or an external device. The interface (150) can receive image information displayed on the display (110). The interface (150) may include a communication interface, an operation interface, and an input / output interface.
[0054] For example, a communication interface is a configuration for performing communication with at least one external device. The communication interface may include at least one wireless communication module, at least one wired communication module, etc. Each communication module may be implemented in the form of at least one hardware chip. The wireless communication module may include at least one module among a Wi-Fi module, a Bluetooth module, an infrared communication module, or other communication modules. In addition, the communication interface may include at least one communication chip that performs communication according to various wireless communication standards such as Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), LTE-A (LTE Advanced), 4G (4th Generation), 5G (5th Generation), etc. The wired communication module may include, for example, at least one among a LAN (Local Area Network) module, an Ethernet module, a pair cable, a coaxial cable, a fiber optic cable, or a UWB (Ultra Wide-Band) module. The communication interface is implemented in various forms like this, and by performing communication with an external device, it is possible to receive video data containing content from the external device.
[0055] The operation interface is a configuration for receiving user input. The operation interface may include various buttons, a touch screen, etc. provided on the main body of the display device (100). Using the operation interface, the user can change the content screen or directly input various data related to multi-view screen settings, etc., into the display device (100).
[0056] The input / output interface is a configuration for inputting and outputting various external signals. The input / output interface can be connected to various external memories or external sources (e.g., web servers, user terminal devices, etc.) and can input various data. The input / output interface can be implemented as at least one interface among HDMI (High Definition Multimedia Interface), MHL (Mobile High-Definition Link), USB (Universal Serial Bus), USB C-type, DP (Display Port), Thunderbolt, VGA (Video Graphics Array) port, RGB port, D-SUB (Dsubminiature), and DVI (Digital Visual Interface). At least some of the input / output interfaces can be connected to a communication interface. For example, the input / output interface can transmit information received from an external device to the communication interface or transmit information received through the communication interface to the external device. The display device (100) can directly read or receive various image data stored in an external memory or external source connected through the input / output interface.
[0057] FIGS. 3 and 4 are diagrams illustrating the operation of a processor according to various embodiments of the present disclosure. Specifically, FIGS. 3 and 4 are diagrams illustrating the operation of the processor (130) according to image content received via the interface (150).
[0058] If the screen displayed on the display (110) is a multi-view screen composed of multiple partial screens, the processor (130) may determine whether to detect a text area for the image displayed on the display (110) based on the type of image information input through the interface (150). Specifically, the processor (130) may determine the analysis target screen for converting subtitle data into voice based on the type of image content input through the interface (150) or whether the character of the image content can be recognized.
[0059] If the video content input through the interface (150) does not include foreign language audio or subtitles, the processor (150) does not need to perform analysis to provide audio subtitles for the video content. In addition, among the video contents input to the display device (100), there may be content for which video analysis, such as character recognition, is restricted. For example, a specific OTT channel (Over-The-Top) provides subtitles for foreign language audio but does not allow video analysis of the video itself. If video contents for which video analysis is not permitted are input through the interface (150), the processor (130) may omit video analysis, such as an operation for detecting text areas in the video.
[0060] According to FIG. 3, when the screen displayed on the display (110) is a multi-view screen composed of two partial screens, a foreign video may be displayed on the first partial screen (310), and an OTT video may be displayed on the second partial screen (320). In the following description, the OTT video displayed on the second partial screen (320) is based on the case where the video analysis is not permitted.
[0061] In this case, the processor (130) may perform an image analysis process, such as detecting a text area for a foreign language video displayed on the first partial screen, identifying a subtitle area among the text areas, and performing character recognition. However, the processor (130) may not perform image analysis for an OTT video displayed on the second partial screen.
[0062] In this way, the display device (100) according to the present disclosure can reduce resource consumption such as a processor and memory for image analysis by determining whether to perform image analysis on an image displayed on the display (110) depending on the type of image content input through the interface (150).
[0063] FIG. 4 illustrates a case where a screen displayed on a display (110) is a multi-view screen composed of two partial screens, a first foreign language video is displayed on a first partial screen (410), and a second foreign language video is displayed on a second partial screen (420). In this case, since both the first foreign language video and the second foreign language video displayed on the display (110) are video contents that allow video analysis, the processor (130) must perform video analysis of the first foreign language video and the second foreign language video.
[0064] The processor (130) detects text areas for the first foreign language video and the second foreign language video, obtains information about the detected text areas, and inputs information about each detected text area into a neural network model to identify subtitle areas for the first foreign language video and the second foreign language video.
[0065] When the user selects the first partial screen (410) as the focus screen while performing image analysis on the first foreign language video and the second foreign language video, the processor (130) may convert the subtitle area of the first foreign language video into audio and output an audio subtitle signal for the first foreign language video through the speaker (140). When the user moves the focus screen from the first partial screen (410) to the second partial screen (420), the processor (130) may stop outputting the audio subtitle signal for the first foreign language video, convert the subtitle area of the second foreign language video on which image analysis was being performed into audio and output an audio subtitle signal for the second foreign language video through the speaker (140).
[0066] When the processor (130) performs image analysis on a plurality of partial screens and an image displayed on at least one of the plurality of partial screens is changed, the processor (130) can identify the continuity of the image for the changed partial screen and determine whether to analyze the image for the changed partial screen based on the continuity of the identified image.
[0067] Referring to FIG. 4, when a user selects a first partial screen (410) to watch a first foreign language video, and the second foreign language video displayed on a second partial screen (420) changes to an advertisement video, the processor (130) can identify the continuity of the video for the second partial screen (420) and determine whether to analyze the video for the second partial screen (420).
[0068] The method for determining whether to analyze an image by identifying the continuity of the image is described in detail through Figures 5 and 6 below.
[0069] Figures 5 and 6 are diagrams illustrating the continuity of images according to various embodiments of the present disclosure. Figure 5 illustrates a text area for a preset initial image, and Figure 6 illustrates a text area for a modified image. In this case, the initial image may be determined based on the start time at which the image content is displayed. For example, the processor (130) may determine an image displayed one minute after the start time at which the foreign image is displayed as the initial image.
[0070] Continuity of the image is used to determine whether the image displayed on the display (110) is an image of the same content as the previous image. If the image displayed on the display (110) is an image of the same content as the initial image or the previous image, the processor (130) can determine that there is continuity of the image. However, if the image displayed on the display (110) is an image of different content as compared to the initial image or the previous image, the processor (130) can determine that there is no continuity of the image.
[0071] For example, when a foreign video is displayed on the display (110), if the video is connected according to the progression of the video content, it can be identified that there is video continuity. However, if the foreign video is changed to an advertisement video while being displayed on the display (110), the processor (130) can identify that there is no video continuity because the content video has different content compared to the initial video or previous video.
[0072] If the screen displayed on the display (110) is a multi-view screen composed of multiple partial screens, the processor (130) can identify the continuity of images for the multiple partial screens based on information about the text area. If the processor (130) determines that there is image continuity for the changed image, the processor (130) can continue to perform image analysis for the changed image. For example, the processor (130) can input information about at least one text area detected in the changed image into a neural network model to identify a subtitle area among the at least one text area, and obtain subtitle data for the text in the subtitle area.
[0073] However, if the processor (130) determines that there is no video continuity for the changed video, it may stop the video analysis for the changed video. In this case, the processor (130) may perform the video analysis by extending the video analysis time for a preset delay time based on the time at which it is determined that there is no video continuity. For example, the delay time may be set to 1 to 10 minutes, which is the display time of an advertisement video inserted in the middle of a foreign video. However, the delay time is not limited thereto, and the delay time may be set in consideration of the user's selection or the resource aspects of the display device (100).
[0074] Considering the resource aspect of the display device (100), it is more efficient to continue performing image analysis than to stop and then restart the image analysis within a certain period of time, so the processor (130) can extend the image analysis time during the delay time even if it is determined that there is no continuity of the image.
[0075] The processor (130) may compare the similarity between a preset initial image for multiple partial screens and an image displayed in real time based on information about at least one text area detected in the image. For example, if the initial image of the image content displayed on the second partial screen on which the processor (130) performs image analysis is configured as in FIG. 5 and the image of the second partial screen is changed as in FIG. 6, the processor (130) may compare information about the text area of the initial image with information about the text area of the changed image to calculate the similarity.
[0076] The processor (130) determines that there is continuity in the image of a partial screen where the calculated similarity is greater than or equal to a preset reference value, and can identify a subtitle area among at least one text area for the partial screen where the continuity of the image is determined. However, if the similarity of the changed image is identified as being less than or equal to the reference value compared to the initial image as shown in FIG. 6, the processor (130) can determine that there is no continuity in the image. If the processor (130) determines that there is no continuity in the image, it can perform image analysis during a delay time. If the similarity of the changed image is less than or equal to the reference value even after the delay time, the processor (130) can stop image analysis on the changed image.
[0077] FIG. 7 is a diagram illustrating advertisement identification factors according to various embodiments of the present disclosure. In FIG. 7, the left diagram represents an advertisement video, and the right diagram represents video content provided by a broadcaster.
[0078] Because visually impaired users have difficulty utilizing the benefits of the multiview feature when using a multiview screen, the usage rate of the multiview screen for the visually impaired is low. For example, when a non-visually impaired user is watching a video on the first screen using the multiview screen, if an advertisement is displayed in the middle of the video being watched, the non-visual user can move the focus screen to the second screen of the multiview screen that does not display the advertisement and watch a different video. Furthermore, when the advertisement video ends on the previous first screen, the non-visual user can visually confirm the end of the advertisement on the first screen and move the focus screen from the second screen to the first screen.
[0079] However, when a visually impaired person is watching a video using a multi-view screen and an advertisement is displayed in the middle of the video being watched, the visually impaired person cannot identify whether the video is an advertisement or not, and therefore cannot move the focus screen to another part of the screen that does not display the advertisement, which may cause confusion in the content of the video being watched. Even if the visually impaired person recognizes that an advertisement is displayed in the middle of the video being watched on the multi-view screen, if the focus screen is moved to another part of the screen, the state of the previous part of the screen cannot be confirmed, and thus the focus screen cannot be moved to another part of the screen. For example, if the visually impaired person moves the focus screen from the first part of the screen being watched to the second part of the screen, the person can only hear the sound of the video displayed in the second part of the screen, and therefore the state of the first part of the screen cannot be confirmed.
[0080] According to FIG. 7, the advertisement video displayed on the first partial screen (710) may not display the broadcaster logo (721), but may display an advertisement mark (711) indicating that it is an advertisement broadcast or an advertisement subtitle indicating that it is an advertisement. In addition, the video content provided by the broadcaster, such as the second partial screen (720), displays the broadcaster logo (721) of the broadcaster providing the video content.
[0081] If the screen displayed on the display (110) is a multi-view screen composed of multiple partial screens, the processor (130) can detect text from multiple images displayed on the multiple partial screens. For example, the processor (130) can detect text from the images using an optical character recognition (OCR) technique. The processor (130) can detect text from the images by periodically performing OCR (Optical Character Recognition), which recognizes characters from the images displayed on the display (110).
[0082] The processor (130) can acquire advertisement identification factors included in multiple partial screens based on the detected text. In this case, the advertisement identification factors may include advertisement marks (711) appearing in commercial broadcasts, a broadcaster's logo (721), a watermark provided by the producer of the video content, etc.
[0083] A watermark is a logo or text inserted into documents, photos, or other creative files to indicate the copyright of the original image without damaging it. The term watermark originates from the ciphers used in medieval churches. When sending encrypted messages, the church would insert transparent text or images into them. This method of insertion made it possible to read only when the paper was soaked in water or under a light, hence the term "watermark."
[0084] However, with the recent advancement of the Internet, watermarks primarily represent digital watermarks used in digital environments. For example, video producers insert watermarks into video content to prevent redistribution, while television broadcasters insert channel names, broadcaster logos, and program titles into the corners of their video content.
[0085] The processor (130) can identify a partial screen on which an advertisement image is displayed among multiple partial screens based on the position of the acquired advertisement identification factor. For example, referring to FIG. 7, it can be confirmed that an advertisement mark (711) appears on a first partial screen (710) on which an advertisement image is displayed. If the processor (130) acquires the advertisement mark (711) from the image, it can identify the image of the partial screen on which the advertisement mark (711) appears as an advertisement image. In addition, if the processor (130) acquires the broadcasting station logo (721) from the image, it can identify the image of the partial screen on which the broadcasting station logo (721) appears as video content rather than an advertisement image.
[0086] According to the Broadcasting Act of the Korea Communications Commission, broadcasting programs provided by broadcasters must be distinguished from commercials, and only video content produced by broadcasters may display the broadcaster's logo. Therefore, broadcaster logos cannot be displayed on commercials.
[0087] Accordingly, the processor (130) can identify whether an advertisement is broadcast for each video by using advertisement identification factors such as an advertisement display, a broadcaster's logo, a watermark provided by the producer of the video content, and can identify a partial screen on which an advertisement video is displayed among a plurality of partial screens.
[0088] When an advertisement video is displayed on a focus screen selected by the user among a plurality of partial screens, the processor (130) can move the focus screen to another partial screen among the plurality of partial screens where the advertisement video is not displayed. For example, when a visually impaired person watches a video using a multi-view screen, if an advertisement video is displayed on the first partial screen being watched, the processor (130) can automatically move the focus screen to another partial screen among the plurality of partial screens where the advertisement video is not displayed. In addition, when the advertisement video on the first partial screen ends and the video being watched by the user is displayed again on the first partial screen, the processor (130) can move the focus screen to the first partial screen.
[0089] When an advertisement video is displayed on a focus screen selected from among multiple partial screens, the processor (130) can output voice guidance information regarding whether or not an advertisement will be broadcast through the speaker (140). If the user does not want voice guidance information regarding whether or not an advertisement will be broadcast, the user can select not to provide voice guidance information using a user option of the display device (100). In this case, when an advertisement video is displayed on a focus screen selected from among multiple partial screens, the processor (130) can maintain the current screen state and display the advertisement video as is.
[0090] FIGS. 8 to 10 are flowcharts for explaining a method of controlling a display device according to various embodiments of the present disclosure.
[0091] According to FIG. 8, the display device detects at least one text area of an image displayed on the display device (S810). The text area can be detected as a minimum bounding rectangle (MBR) that surrounds the text displayed on the image. The minimum bounding rectangle (MBR) can be represented as a bounding box (BBOX) specified using two coordinates for the minimum and maximum of the x and y coordinates. For example, the text area can be represented as a bounding box (BBOX) specified as A (minimum x coordinate, minimum y coordinate) and B (maximum x coordinate, maximum y coordinate).
[0092] The display device analyzes at least one text area to obtain information about at least one text area (S820). The information about the text area can be obtained based on a minimum rectangle (MBR) enclosing the text detected from the image. For example, if each text area is detected as a minimum rectangle (MBR) enclosing the text, the information about the text area can be obtained using the size of the minimum rectangle (MBR), the number of the minimum rectangles (MBR), the relative position of the minimum rectangle (MBR) in the image, and the width / height ratio of the minimum rectangle (MBR).
[0093] A display device inputs information about at least one text region acquired into a neural network model to identify a subtitle region among the at least one text region (S830). The neural network model may be a model trained using a plurality of previously collected video contents. The neural network model may include a plurality of weights for calculating information about the text region. When information about a text region of an image is acquired, the display device may input the information about at least one text region acquired into the neural network model to calculate a weight of the information about the text region. The display device may identify a subtitle region among the at least one text region based on the calculated weight result.
[0094] The display device performs character recognition on text in a subtitle area to obtain subtitle data (S840). If a subtitle area is identified among at least one text area, the display device can perform character recognition on text in the identified subtitle area. For example, the display device can perform character recognition on text in the subtitle area using OCR (Optical Character Recognition) and obtain subtitle data on text in the subtitle area displayed on the image.
[0095] FIG. 9 illustrates a method for identifying a subtitle area among at least one text area using information about at least one text area acquired by a display device.
[0096] According to FIG. 9, the display device can identify whether the screen displayed on the display device is a multi-view screen composed of multiple partial screens (S910).
[0097] If the screen displayed on the display device is a multi-view screen, the display device can identify the continuity of images for multiple partial screens based on information about the text area (S920). The display device can compare the similarity between the displayed image and a preset initial image for multiple partial screens in real time based on the information about at least one text area acquired. The display device can determine that the image of a partial screen having a similarity greater than a preset reference value has continuity of images.
[0098] The display device can identify a subtitle area among at least one text area based on the continuity of the image (S930). If the display device determines that the image has continuity, the display device can continue the analysis process for the image being analyzed. For example, the display device can identify a subtitle area for a subscreen among multiple subscreens in which the continuity of the image has been identified, and obtain subtitle data for the text in the subtitle area.
[0099] Additionally, the display device can output subtitle data by converting it into audio when a screen where a subtitle area is located is selected. For example, when the screen displayed on the display device is a multi-view screen, if a sub-screen where a subtitle area is located among multiple sub-screens is selected as the focus screen, the display device can convert the subtitle data of the selected sub-screen into audio and output it through the speaker.
[0100] Figure 10 illustrates a method for a display device to identify a screen on which an advertising image is displayed.
[0101] According to FIG. 10, the display device can identify whether the screen displayed on the display device is a multi-view screen composed of multiple partial screens (S1010).
[0102] If the screen displayed on the display device is a multi-view screen, the display device can detect text in multiple images displayed on multiple partial screens (S1020). For example, the display device can detect text in the images using an optical character recognition (OCR) technique.
[0103] The display device can acquire advertisement identification factors included in multiple partial screens based on the detected text (S1030). Advertisement identification factors may include advertisement signs appearing in commercial broadcasts, broadcaster logos, watermarks provided by video content producers, and the like.
[0104] The display device can identify a sub-screen where an advertisement video is played among multiple sub-screens based on the position of the acquired advertisement identification factor (S1040). For example, if an advertisement appears on the displayed video, the display device can identify the video with the advertisement as an advertisement video. However, if a broadcasting company logo appears on the displayed video, the display device can identify the video with the broadcasting company logo as video content, not an advertisement video.
[0105] The display device can identify whether an advertisement video is being played on a focus screen selected from among multiple sub-screens based on the result of identifying the sub-screen where the advertisement video is being played. If an advertisement video is displayed on the focus screen selected by the user, the display device can move the focus screen to another sub-screen among the multiple sub-screens where the advertisement video is not being played.
[0106] When an advertisement video is displayed on a focus screen selected by the user among multiple partial screens, the display device can output audio guidance information regarding whether the advertisement is being broadcast through the speaker. If the user does not want to receive audio guidance regarding whether the advertisement is being broadcast, the user can select not to provide audio guidance using the display device's user options. In this case, when an advertisement video is displayed on a focus screen selected among multiple partial screens, the display device can maintain the current screen state and display the advertisement video as is.
[0107] As such, the display device and its control method according to various embodiments of the present disclosure can provide an audible subtitle function for the visually impaired, thereby improving their understanding of the video when watching a foreign language video or a non-dubbed video. Furthermore, when the visually impaired person watches a video using a multi-view screen, the screen on which an advertisement video is displayed can be identified to provide guidance information, or the focus screen can be moved to another part of the screen where the advertisement video is not displayed.
[0108] Through this, the display device and the control method thereof according to various embodiments of the present disclosure can prevent confusion in the content of the video being viewed even if an advertisement is displayed in the middle of the video being viewed when the visually impaired person watches the video using a multi-view screen.
[0109] Meanwhile, according to an embodiment of the present disclosure, the various embodiments described above may be implemented as software including commands stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call commands stored from the storage medium and operate according to the called commands, and may include a display device according to the disclosed embodiments. When a command is executed by a processor, the processor can perform a function corresponding to the command directly or under the control of the processor by using other components. The command may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' means that the storage medium does not contain a signal and is tangible, but does not distinguish between data being stored semi-permanently or temporarily in the storage medium.
[0110] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or online through an application store (e.g., Play Store™). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0111] In addition, each of the components (e.g., modules or programs) according to the various embodiments described above may be composed of a single or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into a single entity, which may perform the same or similar functions as those performed by each of the corresponding components prior to integration. Operations performed by modules, programs or other components according to various embodiments may be executed sequentially, in parallel, iteratively or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.
[0112] Although the preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present disclosure.
Claims
1. In the display device, display; Memory for storing neural network models; Processor; including; The above processor, Detecting at least one text area of an image displayed on the display, analyzing at least one text area to obtain information about the at least one text area, Inputting information about at least one text area into the neural network model to identify a subtitle area among the at least one text area, A display device that obtains subtitle data by performing character recognition on text in the above subtitle area.
2. In paragraph 1, The above processor, If the screen displayed on the above display is a multi-view screen composed of multiple partial screens, the continuity of the image for the multiple partial screens is identified based on the information about the text area, A display device that identifies a subtitle area among at least one text area based on the continuity of the image.
3. In paragraph 2, The above processor, Calculating the similarity between a preset initial image for a plurality of partial screens and an image displayed in real time based on information about at least one text area, A display device that identifies a subtitle area among at least one text area in an image of a partial screen having a similarity level greater than a preset reference value.
4. In paragraph 1, further comprising a speaker for outputting a voice signal; The above processor, A display device that converts the subtitle data into voice and outputs it through the speaker when the screen where the subtitle area is located is selected.
5. In paragraph 1, Information about the above text area is: A display device comprising at least one of the size, number, position, and width / height ratio of the above text area.
6. In paragraph 1, Further comprising an interface for receiving image information displayed on the display; The above processor, A display device that determines whether to detect the text area of the image displayed on the display based on the type of image information input through the interface, when the screen displayed on the display is a multi-view screen composed of a plurality of partial screens.
7. In paragraph 1, The above processor, If the screen displayed on the above display is a multi-view screen composed of multiple partial screens, detecting text of multiple images displayed on the multiple partial screens, Obtaining an advertisement identification factor included in the plurality of partial screens based on the detected text; A display device that identifies a partial screen on which an advertisement image is displayed among the plurality of partial screens based on the position of the advertisement identification factor.
8. In paragraph 7, The above processor, A display device that, when an advertisement image is displayed on a focus screen selected from among the plurality of partial screens, moves the focus screen to another partial screen from among the plurality of partial screens where an advertisement image is not displayed.
9. In paragraph 7, further comprising a speaker for outputting a voice signal; The above processor, A display device that outputs voice guidance information regarding whether an advertisement is broadcast through the speaker when an advertisement video is displayed on a focus screen selected from among the plurality of partial screens.
10. In a method for controlling a display device, A step of detecting at least one text area of an image displayed on the display device; A step of analyzing at least one text area to obtain information about the at least one text area; A step of inputting information about at least one text region obtained above into a neural network model to identify a subtitle region among the at least one text region; and A control method comprising: a step of obtaining subtitle data by performing character recognition on text in the subtitle area; 11. In paragraph 10, The step of identifying a subtitle area among at least one text area above comprises: A step of identifying whether the screen displayed on the display device is a multi-view screen composed of multiple partial screens; If the screen displayed on the display device is a multi-view screen, a step of identifying the continuity of images for a plurality of partial screens based on information about the text area; and A control method comprising: a step of identifying a subtitle area among the at least one text area based on the continuity of the image; 12. In paragraph 10, A control method further comprising: a step of converting the subtitle data into voice and outputting it when the screen on which the subtitle area is located is selected.
13. In paragraph 10, A step of identifying whether the screen displayed on the display device is a multi-view screen composed of multiple partial screens; If the screen displayed on the display device is a multi-view screen, a step of detecting text of a plurality of images displayed on the plurality of partial screens; A step of obtaining an advertisement identification factor included in the plurality of partial screens based on the detected text; and A control method further comprising: a step of identifying a partial screen on which an advertisement image is displayed among the plurality of partial screens based on the position of the advertisement identification factor.
14. In paragraph 13, A step of identifying whether an advertisement video is displayed on a selected focus screen among the plurality of partial screens; and A control method further comprising: when an advertisement image is displayed on the selected focus screen, a step of moving the focus screen to another partial screen among the plurality of partial screens in which the advertisement image is not displayed.
15. In paragraph 10, Information about the above text area is: A control method comprising at least one of the size, number, position, and ratio of width / height of the above text area.
Citation Information
Patent Citations
Telop information processor and telop information display device
JP2001285716A
Advertisement broadcasting avoidance watching control device in a multi-screen television, especially for reproducing a broadcasting channel on a sub screen if it is detected that advertisement broadcasting is reproduced, and reprod ...
KR1019990050596A
Roadcasting receiving apparatus and control method thereof
KR1020160104493A
Plasma apparatus and reformer system including the same
KR1020200131924A
Apparatus for diagnosing battery and operating method thereof
KR1020250107447A