Display apparatus and control method thereof

The display device analyzes image frames to identify object changes and adjusts voice output, allowing visually impaired users to comprehend image modifications effectively.

WO2025170199A1PCT designated stage Publication Date: 2025-08-14SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/096767
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-08
Filing Date
2024-12-12
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Visually impaired users face difficulty in recognizing changes in images displayed on display devices, making it challenging to utilize these devices effectively.

Method used

A display device and control method that analyze image frames to identify changes in objects within the image by determining the number, distance, and type of objects, and adjust voice output accordingly to inform users of these changes.

Benefits of technology

Enables visually impaired individuals to detect changes in images through voice feedback, enhancing their ability to understand and interact with displayed content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096767_14082025_PF_FP_ABST
    Figure KR2024096767_14082025_PF_FP_ABST
Patent Text Reader

Abstract

This display apparatus comprises: a display unit on which a video including a first video frame and a second video frame is displayed; and at least one processor for acquiring first video frame object information that includes at least one from among the number of objects included in the first video frame, the distance between the objects, and the types of the objects, acquiring second video frame object information including at least one from among the number of objects included in the second video frame, the distance between the objects, and the types of the objects, and recognizing, on the basis of the difference between the first video frame object information and the second video frame object information, whether the object displayed on the display unit has changed.
Need to check novelty before this filing date? Find Prior Art

Description

Display device and method of controlling the display device

[0001] The present disclosure relates to a display device and a method for controlling the display device.

[0002] In general, a display device is a type of output device that converts acquired or stored electrical information into visual information and displays it to the user, and is used in various fields such as homes and businesses.

[0003] Display devices include monitor devices connected to personal computers or server computers, portable computer devices, navigation terminal devices, general television devices, Internet Protocol television (IPTV) devices, portable terminal devices such as smart phones, tablet PCs, personal digital assistants (PDAs), or cellular phones, various display devices used to play images such as advertisements or movies in industrial settings, transparent displays, projectors, or various types of audio / video systems.

[0004] The display device includes a light source module for converting electrical information into visual information, and the light source module includes a plurality of light sources for independently emitting light. Each of the plurality of light sources includes, for example, a light emitting diode (LED), an organic light emitting diode (OLED), a micro LED, a mini LED, or a QD LED.

[0005] Hearing-impaired users of display devices can understand the content of a video through subtitles displayed on the device. However, visually impaired users of display devices often have difficulty visually recognizing the video displayed on the device, making it difficult to use.

[0006] The present invention can provide a display device and a control method of the display device that can recognize that an object in an image changes according to object information displayed in the image of the display device.

[0007] A display device according to the invention may include a display unit on which an image including a first image frame and a second image frame is displayed; and at least one processor that obtains first image frame object information including at least one of the number of objects included in the first image frame, the distance between the objects, and the type of the objects, obtains second image frame object information including at least one of the number of objects included in the second image frame, the distance between the objects, and the type of the objects, and recognizes whether an object displayed on the display unit has changed based on a difference between the first image frame object information and the second image frame object information.

[0008] A method for controlling a display device according to the invention includes a display unit on which an image including a first image frame and a second image frame is displayed, the method comprising: obtaining first image frame object information including at least one of the number of objects included in the first image frame, the distance between the objects, and the type of the objects; obtaining second image frame object information including at least one of the number of objects included in the second image frame, the distance between the objects, and the type of the objects; and recognizing whether an object displayed on the display unit has changed based on a difference between the first image frame object information and the second image frame object information.

[0009] According to one aspect of the disclosed invention, it is possible to determine that a situation within an image has changed by recognizing that an object within an image has changed according to object information displayed within an image of a display device.

[0010] According to one aspect of the disclosed invention, a user of a display device can easily recognize that a situation within an image has changed by changing a voice based on recognizing that an object within an image of the display has changed.

[0011] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.

[0012] FIG. 1 illustrates a display device according to one embodiment.

[0013] FIG. 2 is an exploded view of a display device according to one embodiment.

[0014] Figure 3 is a control block diagram of a display device according to one embodiment.

[0015] FIG. 4 is a diagram for explaining a method of obtaining an image frame from an image according to one embodiment.

[0016] FIGS. 5 and 6 illustrate image frames obtained from an image, according to one embodiment.

[0017] Fig. 7 illustrates a flowchart of a method for controlling a display device according to one embodiment.

[0018] FIG. 8 shows a table regarding weights assigned according to the type of object in an image frame according to one embodiment.

[0019] FIG. 9 illustrates an image frame obtained from an image, according to one embodiment.

[0020] It should be understood that the various embodiments and terms used in this document are not intended to limit the technical features described in this document to specific embodiments, but rather to include various modifications, equivalents, or substitutes of the embodiments.

[0021] In connection with the description of the drawings, similar reference numerals may be used for similar or related components.

[0022] The singular form of a noun corresponding to an item may include one or more of said items, unless the relevant context clearly indicates otherwise.

[0023] In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in that phrase, or all possible combinations thereof.

[0024] The term “and / or” includes any combination of a plurality of related described elements or any one of a plurality of related described elements.

[0025] The terms "part," "module," and "member" may be implemented in hardware or software. Depending on the embodiments, multiple "parts," "modules," or "members" may be implemented as a single component, or a single "part," "module," or "member" may include multiple components.

[0026] Terms such as "first," "second," or "first" or "second" may be used simply to distinguish one component from another and do not qualify the components in any other respect (e.g., importance or order).

[0027] When a component (e.g., a first component) is referred to as being "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0028] The terms “include” or “have” are intended to specify the presence of a feature, number, step, operation, component, part or combination thereof described in this document, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts or combinations thereof.

[0029] When a component is said to be “connected,” “coupled,” “supported,” or “in contact with” another component, this includes not only cases where the components are directly connected, coupled, supported, or in contact, but also cases where the components are indirectly connected, coupled, supported, or in contact through a third component.

[0030] When we say that a component is “on” another component, this includes not only cases where the component is in contact with the other component, but also cases where there is another component between the two components.

[0031] Meanwhile, the terms “front,” “back,” “left,” “right,” “upper,” and “lower” used in the following description are defined based on the drawing, but the shape and position of each component are not limited by the above terms. For example, the front side may be defined as the +X side, and the rear side may be defined as the -X side. For example, based on the drawing, the right side may be defined as the +Y side, and the left side may be defined as the -Y side. For example, based on the drawing, the upper side may be defined as the +Z side, and the lower side may be defined as the -Z side.

[0032] Hereinafter, embodiments according to the present invention will be described in detail with reference to the attached drawings.

[0033] FIG. 1 illustrates a display device according to one embodiment.

[0034] Referring to FIG. 1, the display device (1) may include a main body (11) and / or a display area (12) on which an image (I) is displayed.

[0035] The main body (11) forms the outer shape of the display device (1), and components for displaying an image (I) or performing various functions may be provided inside the main body (11). The main body (11) illustrated in Fig. 1 has a flat plate shape, but the shape of the main body (11) is not limited to that illustrated in Fig. 1. For example, the main body (11) may have a curved plate shape.

[0036] The display area (12) is formed on the front of the main body (11) and can display an image (I). For example, the display area (12) can display a still image or a moving image. In addition, the display area (12) can display a two-dimensional flat image or a three-dimensional stereoscopic image utilizing the parallax of the user's two eyes.

[0037] FIG. 2 is an exploded view of a display device according to one embodiment.

[0038] As shown in Fig. 2, various components for generating an image (I) in a display area (12) may be provided inside the main body (11).

[0039] For example, the main body (11) is provided with a light source device (100) which is a surface light source, a liquid crystal panel (20) which blocks or passes light emitted from the light source device (100), a control assembly (50) which controls the operation of the light source device (100) and / or the liquid crystal panel (20), or a power assembly (60) which supplies power to the light source device (100) and / or the liquid crystal panel (20). In addition, the main body (11) may include a bezel (13) for supporting and fixing the liquid crystal panel (20), the light source device (100), the control assembly (50) and / or the power assembly (60), a frame middle mold (14), a bottom chassis (15), and a rear cover (16).

[0040] The light source device (100) may include a point light source that emits monochromatic light or white light, and may refract, reflect, and / or scatter light to convert light emitted from the point light source into uniform surface light. In this way, the light source device (100) may emit uniform surface light toward the front by refracting, reflecting, and / or scattering light emitted from the point light source.

[0041] The liquid crystal panel (20) can be provided in front of the light source device (100) and can block or allow light emitted from the light source device (100) to pass through in order to form an image (I).

[0042] The front surface of the liquid crystal panel (20) can form a display area (12) of the display device (1) described above.

[0043] The display unit (500, see FIG. 3) may include a display area (12) and / or various components for generating an image (I) in the display area (12).

[0044] Displaying an image (I) in the display area (12) may include displaying an image (I) in the display unit (500). Hereinafter, displaying an image (I) in the display area (12) is described as displaying an image (I) in the display unit (500).

[0045] Figure 3 is a control block diagram of a display device according to one embodiment.

[0046] Referring to FIG. 3, the display device (1) may include a communication interface (300), a content receiving unit (80), a control unit (90), an input interface (400), a display unit (500), and / or a speaker (600).

[0047] The communication interface (300) may include at least one of a short-range communication module or a long-range communication module.

[0048] The communication interface (300) can transmit data to an external device or receive data from an external device. For example, the communication interface (300) can establish communication with a server, an external device (e.g., a keyboard, a mouse, a remote control, a touchscreen, a touchpad, a VR controller (Virtual Reality Controller), a user device), and / or other home appliances, and transmit and receive various data.

[0049] User devices may include, but are not limited to, personal computers, terminals, mobile phones, smart phones, handheld devices, wearable devices, display devices, etc.

[0050] To this end, the communication interface (300) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between external devices, and the performance of communication through the established communication channel. According to one embodiment, the communication interface (300) can include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with the external device through a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These different types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips).

[0051] The short-range wireless communication module may include, but is not limited to, a Bluetooth communication module, a BLE (Bluetooth Low Energy) communication module, a near field communication module, a WLAN (Wi-Fi) communication module, a Zigbee communication module, an infrared (IrDA, infrared Data Association) communication module, a WFD (Wi-Fi Direct) communication module, an UWB (ultrawideband) communication module, an Ant+ communication module, a microwave (uWave) communication module, etc.

[0052] The remote communication module may include a communication module that performs various types of remote communication and may include a mobile communication unit. The mobile communication unit may transmit and receive wireless signals with at least one of a base station, an external terminal, and a server on a mobile communication network.

[0053] In one embodiment, the communication interface (300) can communicate with external devices such as servers, user devices, and other home appliances through a surrounding access point (AP).

[0054] The communication interface (300) can transmit data received from a server, external device, and / or other home appliance to the control unit (90).

[0055] The content receiving unit (80) may include a receiving terminal (81) and a tuner (82) that receive content including video signals and / or audio signals from content sources.

[0056] The receiving terminal (81) can receive video signals and audio signals from content sources via a cable. For example, the receiving terminal (81) can include a component (YPbPr / RGB) terminal, a composite video blanking and sync (CVBS) terminal, an audio terminal, a High Definition Multimedia Interface (HDMI) terminal, a Universal Serial Bus (USB) terminal, etc.

[0057] The tuner (82) can receive broadcast signals from a broadcast reception antenna or a wired cable. In addition, the tuner (82) can extract broadcast signals of a channel selected by the user from among the broadcast signals. For example, the tuner (82) can pass broadcast signals having a frequency corresponding to the channel selected by the user among a plurality of broadcast signals received through a broadcast reception antenna or a wired cable, and block broadcast signals having other frequencies.

[0058] In this way, the content receiving unit (80) can receive video signals and audio signals from content sources through the receiving terminal (81) and / or the tuner (82). The content receiving unit (80) can transmit the video signals and / or audio signals received through the receiving terminal (81) and / or the tuner (82) to the control unit (90).

[0059] The control unit (90) can generate image data from a video signal received from the content receiving unit (80). The control unit (90) can control the display unit (500) to output an image (I) corresponding to the image data.

[0060] The control unit (90) can generate voice data from an audio signal received from the content receiving unit (80). The control unit (90) can control the speaker (600) to output voice information corresponding to the voice data.

[0061] The control unit (90) may include hardware such as at least one processor (91) and at least one memory (92).

[0062] At least one memory (92) can store data in the form of an algorithm or program for controlling the operation of components within the display device (1). In addition, at least one memory (92) can store an application program. For example, at least one memory (92) can store an application program for the display unit (500) to output a predetermined image (I) or for the speaker (600) to output a predetermined sound.

[0063] At least one memory (92) can store an algorithm capable of obtaining an image frame from an image (I).

[0064] The memory (92) and the processor (91) may each be implemented as separate chips. The processor (91) may include one or more processor chips or one or more processing cores. The memory (92) may include one or more memory chips or one or more memory blocks. In addition, the memory (92) and the processor (91) may also be implemented as a single chip.

[0065] At least one processor (91) may include an image processor such as a CPU or a graphics card. At least one processor (91) may perform the operations described above and the operations described below using data stored in at least one memory (92).

[0066] For example, at least one processor (91) can execute an application program stored in a memory (92) to cause the display unit (500) to output a predetermined image (I) or the speaker (600) to output a predetermined sound.

[0067] At least one processor (91) can acquire predetermined information using a learned machine learning model. The learned machine learning model may include one learned using multiple video frames, multiple images, etc. as learning data. The learned machine learning model may be implemented using various machine learning networks, such as an artificial neural network (ANN) or a convolutional neural network (CNN).

[0068] At least one processor (91) is electrically connected to various components mounted on the communication interface (300), the input interface (400), the display unit (500), the speaker (600), and / or the display device (1), and can control various components mounted on the communication interface (300), the input interface (400), the display unit (500), the speaker (600), and / or the display device (1).

[0069] The display unit (500) can display an image (I) including a plurality of image frames.

[0070] The display unit (500) may include at least one display. The at least one display may be a light emitting diode (LED) panel, an organic light emitting diode (OLED) panel, and / or a liquid crystal display (LCD) panel and / or an indicator. The display unit (500) may also include a touch screen.

[0071] The input interface (400) may include buttons (e.g., push buttons, touch buttons), dials, touchpads, touch screens, and / or microphones provided at various locations on the display device (1).

[0072] The input interface (400) can receive a user's input to control the image (I) output to the display unit (500).

[0073] The speaker (600) can output sound based on an audio output signal received from the control unit (90).

[0074] The speaker (600) may be provided in various locations of the display device (1). For example, the speaker (600) may be provided inside the display device (1) or in the main body (11).

[0075] The speaker (600) can be connected to the display device (1) via a wire (e.g., USB connection or High Definition Multimedia Interface; HDMI) or wirelessly (e.g., WIFI, Bluetooth connection).

[0076] FIG. 4 is a diagram illustrating a method for obtaining an image frame from an image according to one embodiment. FIGS. 5 and 6 illustrate image frames obtained from an image according to one embodiment.

[0077] Referring to FIG. 4, at least one processor (91) can obtain multiple image frames from the image (I). For example, at least one processor (91) can obtain multiple image frames from the image (I) using an algorithm stored in the memory (92).

[0078] The image (I) may include multiple image frames. For example, the image (I) may include a first image frame (F1), a second image frame (F2), and / or a third image frame (F3). The number of image frames according to the present disclosure is not limited thereto, and the number of image frames may vary according to various embodiments.

[0079] The display of the image (I) on the display unit (500) may include a plurality of image frames being displayed sequentially on the display unit (500).

[0080] For example, displaying an image (I) on a display unit (500) may include sequentially displaying a first image frame (F1), a second image frame (F2), and a third image frame (F3) on the display unit (500).

[0081] Referring to FIGS. 5 and 6, in one embodiment, at least one processor (91) can determine an area occupied by an object included in an image frame.

[0082] For example, at least one processor (91) can input a first image frame (F1) into a learned machine learning model to determine an area (A1) occupied by a first object (O2) included in the first image frame (F1), an area (A2) occupied by a second object (O2) included in the first image frame (F1), and an area (A3) occupied by a third object (O3) included in the first image frame (F1).

[0083] For another example, at least one processor (91) can determine an area (A4) occupied by a fourth object (O11) included in a second image frame (F2), an area (A5) occupied by a fifth object (O12) included in a second image frame (F2), an area (A6) occupied by a sixth object (O41) included in a second image frame (F2), and / or an area (A7) occupied by a seventh object (O42) included in a second image frame (F2).

[0084] In one embodiment, at least one processor (91) can determine the type of object included in the image frame.

[0085] For example, at least one processor (91) may input a first image frame (F1) into a learned machine learning model to determine that a first object (O1) is a person, a second object (O2) is a spoon, and a third object (O3) is a bowl.

[0086] As another example, at least one processor (91) may input the second image frame (F2) into a learned machine learning model to determine that the fourth object (O11) is a male person, the fifth object (O12) is a female person, and the sixth object (041) and the seventh object (042) are cups.

[0087] In one embodiment, at least one processor (91) can determine the distance between objects based on the area occupied by the objects included in the image frame.

[0088] For example, at least one processor (91) can determine the distance (d2) between the first object (O1) and the second object (O2) as the distance from the center of the area (A1) occupied by the first object (O1) to the center of the area (A2) occupied by the second object (O2), can determine the distance (d1) between the first object (O1) and the third object (O3) as the distance from the center of the area (A1) occupied by the first object (O2) to the center of the area (A3) occupied by the third object (O3), and can determine the distance (d3) between the second object (O2) and the third object (O3) as the distance from the center of the area (A2) occupied by the second object (O2) to the center of the area (A3) occupied by the third object (O3).

[0089] For another example, the distance (d4) between the fourth object (O11) and the fifth object (O12) can be determined as the distance from the center of the area (A5) occupied by the fourth object (O11) to the center of the area (A6) occupied by the fifth object (O12).

[0090] In one embodiment, at least one processor (91) can obtain subtitle data from subtitles included in an image (I).

[0091] For example, at least one processor (91) can obtain subtitle data from a first subtitle (S1) included in a first image frame (F1).

[0092] As another example, at least one processor (91) can obtain subtitle data from a second subtitle (S2) included in a second image frame (F2).

[0093] In one embodiment, at least one processor (91) can convert subtitle data into text data and obtain text information based on the text data.

[0094] For example, at least one processor (91) can obtain subtitle data from the first subtitle (S1), convert the obtained subtitle data into text data, and obtain text information (e.g., “It looks delicious”) corresponding to the first subtitle (S1) based on the text data.

[0095] As another example, at least one processor (91) can obtain subtitle data from the second subtitle (S2), convert the obtained subtitle data into text data, and obtain text information corresponding to the second subtitle (S2) (e.g., “Let’s go out in 10 minutes”) based on the text data.

[0096] At least one processor (91) can perform character recognition (Optical Character Reader; OCR) on subtitles included in a video frame to obtain subtitle data.

[0097] At least one processor (91) can separately obtain subtitle data from the image (I) if the subtitle included in the image frame is a closed subtitle.

[0098] In one embodiment, at least one processor (91) can obtain voice information based on the acquired subtitle data and control the speaker (600) to output the voice information.

[0099] For example, at least one processor (91) can obtain subtitle data from a first subtitle (S1) included in a first image frame (F1), obtain text information (e.g., “It looks delicious”) based on the obtained subtitle data, obtain voice information using TTS (Text to Speech) from the obtained text information, and control a speaker (600) to output the obtained voice information.

[0100] For another example, at least one processor (91) may obtain subtitle data from a second subtitle (S2) included in a second image frame (F2), obtain text information (e.g., “Let’s go out in 10 minutes”) based on the obtained subtitle data, obtain voice information using TTS from the obtained text information, and control a speaker (600) to output the obtained voice information.

[0101] In one embodiment, at least one processor (91) can control a speaker (600) to obtain audio information based on subtitle data and output the audio information when the user selects a mode for reading subtitles included in the image (I).

[0102] For example, at least one processor (91) may control a speaker (600) to obtain voice information based on subtitle data and output the voice information when the communication interface (400) receives a mode selection command for reading subtitles received from an external device (e.g., a user device).

[0103] The mode for reading subtitles may include a mode for outputting voice information obtained based on subtitle data obtained from subtitles included in an image (I) rather than an audio signal received from a content receiving unit (80).

[0104] The mode for reading subtitles may include a mode for outputting a voice in a language different from the language of the voice corresponding to the audio signal received from the content receiving unit (80).

[0105] Fig. 7 illustrates a flowchart of a control method for a display device according to one embodiment. Fig. 8 illustrates a table regarding weights assigned according to the type of object within an image frame according to one embodiment.

[0106] In one embodiment, at least one processor (91) can obtain image frame object information including at least one of the number of objects included in the image frame, the distance between objects, and the type of objects.

[0107] For example, referring to FIG. 5, at least one processor (91) can obtain first image frame object information including at least one of the number of objects included in the first image frame (F1) (e.g., 3), the distance between objects (e.g., d1, d2, d3), and the type of object (e.g., person, spoon, bowl).

[0108] For another example, referring to FIG. 6, at least one processor (91) can obtain second image frame object information including at least one of the number of objects included in the second image frame (F2) (e.g., 4), the distance between objects (e.g., d4, d5, d6, d7), and the type of object (e.g., male person, female person, cup).

[0109] According to various embodiments, at least one processor (91) can recognize whether an object displayed on the display unit (500) has changed based on a difference between the first image frame object information and the second image frame object information.

[0110] In one embodiment, at least one processor (91) can recognize whether an object displayed on the display unit has changed based on whether a difference between the number of objects included in the first image frame and the number of objects included in the second image frame is greater than or equal to a first reference value (examples 1000 and 1300). A change in an object displayed on the display unit may include a change in a scene of the image (I), a change in a position of an object within the image (I), an object displayed within the image (I) not being displayed, or an object not displayed within the image (I) being displayed.

[0111] For example, referring to FIGS. 5 and 6, at least one processor (91) can recognize whether an object displayed on the display unit (500) has changed based on whether the difference between the number of objects (O1, O2, O3) included in the first image frame (F1) and the number of objects (O11, O12, O41, O42) included in the second image frame (F2) is greater than or equal to a first reference value.

[0112] In one embodiment, at least one processor (91) can recognize whether an object displayed on the display unit (500) has changed based on whether a difference between a distance between objects included in a first image frame and a distance between objects included in a second image frame is greater than or equal to a second reference value (examples 1100 and 1300). The distance between objects included in the image frames may include a sum total of distances between objects included in the image frames.

[0113] For example, referring to FIGS. 5 and 6, at least one processor (91) can recognize whether an object displayed on the display unit (500) has changed based on whether the difference between the sum total of distances (d1+d2+d3) between objects (O1, O2, O3) included in the first image frame (F1) and the sum total of distances (d4+d5+d6+d7) between objects (O11, O12, O41, O42) included in the second image frame (F2) is greater than or equal to a second reference value.

[0114] Referring to FIG. 8, at least one memory (92) can store information about weights according to the type of object.

[0115] For example, the weight according to the type of object may include a first weight (W1) assigned to the object if the object is a person, a second weight (W2) assigned to the object if the object is a spoon, a third weight (W3) assigned to the object if the object is a bowl, and / or a fourth weight (W4) assigned to the object if the object is a cup. However, the weight according to the present disclosure is not limited thereto, and may include a type of object and a weight according to the type of object according to various embodiments.

[0116] The first image frame object information may include a first weight assignment result value that assigns a weight to each object included in the first image frame according to the type of the object. For example, referring to FIG. 5, the first image frame object information may include a first weight assignment result value (O1*W1+ O2*W2+ O3*W3) that adds together a value (O1*W1) that assigns a weight (W1) to the first object (O1), a value (O2*W2) that assigns a weight (W2) to the second object (O2), and a value (O3*W3) that assigns a weight (W3) to the third object (O3).

[0117] The second video frame object information may include a second weight assignment result value that assigns a weight according to the type of object to each object included in the second video frame. For example, referring to FIG. 6, the second video frame object information may include a second weight assignment result value (O11*W1+ O12*W1+ O41*W4+O42*W4) that adds together a value (O11*W1) assigning a weight (W1) to the fourth object (O11), a value (O12*W1) assigning a weight (W1) to the fifth object (O12), a value (O41*W4) assigning a weight (W4) to the sixth object (O41), and a value (O42*W4) assigning a weight (W4) to the seventh object (O42).

[0118] Referring again to FIG. 7, in one embodiment, at least one processor (91) can recognize whether an object displayed on the display unit (500) has changed based on whether a difference between a first weight assignment result value according to the type of object included in the first image frame and a second weight assignment value according to the type of object included in the second image frame is greater than or equal to a third reference value (examples 1200 and 1300).

[0119] In one embodiment, at least one processor (91) can change at least one of the pitch, speed, and size of the voice output by the speaker (600) based on whether it recognizes a change in an object displayed on the display unit (500) (1300 and 1400).

[0120] For example, referring to FIGS. 5 and 6, when at least one processor (91) recognizes whether an object displayed on the display unit (500) has changed based on a difference between the first image frame object information included in the first image frame (F1) and the second image frame object information included in the second image frame (F2), the processor (91) may control the speaker (600) to output a voice corresponding to the first subtitle (S1) included in the first image frame (F1), and then control the speaker (600) to output a voice corresponding to the second subtitle (S2) included in the second image frame (F2) in which at least one of a pitch, a speed, and a size of the voice is different from the voice corresponding to the first subtitle (S1).

[0121] According to the present disclosure, a visually impaired person who cannot visually see an image displayed on a display device can detect a change in an object in the image by voice.

[0122] FIG. 9 illustrates an image frame obtained from an image, according to one embodiment.

[0123] Referring to FIGS. 6 and 9, in one embodiment, at least one processor (91) can recognize whether an object displayed on the display unit (500) has changed even if a person is not included among the objects included in the image frame.

[0124] For example, at least one processor (91) can recognize a change in an object displayed on the display unit (500) based on the difference between the number of objects (O11, O12, O41, O42) included in the second image frame (F2) and the number of objects (O41, O42) included in the third image frame (F3) being greater than or equal to a first reference value.

[0125] For another example, at least one processor (91) can recognize a change in an object displayed on the display unit (500) based on the difference between the sum total (d4+d5+d6+d7) of distances between objects (O11, O12, O41, O42) included in the second image frame (F2) and the sum total (d7) of distances between objects (O41, O42) included in the third image frame (F3) being greater than or equal to a second reference value.

[0126] For another example, at least one processor (91) can recognize whether an object displayed on the display unit (500) has changed based on whether the difference between the second weight assignment result value (O11*W1+ O12*W1+ O41*W4+O42*W4) in which weights are assigned to each object (O11, O12, O41, O42) included in the second image frame (F2) according to the type of the object and the third weight assignment result value (O41*W4+O42*W4) in which weights are assigned to each object (O41, O42) included in the third image frame (F3) according to the type of the object is greater than or equal to the third reference value.

[0127] A display device according to one embodiment of the present disclosure may include a display unit on which an image including a first image frame and a second image frame is displayed; and at least one processor that obtains first image frame object information including at least one of the number of objects included in the first image frame, the distance between the objects, and the type of the objects, obtains second image frame object information including at least one of the number of objects included in the second image frame, the distance between the objects, and the type of the objects, and recognizes whether an object displayed on the display unit has changed based on a difference between the first image frame object information and the second image frame object information.

[0128] The at least one processor can recognize whether an object displayed on the display unit has changed based on a difference between the number of objects included in the first image frame and the number of objects included in the second image frame being greater than or equal to a first reference value.

[0129] The distance between objects included in the first image frame includes a sum of distances between objects included in the first image frame, the distance between objects included in the second image frame includes a sum of distances between objects included in the second image frame, and the at least one processor can recognize whether an object displayed on the display unit has changed based on a difference between the sum of distances between objects included in the first image frame and the sum of distances between objects included in the second image frame being greater than or equal to a second reference value.

[0130] A memory storing information on weights according to the type of object is further included, wherein the first image frame object information includes a first weight assignment result value that assigns a weight according to the type of object to each object included in the first image frame, and the second image frame object information includes a second weight assignment result value that assigns a weight according to the type of object to each object included in the first image frame.

[0131] The at least one processor can recognize whether an object displayed on the display unit has changed based on whether a difference between the first weight assignment result value and the second weight assignment result value is greater than or equal to a third reference value.

[0132] The at least one processor may input the first image frame to a learned machine learning model to determine an area occupied by an object included in the first image frame and a type of object included in the first image frame, and input the second image frame to the learned machine learning model to determine an area occupied by an object included in the second image frame and a type of object included in the second image frame.

[0133] The at least one processor can determine a distance between objects included in the first image frame based on an area occupied by the objects included in the first image frame, and can determine a distance between objects included in the first image frame based on an area occupied by the objects included in the second image frame.

[0134] A speaker; further comprising at least one processor, wherein the at least one processor can control the speaker to obtain subtitle data from subtitles included in the image, obtain audio information based on the subtitle data, and output the audio information.

[0135] The at least one processor can change at least one of the pitch, speed, and size of the voice based on recognition of whether an object displayed on the display unit has changed.

[0136] The at least one processor may control the speaker to obtain the audio information based on the subtitle data and output the audio information when the user selects a mode for reading the subtitles included in the video.

[0137] A method for controlling a display device according to one embodiment of the present disclosure may include: a method for controlling a display device including a display unit on which an image including a first image frame and a second image frame is displayed; obtaining first image frame object information including at least one of the number of objects included in the first image frame, a distance between the objects, and a type of the objects; obtaining second image frame object information including at least one of the number of objects included in the second image frame, a distance between the objects, and a type of the objects; and recognizing whether an object displayed on the display unit has changed based on a difference between the first image frame object information and the second image frame object information.

[0138] Recognizing whether an object displayed on the display unit has changed may include recognizing whether an object displayed on the display unit has changed based on a difference between the number of objects included in the first image frame and the number of objects included in the second image frame being greater than or equal to a first reference value.

[0139] The distance between objects included in the first image frame includes a sum of distances between objects included in the first image frame, the distance between objects included in the second image frame includes a sum of distances between objects included in the second image frame, and recognizing whether objects displayed on the display unit have changed may include recognizing whether objects displayed on the display unit have changed based on a difference between the sum of distances between objects included in the first image frame and the sum of distances between objects included in the second image frame being greater than or equal to a second reference value.

[0140] It further includes storing information on weights according to the type of object; and the first image frame object information may include a first weight assignment result value that assigns a weight according to the type of object to each object included in the first image frame, and the second image frame object information may include a second weight assignment result value that assigns a weight according to the type of object to each object included in the first image frame.

[0141] Recognizing whether an object displayed on the display unit has changed may include recognizing whether an object displayed on the display unit has changed based on a difference between the first weight assignment result value and the second weight assignment result value being greater than or equal to a third reference value.

[0142] The method may further include inputting the first image frame to a learned machine learning model to determine an area occupied by an object included in the first image frame and a type of object included in the first image frame, and inputting the second image frame to the learned machine learning model to determine an area occupied by an object included in the second image frame and a type of object included in the second image frame.

[0143] It may further include determining a distance between objects included in the first image frame based on an area occupied by the objects included in the first image frame, and determining a distance between objects included in the first image frame based on an area occupied by the objects included in the second image frame.

[0144] It may further include obtaining subtitle data from subtitles included in the above video, obtaining audio information based on the subtitle data, and outputting the audio information.

[0145] Outputting the voice information may include changing at least one of the pitch, speed, and size of the voice based on recognizing whether an object displayed on the display unit has changed.

[0146] Outputting the above audio information may include obtaining the audio information based on the subtitle data and outputting the audio information when the user selects a mode for reading subtitles included in the video.

[0147] A non-transitory storage medium storing computer-readable commands may include a non-transitory storage medium, wherein the commands, when executed by a processor, cause the processor to: control a display unit to display an image including a first image frame and a second image frame; obtain first image frame object information including at least one of the number of objects included in the first image frame, a distance between the objects, and a type of the objects; obtain second image frame object information including at least one of the number of objects included in the second image frame, a distance between the objects, and a type of the objects; and recognize whether an object displayed on the display unit has changed based on a difference between the first image frame object information and the second image frame object information.

[0148] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0149] Computer-readable storage media include all types of storage media that store instructions that can be deciphered by a computer. Examples include read-only memory (ROM), random access memory (RAM), magnetic tape, magnetic disks, flash memory, and optical data storage devices.

[0150] Additionally, a computer-readable recording medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0151] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable recording medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated on a machine-readable recording medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0152] The disclosed embodiments have been described with reference to the attached drawings as described above. Those skilled in the art will understand that the present invention can be implemented in forms other than the disclosed embodiments without altering the technical spirit or essential features of the present invention. The disclosed embodiments are illustrative and should not be construed as limiting.

Claims

1. A display unit on which an image including a first image frame and a second image frame is displayed; and Obtain first image frame object information including at least one of the number of objects included in the first image frame, the distance between the objects, and the type of the objects, Obtain second image frame object information including at least one of the number of objects included in the second image frame, the distance between the objects, and the type of the objects, A display device comprising at least one processor that recognizes whether an object displayed on the display unit has changed based on a difference between the first image frame object information and the second image frame object information.

2. In paragraph 1, At least one processor, A display device that recognizes whether an object displayed on the display unit has changed based on a difference between the number of objects included in the first image frame and the number of objects included in the second image frame being greater than or equal to a first reference value.

3. In paragraph 1, The distance between objects included in the first image frame includes the sum of the distances between objects included in the first image frame, and the distance between objects included in the second image frame includes the sum of the distances between objects included in the second image frame. At least one processor, A display device that recognizes whether an object displayed on the display unit has changed based on whether the difference between the sum of distances between objects included in the first image frame and the sum of distances between objects included in the second image frame is greater than or equal to a second reference value.

4. In paragraph 1, Further comprising a memory for storing information about weights according to the type of object; The above first video frame object information is, Includes a first weight assignment result value that assigns a weight to each object included in the first image frame according to the type of the object, The above second video frame object information is, A display device including a second weight assignment result value that assigns a weight according to the type of object to each object included in the first image frame.

5. In paragraph 4, At least one processor, A display device that recognizes whether an object displayed on the display unit has changed based on whether the difference between the first weight assignment result value and the second weight assignment result value is greater than or equal to a third reference value.

6. In paragraph 1, At least one processor, A display device that inputs the first image frame to a learned machine learning model to determine an area occupied by an object included in the first image frame and a type of object included in the first image frame, and inputs the second image frame to the learned machine learning model to determine an area occupied by an object included in the second image frame and a type of object included in the second image frame.

7. In paragraph 6, At least one processor, A display device that determines the distance between objects included in the first image frame based on the area occupied by the objects included in the first image frame, and determines the distance between objects included in the first image frame based on the area occupied by the objects included in the second image frame.

8. In paragraph 1, including speakers; At least one processor, A display device that obtains subtitle data from subtitles included in the image, obtains audio information based on the subtitle data, and controls the speaker to output the audio information.

9. In paragraph 8, At least one processor, A display device that changes at least one of the pitch, speed, and size of the voice based on recognition of whether an object displayed on the display unit has changed.

10. In paragraph 8, At least one processor, A display device that obtains the voice information based on the subtitle data and controls the speaker to output the voice information when the user selects a mode for reading the subtitles included in the video.

11. A method for controlling a display device including a display unit on which an image including a first image frame and a second image frame is displayed, Obtain first image frame object information including at least one of the number of objects included in the first image frame, the distance between the objects, and the type of the objects; Obtain second image frame object information including at least one of the number of objects included in the second image frame, the distance between the objects, and the type of the objects; A control method for a display device, comprising: recognizing whether an object displayed on the display unit has changed based on a difference between the first image frame object information and the second image frame object information.

12. In paragraph 11, Recognizing whether the object displayed on the above display unit has changed is as follows: A control method for a display device, comprising: recognizing whether an object displayed on the display unit has changed based on a difference between the number of objects included in the first image frame and the number of objects included in the second image frame being greater than or equal to a first reference value.

13. In paragraph 11, The distance between objects included in the first image frame includes the sum of the distances between objects included in the first image frame, and the distance between objects included in the second image frame includes the sum of the distances between objects included in the second image frame. Recognizing whether the object displayed on the above display unit has changed is as follows: A control method for a display device, comprising: recognizing whether an object displayed on the display unit has changed based on a difference between a sum of distances between objects included in the first image frame and a sum of distances between objects included in the second image frame being greater than or equal to a second reference value.

14. In paragraph 11, further comprising storing information about weights according to the type of object; The above first video frame object information is, Includes a first weight assignment result value that assigns a weight to each object included in the first image frame according to the type of the object, The above second video frame object information is, A control method for a display device including a second weight assignment result value in which weights are assigned to each object included in the first image frame according to the type of the object.

15. In paragraph 14, Recognizing whether the object displayed on the above display unit has changed is as follows: A control method for a display device, comprising: recognizing whether an object displayed on the display unit has changed based on a difference between the first weight assignment result value and the second weight assignment result value being greater than or equal to a third reference value.

Citation Information

Patent Citations

  • Terminal device and method for providing object information

    KR1020120017870A

  • Shadow Cinema System Using Smart Device

    KR1020160133096A

  • System and Method for Processing Image Informaion

    KR102391853B1

  • Connected interactive content data creation, organization, distribution and analysis

    US20220141534A1

  • KR20220124016A