Display device and screenshot method of video frame

By automatically recognizing the content of video frames on the display device and taking screenshots, the problem of users having to input commands frequently is solved, and efficient and accurate video frame capture is achieved.

CN117676254BActive Publication Date: 2026-08-04VIDAA INT HLDG (NETHERLANDS) CO
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VIDAA INT HLDG (NETHERLANDS) CO
Filing Date
2022-08-31
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

When users capture video frames using display devices, they need to frequently input screenshot commands, and the captured video frames are easily inaccurate due to time differences or lack of concentration.

Method used

Display devices automatically identify target video frames and take screenshots by recognizing the content of video frames. This includes recognizing content areas and background areas, judging changes in the image, adapting to the characteristics of different video types, and achieving automatic or semi-automatic screenshotting.

Benefits of technology

It simplifies user operations, improves the accuracy and efficiency of video frame extraction, and ensures that all required video frames are extracted without omission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117676254B_ABST
    Figure CN117676254B_ABST
Patent Text Reader

Abstract

The application provides a display device and a screenshot method of a video frame. After a first screenshot function is turned on, the display device acquires picture content of each video frame of a played video, and determines a first target video frame to be captured according to the picture content of two adjacent video frames. After the first target video frame is determined, the display device automatically captures the first target video frame. In the above screenshot process, the display device can automatically capture a video picture required by a user without inputting a screenshot instruction by the user for each video picture to be captured, so that the user operation can be effectively simplified. Moreover, the display device can accurately and without omission determine the first target video frame through picture recognition, so that the screenshot accuracy of the video picture can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent display device technology, and in particular to a display device and a method for taking screenshots of video frames. Background Technology

[0002] Display devices refer to terminal devices capable of outputting specific display images, such as smart TVs, mobile terminals, smart advertising screens, and projectors. Taking smart TVs as an example, smart TVs are television products based on Internet application technologies, possessing open operating systems and chips, and having open application platforms. They enable two-way human-computer interaction and integrate multiple functions such as audio-visual, entertainment, and data to meet diverse and personalized user needs.

[0003] When users play videos on a display device, they can send a screenshot command to the device to instruct it to capture the desired video frame. Each time a user views a desired frame, they need to send a screenshot command for that frame to instruct the display device to capture that frame. Therefore, if a user needs to capture multiple frames, they need to send multiple screenshot commands, which is cumbersome. There can be a time difference between when the user sends the screenshot command and when the desired video frame is captured, resulting in the display device capturing a different frame than the user expects. Furthermore, if the user is focused on watching the video, they may forget to send the screenshot command, thus failing to capture the desired frame. Summary of the Invention

[0004] This application provides a display device and a method for capturing video frames. When the display device is playing a video, it identifies the content of the screen, determines the target video frame to be captured, and automatically captures the target video frame, thereby effectively reducing user operations and improving the accuracy of the captured video frames.

[0005] In a first aspect, this application provides a display device, comprising:

[0006] The monitor is configured to display the video feed.

[0007] The controller is configured as follows:

[0008] After activating the first screenshot function, the screen content of each video frame of the video is obtained, and the screen content includes a content area and a background area;

[0009] A first target video frame is determined based on the content of two adjacent video frames. The two adjacent video frames include a preceding video frame and a following video frame in playback order. If the content area of ​​the following video frame is partially or completely different from the content area of ​​the preceding video frame, the first target video frame is the following video frame. If the content area of ​​the following video frame is different from the content area of ​​the preceding video frame, and the content area of ​​the preceding video frame includes the entire content area of ​​the following video frame, the first target video frame is the preceding video frame.

[0010] Capture the first target video frame.

[0011] In some embodiments of this application, the controller is configured to determine a first target video frame based on the content of two adjacent video frames, as follows:

[0012] Identify the content type of the video;

[0013] The first target video frame is determined based on the content type and the content of the two adjacent video frames;

[0014] Wherein, if the video is a first type of video, the first target video frame is the next video frame among the two adjacent video frames, and the first type of video includes videos presented in the form of playing a slideshow;

[0015] If the video is a second type of video, the first target video frame is the preceding video frame among the two adjacent video frames, and the second type of video includes videos presented in the form of written blackboard writing.

[0016] In some embodiments of this application, if the video is a first type of video, the controller is configured to determine the first target video frame based on the content type and the content of the two adjacent video frames, as follows:

[0017] Identify the content of two adjacent video frames;

[0018] If the content of two adjacent video frames is different, the latter video frame among the two adjacent video frames is determined to be the first target video frame.

[0019] In some embodiments of this application, the controller is configured to identify whether the content of two adjacent video frames is different:

[0020] Obtain the image features of the two adjacent video frames;

[0021] Calculate the matching degree of image features between two adjacent video frames;

[0022] If the matching degree is greater than or equal to the matching degree threshold, it is determined that the content of the two adjacent video frames is the same; if the matching degree is less than the matching degree threshold, it is determined that the content of the two adjacent video frames is different.

[0023] In some embodiments of this application, if the video is a second type of video, the controller is configured to determine the first target video frame based on the content type and the content of the two adjacent video frames, as follows:

[0024] Store a specified number of video frames and identify content-erased videos, wherein the preceding video frame of the content-erased video frame includes all the screen content of the content-erased video frame, and the screen content of the content-erased video frame is less than the screen content of the preceding video frame of the content-erased video frame.

[0025] If the content-erased video frame is identified, the video frame preceding the content-erased video frame in the specified number of stored video frames is determined as the first target video frame.

[0026] In some embodiments of this application, the controller is configured to perform the function of identifying whether a content-wiped video has appeared:

[0027] Obtain the area ratio of the background region to the content area of ​​the adjacent video frames;

[0028] If the area ratio of the background region to the screen content of the adjacent video frames is different, determine whether the area ratio of the background region to the screen content of the next video frame in the adjacent video frames is greater than or equal to a preset ratio.

[0029] If the ratio of the area of ​​the background region to the area of ​​the content in the next video frame is greater than or equal to the preset ratio, the next video frame is determined to be the content-erasing video frame.

[0030] In some embodiments of this application, the controller is configured to determine a first target video frame based on the content of two adjacent video frames, as follows:

[0031] Identify human images in the content of the video frame;

[0032] Based on the positional relationship between the person image and the content area of ​​the video frame, a valid video frame is determined, wherein the valid video frame is a video frame in which the content area of ​​the person image and the video frame do not overlap.

[0033] Remove the image of the person from the content of the valid video frames;

[0034] The first target video frame is determined based on the content of the two adjacent valid video frames after removing the images of the people.

[0035] In some embodiments of this application, the controller is further configured to:

[0036] Receive a screenshot command input by the user, wherein the screenshot command indicates the second target video frame to be captured;

[0037] In response to the screenshot command, the second target video frame is captured.

[0038] In some embodiments of this application, the controller is further configured to:

[0039] Obtain the link information of the target video frame to be stored. The link information includes the address information of the target video to which the target video frame to be stored belongs and the frame number of the target video frame to be stored in the video frame sequence. The video frame sequence is composed of the video frames of the target video arranged in the playback order.

[0040] Generate corresponding links based on the link information of the target video frames to be stored;

[0041] The target video frame to be stored is stored at a specified address, and a mapping relationship is established between the target video frame to be stored and the corresponding link, so that after the target video frame to be stored is selected, the target video is played starting from the target video frame to be stored according to the corresponding link.

[0042] Secondly, this application provides a method for capturing a video frame, applied to a display device, wherein the display device displays the video frame of a video, the method comprising:

[0043] After activating the first screenshot function, the screen content of each video frame of the video is obtained, and the screen content includes a content area and a background area;

[0044] A first target video frame is determined based on the content of two adjacent video frames. The two adjacent video frames include a preceding video frame and a following video frame in playback order. If the content area of ​​the following video frame is partially or completely different from the content area of ​​the preceding video frame, the first target video frame is the following video frame. If the content area of ​​the following video frame is different from the content area of ​​the preceding video frame, and the content area of ​​the preceding video frame includes the entire content area of ​​the following video frame, the first target video frame is the preceding video frame.

[0045] Capture the first target video frame.

[0046] After enabling the first screenshot function, the display device acquires the content of each video frame of the playing video and determines the first target video frame to be captured based on the content of two adjacent video frames. Once the first target video frame is determined, the display device automatically captures it. During this screenshot process, the display device can automatically capture the video frame desired by the user, eliminating the need for the user to input screenshot commands for each individual frame, thus effectively simplifying user operations. Furthermore, through image recognition, the display device can accurately and completely determine the first target video frame, thereby significantly improving the accuracy of video screenshots. Attached Figure Description

[0047] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This application illustrates the usage scenario of the display device.

[0049] Figure 2 This is a block diagram showing the configuration of the control device in an embodiment of this application;

[0050] Figure 3 This is a configuration diagram of the display device in the embodiments of this application;

[0051] Figure 4 This is a diagram showing the operating system configuration of the display device in the embodiments of this application;

[0052] Figure 5 This is a schematic diagram of a user manually taking a screenshot of video frame A in an embodiment of this application;

[0053] Figure 6 This is a schematic diagram of a user manually capturing video frame B in an embodiment of this application;

[0054] Figure 7 This is a schematic diagram of a user manually capturing video frame B in an embodiment of this application;

[0055] Figure 8 This is a schematic diagram of the screenshot function settings interface in an embodiment of this application;

[0056] Figure 9 This is a schematic diagram illustrating the process of the display device automatically capturing video frames in an embodiment of this application;

[0057] Figure 10 This is a schematic diagram of video frame D in an embodiment of this application;

[0058] Figure 11This is a schematic diagram illustrating the process by which a display device determines a first target video frame in an embodiment of this application.

[0059] Figure 12 This is a schematic diagram illustrating the process by which a display device determines a first target video frame for a first type of video in an embodiment of this application.

[0060] Figure 13 This is a schematic diagram illustrating the process by which a display device identifies whether the content of two adjacent video frames is different in an embodiment of this application.

[0061] Figure 14 This is a schematic diagram showing the display device automatically taking a screenshot during the playback of video 2 in an embodiment of this application;

[0062] Figure 15 This is a schematic diagram illustrating the process by which a display device determines a first target video frame for a second type of video in an embodiment of this application.

[0063] Figure 16 This is a schematic diagram illustrating the process of the display device determining content to erase video frames in an embodiment of this application;

[0064] Figure 17 This is a schematic diagram of video frame segmentation in an embodiment of this application;

[0065] Figure 18 This is a schematic diagram showing the display device automatically taking a screenshot during the playback of video 3 in an embodiment of this application;

[0066] Figure 19 This is a schematic diagram illustrating the process by which a display device acquires the image content of a video frame in an embodiment of this application.

[0067] Figure 20 This is a schematic diagram of invalid video frames in an embodiment of this application;

[0068] Figure 21 This is a schematic diagram of valid video frames in the embodiments of this application;

[0069] Figure 22 This is a schematic diagram illustrating the process of a display device storing a first target video frame in an embodiment of this application;

[0070] Figure 23 This is a schematic diagram of the prompt interface in an embodiment of this application;

[0071] Figure 24 This is a schematic diagram illustrating the process of a display device capturing a second target video frame in response to a user's instruction, as shown in an embodiment of this application.

[0072] Figure 25 This is a schematic diagram of a semi-automatic screenshot taken during the playback of video 2 in an embodiment of this application. Detailed Implementation

[0073] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0074] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0075] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0076] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0077] The display device provided in this application can have various implementation forms, such as a smart TV, laser projection device, monitor, electronic bulletin board, electronic table, etc., or a device with a display screen such as a mobile phone, tablet computer, or smartwatch. Figure 1 and Figure 2 This is one specific embodiment of the display device of this application.

[0078] Figure 1 This is a schematic diagram illustrating a usage scenario of the display device according to the embodiments. For example... Figure 1 As shown, users can operate the display device 200 through the mobile terminal 300 or the control device 100. The display device 200 can obtain network data through the server 400 or obtain live broadcast signals through satellite.

[0079] Figure 2This is a configuration block diagram of the control device 100. In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device 200 includes infrared protocol communication, Bluetooth protocol communication, and at least one of other short-range communication methods, controlling the display device 200 wirelessly or via a wired connection. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.

[0080] In some embodiments, a mobile terminal 300 (such as a mobile phone, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the mobile terminal 300 may be used to control the display device 200.

[0081] Figure 3 A configuration block diagram of a display device 200 according to an exemplary embodiment is shown.

[0082] Display device 200 includes at least one of tuner / demodulator 210, communicator 220, detector 230, external device interface 240, controller 250, display 260, audio output interface 270, memory, power supply, and user interface 280.

[0083] In some embodiments, the display device 200 can establish the transmission and reception of control signals and data signals with the control device 100 or the server 400 via the communicator 220. In some embodiments, the controller 250 and the tuner 210 can be located in different separate devices, that is, the tuner 210 can also be in an external device of the main device where the controller 250 is located, such as an external set-top box. In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200. In some embodiments, the controller 250 includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (Random Access Memory), ROM (Read-Only Memory), a first to an nth interface for input / output, a communication bus, etc. In some embodiments, the display 260 includes a display screen component for presenting images, a driving component for driving image display, a component for receiving image signals output from the controller 250, and a user interface (UI) for displaying video content, image content, and a menu control interface. In some embodiments, the display 260 can be a liquid crystal display, an OLED display, or a projection display, and can also be a projection device and a projection screen. In some embodiments, the user can input user commands through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the GUI. Alternatively, the user can input user commands by inputting specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors. In some embodiments, a "user interface" is a medium interface for interaction and information exchange between an application or operating system and the user; it realizes the conversion between the internal form of information and a form acceptable to the user. A commonly used form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an interface element such as an icon, window, or control displayed on the screen of an electronic device. The control can include at least one of the visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0084] In some embodiments, the display device 200 is based on the VIDAA software platform, such as Figure 4As shown, the operating system is divided into three layers, from top to bottom: the application layer, the middleware layer, and the hardware layer. The application layer includes various applications and application frameworks. The middleware layer includes various protocols and components. The hardware layer includes various interfaces, hardware, and drivers.

[0085] When playing video, display device 200 displays each video frame in the order they are played. If a user is interested in a particular video frame while watching the video, they may want to capture that frame. In this case, the user can input a screenshot command into display device 200 to instruct it to capture that video frame. This explanation will be illustrated using an example where display device 200 is a television, and the user controls display device 200 through control device 100, which is a remote control. Figure 5 As shown, when the TV is playing video frame A, if the user needs to take a screenshot, the user can send a first screenshot command to the TV by pressing the screenshot button on the remote control. After receiving the first screenshot command, the TV responds by capturing the currently displayed video frame. If the currently displayed video frame is video frame A, the TV can capture video frame A. Figure 6 As shown, when the TV is playing video frame B, if the user needs to take a screenshot again, the user still needs to send a second screenshot command to the TV by pressing the screenshot button on the remote control. After receiving the second screenshot command, the TV responds by capturing the currently displayed video frame. If the currently displayed video frame is video frame B, the TV can capture video frame B.

[0086] It is evident that for each video frame captured, the user needs to input a screenshot command into the display device 200, making the screenshot operation rather cumbersome.

[0087] If a user misses the opportunity to take a screenshot while watching a video—for example, if the user forgets to send the screenshot command, or if the display device 200 displays a different video frame when it receives the screenshot command—the desired video frame will not be captured. Figure 7 As shown, when the TV is playing video frame B, if the user needs to take a screenshot, the user sends a screenshot command to the TV by pressing the screenshot button on the remote control. After receiving the screenshot command, the TV responds by capturing the currently displayed video frame. However, if the currently displayed video frame is the next video frame after video frame B, such as video frame C, the TV will capture video frame C and will not be able to capture video frame B.

[0088] To address the aforementioned issues, the display device 200 provided in this embodiment is configured to have a first screenshot function. After enabling the first screenshot function, the display device 200 can automatically capture video frames during video playback through image recognition.

[0089] The video played on the display device 200 can be live video, online video, etc., such as live course videos or online course videos. This video can be provided by an application on the display device 200 or by a screen-sharing device connected to the display device 200, such as a mobile phone or computer.

[0090] The first screenshot function includes: automatic screenshot mode and / or semi-automatic screenshot mode. Automatic screenshot mode means that the display device 200 automatically captures video frames without responding to user-inputted screenshot commands. Semi-automatic screenshot mode means that the display device 200 automatically captures video frames while also responding to user-inputted screenshot commands.

[0091] Display device 200 can be configured to automatically activate the first screenshot function when playing a video. If the first screenshot function includes an automatic screenshot mode and a semi-automatic screenshot mode, display device 200 will automatically activate one of these modes. Display device 200 can also be configured to determine whether to activate the first screenshot function by identifying the content type of the video being played. For example, if the video content type is identified as a course, display device 200 will automatically activate the first screenshot function. Display device 200 can identify the video content type based on the video name, user settings, etc. If the first screenshot function includes an automatic screenshot mode and a semi-automatic screenshot mode, display device 200 will automatically activate one of these modes. If the video content type is identified as a movie / TV show, display device 200 will deactivate the first screenshot function. Display device 200 can also be configured to activate the first screenshot function in response to a user's command to do so while playing a video. For example: A user controls a display device 200 through a control device 100. The user inputs a setting command to the display device 200 to configure the first screenshot function. The display device 200 responds to this setting command by displaying the screenshot function settings interface. (See reference...) Figure 8 The screenshot function settings interface shown includes a first screenshot function. If the first screenshot function includes an automatic screenshot mode and a semi-automatic screenshot mode, the screenshot function settings interface includes an automatic screenshot option and a semi-automatic screenshot option. The automatic screenshot option corresponds to the automatic screenshot mode, and the semi-automatic screenshot option corresponds to the semi-automatic screenshot mode. The user moves the focus to the option corresponding to the desired screenshot mode by manipulating the control device 100. For example, if the desired screenshot mode is automatic screenshot mode, such as... Figure 8 As shown, move the focus to the automatic screenshot option and select it.

[0092] In some embodiments, the display device 200 is further configured to have a second screenshot function. This second screenshot function means that during video playback, the display device 200 does not automatically capture video frames; instead, it only captures frames in response to a user-inputted screenshot command. This second screenshot function can be one of the screenshot modes in the first screenshot function, for example, such as... Figure 8 The manual screenshot mode shown.

[0093] In some embodiments, the display device 200 can be configured to enable a second screenshot function by default when playing a video, i.e., to enable manual screenshot mode by default. After receiving a user's instruction to enable the first screenshot function, it switches to the first screenshot function. Alternatively, it switches to the first screenshot function only after recognizing that the content type of the played video is a specified type, such as a course.

[0094] Example 1

[0095] Display device 200 activates the automatic screenshot mode in the first screenshot function. In this automatic screenshot mode, display device 200 can perform the following steps: Figure 9 The process shown involves capturing video frames, and the specific steps are as follows:

[0096] S901, acquires the image content of each video frame.

[0097] The video played by the display device 200 consists of N (N is a positive integer greater than 0) video frames. These N video frames form a frame sequence according to the playback order, and each video frame corresponds to a frame number in the frame sequence. For example, video 1 consists of 100 video frames, which form a frame sequence, and these 100 video frames correspond to frame numbers 1 to 100 in the playback order.

[0098] Each video frame consists of multiple pixels, each with a corresponding color value and / or grayscale value. The content of a video frame includes a content area and a background area. The content area displays the content of the video frame, which may include people, patterns, characters, etc. The background area is the area in the video frame excluding the content area.

[0099] This involves capturing the content of a video frame, specifically the content and background areas. Let's take a course as an example to illustrate this. Figure 10 As shown, video frame D includes a content area 1001 (including patterns and characters) and a background area 1002. Obtaining the content of video frame D means obtaining the content area 1001 and the background area 1002. Based on the content area and background area of ​​the video frame, the content of the video frame can be distinguished.

[0100] In some embodiments, the display device 200 may set part or all of the background area in a video frame as an edge area. The display device 200 may determine the region parameters of the edge area, such as its position, size, and shape within the video frame, based on relevant video parameters, such as the video type. For example, if the video type is a course, the content area of ​​each video frame will be rectangular and concentrated in the center of the video frame; therefore, the background areas at the top and bottom, left and right sides, or both the top and bottom and left and right sides of the video frame may be set as edge areas.

[0101] Remove edge regions from video frames. Figure 10 Taking video frame D as an example. If the edge regions are set to be located on the left and right sides and top and bottom of the video frame, where the width of the left and right edge regions is equal and the sum of their widths is 5% of the video frame's width, and the height of the top and bottom edge regions is equal and the sum of their heights is 5% of the video frame's height, then... Figure 10 The edge region 1003 is shown in shaded area. The display device 200 removes the edge region 1003 from the video frame D. Removing the edge region from the video frame not only effectively avoids interference from the edge region to subsequent calculations, but also effectively reduces the size of the image content, thereby effectively reducing the amount of computation required for subsequent calculations based on the image content.

[0102] S902, determine the first target video frame based on the content of two adjacent video frames.

[0103] The display device 200 automatically captures n video frames (where n is a positive integer greater than 0) using the first screenshot function. After detecting a change in the screen content, the display device 200 identifies the first target video frames from the video frames where the change occurred, ensuring that all first target video frames are captured without omission. The n captured first target video frames can include the entire content of the video, allowing the user to grasp the entire content of the video by browsing these n first target video frames.

[0104] The display device 200 determines whether the image content has changed by comparing the currently displayed video frame with the previous video frame. This allows for real-time monitoring of video frames where image changes have occurred. The currently displayed video frame and the previous video frame are considered adjacent video frames. Following the playback order, the previous video frame is the preceding video frame among the two adjacent video frames, and the currently displayed video frame is the following video frame.

[0105] During video playback, it's possible that m consecutive video frames (where m is a positive integer greater than 0, and m is less than or equal to the total number of video frames) have the same content. The content of the (m+1)th video frame differs from that of the mth video frame, meaning the content changes. The first target video frame is determined based on the difference between the content of the (m+1)th video frame and the mth video frame.

[0106] When playing videos of different content types, the display device 200 determines the first target video frame using a method appropriate to that content type. The video content type can be categorized into a first type of video and a second type of video based on the content of the video frame.

[0107] The first type of video includes at least one first type of video segment, in which the content of each video frame within the same first type of video segment is the same, and in two adjacent first type of video segments, the content of the last video frame of the preceding first type of video segment is partially or completely different from the content of the first video frame of the following first type of video segment.

[0108] In some embodiments, if the video type is a course, the first type of video may include: courses presented in the form of PowerPoint presentations (PPT), etc. Each first type of video segment corresponds to displaying one slide of a PPT.

[0109] The second type of video includes second-type video segments and third-type video segments. Within the same second-type video segment, the content of the latter video frame includes the entire content of the former video frame. Similarly, within the same third-type video segment, the content of the former video frame includes the entire content of the latter video frame. The second-type and third-type video segments are distributed alternately. If a second-type video segment precedes a third-type video segment, the content of the last video frame of the second-type video segment includes the entire content of the first video frame of the third-type video segment. If a second-type video segment follows a third-type video segment, the content of the first video frame of the second-type video segment includes the entire content of the last video frame of the third-type video segment.

[0110] In some embodiments, if the video type is a course, the second type of video may include: a course presented in the form of a whiteboard, etc. Video frames in which content is continuously added to the whiteboard belong to the same second type of video segment, and video frames in which content is continuously erased from the whiteboard belong to the same third type of video segment.

[0111] Display device 200 can be in accordance with Figure 11The process shown determines the first target video frame, and the specific steps are as follows:

[0112] S1101 identifies the content type of the video.

[0113] In some embodiments, a user can configure the type of content to be played in the display device 200 before watching the video. See also... Figure 8 The screenshot function settings interface shown includes type options: slideshow and whiteboard. When setting the first screenshot function, the user can simultaneously indicate the content type of the video being played. For example, if the display device 200 receives an instruction to select the slideshow option, it determines that the video being played is a first-type video. If the display device 200 receives an instruction to select the whiteboard option, it determines that the video being played is a second-type video.

[0114] In some embodiments, the display device 200 can determine the content type of a video based on video information such as its name and description. For example, if the display device 200 recognizes that the video description is "a PPT course about graphic explanations," it determines that the video being played is a first-type video. If the display device 200 recognizes that the video description is "a lecture video about graphic explanations," it determines that the video being played is a second-type video.

[0115] S1102, determine the first target video frame based on the content type and the content of two adjacent video frames.

[0116] After determining the content type of the video, the display device 200 determines the first target video frame based on the content type and the content of two adjacent video frames.

[0117] Specifically, for the first type of video, the display device 200 can be configured as follows: Figure 12 The process shown determines the first target video frame, and the specific steps are as follows:

[0118] S1201 identifies the content of two adjacent video frames.

[0119] After acquiring the image content of each video frame, i.e., the content area and blank area of ​​each video frame, the display device 200 can determine whether the image content is different by comparing the content areas and blank areas of two adjacent video frames. Specifically, if the content area of ​​a later video frame is partially or completely different from the content area of ​​a previous video frame, the image content of two adjacent video frames is considered different.

[0120] In some embodiments, the display device 200 may be configured as follows: Figure 13 The process shown identifies whether the content of two adjacent video frames is different. The specific steps are as follows:

[0121] S1301, acquire the image features of two adjacent video frames.

[0122] In some embodiments, the image features of a video frame include pixel features, which characterize the grayscale value and / or color value of a pixel. The content area and background area of ​​a video frame can be distinguished by the grayscale value and / or color value of each pixel. Changes in the content of a video frame, i.e., changes in the content area and background area, are also changes in the pixels of the video frame. Therefore, it can be determined whether the content of a video frame has changed by comparing its pixel features.

[0123] In some embodiments, the image features of a video frame also include character information. For example, the content area of ​​a video frame mainly includes characters, such as text, letters, and symbols, while the content area includes fewer or no patterns. The background area remains unchanged or changes only slightly, or the background area is a solid color. In this case, if only pixel features are used to determine whether the content of the video frame has changed, it is difficult to identify subtle differences, resulting in low accuracy. For example, if the difference between two video frames is only one character in the content area (only the content changes, other parameters, such as color and font size, remain unchanged), comparing only pixel features will easily overlook this change due to its small magnitude, leading to an inaccurate identification of the change in the image content. In this case, character information can be compared, or pixel features and character information can be compared, to determine whether the content of the video frame has changed.

[0124] S1302, calculate the matching degree of image features between two adjacent video frames.

[0125] In the first example, if we want to determine whether the content of a video frame has changed based on the pixel features of two adjacent video frames, we can calculate the matching degree of the pixel features using template matching.

[0126] Template matching refers to comparing the grayscale or color values ​​of pixels at the same position in two video frames and calculating the sum of the differences between the grayscale values ​​or the color values. If the image size of the video frame is Nx×Ny, the matching degree of the pixel features of two adjacent video frames satisfies the following formula (1):

[0127]

[0128] Where Ii represents the i-th video frame, Ij represents the j-th video frame, and the i-th and j-th video frames are adjacent video frames, d(Ii, Ij) represents the matching degree of pixel features between two adjacent video frames, Ii(x, y) represents the grayscale value or color value of the pixel located at (x, y) in the i-th video frame, and Ij(x, y) represents the grayscale value or color value of the pixel located at (x, y) in the j-th video frame.

[0129] In the second example, if we want to determine whether the content of a video frame has changed based on the pixel features of two adjacent video frames, we can calculate the matching degree of the pixel features using the color histogram matching method.

[0130] A color histogram is a statistical result of the number of pixels in a video frame that fall into different color values. If the color value (or gray value) of a video frame I consists of n levels, each color value is represented as qi (i = 1, 2, 3, ..., n), and the number of pixels with the color value qi is represented as hi (i = 1, 2, 3, ..., n), that is, the color histogram of video frame I is composed of h1, h2, ..., hn. The matching degree of pixel features between two adjacent video frames satisfies the following formula (2):

[0131]

[0132] Where Ii represents the i-th video frame, Ij represents the j-th video frame, and the i-th and j-th video frames are adjacent video frames, d(Ii, Ij) represents the matching degree of pixel features between two adjacent video frames, Hi(k) represents the color feature of the k-th level color value in the color histogram of the i-th video frame, and Hj(k) represents the color feature of the k-th level color value in the color histogram of the j-th video frame.

[0133] In the third example, text recognition technology, such as Optical Character Recognition (OCR), can be used to identify characters in the content areas of two adjacent video frames.

[0134] In the fourth example, the template matching method from the first example can be combined with the OCR technology from the third example to compare the content of two adjacent video frames.

[0135] In the fifth example, the color histogram matching method from the second example can be combined with the OCR technology from the third example to compare the content of two adjacent video frames.

[0136] S1303, Based on the matching degree, determine whether the content of two adjacent video frames is different.

[0137] Based on the matching degree calculated in S1302, if the matching degree is greater than or equal to the matching degree threshold, the content of two adjacent video frames is the same. If the matching degree is less than the matching degree threshold, the content of two adjacent video frames is different.

[0138] Taking the first example in S1302 as an example, if the matching degree calculated according to formula (1) is greater than or equal to the matching degree threshold, the content of two adjacent video frames is the same. If the matching degree calculated according to formula (1) is less than the matching degree threshold, the content of two adjacent video frames is different.

[0139] Taking the second example in S1302 as an example, if the matching degree calculated according to formula (2) is greater than or equal to the matching degree threshold, the content of two adjacent video frames is the same. If the matching degree calculated according to formula (2) is less than the matching degree threshold, the content of two adjacent video frames is different.

[0140] Taking the third example in S1302 as an example, if the matching degree calculated according to formula (1) is greater than or equal to the matching degree threshold, and the characters remain unchanged, the content of the two adjacent video frames is the same. If the matching degree calculated according to formula (1) is less than the matching degree threshold, and / or the characters change, the content of the two adjacent video frames is different.

[0141] Taking the fourth example in S1302 as an example, if the matching degree calculated according to formula (2) is greater than or equal to the matching degree threshold, and the characters remain unchanged, the content of the two adjacent video frames is the same. If the matching degree calculated according to formula (2) is less than the matching degree threshold, and / or the characters change, the content of the two adjacent video frames is different.

[0142] In some embodiments, the method for determining whether the content of two adjacent video frames has changed is not limited to the four examples mentioned above, and may also include regression analysis, image transformation methods (such as texture analysis, principal component analysis, change vector analysis, etc.), post-classification comparison, machine learning (such as regression analysis, support vector machine SVM, decision tree, etc.), deep learning (such as various neural networks), etc.

[0143] S1202, If the content of two adjacent video frames is different, determine the latter video frame of the two adjacent video frames as the first target video frame.

[0144] In the first type of video, if the content area of ​​the (m+1)th video frame (the currently displayed video frame) is different from the content area of ​​the mth video frame, the first target video frame is the (m+1)th video frame. That is, if the display device 200 detects that the content of the currently displayed video frame is different from the previous video frame, and the content area of ​​the currently displayed video frame is partially or completely different from the content area of ​​the previous video frame, the currently displayed video frame is determined as the first target video frame. If the content area of ​​the currently displayed video frame is partially or completely different from the content area of ​​the previous video frame, it can be considered that the content displayed by the currently displayed video frame is new content compared to the content of the previous video frame, rather than content obtained based on the content of the previous video frame. In other words, it has no relation to the content of the previous video frame, and the content of the previous video frame cannot cover the content of the currently displayed video frame. To ensure that the captured video frame can include complete video content, the currently displayed video frame needs to be used as the first target video frame, and the currently displayed video frame needs to be captured.

[0145] In some embodiments, to ensure that no first target video frame is omitted, the first target video frame may also include the first video frame of the first type of video.

[0146] In one example, we will illustrate this by showing video 2 playing on a television. Figure 14 As shown, Video 2 is a course presented in PowerPoint format. Video 2 includes frames 1 through 6 in the order they appear on screen. Frames 1 and 2 have the same content, both including a circle. Frames 3 and 4 have the same content, both including a triangle. Frame 5 includes a circle. Frame 6 includes a star. If Video 2 shows a course with 4 slides, frames 1 and 2 show slide 1, frames 3 and 4 show slide 2, frame 5 shows slide 3, and frame 6 shows slide 4.

[0147] The television identifies video 2 as a first-class video and uses template matching to determine whether the content of two adjacent video frames is the same. Taking two adjacent video frames, the second and third video frames, as an example, the television calculates that the matching degree between the pixel features of the third video frame and the pixel features of the second video frame is less than the matching degree threshold. Therefore, the television determines that the content of the third video frame is different from that of the second video frame and identifies the third video frame as the first target video frame. (Referring to identifying the third video frame as the first target video frame...) Figure 14In this context, the first target video frame is represented by a video frame with a gray background. The first target video frame also includes the 5th and 6th video frames, which will not be described in detail here. In some embodiments, the 1st video frame of video 2 is also determined as the first target video frame.

[0148] For the second type of video, the display device 200 can be configured as follows: Figure 15 The process shown determines the first target video frame, and the specific steps are as follows:

[0149] S1501, stores a specified number of video frames and identifies content to erase the video.

[0150] When playing video, the display device 200 always caches a specified number of video frames. Specifically, the display device 200 caches the currently playing video frames and deletes the first cached video frame. For example, if the display device 200 caches 20 video frames, and it has already cached frames 50-69, and is currently displaying the 70th video frame, the display device 200 will store the 70th frame, delete the cached 50th frame, and update the cache by storing frames 51-70.

[0151] The content-erased video frame is the first video frame in the third type of video segment within the second type of video. That is, the video frame preceding the content-erased video frame includes all the screen content of the content-erased video frame, and the screen content of the content-erased video frame is less than the screen content of the video frame preceding it.

[0152] In some embodiments, the display device 200 can determine a content-erased video frame by comparing the content areas in two adjacent video frames. Specifically, if the content area of ​​the later video frame is smaller than that of the earlier video frame (i.e., the background area of ​​the later video frame is larger than that of the earlier video frame), and the content area of ​​the earlier video frame includes the entire content area of ​​the later video frame, it is equivalent to erasing part or all of the content area from the previous video frame to obtain the later video frame. Therefore, the later video frame can be determined as a content-erased video frame.

[0153] In some embodiments, directly using video frames with reduced content areas as content-erasing video frames can result in the capture of video frames with highly repetitive content. For example, consider three consecutive video frames where the content area of ​​the second video frame is only one character less than that of the first, and the content area of ​​the third video frame is only one character more than that of the second, with the extra character located precisely at the position where the second video frame's content area is missing one character compared to the first. Essentially, one character in the first video frame has been modified to obtain the third video frame. If the second video frame is then used as the content-erasing video frame, and the first video frame is further selected as the first target video frame for capture, the captured video frame's content will not meet the user's requirements. To avoid this problem, the content-erasing video frame can be determined based on the area of ​​the background region within the video frame.

[0154] Display device 200 can be in accordance with Figure 16 The process shown involves determining the content to be erased from the video frames. The specific steps are as follows:

[0155] S1601, obtain the ratio of the area of ​​the background region to the area of ​​the content in the adjacent video frame.

[0156] The difference in content between video frames can be determined by comparing changes in the area of ​​the content and background regions. Specifically, if the areas of the content and background regions change between two adjacent video frames, the content of those two frames is different. If the areas of the content and background regions do not change between two adjacent video frames, the content of those two frames is the same.

[0157] Since the sum of the areas of the background region and the content region equals the area of ​​the image content, changes in the background region and the content region inevitably lead to changes in the other region. Therefore, a change in one of the regions can be used to represent a change in the image content. In some embodiments, the background region can be represented by its proportion of the image content (i.e., the ratio of the background region to the image content area). If the ratio of the background region to the image content area changes in a video frame, the image content of the video frame changes.

[0158] In some embodiments, image segmentation methods can be used to accurately segment the content region and background region of a video frame. For example, threshold segmentation can be used to segment the content region and background region. Threshold segmentation calculates one or more grayscale thresholds based on the grayscale features of the video frame, compares the grayscale value of each pixel in the video frame with the grayscale threshold, and classifies the corresponding pixel into the appropriate category (content region and background region) according to the comparison result. The classification of the pixel satisfies the following formula (3):

[0159]

[0160] Alternatively, the classification of pixels satisfies the following formula (4):

[0161]

[0162] like Figure 17 As shown, Figure 17 In the middle, ① represents the content of the video frame before the segmented region. Figure 17 In section ②, the video frame content is shown after the segmented region. The shaded area represents the content area, and the non-shaded area represents the background area.

[0163] In some embodiments, other segmentation algorithms can also be used to segment the content of video frames, such as: region-based image segmentation methods, edge detection-based segmentation methods, image segmentation algorithms combined with specific tools, image segmentation based on genetic algorithms, and segmentation methods based on active contour models.

[0164] In some embodiments, image segmentation methods and image recognition algorithms can be combined to segment the content of video frames. The image recognition algorithms may include: K-Nearest Neighbor (KNN) classification algorithm, Support Vector Machine (SVM) machine learning, Back Propagation Neural Network (BPNN) neural network, Convolutional Neural Network (CNN), etc. Image segmentation methods can be referred to the above embodiments and will not be repeated here.

[0165] After the display device divides the video frame content into 200 segments, it can accurately obtain the background area and the content area, and further determine whether the screen content has changed based on the obtained background area and content area.

[0166] The area ratio is obtained by calculating the ratio of the area of ​​the background region to the area of ​​the content in the image.

[0167] S1602, if the area ratio of the background area to the screen content of adjacent video frames is different, determine whether the area ratio of the background area to the screen content of the next video frame in the adjacent video frames is greater than or equal to a preset ratio.

[0168] S1603, if the ratio of the area of ​​the background region to the area of ​​the content in the next video frame is greater than or equal to a preset ratio, the next video frame is determined to be a content erasure video frame.

[0169] Generally, if the ratio of the background area to the content area of ​​the video frame is greater than or equal to a preset ratio, it can be assumed that the content of that video frame was obtained by extensively erasing the content area of ​​the previous video frame, rather than through minor modifications. This means there is a high probability that subsequent video frames will contain new content. This video frame should be identified as a content-erased video frame. The preset ratio can be set based on the ratio of the background area to the content area of ​​the previous video frame. For example, the preset ratio can be set to the ratio of the background area to the content area of ​​the previous video frame, or it can be set to a ratio that is a specified value smaller than the ratio of the background area to the content area of ​​the previous video frame.

[0170] If the ratio of the background area to the content area is less than a preset ratio, it can be considered that the content of this video frame is not significantly different from the content of the previous video frame. There is a high probability that only minor modifications have been made to the content of the previous video frame, and the content of subsequent video frames will largely overlap with the content of the previous video frame. This video frame should not be identified as a content-erased video frame.

[0171] S1502, if a content-erased video frame is identified, determine the first target video frame from among the specified number of stored video frames that is the video frame preceding the content-erased video frame.

[0172] If a content-erased video frame is detected, it means that the content of the two adjacent video frames has changed. The change in content means that the content area of ​​the content-erased video frame (i.e., the next video frame) is different from the content area of ​​the previous video frame, and the content area of ​​the previous video frame includes the entire content area of ​​the content-erased video frame (i.e., the next video frame). The previous video frame is then identified as the first target video frame.

[0173] Among the N video frames of the first type of video, if the content area of the (m + 1)-th video frame (the currently displayed video frame) is different from that of the m-th video frame, and the content area of the m-th video frame includes the entire content area of the (m + 1)-th video frame, the (m + 1)-th video frame is a content erasure video frame, and the m-th video frame is the first target video frame. That is to say, if the display device 200 recognizes that the currently displayed video frame is different from the previous video frame in terms of the picture content, and the content area of the previous video frame includes the entire content area of the currently displayed video frame, the previous video frame is determined as the first target video frame. If the currently displayed video frame is different from the previous video frame in terms of the picture content, and the content area of the previous video frame includes the entire content area of the currently displayed video frame, it can be considered that the picture content shown in the currently displayed video frame, compared with the picture content of the previous video frame, is the original picture content, and is obtained by means of deletion or other processing based on the picture content of the previous video frame. It can be said that it is related to the picture content of the previous video frame. The picture content of the currently displayed video frame cannot cover the picture content of the previous video frame. To ensure that the intercepted video frame can include the complete video content, the previous video frame needs to be used as the first target video frame, and the previous video frame is intercepted.

[0174] In some embodiments, to ensure that there is no omission of the first target video frame, the first target video frame also includes the last video frame of the video.

[0175] In an example, taking the video 3 played on a TV as an example for illustration. As Figure 18 shown, the video 3 is a course presented in the form of blackboard writing. The video 3 sequentially includes the 1st to the 7th video frames in the playing order. Among them, the picture content of the 1st video frame includes a circle, the picture content of the 2nd video frame includes a circle and a triangle, the picture content of the 3rd video frame is all blank areas, the picture content of the 4th video frame includes a circle, the picture content of the 5th video frame includes a circle and a line of text, the 6th video frame includes a circle and a line of text (one character 'form' less than that in the 5th video frame), and the 7th video frame includes a circle and a line of text (one character 'circle' more than that in the 6th video frame).

[0176] The TV identifies video 3 as a second-type video. It can use image segmentation to obtain the background area of ​​the video frame and determine whether a content-erased video frame exists based on the ratio of the background area to the screen content area. If the TV is configured to buffer three video frames, let's take the currently displayed third video frame as an example. The TV buffers the third video frame; at this point, the TV has buffered video frames 1 through 3. The TV calculates that the ratio of the background area to the screen content area in the third video frame is 0, which is less than a preset ratio. Therefore, the TV determines that the third video frame is a content-erased video frame. From the buffered first through third video frames, the TV determines the preceding video frame, i.e., the second video frame, as the first target video frame. Referring to the method used to determine the 3rd video frame as the first target video frame, although the content area of ​​the 6th video frame is smaller than that of the 5th video frame, and the content area of ​​the 5th video frame includes the entire content area of ​​the 6th video frame, the 6th video frame is not determined as a content-erased video frame because the ratio of the background area to the content area of ​​the 6th video frame is less than a preset ratio. Therefore, the 5th video frame is not determined as the first target video frame. The 7th video frame, as the last video frame of video 3, is also determined as the first target video frame.

[0177] In some embodiments, the first type of video also includes courses presented in the form of live-action PowerPoint presentations, and the second type of video also includes courses presented in the form of live-action handwriting on a whiteboard. The video content also includes images of people that may affect the display of the content area. For example, a person may obscure the display area of ​​the PowerPoint presentation or the display area of ​​the handwriting on the whiteboard. To ensure the effectiveness of capturing the content area within the video content, it is necessary to capture video frames containing the content area that is not obscured by the image of a person.

[0178] Display device 200 can be configured as follows Figure 19 The process shown below involves acquiring the content of video frames. The specific steps are as follows:

[0179] S1901, identifies human images in the content of video frames.

[0180] The display device 200 can accurately identify human images and non-human images in the screen content based on image recognition technology. The non-human images are the images corresponding to the content area and background area in the screen content.

[0181] S1902, determine the valid video frame based on the positional relationship between the person image and the content area of ​​the video frame.

[0182] A valid video frame is a video frame in which the content area of ​​the image of the person does not overlap with that of the video frame.

[0183] For example, a course presented in the form of handwritten blackboard notes by a real person, such as Figure 20As shown, the video frame contains a content area 2001, a background area 2002, and a person image 2003. The person image 2003 overlaps with the content area 2001, making this video frame invalid. Figure 21 As shown in ①, the video frame contains a content area 2101, a background area 2102, and a person image 2103. The person image 2103 does not overlap with the content area 2101, and the video frame is a valid video frame.

[0184] Invalid video frames cannot display a complete content area to the user, meaning they cannot show any valid video content. Therefore, invalid video frames are discarded. Valid video frames can display a complete content area to the user, meaning they can show any valid video content. Therefore, valid video frames are used as the basis for selecting the first target video frame.

[0185] S1903, remove human images from the content of valid video frames.

[0186] by Figure 21 Taking the video frame shown in ① as an example, the area containing the person's image is cropped from the content of the frame, or the area outside the content area is cropped from the content of the frame, to obtain the frame content after removing the person's image, as shown in ①. Figure 21 As shown in ② (the cropped area is indicated by shading). This avoids the need for the human image to participate in the comparison process of subsequent image content, effectively reducing computational load and improving calculation accuracy.

[0187] S1904, determine the first target video frame based on the content of the two adjacent valid video frames after removing the images of people.

[0188] The content of the valid video frames after removing the images of people is used as the image basis for determining the first target video frame. This can be referred to in steps S901-S902, which will not be repeated here.

[0189] S903, capture the first target video frame.

[0190] In this way, during video playback, the display device 200 can automatically determine and capture the first target video frame without requiring the user to repeatedly input screenshot commands. The first target video frame determined by the display device 200 not only includes the complete video content but also effectively reduces the number of first target video frames, thereby reducing the number of screenshots required by the display device 200 and the storage space occupied by the captured video frames.

[0191] In some embodiments, the display device 200 may be configured as follows: Figure 22 The process shown stores the first target video frame, and the specific steps are as follows:

[0192] S2201, Obtain the link information of the first target video frame.

[0193] The link information of the first target video frame includes the address information of the video to which it belongs (such as network storage address, local storage address, etc.) and the frame number of the first target video frame in the video frame sequence.

[0194] S2202, Generate the link of the first target video frame based on the link information of the first target video frame.

[0195] The link to the first target video frame includes the corresponding link information.

[0196] S2203, store the first target video frame to the specified address, and establish a mapping relationship between the first target video frame and the link of the first target video frame.

[0197] like Figure 8 The screenshot function settings interface shown includes a specified address, such as " / User / Local / bin / XXXXXX". This specified address is where each first target video frame is stored, and this address can be customized by the user.

[0198] In some embodiments, the display device 200 may directly store each first target video frame to the designated address in real time.

[0199] In some embodiments, after capturing all the first target video frames, the display device 200 stores all the first target video frames to a specified address. For example, the display device 200 determines that all the first target video frames have been captured when it detects that the video playback has finished. The display device 200 can display, as shown in the example below. Figure 23 The prompt interface shown indicates that these first target video frames have been stored at a specified address. Users can choose to enter the specified address to view these first target video frames, or they can choose to exit the prompt interface.

[0200] When storing the first target video frame, a mapping relationship is established between the first target video frame and its links. Thus, if a user browses the first target video frame in this storage location and inputs an instruction to the display device 200 to select the first target video frame, the display device 200 will respond to the instruction and play the video to which the first target video frame belongs, starting from the link provided.

[0201] For example: the first target video frame is the third video frame of video 2, the local storage address of video 2 is A, and the storage location of the first target video frame in the display device 200 is B. When the user finds the first target video frame according to the storage location B, if the user inputs the command to select the first target video frame into the display device 200, the display device 200 will find video 2 according to the local storage address A and start playing video 2 from the third video frame.

[0202] Example 2

[0203] Display device 200 activates the semi-automatic screenshot mode in the first screenshot function. In this semi-automatic screenshot mode, while automatically capturing video frames, display device 200 also responds to the user's screenshot command and captures the video frame indicated by the user. The process of display device 200 automatically capturing video frames is described in Example 1 and will not be repeated here. Display device 200 can proceed as follows... Figure 24 The process shown captures the video frame specified by the user. The specific steps are as follows:

[0204] S2401 receives screenshot commands input by the user.

[0205] The screenshot command specifies the second target video frame to be captured.

[0206] S2402, in response to the screenshot command, captures a second target video frame.

[0207] The process by which the display device 200 responds to a screenshot command to capture a second target video frame can be referred to the above description. Figure 5 and Figure 6 The relevant parts will not be elaborated here.

[0208] The following explanation uses video 2 from Example 1 as an example. Figure 25 As shown, if the user instructs the TV to capture the 4th video frame, the TV will capture the 3rd, 5th, and 6th video frames, as well as the 4th video frame, which will be used as the second target video frame (the second target video frame is shown with the background as a shadow).

[0209] In some embodiments, after capturing the second target video frame, the display device 200 stores the second target video frame to a designated address. The process of storing the second target video frame can refer to the process of storing the first target video frame in Embodiment 1, and will not be repeated here. Thus, if the user finds the second target video frame at the designated address and inputs an instruction to the display device 200 to select the second target video frame, the display device 200 will respond to the instruction and play the video to which the second target video frame belongs, starting from the link of the second target video frame.

[0210] Example 3

[0211] Display device 200 enables the second screenshot function, which is the manual screenshot mode of the first screenshot function. In this manual screenshot mode, display device 200 only responds to the user's screenshot command and captures the video frame indicated by the user. The process by which display device 200 responds to the user's screenshot command and captures the video frame indicated by the user can be referred to the above description. Figure 5 and Figure 6 The relevant parts will not be elaborated here.

[0212] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the contents of this disclosure, thereby enabling those skilled in the art to better utilize the embodiments.

Claims

1. A display device, characterized in that, include: The monitor is configured to display the video feed. The controller is configured as follows: After activating the first screenshot function, the screen content of each video frame of the video is captured, and the screen content includes a content area and a background area; A first target video frame is determined based on the content of two adjacent video frames. The two adjacent video frames include a preceding video frame and a following video frame in playback order. If the content area of ​​the following video frame is partially or completely different from the content area of ​​the preceding video frame, the first target video frame is the following video frame. If the content area of ​​the following video frame is different from the content area of ​​the preceding video frame, and the content area of ​​the preceding video frame includes the entire content area of ​​the following video frame, the first target video frame is the preceding video frame. Capture the first target video frame; Specifically, the controller determines the first target video frame based on the content of two adjacent video frames, and is configured as follows: Identify the content type of the video; The first target video frame is determined based on the content type and the content of the two adjacent video frames; Wherein, if the video is a first type of video, the first target video frame is the next video frame among the two adjacent video frames, and the first type of video includes videos presented in the form of playing a slideshow; If the video is a second type of video, the first target video frame is the preceding video frame among the two adjacent video frames, and the second type of video includes videos presented in the form of written blackboard writing; Wherein, if the video is a second type of video, the controller performs the following: determining the first target video frame based on the content type and the content of the two adjacent video frames, configured as follows: Store a specified number of video frames and identify content-erased videos, wherein the preceding video frame of the content-erased video frame includes all the screen content of the content-erased video frame, and the screen content of the content-erased video frame is less than the screen content of the preceding video frame of the content-erased video frame. If the content-erased video frame is identified, the video frame preceding the content-erased video frame in the specified number of stored video frames is determined as the first target video frame.

2. The display device of claim 1, wherein, If the video is a first type of video, the controller performs the following: Based on the content type and the content of the two adjacent video frames, it determines the first target video frame, which is configured as follows: Identify the content of two adjacent video frames; If the content of two adjacent video frames is different, the latter video frame among the two adjacent video frames is determined to be the first target video frame.

3. The display device of claim 2, wherein, The controller is configured to identify whether the content of two adjacent video frames is different. Obtain the image features of the two adjacent video frames; Calculate the matching degree of image features between two adjacent video frames; If the matching degree is greater than or equal to the matching degree threshold, it is determined that the content of the two adjacent video frames is the same; if the matching degree is less than the matching degree threshold, it is determined that the content of the two adjacent video frames is different.

4. The display device of claim 1, wherein, The controller is configured to identify whether a video with content erasure has occurred. Obtain the area ratio of the background region to the content area of ​​the adjacent video frames; If the ratio of the area of ​​the background region to the area of ​​the content in the adjacent video frames is different, determine whether the ratio of the area of ​​the background region to the content in the next video frame in the adjacent video frames is greater than or equal to a preset ratio. If the ratio of the area of ​​the background region to the area of ​​the content in the next video frame is greater than or equal to the preset ratio, the next video frame is determined to be the content-erasing video frame.

5. The display device of claim 1, wherein, The controller determines the first target video frame based on the content of two adjacent video frames, and is configured as follows: Identify human images in the content of the video frame; Based on the positional relationship between the person image and the content area of ​​the video frame, a valid video frame is determined, wherein the valid video frame is a video frame in which the content area of ​​the person image and the video frame do not overlap. Remove the image of the person from the content of the valid video frames; The first target video frame is determined based on the content of the two adjacent valid video frames after removing the images of the people.

6. The display device of claim 1, wherein, The controller is also configured to: Receive a screenshot command input by the user, wherein the screenshot command indicates the second target video frame to be captured; In response to the screenshot command, the second target video frame is captured.

7. The display device of claim 1, wherein, The controller is also configured to: Obtain the link information of the target video frame to be stored. The link information includes the address information of the target video to which the target video frame to be stored belongs and the frame number of the target video frame to be stored in the video frame sequence. The video frame sequence is composed of the video frames of the target video arranged in the playback order. Generate corresponding links based on the link information of the target video frames to be stored; The target video frame to be stored is stored at a specified address, and a mapping relationship is established between the target video frame to be stored and the corresponding link, so that after the target video frame to be stored is selected, the target video is played starting from the target video frame to be stored according to the corresponding link.

8. A method for taking a screenshot of a video frame, the method comprising: Applied to a display device that displays video frames, the method includes: After activating the first screenshot function, the screen content of each video frame of the video is obtained, and the screen content includes a content area and a background area; A first target video frame is determined based on the content of two adjacent video frames. The two adjacent video frames include a preceding video frame and a following video frame in playback order. If the content area of ​​the following video frame is partially or completely different from the content area of ​​the preceding video frame, the first target video frame is the following video frame. If the content area of ​​the following video frame is different from the content area of ​​the preceding video frame, and the content area of ​​the preceding video frame includes the entire content area of ​​the following video frame, the first target video frame is the preceding video frame. Capture the first target video frame; Based on the content of two adjacent video frames, the first target video frame is determined to be: Identify the content type of the video; The first target video frame is determined based on the content type and the content of the two adjacent video frames; Wherein, if the video is a first type of video, the first target video frame is the next video frame among the two adjacent video frames, and the first type of video includes videos presented in the form of playing a slideshow; If the video is a second type of video, the first target video frame is the preceding video frame among the two adjacent video frames, and the second type of video includes videos presented in the form of written blackboard writing; Wherein, if the video is a second type of video, the first target video frame is determined based on the content type and the content of the two adjacent video frames, specifically as follows: Store a specified number of video frames and identify content-erased videos, wherein the preceding video frame of the content-erased video frame includes all the screen content of the content-erased video frame, and the screen content of the content-erased video frame is less than the screen content of the preceding video frame of the content-erased video frame. If the content-erased video frame is identified, the video frame preceding the content-erased video frame in the specified number of stored video frames is determined as the first target video frame.