Subtitle display method and related device
By enabling a local player in AR glasses to play unrelated videos and generate subtitle views, the problem of third-party videos not being able to display subtitles when running in the background is solved, achieving the effect of simultaneously handling other applications and viewing subtitles in AR glasses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-10-23
- Publication Date
- 2026-04-24
AI Technical Summary
In AR glasses on the visionOS platform, third-party videos cannot be displayed in picture-in-picture mode when running in the background, preventing users from simultaneously handling other applications and watching subtitles for third-party videos.
By enabling a local second player to play a second video unrelated to the first video, and extracting the target subtitles from the first video to generate a target subtitle view, this view is overlaid on the picture-in-picture window, thus obscuring the second video.
Even when the local platform does not support third-party videos being displayed in picture-in-picture mode in the background, it can still display the subtitles of third-party videos, meeting the user's multitasking needs.
Smart Images

Figure CN121924314A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of subtitle display technology, and in particular to a subtitle display method and related apparatus. Background Technology
[0002] The VisionOS platform is an augmented reality (AR) glasses application platform that extends traditional displays into virtual ultrawide screens, providing users with a superior visual experience. On the VisionOS platform, local videos can be played in picture-in-picture mode after entering background operation mode, allowing users to perform other tasks simultaneously. For example, in AR glasses on the VisionOS platform, users see a virtual ultrawide screen through the glasses. Before entering background operation, the local video is displayed on the entire screen. Once in background operation (allowing users to watch local videos while handling other applications), the local video is displayed in picture-in-picture mode on a portion of the screen, while other applications, such as e-books, can be displayed in other areas, allowing users to watch local videos and read e-books at the same time.
[0003] However, when playing third-party videos from third-party applications in AR glasses on the visionOS platform, the visionOS platform does not support picture-in-picture display when the third-party video is running in the background. Therefore, users either have to exit the third-party video or watch it in full screen. However, users often need third-party videos to run in the background so that they can at least see the subtitles of the third-party video while handling other applications on the screen. Summary of the Invention
[0004] This disclosure provides a subtitle display method and related apparatus that can still display subtitles of third-party videos when the third-party video is running in the background, even when the local platform does not support the display of third-party videos in picture-in-picture mode.
[0005] According to one aspect of this disclosure, a method for displaying subtitles is provided, comprising:
[0006] When the first video of a third-party application starts playing through the first player of the third-party application, the local second player is enabled to play the second video. The second player cannot play the first video of the third-party application.
[0007] In response to the first player entering background running mode, instruct the second player to enter background running mode in order to display the target picture-in-picture window;
[0008] Obtain the target subtitle from the first video, generate a target subtitle view based on the target subtitle, and overlay the target subtitle view onto the target picture-in-picture window to cover the second video.
[0009] According to one aspect of this disclosure, a subtitle display device is provided, comprising:
[0010] The startup unit is used to enable a local second player to play a second video when the first video of the third-party application starts playing through the first player of the third-party application, wherein the second player cannot play the first video of the third-party application;
[0011] The indicator unit is used to instruct the second player to enter the background running mode in response to the first player entering the background running mode, so as to display the target picture-in-picture window;
[0012] The window overlay unit is used to obtain the target subtitle from the first video, generate a target subtitle view based on the target subtitle, and overlay the target subtitle view onto the target picture-in-picture window to cover the second video.
[0013] Optionally, the indicating unit is specifically used for:
[0014] Use a second player to generate the target picture-in-picture controller;
[0015] In response to the first player entering background running mode, instruct the second player to enter background running mode;
[0016] The second player invokes the target picture-in-picture controller to display the target picture-in-picture window.
[0017] Optionally, the indicating unit is specifically used for:
[0018] The initial picture-in-picture controller is generated using the second player. The initial picture-in-picture controller has display status parameters for the control component view. The initial value of the display status parameters indicates each control component in the display control component view.
[0019] Change the initial value of the displayed status parameter in the initial picture-in-picture controller to the target value. The target value indicates the control components in the hidden control component view, thereby turning the initial picture-in-picture controller into the target picture-in-picture controller.
[0020] Optionally, the window overlay unit is specifically used for:
[0021] Use the target picture-in-picture controller to generate a view of control components that hides each control component;
[0022] Add the target caption to the control component view where all control components are hidden, and you will get the target caption view.
[0023] Optionally, the indicating unit is specifically used for:
[0024] The second player calls the target picture-in-picture controller to obtain the global picture-in-picture window list, which includes multiple candidate picture-in-picture windows.
[0025] From multiple candidate picture-in-picture windows, select the target picture-in-picture window that matches the target subtitle;
[0026] Display the target picture-in-picture window.
[0027] Optionally, the indicating unit is specifically used for:
[0028] Get the size of the first window of the target application window that is to be displayed in parallel with the picture-in-picture window;
[0029] Get the number of characters in the target subtitle;
[0030] Based on the size of the first window and the number of characters, determine the size of the second window of the target picture-in-picture window;
[0031] Based on the first display style of the target application window, determine the second display style of the target picture-in-picture window;
[0032] The target picture-in-picture window is determined from multiple candidate picture-in-picture windows based on the second window size and the second display style.
[0033] Optionally, the aspect ratio of the target caption view is consistent with the aspect ratio of the target picture-in-picture window, and the window overlay unit is specifically used for:
[0034] The target caption view is scaled up so that it matches the size of the target picture-in-picture window, and then the scaled target caption view is overlaid on the target picture-in-picture window to completely cover the second video.
[0035] Optionally, the target caption view is smaller than the target picture-in-picture window, and the background color of the target caption view is the same as the background color of the second video.
[0036] Optionally, the window overlay unit is specifically used for:
[0037] Obtain video frames from the first video and generate a video frame view based on the video frames;
[0038] Overlay the video frame view onto the target picture-in-picture window to cover the second video;
[0039] Obtain the target subtitle from the first video, and generate a target subtitle view based on the target subtitle;
[0040] Overlay the target caption view onto the target picture-in-picture window to partially obscure the video frame view.
[0041] Optionally, the startup unit is specifically used for:
[0042] Retrieve the local video library, which includes multiple local candidate videos;
[0043] Get the primary color tone and duration of the first video;
[0044] Obtain the second primary color tone and second duration of the local candidate videos;
[0045] Based on the first primary color tone, the second primary color tone, the first duration, and the second duration, the second video is determined from multiple local candidate videos;
[0046] Enable the second player to play the second video.
[0047] Optionally, the startup unit is specifically used for:
[0048] Based on the first primary color and the second primary color, determine the first matching degree between the local candidate video and the first video;
[0049] Based on the first duration and the second duration, determine the second matching degree between the local candidate video and the first video;
[0050] Based on the first and second matching scores, the total matching score between the local candidate video and the first video is determined.
[0051] Based on the overall matching degree, the second video is determined from multiple local candidate videos.
[0052] Optionally, the startup unit is specifically used for:
[0053] Obtain multiple first keyframes from the first video;
[0054] Determine the first background color for each first keyframe;
[0055] Determine the first primary color tone based on the first background color of each first keyframe;
[0056] Get the first duration of the first video.
[0057] Optionally, the startup unit is specifically used for:
[0058] Obtain multiple second keyframes from the local candidate video;
[0059] Determine the second background color for each second keyframe;
[0060] Determine the second primary color tone based on the second background color of each second keyframe;
[0061] Get the second duration of the local candidate video.
[0062] Optionally, a second video corresponding to the first video is stored in a third-party application, and the startup unit is specifically used for:
[0063] Download a second video from a third-party application to your local device;
[0064] Enable the second player to play the second video.
[0065] Optionally, the subtitle display device further includes:
[0066] The detection unit is used to detect the first control state of the first player;
[0067] The synchronization unit is used to keep the second control state of the second player synchronized with the first control state.
[0068] Optionally, the subtitle display device further includes:
[0069] The rewind unit is used to obtain the first unplayed duration of the first video if the second player has finished playing the second video but the first player has not finished playing the first video, and to make the second player start playing from the position of the first unplayed duration after rewinding from the end of the second video.
[0070] The skip unit is used to allow the second player to skip the unfinished second video if the second player has not finished playing the second video while the first player has finished playing the first video.
[0071] Optionally, the indicating unit is specifically used for:
[0072] In response to the triggering of the background running control in the first player, instruct the first player to enter background running mode;
[0073] Instruct the second player to enter background operation mode;
[0074] Hide the first player and show the target picture-in-picture window.
[0075] Optionally, the subtitle display device further includes:
[0076] The close unit is used to close the target picture-in-picture window and eliminate the target caption view in response to the triggering of the restore foreground running control associated with the target caption view;
[0077] The display playback unit is used to display the first player and play the first video in the first player.
[0078] Optionally, the display playback unit is specifically used for:
[0079] The first screen where the first video is playing when the first player enters background running mode;
[0080] Get the second interface of the first player before the first interface;
[0081] Display the first player, and then display the second interface within the first player;
[0082] In response to the selection of the first video on the second interface, the first video is played.
[0083] According to one aspect of this disclosure, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described subtitle display method.
[0084] According to one aspect of this disclosure, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the above-described subtitle display method.
[0085] According to one aspect of this disclosure, a computer program product is provided, the computer program product including a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the above-described subtitle display method.
[0086] Although the local platform does not support displaying the first video of a third-party application in picture-in-picture mode when running in the background, it does support displaying the second video played locally in picture-in-picture mode when running in the background. Therefore, in this embodiment, when the first video of a third-party application starts playing through the first player of the third-party application, the local second player is activated to play the local second video, which is synchronized with the first video. The content of this second video can be unrelated to the first video. When the first player playing the first video enters background running mode, the second player playing the second video also enters background running mode. At this time, the local platform will display the local second video in picture-in-picture mode. Although the content of this second video is not the desired content of the first video, this embodiment obtains the target subtitles from the first video, generates a target subtitle view accordingly, and overlays the target subtitle view onto the picture-in-picture, thus covering the second video. Therefore, although the content of the second video is unrelated to the first video, it is covered up, and the subtitles of the first video are still displayed in the picture-in-picture mode. This achieves the effect of displaying the subtitles of the third-party video when it runs in the background, even when the local platform does not support displaying the third-party video played in picture-in-picture mode when running in the background. Attached Figure Description
[0087] Figure 1 This is a system framework diagram showing the application of the subtitle display method in the embodiments of this disclosure;
[0088] Figures 2A-2BThe caption display method according to this disclosure is applied to a scenario where gesture-controlled extended reality (XR) glasses are used.
[0089] Figures 3A-3B The subtitle display method disclosed herein is applied to scenarios where XR glasses are controlled by a mouse or keyboard;
[0090] Figures 4A-4B This describes the application of the subtitle display method according to this disclosure to scenarios where XR glasses are controlled by the eyes;
[0091] Figure 5 This is a flowchart of a subtitle display method according to an embodiment of the present disclosure;
[0092] Figure 6 yes Figure 5 A detailed flowchart of step 510;
[0093] Figure 7 yes Figure 5 Another specific flowchart for step 510;
[0094] Figure 8 yes Figure 5 Another specific flowchart for step 510;
[0095] Figure 9 This is a flowchart of a caption display method according to another embodiment of the present disclosure;
[0096] Figure 10 This is a flowchart of a caption display method according to another embodiment of the present disclosure;
[0097] Figure 11 yes Figure 5 A detailed flowchart of step 520;
[0098] Figure 12 yes Figure 11 A detailed flowchart of step 1110;
[0099] Figure 13 yes Figure 11 A detailed flowchart of step 1130;
[0100] Figure 14 This is a diagram illustrating a global list of picture-in-picture windows;
[0101] Figure 15 yes Figure 13 Another specific flowchart for step 1320;
[0102] Figure 16 yes Figure 13 Another specific flowchart for step 1320;
[0103] Figure 17 yes Figure 13 Another specific flowchart for step 1320;
[0104] Figure 18 yes Figure 5 Another specific flowchart for step 520;
[0105] Figure 19 yes Figure 5 A detailed flowchart of step 530;
[0106] Figure 20 It is a manifestation Figure 5 Flowchart of another specific process for subtitle display in step 530;
[0107] Figure 21 yes Figure 5 Another specific flowchart for step 530;
[0108] Figure 22 This is a flowchart of a caption display method according to another embodiment of the present disclosure;
[0109] Figure 23 yes Figure 22 A detailed flowchart of step 550;
[0110] Figures 24A-24F It shows Figure 23 An example interface change diagram during the process;
[0111] Figure 25 A flowchart illustrating a detailed embodiment of this disclosure is shown;
[0112] Figure 26 This is a block diagram of a subtitle display device according to an embodiment of the present disclosure;
[0113] Figure 27 A structural diagram of XR glasses applying a caption display method according to an embodiment of the present disclosure is shown. Detailed Implementation
[0114] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.
[0115] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:
[0116] Third-party applications: Third-party applications refer to applications developed by third-party developers or teams, typically not by the operating system provider. They can be plugins, tools, applets, etc., used to enhance or extend the functionality of existing software. A key characteristic of third-party applications is that they are usually created by non-official or non-native developers, meaning they may not possess all the functionality and security features of the original software.
[0117] A media player is software that can play video or audio files stored in digital signal form. Most media players have built-in decoders to restore compressed media files and usually include a complete set of algorithms for frequency conversion and buffering to ensure a smooth playback experience.
[0118] Picture-in-Picture (PiP): Picture-in-Picture (PiP) is a common feature in modern web pages and applications that allows users to watch video content in a small floating window while performing other tasks in the main window. This feature is especially suitable for users who need to multitask, such as browsing the web while watching videos. The floating window can contain another video, image, or animation, which can be overlaid, blended, or separated from the main window, thus providing viewers with a richer visual experience.
[0119] Currently, when playing third-party videos from third-party applications in AR glasses on the visionOS platform, the visionOS platform does not support picture-in-picture display of the third-party video while it is running in the background. Therefore, users either have to exit the third-party video or watch it in full screen. However, users often need third-party videos to run in the background so that they can see at least the video's subtitles while handling other applications on the screen. To solve this problem, this disclosure provides a subtitle display method that can still display the subtitles of third-party videos when they are running in the background, even when the local visionOS platform does not support picture-in-picture display of third-party videos in the background.
[0120] System architecture and scenario description of the embodiments disclosed herein
[0121] Figure 1 This is a system architecture diagram of the caption display method according to an embodiment of the present disclosure. The system architecture includes extended reality (XR) glasses 110, Internet 120, gateway 130, XR glasses application platform (such as visionOS platform) server 140, and third-party application server 150, etc.
[0122] XR Glasses 110 is a general term for virtual reality (VR) glasses, augmented reality (AR) glasses, and mixed reality (MR) glasses.
[0123] AR glasses are devices used to overlay virtual images onto the real world, allowing users to see a mixed view of reality and virtuality simultaneously. The relative position of the virtual images moves with the device. For example, when watching a video through a third-party app using AR glasses, the user sees the real world through the glasses like with normal glasses, while simultaneously viewing a virtual screen on which the third-party app plays the video. The AR glasses see the effect of the real world overlaid with the video played through the third-party app. As the user moves, the AR glasses move with them, and the relative positions of the virtual screen and real objects change accordingly. For example, a user might see a virtual screen and a real chair through AR glasses. When the user moves away from the chair, they might no longer see the chair through the AR glasses, but instead see a real window and the virtual screen.
[0124] VR glasses are devices that allow users to see only the virtual world and not the real world. For example, when a user wears VR glasses, they only see the virtual world of a game through the glasses, without any real-world scenes superimposed on them.
[0125] MR (Mixed Reality) glasses are used to overlay virtual images onto the real world, allowing users to simultaneously see a blend of reality and virtuality, with the relative positions of the virtual images remaining constant regardless of device movement. For example, a virtual cat might be positioned around a real stove, and a virtual dog might be lying on a real chair. When anyone wearing MR glasses faces the real stove, the virtual cat will remain beside it, their relative positions constant. Similarly, when anyone wearing MR glasses faces a real chair, the virtual dog will remain on it, their relative positions constant.
[0126] XR glasses 110 can communicate with the Internet 120 via wired or wireless means to exchange data.
[0127] A server refers to a computer system that can provide certain services to XR glasses 110. Third-party application server 150 provides third-party videos, allowing users to download these videos from XR glasses 110 and watch them locally. XR glasses application platform server 140 supports displaying local videos from the XR glasses application platform in picture-in-picture mode on XR glasses 110 when running in the background, but does not support displaying third-party videos from third-party application server 150 in picture-in-picture mode on XR glasses 110 when running in the background.
[0128] Compared to ordinary XR glasses, servers have higher requirements in terms of stability, security, and performance. A server can be a single high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a single high-performance computer (such as a virtual machine), or a combination of portions of multiple high-performance computers (such as virtual machines).
[0129] Gateway 130, also known as an internetwork connector or protocol converter, is a computer system or device that enables network interconnection at the transport layer and acts as a translator. It bridges the gap between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateway 130 can also provide filtering and security functions. Messages sent from XR glasses 110 to XR glasses application platform server 140 must be routed through gateway 130 to the corresponding XR glasses application platform server 140. Similarly, messages sent from XR glasses application platform server 140 to XR glasses 110 must also be routed through gateway 130 to the corresponding XR glasses 110.
[0130] The embodiments disclosed herein can be applied in various scenarios, such as Figures 2A-2B The scene shown illustrates gesture control of XR glasses. Figures 3A-3B The scene shown illustrates how to control XR glasses with a mouse or keyboard. Figures 4A-4B The scene shown depicts eyes controlling XR glasses.
[0131] (a) Scenarios of controlling XR glasses with gestures
[0132] When a user wears XR glasses to watch a third-party video, the user sees a superposition of the virtual screen and the real world through the XR glasses 110. Because the user sees a superposition of the virtual screen and the real world, the user can see their own hands, therefore... Figure 2A As shown, users can launch third-party applications and play third-party videos in full-screen mode using gestures on the virtual screen. If other operations need to be performed during the playback of a third-party video, such as reading an e-book or listening to music, a need arises where users can perform other operations on the screen while continuing to access information about the third-party video.
[0133] If third-party videos are displayed in picture-in-picture mode, users can perform other operations on the area of the screen where the third-party video is not displayed. However, the XR glasses application platform 140 does not support the display of third-party videos from the third-party application server in picture-in-picture mode on the XR glasses 110 when running in the background. Therefore, how to allow users to perform other operations on the screen while still being able to view information about the third-party video becomes a problem. Figure 2B In this embodiment of the disclosure, users can see both subtitle information of third-party videos and an area for performing other operations on the screen. How this is achieved will be described in detail below.
[0134] (ii) Scenarios where XR glasses are controlled by mouse or keyboard
[0135] Figures 3A-3B This refers to scenarios where users control XR glasses using a mouse or keyboard instead of gestures. Figures 3A-3B and Figures 2A-2B Apart from the different control methods, they are otherwise similar. Figures 2A-2B They are basically the same. To save space, I will not go into details.
[0136] (III) Scenarios where XR glasses are controlled by the eyes
[0137] Figures 4A-4B This describes a scenario where the user controls XR glasses through the glasses themselves. A pupil position detector is located on the side of the XR glasses facing the user's eyes. This detector senses movement in the pupil position when the user's gaze changes. Simultaneously, the XR glasses have multiple cameras positioned at different locations to capture 2D images from their respective angles. An internal processor in the XR glasses generates a 3D map of the environment in front of the user based on the 2D images captured by the various cameras. This processor maps the pupil position movements sensed by the pupil position detector onto the 3D map, thereby determining the position where the user's gaze is focused on the screen, and thus generating the corresponding action. Figures 4A-4B and Figures 2A-2B Apart from the different control methods, they are otherwise similar. Figures 2A-2B They are basically the same. To save space, I will not go into details.
[0138] General Description of Embodiments in this Disclosure
[0139] According to one embodiment of this disclosure, a method for displaying subtitles is provided.
[0140] The subtitle display method refers to the method used to display subtitles in a video. It can be performed by devices such as XR glasses.
[0141] like Figure 5As shown, a subtitle display method according to an embodiment of this disclosure includes:
[0142] 510. When the first video of a third-party application starts playing through the first player of the third-party application, the local second player is enabled to play the second video. The second player cannot play the first video of the third-party application.
[0143] 520. In response to the first player entering background running mode, instruct the second player to enter background running mode in order to display the target picture-in-picture window;
[0144] 530. Obtain the target subtitle from the first video, generate a target subtitle view based on the target subtitle, and overlay the target subtitle view onto the target picture-in-picture window to cover the second video.
[0145] First, a brief description of steps 510-530 above will be given.
[0146] In step 510, "third-party application" refers to an application provided by a platform other than the XR glasses application platform, such as the A video application on the A video platform. "First player" refers to the player used to play the A video application on the A video platform. When a third-party application is downloaded to the XR glasses, the first player included with the third-party application is also downloaded to the XR glasses by default. When a user selects a video in the XR glasses, that video is played by default using the first player of the third-party application. "First video" refers to a third-party video within a third-party application. When a user enters a third-party application and selects to watch a specific first video, the third-party application sends a video request to its server. The server then sends the video frames of the first video to the third-party application, which plays them using the first player.
[0147] The second player refers to the player native to the XR glasses, downloaded from the XR glasses application platform server. This second player is pre-installed when the XR glasses are factory-configured. It plays local videos downloaded from the server, i.e., the second video. During initialization, it is specified that it cannot play the first video from third-party applications downloaded from third-party application servers. A key feature of the second player is that even when it enters background mode, the second video can still play in a picture-in-picture window. This allows the user to perform other operations on the virtual screen of the XR glasses, such as reading ebooks, while the second video is displayed in picture-in-picture mode.
[0148] In this step, once the first video from a third-party application starts playing on the XR glasses using the first player, a second local video is immediately launched using the second player. This second video can be different from the first video, such as a completely black video. At this time, because the first player is still running in foreground mode, the first video covers the entire virtual screen and is not displayed in picture-in-picture mode. Therefore, although the second video is playing, it is not visible on the virtual screen.
[0149] In step 520, the user instructs the first player playing the first video to enter background running mode.
[0150] Background operation mode refers to a mode where the video runs in the background while other applications are displayed in the foreground. The reason a video enters background operation mode is that the user opens other applications while the video is playing, forcing the video to run in the background. For example, when playing a third-party video, if the user opens an e-book, the e-book enters foreground operation mode, and the third-party video enters background operation mode. In related technologies, on XR glasses application platforms, once a third-party video enters background operation, it cannot be displayed in picture-in-picture mode.
[0151] like Figures 2A-2B , Figures 3A-3B , Figures 4A-4B As shown, users can open another application for other operations, such as an e-book reader, through gestures, mouse, keyboard, or eye movements. Figure 2B , Figure 3B and Figure 4B (On the left side of the virtual screen), the first player enters background mode, while another application enters foreground mode. Due to XR glasses platform regulations, the first video from a third-party application cannot be displayed in picture-in-picture mode after entering background mode, and therefore it is no longer displayed. At this point, the user cannot see any information about the first video on the virtual screen. However, it is possible to have a second player playing a locally sourced second video also enter background mode when the first player enters background mode. Since the locally sourced second video can be displayed in picture-in-picture mode when in background mode, although the first video cannot be seen on the virtual screen, the second video can be seen in picture-in-picture mode. Figure 2B , Figure 3B and Figure 4B (Right side of the virtual screen). The target picture-in-picture window is a window that displays the second video in picture-in-picture format.
[0152] In step 530, the target subtitle is the subtitle in the first video. The content of the first video needs to be parsed to obtain the target subtitle. Subtitles can be divided into hard subtitles and soft subtitles. Hard subtitles, also known as embedded subtitles, refer to subtitles that are compressed into the same video file along with the video and audio content, so the subtitle information is naturally included in the video frame during playback. Soft subtitles, on the other hand, are subtitles that are separated from the video file and are rendered separately during playback; the subtitles are not displayed in the video frame itself. In this embodiment, when the target subtitle of the first video is a soft subtitle, it is directly extracted from the subtitle file outside the video file. When the target subtitle of the first video is a hard subtitle, the content of the first video needs to be parsed, for example, using a video-to-text tool or an Optical Character Recognition (OCR) tool to recognize the video content and extract the target subtitle.
[0153] The target caption view is a view that contains the target caption. Its main body is the target caption, and its background can be transparent, a certain color such as pure black, or an image such as a picture of blue sky and white clouds.
[0154] When the target caption view has a transparent background, the target caption is overlaid on top of the transparent background view frame, thus generating the target caption view. Then, the target caption view is overlaid on the target picture-in-picture window. Because the background of the target caption view is transparent, the portion of the target picture-in-picture window without the target characters will be fully exposed to the second video.
[0155] When the target caption view has a background of a certain color, the target caption is overlaid on top of the background view frame of that color, thus generating the target caption view. At this time, because the background of the target caption video is colored, it will obscure part of the target picture-in-picture window it covers.
[0156] When the target caption view is a view of an image, the target caption is overlaid on top of the background view frame of that image, thus generating the target caption view. At this point, because the background of the target caption video is an image, it will obscure part of the target picture-in-picture window it covers.
[0157] In one embodiment, the size of the target caption view can be the same as the size of the target picture-in-picture window. This way, the target caption view will completely obscure the second video. For example, if the caption in the first video is "Have you gone out yet?" and the second video is a completely black video, the target caption view is a box containing the caption "Have you gone out yet?", which will completely obscure the black video. Only the box containing "Have you gone out yet?" will be displayed on the virtual screen, and the black video will not be shown at all.
[0158] In another embodiment, the target caption view is smaller than the target picture-in-picture window. This way, the target caption view will only obscure a portion of the second video. In the example above, assuming the target caption video only obscures the lower rectangular area of the second video, a box containing "Have you gone out?" will be displayed in the lower rectangular area of the pure black video, while the upper part remains pure black.
[0159] Because the target caption view is overlaid on the target picture-in-picture window in this step, even though the content of the second video in the target picture-in-picture window may be unrelated to the first video, the second video is covered by the target caption view of the target caption extracted from the first video. The picture-in-picture window still displays the target caption of the first video. This achieves the effect of displaying the caption of the first video when the first video of a third-party application is running in the background, even when the XR glasses application platform does not support the first video of a third-party application running in the background in picture-in-picture mode.
[0160] Steps 510-530 are described in detail below.
[0161] Detailed description of step 510
[0162] In step 510, when the first video of the third-party application starts playing through the first player of the third-party application, the local second player is enabled to play the local second video.
[0163] Playing the local second video immediately after starting the first video from a third-party application is to ensure synchronization between the two videos. If the first video finishes playing but the second video doesn't start immediately, the target subtitle view has no basis for being attached during the time interval between the first and second video playback. This prevents the subtitles from still being displayed in the first video even when the third-party application is running in the background.
[0164] In one embodiment, the first video of a third-party application that the user wants to play may be different, but a unique local second video is set for each different first video, such as a black video, that is, the target subtitle view with a black background is used to cover the target picture-in-picture window.
[0165] In another embodiment, considering the aforementioned possibility that the target caption view may have a transparent background, in which case the second video will be exposed in the target picture-in-picture window. In this case, if the main color tone of the second video is similar to that of the first video, it will give the user the feeling that the target captions of the first video are being displayed while the first video is playing, improving user-friendliness and increasing user engagement.
[0166] In this scenario, XR glasses can store multiple local candidate videos and select a second video from the multiple local candidate videos that matches the dominant color tone of the first video, based on the dominant color tone of the first video.
[0167] In this embodiment, such as Figure 6 As shown, step 510 may include:
[0168] 610. Obtain the local video library, which includes multiple local candidate videos;
[0169] 620. Obtain the primary color tone of the first video;
[0170] 630. Obtain the second primary color tone of the local candidate video;
[0171] 640. Based on the first primary color and the second primary color, determine the second video from multiple local candidate videos;
[0172] 650. Enable the second player to play the second video.
[0173] The following is a detailed description of steps 610-650 above.
[0174] In step 610, the local video library refers to the video library used to provide the second video. The local candidate video refers to a video that can be used as the second video.
[0175] In one embodiment, the local video library is a video library pre-configured when the XR glasses are initialized.
[0176] In one embodiment, the local video library is a collection of locally saved videos that the user has historically stored on the XR glasses. These local videos were downloaded from third-party application servers and stored on the XR glasses to form the local video library.
[0177] In one embodiment, the local video library is a library formed by videos imported into the XR glasses by the user using an import tool. For example, the user imports a completely black video into the XR glasses' memory, then imports a completely white video into the XR glasses' memory, and so on. The collection of completely black videos, completely white videos, etc., constitutes the local video library.
[0178] In step 620, the first primary color tone of the first video is obtained.
[0179] The primary color tone refers to the dominant color in the first video.
[0180] In one embodiment, the first video itself contains an attribute field, and the attribute field contains a first primary color tone. Therefore, the first primary color tone can be obtained by reading the attribute field.
[0181] In one embodiment, the primary color tone can be determined using keyframes.
[0182] Specifically, in one implementation, multiple first keyframes of the first video can be acquired first, then the first background color of each first keyframe can be determined, and finally the first primary color tone of the first video can be determined based on the first background color of each first keyframe. A first keyframe is a representative frame in the first video. First keyframes can be extracted at equal time intervals at different points in time. For example, for a 20-second first video, a first keyframe can be extracted every 2 seconds, resulting in 10 first keyframes. Alternatively, first keyframes can be extracted based on factors such as the movement of people or changes in the background within the video frame. For example, the video frame where the movement of people is first detected can be used as a first keyframe, and the video frame where the movement of people stops can also be used as a first keyframe, and so on.
[0183] The primary background color is the most prominent background color in the first keyframe. For example, if the background in the first keyframe is a blue sky with light clouds, then the primary background color is blue.
[0184] The implementation of determining the first primary color tone of the first video based on the first background color of each first keyframe may include: determining the number of keyframes corresponding to each first background color based on the first background color of each keyframe; and determining the first background color with the most keyframes as the primary color tone of the first video.
[0185] For example, if the first video has 20 keyframes, of which 15 are for black, 4 are for white, and 1 is for red, then the background color with the most occurrences, namely black, is determined as the primary color of the first video.
[0186] The advantage of determining the primary color tone through multiple keyframes is that the primary color tone can be determined comprehensively based on multiple keyframes located in different positions in the first video, so that the obtained primary color tone can truly reflect the overall situation of the first video, improving the accuracy and flexibility of determining the primary color tone.
[0187] Another way to determine the first primary color tone using keyframes is to use either the start keyframe or the end keyframe. Specifically, determine the first background color of the start keyframe in the first video and use it as the first primary color tone of the first video; or, determine the first background color of the end keyframe in the first video and use it as the first primary color tone of the first video.
[0188] The start keyframe is the first keyframe in the first video, and the end keyframe is the last keyframe in the first video. The advantage of determining the primary color tone solely through the start or end keyframe is that the processing is simpler, has less processing overhead, and the primary color tone of the entire first video can be roughly determined from the start or end keyframe, making the entire processing process roughly accurate.
[0189] In step 630, the second primary color refers to the dominant color in the local candidate video. The implementation process of this step is similar to that of step 620, so it will not be described in detail.
[0190] In step 640, in one embodiment, the first primary color and the second primary color are a certain color. In this case, a local candidate video whose second primary color is exactly the same as the first primary color can be used as the second video. For example, if the first primary color is red, and the second primary color of a certain local candidate video is also red, that local candidate video can be used as the second video.
[0191] In one embodiment, the first primary chromaticity and the second primary hue are colors with chromaticity. For example, light red and dark red are not the same color. The matching degree between each second primary hue and the first primary hue can be determined separately; the local candidate video corresponding to the second primary hue with the highest matching degree is determined as the second video.
[0192] Matching degree refers to the degree to which hues match, and it can be expressed as similarity or color difference. Similarity is the degree of similarity between hues. Color difference is the distance between hues on the color spectrum.
[0193] When matching is expressed as similarity, a preset similarity table can be consulted to obtain the similarity between two color tones. The local candidate video corresponding to the second dominant color tone with the highest similarity is determined as the second video. For example, if the first dominant color tone is dark red, and the second dominant colors of the three local candidate videos are bright red, rooster comb red, and purplish-red, consulting the similarity table shows that the similarity between bright red and dark red is 0.9, the similarity between rooster comb red and dark red is 0.8, and the similarity between purplish-red and dark red is 0.6. Therefore, the local candidate video with bright red as the second dominant color tone is determined as the second video.
[0194] When matching is expressed as color difference, the distance between two hues can be obtained from the color spectrum as the color difference. The local candidate video corresponding to the second dominant hue with the smallest color difference is determined as the second video. For example, if the first dominant hue is deep red, and the second dominant hues of the three local candidate videos are bright red, rooster comb red, and magenta, respectively, the distance between bright red and deep red is the smallest according to the color spectrum. Therefore, the local candidate video with bright red as the second dominant hue is determined as the second video.
[0195] The advantage of determining the second video based on the matching degree between the second primary color and the first primary color is that the second video determined in this way is more accurate because the calculation of the matching degree between the colors better reflects the similarity of the viewing experience of different videos.
[0196] In step 650, the second player is activated to play the second video. Since the first player is currently playing the first video in full-screen mode in the foreground, the second player can play silently in an invisible area, such as hidden behind the first video.
[0197] The advantage of steps 610-650 is that the second video can be selected based on the second primary color and the first primary color, which reduces the difference in overall visual appearance between the second video and the first video, and makes the final generated target subtitles blend more naturally and smoothly with the second video.
[0198] Besides comparing the primary color tone to determine the second video, another method is to compare its duration. Since once the first video enters background operation mode, the target subtitle view based on the first video should be overlaid on the second video, obscuring it. If the second video has already finished playing, this overlay cannot be achieved, requiring post-processing as described later. Therefore, it's best to select a second video with a duration approximately similar to the first video.
[0199] In another embodiment, such as Figure 7 As shown, step 510 may include:
[0200] 710. Obtain the local video library, which includes multiple local candidate videos;
[0201] 720. Obtain the first duration of the first video;
[0202] 730. Obtain the second duration of the local candidate videos;
[0203] 740. Based on the first duration and the second duration, determine the second video from multiple local candidate videos;
[0204] 750. Enable the second player to play the second video.
[0205] The following is a detailed description of steps 710-750 above.
[0206] Step 710 is the same as step 610, and will not be repeated here.
[0207] In step 720, the first duration is the total playback duration of the first video.
[0208] In one embodiment, the first video includes an attribute field containing the total playback duration, and the first duration can be obtained by reading the attribute field.
[0209] In one embodiment, the first duration of the first video can be estimated by the size of the first video. The size of the first video is substantially proportional to the first duration. The conversion relationship between the size of the first video and the first duration can be obtained in advance, and the first duration can be calculated based on this conversion relationship and the size of the first video.
[0210] Step 730 is similar to step 720, so it will not be described in detail.
[0211] In step 740, in one embodiment, a local candidate video with a second duration equal to the first duration can be used as the second video. For example, if the first duration is 51.7s, there are three local candidate videos with second durations of 49.2s, 20s, and 51.7s respectively, then the third local candidate video can be used as the second video.
[0212] In one embodiment, a second video can be randomly selected from local candidate videos whose absolute difference between the second and first durations is less than a predetermined absolute value threshold. Continuing with the previous example, if the predetermined absolute value threshold is 3 seconds, either the first or third local candidate video can be used as the second video.
[0213] In one embodiment, the local candidate video with the smallest absolute difference between the second duration and the first duration can be selected as the second video. For example, if the first duration is 51.7s, there are three local candidate videos with second durations of 49.2s, 20s, and 55.7s, then the first local candidate video with the smallest absolute difference between its second duration and the first duration can be selected as the second video.
[0214] Step 750 is similar to step 650, so it will not be described in detail.
[0215] The advantage of steps 710-750 is that it selects a second video whose duration is as close as possible to that of the first video, thus reducing the chance that the second video will have ended when the target subtitle view of the first video needs to cover the second video.
[0216] The above embodiments filter local candidate videos based on two aspects: dominant color tone and duration. Alternatively, both dominant color tone and duration can be considered for filtering. The specific process is as follows:
[0217] In another embodiment, such as Figure 8 As shown, step 510 may include:
[0218] 810. Obtain the local video library, which includes multiple local candidate videos;
[0219] 820. Obtain the primary color tone and primary duration of the first video;
[0220] 830. Obtain the second primary color tone and second duration of the local candidate videos;
[0221] 840. Based on the first primary color tone, the second primary color tone, the first duration, and the second duration, determine the second video from multiple local candidate videos;
[0222] 850. Enable the second player to play the second video.
[0223] The following is a detailed description of steps 810-850 above.
[0224] Step 810 is similar to step 610 and will not be repeated.
[0225] Step 820 is equivalent to a combination of steps 620 and 720, and will not be described in detail here.
[0226] Step 830 is equivalent to a combination of steps 630 and 730, and will not be described in detail here.
[0227] One implementation of step 840 is to first filter out a selected video from multiple local candidate videos based on a first primary color tone and a second primary color tone, and then determine a second video from the selected videos based on a first duration and a second duration.
[0228] When selecting the final video from multiple local candidate videos based on a first and second primary color tone, as mentioned earlier, one can select local candidate videos whose second primary color tone is completely identical to the first primary color tone, and use these as the final video. Alternatively, one can select the final video from multiple local candidate videos based on the matching degree between the second and first primary color tones. These methods have been discussed earlier and will not be repeated here.
[0229] When determining the second video based on the first and second durations from the filtered videos, as mentioned earlier, one can select the video whose second duration is exactly equal to the first duration from the filtered videos as the second video. Alternatively, one can select the video with the smallest absolute value of the difference between the second and first durations from the filtered videos as the second video. Another option is to randomly select any video from the filtered videos whose absolute value of the difference between the second and first durations is less than a predetermined absolute value threshold as the second video. These methods have been discussed above and will not be repeated here.
[0230] The advantage of this embodiment is that by considering the influence of the main color tone and duration on the determination of the second video separately, the determined second video meets the conditions in terms of both main color tone and duration, making the factors considered in the determination of the second video more comprehensive and improving the quality of the determined second video.
[0231] In another embodiment, the process of determining the second video may specifically include:
[0232] Based on the first primary color and the second primary color, determine the first matching degree between the local candidate video and the first video;
[0233] Based on the first duration and the second duration, determine the second matching degree between the local candidate video and the first video;
[0234] Based on the first and second matching scores, the total matching score between the local candidate video and the first video is determined.
[0235] Based on the overall matching degree, the second video is determined from multiple local candidate videos.
[0236] The closer the primary color tone is to the secondary color tone, the smaller the difference between the local candidate video and the primary video, and the higher the match. Conversely, the greater the difference between the primary and secondary color tones, the greater the difference between the local candidate video and the primary video, and the lower the match. Similarly, the closer the primary and secondary video durations are to each other, the smaller the difference between the local candidate video and the primary video, and the higher the match. Conversely, the greater the difference between the primary and secondary durations, the greater the difference between the local candidate video and the primary video, and the lower the match.
[0237] The first matching degree refers to the matching degree between the local candidate video and the first video mentioned in step 640. Its calculation method has been discussed previously; it is a matching degree from the perspective of the dominant color tone. The second matching degree is the matching degree between the local candidate video and the first video from the perspective of duration. The second matching degree can be determined based on the absolute value of the duration difference between the local candidate video and the first video, specifically using a lookup table or formula method. The smaller the absolute value of the duration difference between the local candidate video and the first video, the greater the second matching degree. Table 1 shows the correspondence between the absolute value of the duration difference and the second matching degree.
[0238] Range of absolute values of time difference Second matching degree 0-5 seconds 100 5-10 seconds 90 10-15 seconds 80 15-20 seconds 70 …… ……
[0239] Table 1
[0240] For example, the first video has a duration of 300 seconds, and the second candidate video has a duration of 282 seconds. The absolute value of the duration difference is 18 seconds. According to the table, the corresponding second matching degree is 70.
[0241] When using the formula method, Formula 1 shows the formula for how the second degree of matching changes with the absolute value of the time difference:
[0242] W=ba*ΔS Formula 1
[0243] In the above formula, W represents the second matching degree, ΔS is the absolute value of the duration difference between the first and second durations, a represents the first coefficient that changes the second matching degree with the absolute value of the duration difference, and b is a constant. For example, a is 1.5 and b is 100. If ΔS is 20 seconds, we can obtain W = 70.
[0244] Understandably, when determining the total match between the local candidate video and the first video based on the first and second match scores, an averaging or weighted averaging method can be used. When using the weighted averaging method, different weights are set for the first and second match scores. These weights reflect the different impacts of the dominant color tone and duration on the selection of local candidate videos, making the selection of local candidate videos more flexible and allowing for the selection of a suitable second video based on the actual situation, thus improving the universality of the scenario.
[0245] Step 850 is similar to step 750, so it will not be described in detail.
[0246] The advantage of steps 810-850 is that it can combine the matching degree of the two dimensions of duration and main color tone to calculate the total matching degree, thereby determining the second video. Compared with considering duration and main color tone separately, it pays more attention to the intrinsic relationship between the two dimensions, thus improving the accuracy of determining the second video.
[0247] In the above embodiment, the second video is selected from locally stored candidate videos in a local video library. However, in another embodiment, the second video may not be obtained locally, but downloaded from a third-party application. In this embodiment, step 510 includes: downloading the second video from the third-party application to the local device; and enabling a second player to play the second video.
[0248] In one embodiment, the second video is downloaded along with the third-party application from the third-party application server when the third-party application is downloaded to the XR glasses. This second video is suitable for all first videos on the third-party application, such as a completely black video. When a user starts playing any of the first videos from the third-party application on the first player, the second video can be used as the object covered by the subsequently generated target subtitle video. Note that once the second video is downloaded to and stored in the XR glasses, it becomes a local video of the XR glasses, and when played using a local second player and running in the background, it can be displayed in picture-in-picture mode.
[0249] In another embodiment, the second video is downloaded to the XR glasses before the first video is played. In this case, each first video can correspond to a unique second video. The first and second videos are stored correspondingly in a third-party application. For example, if the primary color of the first video A is black, the corresponding second video is designated as a pure black video. The first video A and the pure black video are stored correspondingly in the third-party application. After the user selects the first video on the XR glasses through the third-party application, the XR glasses send a notification to the third-party application server. The third-party application server first sends the stored second video corresponding to the first video to the XR glasses for download, and then sends the video frames of the first video for the user to watch. It should be noted that after the second video is downloaded to and stored in the XR glasses, it becomes a local video of the XR glasses. When it is played using a local second player and enters background running mode, it is allowed to be displayed in picture-in-picture mode.
[0250] The advantage of downloading the second video from a third-party application to the local device is that, since the second video is downloaded from the third-party application, it does not occupy the usual storage space of the XR glasses, and the second video may be bound to the first video in the third-party application, making it easy to personalize.
[0251] The above embodiments have described step 510 in detail, such as... Figure 9 As shown, after step 510, the subtitle display method of this embodiment may further include:
[0252] 511. Detect the first control state of the first player;
[0253] 512. Synchronize the second control state of the second player with the first control state.
[0254] In step 511, the first control state refers to the state in which the first player plays the first video, such as paused playback, speed-up playback, slow-down playback, or stopped playback.
[0255] In one embodiment, the first player has a first control state indicator. The first control state indicator indicates a first control state and updates in real time based on the state of the first player playing the first video. Therefore, the first control state of the first player can be detected by reading the first control state indicator. For example, when the first player pauses playback of the first video, the first control state indicator will indicate pause playback. Reading the first control state indicator allows detection that the first control state is paused playback.
[0256] In another embodiment, the first control state of the first player can be detected by the second player periodically sending a first control state polling signal to the first player. For example, the second player sends the first control state polling signal to the first player periodically every 0.5 seconds. After receiving the first control state polling signal, the first player sends a first control state response signal to the second player, which carries the first control state. For example, when the first player is in pause playback, it sends a first control state response signal to the second player, which carries the pause playback status.
[0257] In step 512, the second control state refers to the state in which the second player plays the second video, such as paused playback, speed-up playback, slow-down playback, or stopped playback.
[0258] In one embodiment, a second control state indicator is provided in the second player, corresponding to the first control state indicator. The second control state indicator indicates a second control state, and the second player plays a second video according to this second control state. If the first control state of the first player is detected via the first control state indicator, once the first control state is detected, the second control state indicated by the second control state indicator is immediately changed to match the first control state, and the second player plays the second video according to this second control state. For example, if the first control state is detected as paused playback, the second control state indicated by the second control state indicator in the second player is immediately changed to paused playback, and the second player pauses playback.
[0259] In another embodiment, when the second player periodically sends a first control state polling signal to the first player, the second player obtains the first control state from the first control state response signal sent by the first player, and makes the second control state of the second player consistent with the first control state, so that the second player plays the second video according to the second control state. For example, if the first control state obtained from the first control state response signal is pause playback, then the second player changes the second control state to pause playback, thereby pausing the playback of the second video.
[0260] The purpose of steps 511-512 is that, although there is an intentional limit to ensure the duration of the first and second videos is as consistent as possible, it's still possible that one video plays faster than the other, or one video pauses or freezes while the other continues playing, preventing the two videos from completing simultaneously. If the first video hasn't finished playing while the second has, and the first player playing the first video enters background mode, requiring the second player to also enter background mode to display the target picture-in-picture window, the second player cannot enter background mode because the second video has already finished playing. The synchronization in steps 511-512 can reduce the occurrence of this situation.
[0261] In one embodiment, such as Figure 10 As shown, the subtitle display method may also include the following steps:
[0262] 513. If the second player has finished playing the second video, but the first player has not finished playing the first video, obtain the first unplayed duration of the unplayed first video, and make the second player start playing from the position of the first unplayed duration after rewinding from the end of the second video.
[0263] 514. If the second player has not finished playing the second video, while the first player has finished playing the first video, make the second player skip the second video that has not finished playing.
[0264] In step 513, either because the second duration of the second video is less than the first duration of the first video, or because the synchronization mechanism in steps 511-512 fails, a situation arises where the second player has finished playing the second video, but the first player has not yet finished playing the first video. In this case, as mentioned above, if the first player playing the first video enters background mode again, the second player cannot follow suit because the second video has finished playing. To maintain the proper operation of this embodiment, the first unplayed duration of the unfinished first video can be obtained, and the second player can start playing from the end of the second video, rewinding to the position of the first unplayed duration. This way, the first player playing the first video enters background mode again, and the second player can also follow suit because there is still content to be played.
[0265] The first unplayed duration of the first video can be determined based on its first duration and the already played duration. The first duration of the first video is known. The first player has a played duration indicator to show the already played duration of the currently playing video. By reading the played duration indicator of the first player, the already played duration can be obtained. Subtracting the already played duration from the first duration gives the first unplayed duration.
[0266] The first unplayed duration of the first unplayed video can also be obtained by sending a first unplayed duration polling request to the first player. After receiving the first unplayed duration polling request, the first player sends a first unplayed duration response to the second player, which contains the first unplayed duration. This is similar to the first control state polling signal and the first control state response signal mentioned above, so it will not be described in detail.
[0267] For example, the first video has 5 seconds left to play, but the second video has already finished. The second video is 20 seconds long, and the second player rewinds to the 15-second mark of the second video to continue playing it.
[0268] In step 514, if the synchronization mechanism in steps 511-512 detects that the second player has entered the playback end state, and the first player still has a portion of the second video that has not finished playing, then in order to synchronize with the second player, the first player will skip the unfinished portion of the second video and stop playing it. For example, if the second video has 5 seconds left to play, but the first video has already finished playing, and the second video is 20 seconds long, and the first player is currently playing at the 15th second, then the first player will jump directly to the 20th second position, that is, stop playing the first video.
[0269] The advantage of steps 513-514 is that when the synchronization mechanism in steps 511-512 fails, the playback of the first video and the second video can at least remain synchronized at the end, thus improving the robustness of the subtitle display method of this embodiment.
[0270] Detailed description of step 520
[0271] In step 520, in response to the first player entering background running mode, the second player is instructed to enter the background running mode and display the target picture-in-picture window.
[0272] In one embodiment, such as Figure 11 As shown, step 520 may include:
[0273] 1110. Use a second player to generate the target picture-in-picture controller;
[0274] 1120. In response to the first player entering background running mode, instruct the second player to enter background running mode;
[0275] 1130. Use the second player to call the target picture-in-picture controller to display the target picture-in-picture window.
[0276] The following is a detailed description of steps 1110-1130 above.
[0277] In step 1110, the picture-in-picture controller is a component set in the object layer of the second player. It can display the picture-in-picture window and configure the display effect of the picture-in-picture window.
[0278] In one embodiment, such as Figure 12 As shown, step 1110 may include:
[0279] 1210. Use the second player to generate an initial picture-in-picture controller. The initial picture-in-picture controller has display status parameters for the control component view. The initial value of the display status parameters indicates each control component in the display control component view.
[0280] 1220. Change the initial value of the display status parameter in the initial picture-in-picture controller to the target value. The target value indicates the control components in the hidden control component view, thereby turning the initial picture-in-picture controller into the target picture-in-picture controller.
[0281] In step 1210, the initial picture-in-picture controller refers to the picture-in-picture controller that comes pre-installed with the second player. Control components are components that control the display of content in the picture-in-picture window, such as rewind, fast forward, and progress bar buttons. The control component view is a view that carries these control components, such as a view containing rewind, fast forward, and progress bar buttons. Through this control component view, the user can clearly understand the playback status of the picture-in-picture window. This embodiment of the disclosure utilizes this control component view containing control components, so that the control components are no longer displayed, but the target subtitles required by this embodiment are overlaid on them, achieving the effect of still displaying the subtitles of the third-party video when it runs in the background.
[0282] The display status parameter specifies the content displayed in the control component view. If the control component view needs to display the control component normally, it is set to the value that enables the control component view to display the control component normally, i.e., the initial value in step 1210. If the control component view does not display the control component, but instead overlays a target subtitle on it using this embodiment, it is set to the value that prevents the control component view from displaying the control component, i.e., the target value in step 1220.
[0283] Steps 1210-1220 modify the display status parameter value in the picture-in-picture controller to change the content displayed in the control component view, transforming the control component view into the target subtitle view and changing the initial picture-in-picture controller into the target picture-in-picture controller. Through this simple modification of the display status parameter, this embodiment achieves the effect of displaying subtitles of a third-party video even when the video is running in the background, thus improving the efficiency of subtitle display processing in this embodiment.
[0284] In the above embodiment, there is only one target picture-in-picture window (with a unique shape, size, font, etc.). In step 1130, the target picture-in-picture controller is invoked through the second player, and only this unique target picture-in-picture window can be displayed. In other embodiments, multiple candidate picture-in-picture windows can be selected. This selection can be made by the user from multiple candidate windows, or by the XR glasses based on several factors.
[0285] In one embodiment, such as Figure 13 As shown, step 1130 may include:
[0286] 1310. Through the second player, call the target picture-in-picture controller to obtain the global picture-in-picture window list, which includes multiple candidate picture-in-picture windows;
[0287] 1320. From multiple candidate picture-in-picture windows, obtain the target picture-in-picture window that matches the target subtitle;
[0288] 1330. Display the target picture-in-picture window.
[0289] In step 1310, the global picture-in-picture window list is a list of all available picture-in-picture windows. The candidate picture-in-picture window is each picture-in-picture window in the global list. Because there are multiple candidate picture-in-picture windows to choose from, the global list is retrieved instead of a single target picture-in-picture window.
[0290] In step 1320, in one embodiment, a global list of picture-in-picture windows can be displayed to the user, allowing the user to select the target picture-in-picture window that matches the target subtitle. For example... Figure 14 As shown, an example global picture-in-picture window list includes:
[0291] 1) 3cm × 1cm, boldface;
[0292] 2) 16cm × 2cm, boldface;
[0293] 3) 3cm × 1cm, in regular script;
[0294] 4) 16cm×2cm, in regular script.
[0295] Since the target subtitle is in KaiTi font and has a large number of characters, in order to display it properly, the user selects the candidate picture-in-picture window of "16cm×2cm, KaiTi font" as the target picture-in-picture window.
[0296] In another embodiment, a target picture-in-picture window matching the target subtitle is obtained from multiple candidate picture-in-picture windows according to certain criteria.
[0297] In one embodiment, a suitable picture-in-picture window can be selected based on the number of characters in the target subtitle. The more characters in the target subtitle, the larger the picture-in-picture window needs to be.
[0298] In one embodiment, such as Figure 15 As shown, step 1320 may specifically include:
[0299] 1510. Obtain the number of characters in the target subtitle;
[0300] 1520. Determine the size of the second window of the target picture-in-picture window based on the number of characters;
[0301] 1530. Based on the second window size, determine the target picture-in-picture window from multiple candidate picture-in-picture windows.
[0302] In step 1510, as mentioned earlier, the target subtitles are divided into hard subtitles and soft subtitles. In the case of hard subtitles, the compressed file of the first video can be decompressed, the subtitles extracted, and the number of characters in the subtitles counted. In the case of soft subtitles, OCR tools can be used to parse the first video, extract the subtitles, and count the number of characters in the subtitles.
[0303] In step 1520, the size of the second window of the target picture-in-picture window is determined based on the number of characters. This can be done by using a lookup table or a formula.
[0304] Table 2 below shows the correspondence between the number of characters and the size of the second window.
[0305] word count Second window size 0-5 3cm×1cm 6-10 4cm×2cm 11-15 6cm×2cm 16-20 8cm×3cm 21 or more 12cm×3cm
[0306] Table 2
[0307] For example, if the target subtitle has 14 characters, the corresponding second window size is 6cm × 2cm according to Table 2.
[0308] Note that in the example above, the size of the second window is expressed using its width and height. In another embodiment, the size of the second window can also be expressed using its area.
[0309] Formula 2 below shows the relationship between the size of the second window and the number of characters:
[0310] S = B * h Formula 2
[0311] In Formula 2 above, S represents the size of the second window, h represents the number of characters, and B is a normal number.
[0312] For example, if B = 0.6 and h = 5, then S = 0.6 * 5 = 3 (cm) 2 The second window is 3cm in size. 2Note that the size of the second window is expressed as its area.
[0313] In step 1530, based on the second window size determined in step 1520, the best matching picture-in-picture window is selected from multiple candidate picture-in-picture windows.
[0314] In one embodiment, if one of the candidate picture-in-picture windows has the exact same size as the second window, then that candidate picture-in-picture window is the target picture-in-picture window. For example, if the second window size is 12cm × 3cm, and a candidate picture-in-picture window is also 12cm × 3cm, then that candidate picture-in-picture window is used as the target picture-in-picture window. Another example is that the second window size is 36cm. 2 If the size of a candidate picture-in-picture window is also 36cm 2 If so, then the candidate picture-in-picture window will be used as the target picture-in-picture window.
[0315] If none of the candidate picture-in-picture windows is exactly the same size as the second window, then different target picture-in-picture window selection strategies can be adopted according to the representation of the second window size.
[0316] In one embodiment, if the size of the second window is expressed in terms of the area of the second window, then step 1530 may include:
[0317] Get the area of each candidate picture-in-picture window;
[0318] Calculate the difference between the area of each candidate picture-in-picture window and the area of the second window;
[0319] Select the target picture-in-picture window from multiple candidate picture-in-picture windows based on the absolute value of the area difference.
[0320] When selecting a target picture-in-picture window from multiple candidate picture-in-picture windows based on the absolute value of the area difference, one can choose the candidate window with the smallest absolute value as the target window, or randomly select any candidate window whose absolute value is less than a predetermined absolute value threshold. The former has the advantage of making full use of the target picture-in-picture window's space and minimizing space waste. The latter has the advantage of avoiding selecting only one or a few candidate picture-in-picture windows, thus preventing the target picture-in-picture window from being too singular.
[0321] In one embodiment, if the size of the second window is represented by a first width and a first height of the second window, then step 1530 may include:
[0322] Get the second width and second height of each candidate picture-in-picture window;
[0323] Calculate the area of the candidate picture-in-picture window based on the second width and the second height;
[0324] Calculate the area of the second window based on the first width and the first height;
[0325] Calculate the difference between the area of each candidate picture-in-picture window and the area of the second window;
[0326] Select the target picture-in-picture window from multiple candidate picture-in-picture windows based on the absolute value of the area difference.
[0327] The last two steps of this embodiment are the same as those of the previous embodiment, and therefore will not be described in detail. The difference between the first three steps of this embodiment and the first step of the previous embodiment is that the latter includes the process of calculating the area of the second window based on the first width and first height of the second window, and the process of calculating the area of the candidate picture-in-picture window based on the second width and second height of the candidate picture-in-picture window. The principle is largely the same as that of the previous embodiment, and therefore will not be described in detail.
[0328] The advantage of this embodiment is that it converts the width and height of the window into areas for comparison between windows. Since the selection of the window size mainly depends on whether the area of the window can accommodate all the characters of the target subtitle, the accuracy of determining the target picture-in-picture window size in this embodiment of the disclosure is improved.
[0329] Besides selecting a suitable picture-in-picture window based on the number of characters in the target subtitles, the size of the picture-in-picture window can also be determined by considering both the size of other application windows and the number of characters in the target subtitles. The purpose of using a picture-in-picture window when playing the first video is to simultaneously operate the target application, such as reading an e-book. Operating the target application involves a target application window, which is displayed simultaneously with the target picture-in-picture window on the virtual screen seen by the user. If the target application window is large, and the virtual screen area is fixed, then the target picture-in-picture window should be small. If the target application window is small, then the target picture-in-picture window can be large.
[0330] In one embodiment, such as Figure 16 As shown, step 1320 may specifically include:
[0331] 1610. Get the size of the first window of the target application window that is to be displayed in parallel with the picture-in-picture window;
[0332] 1620. Obtain the number of characters in the target subtitle;
[0333] 1630. Based on the size of the first window and the number of characters, determine the size of the second window of the target picture-in-picture window;
[0334] 1640. Based on the second window size, determine the target picture-in-picture window from among multiple candidate picture-in-picture windows.
[0335] In step 1610, since the first player enters background operation mode, it is usually because the target application window is opened. Therefore, the first window size of the opened target application window can be obtained. This first window size can be represented by the width and height of the first window, or by the area of the first window.
[0336] Step 1620 is the same as step 1510, so it will not be described again.
[0337] Step 1630 is similar to step 1520; it can also be done using a table lookup method or a formula method. When using a formula method, the following formula can be used:
[0338] S = B*h + C*(QT) Formula 3
[0339] In Formula 3 above, S represents the size of the second window, h represents the number of characters, B is a positive integer, Q represents the area of the entire virtual screen, T represents the size of the first window, and C is a positive integer. Formula 3 shows that the size of the second window is an increasing function of the number of characters and a decreasing function of the size of the first window.
[0340] Step 1640 is the same as step 1530, so it will not be repeated.
[0341] The above method for determining the target picture-in-picture window takes into account both the size of the first window and the number of characters in the target subtitle, thus improving the accuracy of the target picture-in-picture window determination.
[0342] The above illustrates several embodiments for determining a target picture-in-picture window from multiple candidate picture-in-picture windows based on a second window size. Window size is only one attribute of the target picture-in-picture window. In another embodiment, in addition to determining the second window size of the target picture-in-picture window, a second display style of the target picture-in-picture window is also determined, and the target picture-in-picture window is determined from multiple candidate picture-in-picture windows based on the second window size and the second display style.
[0343] like Figure 17 As shown, in this embodiment, step 1320 may include:
[0344] 1710. Get the size of the first window of the target application window that is to be displayed in parallel with the picture-in-picture window;
[0345] 1720. Obtain the number of characters in the target subtitle;
[0346] 1730. Based on the size of the first window and the number of characters, determine the size of the second window of the target picture-in-picture window;
[0347] 1740. Based on the first display style of the target application window, determine the second display style of the target picture-in-picture window;
[0348] 1750. Based on the second window size and the second display style, determine the target picture-in-picture window from multiple candidate picture-in-picture windows.
[0349] Steps 1710-1730 are similar to steps 1610-1630, and will not be repeated here.
[0350] In step 1740, display style refers to the primary color scheme or window style of a window. The first display style is the primary color scheme or window style of the target application window. The second display style is the primary color scheme or window style of the target picture-in-picture window. For example, if the first display style of the target application window is green, then the second display style of the target picture-in-picture window is also green.
[0351] Since the first display style and the second display style are displayed on the same virtual screen, their consistent display styles help improve the visual experience of the virtual screen, thereby increasing user engagement.
[0352] In step 1750, if the size of the second window is exactly the same as the window size of a single candidate picture-in-picture window, and the second display style is also exactly the same as the display style of the candidate picture-in-picture window, then the candidate picture-in-picture window is determined as the target picture-in-picture window.
[0353] If the size of the second window is not exactly the same as the window size of each candidate picture-in-picture window, and the second display style is not exactly the same as the display style of each candidate picture-in-picture window, in one embodiment, a first degree of matching between the size of the second window and the window size of a single candidate picture-in-picture window can be determined, and a second degree of matching between the second display style and the display style of the candidate picture-in-picture window can be determined. Then, the target picture-in-picture window is determined among the multiple candidate picture-in-picture windows based on the first degree of matching and the second degree of matching.
[0354] In one embodiment, determining the first matching degree between the size of the second window and the window size of a single candidate picture-in-picture window can be achieved by determining the absolute value of the difference between the two window sizes, dividing that absolute value by the second window size to obtain the difference ratio, and then subtracting that difference ratio from 1 to obtain the first matching degree. For example, the second window size is 20cm. 2 The window size of a single candidate picture-in-picture window is 22cm. 2 The absolute value of the difference between the two window sizes is 2cm. 2 The difference ratio obtained by dividing the absolute value by the second window size is 10%, and the first match degree is 90%.
[0355] In one embodiment, determining the second matching degree between the second display style and the display style of the candidate picture-in-picture window can be done by looking up a table. Since the lookup table method has been mentioned several times previously, it will not be elaborated upon further.
[0356] In one embodiment, when determining the target picture-in-picture window among multiple candidate picture-in-picture windows based on a first matching degree and a second matching degree, the sum of the first matching degree and the second matching degree can be calculated, and the target picture-in-picture window can be determined among the multiple candidate picture-in-picture windows based on this sum. In another embodiment, the average of the first matching degree and the second matching degree can be calculated, and the target picture-in-picture window can be determined among the multiple candidate picture-in-picture windows based on this average. In yet another embodiment, a weighted average of the first matching degree and the second matching degree can be calculated, and the target picture-in-picture window can be determined among the multiple candidate picture-in-picture windows based on this weighted average.
[0357] The advantage of steps 1710-1750 is that, based on the various attributes (size, display style) of the target picture-in-picture window, the target picture-in-picture window is determined from multiple candidate picture-in-picture windows, thus improving the accuracy of determining the target picture-in-picture window.
[0358] In one embodiment, such as Figure 18 As shown, step 520 may also include:
[0359] 1810. In response to the triggering of the background running control in the first player, instruct the first player to enter the background running mode;
[0360] 1820. Instruct the second player to enter background operation mode;
[0361] 1830. Hide the first player and show the target picture-in-picture window.
[0362] In step 1810, the background running control is a control within the first player; triggering it causes the first player to enter background running mode. It can be... Figure 2A , Figure 3A and Figure 4A The control in the upper right corner of the virtual screen shown could also be an item in a drop-down menu that appears on the virtual screen upon some trigger, and so on. For example, if a user touches a blank area of the virtual screen, a drop-down menu pops up on the virtual screen, and this drop-down menu contains the background running control. The user triggers this background running control through touch, keyboard, eye contact, etc., and the first player enters background running mode.
[0363] In step 1820, the first player can send a background operation command to the second player. After receiving the background operation command, the second player enters the background operation mode.
[0364] In step 1830, since the first player is playing the first video of a third-party application, it is not allowed to be displayed in the form of a picture-in-picture window, so it needs to be hidden. However, the second video played by the second player is a local video, which is allowed to be displayed in the form of a picture-in-picture window, so the target picture-in-picture window is displayed.
[0365] Steps 1810-1830 provide a convenient triggering method to trigger the first player and the second player to enter the background running mode, and hide the first player after it enters the background running mode, eliminating its interference with the target picture-in-picture window and the target application window, and improving user operation efficiency.
[0366] Detailed description of step 530
[0367] In step 530, target subtitles are obtained from the first video, and a target subtitle view is generated based on the target subtitles. The target subtitle view is then overlaid on the target picture-in-picture window to cover the second video.
[0368] Steps 1210-1220 above describe modifying the display status parameters in the target picture-in-picture controller. In one embodiment, based on steps 1210-1220, such as Figure 19 As shown, step 530 may include:
[0369] 1910. Using the target picture-in-picture controller, generate a view of the control components with each control component hidden.
[0370] 1920. On the control component view where all control components are hidden, add the target caption to obtain the target caption view.
[0371] The detailed description of steps 1210-1220 above has clarified the method for modifying the display status parameters in the picture-in-picture controller, as well as the relationship between the control component view, the control component, and the target caption. Steps 1910-1920 are specific applications of the above control method, and therefore will not be elaborated upon. It provides an efficient way to obtain the target caption view.
[0372] As mentioned earlier, when the target caption view is overlaid on the target picture-in-picture window, it can partially or completely obscure the second video. When completely obscuring the second video, the aspect ratio of the target caption view must be the same as that of the target picture-in-picture window. In one embodiment, such as Figure 20 As shown, step 530 may also include:
[0373] 2010. Scale the target caption view so that the scaled target caption view is the same size as the target picture-in-picture window, and then overlay the scaled target caption view onto the target picture-in-picture window to completely cover the second video.
[0374] Since the aspect ratio of the target caption view is the same as that of the target picture-in-picture window, scaling the target caption view will make it the same size as the target picture-in-picture window. In this way, when the scaled target caption view covers the target picture-in-picture window, it will completely obscure the second video, avoiding any blending issues caused by inconsistencies between the background of the target caption view and the content of the target picture-in-picture window.
[0375] In another embodiment, the target caption view is smaller than the target picture-in-picture window. In this case, the background color of the target caption view is required to match the background color of the second video. Because the target caption view is smaller than the target picture-in-picture window, it does not completely obscure the window. If the color of the portion of the second video not obscured by the target caption view differs from its background color, it creates a visual misalignment and disrupts the user experience. By matching the background color of the target caption view with the background color of the second video, the harmony of the displayed image is improved.
[0376] The above embodiments describe the scenario where only the target caption view is overlaid on the target picture-in-picture window. In another embodiment, the playback content of the first video can also be overlaid on the target picture-in-picture window. In this way, the playback content in the target picture-in-picture window will completely simulate the playback content of the first video, giving the impression that the first video of a third-party application can also be played in picture-in-picture format.
[0377] In this embodiment, such as Figure 21 As shown, step 530 may include:
[0378] 2110. Obtain video frames from the first video and generate a video frame view based on the video frames;
[0379] 2120. Overlay the video frame view onto the target picture-in-picture window to cover the second video;
[0380] 2130. Obtain the target subtitle from the first video and generate a target subtitle view based on the target subtitle;
[0381] 2140. Overlay the target caption view onto the target picture-in-picture window to partially obscure the video frame view.
[0382] In step 2110, a video frame is a frame of the first video at various playback moments. A video frame view is a collection of these video frames arranged in the order they were extracted.
[0383] In one embodiment, video frames are extracted from the first video at fixed periodic intervals. For example, a video frame is extracted from the first video every 1 second. The normal frame rate for PAL images is 25 frames per second, and for NTSC images it is 30 frames per second. The frame rate for extracting video frames needs to be much lower than the image frame rate. If the frame rate for extracting video frames equals the image frame rate, it is equivalent to downloading the first video. Downloading is relatively slow and would affect the real-time performance of covering the target picture-in-picture window with the target caption view. Since only one video frame is extracted from the first video every few frames, the video frame view is actually discontinuous, but due to continuous playback, the persistence of vision makes the user see what appears to be a complete first video.
[0384] In step 2120, after the video frame view is overlaid on the target picture-in-picture window, the second video in the target picture-in-picture window is obscured by the video frame view. The target picture-in-picture window that the user sees will be a window playing video frames. This video frame still only contains video images and no subtitles. In order to obtain subtitles, a target subtitle view is obtained through steps 2130-2140, and the target subtitle view is used to partially cover the video frame view. In this way, what the user sees is a target picture-in-picture window that contains both the video frames extracted from the first video and subtitles, improving the richness of the displayed content of the first video in the third-party application.
[0385] The preceding embodiments described the scenario where the target picture-in-picture window was displayed when the first player, playing the first video from a third-party application, entered background mode. The following embodiments will describe the scenario when the first player exits background mode. Exiting background mode means the first player enters foreground mode, at which point the first video from the third-party application can play normally in the foreground. Therefore, the target picture-in-picture window can be closed, and the first player interface can be displayed normally.
[0386] In one embodiment, such as Figure 22 As shown, after step 530, the subtitle display method includes:
[0387] 540. In response to a triggering of a restore foreground running control associated with the target caption view, close the target picture-in-picture window and remove the target caption view;
[0388] 550. Display the first player and play the first video in the first player.
[0389] In step 540, the "Restore Foreground Running Control" is a control that brings the first player back into the foreground. In one embodiment, it can be the icon of the first player on the desktop. Triggering this icon re-enters the first player interface. In another embodiment, it can be located in a drop-down menu that pops up when triggered on a virtual screen. For example, triggering it in a blank area of the virtual screen displays a drop-down menu, and the "Restore Foreground Running Control" is an item in the drop-down menu.
[0390] When the control to restore foreground operation is triggered, since the first player is restored to foreground operation and does not need to be displayed in picture-in-picture mode, the target picture-in-picture window is closed, the target caption view is cleared, and in step 550, the first player is displayed in the foreground. The first video plays in the first player.
[0391] Steps 540-550 provide a way to exit from the target picture-in-picture window, allowing users to switch between normal display and picture-in-picture display multiple times, improving user operation flexibility.
[0392] In one embodiment, when the first player resumes foreground operation, it does not play the first interface that was played when the first player entered background operation mode, but instead returns to the second interface before the first interface, so as to provide the user with more choices.
[0393] like Figure 23 As shown, in this embodiment, step 550 may specifically include:
[0394] 2310. Obtain the first screen of the first player when it enters background running mode and is playing the first video;
[0395] 2320. Obtain the second interface of the first player before the first interface;
[0396] 2330. Display the first player, and then display the second interface within the first player;
[0397] 2340. In response to the selection of the first video in the second interface, play the first video.
[0398] The following is combined with Figures 24A-24F Exemplary description Figure 23 The process.
[0399] exist Figure 24A When a user opens XX Player, a list of multiple videos from third-party applications is first displayed. The user selects "City Forest" from the list, and then... Figure 24BThe interface shown displays video frames from the "urban forest," featuring a house and a tree, with the narration "My home is in the countryside." To exit the foreground and enter background mode, the user clicks the "-" control in the upper right corner of the XX player. Figure 24C A standard virtual screen interface. Figure 24C Click the "eBook" icon on the normal virtual screen interface. The "eBook" app enters foreground mode, and the XX Player enters background mode. At this point, the first screen playing when entering background mode, as obtained in step 2310, is... Figure 24B The interface. The second interface, obtained before the first interface in step 2320, is... Figure 24A The interface.
[0400] When the "e-book" enters foreground mode, such as Figure 24D As shown, the target caption view, with the target subtitle "My home is in the countryside," overlays the target picture-in-picture window and displays in parallel with the "eBook" page. When the user wants to exit the XX player's background running mode, the user clicks... Figure 24D The control with a "-" in the upper right corner of the target picture-in-picture window displays... Figure 24E The interface. The user clicks... Figure 24E The icon for "XX Player" appears. "XX Player" enters foreground mode, and in step 2330, it displays as shown below. Figure 24F The second interface. This second interface is the interface preceding the first interface, and is related to... Figure 24A Same. In step 2340, the user selects "City Forest" and re-views "City Forest".
[0401] Detailed description of the complete embodiments of this disclosure
[0402] The following is combined with Figure 25 A complete embodiment of the subtitle display method disclosed herein will be described in detail.
[0403] In step 2501, the first player downloads the first video from the third-party application, ready for playback.
[0404] In step 2502, the first player starts playing the first video.
[0405] In step 2503, synchronously with the first player starting to play the first video, the second player prepares and starts playing the second video.
[0406] In step 2504, the second player generates a target picture-in-picture controller. Specifically, it changes the initial value of the display status parameter in the initial picture-in-picture controller to a target value, which indicates the control components in the hidden control component view.
[0407] In step 2505, the first player instructs the target picture-in-picture controller to display the target picture-in-picture window.
[0408] In step 2506, the target picture-in-picture controller displays the target picture-in-picture window, and plays the second video in the target picture-in-picture window.
[0409] In step 2507, the first player obtains the target subtitle from the first video, generates a target subtitle view based on the target subtitle, and overlays the target subtitle view onto the target picture-in-picture window, thus covering the second video.
[0410] In step 2508, the target picture-in-picture window covered by the target caption view is displayed.
[0411] In step 2509, the first player plays the first video, and in step 2510, the second player plays the second video. Throughout this process, the first and second players remain synchronized, ensuring that their playback progress is consistent.
[0412] In steps 2511 and 2512, when the first player pauses playing the first video, the second player simultaneously pauses playing the second video; when the second player pauses playing the second video, the first player simultaneously pauses playing the first video.
[0413] In steps 2513 and 2514, when the first player plays the first video at double speed, the second player plays the second video at the same double speed; when the second player plays the second video at double speed, the first player plays the first video at the same double speed.
[0414] In steps 2515 and 2516, when the first player stops playing the first video, the second player simultaneously stops playing the second video; when the second player stops playing the second video, the first player simultaneously stops playing the first video.
[0415] In step 2517, the target picture-in-picture controller closes the target picture-in-picture window and removes the target caption view.
[0416] In step 2518, the first player displays a second interface, which is the interface before the first interface when the first player is playing the first video when it enters the background running mode.
[0417] Description of apparatus and devices according to embodiments of this disclosure
[0418] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0419] It should be noted that in the various specific embodiments of this disclosure, when processing data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, is required, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when this application embodiment needs to obtain target object attribute information, separate permission or consent from the target object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of this application embodiment be obtained. For example, target object-related data may be data related to the user's eye movements or the user's gesture operation habits.
[0420] Reference Figure 26 , Figure 26 This is a schematic diagram of the structure of a subtitle display device 2600 provided in an embodiment of the present disclosure. The subtitle display device 2600 includes:
[0421] The startup unit 2610 is used to enable a local second player to play a second video when the first video of a third-party application starts playing through the first player of the third-party application, wherein the second player cannot play the first video of the third-party application.
[0422] The instruction unit 2620 is used to instruct the second player to enter the background running mode in response to the first player entering the background running mode, so as to display the target picture-in-picture window;
[0423] The window overlay unit 2630 is used to obtain target subtitles from the first video, generate a target subtitle view based on the target subtitles, and overlay the target subtitle view onto the target picture-in-picture window to cover the second video.
[0424] Optionally, the indicating unit 2620 is specifically used for:
[0425] Use a second player to generate the target picture-in-picture controller;
[0426] In response to the first player entering background running mode, instruct the second player to enter background running mode;
[0427] The second player invokes the target picture-in-picture controller to display the target picture-in-picture window.
[0428] Optionally, the indicating unit 2620 is specifically used for:
[0429] The initial picture-in-picture controller is generated using the second player. The initial picture-in-picture controller has display status parameters for the control component view. The initial value of the display status parameters indicates each control component in the display control component view.
[0430] Change the initial value of the displayed status parameter in the initial picture-in-picture controller to the target value. The target value indicates the control components in the hidden control component view, thereby turning the initial picture-in-picture controller into the target picture-in-picture controller.
[0431] Optionally, the window overlay unit 2630 is specifically used for:
[0432] Use the target picture-in-picture controller to generate a view of control components that hides each control component;
[0433] Add the target caption to the control component view where all control components are hidden, and you will get the target caption view.
[0434] Optionally, the indicating unit 2620 is specifically used for:
[0435] The second player calls the target picture-in-picture controller to obtain the global picture-in-picture window list, which includes multiple candidate picture-in-picture windows.
[0436] From multiple candidate picture-in-picture windows, select the target picture-in-picture window that matches the target subtitle;
[0437] Display the target picture-in-picture window.
[0438] Optionally, the indicating unit 2620 is specifically used for:
[0439] Get the size of the first window of the target application window that is to be displayed in parallel with the picture-in-picture window;
[0440] Get the number of characters in the target subtitle;
[0441] Based on the size of the first window and the number of characters, determine the size of the second window of the target picture-in-picture window;
[0442] Based on the first display style of the target application window, determine the second display style of the target picture-in-picture window;
[0443] The target picture-in-picture window is determined from multiple candidate picture-in-picture windows based on the second window size and the second display style.
[0444] Optionally, the aspect ratio of the target caption view is consistent with the aspect ratio of the target picture-in-picture window, and the window overlay unit 2630 is specifically used for:
[0445] The target caption view is scaled up so that it matches the size of the target picture-in-picture window, and then the scaled target caption view is overlaid on the target picture-in-picture window to completely cover the second video.
[0446] Optionally, the target caption view is smaller than the target picture-in-picture window, and the background color of the target caption view is the same as the background color of the second video.
[0447] Optionally, the window overlay unit 2630 is specifically used for:
[0448] Obtain video frames from the first video and generate a video frame view based on the video frames;
[0449] Overlay the video frame view onto the target picture-in-picture window to cover the second video;
[0450] Obtain the target subtitle from the first video, and generate a target subtitle view based on the target subtitle;
[0451] Overlay the target caption view onto the target picture-in-picture window to partially obscure the video frame view.
[0452] Optionally, the starting unit 2610 is specifically used for:
[0453] Retrieve the local video library, which includes multiple local candidate videos;
[0454] Get the primary color tone and duration of the first video;
[0455] Obtain the second primary color tone and second duration of the local candidate videos;
[0456] Based on the first primary color tone, the second primary color tone, the first duration, and the second duration, the second video is determined from multiple local candidate videos;
[0457] Enable the second player to play the second video.
[0458] Optionally, the starting unit 2610 is specifically used for:
[0459] Based on the first primary color and the second primary color, determine the first matching degree between the local candidate video and the first video;
[0460] Based on the first duration and the second duration, determine the second matching degree between the local candidate video and the first video;
[0461] Based on the first and second matching scores, the total matching score between the local candidate video and the first video is determined.
[0462] Based on the overall matching degree, the second video is determined from multiple local candidate videos.
[0463] Optionally, the starting unit 2610 is specifically used for:
[0464] Obtain multiple first keyframes from the first video;
[0465] Determine the first background color for each first keyframe;
[0466] Determine the first primary color tone based on the first background color of each first keyframe;
[0467] Get the first duration of the first video.
[0468] Optionally, the starting unit 2610 is specifically used for:
[0469] Obtain multiple second keyframes from the local candidate video;
[0470] Determine the second background color for each second keyframe;
[0471] Determine the second primary color tone based on the second background color of each second keyframe;
[0472] Get the second duration of the local candidate video.
[0473] Optionally, a second video corresponding to the first video is stored in a third-party application, and the startup unit 2610 is specifically used for:
[0474] Download a second video from a third-party application to your local device;
[0475] Enable the second player to play the second video.
[0476] Optionally, the subtitle display device further includes:
[0477] A detection unit (not shown) is used to detect the first control state of the first player;
[0478] A synchronization unit (not shown) is used to keep the second control state of the second player synchronized with the first control state.
[0479] Optionally, the subtitle display device further includes:
[0480] The rewind unit (not shown) is used to obtain the first unplayed duration of the first video if the second player has finished playing the second video but the first player has not finished playing the first video, and to make the second player start playing from the position of the first unplayed duration after rewinding from the end of the second video.
[0481] A skip unit (not shown) is used to cause the second player to skip the unfinished second video if the second player has not finished playing the second video while the first player has finished playing the first video.
[0482] Optionally, the indicating unit 2620 is specifically used for:
[0483] In response to the triggering of the background running control in the first player, instruct the first player to enter background running mode;
[0484] Instruct the second player to enter background operation mode;
[0485] Hide the first player and show the target picture-in-picture window.
[0486] Optionally, the subtitle display device further includes:
[0487] The closing unit (not shown) is used to close the target picture-in-picture window and eliminate the target caption view in response to a triggering of a restore foreground running control associated with the target caption view;
[0488] The display playback unit (not shown) is used to display the first player and play the first video in the first player.
[0489] Optionally, the display playback unit is specifically used for:
[0490] The first screen where the first video is playing when the first player enters background running mode;
[0491] Get the second interface of the first player before the first interface;
[0492] Display the first player, and then display the second interface within the first player;
[0493] In response to the selection of the first video on the second interface, the first video is played.
[0494] Reference Figure 27 , Figure 27To illustrate the structural block diagram of the XR glasses 110 used to implement the subtitle display method of this embodiment, the XR glasses 110 includes: a radio frequency (RF) circuit 2710, a memory 2715, an input unit 2730, a display unit 2740, a sensor 2750, an audio circuit 2760, a wireless fidelity (WiFi) module 2770, a processor 2780, and a power supply 2790, among other components. Those skilled in the art will understand that... Figure 27 The structure of the XR glasses 110 shown does not constitute a limitation on the XR glasses 110, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0495] The RF circuit 2710 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 2780; in addition, it transmits uplink data to the base station.
[0496] The memory 2715 can be used to store software programs and modules, and the processor 2780 executes various functional applications and data processing of the XR glasses by running the software programs and modules stored in the memory 2715.
[0497] The input unit 2730 can be used to receive input digital or character information, and to generate key signal inputs related to the settings and function control of the XR glasses. Specifically, the input unit 2730 may include a touch panel 2731 and other input devices 2732.
[0498] Display unit 2740 can be used to display input or provided information, as well as various menus of the terminal. Display unit 2740 may include display panel 2741.
[0499] Audio circuitry 2760, speaker 2761, and microphone 2762 provide an audio interface.
[0500] In this embodiment, the processor 2780 included in the XR glasses 110 can execute the subtitle display method of the previous embodiment.
[0501] This disclosure also provides a computer-readable storage medium for storing program code for executing the subtitle display methods of the foregoing embodiments.
[0502] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the aforementioned subtitle display method.
[0503] This disclosure also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described subtitle display method. This electronic device can be a smart terminal including XR glasses.
[0504] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0505] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0506] Finally, it should be noted that the above embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A method for displaying subtitles, characterized in that, The method includes: When the first video of the third-party application is started playing through the first player of the third-party application, the local second player is enabled to play the second video, wherein the second player cannot play the first video of the third-party application; In response to the first player entering background running mode, the second player is instructed to enter the background running mode to display the target picture-in-picture window; Obtain the target subtitle from the first video, generate a target subtitle view based on the target subtitle, and overlay the target subtitle view onto the target picture-in-picture window to cover the second video.
2. The subtitle display method according to claim 1, characterized in that, The step of instructing the second player to enter the background running mode in response to the first player entering the background running mode, so as to display the target picture-in-picture window, includes: The target picture-in-picture controller is generated using the second player; In response to the first player entering the background running mode, the second player is instructed to enter the background running mode; The second player invokes the target picture-in-picture controller to display the target picture-in-picture window.
3. The subtitle display method according to claim 2, characterized in that, The step of generating a target picture-in-picture controller using the second player includes: The second player is used to generate an initial picture-in-picture controller, which has display status parameters for a control component view. The initial value of the display status parameters indicates the display of each control component in the control component view. The initial value of the display state parameter in the initial picture-in-picture controller is changed to a target value, the target value indicating that each control component of the control component view is hidden, thereby turning the initial picture-in-picture controller into the target picture-in-picture controller.
4. The subtitle display method according to claim 3, characterized in that, The process of generating a target subtitle view based on the target subtitle includes: Using the target picture-in-picture controller, generate a view of the control components with each control component hidden; Add the target subtitle to the control component view where all control components are hidden to obtain the target subtitle view.
5. The subtitle display method according to claim 2, characterized in that, The step of displaying the target picture-in-picture window by calling the target picture-in-picture controller through the second player includes: The second player invokes the target picture-in-picture controller to obtain a global picture-in-picture window list, which includes multiple candidate picture-in-picture windows. From the plurality of candidate picture-in-picture windows, obtain the target picture-in-picture window that matches the target subtitle; Display the target picture-in-picture window.
6. The subtitle display method according to claim 5, characterized in that, The step of obtaining the target picture-in-picture window that matches the target subtitle from the plurality of candidate picture-in-picture windows includes: Obtain the first window size of the target application window to be displayed in parallel with the picture-in-picture window; Obtain the number of characters in the target subtitle; Based on the first window size and the number of characters, determine the second window size of the target picture-in-picture window; Based on the first display style of the target application window, determine the second display style of the target picture-in-picture window; Based on the second window size and the second display style, the target picture-in-picture window is determined from the plurality of candidate picture-in-picture windows.
7. The subtitle display method according to claim 1, characterized in that, The aspect ratio of the target caption view is the same as that of the target picture-in-picture window; The step of overlaying the target subtitle view onto the target picture-in-picture window to cover the second video includes: The target subtitle view is scaled up so that it matches the size of the target picture-in-picture window, and then the scaled target subtitle view is overlaid on the target picture-in-picture window to completely cover the second video.
8. The subtitle display method according to claim 1, characterized in that, The step of obtaining target subtitles from the first video, generating a target subtitle view based on the target subtitles, and overlaying the target subtitle view onto the target picture-in-picture window to cover the second video includes: Obtain video frames from the first video, and generate a video frame view based on the video frames; Overlay the video frame view onto the target picture-in-picture window to cover the second video; Obtain the target subtitle from the first video, and generate a target subtitle view based on the target subtitle; The target caption view is overlaid on the target picture-in-picture window to partially obscure the video frame view.
9. The subtitle display method according to claim 1, characterized in that, The method of enabling a local second player to play a second video includes: Obtain a local video library, which includes multiple local candidate videos; Obtain the first primary color tone and the first duration of the first video; Obtain the second primary color tone and second duration of the local candidate video; Based on the first primary color tone, the second primary color tone, the first duration, and the second duration, the second video is determined from the plurality of local candidate videos; Enable the second player to play the second video.
10. The subtitle display method according to claim 9, characterized in that, The step of determining the second video from the plurality of local candidate videos based on the first primary color tone, the second primary color tone, the first duration, and the second duration includes: Based on the first primary color and the second primary color, a first matching degree between the local candidate video and the first video is determined; Based on the first duration and the second duration, a second matching degree between the local candidate video and the first video is determined; Based on the first matching degree and the second matching degree, the total matching degree between the local candidate video and the first video is determined; Based on the total matching degree, the second video is determined from the plurality of local candidate videos.
11. The subtitle display method according to claim 1, characterized in that, The second video corresponding to the first video is stored in the third-party application; The method of enabling a local second player to play a local second video includes: Download the second video to your local device from the third-party application; Enable the second player to play the second video.
12. The subtitle display method according to claim 1, characterized in that, After enabling the local second player to play the second video, the method further includes: Detect the first control state of the first player; The second control state of the second player is synchronized with the first control state.
13. The subtitle display method according to claim 12, characterized in that, After synchronizing the second control state of the second player with the first control state, the method further includes: If the second player has finished playing the second video, but the first player has not finished playing the first video, obtain the first unplayed duration of the unplayed first video, and cause the second player to start playing from the position of the first unplayed duration after rewinding from the end of the second video; If the second player has not finished playing the second video, while the first player has finished playing the first video, the second player will skip the unfinished second video.
14. The subtitle display method according to claim 1, characterized in that, The step of instructing the second player to enter the background running mode in response to the first player entering the background running mode to display the target picture-in-picture window includes: In response to the triggering of the background running control in the first player, the first player is instructed to enter the background running mode; Instruct the second player to enter the background running mode; Hide the first player and display the target picture-in-picture window.
15. The subtitle display method according to claim 1, characterized in that, After obtaining target subtitles from the first video, generating a target subtitle view based on the target subtitles, and overlaying the target subtitle view onto the target picture-in-picture window to cover the second video, the method further includes: In response to the triggering of the restore foreground running control associated with the target caption view, the target picture-in-picture window is closed and the target caption view is eliminated; The first player is displayed, and the first video is played in the first player.
16. The subtitle display method according to claim 15, characterized in that, The step of displaying the first player and playing the first video in the first player includes: Obtain the first interface of the first player that is playing the first video when the first player enters the background running mode; Obtain the second interface of the first player before the first interface; Display the first player, and display the second interface within the first player; In response to the selection of the first video in the second interface, the first video is played.
17. A subtitle display device, characterized in that, The subtitle display device includes: A startup unit is configured to enable a local second player to play a second video when a first video of a third-party application is started playing through a first player of the third-party application, wherein the second player cannot play the first video of the third-party application; An instruction unit is configured to, in response to the first player entering a background running mode, instruct the second player to enter the background running mode in order to display the target picture-in-picture window; A window overlay unit is used to obtain target subtitles from the first video, generate a target subtitle view based on the target subtitles, and overlay the target subtitle view onto the target picture-in-picture window to cover the second video.
18. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the subtitle display method according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the subtitle display method according to any one of claims 1 to 16.
20. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the subtitle display method according to any one of claims 1 to 16.