Caption display method and related device
The caption display method enhances caption recognition by generating a mask based on the caption's color and surrounding area, improving visibility without altering the caption's color, thus addressing the issue of low caption recognition in video playback.
Patent Information
- Application Number
- JP2023580652
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-30
- Filing Date
- 2022-05-26
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-05-26
AI Technical Summary
In video playback scenarios, captions displayed simultaneously with videos often have low recognition due to matching color and brightness with the video, making them difficult for users to see clearly, especially in high-brightness or snowy conditions.
A caption display method where an electronic device generates a mask based on the color value of the caption or the surrounding area, overlaying the caption with the mask to enhance recognition without changing the caption's color. The method calculates the color value and transparency of the mask based on the caption recognition degree, improving visibility.
The method significantly improves caption recognition by adjusting the mask's color and transparency, ensuring better visibility of captions without obscuring the video content, thereby enhancing user experience.
Smart Images

Figure 0007687765000012 
Figure 0007687765000013 
Figure 0007687765000014
Abstract
Description
Technical Field
[0001] [Technical Field] This application relates to the field of terminal technologies, and in particular, to a caption display method and related devices.
Background Art
[0002] With the rapid development of electronic products, electronic devices such as mobile phones, tablet computers, and smart TVs have become ubiquitous in people's lives, and video playback has also become an important application function of these electronic devices. When an electronic device plays a video, an application scenario in which captions related to the played video are displayed in the video playback window Also common exists. For example, captions synchronized with audio are displayed in the video playback window, or captions input by the user are displayed in the video playback window to improve video interaction.
[0003] However, in the above-mentioned application scenario where captions are displayed simultaneously while the video is being played, when the color and brightness of the video are the same as the color Closer of the captions, or when the color and brightness of the captions highly match the color and brightness of the video at the caption display position. For example, in a high-brightness scenario where some bright-colored captions are displayed, or in a snow scenario where some white captions are displayed, the captions cannot be recognized sufficiently and are difficult for the user to see clearly Yes. Therefore, and the user experience deteriorates.
Summary of the Invention
[0004] An embodiment of this application provides a caption display method and related devices, which solve the problem of low caption recognition in the process of a user watching a video and improve the user experience.
[0005] According to a first aspect, an embodiment of the present application provides a caption display method. The method includes: An electronic device plays a first video. When the electronic device displays a first interface, the first interface includes a first picture and a first caption. The first caption is displayed in a floating manner on a first area of the first picture by using a first mask as a background. The first area is within the first picture and is an area corresponding to the display position of the first caption. The difference value between the color value of the first caption and the color value of the first area is a first value. When the electronic device displays a second interface, the second interface includes a second picture and the first caption. The mask is not displayed for the first caption. The first caption is displayed in a floating manner on a second area of the second picture. The second area is within the second picture and is an area corresponding to the display position of the first caption. The difference value between the color value of the first caption and the color value of the second area is a second value, and the second value is greater than the first value. The first picture is one picture within the first video, and the second picture is another picture within the first video.
[0006] In this embodiment of the present application, by implementing the foregoing caption display method, when the caption recognition degree of the electronic device is low, the electronic device can set a mask for the caption and increase the caption recognition degree without changing the color of the caption.
[0007] In a possible implementation, before the electronic device displays the first picture, the method further includes: The electronic device obtains a first video file and a first caption file, where the first video file and the first caption file carry the same time information. The electronic device generates a first video frame based on the first video file, where the first video frame is used to generate the first picture. The electronic device generates a first caption frame based on the first caption file, obtains the color value and display position of the first caption from the first caption frame, where the time information carried in the first caption frame is the same as the time information carried in the first video frame. The electronic device determines a first area based on the display position of the first caption. The electronic device generates a first mask based on the color value of the first caption or the color value of the first area. The electronic device overlays the first caption on the first mask in the first caption frame to generate a second caption frame, and combines the second caption frame and the first video frame. In this way, the electronic device can obtain the video file to be played and the caption file to be displayed, then decode the video file to obtain the video frame, and decode the caption file to obtain the caption frame. Then, the electronic device can extract caption color gamut information, caption position information, etc. from the caption frame, extract the color gamut information at the caption display position corresponding to the caption in the video frame based on the caption position information, and calculate the caption recognition degree based on the caption color gamut information and the color gamut information at the caption display position corresponding to the caption in the video frame. Further, the electronic device calculates the color value of the mask corresponding to the caption based on the caption recognition degree, generates a caption frame with a mask frame, then combines the video frame and the caption frame with the mask, and renders the combined video frame.
[0008] In a possible implementation form, before the electronic device generates the first mask based on the color value of the first caption or the color value of the first region, the method further includes: The electronic device determines that the first value is less than the first threshold. In this way, the electronic device can further determine that the caption recognition degree is low by determining that the first value is less than the first threshold.
[0009] In a possible implementation form, the electronic device's determination that the first value is less than the first threshold specifically includes: The electronic device divides the first region into N first sub-regions, where N is a positive integer. The electronic device determines that the first value is less than the first threshold based on the color value of the first caption and the color values of the N first sub-regions. In this way, the electronic device can determine that the first value is less than the first threshold based on the color value of the first caption and the color values of the N first sub-regions.
[0010] In a possible implementation form, the electronic device's generation of the first mask based on the color value of the first caption or the color value of the first region specifically includes: The electronic device determines the color value of the first mask based on the color value of the first caption or the color values of the N first sub-regions. The electronic device generates the first mask based on the color value of the first mask. In this way, the electronic device can determine the color value of the first mask based on the color value of the first caption or the color values of the N first sub-regions and further generate the first mask for the first caption.
[0011] In a possible implementation form, for the electronic device to determine that the first value is less than the first threshold specifically includes the following: The electronic device divides the first region into N first sub-regions, where N is a positive integer. The electronic device determines whether to combine adjacent first sub-regions into second sub-regions based on the difference value between the color values of adjacent first sub-regions. If the difference value between the color values of adjacent first sub-regions is less than the second threshold, the electronic device combines the adjacent first sub-regions into second sub-regions. The electronic device determines that the first value is less than the first threshold based on the color value of the first caption and the color value of the second sub-region. In this way, the electronic device combines first sub-regions with close color values to generate second sub-regions, and based on the color value of the first caption and the color value of the second sub-region, it can further determine that the first value is less than the first threshold.
[0012] In a possible implementation form, the first region includes M second sub-regions, M is a positive integer and is less than or equal to N. The second sub-region includes one or more first sub-regions, and the number of first sub-regions included in each second sub-region is the same as or different from the number of first sub-regions included in another second sub-region. In this way, the electronic device can divide the first region into M second sub-regions.
[0013] In a possible implementation form, for the electronic device to generate the first mask based on the color value of the first caption or the color value of the first region specifically includes the following: The electronic device sequentially calculates the color values of M first sub-masks based on the color value of the first caption or the color values of M second sub-regions. The electronic device generates M first sub-masks based on the color values of M first sub-masks, where the M first sub-masks are combined into the first sub-mask. In this way, the electronic device can generate M first sub-masks for the first caption.
[0014] In a possible implementation form, the method further includes: when the electronic device displays a third interface, the third interface includes a third picture and a first caption, the first caption includes at least a first part and a second part, a second sub-mask is displayed for the first part, and a third sub-mask is displayed for the second part or the third sub-mask is not displayed, and the color value of the second sub-mask is different from the color value of the third sub-mask. In this way, the electronic device can display captions corresponding to a plurality of sub-masks.
[0015] In a possible implementation form, the display position of the first mask is determined based on the display position of the first caption. In this way, the display position of the first mask can overlap with the display position of the first caption.
[0016] In a possible implementation form, the difference value between the color value of the first mask and the color value of the first caption is greater than the first value. In this way, the caption recognition degree can be improved.
[0017] In a possible implementation form, in the first picture and the second picture, the display position of the first caption with respect to the display screen of the electronic device is not fixed or is fixed, and the first caption is a segment of continuously displayed characters or symbols. In this way, the first caption can be a bullet comment (bullet barrage) , or a caption synchronized with the sound, and the first caption is not all the captions displayed on the display screen, but one caption.
[0018] In a possible implementation form, before the electronic device displays the first interface, the method includes: the electronic device sets the transparency of the first mask to less than 100%. In this way, it can be ensured that the video frame corresponding to the area where the first mask is located is still visible to a certain extent.
[0019] In a possible implementation, before the electronic device displays the second interface, the method includes: the electronic device generates a second mask based on the color value of the first caption or the color value of the second region, and superimposes the first caption on the second mask, where the color value of the second mask is a preset color value and the transparency of the second mask is 100%. Alternatively, the electronic device skips generating the second mask. In this way, for a caption with high recognition, the electronic device may set a mask with 100% transparency for the caption, or may not set a mask for the caption. Without either.
[0020] According to a second aspect, an embodiment of the present application provides an electronic device. The electronic device includes one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are configured to store computer program code. The computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device can execute the method in any possible implementation of the first aspect.
[0021] According to a third aspect, an embodiment of the present application provides a computer storage medium. The computer storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed on an electronic device, the electronic device can execute the method according to any possible implementation of the first aspect.
[0022] According to a fourth aspect, an embodiment of the present application provides a computer program product. When the computer program product is executed on a computer, the computer can execute the method according to any possible implementation of the first aspect.
Brief Description of the Drawings
[0023]
Figure 1A
Figure 1B
Figure 2A
Figure 2B
Figure 2C
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Figure 8C
Figure 9
Figure 10
Figure 11
Figure 12
MODE FOR CARRYING OUT THE INVENTION
[0024] Hereinafter, with reference to the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be Surely described. In the description of the embodiments of the present application, unless otherwise specified, " / " indicates an "or" relationship. For example, A / B may represent A or B. "And / or" in this specification is merely a related relationship for explaining related objects and represents that three relationships may exist. For example, A and / or B may represent the following three cases: only A exists, both A and B exist, and only B exists. In addition, in the description of the embodiments of the present application, "a plurality of" means "two or more".
[0025] In the specification, claims, and appended drawings of this application, terms such as "first", "second", etc. are intended to distinguish different objects and are not intended to indicate a specific order. Additionally, terms such as "include", "have", and any other variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the enumerated steps or units, and may optionally further include steps or units not enumerated, or may optionally further include other specific steps or units of the process, method, product, or device.
[0026] As used in this application, "embodiment" means that a particular characteristic, structure, or feature described with reference to an embodiment may be included in at least one embodiment of this application. Expressions shown in various places in this specification do not necessarily mean the same embodiment, nor are they exclusive, independent, or any particular embodiment from other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0027] For ease of understanding, some related concepts in the embodiments of this application will be described first below.
[0028] 1. Video decoding:
[0029] A process in which the binary data of a video file is read and interpreted according to the compression algorithm of the video file to obtain the data of image frames (sometimes also called video frames) for video playback.
[0030] 2. Caption:
[0031] Text and symbol information that is displayed in the video playback window during video playback and is independent of the video file.
[0032] 3. Video playback:
[0033] After operations such as video decoding and video rendering are performed on a video file, a process of displaying a group of images and corresponding audio information in chronological order in a video playback window.
[0034] 4. Bullet comments:
[0035] Captions that can be displayed in the video playback window of the input user or in the video playback window of another user on the video playback client (also called a video application), which are input by the user on the video playback client and are based on the position corresponding to the time input by the user for the image frame for video playback.
[0036] With the rapid development of electronic products, electronic devices such as mobile phones, tablet computers, and smart TVs have become ubiquitous in people's lives, and video playback has also become an important application function of these electronic devices. When an electronic device plays a video, an application scenario where captions related to the played video are displayed in the video playback window Also common exists. For example, captions synchronized with the audio are displayed in the video playback window, or captions input by the user (i.e., bullet comments) are displayed in the video playback window to improve video interaction.
[0037] In an application scenario where captions synchronized with the audio are displayed in the video playback window, usually at a lower position in the video playback window, matching is performed between the time stamp of the caption and the time stamp of the image frame played in the video, and the caption and the corresponding image frame played in the video are synthesized. Specifically, the caption is overlaid on the corresponding video frame, and the overlapping position Is solid is determined.
[0038] In an application scenario where captions (i.e., bullet comments) input by a user are displayed in a video playback window, during the video playback process, there are multiple captions flying from left to right or from right to left in the video playback window, and the overlapping position between the captions and the video frames Is solid is not determined.
[0039] In some actual application scenarios, in order to improve the enjoyment of video playback, the video playback platform usually provides the user with the ability to select the color of the captions. In an application scenario where captions synchronized with audio are displayed in a video playback window, the color of the captions is usually the system default color, and when playing the video, the user can select the preferred caption color. In this case, the electronic device displays the captions in the video playback window based on the color selected by the user. In an application scenario where bullet comments are displayed in the video playback window, the user who sends the bullet comments can select the color of the bullet comments to be sent, and the color of the bullet comments seen by another user matches the color of the bullet comments selected by the user who sends the bullet comments. Therefore, when a user views bullet comments, the colors of the bullet comments displayed within the same video frame may be different.
[0040] To implement the two foregoing application scenarios, one embodiment of the present application provides a caption display method. The electronic device may first obtain a video file to be played and a caption file to be displayed in a video playback window, and then may perform video decoding on the video file to obtain video frames and perform caption decoding on the caption file to obtain caption frames, respectively. Next, the electronic device may align and match the video frames and the caption frames based on time series to synthesize the final video frames to be displayed and store them in a video frame queue. Thereafter, the electronic device may read and render the video frames to be displayed based on time series, and finally, display the rendered video frames in the video playback window.
[0041] Hereinafter, the method steps of the foregoing caption display method will be described in detail.
[0042] FIGS. 1A and 1B show an example of the method steps of a caption display method according to an embodiment of the present application.
[0043] As shown in FIGS. 1A and 1B, the present method may be applied to an electronic device 100 having a video playback function. Hereinafter, specific steps of the method will be described in detail.
[0044] Phase 1: Acquisition of Video Information Stream and Caption Information Stream
[0045] S101 and S102: The electronic device 100 may detect an operation of playing a video on a video application by a user, and in response to this operation, the electronic device 100 may obtain a video information stream and a caption information stream.
[0046] Specifically, the video application can be installed on the electronic device 100. After detecting an operation by the user to play a video on the video application, in response to this operation, the electronic device 100 can obtain a video information stream (or what is called a video file) corresponding to the video that the user wants to play and a caption information stream (or what is called a caption file).
[0047] For example, FIG. 2A shows a user interface (UI) provided by the electronic device 100 to display an application installed on the electronic device 100. The electronic device 100 can detect an operation (e.g., a tap operation) performed by the user on the “Video” application option 211 within the user interface 210. In response to this operation, the electronic device 100 can display an exemplary user interface 220 shown in FIG. 2B. The user interface 220 can be the main interface of the “Video” application. The electronic device 100 can detect an operation (e.g., a tap operation) performed by the user on the video playback option 221 within the user interface 220, and in response to this operation, the electronic device 100 can obtain a video information stream and a caption information stream corresponding to the video.
[0048] The video information stream and the caption information stream may be files downloaded by the electronic device 100 from the server of the video application, or may be files obtained from the electronic device 100. Both the video file and the caption file carry time information.
[0049] It can be understood that FIGS. 2A and 2B only show exemplary user interfaces on the electronic device 100 and do not constitute a limitation to this embodiment of the present application.
[0050] Phase 2: Decoding of the Video
[0051] S103: The video application on the electronic device 100 transmits the video information stream to the video decoding module on the electronic device 100.
[0052] Specifically, after obtaining the video information stream, the video application may transmit the video information stream to the video decoding module.
[0053] S104 and S105: The video decoding module on the electronic device 100 decodes the video information stream to generate video frames and transmits the video frames to the video frame synthesis module on the electronic device 100.
[0054] Specifically, after receiving the video information stream transmitted by the video application, the video decoding module may decode the video information stream to generate video frames. The video frames may be all the video frames in the video playback process. The video frames may also be called image frames, and each video frame may carry the time information (i.e., time stamp) of the video frame. Then, the video decoding module transmits the video frames generated through decoding to the video frame synthesis module, and thereafter, may generate the video frames to be displayed.
[0055] The video decoding module may decode the video information stream by using the video decoding method of the prior art. This is not limited in this embodiment of the present application. For the specific implementation of the video decoding method, please refer to the technical documents related to video decoding. Details are not described here.
[0056] Phase 3: Decoding of captions
[0057] S106: The video application on the electronic device 100 transmits the caption information stream to the caption decoding module on the electronic device 100.
[0058] Specifically, after obtaining the caption information stream, the video application may send the caption information stream to the caption decoding module.
[0059] S107 and S108: The caption decoding module on the electronic device 100 decodes the caption information stream to generate a caption frame, and sends the caption frame to the video frame synthesis module on the electronic device 100.
[0060] Specifically, after receiving the caption information stream sent by the video application, the caption decoding module may decode the caption information stream to generate a caption frame. The caption frame may be all the caption frames in the video playback process. Each caption frame may include caption text, the display position of the caption text, the font color of the caption text, the font format of the caption text, etc., and may further carry the time information (i.e., timestamp) of the caption frame. Then, the caption decoding module sends the caption frame generated through decoding to the video frame synthesis module, and thereafter, may generate the video frame to be displayed.
[0061] The caption decoding module may decode the caption information stream by using the caption decoding method of the prior art. This is not limited in this embodiment of the present application. For the specific implementation of the caption decoding method, please refer to the technical documents related to caption decoding. Details are not described here.
[0062] Note that the example where the video decoding step in Phase 2 is executed first and then the caption decoding step in Phase 3 is executed is merely used in this embodiment of the present application. In some embodiments, the caption decoding step in Phase 3 may be executed first, and then the video decoding step in Phase 2 may be executed, or the video decoding step in Phase 2 and the caption decoding step in Phase 3 may be executed simultaneously. This is not limited in this embodiment of the present application.
[0063] Phase 4: Video Frame Composition, Rendering, and Display
[0064] S109 and S110: The video frame composition module on the electronic device 100 overlaps and combines the received video frame and the caption frame to generate a video frame to be displayed, and transmits the video frame to be displayed to the video frame queue on the electronic device 100.
[0065] Specifically, the video frame composition module performs matching based on the time information corresponding to the video frame and the time information corresponding to the caption frame. After the matching is completed, the caption frame is overlapped with the corresponding video frame and combined to generate a video frame to be displayed. Then, the video frame composition module may transmit the video frame to be displayed to the video frame queue.
[0066] S111 to S113: The video rendering module reads the video frame to be displayed from the video frame queue based on the time series, renders the video frame to be displayed based on the time series, and may generate a rendered video frame.
[0067] Specifically, the video rendering module can obtain in real time (or at intervals of a certain time period) the video frames to be displayed from the video frame queue. After the video frame composition module sends the video frames to be displayed to the video frame queue, the video rendering module can read the video frames to be displayed from the video frame queue based on the time series, render the video frames to be displayed, and generate the rendered video frames. Then, the video rendering module can send the rendered video frames to the video application.
[0068] The video rendering module can render the video frames to be displayed by using the video rendering method of the prior art. This is not limited in this embodiment of the present application. For the specific implementation of the video rendering method, please refer to the technical documents related to video rendering. Details are not described here.
[0069] S114: The electronic device 100 displays the rendered video frames.
[0070] Specifically, after receiving the rendered video frames sent by the video rendering module, the video application on the electronic device 100 can display the rendered video frames on the display screen of the electronic device 100 (i.e., the video playback window).
[0071] For example, FIG. 2C can be a picture of a frame within a rendered video frame that is displayed after the electronic device 100 executes the caption display method shown in FIGS. 1A and 1B. The captions "W S Y T K L D G S Y D Z M", "High-recognition caption", and "Caption with unclear color" are all bullet comments, and the display positions of the bullet comments are not fixed with respect to the display screen of the electronic device 100. The display position of the caption "Caption synchronized with audio" is fixed with respect to the display screen of the electronic device 100. As can be easily seen from FIG. 2C, since the color difference between both sides of the caption "W S Y T K L D G S Y D Z M" and the color of the video is small, the caption recognition degree is low, and as a result, the user cannot clearly see this caption. The color differences between the captions "High-recognition caption" and "Caption synchronized with audio" and the color of the video Is large are large, so the caption recognition degree is high and the user can clearly see these captions. Although the color difference between the caption "Caption with unclear color" and the color of the video is not small, due to the brightness of the video Is high in this case, the caption recognition degree is also low and the user cannot clearly see this caption.
[0072] As can be seen from FIG. 2C, when using the caption display method shown in FIGS. 1A and 1B, in an application scenario where captions are displayed while playing a video, if the color of the caption largely overlaps with the color and brightness of the video at the caption display position, the caption recognition degree will be low, making it difficult for the user to clearly see the caption, thus reducing the user experience.
[0073] To solve the foregoing problems, one embodiment of the present application provides another caption display method. The electronic device first obtains a video file to be played and a caption file to be displayed in the video playback window, and then performs video decoding on the video file to obtain video frames, and may perform caption decoding on the caption file to obtain caption frames. Next, the electronic device extracts caption color gamut information, caption position information, etc. from the caption frames, and based on the caption position information, extracts the color gamut information at the caption display position corresponding to the caption within the video frame. Then, based on the caption color gamut information and the color gamut information at the caption display position corresponding to the caption within the video frame, the caption recognition degree can be calculated. Caption recognition degree Is low In this case, the electronic device adds a mask to the caption, calculates the color value and transparency of the mask based on the caption recognition degree, generates a masked caption frame, and performs alignment and matching between the video frame and the caption frame with the mask based on the time series to synthesize the final video frame to be displayed. The electronic device buffers the video frame to be displayed in a video queue, then reads and renders the video frame to be displayed based on the time series, and finally displays the rendered video frame in the video playback window. In this way, the problem of low caption recognition degree can be solved by adjusting the color and transparency of the caption mask without changing the color of the caption selected by the user. In addition, thereby, the shielding of the video content by the caption is reduced, and the specific visibility of the video content can be ensured, thereby improving the user experience.
[0074] Hereinafter, another caption display method provided in one embodiment of the present application will be described.
[0075] FIG. 3A, FIG. 3B, FIG. 3C, and FIG. 3D show an example of the method steps of another caption display method according to an embodiment of the present application.
[0076] As shown in FIG. 3A, FIG. 3B, FIG. 3C, and FIG. 3D, the method can be applied to an electronic device 100 having a video playback function. Hereinafter, specific steps of the method will be described in detail.
[0077] Phase 1: Acquisition of Video Information Stream and Caption Information Stream
[0078] S301 and S302: The electronic device 100 detects an operation of playing a video on a video application by the user, and in response to this operation, the electronic device 100 may acquire a video information stream and a caption information stream.
[0079] For the specific execution process of steps S301 and S302, please refer to the relevant content of steps S101 and S102 in the embodiments shown in FIGS. 1A and 1B. Details will not be described again here.
[0080] Phase 2: Decoding of Video
[0081] S303: The video application on the electronic device 100 transmits the video information stream to the video decoding module on the electronic device 100.
[0082] S304 and S305: The video decoding module on the electronic device 100 decodes the video information stream to generate video frames, and transmits the video frames to the video frame synthesis module on the electronic device 100.
[0083] For the specific execution process from step S303 to step S305, please refer to the relevant content of steps S103 to S105 in the embodiments shown in FIGS. 1A and 1B. Details will not be described again here.
[0084] Phase 3: Decoding of Captions
[0085] S306: The video application on the electronic device 100 transmits the caption information stream to the caption decoding module on the electronic device 100.
[0086] S307: The caption decoding module on the electronic device 100 decodes the caption information stream to generate a caption frame.
[0087] For the specific execution processes of steps S306 and S307, refer to the relevant content of steps S106 and S107 in the embodiments shown in FIGS. 1A and 1B. Details are not described here again.
[0088] FIG. 4 shows an example of a caption frame generated by decoding a caption information stream by a caption decoding module.
[0089] As shown in FIG. 4, the area inside the rectangular solid line box may represent a caption frame display area (or, also referred to as a video playback window area), and may coincide with the video frame display area. In this area, for example, one or more captions such as "W S Y T K L D G S Y D Z M", "High-recognition captions", "Captions with unclear colors", "Captions synchronized with audio" may be displayed. "W S Y T K L D G S Y D Z M", "High-recognition captions", etc. may each be referred to as one caption, and all the captions displayed in this area may be referred to as a caption group. For example, a group of captions such as "W S Y T K L D G S Y D Z M", "High-recognition captions", "Captions with unclear colors", "Captions synchronized with audio" may be referred to as a caption group.
[0090] Each caption shown in FIG. 4Surrounding The rectangular dashed box is merely an auxiliary element used to identify the position of each caption and may not be displayed in the video playback process.
[0091] Based on the above description and the description of captions and caption groups, as shown in Figure 2C, it can be easily understood that there are four captions displayed in the picture shown in Figure 2C, namely "W S Y T K L D G S Y D Z M", "Caption with high recognition", "Caption with unclear color", and "Caption synchronized with audio". These four captions form a caption group.
[0092] S308: The caption decoding module on the electronic device 100 extracts caption position information, caption color gamut information, etc. of each caption from the caption frame to generate caption group information.
[0093] Specifically, after generating the caption frame, the caption decoding module can extract caption position information, caption color gamut information, etc. of each caption from the caption frame to generate caption group information. The caption position information is the display position of each caption in the caption frame display area, and the caption color gamut information may include the color values of each caption. The caption group information may include the caption position information and caption color gamut information of all captions in the caption frame.
[0094] Optionally, the caption color gamut information may also include information such as the brightness of the caption.
[0095] Hereinafter, the process of extracting caption position information and caption color gamut information will be described in detail.
[0096] 1. Process of extracting caption position information:
[0097] The caption display position area can be the inner area of the rectangular dashed box shown in FIG. 4 that can exactly cover the caption, or another inner area of any shape that can cover the caption. This is not limited in this embodiment of the present application.
[0098] In this embodiment of the present application, for the purpose of explaining the process of extracting caption position information, an example where the inner area of the rectangular dashed box is the caption display position area is used.
[0099] Taking the extraction of the caption position information of the caption "W S Y T K L D G S Y D Z M" shown in FIG. 4 as an example. The caption decoding module can first establish an X - O - Y plane rectangular coordinate system within the caption frame display area, and then select a point within the caption frame display area (for example, the vertex at the lower left corner of the solid - line box of the rectangle) as the reference coordinate point O. The coordinates of the reference coordinate point O can be set to (0,0). As can be known from mathematical knowledge, the coordinates (x1,y1), (x2,y2), (x3,y3), and (x4,y4) of the four vertices of the outer rectangular dashed box of the caption "W S Y T K L D G S Y D Z M" can be calculated. In this case, the position information of the caption "W S Y T K L D G S Y D Z M" can include the coordinates of the four vertices of the rectangular dashed box. Alternatively, since the rectangle is a regular shape, the position area of the rectangle can be determined only by determining the coordinates of two vertices on the diagonal of the rectangular dashed box. Therefore, the position information of the caption "W S Y T K L D G S Y D Z M" can include only the coordinates of two vertices on the diagonal of the rectangular dashed box.
[0100] Similarly, the caption position information of other captions shown in FIG. 4 can also be extracted by using the aforementioned caption position extraction method. Details will not be described again here.
[0101] After the caption decoding module determines the position information of all captions in the caption frame, it indicates that the caption decoding module has completed the extraction of the caption position information.
[0102] It should be noted that the above-mentioned caption position information extraction process is only a possible implementation form for extracting caption position information. The implementation for extracting caption position information may alternatively be another implementation form in the prior art. This is not limited in this embodiment of the present application.
[0103] 2. Process for extracting caption color gamut information:
[0104] First, the concepts related to the process for extracting caption color gamut information will be described below.
[0105] Color value:
[0106] A color value is a group of color values corresponding to colors in a color mode. The RGB color mode is used as an example. In the RGB color mode, colors are formed by mixing red, green, and blue, and the color value of each color can be represented by (r, g, b), where r, g, and b represent the values of the three primary colors of red, green, and blue respectively, and the value range is [0, 255]. For example, the color value of red can be represented by (255, 0, 0), the color value of green can be represented by (0, 255, 0), the color value of blue can be represented by (0, 0, 255), the color value of black can be represented by (0, 0, 0), and the color value of white can be represented by (255, 255, 255).
[0107] Color gamut:
[0108] A color gamut is a set of color values, that is, a set of colors that can be generated in a specific color mode. In the RGB color mode, up to 256×256×256 = 16777216 different colors, that is, 2 24 different colors can be generated, and it can be easily understood that the color gamut is [0, 2 24 -1]. 224 A plurality of different colors and color values corresponding to each color can form a color value table, and the color values corresponding to each color can be found within the color value table.
[0109] After extracting the caption position information, the caption decoding module can search the color value table to obtain the color value corresponding to the font color of the caption at the caption position in order to determine the color value of the caption based on the font color of the caption at the caption position.
[0110] After the caption decoding module determines the color values of all the captions within the caption frame, it indicates that the caption decoding module has completed the extraction of the caption color gamut information.
[0111] S309: The caption decoding module on the electronic device 100 sends an instruction to obtain the mask parameter of the caption group to the video frame color gamut interpretation module on the electronic device 100, and this instruction carries the time information of the caption frame, the caption group information, etc.
[0112] Specifically, after generating the caption group information, the caption decoding module can send an instruction to obtain the mask parameter of the caption group to the video frame color gamut interpretation module. The instruction is used to instruct the video frame color gamut interpretation module to send the mask parameter corresponding to the caption group (including the color value and transparency of the mask) to the caption decoding module. The color value and transparency are sometimes referred to as a group of mask parameters. The instruction can carry the time information of the caption frame, the caption group information, etc. The time information of the caption frame can be used to obtain the video frame corresponding to the caption group in subsequent steps, and the caption group information can be used to analyze the caption recognition degree in subsequent steps.
[0113] S310: The video frame color gamut interpretation module on the electronic device 100 sends an instruction to the video decoding module on the electronic device 100 to obtain a video frame corresponding to the caption group. This instruction conveys time information of the caption frame and the like.
[0114] Specifically, after receiving an instruction to obtain the mask parameter of the caption group sent by the caption decoding module, the video frame color gamut interpretation module may send an instruction to the video decoding module to obtain a video frame corresponding to the caption group. The instruction is used to instruct the video decoding module to send the video frame corresponding to the caption group to the video frame color gamut interpretation module. The instruction may convey the time information of the caption frame, and the time information of the caption frame may be used by the video decoding module to find the video frame corresponding to the caption group.
[0115] S311 and S312: The video decoding module on the electronic device 100 searches for a video frame corresponding to the caption group and sends the video frame corresponding to the caption group to the video frame color gamut interpretation module on the electronic device 100.
[0116] Specifically, after the video decoding module receives an instruction to obtain a video frame corresponding to a caption group sent by the video frame color gamut interpretation module, the video decoding module can find the video frame corresponding to the caption group based on the time information of the caption frame carried in the instruction. Since the video decoding module obtains the time information of all video frames through decoding in the video decoding phase, the video decoding module can match the time information of all video frames with the time information of the caption frame. When the matching is successful (i.e., when the time information of the video frame matches the time information of the caption frame), the video frame is the video frame corresponding to the caption group. Thereafter, the video decoding module can send the video frame corresponding to the caption group to the video frame color gamut interpretation module.
[0117] S313: The video frame color gamut interpretation module on the electronic device 100 obtains the color gamut information at the position of each caption in the video frame corresponding to the caption group based on the caption position information in the caption group information.
[0118] Specifically, after obtaining the video frame corresponding to the caption group, the video frame color gamut interpretation module can determine the video frame area corresponding to the position of each caption based on each caption position information in the caption group information. Further, the video frame color gamut interpretation module can calculate the color gamut information of the video frame area corresponding to the position of each caption.
[0119] Hereinafter, the process by which the video frame color gamut interpretation module calculates the color gamut information of the video frame area corresponding to the position of each caption will be described in detail.
[0120] Assume that the caption "W S Y T K L D G S Y D Z M" in the picture shown in FIG. 2C is Caption 1. An example where the video frame color gamut interpretation module calculates the color gamut information of the video frame area corresponding to Caption 1 is used for the purpose of explanation.
[0121] As shown in FIG. 5, the video frame area corresponding to the position of Caption 1 may be the inner area of the solid-line box of the upper rectangle in FIG. 5. Since pixel areas of different color gamuts may exist within one video frame area, one video frame area may be divided into a plurality of sub-areas, and each sub-area may be called a video frame color gamut extraction unit. The division of the sub-areas may be performed based on a preset width, or may be divided based on the width of each character in the caption. For example, Caption 1 has a total of 13 characters. In this case, in FIG. 5, the video frame area corresponding to the position of Caption 1 is divided into 13 sub-areas, that is, 13 video frame color gamut extraction units, based on the width of each character in Caption 1.
[0122] Furthermore, the video frame color gamut interpretation module may sequentially calculate the color gamut information of all sub-areas from left to right (or from right to left). Taking the calculation of the color gamut information of one sub-area within the video frame area as an example. The video frame color gamut interpretation module may obtain the color values of all pixels within the sub-area, and then perform superposition and averaging on the color values of all pixels to obtain the average value of the color values of all pixels within the sub-area. The average value is the color value of the sub-area, and the color value of the sub-area is the color gamut information of the sub-area.
[0123] For example, assuming that the sub-area is m pixels wide and n pixels high, the sub-area has a total of m * n pixels, and the color value x of each pixel can be represented by (r, g, b). In this case, the average value of the color values of all pixels within the sub-area
Equation
Number
[0124] r i is the average red value of all pixels in the sub-region, and g i is the average green value of all pixels in the sub-region, and b i is the average blue value of all pixels in the sub-region.
Number
Number
Number
[0125] Similarly, the video frame color gamut interpretation module can calculate the color gamut information of all sub-regions of the video frame region corresponding to the position of each caption, that is, the color gamut information of the caption position in the video frame corresponding to the caption group.
[0126] It should be understood that the number of sub-regions into which the video frame region corresponding to the caption is divided may be determined according to a preset division rule. This is not limited in this embodiment of the present application.
[0127] Optionally, the color gamut information of the video frame region may also include information such as the luminance of the video frame region.
[0128] It should be noted that the above-mentioned process of calculating the color gamut information of the video frame region corresponding to the position of each caption is only a possible implementation form, and another implementation form may be used. This is not limited in this embodiment of the present application.
[0129] S314: The video frame color gamut interpretation module on the electronic device 100 generates a superimposed caption recognition analysis result based on each caption color gamut information in the caption group information and the color gamut information at each caption position in the video frame corresponding to the caption group.
[0130] Specifically, after calculating the color gamut information at the caption position in the video frame corresponding to the caption group, the video frame color gamut interpretation module may perform a superimposed caption recognition analysis based on the caption color gamut information in the caption group information and the color gamut information at the caption position in the video frame corresponding to the caption group. Furthermore, the superimposed caption recognition analysis result may be generated through the superimposed caption recognition analysis, and the result is used to indicate the magnitude of the recognition degree (also called the magnitude of discrimination) of each caption in the caption group.
[0131] In other words, after the caption group is superimposed on the caption position in the video frame corresponding to the caption group, the video frame color gamut interpretation module may determine the difference between the color of the caption and the color of the video frame area corresponding to the caption. If the difference is small, it indicates that the recognition degree of the caption is low and it is not easily recognized by the user.
[0132] Hereinafter, the process by which the video frame color gamut interpretation module performs the superimposed caption recognition analysis will be described in detail.
[0133] The video frame color gamut interpretation module may determine a color difference value between the color of the caption and the color of the video frame area corresponding to the caption, and the color difference value is used to indicate the difference between the color of the caption and the color of the video frame area corresponding to the caption. The color difference value may be determined by using related algorithms in the prior art.
[0134] In a possible implementation form, the color difference value Diff can be calculated by using the following formula: [Number]
[0135] k is the number of all sub-regions of the video frame region corresponding to the caption, and r i is the average red value of all pixels in the sub-region, and g i is the average green value of all pixels in the sub-region, and b i is the average blue value of all pixels in the sub-region. r 0 is the red value of the caption, and g 0 is the green value of the caption, and b 0 is the blue value of the caption.
[0136] Furthermore, after obtaining the color difference value through calculation, the video frame color gamut interpretation module can determine the magnitude of caption recognition by determining whether the color difference value is less than a preset color difference threshold value.
[0137] If the color difference value is less than a preset color difference threshold value (which may also be called the first threshold value), it indicates that the caption recognition is low.
[0138] In some embodiments, the caption recognition can be further determined with reference to the luminance of the video frame region corresponding to the caption.
[0139] For example, although the color difference value between the color of the caption "Caption with unclear color" shown in FIG. 2C and the color of the video frame area corresponding to the caption is not so small, since the luminance of the video frame area corresponding to the caption is too high, the problem that the caption recognition degree is low still exists. Therefore, in this case, the caption recognition degree can be further determined with reference to the luminance of the video frame area corresponding to the caption. When the luminance of the video frame area corresponding to the caption is higher than a preset luminance threshold value, it indicates that the caption recognition degree is low.
[0140] In the case of a solid-color caption, the extracted caption color gamut information may include only one parameter, that is, only one color value corresponding to the caption. In the case of a non-solid-color caption, the extracted caption color gamut information may include a plurality of parameters. For example, in the case of a gradient color caption, the extracted caption color gamut information may include a plurality of parameters such as a start point color value, an end point color value, and a gradient direction. In this case, in a possible implementation form, the average value of the start point color value and the end point color value of the caption may be calculated first, and then the average value is used as the color value corresponding to the caption to perform an overlapping caption recognition degree analysis.
[0141] It should be noted that the foregoing process in which the video frame color gamut interpretation module performs an overlapping caption recognition degree analysis is only a possible implementation form, and another implementation form may be used. This is not limited in this embodiment of the present application.
[0142] S315: The video frame color gamut interpretation module on the electronic device 100 calculates the color value and transparency of the mask corresponding to each caption in the caption group based on the overlapping caption recognition degree analysis result.
[0143] Specifically, after generating the overlapping caption recognition degree analysis result, the video frame color gamut interpretation module may calculate the color value and transparency of the mask corresponding to each caption in the caption frame based on the result.
[0144] In the case of a caption with high recognition (for example, the caption "Caption with high recognition" in FIG. 2C or the caption "Caption synchronized with audio"), the color value of the mask corresponding to the caption can be a preset fixed value, and the transparency can be set to 100%.
[0145] In the case of a caption with low recognition (for example, the caption "W S Y T K L D G S Y D Z M" in FIG. 2C or the caption "Caption with unclear color"), the color value and transparency of the mask corresponding to the caption need to be further determined based on the color gamut information of the caption or the color gamut information of the video frame area corresponding to the position of the caption.
[0146] There can be many specific methods for determining the color value and transparency of the mask corresponding to the caption. This is not limited in this embodiment of the present application, and those skilled in the art may select a method according to requirements.
[0147] In a possible implementation form, the color value corresponding to the color with the maximum color difference value from the color value of the caption or the color value of the video frame area corresponding to the caption can be determined as the color value of the mask corresponding to the caption. In this way, the user can see the caption more clearly. Alternatively, the color value corresponding to the color with the intermediate color difference value from the color value of the caption or the color value of the video frame area corresponding to the caption may be determined as the color value of the mask corresponding to the caption, thereby ensuring that the user can see the caption clearly while avoiding the discomfort of the eyes caused by a large color difference.
[0148] For example, the electronic device 100 can calculate the color difference value Diff between the color value corresponding to each color in the color value table and the color value of the caption, and then can select the color value corresponding to the color with the maximum / intermediate color difference value Diff as the color value of the mask. In a possible implementation, the color difference value Diff between the color value corresponding to each color in the color value table and the color value of the caption can be calculated by using the following formula: [Number]
[0149] Assuming that the color value corresponding to a specific color in the color value table is (R 0 , G 0 , B 0 ), R 0 is the red value corresponding to that color, G 0 is the green value corresponding to that color, and B 0 is the blue value corresponding to that color. r 0 is the red value of the caption, g 0 is the green value of the caption, and b 0 is the blue value of the caption.
[0150] As another example, the electronic device 100 can calculate the color difference value Diff between the color value corresponding to each color in the color value table and the color value of the video frame area corresponding to the caption, and then can select the color value corresponding to the color with the maximum / intermediate color difference value Diff as the color value of the mask. In a possible implementation, the color difference value Diff between the color value corresponding to each color in the color value table and the color value of the video frame area corresponding to the caption can be calculated by using the following formula: [Number]
[0151] Assuming that the color value corresponding to a specific color in the color value table is (R 0 , G 0 , B 0 ), R 0is the red value corresponding to the color, and G 0 is the green value corresponding to the color, and B 0 is the blue value corresponding to the color. k is the number of all sub-regions of the video frame region corresponding to the caption, and r i is the average red value of all pixels within the sub-region, and g i is the average green value of all pixels within the sub-region.
[0152] In a possible implementation form, the transparency of the mask corresponding to the caption can be further determined based on the color value of the mask corresponding to the caption. For example, when the color value of the mask corresponding to the caption is significantly different from the color value of the caption, for the transparency of the mask corresponding to the caption Large a threshold value (for example, a value greater than 50%) can be selected to ensure that the user can clearly see the caption while reducing the shielding of the video picture by the caption overlay area.
[0153] S316: The video frame color gamut interpretation module on the electronic device 100 transmits the color value and transparency of the mask corresponding to each caption in the caption group to the caption decoding module on the electronic device 100.
[0154] Specifically, after calculating the color value and transparency of the mask corresponding to each caption in the caption group, the video frame color gamut interpretation module may transmit the color value and transparency of the mask corresponding to each caption in the caption group to the caption decoding module, and the caption position information of the caption corresponding to the mask may also be carried, so that the caption decoding module can map the caption to the mask one-to-one.
[0155] S317: The caption decoding module on the electronic device 100 generates a corresponding mask based on the color value and transparency of the mask corresponding to each caption in the caption group, and overlays each caption in the caption group with the corresponding mask to generate a caption frame with a mask.
[0156] Specifically, after receiving the color value and transparency of the mask corresponding to each caption in the caption group transmitted by the video frame color gamut interpretation module, the caption decoding module can generate a mask corresponding to the caption (for example, the mask corresponding to caption 1 shown in FIG. 5) based on the color value and transparency of the mask corresponding to the caption and the caption position information of the caption. The shape of the mask may be rectangular or any other shape that can cover the caption. This is not limited in this embodiment of the present application.
[0157] Similarly, the caption decoding module can generate a mask corresponding to each caption in the caption group.
[0158] For example, as shown in FIG. 2C, it can be easily seen that there are four captions in the picture. Therefore, the caption decoding module can generate four masks, and one caption corresponds to one mask.
[0159] Furthermore, the caption decoding module can overlay the caption on the upper layer of the mask corresponding to the caption to generate a caption with a mask (for example, caption 1 with a mask shown in FIG. 5).
[0160] Similarly, the caption decoding module can overlay each caption in the caption group on the corresponding mask to generate a caption frame with a mask.
[0161] FIG. 6A shows an example of a caption frame with a mask. It can be seen that a mask is superimposed on each caption. The transparency of the mask corresponding to a caption with high recognition (e.g., "caption with high recognition" or "caption synchronized with audio") is 100%, and the transparency of the mask corresponding to a caption with low recognition (e.g., "W S Y T K L D G S Y D Z M" or "caption with unclear color") is less than 100%, and there are specific color values.
[0162] S318: The caption decoding module on the electronic device 100 transmits the caption frame with a mask to the video frame synthesis module on the electronic device 100.
[0163] Specifically, after generating the caption frame with a mask, the caption decoding module transmits the caption frame with a mask to the video frame synthesis module, and then a video frame to be displayed can be generated.
[0164] Phase 4: Synthesis, rendering, and display of video frames
[0165] S319 and S320: The video frame synthesis module on the electronic device 100 overlays and combines the received video frame and the caption frame with a mask to generate a video frame to be displayed, and transmits the video frame to be displayed to the video frame queue on the electronic device 100.
[0166] S321 to S323: The video rendering module reads the video frame to be displayed from the video frame queue based on the time series, renders the video frame to be displayed based on the time series, and can generate a rendered video frame.
[0167] S324: The electronic device 100 displays the rendered video frame.
[0168] For the specific execution process from step S319 to step S324, refer to the relevant content of steps S109 to S114 in the embodiments shown in FIGS. 1A and 1B. Details will not be described again here.
[0169] Note that in some embodiments, the video decoding module, caption decoding module, video frame color gamut interpretation module, video frame composition module, video frame queue, and video rendering module may alternatively be integrated into the video application to execute the caption display method provided in this embodiment of the present application. This is not limited in this embodiment of the present application.
[0170] For example, FIG. 6B may be a picture of a frame in a rendered video frame displayed after the electronic device 100 executes the caption display method shown in FIGS. 3A, 3B, 3C, and 3D (one caption may correspond to one mask). Compared with the picture shown in FIG. 2C, it can be easily seen that after the corresponding mask is added to the caption group, the recognition degree of the caption "W S Y T K L D G S Y D Z M" and the recognition degree of the caption "Caption with unclear color" are significantly improved. In addition, since the mask corresponding to the caption has a specific transparency, the caption overlay area does not completely obscure the video picture. In this way, the effects of video display and caption display are comprehensively considered, thereby ensuring that the user can clearly see the caption while ensuring a specific visibility of the video picture without changing the color of the caption selected by the user, thereby improving the user experience.
[0171] Furthermore, during the entire video playback process, the position of the caption, the color of the video background, etc. may change. Therefore, the above-described caption display method can always be executed so that the user can clearly see the caption during the entire video playback process. For example, FIG. 6B may be a schematic diagram of a first user interface when the video playback progress is at moment 8:00, and FIG. 6C may be a schematic diagram of a second user interface when the video playback progress is at moment 8:02. The video frame included in the first user interface is different from the video frame included in the second user interface. As shown in FIG. 6C, it can be seen that compared with FIG. 6B, the captions "W S Y T K L D G S Y D Z M", "Caption with high recognition", and "Caption with unclear color" have all moved to the left side of the display screen. The electronic device 100 recalculates the color value and transparency of the mask corresponding to the caption based on the color value of the caption and the color value of the current video frame area corresponding to the caption, and generates a mask corresponding to the caption. In the second user interface, it can be easily seen that the video background color of the caption "W S Y T K L D G S Y D Z M" corresponding to the current video frame area changes, and the recognition degree of the caption also increases. Therefore, compared with FIG. 6B, the mask corresponding to the caption "W S Y T K L D G S Y D Z M" also changes. It can be seen that no mask is displayed for the caption. Specifically, the transparency of the mask corresponding to the caption is changed to 100%, or there is no mask for the caption.
[0172] The video playback pictures shown in FIGS. 6B and 6C may be displayed in full-screen mode or in partial-screen mode. This is not limited in this embodiment of the present application.
[0173] The mask corresponding to the caption shown in FIG. 6B is a mask that spans the area where the entire caption is located, that is, one caption corresponds to only one mask. In some actual application scenarios, one caption may Large span multiple regions with significant color gamut differences. As a result, the recognition rate of some parts of the caption Is high decreases, and the recognition rate of other parts of the caption Is low decreases. In this case, multiple corresponding masks can be generated for one caption. For example, for the caption "W S Y T K L D G S Y D Z M" shown in FIG. 2C, the caption recognition rate of the beginning part of the caption area is low (that is, the four characters "W S Y T" are difficult for the user to recognize), and the caption recognition rate of the end part of the caption area is also low (that is, the four characters "Y D Z M" are difficult for the user to recognize). The caption recognition rate of the middle part of the caption area is high (that is, the four characters "K L D G S" are easy for the user to recognize). Therefore, in this case, one corresponding mask can be generated for each of the beginning part, middle part, and end part of the caption area, that is, the caption can have three corresponding masks.
[0174] In the case of the foregoing application scenario where one caption corresponds to multiple masks, in this embodiment of the present application, several corresponding improvements can be made to steps S313 to S317 based on the methods shown in FIGS. 3A, 3B, 3C, and 3D so that one caption corresponds to multiple masks. The other steps do not need to be changed.
[0175] Hereinafter, the process in which one caption corresponds to multiple masks will be described in detail.
[0176] In the process of generating color gamut information at the caption position within the video frame corresponding to the caption group, the video frame color gamut interpretation module may sequentially calculate the color values of all sub-regions in order from left to right (or from right to left). In the aforementioned application scenario where one caption needs to correspond to multiple masks, that is, one caption Is large In the application scenario spanning multiple regions with a significant color gamut difference, the video frame color gamut interpretation module may compare the color values of adjacent sub-regions. If the color values of adjacent sub-regions are close, the adjacent sub-regions are combined into one region, and the combined region corresponds to one mask. If the color values of adjacent sub-regions are significantly different, the adjacent sub-regions are not combined, and the two uncombined regions correspond to their respective masks. Therefore, one caption may correspond to multiple masks.
[0177] As shown in FIG. 7A, when one caption may correspond to multiple masks, steps S313 to S317 may specifically be executed based on the following steps. Hereinafter, the case where caption 1 shown in FIG. 7B is the caption "W S Y T K L D G S Y D Z M" shown in FIG. 2C will be used for explanation.
[0178] S701: The video frame color gamut interpretation module sequentially calculates the color values of all sub-regions of the video frame region corresponding to the position of the caption, combines the sub-regions with close color values, and obtains M second sub-regions.
[0179] Specifically, based on step S313, after calculating the color values of all sub-regions sequentially from left to right (or from right to left), the video frame color gamut interpretation module further needs to compare the color values of adjacent sub-regions and combine the sub-regions with similar color values to obtain M second sub-regions, where M is a positive integer. As shown in FIG. 7B, after comparing the color values of adjacent sub-regions and combining the sub-regions with similar color values, the video frame color gamut interpretation module divides the video frame region corresponding to the caption position into three regions, namely region A, region B, and region C (i.e., three second sub-regions). Region A is a pieces of formed by combining sub-regions, region B is formed by combining b sub-regions, and region C is assumed to be formed by combining c sub-regions.
[0180] Saying that the color values are close may mean that the difference value between the color values of two sub-regions is less than a second threshold value, and the second threshold value is preset.
[0181] S702: The video frame color gamut interpretation module separately performs an analysis of the superimposed caption recognition degree on the M second sub-regions to generate the analysis results of the superimposed caption recognition degree of the M second sub-regions.
[0182] Specifically, the video frame color gamut interpretation module does not directly perform an analysis of the superimposed caption recognition degree on the entire video frame region, but needs to separately perform an analysis of the superimposed caption recognition degree on region A, region B, and region C. Similarly, the video frame color gamut interpretation module can separately perform an analysis of the superimposed caption recognition degree on region A, region B, and region C using the color difference value in step S314. The process is as follows.
[0183] Color difference value Diff1 of region A:
Number
[0184] a is the number of sub - regions included in region A, r i is the average red value of all pixels within the sub - regions in region A, g i is the average green value of all pixels within the sub - regions in region A, b i is the average blue value of all pixels within the sub - regions in region A. r 0 is the red value of the caption in region A, g 0 is the green value of the caption in region A, b 0 is the blue value of the caption in region A.
[0185] Color difference value Diff2 of region B:
Number
[0186] b is the number of sub - regions included in region B, r i is the average red value of all pixels within the sub - regions in region B, g i is the average green value of all pixels within the sub - regions in region B, b i is the average blue value of all pixels within the sub - regions in region B. r 0 is the red value of the caption in region B, g 0 is the green value of the caption in region B, b 0 is the blue value of the caption in region B.
[0187] Color difference value Diff3 of region C:
Number
[0188] c is the number of sub - regions included in region C, r i is the average red value of all pixels within the sub - regions in region C, g i is the average green value of all pixels within the sub - regions in region C, b iis the average blue value of all pixels in the sub-region within region C. r 0 is the red value of the caption in region C, and g 0 is the green value of the caption in region C, and b 0 is the blue value of the caption in region C.
[0189] After separately obtaining the color difference values of region A, region B, and region C through calculation, the video frame color gamut interpretation module can determine whether the color difference values of the three regions are less than a preset color difference threshold. If the color difference value of a region is less than the preset color difference threshold, it indicates that the caption recognition degree of the region is low.
[0190] S703: The video frame color gamut interpretation module determines the color value and transparency of the mask corresponding to each of the M second sub-regions based on the superimposed caption recognition degree analysis results of the M second sub-regions and the caption color gamut information.
[0191] Specifically, the video frame color gamut interpretation module needs to determine the color value and transparency of the mask corresponding to region A, the color value and transparency of the mask corresponding to region B, and the color value and transparency of the mask corresponding to region C based on the caption color gamut information and the superimposed caption recognition degree analysis results of region A, region B, and region C. The specific process of determining the color value and transparency of the mask corresponding to each second sub-region is the same as the process of determining the color value and transparency of the mask corresponding to the entire video frame region corresponding to the position of the caption in step S315. Please refer to the relevant content mentioned above. Details will not be explained again here.
[0192] S704: The video frame color gamut interpretation module sends the color value, transparency, and position information of the mask corresponding to the M second sub-regions to the caption decoding module.
[0193] Specifically, since one caption can correspond to multiple masks, the video frame color gamut interpretation module not only needs to send the color values and transparencies of the masks corresponding to each caption in the caption group to the caption decoding module, but also needs to send the position information of each mask (or the position information of each mask relative to the corresponding caption) to the caption decoding module. The position information of each mask can be obtained based on the caption position information. Specifically, when one caption corresponds to multiple masks, since the caption position information is known, the position information of all sub-regions within the video frame region at the caption position can be estimated. Further, the position information of the mask corresponding to each second sub-region can be estimated.
[0194] S705: The caption decoding module generates a mask corresponding to the caption based on the color values, transparencies, and position information of the masks corresponding to the M second sub-regions, and superimposes the caption on this mask to generate a caption with a mask.
[0195] Specifically, in the case of one caption corresponding to multiple masks, the caption decoding module can generate three masks corresponding to the caption (for example, the masks corresponding to caption 1 shown in FIG. 7B) based on the color values, transparencies, and mask position information of the masks of each second sub-region corresponding to the caption. Then, the caption decoding module can superimpose the caption on the upper layer of the mask corresponding to the caption to generate a caption with a mask (for example, caption 1 with a mask shown in FIG. 7B).
[0196] As shown in FIG. 2C, the three captions "caption with high recognition", "caption with unclear color", and "caption synchronized with audio" Large do not span multiple regions with significant color gamut differences, so each of the three captions still corresponds to one mask.
[0197] The caption decoding module can generate a caption frame with a mask by superimposing each caption in the caption group on the corresponding mask.
[0198] FIG. 8A shows an example of a caption frame with a mask. It can be seen that three masks are superimposed on the caption "W S Y T K L D G S Y D Z M". Since the recognition degrees of "W S Y T" and "Y D Z M" are low, the transparency of the corresponding masks is less than 100%, and specific color values exist. Since the recognition degree of "K L D G S" is high, the transparency of the corresponding mask is 100%. One mask is superimposed on each of the other three captions. Since the recognition degrees of the caption "caption with high recognition degree" and the caption "caption synchronized with audio" are high, the transparency of the corresponding masks is 100%. Since the recognition degree of the caption "caption with unclear color" is low, the transparency of the corresponding mask is less than 100%, and specific color values exist.
[0199] For example, FIG. 8B can be a picture of a frame in a rendered video frame displayed after the electronic device 100 executes the improved caption display method shown in FIGS. 3A, 3B, 3C, and 3D. (Large Captions spanning multiple regions with significant color gamut differences can correspond to multiple masks. Compared with the picture shown in FIG. 6B, the caption "W S Y T K L D G S Y D Z M" LargeSince it spans multiple regions with different color gamut differences, the mask corresponding to the caption is changing. Since the middle part of the caption area (i.e., the part "K L D G S") has a high caption recognition rate, it can be easily seen that the transparency of the mask corresponding to that part is set to 100% (i.e., completely transparent), or the mask may not be set. Since the start part (i.e., the part "W S Y T") and the end part (i.e., the part "Y D Z M") of the caption area have a low caption recognition rate, the color value and transparency of the masks corresponding to these two parts are calculated based on the caption color gamut information and the color gamut information of the regions where the two parts are located, respectively. In this way, since the transparency of the mask corresponding to the middle part of the area of the caption "W S Y T K L D G S Y D Z M" is 100% or the mask may not be set, based on achieving the beneficial effect shown in FIG. 6B, the shielding of the video picture by the mask is further reduced, and the user experience is further improved.
[0200] Furthermore, during the entire video playback process, the position of the caption, the color of the video background, etc. can change. Therefore, the aforementioned caption display method can always be executed so that the user can clearly see the caption during the entire video playback process. For example, FIG. 8B can be a schematic diagram of a user interface where the video playback progress is at moment 8:00 and includes a first video frame. FIG. 8C can be a schematic diagram of a user interface where the video playback progress is at moment 8:01 and includes a second video frame. The first video frame is the same as the second video frame. As shown in FIG. 8C, it can be seen that compared with FIG. 8B, the captions "W S Y T K L D G S Y D Z M", "Caption with high recognition", and "Caption with unclear color" have all moved to the left side of the display screen. The electronic device 100 recalculates the color value and transparency of the mask corresponding to the caption based on the color value of the caption and the color value of the current video frame area corresponding to the caption, and generates a mask corresponding to the caption. It can be easily seen that compared with FIG. 8B, the mask corresponding to the caption "W S Y T K L D G S Y D Z M" in FIG. 8C has changed significantly. In FIG. 8B, the parts with low caption recognition are "W S Y T" and "Y D Z M". Therefore, the masks corresponding to these two parts have specific color values, and the transparency of the corresponding masks is less than 100%. The part with high caption recognition is "K L D G S", so no mask is displayed for this part. Specifically, the transparency of the mask corresponding to the caption may be set to 100%, or the mask may not be set. However, in FIG. 8C, the parts with low caption recognition have changed to "W S Y T K" and "D Z M". Therefore, the electronic device 100 recalculates the color value and transparency of the masks corresponding to these two parts based on the color value of the caption and the color value of the current video frame area corresponding to the caption. The recognition of the two parts is low, so the masks corresponding to the two parts have specific color values. In addition, the transparency of the corresponding masks is less than 100%.The parts with high caption recognition have changed to "L D G S Y". Therefore, no mask is displayed for these parts. Specifically, the transparency of the mask corresponding to the caption may be set to 100%, or no mask may be set. The process of generating the mask corresponding to the caption in FIG. 8C is the same as the process of generating the mask corresponding to the caption in FIG. 8B. Details will not be described again here.
[0201] The video playback pictures shown in FIGS. 8B and 8C may be displayed in full-screen mode or in partial-screen mode. This is not limited in this embodiment of the present application.
[0202] In this embodiment of the present application, in the case of a caption with high recognition, the electronic device 100 generates a mask for the caption, where the color value of the mask can be a preset color value and the transparency of the mask is 100%. In some embodiments, in the case of a caption with high recognition, the electronic device 100 alternatively may not generate a mask for the caption. Specifically, when the electronic device 100 determines that the caption has high recognition, the electronic device 100 may not perform further processing on the caption. Therefore, the caption has no corresponding mask, that is, no mask is set for the caption.
[0203] In this embodiment of the present application, the fact that one caption corresponds to one mask (i.e., one caption corresponds to one group of mask parameters) may mean that one caption corresponds to one mask including one color value and transparency. The fact that one caption corresponds to a plurality of masks (i.e., one caption corresponds to a plurality of groups of mask parameters) may mean that one caption corresponds to a plurality of masks having different color values and different transparencies, and one caption corresponds to one mask having different color values and different transparencies (i.e., a plurality of masks having different color values and different transparencies are combined into one mask having different color values and different transparencies).
[0204] In this embodiment of the present application, a mobile phone is used as an example of the electronic device 100. Alternatively, the electronic device 100 may be a portable electronic device such as a tablet computer (Pad), a personal digital assistant (Personal Digital Assistant, PDA), or a laptop computer (Laptop). The type, physical form, and size of the electronic device 100 are not limited in this embodiment of the present application.
[0205] In an embodiment of the present application, the first video may be a video played by the electronic device 100 after the user taps the video playback option 221 shown in FIG. 2B. The first interface may be the user interface shown in FIG. 6B. The first picture may be a picture of the video frame shown in FIG. 6B. The first caption may be the caption "W S Y T K L D G S Y D Z M". The first region is within the first picture and corresponds to the display position of the first caption. The first value may be a color difference value between the color of the first caption and the color of the first picture region corresponding to the display position of the first caption. The second interface may be the user interface shown in FIG. 6C. The second picture may be a picture of the video frame shown in FIG. 6C. The second region is within the second picture and corresponds to the display position of the first caption. The second value may be a color difference value between the color of the first caption and the color of the second picture region corresponding to the display position of the first caption. The first video file may be a video file corresponding to the first video, and the first caption file may be a caption file corresponding to the first video. The first video frame is a video frame used to generate the first picture. The first caption frame includes the first caption and is a caption frame that carries the same time information as the first video frame. The second caption frame is a caption frame generated after the first caption is superimposed on the first mask (i.e., a caption frame with a mask). The first sub-region may be a video frame color gamut extraction unit. The second sub-region may be a region obtained after adjacent first sub-regions having similar color values are combined (e.g., region A, region B, and region C). The first sub-mask may be a mask corresponding to each second sub-region.The first mask may be a mask corresponding to the caption "W S Y T K L D G S Y D Z M" shown in FIG. 6B, or may be a mask corresponding to the caption "W S Y T K L D G S Y D Z M" shown in FIG. 8B. The third interface may be the user interface shown in FIG. 8B, and the third picture may be the picture of the video frame shown in FIG. 8B. The first part may be "W S Y T" in the caption "W S Y T K L D G S Y D Z M", and the second part may be "K L D G S" in the caption "W S Y T K L D G S Y D Z M". The second sub-mask may be a mask corresponding to "W S Y T" (i.e., the mask of region A shown in FIG. 7B). The third sub-mask may be a mask corresponding to "K L D G S" (i.e., the mask of region B shown in FIG. 7B). The second mask may be a mask corresponding to the caption "W S Y T K L D G S Y D Z M" shown in FIG. 6C.
[0206] Hereinafter, the structure of the electronic device 100 according to an embodiment of the present application will be described.
[0207] FIG. 9 shows an example of the configuration of the electronic device 100 according to an embodiment of the present application.
[0208] As shown in FIG. 9, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identity module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyro sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, an optical proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0209] It can be understood that the structure shown in this embodiment of the present application does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, some components may be combined, some components may be divided, or different component arrangements may be used. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0210] Processor 110 may include one or more processing units. For example, processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be separate devices or may be integrated into one or more processors.
[0211] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal based on an instruction operation code and a timing signal to execute control related to instruction fetching and execution.
[0212] The memory may be further disposed in the processor 110 and is configured to store instructions and data. In some embodiments, the memory within the processor 110 is a cache. The memory may store instructions or data that have just been used by the processor 110 or are used periodically. When the processor 110 needs to reuse an instruction or data, the processor 110 may directly call the instruction or data from the memory. This avoids repeated accesses, reduces the latency of the processor 110, and improves system efficiency.
[0213] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, a universal serial bus (USB) port, and / or the like.
[0214] I 2 The I2C interface is a bidirectional synchronous serial bus including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include a plurality of I2C buses. The processor 110 may be coupled to a touch sensor 180K, a charger, a camera flash, a camera 193, etc. via different I2C bus interfaces respectively. For example, the processor 110 may be coupled to the touch sensor 180K via an I2C interface, and the processor 110 and the touch sensor 180K may communicate with each other via the I2C bus interface, thereby implementing the touch function of the electronic device 100. 2 C buses. The processor 110 may be coupled to a touch sensor 180K, a charger, a camera flash, a camera 193, etc. via different I2C bus interfaces respectively. For example, the processor 110 may be coupled to the touch sensor 180K via an I2C interface, and the processor 110 and the touch sensor 180K may communicate with each other via the I2C bus interface, thereby implementing the touch function of the electronic device 100. 2 C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K via an I2C interface, and the processor 110 and the touch sensor 180K may communicate with each other via the I2C bus interface, thereby implementing the touch function of the electronic device 100.
[0215] I 2 The I2S interface may be used for audio communication. In some embodiments, the processor 110 may include a plurality of I2S2 It may include an I2S bus. The processor 110 may be coupled to the audio module 170 via the I2S bus to realize communication between the processor 110 and the audio module 170. 2 In some embodiments, the audio module 170 may be coupled to the wireless communication module 160 via the I2S interface to implement the function of answering calls via a Bluetooth headset. 2 The audio module 170 may transmit an audio signal to the wireless communication module 160 via the I2S interface.
[0216] The PCM interface may also be used for audio communication and may perform sampling, quantization, and encoding on analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 may be coupled via the PC M type interface. In some embodiments, the audio module 170 may also transmit an audio signal to the wireless communication module 160 via the PCM interface to implement the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface may be used for audio communication.
[0217] The UART interface is a universal serial data bus and is used for asynchronous communication. The bus may be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically configured to connect to the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement the Bluetooth function. In some embodiments, the audio module 170 may transmit an audio signal to the wireless communication module 160 via the UART interface to implement the function of playing music via a Bluetooth headset.
[0218] The MIPI interface may be configured to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. The MIPI interface may include a camera serial interface (CSI), a display serial interface (DSI), and the like. In some embodiments, the processor 110 communicates with the camera 193 via the CSI interface to implement the photographing function of the electronic device 100. The processor 110 communicates with the display 194 via the DSI interface to implement the display function of the electronic device 100.
[0219] The GPIO interface may be configured by software. The GPIO interface may be configured as a control signal or a data signal. In some embodiments, the GPIO interface may be configured to connect the processor 110 to the camera 193, the display 194, the wireless communication module 160, the audio module 170, the sensor module 180, and the like. The GPIO interface may be further configured as an I 2 C interface, an I 2 S interface, a UART interface, a MIPI interface, or the like.
[0220] The USB interface 130 is an interface compliant with the USB standard specification, and specifically may be a mini USB interface, a micro USB interface, a USB Type-C interface, or the like. The USB interface 130 may be configured to connect to a charger to charge the electronic device 100, and may also be configured to transmit data between the electronic device 100 and peripheral devices. It may also be configured to connect to a headset to play audio via the headset. The interface may also be configured to connect to another terminal device such as an AR device.
[0221] The interface connection relationship between the modules shown in this embodiment of the present application is merely an exemplary illustration and should not be construed as a limitation on the structure of the electronic device 100. In some other embodiments of the present application, the electronic device 100 may alternatively use an interface connection method different from the foregoing embodiments, or a combination of multiple interface connection methods.
[0222] The charging management module 140 is configured to receive a charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive a charging input from a wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive a wireless charging input via the wireless charging coil of the electronic device 100. The charging management module 140 can further supply power to the electronic device 100 via the power management module 141 while charging the battery 142.
[0223] The power management module 141 is configured to connect to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can be further configured to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage and impedance). In some other embodiments, the power management module 141 can also be provided within the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can be provided within the same device.
[0224] The wireless communication function of the electronic device 100 can be implemented via antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, baseband processor, etc.
[0225] Antenna 1 and antenna 2 are configured to transmit and receive electromagnetic wave signals. Each antenna within the electronic device 100 can be configured to cover one or more communication frequency bands. Different antennas can be further multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna in a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0226] The mobile communication module 150 can provide solutions applied to the electronic device 100 for wireless communication including 2G, 3G, 4G, 5G, etc. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, perform processing such as filtering and amplification on the received electromagnetic waves, and transmit the processed electromagnetic waves to the modem processor for demodulation. The mobile communication module 150 can further amplify the signal modulated by the modem processor and convert the signal into an electromagnetic wave by using antenna 1 for radiation. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be arranged in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be arranged in the same device as at least some of the modules of the processor 110.
[0227] The modem processor may include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into an intermediate-frequency signal and a high-frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. Then, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs an audio signal via an audio device (not limited to, for example, loudspeaker 170A and receiver 170B), or displays an image or video via the display 194. In some embodiments, the modem processor may be a separate device. In some other embodiments, the modem processor may be provided within the same device as the mobile communication module 150 or another functional module and is independent of the processor 110.
[0228] The wireless communication module 160 is a wireless communication solution applied to the electronic device 100, and can provide a wireless communication solution including wireless local area networks (WLAN) (for example, wireless fidelity (Wi-Fi) network), Bluetooth (registered trademark) (Bluetooth (registered trademark), BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC) technology, infrared (IR) technology, etc. The wireless communication module 160 can be one or more components integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 can further receive the signal to be transmitted from the processor 110, perform frequency modulation and amplification on the signal, and convert the signal into an electromagnetic wave for radiation via the antenna 2.
[0229] In some embodiments, in the electronic device 100, the antenna 1 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, whereby the electronic device 100 can communicate with a network and other devices by using wireless communication technologies. The wireless communication technologies may include the global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA (registered trademark)), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), Bluetooth (registered trademark), BT ) , GNSS, WLAN, NFC, FM, IR technology, and / or the like. GNSS may include the global positioning system (GPS), global navigation satellite system (GLONASS), BeiDou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0230] The electronic device 100 implements a display function by using a GPU, a display 194, an application processor, etc. The GPU is a microprocessor for image processing and is connected to the display 194 and the application processor. The GPU is configured to execute mathematical and geometric calculations and render images. The processor 110 may include one or more GPUs that execute program instructions for generating or changing display information.
[0231] The display 194 is configured to display images, videos, etc. The display 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini LED, a micro LED, a micro OLED, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0232] The electronic device 100 may implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display 194, an application processor, etc.
[0233] The ISP is configured to process the data fed back by camera 193. For example, during shooting, the shutter is opened and light rays are transmitted through the lens to the photosensitive element of the camera. The optical signal is converted into an electrical signal. The photosensitive element of the camera transmits the electrical signal to the ISP for processing and converts the electrical signal into a visible image. The ISP may further perform algorithm optimization for image noise, brightness, and skin color. The ISP can further optimize parameters such as exposure and color temperature in the shooting scenario. In some embodiments, the ISP may be disposed within camera 193.
[0234] Camera 193 is configured to capture still images or videos. The optical image of an object is generated through the lens and projected onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0235] The digital signal processor is configured to process digital signals and may also process another digital signal in addition to the digital image signal. For example, when electronic device 100 selects a frequency, the digital signal processor is configured to perform operations such as Fourier transform on the frequency energy.
[0236] The video codec is configured to compress or decompress digital video. The electronic device 100 may support one or more video codecs. In this way, the electronic device 100 can play or record video in multiple encoding formats, such as MPEG (Moving Picture Experts Group)-1, MPEG-2, MPEG-3, and MPEG-4.
[0237] The NPU is a neural-network (NN) computing processor that can simulate the biological neural network structure such as the transmission mode between neurons in the human brain, perform high-speed processing on input information, and perform continuous self-learning. The NPU can implement applications for the intelligent recognition of the electronic device 100, such as image recognition, face recognition, voice recognition, and text understanding.
[0238] The external memory interface 120 may be configured to connect to an external memory card, such as a micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 by using the external memory interface 120 to implement a data storage function. Files such as music and video are stored in the external storage card.
[0239] The internal memory 121 can be configured to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes the instructions stored in the internal memory 121 to perform various functional applications and data processing of the electronic device 100. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system, applications required for at least one function (such as a voice playback function and an image playback function), etc. The data storage area may store data created based on the use of the electronic device 100 (such as voice data and a phone book), etc. In addition, the internal memory 121 may include a high-speed random access memory and may further include a non-volatile memory, for example, at least one magnetic storage device, a flash memory device, and a universal flash storage (UFS).
[0240] The electronic device 100 may implement audio functions such as music playback and recording through an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, an application processor, etc.
[0241] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and is also configured to convert an analog audio input into a digital audio signal. The audio module 170 may be further configured to encode and decode audio signals. In some embodiments, the audio module 170 may be disposed within the processor 110, or some functional modules within the audio module 170 may be disposed within the processor 110.
[0242] Loudspeaker 170A, also called a "loudspeaker", is configured to convert an audio electrical signal into an audio signal. The electronic device 100 can listen to music or respond to calls in hands-free mode through the loudspeaker 170A.
[0243] Receiver 170B, also called an "earpiece", is configured to convert an audio electrical signal into an audio signal. When responding to a call or receiving a voice message via the electronic device 100, the receiver 170B can be brought close to a person's ear to listen to the voice.
[0244] Microphone 170C, also called a "mic", is configured to convert an audio signal into an electrical signal. When making a call or sending a voice message, the user can bring their mouth close to the microphone 170C to make a sound and input the audio signal into the microphone 170C. At least one microphone 170C can be disposed within the electronic device 100. In some other embodiments, in addition to collecting audio signals, two microphones 170C can be disposed within the electronic device 100 to implement a noise reduction function. In some other embodiments, alternatively, three, four, or more microphones 170C can be disposed within the electronic device 100 to collect audio signals, implement noise reduction, identify the sound source, and thereby implement functions such as a directional recording function.
[0245] The headset jack 170D is configured to connect to a wired headset. The headset jack 170D may be a USB interface 130, or may be a 3.5 mm open mobile terminal platform (OMTP) standard interface or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface.
[0246] The pressure sensor 180A is configured to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A may be disposed within the display 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. The capacitive pressure sensor may include at least two parallel plates having a conductive material. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is performed on the display 194, the electronic device 100 detects the intensity of the touch operation by using the pressure sensor 180A. The electronic device 100 may also calculate the touch position based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations that are performed on the same touch position but have different touch operation intensities may correspond to different operation commands. For example, when a touch operation having a touch operation intensity less than a first pressure threshold is performed on the short message application icon, a command to view the short message is executed. When a touch operation having a touch operation intensity greater than or equal to the first pressure threshold is performed on the short message application icon, a command to create a new short message is executed.
[0247] The gyro sensor 180B may be configured to determine the movement posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (x, y, and z axes) may be determined by using the gyro sensor 180B. The gyro sensor 180B may be used for shake correction. For example, when the shutter is pressed, the gyro sensor 180B detects the angle by which the electronic device 100 jitters, calculates the distance that the lens module needs to compensate based on this angle, and enables the lens to cancel out the shake of the electronic device 100 through reverse movement, thereby realizing image stabilization. The gyro sensor 180B may be further used in navigation scenarios and motion sensing game scenarios.
[0248] The barometric pressure sensor 180C is configured to measure barometric pressure. In some embodiments, the electronic device 100 calculates altitude by using the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0249] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can detect the opening and closing of a flip holster by using the magnetic sensor 180D. In some embodiments, when the electronic device 100 is a flip device, the electronic device 100 can detect the opening and closing of the flip by using the magnetic sensor 180D. Further, functions such as automatically unlocking when the flip cover is opened are set based on the detected opening and closing state of the leather case or the flip cover.
[0250] The acceleration sensor 180E can detect the acceleration of the electronic device 100 in various directions (usually three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. The acceleration sensor 180E may be further configured to identify the posture of the electronic device 100 and is applied to uses such as a pedometer and switching between landscape mode and portrait mode.
[0251] The distance sensor 180F is configured to measure distance. The electronic device 100 can measure distance by an infrared method or a laser method. In some embodiments, in a shooting scenario, the electronic device 100 uses the distance sensor 180F to measure distance and can achieve high-speed focusing.
[0252] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector such as a photodiode. The light-emitting diode may be an infrared light-emitting diode. The electronic device 100 emits infrared light externally by using the light-emitting diode. The electronic device 100 uses the photodiode to detect infrared reflected light from surrounding objects. When abundant reflected light is detected, it may be determined that an object is present in the vicinity of the electronic device 100. When the reflected light is not sufficiently detected, the electronic device 100 may determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity light sensor 180G to detect that the user is holding the electronic device 100 near the ear for a call, and automatically turn off the display to save power. The proximity light sensor 180G may also be used in a holster mode or a pocket mode for automatic unlocking and screen locking.
[0253] The ambient light sensor 180L is configured to sense the brightness of ambient light. The electronic device 100 may adaptively adjust the brightness of the display 194 based on the sensed brightness of the ambient light. The ambient light sensor 180L may be further configured to automatically adjust the white balance during photography. The ambient light sensor 180L may further cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is in a pocket and prevent accidental touches.
[0254] The fingerprint sensor 180H is configured to collect fingerprints. The electronic device 100 may use the characteristics of the collected fingerprints to perform fingerprint-based unlocking, application lock access, fingerprint-based photography, fingerprint-based incoming call response, etc.
[0255] The temperature sensor 180J is configured to detect temperature. In some embodiments, the electronic device 100 executes a temperature processing policy based on the temperature detected by the temperature sensor 180J. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device 100 reduces the performance of the processor located near the temperature sensor 180J to reduce power consumption and implement thermal protection. In some other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery 142 to avoid abnormal shutdown of the electronic device 100 caused by low temperature. In some other embodiments, when the temperature is lower than yet another threshold, the electronic device 100 boosts the output voltage of the battery 142 to avoid abnormal shutdown caused by low temperature.
[0256] The touch sensor 180K is also referred to as a "touch panel". The touch sensor 180K can be disposed on the display 194, and the touch sensor 180K and the display 194 form a touch screen, which is also referred to as a "touch screen". The touch sensor 180K is configured to detect a touch operation performed on or near the touch sensor. The touch sensor can transfer the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided on the display 194. In some other embodiments, the touch sensor 180K may alternatively be disposed on the surface of the electronic device 100 at a position different from the position of the display 194.
[0257] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the sound vibration bones of the human body. Also, the bone conduction sensor 180M can receive a blood pressure pulsation signal by contacting the pulse of the human body. In some embodiments, the bone conduction sensor 180M can be disposed on the headset so as to be integrated with the bone conduction headset. The audio module 170 can analyze an audio signal based on the vibration signal of the sound vibration bone acquired by the bone conduction sensor 180M to implement an audio function. The application processor can analyze heart rate information based on the blood pressure pulsation signal acquired by the bone conduction sensor 180M to implement a heart rate detection function.
[0258] The button 190 includes a power button, a volume button, etc. The button 190 may be a mechanical button or a touch sensor type button. The electronic device 100 can receive a button input and generate a button signal input related to user settings and function control of the electronic device 100.
[0259] The motor 191 can generate a vibration alert. The motor 191 can be used for a vibration alert for an incoming call and can also be used for touch vibration feedback. For example, touch operations on different applications (such as shooting and audio playback) can correspond to different vibration feedback effects. For touch operations on different regions of the display 194, the motor 191 can also correspondingly generate different vibration feedback effects. Different application scenarios (such as a time reminder, information reception, an alarm clock, and a game, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can be further customized.
[0260] The indicator 192 can be an indicator and can be configured to indicate a charging status and a power change, or can be configured to indicate a message, a missed call, a notification, etc.
[0261] The SIM card interface 195 is configured to connect to a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to effect contact or separation from the electronic device 100. The electronic device 100 may support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 may support a nano SIM card, a micro SIM card, a SIM card, etc. Multiple cards may be inserted into the same SIM card interface 195. The types of the multiple cards may be the same or different. The SIM card interface 195 may also be compatible with different types of SIM cards. The SIM card interface 195 may also be compatible with an external memory card. The electronic device 100 interacts with a network via the SIM card to implement functions such as calls and data communications. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card may be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0262] The electronic device 100 shown in FIG. 9 is merely an example, and it should be understood that the electronic device 100 may have more or fewer components than those shown in FIG. 9, may combine two or more components, or may have a different component configuration. The various components shown in FIG. 9 may be implemented by using hardware, software, or a combination of hardware and software, including one or more signal processors and / or application-specific integrated circuits.
[0263] Hereinafter, the software structure of the electronic device 100 according to an embodiment of the present application will be described.
[0264] FIG. 10 shows an example of the software structure of the electronic device 100 according to an embodiment of the present application.
[0265] As shown in FIG. 10, the software system of the electronic device 100 may use a hierarchical architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. Hereinafter, an example will be used to describe the software structure of the electronic device 100.
[0266] The software is divided by using a hierarchical architecture Layer and each layer has a clear role and task. These layers communicate with each other by using software interface connections. In some embodiments, the software structure of the electronic device 100 is divided into three layers from top to bottom: an application layer, an application framework layer, and a kernel layer.
[0267] The application layer may include a series of application packages.
[0268] As shown in FIG. 10, the application packages may include applications such as a camera, a gallery, a calendar, a phone, a map, a navigation, a WLAN, Bluetooth (registered trademark), music, a video, and a message. The video may be the video application mentioned in the embodiments of the present application.
[0269] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0270] As shown in FIG. 10, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, a video processing system, etc.
[0271] The window manager is configured to manage window programs. The window manager may obtain a display size, determine whether there is a status bar, lock the screen, take a screenshot, and the like.
[0272] The content provider is configured to store and obtain data so that applications can access these data. The data may include videos, images, audio, outgoing and incoming calls, browsing history and bookmarks, address books, and the like.
[0273] The view system includes visual controls such as controls for displaying text and controls for displaying pictures. The view system may be configured to build an application. The display interface may include one or more views. For example, a display interface including a message notification icon may include a view for displaying text and a view for displaying a picture.
[0274] The phone manager is configured to provide management of the communication functions of the electronic device 100, for example, the call status (including answering, rejecting, etc.).
[0275] The resource manager provides resources for applications such as localized strings, icons, images, layout files, and video files.
[0276] The notification manager can be configured to enable an application to display notification information in the status bar and convey notification messages. The notification manager can automatically disappear after a short pause without requiring user interaction. For example, the notification manager can be configured to notify of download completion, provide message reminders, etc. Alternatively, the notification manager can be a notification that appears in the status bar at the top of the system in the form of a graph or scroll bar text, such as a notification for an application running in the background or a notification that appears on the screen in the form of a dialog window. For example, character information is displayed in the status bar, an announcement is made, the electronic device vibrates, or the indicator light blinks.
[0277] The video processing system can be configured to execute the caption display method provided in the embodiments of the present application. The video processing system can include a caption decoding module, a video frame color gamut interpretation module, a video frame composition module, a video frame queue, and a video rendering module. For specific functions of each module, refer to the relevant content in the foregoing embodiments. Details are not described herein again.
[0278] The kernel layer is a layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, a Bluetooth (registered trademark) driver, and a sensor driver.
[0279] Hereinafter, with reference to a shooting scenario, an example of the operation process of the software and hardware of the electronic device 100 will be described.
[0280] When the touch sensor 180K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into the original input event (including information such as the touch coordinates and timestamp of the touch operation). The original input event is stored in the kernel layer. The application framework layer obtains the original input event from the kernel layer and identifies the control corresponding to the input event. For example, the touch operation is a single tap operation, and the control corresponding to the single tap operation is the control of the camera application icon. The camera application calls an interface in the application framework layer to start the camera application. Then, it calls the kernel layer to start the camera driver and uses the camera 193 to capture a still image or a video.
[0281] Hereinafter, the structure of another electronic device 100 according to an embodiment of the present application will be described.
[0282] FIG. 11 shows an example of the configuration of another electronic device 100 according to an embodiment of the present application.
[0283] As shown in FIG. 11, the electronic device 100 may include a video application 1100 and a video processing system 1110.
[0284] The video application 1100 may be a system application installed in the electronic device 100 (for example, the "Video" application shown in FIG. 2A), or an application provided by a third party, installed in the electronic device 100, and having a video playback function. The video application 1100 is mainly configured to play videos.
[0285] The video processing system 1110 may include a video decoding module 1111, a caption decoding module 1112, a video frame color gamut interpretation module 1113, a video frame composition module 1114, a video frame queue 1115, and a video rendering module 1116.
[0286] The video decoding module 1111 may receive a video information stream transmitted by the video application 1100 and decode the video information stream to generate video frames.
[0287] The caption decoding module 1112 may receive a caption information stream transmitted by the video application 1100, decode the caption information stream to generate caption frames, and transmit caption frames with a mask based on the mask parameters transmitted by the video frame color gamut interpretation module 1113 to enhance caption recognition.
[0288] The video frame color gamut interpretation module 1113 may analyze caption recognition to generate a caption recognition analysis result and calculate mask parameters (mask color values and transparency) corresponding to the caption based on the caption recognition analysis result.
[0289] The video frame composition module 1114 may overlay and combine a video frame and a caption frame to generate a video frame to be displayed.
[0290] The video frame queue 1115 may store the video frames to be displayed transmitted by the video frame composition module 1114.
[0291] The video rendering module 1116 may render the video frames to be displayed based on the time series to generate rendered video frames and transmit the rendered video frames to the video application 1100 for video playback.
[0292] For further details regarding the functions and operating principles of the electronic device 100, refer to the relevant content in the foregoing embodiments. Details will not be described again here.
[0293] The electronic device 100 shown in FIG. 11 is merely an example, and it should be understood that the electronic device 100 may have more or fewer components than those shown in FIG. 11, may combine two or more components, or may have different component configurations. The various components shown in FIG. 11 may be implemented in hardware, software, or a combination of hardware and software.
[0294] The foregoing modules may be divided by function. In an actual product, the modules may be different functions executed by the same software module.
[0295] Hereinafter, the structure of another electronic device 100 according to an embodiment of the present application will be described.
[0296] FIG. 12 shows an example of the configuration of another electronic device 100 according to an embodiment of the present application.
[0297] As shown in FIG. 12, the electronic device 100 may include a video application 1200. The video application 1200 may include a video decoding module 1211, a caption decoding module 1212, a video frame color gamut interpretation module 1213, a video frame synthesis module 1214, a video frame queue 1215, and a video rendering module 1216.
[0298] The video application 1200 can be a system application installed on the electronic device 100 (e.g., the "Video" application shown in FIG. 2A), or it can be an application provided by a third party, installed on the electronic device 100, and having a video playback function. The video application 1200 is mainly configured to play videos.
[0299] The acquisition and display module 1210 can acquire a video information stream and a caption information stream, and display the rendered video frames transmitted by, for example, the video rendering module 1216.
[0300] The video decoding module 1211 can receive the video information stream transmitted by the acquisition and display module 1210, decode the video information stream, and generate video frames.
[0301] The caption decoding module 1212 can receive the caption information stream transmitted by the acquisition and display module 1210, decode the caption information stream to generate caption frames, and generate caption frames with masks based on the mask parameters transmitted by the video frame color gamut interpretation module 1213 to enhance the caption recognition degree.
[0302] The video frame color gamut interpretation module 1213 can analyze the caption recognition degree to generate a caption recognition degree analysis result, and calculate mask parameters (the color value and transparency of the mask) corresponding to the caption based on the caption recognition degree analysis result.
[0303] The video frame composition module 1214 can overlay and combine the video frames and the caption frames to generate the video frames to be displayed.
[0304] The video frame queue 1215 may store the video frames to be displayed transmitted by the video frame synthesis module 1214.
[0305] The video rendering module 1216 may render the video frames to be displayed based on the time series, generate the rendered video frames, and transmit the rendered video frames to the acquisition and display module 1210 for video playback.
[0306] For more details regarding the functions and operating principles of the electronic device 100, refer to the relevant content in the foregoing embodiments. Details will not be described again here.
[0307] It should be understood that the electronic device 100 shown in FIG. 12 is only an example, and the electronic device 100 may have more or fewer components than those shown in FIG. 12, may combine two or more components, or may have different component configurations. The various components shown in FIG. 12 may be implemented in hardware, software, or a combination of hardware and software.
[0308] The foregoing embodiments are only intended to illustrate the technical solutions of this application and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions of the embodiments of this application Range of Without departing from the scope, the technical solutions described in the foregoing embodiments may be further modified, or equivalent substitutions may be made for some of their technical features.
Claims
Claim 1 A caption display method, comprising: a step of an electronic device playing a first video ; when the electronic device displays a first interface, the first interface includes a first picture and a first caption, and the first caption is displayed on a first area of the first picture in a floating manner by using a first mask as a background to overlay the first caption on the first picture. The first area is an area within the first picture and corresponding to the display position of the first caption. A difference value between a color value of the first caption and a color value of the first area is a first value; when the electronic device displays a second interface different from the first interface after displaying the first interface, the second interface includes a second picture and the first caption, no mask is displayed for the first caption, and the first caption is displayed on a second area of the second picture in a floating manner by overlaying the first caption on the second picture. The second area is an area within the second picture and corresponding to the display position of the first caption. A difference value between the color value of the first caption and a color value of the second area is a second value, and the second value is greater than the first value; the first picture is one picture in the first video, and the second picture is another picture in the first video; a difference value between a color value of the first mask and the color value of the first caption is greater than the first value; a method. Claim 2 Before the electronic device displays the first interface, the method further comprises: a step of the electronic device obtaining a first video file and a first caption file, wherein the first video file and the first caption file carry the same time information; a step of the electronic device generating a first video frame based on the first video file, wherein the first video frame is used to generate the first picture; The step in which the electronic device generates a first caption frame based on the first caption file and obtains the color value and the display position of the first caption from the first caption frame, wherein the time information carried in the first caption frame is the same as the time information carried in the first video frame. The step in which the electronic device determines the first region based on the display position of the first caption. The step in which the electronic device generates the first mask based on the color value of the first caption or the color value of the first region. The step in which the electronic device overlays the first caption on the first mask in the first caption frame to generate a second caption frame, and combines the second caption frame and the first video frame. The method according to claim 1, further comprising the above steps.
3. Before the step in which the electronic device generates the first mask based on the color value of the first caption or the color value of the first region, the method further comprises: The step in which the electronic device determines that the first value is less than a first threshold. The method according to claim 2, further comprising the above step.
4. The step in which the electronic device determines that the first value is less than a first threshold is specifically: The step in which the electronic device divides the first region into N first sub-regions, where N is a positive integer. The step in which the electronic device determines that the first value is less than the first threshold based on the color value of the first caption and the color values of the N first sub-regions. The method according to claim 3, comprising the above steps.
5. The step in which the electronic device generates the first mask based on the color value of the first caption or the color value of the first region is specifically: The step in which the electronic device determines the color value of the first mask based on the color value of the first caption or the color values of the N first sub-regions. The step in which the electronic device generates the first mask based on the color value of the first mask. The method according to claim 4, comprising the above steps.
6. The step in which the electronic device determines that the first value is less than a first threshold value specifically includes: The step in which the electronic device divides the first region into N first sub-regions, where N is a positive integer; The step in which the electronic device determines whether to combine adjacent first sub-regions into second sub-regions based on a difference value between color values of adjacent first sub-regions; The step in which the electronic device combines adjacent first sub-regions into second sub-regions when the difference value between the color values of the adjacent first sub-regions is less than a second threshold value; The step in which the electronic device determines that the first value is less than the first threshold value based on the color value of the first caption and the color value of the second sub-region; The method according to claim 3, comprising the above steps. Claim 7 The method according to claim 6, wherein the first region includes M second sub-regions, M is a positive integer and M is less than or equal to N, each second sub-region includes one or more first sub-regions, and the number of first sub-regions included in each second sub-region is the same as or different from the number of first sub-regions included in another second sub-region. Claim 8 The step in which the electronic device generates the first mask based on the color value of the first caption or the color value of the first region specifically includes: The step in which the electronic device sequentially calculates color values of M first sub-masks, which are masks corresponding to the M second sub-regions, based on the color value of the first caption or the color values of the M second sub-regions; The step in which the electronic device generates the M first sub-masks based on the color values of the M first sub-masks, and the M first sub-masks are combined into the first mask; The method according to claim 7, comprising the above steps. Claim 9 The method further includes: When the electronic device displays a third interface, the third interface includes a third picture and the first caption, the first caption includes at least a first part and a second part, a second sub-mask is displayed for the first part, and a third sub-mask is displayed for the second part or the third sub-mask is not displayed, and a color value of the second sub-mask is different from a color value of the third sub-mask. The method according to claim 1, further comprising.
10. The method according to claim 1, wherein a display position of the first mask is determined based on the display position of the first caption.
11. In the first picture and the second picture, the display position of the first caption with respect to a display screen of the electronic device is not fixed or is fixed, and the first caption is a segment of continuously displayed characters or symbols. The method according to claim 1.
12. Before the electronic device displays the first interface, the method is The step of the electronic device setting the transparency of the first mask to less than 100% The method according to claim 1, further comprising.
13. Before the electronic device displays the second interface, the method is The step of the electronic device generating a second mask based on the color value of the first caption or the color value of the second region and superimposing the first caption on the second mask, wherein the color value of the second mask is a preset color value and the transparency of the second mask is 100%, or or The step of the electronic device skipping the generation of the second mask The method according to claim 1, further comprising.
14. An electronic device comprising one or more processors and one or more memories, the one or more memories being coupled to the one or more processors, the one or more memories being configured to store computer program code, the computer program code including computer instructions, and when the one or more processors execute the computer instructions, the electronic device is capable of executing the method according to any one of claims 1 to 13.
15. A computer program, the computer program including program instructions, wherein when the program instructions are executed on an electronic device, the electronic device is capable of executing the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Video image processing method and device and electronic equipment
CN112511890A
Display control apparatus and program for display control processing
JP2005159955A
Image signal output device and image signal output method
JP2006295746A
Optical disk device
JP2009027605A
Image processing device, electronic camera, and program
JP2012074812A