Video display method, device, equipment and storage medium
By dynamically adjusting the font, position and color of subtitles in the video according to the character's emotions, the problem that users in the prior art is difficult to perceive the character's emotions, and the ability and fun of videos to convey information is improved.
Patent Information
- Application Number
- CN202310094081.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-03
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2043-02-03
AI Technical Summary
In the prior art, the fixed display method of video subtitles makes it difficult for users to directly perceive the emotions of characters, and the video has insufficient ability to convey information.
By determining the emotional information of the characters in the video, select the appropriate subtitle fonts based on the emotions type and intensity level, and dynamically adjust the position and color of the subtitles in the video display area to reflect the emotional changes of the characters.
It improves the fun and effectiveness of the video's information, enables users to feel the emotions of the characters more intuitively, and enhances the viewing experience.
Smart Images

Figure CN116055792B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video, and in particular to a video display method, apparatus, device and storage medium. Background Art
[0002] Currently, when playing a video, the subtitles, or the words spoken by the characters in the video, are fixedly displayed in a preset area and a single, preset font. This fixed subtitle display method results in poor viewing experience for users. Characters' emotions vary, and users need to carefully read and understand the subtitles to understand the content of the video. This means that users' perception of the characters' emotions is not direct, indicating that existing technologies currently lack the ability to convey information.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of the present invention is to provide a video display method, aiming to solve the problem in the prior art that the current video has insufficient ability to transmit information.
[0005] To achieve the above objectives, the present application provides a video display method, applied to a video display device, the method comprising:
[0006] Determine the emotional information of the corresponding character in the video;
[0007] Determining a target font for subtitles to be displayed based on the character's emotional information;
[0008] The subtitle to be displayed is displayed in the target font.
[0009] In a possible implementation of the present application, the character's emotional information includes an emotion type, and the step of determining a target font for displaying subtitles based on the character's emotional information includes:
[0010] The target font of the subtitles to be displayed is determined according to the emotion type of the character.
[0011] In a possible implementation of the present application, the character's emotional information includes an emotion intensity level, the emotion types include voice emotion types and facial expression emotion types, the emotion intensity level includes a voice emotion intensity level and a facial expression emotion intensity level, the voice emotion intensity level is determined based on the character's voice in the audio information, and the facial expression emotion intensity level is determined based on the character's facial expression distortion. The step of determining the target font for the subtitles to be displayed based on the character's emotion type includes:
[0012] Determining a target font type for the subtitles to be displayed based on the voice emotion type of the character and / or the facial expression emotion type of the character;
[0013] If the target font type of the subtitles to be displayed is determined based on the voice emotion type of the character, then the target font of the subtitles to be displayed is determined from a candidate font set corresponding to the target font type based on the voice emotion intensity level;
[0014] If the target font type of the subtitles to be displayed is determined based on the type of the facial expression of the character, then the target font of the subtitles to be displayed is determined from a set of candidate fonts corresponding to the target font type based on the intensity level of the facial expression;
[0015] If the target font type of the subtitles to be displayed is determined based on the voice emotion type and the facial expression emotion type of the character, then the target font of the subtitles to be displayed is determined from the candidate font set corresponding to the target font type based on the voice emotion intensity level and the facial expression emotion intensity level.
[0016] In a possible implementation of the present application, the step of determining the target font of the subtitles to be displayed according to the emotion type of the character further includes:
[0017] Determine whether the voice emotion type and the facial expression emotion type belong to the same emotion. If so, determine the target font type of the subtitles to be displayed based on the voice emotion type of the character and the facial expression emotion type of the character; if not, determine the target font type of the subtitles to be displayed based on one of the voice emotion type of the character and the facial expression emotion type of the character.
[0018] In a possible implementation of the present application, the step of displaying the subtitles to be displayed in the target font includes:
[0019] Obtaining a blank ratio between the subtitles to be displayed and a preset edge of the video display area;
[0020] Determining a target display position of the subtitle to be displayed in the video display area based on the blank ratio;
[0021] The subtitle to be displayed is displayed at the target display position in the target font.
[0022] In a possible implementation of the present application, the step of displaying the subtitles to be displayed in the target font includes:
[0023] Determining the outline coordinates of the mouth outline of the character in the current video frame;
[0024] determining an offset of the subtitle to be displayed relative to the mouth contour according to the emotion intensity level;
[0025] determining a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset;
[0026] The subtitle to be displayed is displayed at the target display position in the target font.
[0027] In a possible implementation manner of the present application, the step of determining a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset includes:
[0028] Determining the pixel quantity occupied by a single subtitle based on the number of characters corresponding to the single subtitle and the pixel quantity occupied by each character;
[0029] Determining a direction in which the mouth outline is offset in the video display area based on the outline coordinates and the screen center coordinates of the video display area;
[0030] The target display position of the to-be-displayed subtitle in the video display area is determined based on the pixel quantity occupied by the single subtitle, the offset, and the direction.
[0031] In addition, to achieve the above-mentioned purpose, the present application also provides a video display device, which includes:
[0032] A first determination module is used to determine the emotional information of the corresponding character in the video;
[0033] A second determining module is used to determine a target font for displaying subtitles based on the emotional information of the character;
[0034] The first display module is configured to display the subtitles to be displayed in the target font.
[0035] In addition, to achieve the above-mentioned purpose, the present application also provides a video display device, which is a physical node device, and the video display device includes: a memory, a processor, and a video display program stored on the memory and runnable on the processor, and the processor executes the video display program to implement the steps of the video display method.
[0036] In addition, to achieve the above-mentioned purpose, the present application also provides a storage medium, on which a program for implementing the video display method is stored. When the video display program is executed by a processor, the steps of the video display method described above are implemented.
[0037] This application provides a video display method, apparatus, device, and storage medium. Compared to the existing technology, which currently suffers from insufficient video information transmission capabilities, this application determines the emotional information of a corresponding character in a video; determines a target font for subtitles to be displayed based on the character's emotional information; and displays the subtitles in the target font. In this application, the subtitle font is determined based on the character's emotion. Different fonts allow users to intuitively sense the character's emotions, increasing the interest of video transmission and improving the current video's ability to transmit information. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a flow chart of an embodiment of the video display method of the present application;
[0039] Figure 2 A partial schematic diagram of a video display screen in an embodiment of the video display method of the present application;
[0040] Figure 3 A partial schematic diagram of a video display screen in an embodiment of the video display method of the present application;
[0041] Figure 4 Schematic diagram of a video display device in an embodiment of the video display method of the present application;
[0042] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the video display method of this application. DETAILED DESCRIPTION
[0043] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0044] Example 1
[0045] The embodiment of the present application provides a video display method. In the first embodiment of the video display method of the present application, referring to Figure 1 , applied to a video display device, the method comprising:
[0046] Step S10, determining the emotional information of the corresponding character in the video;
[0047] Step S20, determining a target font for subtitles to be displayed based on the character's emotional information;
[0048] Step S30: displaying the subtitles to be displayed in the target font.
[0049] This embodiment targets an application scenario where, during video playback, the subtitles (the words spoken by the characters in the video) are fixedly displayed in a preset area and a single, preset font. This fixed subtitle display method results in a poor viewing experience for users. Characters exhibit diverse emotions, requiring users to carefully read and understand the subtitles before they can comprehend the content of the video. This results in an indirect perception of the characters' emotions, implying a problem with existing video technologies, namely the inadequate ability of videos to convey information.
[0050] This embodiment aims to improve the problem that current videos are insufficient in transmitting information.
[0051] In this embodiment, the video display method is applied to a video display device.
[0052] In this embodiment, the current video frame may be a video frame in a preset video file to be played, or a video frame in a real-time live broadcast, which is not specifically limited here.
[0053] As an example, the corresponding character in the current video frame can be a real person or a virtual character, and no specific limitation is made here.
[0054] In this embodiment, the current character may have a variety of emotions, such as joy, anger, sadness, happiness, tension, or surprise.
[0055] In this embodiment, the subtitles to be displayed are text information to be displayed in the video frame.
[0056] In this embodiment, the target display position is the specific position of the subtitle to be displayed in the video display area.
[0057] In this embodiment, the target font is the font actually displayed in the video display area of the subtitles to be displayed. Different target fonts will result in different shapes of each text in the subtitles.
[0058] In this embodiment, the target font of the subtitles to be displayed is determined based on the emotional information of the character. The display font of the subtitles is different depending on the character's emotions. For example, if the character's emotion is happy, the target display font is the preset font No. 1; if the character's emotion is angry, the target display font is the preset font No. 2.
[0059] In this embodiment, after the target display position and target font of the subtitles to be displayed are determined, the subtitles to be displayed are displayed in the target font.
[0060] The specific steps are as follows:
[0061] Step S10, determining the emotional information of the corresponding character in the video;
[0062] As an example, when a user is watching a live video, a character in the live video is giving a narration, and the character is in high spirits during the narration, and appears very happy, and even laughs during the narration. The video display device determines that the character in the current live video is in a happy mood.
[0063] Step S20, determining a target font for subtitles to be displayed based on the character's emotional information;
[0064] As an example, the video display device determines the target font for the content explained by the character (ie, subtitles to be displayed) according to the emotions of the character in the live video.
[0065] The character's emotional information includes the type of emotion. Step S20, determining a target font for subtitles to be displayed based on the character's emotional information, includes step S21:
[0066] Step S21 : determining the target font of the subtitles to be displayed according to the emotion type of the character.
[0067] As an example, the font of the subtitle display can be changed instead of being fixed, and the target font of the subtitle display is related to the type of emotion of the character.
[0068] As an example, when the emotion type is happy, the target font is font size 1. Figure 2 (1) is a schematic diagram of a portion of the video display screen. The character in the video is in a happy mood, and the subtitles in the video are displayed in font size 1.
[0069] As an example, in addition to the type of emotion influencing the target font, the intensity of the emotion also influences the target font. Different emotion types correspond to different target font types, and there are significant differences between the target fonts of different target font types. While the same emotion type corresponds to the same target font type, there are subtle differences between the target fonts corresponding to different intensity levels of the same emotion type.
[0070] As an example, the character's emotions include an emotion intensity level and an emotion type. For example, the character's emotion type is joy, and the degree to which the character expresses joy varies. Therefore, the emotion intensity level is used to measure the character's current emotion level.
[0071] As an example, the emotion intensity level corresponding to each emotion type is graded into 10 levels, from level 1 to level 10, where level 1 indicates the lowest emotion intensity and level 10 indicates the highest emotion intensity. Specifically, for example, "joy 10" indicates that the emotion type is joy, and the corresponding emotion intensity level is level 10.
[0072] As an example, the character's emotional information includes an emotion intensity level, the emotion types include voice emotion types and expression emotion types, the emotion intensity level includes voice emotion intensity level and expression emotion intensity level, the voice emotion intensity level is determined based on the character's voice in the audio information, and the expression emotion intensity level is determined based on the distortion of the character's expression.
[0073] Step S21, determining the target font of the subtitles to be displayed according to the emotion type of the character, includes steps A1 to A4:
[0074] Step A1, determining a target font type for the subtitles to be displayed based on the voice emotion type of the character and / or the facial expression type of the character;
[0075] As an example, the emotion category can be determined not only from the character's voice, ie, the voice emotion category, but also from the character's facial expression, ie, the expression emotion category.
[0076] As an example, based on the character's voice in the audio information, the character's voice emotion type is determined.
[0077] As an example, an image of a character in a video frame is obtained and the face is recognized, and the type of the character's facial expression is determined from the character's facial image.
[0078] As an example, the target font type of the subtitles to be displayed is determined to be type A based on the voice emotion type of the character.
[0079] Step A2: if the target font type of the subtitles to be displayed is determined based on the type of the voice emotion of the character, then determining the target font of the subtitles to be displayed from a set of candidate fonts corresponding to the target font type based on the intensity level of the voice emotion;
[0080] As an example, if the target font type of the subtitles to be displayed is determined to be Class A based on the voice emotion type of the character, then based on the voice emotion intensity level being level 5, the target font is determined to be A5 font from the candidate font set corresponding to Class A fonts, such as Figure 2 (2) is a partial schematic diagram of the video display screen. Figure 2 (2) The character's emotion is joy, and the intensity level of joy is level 5. The subtitles in the video are displayed in the joy 5 font. Figure 2 (1) The character's emotion is also happy, but the intensity level of the emotion is level 1, that is, Figure 2 (2) The font size of subtitles is different from that of Figure 2(1) The font of the subtitles is more exaggerated. The characters are both happy, but the degree of their emotions is different. There is still a difference in the font display effect. The subtitles can directly express the emotions of the characters in the video, which optimizes the user's viewing experience.
[0081] Step A3: if the target font type of the subtitles to be displayed is determined based on the type of the facial expression of the character, then the target font of the subtitles to be displayed is determined from a set of candidate fonts corresponding to the target font type based on the intensity level of the facial expression;
[0082] As an example, if the target font type of the subtitles to be displayed is determined to be type B based on the type of facial expression of the character, then based on the intensity level of the facial expression being level 4, the target font is determined to be type B4 from the candidate font set corresponding to type B fonts.
[0083] Step A4: If the target font type of the subtitles to be displayed is determined based on the voice emotion type and the facial expression emotion type of the character, then the target font of the subtitles to be displayed is determined from the candidate font set corresponding to the target font type based on the voice emotion intensity level and the facial expression emotion intensity level.
[0084] As an example, the target font type may be determined jointly based on the voice emotion type and the facial expression emotion type, and then the final target font may be determined from the target font types based on the voice emotion intensity level and the facial expression emotion intensity level.
[0085] As an example, if sound and expression are assigned respective weights, the target font type is determined to be Class A based on the sound emotion type, the weight of the sound emotion type is 0.6, the expression emotion type and the weight of the expression emotion type is 0.4, the sound emotion intensity level, the expression emotion intensity level and their respective weights are statistically valued, and the target font a8 is selected from the candidate font set of Class A fonts based on the statistical values.
[0086] As an example, a method for judging the emotional intensity level of a character from the sound may be to comprehensively judge the character's voice and the background sound in the video frame to obtain the corresponding emotional intensity level of the sound.
[0087] As an example, the vocal emotion type and facial expression type may differ. For example, if the vocal emotion type is identified as joy from the voice, but the facial expression type is identified as sadness from the image, joy and sadness are clearly different. Therefore, when determining the target font based on both the vocal emotion type and the facial expression type, it is necessary to consider whether the vocal emotion type and the facial expression type are the same.
[0088] As an example, the voice emotion type is compared with the facial emotion type to determine whether the voice emotion type is the same as or different from the facial emotion type.
[0089] Step S21, determining the target font of the subtitles to be displayed according to the emotion type of the character, further includes:
[0090] Determine whether the voice emotion type and the facial expression emotion type belong to the same emotion. If so, determine the target font type of the subtitles to be displayed based on the voice emotion type of the character and the facial expression emotion type of the character; if not, determine the target font type of the subtitles to be displayed based on one of the voice emotion type of the character and the facial expression emotion type of the character.
[0091] As an example, if the voice emotion type and the facial expression type belong to the same emotion, the target font type of the subtitles to be displayed is determined based on the voice emotion type and the facial expression type of the character.
[0092] As an example, after determining the target font type based on the voice emotion type and the facial expression emotion type, the target font of the subtitles to be displayed is determined from the candidate font set corresponding to the target font type based on the voice emotion intensity level and the facial expression emotion intensity level. For example, the voice emotion intensity level and the facial expression emotion intensity level are averaged to obtain a comprehensive emotion intensity level, and the target font is determined from the candidate font set corresponding to the target font type based on the comprehensive emotion intensity level.
[0093] As an example, if the voice emotion type and the facial emotion type do not belong to the same emotion, then when determining the target font type, one of the voice emotion type and the facial emotion type may be selected to determine the target font type.
[0094] As an example, when selecting either a vocal emotion category or an emoticon category to determine the target font type, the target font type for the subtitles to be displayed is determined based on the vocal emotion category, with emoticon information not being referenced when determining the font type based on the vocal information. Because audio processing is faster than image processing, prioritizing audio processing speeds up video processing, enhancing the fun of the video display while also increasing the speed of video display. When displaying subtitles based on character emotions, there's no video freeze, optimizing the user experience while watching the video.
[0095] In this embodiment, a step is added to determine the target font from a set of candidate fonts corresponding to the target font type based on the emotion intensity level. Furthermore, a step is added to compare the voice emotion type with the facial expression emotion type to determine whether they are the same. First, the basic target font type is determined based on whether the voice emotion type and the facial expression emotion type are the same. Then, the target font is determined from the set of candidate fonts corresponding to the target font type based on the emotion intensity level. Furthermore, the subtitles displayed based on the character's emotions are changed, enriching the content of the video display and enhancing the interest of the video display.
[0096] Step S30: displaying the subtitles to be displayed in the target font.
[0097] As an example, the target display position of the subtitles to be displayed in the video display area may be fixed or non-fixed.
[0098] As an example, it can be manually set to a fixed mode or a non-fixed mode, and can also be displayed in a fixed or non-fixed position in a default mode without manual setting.
[0099] Step S30, displaying the subtitles to be displayed in the target font, includes steps S31 to S33:
[0100] Step S31, obtaining the blank ratio between the subtitles to be displayed and the edge of a preset video display area;
[0101] As an example, subtitles are displayed at a fixed position. For example, subtitles are displayed at the bottom of a video display area. For example, in existing videos, subtitles are generally displayed in the middle and at the bottom of the video display area. A blanking ratio between the subtitles to be displayed and the edge of the video display area is obtained. The blanking ratio is the blanking ratio between the subtitles to be displayed and the bottom edge of the video display area.
[0102] As an example, the margin ratio m is the ratio of the width of the margin between the subtitles to be displayed and the bottom edge of the video display area to the height of the video display area.
[0103] Step S32, determining a target display position of the subtitle to be displayed in the video display area based on the blank ratio;
[0104] As an example, the size of the video display area is preset, and there may be a variety of different video display areas.
[0105] As an example, the display width of the video display area is VW, and the display height is VH.
[0106] As an example, the pixel width occupied by a single character is FW, and the pixel width and height occupied by a single character is FH. FW and FH are related to the intensity of emotion. The greater the intensity of emotion, the greater the FW and FH.
[0107] As an example, the number of characters corresponding to a single subtitle is FC, then the pixel width occupied by the single subtitle is SUBW=FW*FC, and the pixel height occupied by the single subtitle SUBH is equal to FH.
[0108] As an example, a coordinate system is established with the upper left corner of the video display area as the coordinate origin. The target display position of the subtitles to be displayed in the video display area specifically includes a display horizontal coordinate SUBX and a display vertical coordinate SUBY. The display horizontal coordinate SUBX = (VW - SUBW) / 2, and the display vertical coordinate SUBY = VH - (FH + (VH*m)).
[0109] Step S33: displaying the subtitles to be displayed at the target display position using the target font.
[0110] As an example, the subtitles to be displayed are displayed at the display abscissa SUBX and the display ordinate SUBY in the target font.
[0111] As an example, subtitles are displayed at a non-fixed position. For example, based on the intensity of the character's emotions, the target display position of the character's explanation, i.e., the subtitles to be displayed, in the video display area is determined.
[0112] Step S30, displaying the subtitles to be displayed in the target font, includes steps S34 to S37:
[0113] Step S34, determining the outline coordinates of the mouth outline of the character in the current video frame;
[0114] As an example, the display position of the subtitles changes with the character's mouth, and the subtitles are no longer limited to the bottom of the video display area, that is, when it is a non-fixed position display mode, the contour coordinates of the character's mouth contour in the current video frame are determined.
[0115] As an example, there are countless coordinates for the outline of a character's mouth. Four of these coordinates are selected. These four points are the leftmost, rightmost, topmost, and bottommost endpoints of the mouth outline in the coordinate system corresponding to the video display area. The selected coordinates are the horizontal coordinate value MX1 of the leftmost endpoint, the horizontal coordinate value MX2 of the rightmost endpoint, the vertical coordinate value MY1 of the topmost endpoint, and the vertical coordinate value MY2 of the bottommost endpoint.
[0116] Step S35, determining an offset of the subtitle to be displayed relative to the mouth contour according to the emotion intensity level;
[0117] As an example, the offset of the subtitles to be displayed relative to the mouth contour specifically includes the offset in horizontal pixels and the offset in vertical pixels.
[0118] The higher the level of emotion intensity, the smaller the offset of the subtitles to be displayed relative to the mouth contour. That is, the more intense the character's emotions are, the closer the subtitles are displayed to the character to intuitively express the character's emotions. Users can intuitively feel the emotions of the character in the video from the distance.
[0119] As an example, if the current emotional intensity level is level 3, the horizontal pixel offset of the subtitles to be displayed relative to the mouth contour is (10-3)*DEVX, and the vertical pixel offset is (10-3)*DEVY.
[0120] As can be seen, the higher the emotional intensity, the fewer horizontal and vertical pixels are offset, meaning the closer the subtitles are to the mouth. DEVX is MW / (VW / MW), and DEVY is MH / (VH / MH).
[0121] Step S36, determining a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset;
[0122] As an example, based on the outline coordinates and the offset, the target display position of the subtitle to be displayed in the video display area is determined to be an outline horizontal coordinate offset of (10-3)*DEVX pixels and an outline vertical coordinate offset of (10-3)*DEVY pixels.
[0123] Step S37: displaying the subtitles to be displayed at the target display position using the target font.
[0124] As an example, Figure 3 The figure below shows a partial diagram of the video display screen. The subtitles, or the words spoken by the character in the video, are displayed near the character's mouth, rather than being fixed at the bottom of the video. When the character moves, or when the character's mouth moves, the subtitles follow the character's mouth. The subtitles appear in a different position each time.
[0125] This application provides a video display method, apparatus, device, and storage medium. Compared to the existing technology, which currently suffers from insufficient video information transmission capabilities, this application determines the emotional information of a corresponding character in a video; determines a target font for subtitles to be displayed based on the character's emotional information; and displays the subtitles in the target font. In this application, the subtitle font is determined based on the character's emotion. Different fonts allow users to intuitively sense the character's emotions, increasing the interest of video transmission and improving the current video's ability to transmit information.
[0126] Example 2
[0127] Furthermore, based on the first embodiment of the present application, another embodiment of the present application is provided. In this embodiment, step S36, the step of determining the target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset, includes steps C1 to C3:
[0128] Step C1, determining the pixel quantity occupied by a single subtitle based on the number of characters corresponding to the single subtitle and the pixel quantity occupied by each character;
[0129] As an example, the pixel width occupied by a single character is FW, and the pixel width height occupied by a single character is FH. The number of characters corresponding to a single subtitle is FC, then the pixel width occupied by a single subtitle is SUBW=FW*FC, and the pixel height occupied by a single subtitle is SUBH equal to FH.
[0130] Step C2, determining the direction in which the mouth outline is offset in the video display area based on the outline coordinates and the screen center coordinates of the video display area;
[0131] As an example, the screen center coordinates of the video display area are (CPX, CPY), where CPX=VW / 2 and CPY=VH / 2.
[0132] As an example, if MX1 is greater than or equal to CPX, that is, the character is in the middle right area of the video display area, and the direction in which the mouth outline is offset in the video display area is to the right. If MX1 is less than CPX, that is, the character is in the middle left area of the video display area, and the direction in which the mouth outline is offset in the video display area is to the left.
[0133] Step C3: determining the target display position of the to-be-displayed subtitle in the video display area based on the pixel quantity occupied by the single subtitle, the offset, and the direction.
[0134] As an example, the target display position of the to-be-displayed subtitle in the video display area is determined based on the occupied pixel amounts SUBW and SUBH of the single subtitle, the direction, and the offsets DEVX and DEVY.
[0135] Specifically, when MX1 is greater than or equal to CPX, SUBX = MX1 + DEVX, that is, when the direction in which the mouth outline is offset in the video display area is to the right, subtitles are displayed on the right side of the mouth outline. When MX1 is less than CPX, SUBX = MX1 - SUBW - DEVX, that is, when the direction in which the mouth outline is offset in the video display area is to the left, subtitles are displayed on the left side of the mouth outline.
[0136] In this embodiment, when it is determined that the mouth outline is offset to the right in the video display area, subtitles are displayed to the right of the mouth outline. When it is determined that the mouth outline is offset to the left in the video display area, subtitles are displayed to the left of the mouth outline. That is, the subtitle display area (the blank area without the character) is determined based on the direction in which the mouth outline is offset in the video display area. This prevents subtitles from obscuring the character when displayed on the character, thereby affecting the user's viewing experience. This further improves the current video's ability to convey information.
[0137] Example 3
[0138] Furthermore, based on all the above embodiments of the present application, another embodiment of the present application is provided. In this embodiment, step S30, the step of displaying the subtitles to be displayed in the target font, includes steps D1 and D2:
[0139] Step D1, determining a target color of the target font based on the emotional information of the character;
[0140] In this embodiment, the color of the subtitles can also be changed, and can be variable, rather than a single color such as black.
[0141] As an example, when the character's emotions are joy, anger, sadness, and happiness, the target colors of the target font are pink, red, white, and green, respectively.
[0142] Step D2: displaying the subtitle to be displayed at the target display position in the target font and the target color.
[0143] As an example, the subtitle to be displayed is displayed at the target display position in the target font and the target color.
[0144] In this embodiment, the color of the corresponding font is selected according to the character's emotion, so that the color of the font changes with the character's emotion, enriching the diversity of the font and enhancing the fun, that is, further improving the ability of the current video to convey information.
[0145] Example 4
[0146] Furthermore, based on all the above embodiments, another embodiment of the present application is provided. In this embodiment, as Figure 2 , provides a video display device, the device comprising:
[0147] A first determination module is used to determine the emotional information of the corresponding character in the video;
[0148] A second determining module is used to determine a target font for displaying subtitles based on the emotional information of the character;
[0149] The first display module is configured to display the subtitles to be displayed in the target font.
[0150] In a possible implementation of the present application, the character's emotional information includes an emotion type, and the step of determining a target font for displaying subtitles based on the character's emotional information includes:
[0151] The third determining module is configured to determine the target font of the subtitles to be displayed according to the emotion type of the character.
[0152] In a possible implementation of the present application, the character's emotional information includes an emotion intensity level, the emotion types include voice emotion types and facial expression emotion types, the emotion intensity levels include voice emotion intensity levels and facial expression emotion intensity levels, the voice emotion intensity level is determined based on the character's voice in the audio information, and the facial expression intensity level is determined based on the character's facial expression distortion. The step of determining the target font for the subtitles to be displayed based on the character's emotion type, the apparatus includes:
[0153] a fourth determining module, configured to determine a target font type for the subtitles to be displayed based on the voice emotion type of the character and / or the facial expression type of the character;
[0154] a fifth determining module configured to, if the target font type of the subtitles to be displayed is determined based on the type of the voice emotion of the character, determine the target font of the subtitles to be displayed from a set of candidate fonts corresponding to the target font type based on the intensity level of the voice emotion;
[0155] a sixth determining module, configured to, if the target font type of the subtitles to be displayed is determined based on the type of the facial expression of the character, determine the target font of the subtitles to be displayed from a set of candidate fonts corresponding to the target font type based on the intensity level of the facial expression;
[0156] The seventh determination module is used to determine the target font type of the subtitles to be displayed based on the voice emotion type and the facial expression emotion type of the character, and then determine the target font of the subtitles to be displayed from the candidate font set corresponding to the target font type based on the voice emotion intensity level and the facial expression emotion intensity level.
[0157] In a possible implementation of the present application, in the step of determining the target font of the subtitles to be displayed according to the emotion type of the character, the apparatus further includes:
[0158] The judgment module is used to judge whether the voice emotion type and the facial expression emotion type belong to the same emotion. If so, the target font type of the subtitles to be displayed is determined based on the voice emotion type of the character and the facial expression emotion type of the character; if not, the target font type of the subtitles to be displayed is determined based on one of the voice emotion type of the character and the facial expression emotion type of the character.
[0159] In a possible implementation manner of the present application, the step of displaying the subtitles to be displayed in the target font includes:
[0160] A first acquisition module is used to obtain a blank ratio between the subtitles to be displayed and the edge of a preset video display area;
[0161] An eighth determining module, configured to determine a target display position of the subtitle to be displayed in the video display area based on the blank ratio;
[0162] The second display module is configured to display the subtitles to be displayed at the target display position using the target font.
[0163] In a possible implementation manner of the present application, the step of displaying the subtitles to be displayed in the target font includes:
[0164] a ninth determining module, configured to determine the contour coordinates of the mouth contour of the character in the current video frame;
[0165] a tenth determining module, configured to determine an offset of the subtitle to be displayed relative to the mouth contour according to the emotion intensity level;
[0166] an eleventh determining module, configured to determine a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset;
[0167] A third display module is configured to display the subtitles to be displayed at the target display position using the target font.
[0168] A twelfth determining module is used in a possible implementation manner of the present application, in the step of determining a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset, the device comprising:
[0169] a thirteenth determining module, configured to determine the pixel quantity occupied by a single subtitle based on the number of characters corresponding to the single subtitle and the pixel quantity occupied by each character;
[0170] A fourteenth determining module is configured to determine a direction in which the mouth outline is offset in the video display area based on the outline coordinates and the screen center coordinates of the video display area;
[0171] A fifteenth determining module is configured to determine the target display position of the to-be-displayed subtitle in the video display area based on the number of pixels occupied by the single subtitle, the offset, and the direction.
[0172] The specific implementation of the video display device of the present application is basically the same as the embodiments of the above-mentioned video display method, and will not be repeated here.
[0173] Example 5
[0174] Furthermore, based on all the above embodiments, another embodiment of the present application is provided. In this embodiment, a video display device is provided, which is a physical node device. The video display device includes: a memory, a processor, and a program for implementing the video display method stored in the memory, the memory is used to store the program for implementing the video display method; the processor is used to execute the program for implementing the video display method to implement the steps of the video display method in the above embodiments.
[0175] Reference Figure 3 , Figure 3 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present application.
[0176] like Figure 3As shown, the video display device may include: a processor 1001, such as a CPU, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to implement communication between the processor 1001 and the memory 1005. The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as a disk storage device. The memory 1005 may also be a storage device independent of the processor 1001.
[0177] In a possible implementation of the present application, the video display device may also include a network interface, an audio circuit, a display, a connecting line, a sensor, an input module, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface, a Bluetooth interface), and the input module may optionally include a keyboard, a system soft keyboard, voice input, wireless receiving input, etc.
[0178] Those skilled in the art will appreciate that the structure of the video display device does not limit the video display device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0179] The memory, a computer storage medium, can include an operating system, an information exchange module, and a video display program. The operating system manages and controls the hardware and software resources of the video display device, supporting the operation of the video display program and other software and / or programs. The information exchange module facilitates communication between the various components within the memory, as well as with other hardware and software in the management system.
[0180] In the video display device, the processor is used to execute the video display program stored in the memory to implement the above-mentioned video display steps.
[0181] The specific implementation of the video display device of the present application is basically the same as the embodiments of the above-mentioned video display method, and will not be repeated here.
[0182] Example 6
[0183] An embodiment of the present application provides a storage medium, and the storage medium stores one or more programs, and the one or more programs can also be executed by one or more processors to implement the steps of the video display method in the above embodiment.
[0184] The specific implementation of the storage medium of the present application is basically the same as the embodiments of the above-mentioned video display method, and will not be repeated here.
[0185] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0186] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0187] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM or RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0188] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A video display method, characterized in that: The video display method comprises: Determining emotional information of a corresponding character in the video, wherein the emotional information of the character includes an emotion type and an emotion intensity level, and the emotion type includes a voice emotion type and an expression emotion type; Determining a target font for subtitles to be displayed based on the character's emotion type; The step of determining the target font of the subtitles to be displayed according to the emotion type of the character further includes: Determining whether the voice emotion type and the facial expression emotion type belong to the same emotion; If not, determining a target font type for the subtitles to be displayed based on the voice emotion type of the character; Displaying the subtitles to be displayed in the target font; The step of displaying the subtitles to be displayed in the target font includes: Determining the outline coordinates of the mouth outline of the character in the current video frame; determining an offset of the subtitles to be displayed relative to the mouth contour according to the emotion intensity level, wherein the higher the emotion intensity level, the closer the subtitles to be displayed are to the mouth contour; determining a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset; The subtitle to be displayed is displayed at the target display position in the target font.
2. The video display method according to claim 1, wherein: The emotion intensity level includes a voice emotion intensity level and an expression emotion intensity level, the voice emotion intensity level is determined based on the character's voice in the audio information, and the expression emotion intensity level is determined based on the character's facial expression distortion. The step of determining a target font for displaying subtitles based on the character's emotion type includes: Determining a target font type for the subtitles to be displayed based on the voice emotion type of the character and / or the facial expression emotion type of the character; If the target font type of the subtitles to be displayed is determined based on the voice emotion type of the character, then the target font of the subtitles to be displayed is determined from a candidate font set corresponding to the target font type based on the voice emotion intensity level; If the target font type of the subtitles to be displayed is determined based on the type of the facial expression of the character, then the target font of the subtitles to be displayed is determined from a set of candidate fonts corresponding to the target font type based on the intensity level of the facial expression; If the target font type of the subtitles to be displayed is determined based on the voice emotion type and the facial expression emotion type of the character, then the target font of the subtitles to be displayed is determined from the candidate font set corresponding to the target font type based on the voice emotion intensity level and the facial expression emotion intensity level.
3. The video display method according to claim 2, wherein: After the step of determining whether the voice emotion type and the facial expression emotion type belong to the same emotion, the method further includes: If so, the target font type of the subtitles to be displayed is determined based on the voice emotion type of the character and the facial emotion type of the character.
4. The video display method according to claim 1, wherein: The step of displaying the subtitles to be displayed in the target font includes: Obtaining a blank ratio between the subtitles to be displayed and a preset edge of the video display area; Determining a target display position of the subtitle to be displayed in the video display area based on the blank ratio; The subtitle to be displayed is displayed at the target display position in the target font.
5. The video display method according to claim 4, characterized in that: The step of determining a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset includes: Determining the pixel quantity occupied by a single subtitle based on the number of characters corresponding to the single subtitle and the pixel quantity occupied by each character; Determining a direction in which the mouth outline is offset in the video display area based on the outline coordinates and the screen center coordinates of the video display area; The target display position of the to-be-displayed subtitle in the video display area is determined based on the pixel quantity occupied by the single subtitle, the offset, and the direction.
6. A video display device, characterized in that: A video display device comprising: A first determination module is configured to determine emotional information of a corresponding character in a video, wherein the emotional information of the character includes an emotional type and an emotional intensity level, and the emotional type includes a vocal emotional type and an facial emotional type; A second determining module is used to determine a target font for subtitles to be displayed according to the emotion type of the character; a judgment module, configured to judge whether the voice emotion type and the facial expression emotion type belong to the same emotion; if not, determining a target font type for the subtitles to be displayed based on the voice emotion type of the character; A first display module, configured to display the subtitles to be displayed in the target font; a ninth determining module, configured to determine the contour coordinates of the mouth contour of the character in the current video frame; a tenth determining module, configured to determine an offset of the subtitles to be displayed relative to the mouth contour according to an emotion intensity level, wherein the higher the emotion intensity level, the closer the subtitles to be displayed are to the mouth contour; an eleventh determining module, configured to determine a target display position of the subtitle to be displayed in the video display area based on the outline coordinates and the offset; A third display module is configured to display the subtitles to be displayed at the target display position using the target font.
7. A video display device, characterized in that: The method comprises a memory, a processor and a video display program stored in the memory and executable on the processor, wherein the processor executes the video display program to implement the steps of the video display method according to any one of claims 1 to 5.
8. A storage medium, characterized in that: The storage medium stores a program for implementing the video display method, and the program for implementing the video display method is executed by a processor to implement the steps of the video display method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Information transmission method, receiving terminal device and sending terminal device
CN108521369A
Information display method and device and electronic equipment
CN113794927A
Subtitle generation system using graphic object
WO2020091431A1