Subtitle image generation method and apparatus, electronic device and storage medium

By dividing the subtitle image into layers and transforming the image, rich subtitle animation effects are generated, which solves the problem of the monotonous effects of traditional subtitle animation and improves the rendering efficiency and real-time performance of subtitle animation.

WO2026061166A1PCT designated stage Publication Date: 2026-03-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Traditional subtitle animation effects are monotonous and lack diversity, resulting in a poor viewing experience.

Method used

By dividing the subtitle image into multiple image layers, obtaining initial and target attribute values ​​based on different types of display content, performing text layout drawing and image transformation, and generating rich subtitle animation effects.

Benefits of technology

It enables fine-grained control over subtitle images, enriches subtitle animation effects, and improves the rendering efficiency and real-time performance of subtitle animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025113461_26032026_PF_FP_ABST
    Figure CN2025113461_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present disclosure are a subtitle image generation method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring initial attribute values of image levels in a subtitle image corresponding to a subtitle to be displayed, wherein the image levels are obtained by division on the basis of different types of display content in the subtitle image; acquiring a subtitle text of said subtitle, and for the image levels, performing text layout and rendering on the basis of the subtitle text and the initial attribute values to obtain initial subtitle layers of the image levels at an initial time point; acquiring target attribute values of the image levels, and for each image level, performing image transformation on the initial subtitle layer of the image level on the basis of the corresponding target attribute value to obtain a target subtitle layer of the image level at a target time point; and overlaying the plurality of target subtitle layers to generate a target subtitle image of said subtitle at the target time point. The embodiments of the present disclosure can enrich the animation effect of subtitles to be displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Subtitle image generation method and device, electronic equipment and storage medium

[0001] The present application claims priority to the Chinese patent application No. 2024113052722, filed on September 19, 2024, and entitled "Subtitle image generation method and device, electronic equipment and storage medium", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present disclosure relates to the technical field of computer, and particularly relates to a subtitle image generation method and device, electronic equipment and storage medium. BACKGROUND

[0003] In audio and video playing, subtitles can provide text information corresponding to audio and video content, helping viewers better understand the plot and dialogue in the audio and video content, and carrying animation subtitles can further highlight key information and enhance the atmosphere of the audio and video content, which plays a very important role in improving the viewing experience. At present, all display contents in traditional subtitle animation present the same animation effect, resulting in relatively single animation effect of the subtitle animation. SUMMARY

[0004] The following is a summary of the subject matter of the detailed description of the present disclosure. This summary is not intended to limit the protection scope of the claims.

[0005] The embodiments of the present disclosure provide a subtitle image generation method and device, electronic equipment and storage medium, which can enrich the animation effect of the to-be-displayed subtitles.

[0006] In one aspect, the embodiments of the present disclosure provide a subtitle image generation method, which is executed by an electronic device and includes:

[0007] Obtaining initial attribute values of each image level in a subtitle image corresponding to a to-be-displayed subtitle, wherein the image levels are divided based on different types of display contents in the subtitle image, and the initial attribute values are attribute values of level attributes of the image levels at an initial time point;

[0008] Obtaining a subtitle text of the to-be-displayed subtitle, and for each image level, performing text layout and drawing based on the subtitle text and the initial attribute values of each image level to obtain an initial subtitle layer of each image level at the initial time point;

[0009] obtaining a target attribute value of each of the image levels, for each of the image levels, performing image transformation on the initial subtitle layer of the image level based on the target attribute value of the image level, to obtain a target subtitle layer of the image level at a target time point, wherein the target attribute value is an attribute value of a level attribute of the image level at the target time point, and the target time point is a time point after the initial time point;

[0010] superimposing the target subtitle layer of each of the image levels to generate a target subtitle image of the to-be-displayed subtitle at the target time point.

[0011] In another aspect, the embodiments of the present disclosure further provide a subtitle image generation device, comprising:

[0012] an obtaining module, configured to obtain an initial attribute value of each of image levels in a subtitle image corresponding to a to-be-displayed subtitle, wherein the image levels are divided based on different types of display content in the subtitle image, and the initial attribute value is an attribute value of a level attribute of the image level at an initial time point;

[0013] a layout and drawing module, configured to obtain a subtitle text of the to-be-displayed subtitle, and for each of the image levels, perform text layout and drawing based on the subtitle text and the initial attribute value of each of the image levels, to obtain an initial subtitle layer of each of the image levels at the initial time point;

[0014] an image transformation module, configured to obtain a target attribute value of each of the image levels, for each of the image levels, perform image transformation on the initial subtitle layer of the image level based on the target attribute value of the image level, to obtain a target subtitle layer of the image level at a target time point, wherein the target attribute value is an attribute value of a level attribute of the image level at the target time point, and the target time point is a time point after the initial time point;

[0015] an image generation module, configured to superimpose the target subtitle layer of each of the image levels to generate a target subtitle image of the to-be-displayed subtitle at the target time point.

[0016] In another aspect, the embodiments of the present disclosure further provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned subtitle image generation method when executing the computer program.

[0017] In another aspect, the embodiments of the present disclosure further provide a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned subtitle image generation method.

[0018] In another aspect, the present disclosure also provides a computer program product comprising a computer program stored in a computer readable storage medium. A processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program to cause the computer device to perform the above-described subtitle image generation method.

[0019] The present disclosure has at least the following beneficial effects: obtaining initial attribute values of hierarchical attributes of each image hierarchy at an initial time point, and obtaining subtitle text of a to-be-displayed subtitle. Since the image hierarchy is based on different types of display content in a subtitle image corresponding to the to-be-displayed subtitle, text layout and drawing can be performed based on the subtitle text and the initial attribute values to obtain an initial subtitle layer of the image hierarchy at the initial time point, thereby achieving division of the subtitle layer for different display content. Then, based on target attribute values of the hierarchical attributes at a target time point, image transformation is performed on the initial subtitle layer to obtain a target subtitle layer of the image hierarchy at the target time point, and then the multiple target subtitle layers are superimposed to generate a target subtitle image of the to-be-displayed subtitle at the target time point. In the image transformation, each target attribute value controls the image transformation process of the corresponding initial subtitle layer, thereby achieving fine control of each image hierarchy, so as to finely control different types of display content in the subtitle image corresponding to the to-be-displayed subtitle, and enrich the animation effect of the to-be-displayed subtitle. In addition, for multiple frame images of the to-be-displayed subtitle, the initial subtitle layer can be regarded as a layer in an initial frame image displayed by the to-be-displayed subtitle at the initial time point, and the target subtitle layer can be regarded as a layer in a target frame image displayed by the to-be-displayed subtitle at the target time point. In the subtitle animation rendering process, only text layout and drawing are needed when generating the initial subtitle layer, and no text layout and drawing are needed when generating the target subtitle layer, but the target subtitle layer is obtained by image transformation based on the initial subtitle layer. This can avoid frequent text layout and drawing with large computational complexity, thereby improving the rendering efficiency of the subtitle animation and enhancing the real-time performance and smoothness of the subtitle animation.

[0020] Other features and advantages of the present disclosure will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings are included to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification, and are used together with the embodiments of the present disclosure to explain the technical solutions of the present disclosure, and do not constitute a limitation on the technical solutions of the present disclosure.

[0022] FIG. 1 is a schematic diagram of an optional implementation environment provided by the present disclosure;

[0023] FIG. 2 is an optional flow diagram of a method for generating a subtitle image according to an embodiment of the present disclosure;

[0024] FIG. 3 is an optional flow diagram of a first method for determining a target subtitle layer according to an embodiment of the present disclosure;

[0025] FIG. 4 is an optional flow diagram of a second method for determining a target subtitle layer according to an embodiment of the present disclosure;

[0026] FIG. 5 is an optional flow diagram of a third method for determining a target subtitle layer according to an embodiment of the present disclosure;

[0027] FIG. 6 is an optional flow diagram of a fourth method for determining a target subtitle layer according to an embodiment of the present disclosure;

[0028] FIG. 7 is an optional flow diagram of a fifth method for determining a target subtitle layer according to an embodiment of the present disclosure;

[0029] FIG. 8 is an optional flow diagram of a method for performing a mask transform according to an embodiment of the present disclosure;

[0030] FIG. 9 is an optional grouping diagram of a set of dynamic properties according to an embodiment of the present disclosure;

[0031] FIG. 10 is an optional flow diagram of a method for determining a target property value according to an embodiment of the present disclosure;

[0032] FIG. 11 is an optional architecture diagram of a method for generating a subtitle image according to an embodiment of the present disclosure;

[0033] FIG. 12 is an optional structure diagram of a subtitle image generation apparatus according to an embodiment of the present disclosure;

[0034] FIG. 13 is a partial structure block diagram of a terminal according to an embodiment of the present disclosure;

[0035] FIG. 14 is a partial structure block diagram of a server according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0036] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, further detailed description will be given to the present disclosure in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.

[0037] It should be noted that in various specific embodiments of the present disclosure, when relevant processing needs to be performed on data related to the characteristics of the target object, such as target object attribute information or attribute information set, permission or consent of the target object is obtained first, and the collection, use and processing of the data comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when the embodiments of the present disclosure need to obtain target object attribute information, the separate permission or separate consent of the target object is obtained through a pop-up window or by jumping to a confirmation page, and after obtaining the separate permission or separate consent of the target object, the necessary target object related data for enabling the embodiments of the present disclosure to normally run is obtained.

[0038] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.

[0039] In order to facilitate understanding of the technical solutions provided by the embodiments of the present disclosure, some key terms used by the embodiments of the present disclosure are explained first:

[0040] Subtitles: used to display non-image content such as dialogues in television, film, stage works in the form of text, and also refers to the text processed in the later stage of video and film works. The dialogue subtitles of video and film works generally appear at the bottom of the screen, while the subtitles of drama works may be displayed on both sides or above the stage. Its function is to display the voice content of the program in the form of subtitles, which can help hearing-impaired audiences understand the program content. In addition, subtitles can also be used to translate foreign language programs, so that audiences who do not understand the language can not only hear the original voice, but also understand the program content. Therefore, in audio and video playback, subtitles can provide text information corresponding to the audio and video content, helping the audience better understand the plot and dialogue of the film.

[0041] At present, all display contents in traditional subtitle animation will present the same animation effect, resulting in a single animation effect of the subtitle animation.

[0042] Based on this, the embodiments of the present disclosure provide a subtitle image generation method and device, electronic equipment and storage medium, which can enrich the animation effect of the to-be-displayed subtitles.

[0043] Referring to FIG. 1, FIG. 1 is a schematic diagram of an optional implementation environment provided by an embodiment of the present disclosure, which includes a terminal 101 and a server 102, wherein the terminal 101 and the server 102 are connected through a communication network.

[0044] Exemplarily, the server 102 can acquire subtitle information input by the terminal 101, wherein the subtitle information is information related to a to-be-displayed subtitle, and can include information related to a subtitle attribute, a subtitle text, a subtitle display time point, and the like. An initial attribute value of each image level in a subtitle image corresponding to the to-be-displayed subtitle is acquired from the subtitle information, wherein the image level is divided based on different types of display content in the subtitle image, and the initial attribute value is an attribute value of a level attribute of the image level at an initial time point. A subtitle text of the to-be-displayed subtitle is acquired from the subtitle information, and for each image level, text layout and drawing are performed based on the subtitle text and the initial attribute value of each image level, to obtain an initial subtitle layer of each image level at the initial time point. A target attribute value of each image level is acquired, and for each image level, image transformation is performed on the initial subtitle layer of the image level based on the target attribute value of the image level, to obtain a target subtitle layer of the image level at a target time point, wherein the target attribute value is an attribute value of the level attribute of the image level at the target time point, and the target time point is a time point after the initial time point. The target subtitle layer of each image level is superimposed, to generate a target subtitle image of the to-be-displayed subtitle at the target time point. The server 102 sends the target subtitle image to the terminal 101.

[0045] The server 102 obtains initial attribute values of the hierarchical attributes of the respective image hierarchies at an initial time point, and obtains a subtitle text of a to-be-displayed subtitle. Since the image hierarchies are based on different types of display content in a subtitle image corresponding to the to-be-displayed subtitle, the initial subtitle layers of the respective image hierarchies at the initial time point can be obtained based on the subtitle text and the initial attribute values, so as to realize the division of the subtitle layers for different display content. Then, the initial subtitle layers are image-transformed based on target attribute values of the hierarchical attributes of the image hierarchies at a target time point, to obtain target subtitle layers of the image hierarchies at the target time point, and then the target subtitle layers are superimposed to generate a target subtitle image of the to-be-displayed subtitle at the target time point. In the image transformation, the image transformation process of the corresponding initial subtitle layers is respectively controlled by the respective target attribute values, so as to realize the fine control of the respective image hierarchies, thereby being capable of fine control of different types of display content in the subtitle image of the to-be-displayed subtitle, and enriching the animation effect of the to-be-displayed subtitle. In addition, for multiple frame images of the to-be-displayed subtitle, the initial subtitle layers can be regarded as layers in an initial frame image displayed by the to-be-displayed subtitle at the initial time point, and the target subtitle layers can be regarded as layers in a target frame image displayed by the to-be-displayed subtitle at the target time point. In the subtitle animation rendering process, only the text layout and drawing are needed when the initial subtitle layers are generated, and the text layout and drawing are not needed when the target subtitle layers are generated, but the target subtitle layers are obtained based on the initial subtitle layers through image transformation, so as to avoid frequent text layout and drawing with large amount of calculation, thereby improving the rendering efficiency of the subtitle animation, and enhancing the real-time performance and smoothness of the subtitle animation.

[0046] The server 102 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. In addition, the server 102 can also be a node server in a blockchain network.

[0047] The terminal 101 can be a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a smart wearable device, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be connected directly or indirectly through wired or wireless communication, and the present disclosure is not limited thereto.

[0048] Referring to FIG. 2, FIG. 2 is an optional flowchart of a subtitle image generation method provided by the embodiment of the present disclosure, which can be executed by an electronic device, specifically, a server, or a terminal, or a server in cooperation with a terminal. The subtitle image generation method includes, but is not limited to, the following steps 201 to 204.

[0049] Step 201: Obtain initial attribute values of each image level in a subtitle image corresponding to a to-be-displayed subtitle.

[0050] The to-be-displayed subtitle refers to a subtitle that needs to be displayed on a video, which can be obtained from a movie, an animation, a photographic work, etc. The embodiment of the present disclosure does not limit the source of the video. Since the subtitle is usually displayed for a period of time, the subtitle animation of the to-be-displayed subtitle can be displayed in multiple video frames. Displaying the to-be-displayed subtitle in the video is equivalent to rendering the corresponding subtitle image in multiple video frames, and each subtitle image is a frame image of the subtitle animation.

[0051] The image level is based on different types of display content in the subtitle image corresponding to the to-be-displayed subtitle. The subtitle image can include multiple types of display content, such as text, a border, a shadow, and a background. The text of the subtitle image refers to the textual content displayed on the video screen, which can include dialogue, a voiceover, or other information. The border of the subtitle image refers to a line around the text, which is used to increase the readability of the text. The shadow of the subtitle image refers to a shadow effect added below or around the text, which is used to increase the three-dimensionality of the subtitle. The background of the subtitle image refers to a color block or a pattern behind the text, which is used to enhance the readability of the subtitle. Different types of display content in the subtitle image corresponding to the to-be-displayed subtitle are divided into different image levels. For example, the background of the subtitle image is divided into a first image level, the shadow of the subtitle image is divided into a second image level, the border of the subtitle image is divided into a third image level, and the text of the subtitle image is divided into a fourth image level.

[0052] It should be noted that the subtitle image can have multiple level attributes. The level attribute refers to an attribute used to describe the characteristics of the subtitle image. The level attribute can be added to the subtitle to obtain a corresponding subtitle animation. Since the subtitle animation represents a process in which some level attributes change over time, that is, the level attribute of the subtitle has at least one dynamic attribute, the level attribute can be divided into a static attribute and a dynamic attribute according to whether the animation effect can be realized. The static attribute refers to an attribute that cannot realize the animation effect, and the dynamic attribute refers to an attribute used to realize the animation effect.

[0053] For example, the static attributes can include a font name attribute, a font size attribute, a bold attribute, an italic attribute, an underline attribute, a strikeout attribute, and the like; and the dynamic attributes can include a text color attribute, a character spacing attribute, a border width attribute, a border color attribute, a shadow distance attribute, a shadow color attribute, a background color attribute, a blur attribute, a rotation attribute, a scale attribute, a skew attribute, a position attribute, a color change ratio attribute, and a mask coordinate attribute, and the like.

[0054] The attribute value type of the font name attribute is a string, the attribute value types of the font size attribute, the character spacing attribute, the border width attribute, the shadow distance attribute, the blur attribute, the rotation attribute, the scale attribute, the skew attribute, the position attribute, and the color change ratio attribute are all floating-point numbers, the attribute value types of the text color attribute, the border color attribute, the shadow color attribute, and the background color attribute are all color values, the attribute value types of the bold attribute, the italic attribute, the underline attribute, and the strikeout attribute are all Boolean values, and the attribute value type of the mask coordinate attribute is a coordinate.

[0055] Specifically, the hierarchical attributes can be represented by multi-tuples. Taking the dynamic attributes as an example, the representation formula of the hierarchical attributes in the subtitle image is as follows: p i = <v i,0 ,v i,1 ,t i,0 ,t i,1 ,c i (t) | i ∈ N>

[0056] where p i represents the i-th hierarchical attribute, v i,0 represents an initial attribute value of the i-th hierarchical attribute when the attribute starts to change, v i,1 represents a final attribute value of the i-th hierarchical attribute when the attribute ends to change, t i,0 represents a change start time point of the i-th hierarchical attribute when the attribute starts to change, t i,1 represents a change end time point of the i-th hierarchical attribute when the attribute ends to change, c i (t) represents a transition curve expression of the change of the i-th hierarchical attribute, the transition curve expression is equivalent to an attribute change function, and N represents a set of identifiers of the hierarchical attributes. v i,0 , v i,1 , t i,0 , t i,1 , and c i (t) can all be obtained from the subtitle information input by the relevant personnel.

[0057] It is worth noting that, for the dynamic attributes, a basic animation effect corresponds to one dynamic attribute. For example, the y-axis scale factor of the subtitle changes from 1 to 1.5 in a linear transition manner within 4 to 6 seconds, which can be represented as p fscy= <1, 1.5, 4, 6, 0.25t>, at the display start time point t of the subtitle s to the change end time point t i,0 , the attribute value of the hierarchical attribute remains the attribute initial value, between the change end time point t i,1 to the display end time point t of the subtitle e , the attribute value of the hierarchical attribute remains the attribute final value, the display start time point t s and the display end time point t e belong to the display time point of the subtitle, the display start time point t s and the display end time point t e can be obtained from the subtitle information input by the relevant personnel.

[0058] In addition, for the static attribute, the static attribute can be regarded as a special case of the dynamic attribute, therefore, the static attribute can adopt the same representation formula as the dynamic attribute, and the representation formula of the static attribute generally satisfies the following conditions: i (t) = v i,0 = v i,1 , t i,0 = t s and t i,1 = t e , t s refers to the display start time point of the subtitle, t e refers to the display end time point of the subtitle, that is, the display start time point is the change start time point, the display end time point is the change end time point, the attribute initial value and the attribute final value are the same, which is equivalent to the attribute value of the static attribute remaining unchanged during the display of the subtitle.

[0059] wherein the initial attribute value is the attribute value of the hierarchical attribute of the image level at the initial time point, for any kind of hierarchical attribute, when the representation formula of the hierarchical attribute is not an empty set, it means that the subtitle image possesses the hierarchical attribute, the kind of the hierarchical attribute possessed by the subtitle image can be one or more, and the kind of the hierarchical attribute of each image level can be the same as the kind of the hierarchical attribute possessed by the subtitle image. Since different image levels are used to represent different types of subtitle content, there can be differences between the attribute values of part of the hierarchical attributes of different image levels. Specifically, the number of initial attribute values is the same as the number of kinds of hierarchical attributes of the image level, therefore, the attribute value of each kind of hierarchical attribute of the image level at the initial time point needs to be obtained. Assuming that there are 10 kinds of hierarchical attributes of the image level, the attribute values of the 10 kinds of hierarchical attributes at the initial time point need to be obtained.

[0060] Specifically, when the first frame image of the plurality of subtitle images of the to-be-displayed subtitle is selected as the initial frame image, the initial time point is the playing time point corresponding to the first frame image, that is, the initial time point is the display start time point of the subtitle. When other frame images of the to-be-displayed subtitle are selected as the initial frame image, the initial time point is the playing time point corresponding to the selected frame image. Therefore, the initial time point is specifically the playing time point of the initial frame image, and when the initial frame image is the first frame image, the v i,0 determined as the initial attribute value.

[0061] Based on this, since different image levels correspond to different types of display content in the subtitle image corresponding to the to-be-displayed subtitle, the initial attribute value of each image level is obtained, which is equivalent to obtaining the initial attribute value of different types of display content in the subtitle image. This realizes separate processing of different types of display content in the subtitle image, and subsequent accurate and personalized adjustment of the subtitle effect of different types of display content in the subtitle image, thereby improving the quality and richness of the subtitle animation effect corresponding to the to-be-displayed subtitle.

[0062] Step 202: Obtain the subtitle text of the to-be-displayed subtitle. For each image level, perform text layout and drawing based on the subtitle text and the initial attribute value of each image level, to obtain the initial subtitle layer of each image level at the initial time point.

[0063] The subtitle text refers to the text content displayed on the video picture. The subtitle text can be obtained from the subtitle information input by relevant personnel. For the image level where the text of the subtitle is located, since the type corresponding to the image level is the text type, in the initial subtitle layer corresponding to the image level, the pixel points at the position of the text content are non-transparent pixel points, and the pixel points at other positions are transparent pixel points, which can ensure that the initial subtitle layer can display the text content. The transparent pixel point refers to a pixel point with a transparent color value, and the transparent pixel point is invisible. For other image levels, since the type corresponding to the image level is not the text type, in the initial subtitle layer corresponding to the image level, the pixel points at the position of the text content are transparent pixel points, which can ensure that the initial subtitle layer corresponding to the image level will not display the text content, but only display the display content of the type corresponding to the image level.

[0064] It should be noted that different subtitle layers usually contain different display content. By superimposing a plurality of subtitle layers, the display content of each subtitle layer can be combined and displayed in the same picture. Therefore, superimposing the initial subtitle layers of each image level can obtain the initial frame image displayed by the to-be-displayed subtitle at the initial time point.

[0065] Based on this, the initial subtitle layer is used to display the display content of the corresponding type, the initial subtitle layer is obtained by text layout and drawing based on the subtitle text and the initial attribute value, and through appropriate text layout and drawing, the clarity and readability of the content displayed by the to-be-displayed subtitle at the initial time point can be effectively improved. The text layout and drawing process usually has a large amount of calculation, and especially when the subtitle animation involves many dynamic changing hierarchical attributes, the time consumption of the text layout and drawing process is usually long. By performing layout and drawing only on the subtitle image of the to-be-displayed subtitle at the initial time point, frequent calculation of a large amount of text layout and drawing can be avoided, thereby improving the rendering efficiency of the subtitle animation and improving the real-time performance and smoothness of the subtitle animation.

[0066] Specifically, the text layout and drawing is performed based on the subtitle text and the initial attribute value, and specifically, the subtitle text and the initial attribute value can be input into a text layout and drawing function, and then the initial subtitle layer is generated. The text layout and drawing function can be a function provided in related technologies or a function optimized based on the current video scene, and the embodiments of the present disclosure do not limit this. The determination formula of the initial subtitle layer is as follows: k,0 = typo(text, V k,0 )

[0067] wherein image k,0 is the initial subtitle layer corresponding to the kth image level, typo() is the text layout and drawing function, text is the subtitle text, V k,0 is a set of initial attribute values of all hierarchical attributes corresponding to the kth image level.

[0068] Exemplarily, in the text layout and drawing, in addition to the initial attribute value of the dynamic attribute, the attribute value of the static attribute also needs to be used. When a hierarchical attribute of a certain type is an empty set, the attribute value of the hierarchical attribute is usually set to a preset attribute value in a preset style. Specifically, the font size, text alignment mode, line spacing, word spacing, interface size of the display interface, and the attribute value of the subtitle text can be input into a text wrapping algorithm, the text layout is performed through the text wrapping algorithm to realize reasonable layout of the subtitle text, and then image drawing is performed based on the text layout result to obtain the initial subtitle layer. The font size belongs to the hierarchical attribute, and the text alignment mode, line spacing, word spacing, and interface size of the display interface can be obtained through pre-setting or real-time input, and the text wrapping algorithm can adopt a greedy algorithm, a dynamic programming algorithm, etc., and the embodiments of the present disclosure do not limit this.

[0069] Step 203: obtaining a target attribute value of each image level, and performing image transformation on the initial subtitle layer of each image level based on the target attribute value of the image level to obtain a target subtitle layer of the image level at the target time point.

[0070] The target attribute value is an attribute value of a level attribute of the image level at a target time point, and the target time point is a time point after an initial time point. In general, the initial time point is a display start time point of the subtitle, and the target time point is another display time point of the subtitle. The target time point can be multiple, and thus the target subtitle layer can also be multiple. For example, all display time points after the display start time point of the subtitle can be determined as the target time point. The initial time point can also be another display time point of the subtitle, which is not limited in the embodiments of the present disclosure.

[0071] Specifically, the initial subtitle layer is a two-dimensional digital matrix composed of pixels, and the image is also a two-dimensional digital matrix composed of pixels. Therefore, the initial subtitle layer can be essentially regarded as a special image, and thus image transformation can be performed on the initial subtitle layer to obtain a new subtitle layer, so as to add a dynamic effect to the subtitle, such as color transformation or affine transformation, and the like.

[0072] Based on this, by obtaining the target attribute value of the level attribute at the target time point, the corresponding initial subtitle layer can be transformed based on the target attribute value to obtain the target subtitle layer of the image level at the target time point. Therefore, the target subtitle layer is not obtained by text layout and drawing, but is obtained by image transformation based on the initial subtitle layer. The target subtitle layer is obtained by image transformation with small computational complexity, which can effectively improve the generation efficiency of the target subtitle layer.

[0073] Step 204: superimposing the target subtitle layers of the image levels to generate a target subtitle image of the to-be-displayed subtitle at the target time point.

[0074] The target subtitle layers are used to display the display content of the subtitle image of the to-be-displayed subtitle at the target time point. Different target subtitle layers display different types of display content. Superimposing multiple target subtitle layers is equivalent to superimposing different types of display content, which can generate a target subtitle image that combines different types of display content, which is equivalent to obtaining a target frame image displayed by the to-be-displayed subtitle at the target time point.

[0075] Based on this, the initial attribute value of the hierarchical attribute of each image level at the initial time point is obtained, and the subtitle text of the to-be-displayed subtitle is obtained. Since the image level is based on different types of display content in the subtitle image corresponding to the to-be-displayed subtitle, text layout and drawing can be performed based on the subtitle text and the initial attribute value to obtain the initial subtitle layer of the image level at the initial time point, thereby realizing the division of the subtitle layer for different display content. Then, based on the target attribute value of the hierarchical attribute of the image level at the target time point, the initial subtitle layer is image-transformed to obtain the target subtitle layer of the image level at the target time point, and then the target subtitle layers of each image level are superimposed to generate the target subtitle image of the to-be-displayed subtitle at the target time point. In the image transformation, the image transformation process of the corresponding initial subtitle layer is controlled by each target attribute value, realizing the fine control of each image level, so as to finely control different types of display content in the subtitle image corresponding to the to-be-displayed subtitle, and enrich the animation effect of the to-be-displayed subtitle. In addition, for multiple frame images of the to-be-displayed subtitle, the initial subtitle layer can be regarded as a layer in the initial frame image displayed by the to-be-displayed subtitle at the initial time point, and the target subtitle layer can be regarded as a layer in the target frame image displayed by the to-be-displayed subtitle at the target time point. In the subtitle animation rendering process, only text layout and drawing are needed when generating the initial subtitle layer, and no text layout and drawing are needed when generating the target subtitle layer. Instead, the target subtitle layer is obtained by image transformation based on the initial subtitle layer, which can avoid frequent text layout and drawing with large computational complexity, thereby improving the rendering efficiency of the subtitle animation and enhancing the real-time performance and smoothness of the subtitle animation.

[0076] Specifically, in the subtitle animation rendering process, only one initial text layout and drawing of the subtitle are needed to obtain the initial subtitle layer corresponding to each image level, and then the initial subtitle layer is image-transformed to obtain the target subtitle layer, and the target subtitle image is generated by superimposing multiple target subtitle layers, thereby realizing the dynamic transformation process of converting the subtitle animation into the subtitle image, and improving the rendering efficiency.

[0077] In a possible implementation, based on the target attribute value of the image level, the initial subtitle layer of the image level is image-transformed to obtain the target subtitle layer of the image level at the target time point. Specifically, the attribute change amount between the target attribute value of the image level and the initial attribute value of the image level can be determined, and based on the attribute change amount, the initial subtitle layer of the image level is image-transformed to obtain the target subtitle layer of the image level at the target time point.

[0078] Based on this, the attribute change amount is determined according to the target attribute value and the initial attribute value, and then the initial subtitle layer corresponding to the attribute change amount is subjected to image transformation to obtain the target subtitle layer of the image level at the target time point. Since the initial attribute value is the attribute value of the level attribute of the image level at the initial time point, and the target attribute value is the attribute value of the level attribute of the image level at the target time point, the attribute change amount can be used to determine whether the attribute value of the level attribute changes and the specific change amount when the target time point is reached. In other words, the attribute change amount indicates the difference between the level attribute values corresponding to the initial time point and the target time point. Using the attribute change amount as a reference factor for image transformation of the initial subtitle layer can improve the reliability of image transformation and further improve the accuracy of the image transformation result, thereby obtaining an accurate target subtitle layer.

[0079] In a possible implementation, referring to FIG. 3, which is a first optional flowchart for determining a target subtitle layer provided by an embodiment of the present disclosure, the level attribute includes a color attribute, the target attribute value includes a target color value of the color attribute, and the attribute change amount includes a color change amount of the color attribute. The initial subtitle layer of the image level is subjected to image transformation based on the attribute change amount to obtain the target subtitle layer of the image level at the target time point. Specifically, when the color change amount indicates that the image level changes in color, the color value of a non-transparent pixel point in the initial subtitle layer is transformed into the target color value to obtain the target subtitle layer of the image level at the target time point.

[0080] The color attribute is used to indicate the color state of the display content of the corresponding type of the image level, and can be regarded as an inline attribute that only acts on part of the image level of the to-be-displayed subtitle. The inline attribute can be denoted as S line The color change amount is the difference between the target color value and the initial color value. The target color value is the color value of the color attribute at the target time point, and the initial color value is the color value of the color attribute at the initial time point.

[0081] The non-transparent pixel point in the initial subtitle layer belongs to the part that needs to be displayed, and the transparent pixel point in the initial subtitle layer belongs to the part that does not need to be displayed. The non-transparent pixel point in the initial subtitle layer can be used to determine the display content of the corresponding type in the to-be-displayed subtitle. For example, for the image level where the text of the subtitle is located, the non-transparent pixel point in the initial subtitle layer can be used to determine the text content of the subtitle.

[0082] Specifically, the value of the color attribute can be represented by a numerical value, when the color change amount indicates that the image level changes in color, the color change amount is not 0, and when the color change amount indicates that the image level does not change in color, the color change amount is 0, therefore, by judging whether the color change amount corresponding to each image level is 0, it can be accurately determined whether the color change amount indicates that the image level changes in color.

[0083] Based on this, the image transformation includes color transformation, for any one image level, when the color change amount indicates that the image level changes in color, the initial subtitle layer representing the image level needs to be color transformed, specifically the color value of the non-transparent pixel point in the corresponding initial subtitle layer is transformed into the target color value, which is equivalent to transforming the initial color value into the target color value, obtaining the target subtitle layer of the image level at the target time point, after determining the subtitle animation of the to-be-displayed subtitle through the initial subtitle layer and the target subtitle layer, in the display process of the to-be-displayed subtitle, the dynamic effect generated by the color transformation can increase the visual dynamic effect, improve the information transmission efficiency and improve the viewing experience of the audience.

[0084] Taking the image transformation including only color transformation as an example, the determination formula of the target subtitle layer is:

[0085] Wherein, is the target subtitle layer corresponding to the kth image level at the target time point t, image k,0 is the initial subtitle layer corresponding to the kth image level at the initial time point, v k,t is the target color value of the color attribute corresponding to the kth image level at the target time point t, color() is a color transformation function, which is used to transform the color value of the non-transparent pixel point in the initial subtitle layer into the target color value.

[0086] Further, the color value of the pixel point in the target subtitle layer can be represented as:

[0087] Wherein, is the color value of the pixel point (x, y) in the target subtitle layer corresponding to the kth image level at the target time point t, f k,0 (x, y) is the color value of the pixel point (x, y) in the initial subtitle layer corresponding to the kth image level at the initial time point, f k,0 (x, y)≠0 means non-transparent pixel point, and the transparent color value is 0, v k,t is the target color value of the color attribute corresponding to the kth image level at the target time point t.

[0088] In a possible implementation, referring to FIG. 4, which is a second optional flowchart for determining a target subtitle layer provided by an embodiment of the present disclosure, the hierarchical attribute includes a geometric attribute, the attribute change amount includes a geometric change amount of the geometric attribute, and the initial subtitle layer of the image hierarchy is subjected to image transformation based on the attribute change amount to obtain the target subtitle layer of the image hierarchy at the target time point. Specifically, affine transformation parameters can be determined based on the geometric change amount, and an affine transformation matrix can be constructed according to the affine transformation parameters. The initial subtitle layer is subjected to affine transformation based on the affine transformation matrix to obtain the target subtitle layer of the image hierarchy at the target time point.

[0089] The geometric attribute is used to indicate a geometric state of the display content of the corresponding type of the image hierarchy, for example, the geometric state can include a position, a size, a shape, and the like. The geometric attribute can be regarded as a paragraph attribute acting on the entire to-be-displayed subtitle. The paragraph attribute can be denoted as S para The geometric change amount is a difference between a target geometric value and an initial geometric value. The target geometric value is a geometric value of the geometric attribute at the target time point. The initial geometric value is a geometric value of the geometric attribute at the initial time point.

[0090] Specifically, the affine transformation can generally involve multiple types of geometric change amounts. For example, the geometric attribute can include a rotation attribute, a scaling attribute, a shear attribute, and a position attribute. The geometric change amount can include a rotation change amount of the rotation attribute, a scaling change amount of the scaling attribute, a shear change amount of the shear attribute, and a movement change amount of the position attribute. Specifically, affine transformation parameters are determined based on the geometric change amount. Each type of geometric change amount can determine corresponding affine transformation parameters.

[0091] Based on this, the image transformation includes the affine transformation. The affine transformation refers to the transformation of the position, the size, or the shape of the subtitle layer. When the hierarchical attribute includes the geometric attribute, that is, the geometric attribute is not an empty set, the subtitle layer corresponding to all image hierarchies needs to be subjected to the affine transformation. The affine transformation matrix is constructed according to the affine transformation parameters. Then, the corresponding initial subtitle layer is subjected to the affine transformation based on the affine transformation matrix to obtain the target subtitle layer of the image hierarchy at the target time point. Specifically, the target subtitle layer is obtained according to the product of the affine transformation matrix and the matrix corresponding to the initial subtitle layer. After the subtitle animation of the to-be-displayed subtitle is determined through the initial subtitle layer and the target subtitle layer, the dynamic effect generated by the affine transformation can increase the visual appeal in the display process of the to-be-displayed subtitle.

[0092] Taking the image transformation as an example, the determination formula of the target subtitle layer is as follows:

[0093] V k,Δ = {v i,t -v i,0 |vi,t ∈V k,t ,v i,0 ∈V k,0}, is the target subtitle layer corresponding to the kth image level at the target time point t, image k,0 is the initial subtitle layer corresponding to the kth image level at the initial time point, V k,Δ is a set of geometric changes of all geometric properties corresponding to the kth image level, V k,t is a set of target geometric values of all geometric properties corresponding to the kth image level, V k,0 is a set of initial geometric values of all geometric properties corresponding to the kth image level, v i,t is the target geometric value of the i-th geometric property, v i,0 is the initial geometric value of the i-th geometric property, affine() is an affine transformation function, which is used to transform a pixel point in the initial subtitle layer to a new coordinate to obtain a target subtitle layer. The affine transformation function can be specifically used to construct an affine transformation matrix according to affine transformation parameters, and perform affine transformation on the corresponding initial subtitle layer based on the affine transformation matrix.

[0094] Further, a pixel point in the target subtitle layer can be represented as:

[0095] wherein, is the color value of the pixel point (x', y') in the target subtitle layer corresponding to the kth image level at the target time point t, f k,0 (x,y) is the color value of the pixel point (x,y) in the initial subtitle layer corresponding to the kth image level at the initial time point, which is equivalent to transforming the pixel point (x,y) in the initial subtitle layer to a new coordinate (x',y').

[0096] Specifically, for any type of affine transformation, the transformation formula of the coordinate is as follows:

[0097] wherein, (x,y) is the coordinate of the pixel point in the initial subtitle layer, (x',y') is the coordinate of the pixel point in the target subtitle layer, is a reference transformation matrix, a, b, c, d, e, f are all affine transformation parameters, a is a horizontal scaling factor, b is a horizontal shear factor, c is a horizontal translation, d is a vertical shear factor, e is a vertical scaling factor, and f is a vertical translation.

[0098] It should be noted that when the transformation type of the affine transformation is a scaling transformation, a and e are determined by the scaling variation amount; when the transformation type of the affine transformation is a skew transformation, b and d are determined by the skew variation amount; when the transformation type of the affine transformation is a movement transformation, c and f are determined by the skew variation amount; and when the transformation type of the affine transformation is a rotation transformation, a, b, d, and e are determined by the rotation variation amount.

[0099] Therefore, when multiple types of affine transformations need to be performed, the reference transformation matrix corresponding to each type of affine transformation needs to be determined first, then the reference transformation matrices are multiplied in sequence to obtain the affine transformation matrix, and finally the affine transformation matrix is multiplied with the matrix corresponding to the initial subtitle layer to obtain the target subtitle layer.

[0100] In a possible implementation, the hierarchical attribute includes a mask attribute, the target attribute value includes a target mask value of the mask attribute, and the initial subtitle layer of the image hierarchy is image-transformed based on the target attribute value of the image hierarchy to obtain the target subtitle layer of the image hierarchy at the target time point. Specifically, the mask region can be determined in the initial subtitle layer based on the target mask value; and the color value of a non-transparent pixel point located in the mask region in the initial subtitle layer is transformed into a preset mask color value to obtain the target subtitle layer of the image hierarchy at the target time point.

[0101] The mask attribute is used to indicate a mask element of a display content of the image hierarchy of the corresponding type, and the mask attribute can be regarded as a special attribute that does not directly act on the to-be-displayed subtitle. The special attribute can be denoted as S spec The target mask value is a mask value of the mask attribute at the target time point, and the target mask value is used to determine a mask region of the to-be-displayed subtitle at the target time point. The mask region is a specific region in the initial subtitle layer whose display effect is controlled. The display effect controlled in the mask region can be that the color value of a non-transparent pixel point is transformed into a mask color value. The mask color value can be preset as a transparent color value or a non-transparent color value, which is not limited in the embodiment of the present disclosure.

[0102] Based on this, the image transformation includes a mask transformation, the mask transformation refers to transforming a mask attribute, when the hierarchical attribute includes the mask attribute, that is, the mask attribute is not an empty set, all subtitle layers corresponding to the image hierarchy need to be subjected to the mask transformation, based on a target mask value, a mask region is determined in the corresponding initial subtitle layer, then in the corresponding initial subtitle layer, the color value of the transparent pixel point in the initial subtitle layer is controlled to remain unchanged, which can ensure that the display content of the image hierarchy corresponding type remains unchanged, while the color value of the non-transparent pixel point located in the mask region is transformed into the mask color value, to obtain a target subtitle layer, which realizes accurate control of the display effect of the mask region. After determining the subtitle animation of the to-be-displayed subtitle through the initial subtitle layer and the target subtitle layer, in the display process of the to-be-displayed subtitle, the mask transformation will make the display content in the to-be-displayed subtitle gradually increase, decrease or move in the region that is masked, and through the dynamically changing mask region, part of the display content in the to-be-displayed subtitle is processed, which can highlight the key part in the subtitle, thereby increasing the visual attraction, improving the information transmission efficiency and improving the viewing experience of the audience.

[0103] In a possible implementation, referring to FIG. 5, FIG. 5 is a third optional flowchart for determining a target subtitle layer provided by an embodiment of the present disclosure, the mask attribute includes a color change proportion attribute, the target mask value includes a target color change proportion value of the color change proportion attribute, and the mask region is determined in the initial subtitle layer based on the target mask value. Specifically, the layer width of the target subtitle layer can be obtained, and the width threshold can be determined according to the product of the target color change proportion value and the layer width. Then, in the initial subtitle layer, the region with a horizontal coordinate less than the width threshold is determined as the mask region.

[0104] In the above implementation, the mask transformation includes a color mask transformation, the mask attribute includes a color change proportion attribute corresponding to the color mask transformation, the target color change proportion value is an attribute value of the color change proportion attribute at the target time point, and the layer width of the target subtitle layer is specifically a distance between a left boundary of the target subtitle layer and a right boundary of the target subtitle layer.

[0105] Therefore, the width threshold value is determined by the product of the target color change ratio value and the layer width, and then the region in the initial subtitle layer whose horizontal coordinates belong to the width threshold value is determined as the mask region, and the horizontal coordinates of each pixel point in the mask region are all less than the width threshold value; when the target color change ratio value is larger, the width threshold value is larger, and the mask region is larger, and vice versa, when the target color change ratio value is smaller, the width threshold value is smaller, and the mask region is smaller. After the target color change ratio value of the to-be-displayed subtitle at the target time point is determined, the mask region can be accurately determined by the target color change ratio value, so that the color value of the non-transparent pixel point located in the mask region is transformed into the mask color value, the mask color value is usually a non-transparent color value, the dynamic effect of the display content of the to-be-displayed subtitle changing character by character can be realized, so that the visual effect is enriched, the uncolored part in the to-be-displayed subtitle and the junction of the colored part and the uncolored part can be further highlighted, so that the visual attraction is increased, the information transmission efficiency is improved, and the viewing experience of the audience is improved.

[0106] Specifically, generally, the horizontal coordinates of the pixel points on the left boundary in the initial subtitle layer can be set as 0, and the positive direction of the x-axis is horizontally to the right, therefore, the horizontal coordinates of the pixel points at other positions are all greater than 0, and the mask region is located on the left side of the non-mask region; or the horizontal coordinates of the pixel points on the left boundary in the initial subtitle layer can be set as other numerical values, assuming that the horizontal coordinates of the pixel points on the left boundary are a first preset value, at this time, the first preset value can be regarded as the offset of the region, and the product of the target color change ratio value and the layer width needs to be added to the first preset value to obtain the width threshold value.

[0107] Exemplarily, by setting the color change ratio attribute, the to-be-displayed subtitle can add a karaoke effect, that is, at a certain display time point, the front half of the to-be-displayed subtitle displays one color, and the back half of the to-be-displayed subtitle displays another color.

[0108] Taking the image transformation including only color mask transformation as an example, the determination formula of the pixel point in the target subtitle layer is:

[0109] wherein, is the color value of the pixel point (x, y) in the target subtitle layer corresponding to the kth image level at the target time point t, f k,0 (x, y) is the color value of the pixel point (x, y) in the initial subtitle layer corresponding to the kth image level at the initial time point, (x, y) is the coordinates of the pixel point, Color kf is the mask color value, x is the horizontal coordinates of the pixel point in the initial subtitle layer, w is the layer width, Ckf(t) is the target color change ratio value, w* Ckf(t) is the width threshold value.

[0110] In a possible implementation, referring to FIG. 6, FIG. 6 is a fourth optional flowchart for determining a target subtitle layer provided by an embodiment of the present disclosure. The mask attribute includes a plurality of mask coordinate attributes, the target mask value includes target mask coordinates of each mask coordinate attribute, and the mask region is determined in the initial subtitle layer based on the target mask value. Specifically, the visible region can be determined in the initial subtitle layer according to the target mask coordinates, and the region outside the visible region is determined as the mask region in the initial subtitle layer.

[0111] In the formula, the mask transformation includes a visibility mask transformation, the mask attribute includes a visibility mask attribute corresponding to the visibility mask transformation, the target mask coordinate is an attribute value of the mask coordinate attribute at the target time point, and the target mask coordinate is located on the region boundary of the visible region.

[0112] Specifically, the shape of the visible region can be a rectangle, a circle, an ellipse, or the like, which is not limited in the embodiment of the present disclosure. Taking the shape of the visible region as a rectangle as an example, the number of mask coordinate attributes can be two, one target mask coordinate of one mask coordinate attribute is the coordinate of the upper left corner of the visible region, and the target mask coordinate of the other mask coordinate attribute is the coordinate of the lower right corner of the visible region. In addition, the number of mask coordinate attributes can also be other numbers, and the target mask coordinate can also be the coordinate of other positions of the visible region, which is not limited in the embodiment of the present disclosure.

[0113] Based on this, the visible region is determined in the corresponding initial subtitle layer through each target mask coordinate, and then the region outside the visible region is determined as the mask region, so that the color value of the non-transparent pixel point located in the mask region is transformed into the mask color value, which can be a transparent color value. The dynamic effect of displaying the display content of the to-be-displayed subtitle word by word can be realized, so as to realize rich visual effects, further highlight the display part in the to-be-displayed subtitle, increase the visual attraction, improve the information transmission efficiency, and improve the viewing experience of the audience.

[0114] Taking the image transformation including only the visibility mask transformation as an example, the determination formula of the pixel point in the target subtitle layer is as follows:

[0115] In the formula, x and y are coordinates of the pixel point, and f is the color value of the pixel point (x, y) in the target subtitle layer corresponding to the kth image level at the target time point t, f k,0 (x, y) is the color value of the pixel point (x, y) in the initial subtitle layer corresponding to the kth image level at the initial time point, (x s,0 ,y s,0 ) is the first target mask coordinate, (x e,0 ,y e,0) the second target mask coordinates, at this time, the visible region is a matrix, the first target mask coordinates are located at the upper left corner of the visible region, the second target mask coordinates are located at the lower right corner of the visible region, and the transparent color value is 0.

[0116] In a possible implementation, referring to FIG. 7, FIG. 7 is a fifth optional flow diagram for determining a target subtitle layer provided by an embodiment of the present disclosure. The image transformation can include one or more of color transformation, affine transformation, and mask transformation. In the image transformation process, it is necessary to determine whether color transformation, affine transformation, and mask transformation are needed in sequence, and the corresponding transformation is performed only when needed, which can improve the flexibility and reliability of image transformation.

[0117] It should be noted that when the image transformation includes multiple transformations, the determination of the target subtitle layer needs to be adjusted. The following will be described in detail taking the image transformation including color transformation, affine transformation, and mask transformation as an example.

[0118] First, the color value of the non-transparent pixel point in the initial subtitle layer of the image level is transformed into a target color value to obtain the target subtitle layer of the image level at the target time point, which can be adjusted to: the color value of the non-transparent pixel point in the initial subtitle layer of the image level is transformed into a target color value to obtain the first subtitle layer of the image level at the target time point.

[0119] Then, the initial subtitle layer of the image level is subjected to affine transformation based on the affine transformation matrix to obtain the target subtitle layer of the image level at the target time point, which can be adjusted to: the first subtitle layer of the image level is subjected to affine transformation based on the affine transformation matrix to obtain the second subtitle layer of the image level at the target time point.

[0120] Then, the color value of the non-transparent pixel point located in the mask region in the initial subtitle layer of the image level is transformed into a preset mask color value to obtain the target subtitle layer of the image level at the target time point, which can be adjusted to: the color value of the non-transparent pixel point located in the mask region in the second subtitle layer of the image level is transformed into a preset mask color value to obtain the target subtitle layer of the image level at the target time point.

[0121] It should be noted that referring to FIG. 8, FIG. 8 is an optional flow diagram for mask transformation provided by an embodiment of the present disclosure. The mask transformation can include one or more of color mask transformation and visibility mask transformation. In the mask transformation process, it is necessary to determine whether color mask transformation and visibility mask transformation are needed in sequence, and the corresponding transformation is performed only when needed, which can improve the flexibility and reliability of mask transformation.

[0122] It should be noted that when the mask transformation includes multiple transformations, the determination step of the target subtitle layer needs to be further adjusted. Taking the case that the mask transformation includes color mask transformation and visibility mask transformation as an example, the determination step of the target subtitle layer is described in detail.

[0123] Firstly, the first mask region needs to be determined based on the target color change ratio value when the color mask transformation is performed. The mask color value corresponding to the color mask transformation can be set as the color value Color kf In the initial subtitle layer of the image level, the color value of the non-transparent pixel point located in the mask region is transformed into the preset mask color value, and the target subtitle layer of the image level at the target time point is obtained. It can be adjusted that in the corresponding second subtitle layer, the color value of the non-transparent pixel point located in the first mask region is transformed into the color value Color kf , and the third subtitle layer of the image level at the target time point is obtained.

[0124] Then, the second mask region needs to be determined based on the target mask coordinate value when the visibility mask transformation is performed. The mask color value corresponding to the color mask transformation can be set as the transparent color value. In the initial subtitle layer of the image level, the color value of the non-transparent pixel point located in the mask region is transformed into the preset mask color value, and the target subtitle layer of the image level at the target time point is obtained. It can be adjusted that in the third subtitle layer of the image level, the color value of the non-transparent pixel point located in the second mask region is transformed into the transparent color value, and the target subtitle layer of the image level at the target time point is obtained.

[0125] Based on this, when the image transformation includes color transformation, affine transformation, color mask transformation and visibility mask transformation, the determination formula of the first subtitle layer can be specifically:

[0126] Wherein, is the first subtitle layer corresponding to the kth image level at the target time point t, image k,0 is the initial subtitle layer corresponding to the kth image level at the initial time point, v k,t is the target color value of the color attribute at the target time point t, color() is a color transformation function, which is used to transform the color value of the non-transparent pixel point in the initial subtitle layer into the target color value.

[0127] Further, the color value of the pixel point in the first subtitle layer can be specifically represented as:

[0128] Wherein, is the color value of the pixel point (x, y) in the first subtitle layer corresponding to the kth image level at the target time point t, f k,0(x,y) is the color value of the pixel point (x,y) in the initial subtitle layer corresponding to the kth image level at the initial time point, f k,0 (x,y)≠0 means a non-transparent pixel point, and the transparent color value is 0, v k,t is the target color value of the color attribute corresponding to the kth image level at the target time point t.

[0129] Then, the determination formula of the second subtitle layer can be specifically as follows:

[0130] wherein, V k,Δ ={v i,t -v i,0 |v i,t ∈V k,t ,v i,0 ∈V k,0}, is the second subtitle layer corresponding to the kth image level at the target time point t, image k,0 is the first subtitle layer corresponding to the kth image level at the target time point t, V k,Δ is a set of geometric variation amounts of all geometric attributes corresponding to the kth image level, V k,t is a set of target geometric values of all geometric attributes corresponding to the kth image level, V k,0 is a set of initial geometric values of all geometric attributes corresponding to the kth image level, v i,t is the target geometric value of the ith geometric attribute, v i,0 is the initial geometric value of the ith geometric attribute, and affine() is an affine transformation function, which is used to transform the pixel point in the first subtitle layer to a new coordinate to obtain the second subtitle layer.

[0131] Further, the pixel point in the second subtitle layer can be specifically represented as:

[0132] wherein, is the color value of the pixel point (x',y') in the second subtitle layer corresponding to the kth image level at the target time point t, is the color value of the pixel point (x,y) in the first subtitle layer corresponding to the kth image level at the initial time point, which is equivalent to transforming the pixel point (x,y) in the first subtitle layer to a new coordinate (x',y'), and the affine transformation function can be specifically used to construct an affine transformation matrix according to the affine transformation parameters, and perform affine transformation on the corresponding first subtitle layer based on the affine transformation matrix.

[0133] Then, the determination formula of the pixel point in the third subtitle layer can be specifically as follows:

[0134] wherein, Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, kf Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, kf (t) is a target color change ratio value, w Ckf(t) is a width threshold.

[0135] Then, the determination formula of the pixel point in the target subtitle layer can be specifically:

[0136] wherein, Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, s,0 Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, s,0 Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, e,0 Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t, e,0 Color k (x, y, t) is a color value of a pixel point (x, y) in the third subtitle layer corresponding to the kth image level at a target time point t,

[0137] In a possible implementation, initial attribute values of each image level in a subtitle image corresponding to a to-be-displayed subtitle are obtained, which can be specifically: a dynamic attribute set of the subtitle image corresponding to the to-be-displayed subtitle is obtained, the dynamic attribute set is divided into an equivalent attribute set and a non-equivalent attribute set; when the non-equivalent attribute set is an empty set, initial attribute values of each image level in the subtitle image corresponding to the to-be-displayed subtitle are obtained from the equivalent attribute set.

[0138] wherein, the dynamic attribute set includes level attributes used to realize an animation effect, the equivalent attribute set includes level attributes used to realize the animation effect through image transformation, the non-equivalent attribute set includes level attributes that cannot realize the animation effect through image transformation, the level attributes included in the dynamic attribute set are all dynamic attributes, the level attributes included in the equivalent attribute set are all equivalent attributes, the level attributes included in the non-equivalent attribute set are all non-equivalent attributes, and the non-equivalent attribute set can be denoted as S not .

[0139] Exemplarily, referring to FIG. 9, FIG. 9 is a schematic diagram of an optional grouping of the dynamic attribute set according to an embodiment of the present disclosure.

[0140] The equivalent attribute set can include the color attribute, the geometry attribute and the mask attribute as described above, the color attribute can include the text color attribute, the border color attribute, the shadow color attribute and the background color attribute, the geometry attribute includes the rotation attribute, the scaling attribute, the skew attribute and the position attribute, the mask attribute includes the color change ratio attribute and the mask coordinate attribute, the non-equivalent attribute set can specifically include the character spacing attribute, the border width attribute, the shadow distance attribute and the edge blur attribute, in addition to the above, the dynamic attribute set can further include other hierarchical attributes, which are not limited herein according to the embodiments of the present disclosure.

[0141] It is worth noting that when the color attribute of the subtitle image corresponding to the to-be-displayed subtitle is an empty set, i.e., the text color attribute, the border color attribute, the shadow color attribute and the background color attribute are all empty sets, the subtitle image corresponding to the to-be-displayed subtitle can be processed as a whole hierarchical layer, denoted as L a At this time, the to-be-displayed subtitle does not need to be split into multiple image hierarchical layers, and the preset color value in the preset style can be used to configure the color of each type of display content in the subtitle image corresponding to the to-be-displayed subtitle; when the color attribute of the subtitle image corresponding to the to-be-displayed subtitle is not an empty set, i.e., any one of the text color attribute, the border color attribute, the shadow color attribute and the background color attribute is not an empty set, the subtitle image corresponding to the to-be-displayed subtitle needs to be split into multiple image hierarchical layers, therefore, the list of image hierarchical layers is as follows:

[0142] wherein H is the hierarchical level of the image hierarchical layer, L k is one of the image hierarchical layers, L a is the whole hierarchical layer, L d is the image hierarchical layer corresponding to the background of the subtitle image, L s is the image hierarchical layer corresponding to the shadow of the subtitle image, L b is the image hierarchical layer corresponding to the border of the subtitle image, L w is the image hierarchical layer corresponding to the text of the subtitle image, S line is the color attribute, is an empty set.

[0143] Based on this, when the non-equivalent attribute set is an empty set, the attribute value representing the non-equivalent attribute does not change, and the attribute value of the non-equivalent attribute is usually set to the preset attribute value in the preset style. At this time, the initial attribute value of each image level in the subtitle image corresponding to the to-be-displayed subtitle is obtained from the equivalent attribute set, and the initial attribute value belongs to the attribute value of the dynamic attribute. Subsequently, in the subtitle animation rendering process, only the dynamic effect corresponding to the equivalent attribute needs to be added. Therefore, when generating the target subtitle layer, text layout drawing is not needed, but the target subtitle layer is obtained by image transformation based on the initial subtitle layer, which can avoid frequent text layout drawing with large computational complexity, thereby improving the rendering efficiency of the subtitle animation and enhancing the real-time performance and smoothness of the subtitle animation.

[0144] In addition, since the non-equivalent attribute is closely related to the glyph, when the non-equivalent attribute set is not an empty set, the attribute value representing the non-equivalent attribute changes. Since image transformation cannot effectively process dynamic transformation of the glyph, in the subsequent subtitle animation rendering process, image transformation is no longer used, but the hierarchical attribute of the to-be-displayed subtitle at each display time point is obtained. Then, for each display time point, text layout drawing is performed based on the subtitle text and the hierarchical attribute at the display time point to obtain the real subtitle image at the display time point. This is equivalent to re-layout drawing for each frame of the to-be-displayed subtitle animation, and the glyph dynamic change subtitle animation can be generated by the real subtitle image at each display time point, thereby ensuring the quality of the subtitle animation.

[0145] Specifically, the initial attribute value can be divided into the attribute value of the dynamic attribute and the attribute value of the static attribute. In addition to obtaining the dynamic attribute set of the to-be-displayed subtitle, the static attribute set of the to-be-displayed subtitle also needs to be obtained. The hierarchical attributes included in the static attribute set are all static attributes. When the static attribute set is an empty set, the initial attribute value of each image level in the subtitle image corresponding to the to-be-displayed subtitle does not need to be obtained from the static attribute set. When the static attribute set is not an empty set, the initial attribute value of each image level in the subtitle image corresponding to the to-be-displayed subtitle also needs to be obtained from the static attribute set, and the initial attribute value belongs to the attribute value of the static attribute.

[0146] In a possible implementation, the hierarchical attribute in the equivalent attribute set includes a color attribute of each image layer, and when the nonequivalent attribute set is an empty set, the initial attribute value of each image layer in the subtitle image corresponding to the to-be-displayed subtitle is obtained from the equivalent attribute set. Specifically, when the nonequivalent attribute set is an empty set, for each image layer, the color attribute corresponding to the current image layer in the equivalent attribute set is kept unchanged, and the color attributes corresponding to the remaining image layers in the equivalent attribute set are adjusted to be transparent, to obtain the hierarchical attribute set corresponding to the current image layer; and the initial attribute value of the image layer in the subtitle image corresponding to the to-be-displayed subtitle is obtained from each hierarchical attribute set.

[0147] In the formula, adjusting the color attribute to be transparent is equivalent to adjusting the color value of the color attribute to be a transparent color value, and the color value can be represented by a numerical value. For example, the transparent color value can be 0. Each image layer has a corresponding color attribute. For example, assuming that the number of image layers is four, the first image layer corresponds to the background of the subtitle image, the second image layer corresponds to the shadow of the subtitle image, the third image layer corresponds to the frame of the subtitle image, and the fourth image layer corresponds to the text of the subtitle image. Then, the color attribute corresponding to the first image layer can be a background color attribute, the color attribute corresponding to the second image layer can be a shadow color attribute, the color attribute corresponding to the third image layer can be a frame color attribute, and the color attribute corresponding to the fourth image layer can be a text color attribute.

[0148] Based on this, when the nonequivalent attribute set is an empty set, the hierarchical attribute set corresponding to each image layer is generated from the equivalent attribute set. Specifically, the color attribute corresponding to the current image layer in the equivalent attribute set is kept unchanged, and the color attributes corresponding to the remaining image layers in the equivalent attribute set are adjusted to be transparent. Then, the initial attribute value of the image layer corresponding to the to-be-displayed subtitle is obtained from the hierarchical attribute set, so that the initial attribute values of different image layers are different only in initial color values, and other attribute values are the same. Subsequently, when the initial subtitle layer is generated, only the color value of the display content of the type corresponding to the current image layer is retained in the initial subtitle layer of each image layer, and the color value of the display content of other types is transparent, so that the initial subtitle layer only displays the display content of the type corresponding to the current image layer.

[0149] Specifically, the hierarchical attribute set of the to-be-displayed subtitle can include a static attribute set and a dynamic attribute set. The static attribute set includes hierarchical attributes that cannot realize animation effects, and the dynamic attribute set includes hierarchical attributes used to realize animation effects. For example, the image layer L w For example, the image layer L w The construction formula of the corresponding hierarchical attribute set is as follows: Pw =(P-{p i |i∈N w-})∪{p′ i =<0,0,t s ,t e ,0>|i∈N w-}

[0150] Among them, P w For image level L w The corresponding hierarchical attribute set, P is the hierarchical attribute set, and N is the set of identifiers for the hierarchical attributes. w- In addition to image level L w The set of color attributes corresponding to other image levels besides N w- ={2c,3c,4c}, where 2c is the identifier for the background color attribute, 3c is the identifier for the border color attribute, 4c is the identifier for the shadow color attribute, and p i N refers to w- The i-th color attribute in the array, where the value of the color attribute is represented by a number, with 0 being the transparent color value, p′ i N refers to w- The adjusted i-th color attribute, <0,0,t s ,t e The first element in ,0> refers to N w- The initial value of the i-th color attribute in the array is the transparent color value, <0,0,t. s ,t e The second element in ,0> refers to N w- The final value of the i-th color attribute in t is the transparent color value. s This refers to the start time of the subtitle display, t e This refers to the end time of the subtitle display, <0,0,t s ,t e The fifth element in ,0> refers to N w- The color value of the i-th color attribute in the text remains a transparent color value during the subtitle display.

[0151] It should be noted that in P w During the construction process, when the static attribute set is empty and the non-equivalent attribute set is also empty, P specifically becomes the equivalent attribute set. This achieves the retention of dynamic attributes for text color based on the equivalent attribute set, while treating background color, shadow color, and border color attributes as transparent static attributes, thus making P... w Excluding static attributes, the values ​​of static attributes can be set to the default values ​​in the preset styles; however, when the static attribute set is not empty, the hierarchical attribute set P includes static attributes, such that P... wStatic properties will also be included.

[0152] It is worth noting that P w The construction formula similar to the construction formula of P d The hierarchical attribute set P d The construction formula similar to the construction formula of P s The hierarchical attribute set P s The construction formula similar to the construction formula of P b The hierarchical attribute set P b .

[0153] In a possible implementation, the hierarchical attribute set further includes an attribute change function corresponding to the initial attribute value, and the target attribute value of each image level is obtained, which can be specifically: obtaining a target time point; inputting the target time point into the attribute change function in each hierarchical attribute set for operation to obtain the target attribute value of each image level.

[0154] Wherein, each hierarchical attribute in the hierarchical attribute set has a corresponding attribute change function, the attribute change function is a transition curve expression of the hierarchical attribute change, the independent variable of the attribute change function is the time point, and the dependent variable of the attribute change function is the attribute value, for example, the attribute change function can be a linear function, a quadratic function or an exponential function, etc., which is not limited in the embodiment of the present disclosure; The specific formula for determining the target attribute value is as follows: V k,t = {v i,t = c i (t) | p i,k ∈ P k}

[0155] Wherein, V k,t is the set of target attribute values of all hierarchical attributes corresponding to the kth image level, v i,t is the target attribute value of the ith hierarchical attribute in P k at the target time point t, c i (t) is the attribute change function of the ith hierarchical attribute in P k , p i,k is the ith hierarchical attribute in P k , P k is the set of all hierarchical attributes corresponding to the kth image level.

[0156] Specifically, the attribute change function can be located in the multi-element set corresponding to the hierarchical attribute, and the attribute change function can be determined in various ways, for example, a relevant person can directly input a specific attribute change function, and for another example, the relevant person can input a plurality of sample data combinations, wherein the sample data combination includes a sample attribute value and a sample time point corresponding to the sample attribute value, and the sample time point is a display time point of the subtitle, and then an interpolation operation is performed based on the plurality of sample data combinations to determine the attribute change function. The determination manner of the attribute change function is not limited in the embodiments of the present disclosure.

[0157] Based on this, since each hierarchical attribute in the hierarchical attribute set has a corresponding attribute change function, by inputting the target time point into the attribute change function for operation, the accurate target attribute value of each hierarchical attribute can be obtained, thereby improving the reliability of the image transformation process.

[0158] In a possible implementation, referring to FIG. 10, FIG. 10 is an optional flowchart for determining a target attribute value provided by the embodiments of the present disclosure. The target time point is input into the attribute change function in each hierarchical attribute set for operation to obtain the target attribute value of each image level. Specifically, the target time point can be input into the attribute change function in each hierarchical attribute set for operation to obtain the reference attribute value of each image level. The subtitle text is input into the large language model for sentiment recognition to obtain the target sentiment information of the subtitle text. The target sentiment information and the reference attribute value are spliced and input into the regression model for regression to obtain the target attribute value of each image level.

[0159] The target sentiment information can represent the emotional atmosphere of the subtitle text, for example, the emotional atmosphere can include happiness, sadness, disappointment, surprise, and the like. The target sentiment information can be one of a plurality of emotion identifiers. Different emotion identifiers are used to indicate corresponding sentiment information. For example, the emotion identifier can include 1, 2, and 3, and the like.

[0160] Among them, the large language model (Large Language Model, LLM) is a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text; the large language model generally uses recurrent neural network (Recurrent Neural Network, RNN) or variants such as long short-term memory network (Long Short-Term Memory, LSTM) and gated recurrent unit (Gated Recurrent Unit, GRU) to capture the context information in the text sequence, so as to realize the generation of natural language text, language model evaluation, text classification, sentiment analysis and other tasks; in the field of natural language processing, large language models have been widely used, such as speech recognition, machine translation, automatic abstract, dialogue system, intelligent question and answer, etc.

[0161] Specifically, the large language model can be trained through a plurality of training samples related to the sentiment recognition task, so that the large language model can learn to predict the sentiment of the input text, thereby improving the prediction accuracy of the target sentiment information in the reasoning stage. The regression model can be trained through a plurality of training samples related to the attribute value prediction task, so that the regression model can learn the relationship between the input and the output, thereby improving the prediction accuracy of the target attribute value in the reasoning stage.

[0162] Based on this, the target time point is input into the attribute change function for operation, and the reference attribute values of each level attribute can be obtained, and then the subtitle text is input into the large language model, and the target sentiment information of the subtitle text is predicted through the large language model; then the target sentiment information and the reference attribute value are spliced and input into the regression model for regression, and the regression model can predict the target attribute value of each image level, which is equivalent to optimizing the attribute value of the hierarchical attribute through the target sentiment information of the subtitle text, so that the target attribute value can better fit the emotional atmosphere of the subtitle text, and the display effect and expressiveness of the subtitle can be further enhanced.

[0163] Exemplarily, when the emotional atmosphere is happy, the color value in the target attribute value is a color value for indicating warm color tone under the action of the regression model; when the emotional atmosphere is sad, the color value in the target attribute value is a color value for indicating cold color tone under the action of the regression model; when the emotional atmosphere is surprised, the target attribute value is obtained by adjusting the geometry value in the reference attribute value under the action of the regression model, so that the geometry change amount increases, and the display content of the to-be-displayed subtitle can be highlighted.

[0164] In a possible implementation, the plurality of image levels include a text level, a border level, a shadow level, and a background level, and the target subtitle image of the to-be-displayed subtitle at the target time point is generated by superimposing the target subtitle layers of each image level, specifically, the target subtitle image of the to-be-displayed subtitle at the target time point can be generated by sequentially superimposing the target subtitle layers in the order of the background level, the shadow level, the border level, and the text level.

[0165] Specifically, the order of the background level, the shadow level, the border level, and the text level is the superimposition order, and sequentially superimposing the plurality of target subtitle layers in the order of the background level, the shadow level, the border level, and the text level can sequentially superimpose the plurality of target subtitle layers from bottom to top according to the superimposition order.

[0166] The superimposition order is used to indicate the display priority of each target subtitle layer, and since the display priority of the target subtitle layer located above is higher than that of the target subtitle layer located below, the target subtitle layer located above can cover the target subtitle layer located below, and if the target subtitle layer located above is not completely transparent, the target subtitle layer located above can partially or completely block the target subtitle layer located below.

[0167] Therefore, the display content of the subtitle image corresponding to the to-be-displayed subtitle can include the text, the border, the shadow, and the background of the subtitle, the display content corresponding to the text level is the text of the subtitle image, the display content corresponding to the border level is the border of the subtitle image, the display content corresponding to the shadow level is the shadow of the subtitle image, and the display content corresponding to the background level is the background of the subtitle image. Since the display priority of the text of the subtitle image is higher than that of the border of the subtitle image, the display priority of the border of the subtitle image is higher than that of the shadow of the subtitle image, and the display priority of the shadow of the subtitle image is higher than that of the background of the subtitle image, the plurality of target subtitle layers need to be sequentially superimposed in the order of the background level, the shadow level, the border level, and the text level, so as to appropriately display each type of display content, thereby achieving the expected visual effect of the subtitle and improving the readability of the to-be-displayed subtitle.

[0168] It can be seen that the subtitle image method provided in the embodiments of the present disclosure can be applied to various scenes.

[0169] For example, in the scene of video processing, the to-be-processed video and the subtitle information of the to-be-processed video are obtained, and then the subtitle information is processed by using the subtitle image method provided in the embodiments of the present disclosure, so as to efficiently generate the target subtitle image of each video frame in the to-be-processed video, and then each target subtitle image is added to the corresponding video frame to obtain a target video, thereby effectively improving the efficiency of adding subtitles to the video.

[0170] For another example, in the scenario of audio processing, the subtitle information of the audio to be processed is acquired, and then the subtitle image method provided by the embodiment of the present disclosure is used to process the subtitle information, so that the target subtitle image of each playback time point in the audio to be processed can be efficiently generated. Then, each target subtitle image can be used as a video frame corresponding to the playback time point, and each target subtitle image is combined in the order of the playback time point to generate a subtitle video, so that the generation efficiency of the subtitle video can be effectively improved, and the subtitle video can provide text information to the group with limited hearing or unable to use audio.

[0171] The complete process of the subtitle image generation method will be described in detail below.

[0172] Referring to FIG. 11, FIG. 11 is a schematic diagram of an optional architecture of the subtitle image generation method provided by the embodiment of the present disclosure.

[0173] Firstly, the subtitle information of the subtitle to be displayed is acquired, and it is determined whether the dynamic attribute is contained in the hierarchical attribute of the subtitle information. When the dynamic attribute is not contained in the hierarchical attribute of the subtitle information, the text layout drawing can be performed according to the subtitle information to obtain a static image, and then the static image is used as the subtitle image of the subtitle to be displayed, and the subtitle image is output.

[0174] Then, when the dynamic attribute is contained in the hierarchical attribute of the subtitle information, the dynamic attribute set of the subtitle to be displayed is acquired from the subtitle information, and the dynamic attribute set is divided into an equivalent attribute set and a non-equivalent attribute set. The equivalent attribute set includes the hierarchical attribute used to realize the animation effect through image transformation, and the non-equivalent attribute set includes the hierarchical attribute that cannot realize the animation effect through image transformation.

[0175] Then, when the non-equivalent attribute set is an empty set, for each image level, the color attribute corresponding to the current image level in the equivalent attribute set is kept unchanged, and the color attributes corresponding to the remaining image levels in the equivalent attribute set are adjusted to be transparent to obtain the hierarchical attribute set corresponding to the current image level.

[0176] Then, the initial attribute value of each image level in the subtitle image corresponding to the subtitle to be displayed is acquired from each hierarchical attribute set, respectively. The image level is divided based on different types of display content in the subtitle image corresponding to the subtitle to be displayed, and the initial attribute value is the attribute value of the hierarchical attribute of the image level at the initial time point. The multiple image levels include the text level, the border level, the shadow level, and the background level, so that the image level of the subtitle to be displayed is split.

[0177] Then, the subtitle text of the to-be-displayed subtitle is acquired, and for each image level, text layout and drawing are performed based on the subtitle text and the initial attribute value to obtain an initial subtitle layer of the image level at the initial time point, thereby achieving initial frame layout and drawing.

[0178] Then, the target time point is acquired, and the target time point is input into each attribute change function in the attribute set of each level to perform calculation to obtain the target attribute value of each image level, thereby achieving dynamic attribute interpolation.

[0179] Then, subtitle image transformation is performed, which can be specifically divided into the following transformation cases.

[0180] When the level attribute includes a color attribute, the target attribute value includes a target color value of the color attribute, and the attribute change amount includes a color change amount of the color attribute, when the color change amount indicates that the image level changes in color, the color value of a non-transparent pixel point in the initial subtitle layer of the image level is transformed into the target color value to obtain a target subtitle layer of the image level at the target time point, wherein the target time point is a time point after the initial time point.

[0181] Or, when the level attribute includes a geometric attribute and the attribute change amount includes a geometric change amount of the geometric attribute, an affine transformation parameter is determined based on the geometric change amount, and an affine transformation matrix is constructed according to the affine transformation parameter; the initial subtitle layer of the image level is subjected to affine transformation based on the affine transformation matrix to obtain the target subtitle layer of the image level at the target time point.

[0182] Or, when the mask attribute includes a color change ratio attribute and the target mask value includes a target color change ratio value of the color change ratio attribute, the layer width of the target subtitle layer is acquired, and a width threshold is determined according to the product of the target color change ratio value and the layer width; in the corresponding initial subtitle layer, a region with an abscissa less than the width threshold is determined as a mask region; in the initial subtitle layer of the image level, the color value of a non-transparent pixel point located in the mask region is transformed into a preset mask color value to obtain the target subtitle layer of the image level at the target time point.

[0183] Or, when the mask attribute includes a plurality of mask coordinate attributes and the target mask value includes a target mask coordinate of each mask coordinate attribute, a visible region is determined in the initial subtitle layer of the image level according to each target mask coordinate, respectively, wherein the target mask coordinate is located at the region boundary of the visible region; in the initial subtitle layer of the image level, a region outside the visible region is determined as a mask region; in the initial subtitle layer of the image level, the color value of a non-transparent pixel point located in the mask region is transformed into a preset mask color value to obtain the target subtitle layer of the image level at the target time point.

[0184] Then, the multiple target subtitle layers are superimposed in sequence according to the background level, the shadow level, the border level and the text level to generate the target subtitle image of the to-be-displayed subtitle at the target time point, and the subtitle layer fusion is realized. Finally, the target subtitle image is output.

[0185] Based on this, the initial attribute value of the level attribute of each image level at the initial time point is obtained, and the subtitle text of the to-be-displayed subtitle is obtained. Since the image level is based on different types of display content in the subtitle image corresponding to the to-be-displayed subtitle, the initial subtitle layer of the image level at the initial time point can be obtained by text layout and drawing based on the subtitle text and the initial attribute value, so as to realize the division of the subtitle layer for different display content. Then, the initial subtitle layer is image-transformed based on the target attribute value of the level attribute at the target time point to obtain the target subtitle layer of the image level at the target time point, and then the multiple target subtitle layers are superimposed to generate the target subtitle image of the to-be-displayed subtitle at the target time point. In the image transformation, the image transformation process of the corresponding initial subtitle layer is controlled by each target attribute value, realizing the fine control of each image level, so as to fine control the different types of display content in the to-be-displayed subtitle and enrich the animation effect of the to-be-displayed subtitle. In addition, for multiple frame images of the to-be-displayed subtitle, the initial subtitle layer can be regarded as a layer in the initial frame image displayed by the to-be-displayed subtitle at the initial time point, and the target subtitle layer can be regarded as a layer in the target frame image displayed by the to-be-displayed subtitle at the target time point. In the subtitle animation rendering process, only text layout and drawing are needed when generating the initial subtitle layer, and no text layout and drawing are needed when generating the target subtitle layer, but the target subtitle layer is obtained by image transformation based on the initial subtitle layer. This can avoid frequent text layout and drawing with large amount of calculation, thereby improving the rendering efficiency of the subtitle animation and enhancing the real-time performance and smoothness of the subtitle animation.

[0186] It can be understood that although each step in each of the above flowcharts is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified in this embodiment, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least part of the steps in the above flowcharts can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0187] Referring to FIG. 12, FIG. 12 is a schematic diagram of an optional structure of a subtitle image generation apparatus provided in an embodiment of the present disclosure. The subtitle image generation apparatus 1200 comprises:

[0188] The acquisition module 1201 is configured to acquire initial attribute values of respective image levels in a subtitle image corresponding to a to-be-displayed subtitle, wherein the image levels are divided based on different types of display content in the subtitle image, and the initial attribute values are attribute values of level attributes of the image levels at an initial time point;

[0189] The layout and drawing module 1202 is configured to acquire subtitle text of the to-be-displayed subtitle, and perform text layout and drawing based on the subtitle text and the initial attribute values of the respective image levels, to obtain initial subtitle layers of the respective image levels at the initial time point.

[0190] The image transformation module 1203 is configured to acquire target attribute values of the respective image levels, and perform image transformation on the initial subtitle layers of the image levels based on the target attribute values of the image levels, to obtain target subtitle layers of the image levels at a target time point, wherein the target attribute values are attribute values of the level attributes of the image levels at the target time point, and the target time point is a time point after the initial time point.

[0191] The image generation module 1204 is configured to superimpose the target subtitle layers of the respective image levels, to generate a target subtitle image of the to-be-displayed subtitle at the target time point.

[0192] Further, the image transformation module 1203 is specifically configured to:

[0193] determine an attribute change amount between the target attribute values of the image levels and the initial attribute values of the image levels;

[0194] perform image transformation on the initial subtitle layers of the image levels based on the attribute change amount, to obtain the target subtitle layers of the image levels at the target time point.

[0195] Further, the level attributes comprise color attributes, the target attribute values comprise target color values of the color attributes, and the attribute change amount comprises a color change amount of the color attributes. The image transformation module 1203 is specifically configured to:

[0196] when the color change amount indicates that the image levels change in color, transform color values of non-transparent pixel points in the initial subtitle layers into the target color values, to obtain the target subtitle layers of the image levels at the target time point.

[0197] Further, the level attributes comprise geometric attributes, and the attribute change amount comprises a geometric change amount of the geometric attributes. The image transformation module 1203 is specifically configured to:

[0198] determine the affine transformation parameters based on the geometric change amount, and construct an affine transformation matrix according to the affine transformation parameters;

[0199] perform affine transformation on the initial subtitle layer based on the affine transformation matrix to obtain the target subtitle layer at the target time point.

[0200] Further, the level attribute includes a mask attribute, and the target attribute value includes a target mask value of the mask attribute. The image transformation module 1203 is specifically configured to:

[0201] determine a mask region in the initial subtitle layer based on the target mask value;

[0202] transform a color value of a non-transparent pixel point located in the mask region in the initial subtitle layer into a preset mask color value to obtain the target subtitle layer at the target time point.

[0203] Further, the mask attribute includes a color change ratio attribute, and the target mask value includes a target color change ratio value of the color change ratio attribute. The image transformation module 1203 is specifically configured to:

[0204] obtain a layer width of the target subtitle layer, and determine a width threshold according to a product of the target color change ratio value and the layer width;

[0205] determine a mask region in the initial subtitle layer according to the width threshold.

[0206] Further, the mask attribute includes a plurality of mask coordinate attributes, and the target mask value includes a target mask coordinate of each mask coordinate attribute. The image transformation module 1203 is specifically configured to:

[0207] determine a visible region in the initial subtitle layer according to each target mask coordinate, wherein the target mask coordinate is located at a region boundary of the visible region;

[0208] determine a mask region in the initial subtitle layer according to a region outside the visible region.

[0209] Further, the obtaining module 1201 is specifically configured to:

[0210] obtain a dynamic attribute set of a subtitle image corresponding to a to-be-displayed subtitle, and divide the dynamic attribute set into an equivalent attribute set and a non-equivalent attribute set, wherein the equivalent attribute set includes a level attribute used to realize an animation effect through image transformation, and the non-equivalent attribute set includes a level attribute that cannot realize an animation effect through image transformation;

[0211] when the non-equivalent attribute set is an empty set, obtain initial attribute values of each image level in the subtitle image corresponding to the to-be-displayed subtitle from the equivalent attribute set.

[0212] Further, the hierarchical attribute in the equivalent attribute set comprises a color attribute of each image hierarchy, and the acquisition module 1201 is specifically configured to:

[0213] When the non-equivalent attribute set is an empty set, for each image hierarchy, the color attribute corresponding to the current image hierarchy in the equivalent attribute set is kept unchanged, and the color attributes corresponding to the remaining image hierarchies in the equivalent attribute set are adjusted to be transparent, to obtain a hierarchical attribute set corresponding to the current image hierarchy.

[0214] The initial attribute values of the image hierarchies in the subtitle image corresponding to the to-be-displayed subtitle are acquired from each hierarchical attribute set respectively.

[0215] Further, the hierarchical attribute set further comprises an attribute change function corresponding to the initial attribute value, and the image transformation module 1203 is specifically configured to:

[0216] Acquire a target time point;

[0217] The target time point is input into the attribute change function in each hierarchical attribute set respectively to perform operation, to obtain target attribute values of the image hierarchies.

[0218] Further, the image transformation module 1203 is specifically configured to:

[0219] The target time point is input into the attribute change function in each hierarchical attribute set respectively to perform operation, to obtain reference attribute values of the image hierarchies;

[0220] The subtitle text is input into the large language model to perform sentiment recognition, to obtain target sentiment information of the subtitle text;

[0221] The target sentiment information and the reference attribute values are spliced and input into the regression model to perform regression, to obtain the target attribute values of the image hierarchies.

[0222] Further, the multiple image hierarchies comprise a text hierarchy, a border hierarchy, a shadow hierarchy and a background hierarchy, and the image generation module 1204 is specifically configured to:

[0223] Each target subtitle layer is sequentially superimposed in the order of the background hierarchy, the shadow hierarchy, the border hierarchy and the text hierarchy, to generate a target subtitle image of the to-be-displayed subtitle at the target time point.

[0224] The electronic device for performing the above-mentioned subtitle image generation method provided by the embodiments of the present disclosure can be a terminal. Referring to FIG. 13, FIG. 13 is a partial structural block diagram of a terminal provided by the embodiments of the present disclosure, which includes a camera assembly 1310, a first storage 1320, an input unit 1330, a display unit 1340, a sensor 1350, an audio circuit 1360, a wireless fidelity (Wi-Fi) module 1370, a first processor 1380, and a first power supply 1390, and the like. Those skilled in the art can understand that the terminal structure shown in FIG. 13 does not constitute a limitation on the terminal, and can include more or fewer components than those shown, or combine certain components, or different component arrangements.

[0225] The camera assembly 1310 can be used to collect images or videos. Optionally, the camera assembly 1310 includes a front camera and a rear camera. Generally, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, the rear camera is at least two, which is any one of a main camera, a depth-of-field camera, a wide-angle camera, and a long-focus camera, to realize the background blur function of the main camera and the depth-of-field camera, the panorama shooting and VR (Virtual Reality) shooting function of the main camera and the wide-angle camera, or other fusion shooting functions.

[0226] The first storage 1320 can be used to store software programs and modules, and the first processor 1380 executes various function applications and data processing of the terminal by running the software programs and modules stored in the first storage 1320.

[0227] The input unit 1330 can be used to receive input digital or character information, and generate key signal input related to the setting and function control of the terminal. Specifically, the input unit 1330 can include a touch panel 1331 and other input devices 1332.

[0228] The display unit 1340 can be used to display input information or provided information and various menus of the terminal. The display unit 1340 can include a display panel 1341.

[0229] The audio circuit 1360, the speaker 1361, and the microphone 1362 can provide an audio interface.

[0230] The first power supply 1390 can be alternating current, direct current, disposable batteries, or rechargeable batteries.

[0231] The number of sensors 1350 can be one or more, which includes but is not limited to an acceleration sensor, a gyroscope sensor, a pressure sensor, an optical sensor, and the like.

[0232] The acceleration sensor can detect the acceleration magnitude in three coordinate axes of a coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration in three coordinate axes. The first processor 1380 can control the display unit 1340 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signals collected by the acceleration sensor. The acceleration sensor can also be used for game or user motion data collection.

[0233] The gyro sensor can detect the body direction and rotation angle of the terminal. The gyro sensor can be used in cooperation with the acceleration sensor to collect the 3D motion of the user to the terminal. The first processor 1380 can implement the following functions according to the data collected by the gyro sensor: motion sensing (e.g., changing the UI according to the user's tilt operation), image stabilization when shooting, game control, and inertial navigation.

[0234] The pressure sensor can be arranged on the side frame of the terminal and / or the lower layer of the display unit 1340. When the pressure sensor is arranged on the side frame of the terminal, the user's holding signal to the terminal can be detected, and the first processor 1380 can perform left-hand or right-hand recognition or shortcut operation according to the holding signal collected by the pressure sensor. When the pressure sensor is arranged on the lower layer of the display unit 1340, the first processor 1380 can control the operable control on the UI interface according to the user's pressure operation on the display unit 1340. The operable control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0235] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 1380 can control the display brightness of the display unit 1340 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1340 is increased; when the ambient light intensity is low, the display brightness of the display unit 1340 is decreased. In another embodiment, the first processor 1380 can also dynamically adjust the shooting parameters of the camera assembly 1310 according to the ambient light intensity collected by the optical sensor.

[0236] In the embodiment, the first processor 1380 included in the terminal can perform the subtitle image generation method of the previous embodiment.

[0237] The electronic device for performing the above-mentioned subtitle image generation method provided by the embodiments of the present disclosure can also be a server. Referring to FIG. 14, FIG. 14 is a partial structure block diagram of a server provided by the embodiments of the present disclosure. The server can have relatively large differences due to different configurations or performances, and can include one or more second processors 1410 and a second memory 1430, one or more storage media 1440 (for example, one or more mass storage devices) storing application programs 1443 or data 1442. Among them, the second memory 1430 and the storage medium 1440 can be temporary storage or persistent storage. The programs stored in the storage medium 1440 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the server. Furthermore, the second processor 1410 can be configured to communicate with the storage medium 1440 and execute the series of instruction operations in the storage medium 1440 on the server.

[0238] The server can also include one or more second power supplies 1420, one or more wired or wireless network interfaces 1450, one or more input and output interfaces 1460, and / or one or more operating systems 1441, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM , etc.

[0239] The second processor 1410 in the server can be configured to perform the subtitle image generation method.

[0240] The embodiments of the present disclosure also provide a computer readable storage medium for storing a computer program, and the computer program is configured to perform the subtitle image generation method of the above-mentioned various embodiments.

[0241] The embodiments of the present disclosure also provide a computer program product, which includes a computer program stored in a computer readable storage medium. The processor of the computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program, so that the computer device performs the subtitle image generation method as described above.

[0242] The terms "first", "second", "third", "fourth", and the like in the description of the disclosure and the above drawings, if any, are used to distinguish similar objects, and do not necessarily have to be used to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0243] It should be understood that in the present disclosure, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases: only A, only B, and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0244] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is two or more, greater than, less than, more than, etc. are not included in the number, above, below, etc. are understood to include the number.

[0245] In several embodiments provided by the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division, and actual implementation can have another division manner. For example, multiple units or components can be combined or integrated into another system, or some features can be omitted or not implemented. In addition, the coupling or direct coupling or communication connection between the displayed or discussed objects can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0246] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e. may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0247] In addition, each functional unit in various embodiments of the present disclosure can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0248] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical scheme of the present disclosure essentially or the part that contributes to the prior art, or all or part of the technical scheme can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the various embodiments of the method of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0249] It should also be understood that various embodiments provided by the present disclosure can be combined in any way to achieve different technical effects.

[0250] The above is a specific description of the preferred implementation of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the present disclosure, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present disclosure.

Claims

1. A subtitle image generation method performed by an electronic device, comprising: obtaining initial attribute values of respective image levels in a subtitle image corresponding to a to-be-displayed subtitle, wherein the image levels are divided based on different types of display content in the subtitle image, and the initial attribute values are attribute values of level attributes of the image levels at an initial time point; obtaining subtitle text of the to-be-displayed subtitle, and for each of the image levels, performing text layout and drawing based on the subtitle text and the initial attribute values of the respective image levels to obtain initial subtitle layers of the respective image levels at the initial time point; obtaining target attribute values of the respective image levels, and for each of the image levels, performing image transformation on the initial subtitle layer of the image level based on the target attribute values of the image level to obtain a target subtitle layer of the image level at a target time point, wherein the target attribute values are attribute values of the level attributes of the image levels at the target time point, and the target time point is a time point after the initial time point; superimposing the target subtitle layers of the respective image levels to generate a target subtitle image of the to-be-displayed subtitle at the target time point.

2. The subtitle image generating method according to claim 1, wherein The image transformation on the initial subtitle layer of the image level based on the target attribute values of the image level to obtain the target subtitle layer of the image level at the target time point comprises: determining an attribute change amount between the target attribute values of the image level and the initial attribute values of the image level; and performing image transformation on the initial subtitle layer of the image level based on the attribute change amount to obtain the target subtitle layer of the image level at the target time point.

3. The subtitle image generating method according to claim 2, wherein The level attributes include color attributes, the target attribute values include target color values of the color attributes, the attribute change amount includes a color change amount of the color attributes, and the image transformation on the initial subtitle layer of the image level based on the attribute change amount to obtain the target subtitle layer of the image level at the target time point comprises: when the color change amount indicates that the image level has a color change, transforming color values of non-transparent pixel points in the initial subtitle layer into the target color values to obtain the target subtitle layer of the image level at the target time point.

4. The subtitle image generating method according to claim 2 or 3, wherein The level attributes include geometric attributes, the attribute change amount includes a geometric change amount of the geometric attributes, and the image transformation on the initial subtitle layer of the image level based on the attribute change amount to obtain the target subtitle layer of the image level at the target time point comprises: determining an affine transformation parameter based on the geometric change amount, and constructing an affine transformation matrix according to the affine transformation parameter; and performing affine transformation on the initial subtitle layer based on the affine transformation matrix to obtain the target subtitle layer of the image level at the target time point.

5. The subtitle image generating method according to any one of claims 1 to 4, wherein The subtitle attribute comprises a mask attribute, the target attribute value comprises a target mask value of the mask attribute, and the image transformation of the initial subtitle layer of the image hierarchy based on the target attribute value of the image hierarchy comprises the following steps: determining a mask region in the initial subtitle layer based on the target mask value; transforming color values of non-transparent pixel points in the mask region in the initial subtitle layer into preset mask color values to obtain the target subtitle layer of the image hierarchy at the target time point.

6. The subtitle image generating method according to claim 5, wherein The mask attribute comprises a color change proportion attribute, and the target mask value comprises a target color change proportion value of the color change proportion attribute. The step of determining a mask region in the initial subtitle layer based on the target mask value comprises the following steps: obtaining a layer width of the target subtitle layer, and determining a width threshold value according to a product of the target color change proportion value and the layer width; determining a region with an abscissa less than the width threshold value as the mask region in the initial subtitle layer.

7. The subtitle image generating method of claim 5, wherein, The mask attribute comprises a plurality of mask coordinate attributes, and the target mask value comprises target mask coordinates of each mask coordinate attribute. The step of determining a mask region in the initial subtitle layer based on the target mask value comprises the following steps: determining a visible region in the initial subtitle layer according to each target mask coordinate, wherein the target mask coordinate is located at a region boundary of the visible region; determining a region outside the visible region as a mask region in the initial subtitle layer.

8. The subtitle image generating method according to any one of claims 1 to 7, wherein The step of obtaining initial attribute values of each image hierarchy in a subtitle image corresponding to the to-be-displayed subtitle comprises the following steps: obtaining a dynamic attribute set of the subtitle image corresponding to the to-be-displayed subtitle, and dividing the dynamic attribute set into an equivalent attribute set and a non-equivalent attribute set, wherein the equivalent attribute set comprises hierarchy attributes used to realize an animation effect through image transformation, and the non-equivalent attribute set comprises hierarchy attributes that cannot realize the animation effect through image transformation; when the non-equivalent attribute set is an empty set, obtaining the initial attribute values of each image hierarchy in the subtitle image corresponding to the to-be-displayed subtitle from the equivalent attribute set.

9. The subtitle image generating method according to claim 8, wherein The hierarchy attributes in the equivalent attribute set comprise color attributes of each image hierarchy. When the non-equivalent attribute set is an empty set, the initial attribute values of each image hierarchy in the subtitle image corresponding to the to-be-displayed subtitle are obtained from the equivalent attribute set, comprising the following steps: when the non-equivalent attribute set is an empty set, for each image hierarchy, keeping the current color attribute corresponding to the image hierarchy unchanged in the equivalent attribute set, adjusting the color attributes corresponding to the remaining image hierarchies in the equivalent attribute set to be transparent to obtain a hierarchy attribute set corresponding to the current image hierarchy; respectively obtaining initial attribute values of the image hierarchy in the subtitle image corresponding to the to-be-displayed subtitle from each hierarchy attribute set.

10. The subtitle image generating method of claim 9, wherein, The hierarchical attribute set also includes an attribute change function corresponding to the initial attribute value, and the obtaining of the target attribute value of each image level includes: obtaining the target time point; inputting the target time point into the attribute change function in each hierarchical attribute set respectively to obtain the target attribute value of each image level.

11. The subtitle image generating method of claim 10, wherein, The inputting of the target time point into the attribute change function in each hierarchical attribute set respectively to obtain the target attribute value of each image level includes: inputting the target time point into the attribute change function in each hierarchical attribute set respectively to obtain a reference attribute value of each image level; inputting the subtitle text into a large language model to perform sentiment recognition to obtain target sentiment information of the subtitle text; concatenating the target sentiment information and the reference attribute value and inputting the same into a regression model to perform regression to obtain the target attribute value of each image level.

12. The subtitle image generating method according to any one of claims 1 to 11, wherein The plurality of image levels include a text level, a border level, a shadow level and a background level, and the superimposition of the target subtitle layers of each image level generates a target subtitle image of the to-be-displayed subtitle at the target time point, including: superimposing each target subtitle layer in sequence according to the order of the background level, the shadow level, the border level and the text level to generate the target subtitle image of the to-be-displayed subtitle at the target time point.

13. A subtitle image generating apparatus, wherein, including: an acquisition module, configured to acquire initial attribute values of each image level in a subtitle image corresponding to to-be-displayed subtitles, wherein the image levels are divided based on different types of display content in the subtitle image, and the initial attribute values are attribute values of hierarchical attributes of the image levels at an initial time point; a layout and drawing module, configured to acquire subtitle text of the to-be-displayed subtitles, and perform text layout and drawing based on the subtitle text and the initial attribute values of each image level to obtain initial subtitle layers of each image level at the initial time point; an image transformation module, configured to acquire target attribute values of each image level, and perform image transformation on the initial subtitle layers of each image level based on the target attribute values of the image level to obtain target subtitle layers of the image level at a target time point, wherein the target attribute values are attribute values of the hierarchical attributes of the image level at the target time point, and the target time point is a time point after the initial time point; an image generation module, configured to superimpose the target subtitle layers of each image level to generate a target subtitle image of the to-be-displayed subtitles at the target time point.

14. An electronic device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the computer program to implement the subtitle image generation method in any one of claims 1 to 12.

15. A computer readable storage medium, the storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to implement the subtitle image generation method in any one of claims 1 to 12.

16. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the subtitle image generation method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method for painting subtitle by using pixel as unit

    CN101552877A

  • Subtitle display method, subtitle display device and terminal

    CN110248255A

  • Subtitle generation method, computer equipment and computer readable storage medium

    CN114697573A

  • Subtitle image generation method and device, electronic equipment and storage medium

    CN118864664A

  • Subtitle processing method, device, equipment and storage medium

    JP2024513380A