Text video generation method, apparatus, electronic device, and storage medium
By dynamically adjusting font size and line spacing based on text length, the method generates high-quality text videos, addressing inefficiencies in conventional platforms and improving user experience.
Patent Information
- Application Number
- JP2025533437
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-12-05
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-12-05
AI Technical Summary
Conventional short video platforms lack the ability to automatically generate text videos with good display effects, as they typically require users to manually create videos from plain text, leading to inefficiencies and poor quality due to fixed font sizes and line spacing that can result in unreadable or incomplete text display.
A method and device that dynamically adjusts font size and line spacing based on the length of the text input, generating a video that accurately displays the entire content, including background images and music, to enhance the visual and auditory experience.
The method improves video generation efficiency and quality by ensuring clear and comprehensive text display, enhancing the user experience and diversity of video content on platforms.
Smart Images

Figure 2026500228000001_ABST
Abstract
Description
[Technical Field]
[0001] [Cross-Citation of Related Applications] This application claims priority to a Chinese invention patent application filed on December 7, 2022, entitled "Text video generation method, apparatus, electronic device, and storage medium," with application number 202211567282.4, the entire contents of which are incorporated herein by reference.
[0002] [Technical field] FIELD OF THE INVENTION The present invention relates to the field of Internet technology, and more particularly to a text-video generating method, apparatus, electronic device, and storage medium. [Background technology]
[0003] Currently, short video platforms are increasingly popular with users due to their rich and diverse content. Short video platform clients have submission portals open to the general public, allowing creators to shoot and upload videos, and the short video platform service pushes the content uploaded by creators to viewers for consumption.
[0004] In conventional technology, short video platforms usually only receive video works created by users and cannot edit or generate text videos. As a result, users have to manually create videos of their plain text works and upload them to the short video platform, which reduces the efficiency and quality of video generation in the video creation process. Summary of the Invention [Problem to be solved by the invention]
[0005] SUMMARY OF THE INVENTION Embodiments of the present invention provide a text video generation method, apparatus, electronic device, and storage medium to overcome the problem of not being able to generate video in the form of plain text. [Means for solving the problem]
[0006] In a first aspect, an embodiment of the present invention provides a text video generation method, the method comprising: displaying a text editing page, the text editing page including a text entry area; displaying target text in the text entry area in response to a first input command on the text editing page, the target text having a first font state in the text entry area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text; and generating a target video to exhibit the target text in the text entry area.
[0007] In a second aspect, an embodiment of the present invention provides a text-video generating device, the device comprising: an editing module for displaying a text editing page, the text editing page including a text entry area; a display module for displaying target text in the text entry area in response to a first input command on a text editing page, the target text having a first font state in the text entry area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text; a generation module for generating a target video for displaying the target text in the text input area.
[0008] In a third aspect, embodiments of the present invention provide an electronic device, the electronic device comprising: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; The processor executes computer-executable instructions stored in the memory to implement the text-video generation method according to the first aspect and various possible designs thereof discussed above.
[0009] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having stored thereon computer-executable instructions, which, when executed by a processor, realizes the text-video generation method as set forth in the first aspect and various possible designs of the first aspect described above.
[0010] In a fifth aspect, an embodiment of the present invention provides a computer program product including a computer program which, when executed by a processor, realizes the text video generation method as set forth in the first aspect and various possible designs of the first aspect described above. [Effects of the Invention]
[0011] The text video generation method, apparatus, electronic device, and storage medium provided in this embodiment display a text editing page including a text input area, and in response to a first input command on the text editing page, display target text in the text input area and generate a target video to display the target text in the text input area, where the target text has a first font state in the text input area, the first font state characterizing a font size and / or line spacing of the target text, and the first font state being determined by the length of the target text. By configuring the text input area and dynamically changing the font size and / or line spacing of the target text edited in the text input area according to the length of the text, and then converting the target text in the text input area, the generated target video can clearly and comprehensively display the entire content of the target text, thereby achieving the purpose of generating a text video based on plain text, and improving the efficiency and quality of video generation in the video creation process. [Brief explanation of the drawings]
[0012] In the following, in order to more clearly explain the embodiments of the present invention or the technical solutions of the prior art, the drawings used in the description of the embodiments or the prior art will be briefly described. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative efforts. [Figure 1] FIG. 1 is an application scenario diagram of the text video generation method provided by an embodiment of the present invention; [Figure 2] 1 is a flow diagram of a text video generation method provided by an embodiment of the present invention; [Figure 3] 1 is a schematic diagram of displaying a target text in a text input area provided by an embodiment of the present invention; [Figure 4] 3 is a flowchart of a specific implementation of step S103 in the embodiment shown in FIG. 2. [Figure 5] FIG. 4 is another flow schematic diagram of a text video generating method provided by an embodiment of the present invention; [Figure 6] 6 is a flowchart of a specific implementation of step S203 in the embodiment shown in FIG. 5. [Figure 7] 3 is a schematic diagram of the area size of a text input area provided by an embodiment of the present invention; FIG. [Figure 8] 7 is a flowchart of a specific implementation of step S2033 in the embodiment shown in FIG. 6. [Figure 9] 3 is a schematic diagram of a background image provided by an embodiment of the present invention; [Figure 10] 6 is a flowchart of a specific implementation of step S205 in the embodiment shown in FIG. 5. [Figure 11] 1 is a block diagram of a text video generating device according to an embodiment of the present invention; [Figure 12] 1 is a schematic diagram of the configuration of an electronic device provided by an embodiment of the present invention; [Figure 13] FIG. 2 is a schematic diagram of the hardware configuration of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] In the following, in order to clarify the objectives, technical solutions and advantages of the embodiments of the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described in conjunction with the drawings of the embodiments of the present invention, and it is obvious that the described embodiments are only some embodiments of the present invention, and not all embodiments, and all other embodiments that can be obtained by those skilled in the art based on the embodiments of the present invention without any creative efforts belong to the protection scope of the present invention.
[0014] It should be noted that all user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are information and data approved by the user or fully approved by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and a corresponding operation portal will be provided to allow users to choose to approve or reject.
[0015] The following describes application scenarios of the embodiments of the present invention.
[0016] FIG. 1 is an application scenario diagram of a text video generation method provided by an embodiment of the present invention. The text video generation method provided by an embodiment of the present invention can be applied to an application scenario in which a client of a short video platform edits and uploads videos. Specifically, as shown in FIG. 1, the method provided by an embodiment of the present invention can be applied to a terminal device such as a smartphone. A short video platform client runs in the terminal device. A text editing page is set up in the client. The terminal device triggers a corresponding component in response to a user operation to enter the text editing page. A text input area is set up in the file editing page, and the user edits text in the text input area using an input method. Specifically, as shown in the drawing, for example, when a user clicks on the text input area, the text input area enters an editing state and a text input cursor is displayed. At the same time, an input method interface pops up in the text editing page. The user operates the input method interface to enter text in the text input area and generate target text (in the drawing, an "X" represents a letter). After completing the target text entry, the input method interface disappears, and the complete text input area is displayed in the text editing page. Then, by clicking the "Done" button, the target text in the text input area will be rendered, and the video material will be generated and uploaded to the service side of the short video platform (shown as the platform server in the drawing), thereby completing the process of generating a video work from a plain text work and uploading it to the short video platform. Meanwhile, by clicking the "Back" button, the text editing page can be exited, and the details will not be repeated.
[0017] In conventional technology, a short video platform client typically receives a user-created video and implements necessary steps such as transcoding and compression before uploading it to a service's video pool. The service then pushes the video to different users based on the videos in the video pool. However, for videos with plain text content, a short video platform client typically does not provide a function page for editing. This is because, in the plain text editing process, the length of the text is determined by the text content edited by the user. Therefore, when the text is displayed with a fixed font size and line spacing, the font and line spacing can be too small or too large. For example, taking font as an example, if the target text edited by a user contains five Chinese characters, the font size is appropriate when displayed with a fixed font size of 3 (example). However, if the target text edited by a user contains 500 Chinese characters, the font size is too large when displayed with a fixed font size of 3, and the entire content of the target text cannot be displayed on the same screen. On the other hand, because text videos are videos that statically display text (i.e., the video always displays the same frame of content), this can lead to a problem of missing content in the text video generated based on the rendering of the target text. Conversely, in the above example, when displayed at a font size of 5 (example), if the target text edited by the user contains five Chinese characters, the font will be too small, resulting in a problem of the font being too small in the text video generated based on the rendering of the target text, which affects the video display effect. Because the font and line spacing cannot be automatically adapted, the generated text video will not have normal text content, making it difficult to achieve the purpose of video display of the text video.Therefore, in the prior art, each short video platform can usually receive text videos created by users, but cannot automatically generate text videos with good display effects.
[0018] The embodiment of the present invention provides a text video generation method to solve the above problems.
[0019] Referring to Figure 2, Figure 2 is a flow diagram of a text video generation method provided by an embodiment of the present invention. The method of this embodiment can be applied to a terminal device, and the text video generation method includes the following steps:
[0020] In step S101, a text editing page is displayed, where the text editing page includes a text input area.
[0021] For example, the schematic diagram of the application scenario shown in FIG. 1 shows that a short video platform client (hereinafter referred to as the client) launches a text editing page in response to a trigger operation. Specifically, the functional component for triggering the text editing page may be set on the client's video playback page (i.e., the default page for short video applications) or on the video shooting page for uploading video works, i.e., the frame page. After receiving a user's trigger operation on the functional component, the text editing page is displayed.
[0022] Furthermore, a text input area is set in the text editing page. For example, after receiving a user's click operation on the text input area, the terminal device activates the text input area to obtain input focus, and then the user can input character information such as text or symbols into the text input area using an input method. Here, the position and size of the text input area can be set as needed and are not limited here.
[0023] In step S102, in response to a first input command on the text editing page, display target text in a text input area, where the target text has a first font state in the text input area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text.
[0024] For example, after the text input area has the input focus, the terminal device receives a first input command input by the user, where the first input command is used to generate text information of the target text, and the first input command may include an identifier corresponding to a specific character or symbol, or an identifier for forming the pronunciation element of the character, such as Chinese pinyin, alphabet, etc., which are not repeated here. After receiving the first input command, the terminal device converts the first input command into corresponding text or symbols according to an input method in the system, and displays it in the text input area to form the target text, which is a text to be finally released.
[0025] Here, the target text may include one or more characters or symbols. Here, symbols input and displayed in the text input area include visible symbols and invisible symbols, and visible symbols include symbols for text editing such as commas and periods, and details will not be repeated. It should be noted that invisible symbols include spaces, blank lines, line breaks, etc. The target text composed of characters and symbols has a paragraph structure, for example, the target text is divided into multiple paragraphs and includes blank lines, etc., thereby making the target text in the text input area visually more readable.
[0026] Furthermore, the target text in the displayed text input area has a first font state, where the first font state characterizes the font size and / or line spacing of the target text, and the first font state is determined by the length of the target text. Specifically, the longer the length of the target text, the smaller the font size of the text and symbols in the target text and / or the smaller the line spacing between lines in the target text. Conversely, the shorter the length of the target text, the larger the font size of the text and symbols in the target text and / or the larger the line spacing between lines in the target text. In short, when the length of the target text is longer, the terminal device compresses the target text in the text input area so that the text input area can contain more characters and symbols. When the length of the target text is shorter, the font is enlarged to make the content of the target text more prominent and improve the visual display effect. Of course, it is understood that there is a certain tolerance for adjusting the first font state (i.e., the font size and / or line spacing of the target text), and that there is a predetermined non-linear correspondence between the first font state and the length of the target text, the specific implementation of which will be described in subsequent examples and will not be repeated here.
[0027] Further, based on the above description of the target text implementation, the length of the target text may be determined by the number of characters in the target text, the total number of characters and symbols in the target text, or the overall occupancy length of the target text in the text input area. For example, if the target text includes a blank line character, one blank line character occupies one blank line in the text input area. Therefore, if the target text includes a blank line, determining the length of the target text based on the overall occupancy length of the target text in the text input area can more accurately measure the actual length of the target text, thereby improving the display effect of the final target video.
[0028] 3 is a schematic diagram of displaying target text in a text input area according to an embodiment of the present invention. As shown in FIG. 3, in this embodiment, the total number of characters in the target text is the length of the target text. The terminal device continuously displays the corresponding target text in the text input area based on a first input command. When the length of the target text is N=10 (time 1), the first font state of the target text is Info_1, and the font size characterizing each character and symbol in the target text is #4. With successive input of the first input command, the length of the target text continues to increase. When the length of the target text is N=50 (time 2), the first font state of the target text is Info_2, and the font size characterizing each character and symbol in the target text is #5 (one size smaller than #4). This achieves the purpose of dynamically displaying the font size of the target text in the text input area.
[0029] Further, based on the above-mentioned embodiment, the first font state further includes information characterizing a line spacing, and the line spacing of the target text can be further adjusted according to the first font state. For example, when the length of the target text N is 10, the first font state of the target text is Info_1, the font size characterizing each character and symbol in the target text is #4 size, and the line spacing is 1; when the length of the target text N is 50, the first font state of the target text is Info_2, the font size characterizing each character and symbol in the target text is #5 size, and the line spacing is 0.8.
[0030] Alternatively, in another possible implementation, the line spacing of the target text can also be determined based solely on the length of the target text. For example, as shown in Figure 3, when the length of the target text N is 10, the first font state of the target text is Info_1 and the line spacing of the target text is characterized as 1, and when the length of the target text N is 50, the first font state of the target text is Info_2 and the line spacing of the target text is characterized as 0.8.
[0031] In step S103, a target video is generated to display the target text in the text input area.
[0032] For example, target text in a text input area has a corresponding first font state, and then the text input area is rendered to generate a video having a predetermined duration, i.e., a target video. Here, since the target video corresponds to a restored display of the target text in the text input area, the target video not only displays the text content of the target text but also restores the first font state of the target text, i.e., the character size and / or line spacing of the target text. As a result, the target text displayed in the target video can visually match the target text displayed in the text input area. Furthermore, during the process of a user editing the text content in the text input area, the display effect of the final target video (i.e., text video) can be accurately predicted, and the video quality of the final target video can be improved. This avoids problems such as text being too small to view or too large to display the entire text in the text video.
[0033] In one possible implementation, as shown in FIG. 4, the specific implementation of step S103 is as follows: Step S1031: generating a rendered image including the target text having a first font state based on the target text; Step S1032, determining a video duration based on the length of the target text; Step S1033 includes generating a target video based on the video duration and the rendered image.
[0034] For example, after inputting a target text in the text input area, the target text in the text input area with a first font state is converted into an image, i.e., a rendering image, based on the target text in the text input area. There are various ways to convert text into an image, such as taking a picture of the text input area to generate a rendering image, or inputting the target text and the corresponding first font state as input parameters into an image converter to generate a corresponding rendering image. The specific steps for converting text into an image are known to those skilled in the art and will not be repeated here.
[0035] Furthermore, a corresponding video duration is determined based on the length of the target text, such as the number of characters in the target text, and the finally generated target video needs to match the user's viewing time when displaying the target text, so the longer the target text, the longer the time required for the user to view the target text in the target video. Therefore, by setting a video duration that matches the length of the target text, the display effect of the target video can be improved. Then, using the rendering image as a material and the video duration as a parameter, video conversion is performed to generate a static video, i.e., the target video.
[0036] In this embodiment, a text editing page including a text input area is displayed, and in response to a first input command on the text editing page, target text is displayed in the text input area, and a target video is generated to display the target text in the text input area, where the target text has a first font state in the text input area, the first font state characterizing the font size and / or line spacing of the target text, and the first font state is determined by the length of the target text. By configuring the text input area and dynamically changing the font size and / or line spacing of the target text edited in the text input area according to the length of the text, and then converting the target text in the text input area, the generated target video can clearly and comprehensively display the entire content of the target text, thereby achieving the purpose of generating text videos based on plain text, improving the video generation efficiency in the video creation process, and increasing the diversity of video content on the video platform.
[0037] 5 is another flow diagram of a text video generating method provided by an embodiment of the present invention. Based on the embodiment shown in FIG. 2, this embodiment further refines step S102 and adds a step of configuring background images and background music for the target video. The text video generating method includes: Step S201: displaying a text editing page, where the text editing page includes a text input area; Step S202: generating a target text in response to a first input command and obtaining a total number of characters in the target text; Step S203 includes determining a first font state based on the total number of characters and the area size of the text input area.
[0038] For example, an editing page can be displayed, and in response to a first input command on the editing page, corresponding characters and symbols can be generated based on the information in the first input command, and a target text can be further generated. Specific implementation steps have been described in detail in the embodiment shown in Fig. 2 and will not be repeated here. Then, the total number of characters in the target text can be obtained by calling a statistical function to process the string corresponding to the target text. The total number of characters in this embodiment may be only the number of characters in the target text, or the sum of the number of characters and the number of symbols.
[0039] For example, the text input area of the text editing page has an area size. For example, if the text input area is rectangular, the area size of the text input area may be the length and width of the text input area. Furthermore, the area size of the text input area may indicate the area of the text input area, and the larger the area of the text input area, the greater the total number of characters that can be displayed. The area size of the text input area is a predetermined fixed value and can be determined according to the screen pixel size of the terminal device, and will not be repeated here.
[0040] Furthermore, after determining the total number of characters and the area size of the text input area, first calculate the total area of the text input area based on the area size of the text input area, then use the ratio between the total area and the total number of characters to obtain the unit area occupied by the font size (and symbol), and further determine the font size and / or corresponding line spacing, i.e., the first font state, based on the unit area and a predetermined correction coefficient. The above-mentioned implementation is suitable for roughly calculating the first font state because it does not take into account the influence of blank areas due to line breaks and blank lines.
[0041] In another possible implementation, the area size includes an area horizontal size and an area vertical size, and as shown in FIG. 6, the specific implementation of step S203 is as follows: Step S2031: determining a number of characters per line based on a font width corresponding to a base font size and an area horizontal size, the number of characters per line characterizing the number of characters that can be displayed in one line of the text input area; Step S2032: determining a first vertical size based on the number of characters per line and the total number of characters; Step S2033 includes determining a first font state based on the first vertical size and the area vertical size.
[0042] FIG. 7 is a schematic diagram of the area size of a text input area provided by an embodiment of the present invention. As shown in FIG. 7, the text input area in the client's text editing page is a vertically oriented rectangular area (suitable for a smartphone display screen). The horizontal size of the area is illustrated as dim_x, and the vertical size of the area is illustrated as dim_y. The user inputs target text into the text input area. The base font size is the font size of a given input character and has a corresponding font width and font height. For example, if the font is a Chinese font, the font width and font height of the text may be the same. The number of characters per line, i.e., the number of characters that can be displayed per line in the text input area, can be determined based on the proportional value between the horizontal size of the area and the font width corresponding to the base font size. For example, 20 characters. The total number of lines can then be obtained by dividing the total number of characters in the target text by the number of characters per line. The average line height is obtained based on the font height and base line spacing, and the product of the average line height and the total number of lines can be calculated to obtain the first vertical size, indicated as dim_Y in the drawing. Here, the base line spacing may be a predetermined value, similar to the base font size, and will not be repeated here with examples.
[0043] Furthermore, after determining the first vertical size, the first vertical size is compared with the area vertical size. If the first vertical size is larger than the area vertical size, it indicates that the font size and / or line spacing are larger at this time, resulting in the text input area being exceeded, making it impossible to display all the content of the target text in a static display manner in the target video to be generated subsequently. In this case, the base font size and / or base line spacing, which affect the first vertical size, are reduced to obtain the first font state. Specifically, for example, the base font size and / or base line spacing are multiplied by a proportional coefficient less than 1 to obtain the first font state. In another case, if the first vertical size is smaller than the area vertical size, it indicates that the font size and / or line spacing are smaller at this time, resulting in the target text not being reasonably laid out in the text input area, affecting the display effect of the target video to be generated subsequently. In this case, the base font size and / or base line spacing, which affect the first vertical size, are increased to obtain the first font state. Specifically, for example, the base font size and / or the base line spacing is multiplied by a proportional coefficient greater than 1 to obtain the first font state.
[0044] Here, for example, the proportional relationship between the first vertical size and the vertical size of the area can be determined by adjusting the base font size and / or the proportional coefficient between the base lines. As shown in FIG. 8, for example, a specific implementation of step S2033 is as follows: Step S2033A: obtaining a proportional value between the first vertical size and the area vertical size; Step S2033B, if the proportional value is less than the first proportional threshold, determining a first font state based on a base font size and / or a base line spacing; Step S2033C includes: if the proportional value is greater than the first proportional threshold, reducing the base font size and / or reducing the base line spacing according to the proportional value to obtain a first font state.
[0045] For example, after determining the first vertical size and the area vertical size, a proportional value between the first vertical size and the area vertical size is calculated. For example, if the first vertical size is 8 (a predetermined unit) and the area vertical size is 10, the proportional value between the first vertical size and the area vertical size is 0.8. The first proportional threshold characterizes the proportion of the text input area, i.e., the screen proportion after generating the target video. When the first proportional threshold is 1, it characterizes the entire area of the text input area, and when the first proportional threshold is 0.8, it characterizes 80% of the area (the ideal display area) of the text input area. Furthermore, in a possible case, if the proportional value is smaller than the first proportional threshold, it indicates that the target text does not exceed the text input area, and the generated target video will not be unable to display all the target text. At the same time, the base font size and / or the base line spacing can be considered ideal parameters set according to specific needs, and therefore the base font size and / or the base line spacing can be set to the first font state.
[0046] In another possible case, if the proportional value is greater than the first proportional threshold, it indicates that the target text exceeds the text input area or the ideal display area of the text input area, which may result in the generated target video being unable to display all of the target text. In this case, the base font size or the base line spacing is adjusted based on the proportional value, where the first proportional threshold is, for example, 1. Specifically, for example, if the proportional value is greater than or equal to 0.8 and less than 1, the font size is reduced by one level, and if the proportional value is greater than or equal to 0.6 and less than 0.8, the font size is reduced by two levels. By analogy in this way, the target font size and / or the target line spacing obtained by reducing the base font size and / or the base line spacing are determined as the first font state.
[0047] In step S204, display the target text in the text entry area based on the first font state.
[0048] In this embodiment, a proportional value between the first vertical size and the vertical size of the area is calculated, and the corresponding first font state is dynamically determined by reducing the proportional value according to the base font size and the base line spacing, thereby more accurately controlling the length of the target text in the text input area, improving the content length of the final generated target video, and improving the display effect of the text video and the user's viewing experience.
[0049] In step S205, in response to a second input command to the text edit page, a background image is displayed within the text edit page.
[0050] For example, based on the above steps, a component for configuring a background image is further configured on the text editing page. A user can trigger the component through an operation command to display the background image on the text editing page. Specifically, the display area corresponding to the background image is a background image area, which covers the text input area. FIG. 9 is a schematic diagram of a background image provided by an embodiment of the present invention. As shown in FIG. 9, after receiving a user's click on the "Add Background" control, the terminal device generates and executes a second input command to automatically add the background image to the background image area on the file editing page, where the background image area covers the text input area and the text input area is located at the center of the background image. In a subsequent step, the background image area is rendered as a target area to generate a corresponding target video. The target text in the generated target video is then displayed on the background image, improving the display effect of the target text in the target video.
[0051] Here, for example, the target text further has a second font state, and the second font state characterizes the font color of the target text. As shown in FIG. 10, the specific implementation of step S205 is as follows: step S2051, in response to a second input command, matching a target color based on a second font state, where a color difference between the target color and the font color is greater than a color difference threshold; Step S2052 includes obtaining a background image whose main color is the target color based on the predetermined gallery, and displaying the background image in the text editing page.
[0052] For example, after receiving the second input command, the second font state of the target text, i.e., the font color of the target text, is first obtained. The second font state may be set by an operation command input by a user. The specific implementation is conventional technology, and details will not be repeated. Then, based on the font color characterized by the second font state, another color, i.e., the target color, whose color difference is greater than the color difference threshold is determined. Simply put, a color with a high font color difference of the target text is determined as the target color. For example, if the font of the target text is black, the target color may be white, light green, light blue, etc. (colors whose color difference from black is greater than the threshold). Then, a main color is obtained from a predetermined gallery as a background image of the target color, thereby ensuring that the background image has a higher contrast with the font color of the target text. This avoids the problem of the target text being difficult to distinguish due to the similar colors of the background image and the target text, and improves the display clarity of the target text in the target video.
[0053] In step S206, background music is obtained based on the background image.
[0054] Furthermore, after obtaining a background image, background music that matches the background image is automatically selected based on the background image. For example, if the content of the background image is "landscape," a song tag for music with a matching soothing rhythm is obtained based on the semantic information of the background image. If the content of the background image is "people dancing on a dance floor," a song tag for music with a matching dynamic rhythm is obtained based on the semantic information of the background image. Then, a corresponding target piece of music is selected as background music from a music library based on the song tag. Here, a specific implementation of obtaining the semantic information of a background image and obtaining the corresponding song tag based on the semantic information can be realized using a pre-trained image semantic model, so it will not be repeated here.
[0055] In the steps of this embodiment, by obtaining background music that matches the background image, the background image and the background music have semantic consistency, the video expression of the final generated target video is better, and the quality of the text video is improved.
[0056] In step S207, a target video is generated based on the target text in the text input area, the background image, and the background music.
[0057] The process of obtaining a target text, background image, and background music, and performing rendering based on the target text and background image to obtain a corresponding rendered image has already been described in the embodiment shown in Fig. 2, and will not be repeated here. Then, the rendered image and the background music are merged and converted to obtain a target video. Here, for example, the video duration of the target video in this embodiment may be determined by the music duration of the background music, and the specific implementation steps have already been described in the embodiment shown in Fig. 2, and will not be repeated here.
[0058] Of course, in other possible embodiments, after obtaining the target text and the corresponding first font state, only the background image or background music corresponding to the target text can be obtained, and then the target video can be generated based on the target text and the background image, or the target video can be generated based on the target text and the background music, where it is understood that the background music is determined based on the semantic information of the target text. The specific implementation process is the same as the implementation process in the above-mentioned embodiment steps, and the details will not be repeated.
[0059] Corresponding to the text-video generating method of the above embodiment, Fig. 11 is a block diagram of the text-video generating device provided by the embodiment of the present invention. For convenience of explanation, only the parts related to the embodiment of the present invention are shown. Referring to Fig. 11, the text-video generating device 3 includes: an editing module 31 for displaying a text editing page, the text editing page including a text input area; a display module 32 for displaying target text in a text entry area in response to a first input command on the text editing page, the target text having a first font state in the text entry area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text; A generation module 33 for generating a target video for displaying the target text in the text input area is provided.
[0060] In one embodiment of the present invention, the display module 32 specifically generates a target text in response to a first input command, obtains the total number of characters in the target text, determines a first font state based on the total number of characters and the area size of the text input area, and displays the target text in the text input area based on the first font state.
[0061] In one embodiment of the present invention, the area size includes an area horizontal size and an area vertical size. When determining the first font state based on the total number of characters and the area size of the text input area, the display module 32 specifically determines the number of characters per line based on a font width corresponding to the base font size and the area horizontal size, where the number of characters per line characterizes the number of characters that can be displayed in one line of the text input area. The display module 32 determines a first vertical size based on the number of characters per line and the total number of characters, and determines the first font state based on the first vertical size and the area vertical size.
[0062] In one embodiment of the present invention, when determining the first font state based on the first vertical size and the area vertical size, the display module 32 specifically obtains a proportional value between the first vertical size and the area vertical size, and if the proportional value is smaller than a first proportional threshold, determines the first font state based on the base font size and / or the base line spacing; and if the proportional value is larger than the first proportional threshold, reduces the base font size and / or the base line spacing based on the proportional value to obtain the first font state.
[0063] In one embodiment of the present invention, the generation module 33 specifically generates a rendering image including a target text having a first font state based on the target text, determines a video duration based on the length of the target text, and generates a target video based on the video duration and the rendering image.
[0064] In one embodiment of the present invention, before generating a target video for displaying the target text in the text input area, the display module 32 further displays a background image in the text editing page in response to a second input command for the text editing page, and the generation module 33 specifically generates the target video based on the target text in the text input area and the background image.
[0065] In one embodiment of the present invention, the target text further has a second font state, which characterizes the font color of the target text. When the display module 32 displays a background image within the text editing page in response to a second input command for the text editing page, the display module 32 specifically matches the target color based on the second font state in response to the second input command, where the color difference between the target color and the font color is greater than a color difference threshold. The display module 32 obtains a background image whose main color is the target color based on a predetermined gallery, and displays the background image within the text editing page.
[0066] In one embodiment of the present invention, the generation module 33 further obtains background music based on the background image, and when generating a target video for displaying the target text in the text input area, the generation module 33 specifically generates the target video based on the target text in the text input area and the background music.
[0067] Here, the editing module 31, the display module 32, and the generating module 33 are connected in sequence. The text video generating device 3 provided by this embodiment can implement the technical solutions of the above-mentioned method embodiments, and its realization principles and technical effects are similar, so this embodiment will not be repeated here.
[0068] FIG. 12 is a schematic diagram of an electronic device provided by an embodiment of the present invention. As shown in FIG. 12, the electronic device 4 includes: A processor 41 and a memory 42 communicatively connected to the processor 41, The memory 42 stores computer-executable instructions, The processor 41 executes computer-executable instructions stored in the memory 42 to implement the text video generating method in the embodiment shown in FIGS.
[0069] Here, the processor 41 and the memory 42 are optionally connected via a bus 43 .
[0070] The relevant explanations can be understood by referring to the relevant explanations and effects corresponding to the steps in the embodiments corresponding to FIGS. 2 to 10, and will not be repeated here.
[0071] An embodiment of the present invention provides a computer-readable storage medium having computer-executable instructions stored therein, which, when executed by a processor, realizes the text video generation method provided by any of the embodiments corresponding to Figures 2 to 10 of the present invention.
[0072] 13 is a schematic diagram of the configuration of an electronic device provided by an embodiment of the present invention. As shown in FIG. 13, the electronic device 1300 may be a terminal device or a server. Here, the terminal device may be a mobile phone, a notebook computer, a digital broadcast receiver, a personal digital assistant (PDA), a tablet computer (Portable Digital Assistant), a The electronic devices may include, but are not limited to, mobile terminals such as an Android device (PAD), a portable media player (PMP), an in-vehicle terminal (e.g., an in-vehicle navigation terminal), and fixed terminals such as a digital TV and a desktop computer. The electronic devices shown in FIG. 13 are merely examples and do not limit the functionality and scope of use of the embodiments of the present invention.
[0073] As shown in FIG. 13, the electronic device 1300 includes a read-only memory (Read Only Memory). The program stored in the ROM 1302 or the storage device 1308 is read from the Random Access Memory (Random Access Memory (ROM)). The electronic device 1300 may include a processing unit (e.g., a central processing unit, a graphical processor, etc.) 1301 that can perform various appropriate operations and processes based on programs loaded into a ROM (Read Only Memory, RAM) 1303. The RAM 1303 further stores various programs and data necessary for the operation of the electronic device 1300. The processing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0074] Input devices 1306, typically including touch screens, touch pads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc., liquid crystal displays (LCDs), Devices such as output device 1307 including a Crystal Display (LCD), speaker, vibrator, etc., storage device 1308 including a magnetic tape, hard disk, etc., and communication device 1309 may be connected to I / O interface 1305. Communication device 1309 may enable electronic device 1300 to communicate wirelessly or via wires to exchange data with other devices. While the figures show electronic device 1300 with various devices, it should be understood that it need not implement or include all of the shown devices. More or fewer devices may alternatively be implemented or included.
[0075] In particular, according to embodiments of the present invention, the processes described with reference to the flowcharts above may be implemented as a computer software program. For example, embodiments of the present invention include a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the methods illustrated in the flowcharts. In such embodiments, the computer program may be downloaded and installed from a network via the communication device 1309, or may be installed from the storage device 1308 or the ROM 1302. When the computer program is executed by the processing device 1301, the functions described above, which are specific to the methods of the embodiments of the present invention, are performed.
[0076] It should be noted that the above-mentioned computer-readable medium of the present invention may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, for example, but not limited to, an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium include an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (ERAM), a programmable write-only memory (PGWM), a programmable write-only memory (PGW ... Memory, EPROM or Flash Memory), Fiber Optics, Portable Compact Disc Read-Only Memory This may include, but is not limited to, a computer-readable storage medium (e.g., a memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present invention, a computer-readable storage medium may be a tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. Also, in the present invention, a computer-readable signal medium may include a propagated data signal in baseband or as part of a carrier carrying computer-readable program code. Such a propagated data signal may take various forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium, other than a computer-readable storage medium, that transmits, propagates, or transmits a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium may be transmitted using any suitable medium, including, but not limited to, wire, fiber optic cable, RF (radio frequency), or any suitable combination of the foregoing.
[0077] The computer readable medium described above may be included in the electronic device described above, or may exist separately and not be assembled to said electronic device.
[0078] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method described in the above-described embodiment.
[0079] Computer program code for carrying out the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may run entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. When referring to a remote computer, the remote computer may be connected to a local area network (LAN) or a server. Network (LAN) or Wide Area Network (Wide The user computer may be connected to the network via any type of network, including a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0080] The flowcharts and block diagrams in the figures illustrate possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. From this perspective, each box in a flowchart or block diagram contains one or more executable instructions for implementing the specified logical function(s) and may represent a module, program segment, or portion of code. It should also be noted that in some alternative implementations, the functions depicted in the boxes may occur in a different order than that depicted in the figures. For example, two boxes shown in succession may actually be executed substantially in parallel, or may be executed in the reverse order, depending on the functionality involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs the specified function(s) or operation(s), or in a combination of dedicated hardware and computer instructions.
[0081] The units described as being involved in the embodiments of the present invention may be implemented by software or hardware. The names of the units do not constitute limitations on the units themselves in a particular case. For example, a first acquisition unit may be described as "a unit for acquiring at least two Internet Protocol addresses."
[0082] The functions described herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0083] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program that can be used by or in combination with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0084] In a first aspect, according to one or more embodiments of the present invention, there is provided a text video generation method, said method comprising: displaying a text editing page, the text editing page including a text entry area; displaying target text in the text entry area in response to a first input command on the text editing page, the target text having a first font state in the text entry area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text; and generating a target video to exhibit the target text in the text entry area.
[0085] According to one or more embodiments of the present invention, displaying target text in the text input area in response to a first input command for the text editing page includes generating target text in response to the first input command and obtaining a total number of characters in the target text, determining the first font state based on the total number of characters and an area size of the text input area, and displaying the target text in the text input area based on the first font state.
[0086] According to one or more embodiments of the present invention, the area size includes an area horizontal size and an area vertical size, and determining the first font state based on the total number of characters and the area size of the text input area includes determining a number of characters per line based on a font width corresponding to a base font size and the area horizontal size, where the number of characters per line characterizes a number of characters that can be displayed in one line of the text input area; determining a first vertical size based on the number of characters per line and the total number of characters; and determining the first font state based on the first vertical size and the area vertical size.
[0087] According to one or more embodiments of the present invention, determining the first font state based on the first vertical size and the area vertical size includes: obtaining a proportional value between the first vertical size and the area vertical size; if the proportional value is smaller than a first proportionality threshold, determining the first font state based on the base font size and / or a base line spacing; and if the proportional value is greater than the first proportionality threshold, reducing the base font size and / or reducing the base line spacing based on the proportional value to obtain the first font state.
[0088] According to one or more embodiments of the present invention, generating a target video for displaying target text in the text input area includes generating a rendered image including the target text having the first font state based on the target text, determining a video duration based on a length of the target text, and generating the target video based on the video duration and the rendered image.
[0089] According to one or more embodiments of the present invention, the method further includes displaying a background image on the text editing page in response to a second input command to the text editing page before generating a target video for displaying the target text in the text input area, and generating a target video for displaying the target text in the text input area includes generating the target video based on the target text in the text input area and the background image.
[0090] According to one or more embodiments of the present invention, the target text further has a second font state, the second font state characterizing a font color of the target text, and displaying a background image on the text editing page in response to a second input command to the text editing page includes: matching a target color based on the second font state in response to the second input command, wherein a color difference between the target color and the font color is greater than a color difference threshold; obtaining a background image whose main color is the target color based on a predetermined gallery; and displaying the background image on the text editing page.
[0091] According to one or more embodiments of the present invention, the method further includes obtaining background music based on the background image, and generating a target video for displaying the target text in the text input area includes generating the target video based on the target text in the text input area and the background music.
[0092] In a second aspect, according to one or more embodiments of the present invention, there is provided a text-video generation device, said device comprising: an editing module for displaying a text editing page, the text editing page including a text entry area; a display module for displaying target text in the text entry area in response to a first input command on a text editing page, the target text having a first font state in the text entry area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text; a generation module for generating a target video for displaying the target text in the text input area.
[0093] According to one or more embodiments of the present invention, the display module specifically generates a target text in response to the first input command, obtains a total number of characters in the target text, determines the first font state based on the total number of characters and an area size of the text input area, and displays the target text in the text input area based on the first font state.
[0094] According to one or more embodiments of the present invention, the area size includes an area horizontal size and an area vertical size. When determining the first font state based on the total number of characters and the area size of the text input area, the display module specifically determines the number of characters in one line based on a font width corresponding to a base font size and the area horizontal size, where the number of characters in one line characterizes the number of characters that can be displayed in one line of the text input area. The display module also determines a first vertical size based on the number of characters in one line and the total number of characters, and determines the first font state based on the first vertical size and the area vertical size.
[0095] According to one or more embodiments of the present invention, when determining the first font state based on the first vertical size and the area vertical size, the display module specifically obtains a proportional value between the first vertical size and the area vertical size, and if the proportional value is smaller than a first proportional threshold, determines the first font state based on the base font size and / or base line spacing; and if the proportional value is larger than the first proportional threshold, reduces the base font size and / or reduces the base line spacing based on the proportional value to obtain the first font state.
[0096] According to one or more embodiments of the present invention, the generation module specifically generates a rendered image including a target text having the first font state based on the target text, determines a video duration based on a length of the target text, and generates the target video based on the video duration and the rendered image.
[0097] According to one or more embodiments of the present invention, before generating a target video for displaying the target text in the text input area, the display module further displays a background image on the text editing page in response to a second input command for the text editing page, and the generation module specifically generates the target video based on the target text in the text input area and the background image.
[0098] According to one or more embodiments of the present invention, the target text further has a second font state, and the second font state characterizes a font color of the target text. When the display module displays a background image on the text editing page in response to a second input command for the text editing page, the display module specifically matches a target color based on the second font state in response to the second input command, where a color difference between the target color and the font color is greater than a color difference threshold; and obtains a background image whose main color is the target color based on a predetermined gallery, and displays the background image on the text editing page.
[0099] According to one or more embodiments of the present invention, the generation module further obtains background music based on the background image, and when generating a target video for displaying the target text in the text input area, the generation module specifically generates the target video based on the target text in the text input area and the background music.
[0100] In a third aspect, according to one or more embodiments of the present invention, there is provided an electronic device, the electronic device comprising: a processor; and a memory communicatively coupled to the processor; the memory stores computer-executable instructions; The processor executes computer-executable instructions stored in the memory to implement the text-video generation method according to the first aspect and various possible designs thereof discussed above.
[0101] In a fourth aspect, according to one or more embodiments of the present invention, there is provided a computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, result in the text-video generation method described in the first aspect and various possible designs of the first aspect discussed above.
[0102] In a fifth aspect, an embodiment of the present invention provides a computer program product including a computer program which, when executed by a processor, realizes the text video generation method as set forth in the first aspect and various possible designs of the first aspect described above.
[0103] The above description is merely a preferred embodiment of the present invention and merely describes the applied technical principles. Those skilled in the art should understand that the scope of the present invention is not limited to the technical solution formed by a specific combination of the above technical features, but also includes other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the concept of the present invention. For example, it should be understood that the present invention also includes technical solutions formed by substituting the above features with technical features having similar functions disclosed in the present invention (but not limited to these).
[0104] Also, although operations are depicted in a particular order, this should not be understood as requiring that these operations be performed in the particular order or sequence shown. Multi-task and parallel processing may be advantageous in certain environments. Similarly, although the above discussion includes some specific implementation details, these should not be construed as limiting the scope of the invention. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.
[0105] Although the present subject matter has been described in language specific to structural features and / or logical operations of a method, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the specific features and operations described above are merely example forms of implementing the claims.
Claims
1. displaying a text editing page, the text editing page including a text entry area; displaying target text in the text entry area in response to a first input command on a text editing page, the target text having a first font state in the text entry area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text; generating a target video to display the target text in the text input area. A text video generation method comprising:
2. Displaying target text in the text entry area in response to a first input command on the text editing page includes: generating a target text in response to the first input command and obtaining a total number of characters in the target text; determining the first font state based on the total number of characters and an area size of the text input area; displaying the target text in the text entry area based on the first font state.
2. The text-video generating method according to claim 1.
3. The area size includes an area horizontal size and an area vertical size, and determining the first font state based on the total number of characters and the area size of the text input area includes: determining a number of characters per line based on a font width corresponding to a base font size and a horizontal size of the area, the number of characters per line characterizing the number of characters that can be displayed in one line of the text input area; determining a first vertical size based on the number of characters per line and the total number of characters; determining the first font state based on the first vertical size and the area vertical size; 3. The text-video generating method according to claim 2.
4. Determining the first font state based on the first vertical size and the area vertical size includes: Obtaining a proportional value between the first vertical size and the area vertical size; if the proportional value is less than a first proportional threshold, determining the first font state based on the base font size and / or base line spacing; If the proportional value is greater than a first proportional threshold, reducing the base font size and / or reducing a base line spacing based on the proportional value to obtain the first font state.
4. The text-video generating method according to claim 3.
5. generating a target video for displaying the target text in the text input area; generating a rendered image based on the target text, the rendered image including the target text having the first font state; determining a video duration based on the length of the target text; generating the target video based on the video duration and the rendered image.
2. The text-video generating method according to claim 1.
6. Before generating a target video for displaying the target text in the text input area, and displaying a background image on the text editing page in response to a second input command on the text editing page; generating a target video for displaying the target text in the text input area; generating the target video based on the target text in the text input area and the background image; 2. The text-video generating method according to claim 1.
7. the target text further has a second font state, the second font state characterizing a font color of the target text; displaying a background image on the text editing page in response to a second input command on the text editing page, In response to the second input command, matching a target color based on the second font state, wherein a color difference between the target color and the font color is greater than a color difference threshold; obtaining a background image whose main color is the target color based on a predetermined gallery, and displaying the background image on the text editing page; 7. The text-video generating method according to claim 6.
8. further comprising obtaining background music based on the background image; generating a target video for displaying the target text in the text input area; generating the target video based on the target text in the text input area and the background music; 7. The text-video generating method according to claim 6.
9. an editing module for displaying a text editing page, the text editing page including a text entry area; a display module for displaying target text in the text entry area in response to a first input command to a text editing page, the target text having a first font state in the text entry area, the first font state characterizing a font size and / or line spacing of the target text, the first font state being determined by a length of the target text; a generation module for generating a target video for displaying the target text in the text input area; A text video generation device characterized by:
10. a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; The processor implements the text-video generation method of any one of claims 1 to 8 by executing computer-executable instructions stored in the memory. An electronic device characterized by:
11. A method for generating text-to-video according to any one of claims 1 to 8, wherein computer-executable instructions are stored and, when executed by a processor, the method is implemented. A computer-readable storage medium comprising:
12. A computer program, which, when executed by a processor, implements the text-video generation method according to any one of claims 1 to 8.
1. A computer program product comprising:
Citation Information
Patent Citations
Picture rendering method and device, equipment, storage medium, and program product
CN113989396A
Document processor, and computer-readable recording medium where document processing program is recorded
JP1999096157A
Method and apparatus for adjusting the length of a text string to fit the display size.
JP2012511218A
Devices, methods, and graphical user interfaces for document manipulation
JP2015230732A
Mobile video editing and sharing for social media
US20150139615A1