Picture rendering method, apparatus, device, storage medium and program product

By determining the text area, text type, and pattern type of an image, the image is rendered automatically, solving the inefficiency problem of manual image rendering in existing technologies and achieving fast and beautiful image rendering effects.

CN113989396BActive Publication Date: 2025-12-16DOUYIN VISION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111308496.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-05
Publication Date
2025-12-16
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

In existing technologies, the rendering of recommended images on video websites or applications requires manual operation, which lacks automation and efficiency.

Method used

The text area is determined by processing the image to be rendered. Based on the attribute information of the text area, the target text type and pattern type are determined, and the image is rendered according to these types.

Benefits of technology

It enables fast image rendering and automatically adds given text to images in a harmonious and aesthetically pleasing manner, thus improving rendering efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989396B_ABST
    Figure CN113989396B_ABST
Patent Text Reader

Abstract

The embodiment of the disclosure discloses a picture rendering method, device, equipment, storage medium and program product, the method comprises: processing a picture to be rendered to determine a text region; determining a text target character type based on attribute information of the text region; determining a text target pattern type based on the picture to be rendered; rendering the picture to be rendered based on the text target character type and the text target pattern type. The embodiment of the disclosure determines the text character type based on the obtained text region, determines the text pattern type based on the picture to be rendered, and adds the text in the text region on the picture after rendering the text according to the character type and the pattern type, that is, places the given character and harmonious beauty in the picture, and realizes the rapid rendering of the picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to an image rendering method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the advancement of science and technology, video technology has become increasingly mature. Common video websites and applications recommend videos by displaying suggested images to users.

[0003] However, in related technologies, the recommended images displayed to users all require manual rendering by post-production staff. Summary of the Invention

[0004] To solve the above-mentioned technical problems, or at least partially solve them, embodiments of this disclosure provide an image rendering method, apparatus, device, storage medium, and program product that harmoniously and aesthetically places given text in an image, thereby achieving rapid image rendering.

[0005] In a first aspect, embodiments of this disclosure provide an image rendering method, the method comprising:

[0006] Process the image to be rendered to define the text area;

[0007] The target text type is determined based on the attribute information of the text region;

[0008] The text target pattern type is determined based on the image to be rendered;

[0009] The image to be rendered is rendered based on the target text type and the target pattern type.

[0010] Secondly, embodiments of this disclosure provide an image rendering apparatus, the apparatus comprising:

[0011] The text region determination module is used to process the image to be rendered and determine the text region.

[0012] The target font size determination module is used to determine the target text type based on the attribute information of the text region.

[0013] The target color determination module is used to determine the text target pattern type based on the image to be rendered.

[0014] The rendering module is used to render the image to be rendered based on the target text type and the target text pattern type.

[0015] Thirdly, embodiments of this disclosure provide an electronic device, the electronic device comprising:

[0016] One or more processors;

[0017] Storage device for storing one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the image rendering method as described in any of the first aspects above.

[0019] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the image rendering method as described in any one of the first aspects above.

[0020] Fifthly, embodiments of this disclosure provide a computer program product comprising a computer program or instructions that, when executed by a processor, implement the image rendering method as described in any of the first aspects above.

[0021] This disclosure provides an image rendering method, apparatus, device, storage medium, and program product. The method includes: processing an image to be rendered to determine a text region; determining a target text type based on attribute information of the text region; determining a target text pattern type based on the image to be rendered; and rendering the image based on the target text type and the target text pattern type. This disclosure achieves rapid image rendering by determining the text type based on the obtained text region, determining the text pattern type based on the image to be rendered, rendering the text according to the text type and text pattern type, and then adding the rendered text to the text region on the image. This results in the harmonious and aesthetically pleasing placement of given text within the image. Attached Figure Description

[0022] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0023] Figure 1 This is a flowchart of an image rendering method according to an embodiment of the present disclosure;

[0024] Figure 2 This is a flowchart of an image rendering method according to an embodiment of the present disclosure;

[0025] Figure 3 This is a schematic diagram of the text region in the image to be rendered provided in the embodiments of this disclosure;

[0026] Figure 4This is a schematic diagram of the text color candidate set provided in an embodiment of this disclosure;

[0027] Figure 5 This is a schematic diagram of a rendered image provided in an embodiment of this disclosure;

[0028] Figure 6 This is a schematic diagram of the structure of an image rendering device according to an embodiment of the present disclosure;

[0029] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation

[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0036] The image rendering method proposed in the embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0037] Figure 1 This is a flowchart of an image rendering method according to an embodiment of the present disclosure. This embodiment can be applied to adding text effects to any image. The method can be executed by an image rendering device, which can be implemented in software and / or hardware and can be configured in an electronic device.

[0038] For example, the electronic device may be a mobile terminal, a fixed terminal, or a portable terminal, such as a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.

[0039] For example, the electronic device can be a server, which can be a physical server or a cloud server. The server can be a single server or a server cluster.

[0040] like Figure 1 The image rendering method provided in this embodiment mainly includes the following steps:

[0041] S101. Process the image to be rendered to determine the text area.

[0042] The image to be rendered can be any given image. For example, it could be a photograph to which text needs to be added, or any video frame extracted from a video. In this embodiment, the image to be rendered is described only and not limited to it.

[0043] The text region can be understood as a connected region in the image to be rendered where text can be added. Text can be added to this text region; this text refers to text information related to the image to be rendered. This text information can be determined based on the image information or can be text entered by the user to be added to the image.

[0044] For example, if the image to be rendered is a picture from a movie or video, the text information could be the name of the movie or video. Alternatively, the text information could be the main content of the image to be rendered, such as "mountain peak" or "tree." Another example is user-provided text information. The user-provided text information is input by the user through an input device.

[0045] In one implementation, a connected region is selected at any location in the image to be rendered as the text region. For example, a connected region is selected in the middle of the image, or in the upper left corner. Furthermore, the text region can be any area in the image that will not obscure the main subject of the image after text is added.

[0046] In one implementation, the user's selection operation in the image to be rendered is received, and the area selected by the user in the image to be rendered is used as a text area. For example, the user manually selects a rectangular connected region in the image to be rendered, and the rectangular connected region is used as a text area.

[0047] In one implementation, the image to be rendered is input into a pre-trained segmentation model, and the text region corresponding to the image to be rendered is determined based on the image mask output by the pre-trained segmentation model.

[0048] S102. Determine the target text type based on the attribute information of the text region.

[0049] The attribute information of the text area can be the width and / or height of the bounding rectangle of the text area. The width and height can be represented by length units or pixels. In this embodiment, no specific limitation is made.

[0050] The text type can be understood as information representing the characteristics of the text, such as: font size, font style, character shape, character spacing, and the position of the text relative to the text area. Font size refers to information representing the size of a character, such as initial size, small initial size, size 1, size 2, etc.; font style refers to information representing the shape of the character, such as regular script, Song typeface, boldface, etc.; character shape refers to information representing special effects of a character, such as bolding, italics, etc. Furthermore, the aforementioned text can be any existing writable script, such as Chinese characters, English, Korean, Greek letters, Arabic numerals, etc., and can also be any writable symbol such as "%", "@", "&".

[0051] Furthermore, the target font size of the text is determined based on the width of the bounding box of the text region. This ensures that the text of the target font size can fill the entire text region. Optionally, the bounding box of the text region is a rectangular bounding box, and the width of the bounding box can be understood as the length of the horizontal axis in a two-dimensional coordinate system.

[0052] In one implementation, the target font size is calculated by decreasing the font size sequentially from the largest font size. For each font size, the text width at that font size is calculated, and it is determined whether the text width is less than or equal to the width of the bounding box of the text area. Here, the text width refers to the length of all characters at a given font size. For example, in a 12-point font, each character has a width of 6.3mm, and the text contains 10 characters, so the text width is 63mm.

[0053] In one implementation, the width of a single character at each font size is determined; the ratio of the outer frame width to the number of characters in the text is calculated, and the font size corresponding to the width of the character closest to this ratio is determined as the target font size. For example: for font size 1, the width of each character is 9.8mm; for font size 2, the width of each character is 7.4mm; for font size 2 (small two), the width of each character is 6.3mm; and for font size 3, the width of each character is 5.6mm. If the width of the outer frame is 60mm and the number of characters in the text is 9, the ratio of the outer frame width to the number of characters in the text is 6.67. This ratio is closest to 6.3mm, so font size 2 (small two) corresponding to 6.3mm is selected as the target font size for the text.

[0054] In one implementation, starting with the largest font size and decreasing sequentially, the number of text characters that the outer frame width can accommodate at each font size is calculated until the number of text characters that can be accommodated exceeds the actual number of text characters. The font size corresponding to this accommodated number of text characters is then determined as the target font size. For example: Font size 1 has a character width of 9.8mm, font size 2 has a character width of 7.4mm, font size 2 (small two) has a character width of 6.3mm, and font size 3 has a character width of 5.6mm. If the outer frame width is 70mm and the actual number of characters is 10, font size 1 can accommodate 7.1 characters, font size 2 can accommodate 9.4 characters, and font size 2 (small two) can accommodate 11 characters. Since the number of characters that can be accommodated for font size 2 (small two) is greater than the actual number of characters, font size 2 (small two) is determined as the target font.

[0055] In one implementation, the system default font is used as the target font for the text; alternatively, the target font can be determined in response to a user's input of a font selection action.

[0056] In one implementation, the system default glyph (e.g., regular glyph) is used as the target font for the text, or the target glyph can be determined in response to a user's input of a glyph selection (bold, italic).

[0057] S103. Determine the text target pattern type based on the image to be rendered.

[0058] The pattern type can be understood as a special effect for text fill or border. Optionally, the target pattern type can be any one or more of the following: target color, target texture, target effect, etc. The target color can be a single color value or a gradient color corresponding to multiple color values. The target texture can be understood as a text fill texture; it can be the system default texture, or it can be determined in response to a user's texture selection. The target effect can be one or a combination of several of the following: adding shadows, reflections, adding text borders, glowing effects, 3D effects, etc.

[0059] In one implementation, the target text color can be determined based on the color information of the image to be rendered. This color information can be represented using any of the RGB color system, HSV color space, or HSL color space.

[0060] The RGB color system obtains various colors by varying the three color channels: red (R), green (G), and blue (B) and superimposing them on each other.

[0061] In one implementation, the values ​​corresponding to the three color channels in the RGB color system of the image to be rendered are extracted, and these values ​​are directly determined as the target color of the text.

[0062] In another implementation, the values ​​corresponding to the three color channels in the RGB color system within the text area are extracted, the color corresponding to the value is determined, and the complementary color of the value is determined as the target color of the text. For example, after extracting the text area, the color corresponding to the RGB value is red, and the complementary color of red, green, is determined as the target color of the text.

[0063] The HSV color space represents a color using three parameters: chroma (H), saturation (S), and lightness (V). The HSV color space is a three-dimensional representation of the RGB color system.

[0064] In one implementation, the chromaticity values ​​of the HSV color space are extracted from the image to be rendered, and the average value of the H value of the corresponding text region image is calculated. The color value that differs the most from H_Avg is then used as the text color value.

[0065] In one implementation, any portion of the image to be rendered is extracted as the text target texture.

[0066] S104. Render the image to be rendered based on the target text type and the target pattern type.

[0067] In this embodiment, the text is displayed and rendered within the text area according to certain rules based on the target text type and the target text pattern type. These rules include: centered display, left-aligned display, right-aligned display, etc. The specific display rendering method will not be described in detail in this embodiment.

[0068] This disclosure provides an image rendering method, including: processing an image to be rendered to determine a text region; determining a target text type based on attribute information of the text region; determining a target text pattern type based on the image to be rendered; and rendering the image to be rendered based on the target text type and the target text pattern type. This disclosure achieves rapid image rendering by determining the text type based on the obtained text region, determining the text pattern type based on the image to be rendered, and rendering the text according to the text type and text pattern type before adding it to the text region of the image. In other words, it harmoniously and aesthetically places the given text in the image.

[0069] Based on the above embodiments, the present disclosure further optimizes the image rendering method described above. Figure 2 This is a flowchart of the optimized image rendering method in the embodiments of this disclosure, such as... Figure 2 As shown, the optimized image rendering method provided in this embodiment mainly includes the following steps:

[0070] S201. Select video frames from the video to be processed as images to be rendered.

[0071] In this context, "video" broadly refers to a video composed of multiple video frames, such as short videos, live streams, and movies / TV shows. This application does not limit the specific type of video. The video to be processed can be understood as a video without a cover image.

[0072] In this embodiment, the image rendering method provided in this disclosure embodiment can be executed after receiving the cover generation instruction, i.e., steps S201-S207. The cover generation instruction can be generated and sent in response to a user-inputted cover operation, or it can be automatically generated and sent after receiving a user-uploaded video and detecting that the video does not have a cover.

[0073] In this context, a video cover image refers to an image used to present a summary of a video. A video cover image can be a static image, also known as a static video cover image. Alternatively, it can be a dynamic video clip, also known as a dynamic video cover image. For example, the images displayed as cover images in video listings on video platforms help users get a general idea of ​​the live stream content.

[0074] In one implementation, any frame from the video to be processed is selected as the image to be rendered; or, based on the user's selection operation, the video frame selected by the user is selected as the image to be rendered.

[0075] S202. Input the image to be rendered into the segmentation model to obtain the image mask.

[0076] This example provides a method for training a segmentation model, which mainly includes: collecting data samples, which primarily consist of a base image and an image mask; and inputting the collected data samples into a neural network model for training to obtain the segmentation model.

[0077] The process involves inputting the image to be rendered into the segmentation model, which then processes the image to obtain a mask.

[0078] Figure 3 This is a schematic diagram of the text region in the image to be rendered provided in an embodiment of this disclosure. For example... Figure 3 As shown, Figure 3 The leftmost image to be rendered is input into the segmentation model, which processes it to obtain the grayscale image in the middle. The grayscale image is then binarized to obtain the image mask on the right.

[0079] Furthermore, the purpose of binarization is to classify the target user and the background. The most common method for binarizing grayscale images is the thresholding method. This method utilizes the difference between the target and the background in the image to set the image to two different levels. By selecting an appropriate threshold, it determines whether a pixel is the target or the background, thus obtaining a binarized image.

[0080] In this embodiment, a threshold method is used to... Figure 3 The grayscale image in the middle is binarized to obtain... Figure 3 The binarized image on the right.

[0081] S203. When the foreground region in the image mask is greater than or equal to the first threshold, the text region is set in the region corresponding to the foreground region in the image to be rendered.

[0082] The foreground region can be understood as the area composed of white pixels in a binarized image mask, such as... Figure 3The white area in the right figure on the right. The foreground area can also be referred to as the region of interest. Corresponding to the foreground area is the background area, which refers to the area composed of black pixel points in the binary image mask, such as Figure 3 The black area in the right figure on the right.

[0083] Among them, the first threshold is used to determine whether the size of the foreground area in the image mask is too small. If the size of the foreground area in the image mask is greater than or equal to the set first threshold, it means that the size of the foreground area in the image mask is relatively large and can be set as the text area. If the size of the foreground area in the image mask is less than the set first threshold, it means that the foreground area in the image mask is too small, and setting the text area may cause the main body of the image to be blocked by text, which is not suitable as the text area and other positions need to be reselected as the text area.

[0084] S204. Determine the text target font size based on the width of the bounding box and the number of text characters.

[0085] Among them, the attribute information of the text area includes the width of the bounding box of the text area, and the text target character type includes the text target font size. The bounding box of the text area can be understood as Figure 3 The bounding box of the white pixel points on the right.

[0086] In one embodiment, determining the text target font size based on the width of the bounding box and the number of text characters includes: traversing each font size starting from the largest font size; determining the text width based on the currently traversed font size and the number of text characters; when the text width is less than or equal to the width of the bounding box, determining the currently traversed font size as the text target font size.

[0087] It should be noted that the largest font size and the smallest font size can be set in advance. The largest font size is generally the largest font size provided by the system. For example, the largest font size is the initial size. The smallest font size is the smallest font size provided by the system. For example, the smallest font size is the eighth size.

[0088] In one embodiment, the smallest font size can be set according to the size of the image to be rendered. If the image to be rendered is too large and the text font is too small, it will look unaesthetic and unharmonious, and the too small font also affects the viewing effect of the audience. Therefore, setting the smallest font size according to the size of the image to be rendered can avoid calculating the font size too many times and wasting resources and time.

[0089] In this embodiment, determining the text width based on the currently traversed font size and the number of text characters can include: taking the product of the width of a single font corresponding to the current font size and the number of text characters as the text width.

[0090] Specifically, traverse each font size starting from the largest font size; multiply the single font width corresponding to the current font size by the number of text characters to obtain the text width; when the text width is less than or equal to the width of the bounding box, determine the currently traversed font size as the target font size of the text.

[0091] For example: Take the largest font size, the initial font size, as the current font size, multiply the single font width corresponding to the initial font size by the number of text characters to obtain the text width, compare the text width with the width of the bounding box. If the text width is less than or equal to the width of the bounding box, determine the initial font size as the target font size of the text. If the text width is greater than the width of the bounding box, select the next smaller font size, such as the small initial font size as the current font size, multiply the single font width corresponding to the small initial font size by the number of text characters to obtain the text width, compare the text width with the width of the bounding box. If the text width is less than or equal to the width of the bounding box, determine the small initial font size as the target font size of the text; if the text width is greater than the width of the bounding box, select the next smaller font size, such as the number one font as the current font size, and return to execute the step of multiplying the single font width corresponding to the current font size by the number of text characters as the text width and subsequent steps until the text width is less than or equal to the width of the bounding box, and determine the currently traversed font as the target font size of the text.

[0092] S205. Convert the to-be-rendered picture to the HSV color space.

[0093] Among them, the HSV color space represents a color through three parameters: hue (H), saturation (S), and value (V). The HSV color space is a three-dimensional representation of the RGB color system.

[0094] Among them, the hue (H) component is measured in degrees, with a value range of 0° to 360°, calculated counterclockwise starting from red, where red is 0°, green is 120°, and blue is 240°. Their complementary colors are: yellow is 60°, cyan is 180°, and purple is 300°;

[0095] The saturation (S) component represents the degree to which a color approaches a spectral color. A color can be regarded as the result of mixing a certain spectral color with white. The larger the proportion of the spectral color, the higher the degree to which the color approaches the spectral color, and the higher the saturation of the color. High saturation means the color is deep and vivid. The white component of the spectral color is 0, and the saturation reaches the highest. Usually, the value range is 0% to 100%, and the larger the value, the more saturated the color.

[0096] The value (V) component represents the brightness of a color. For a light source color, the value of value is related to the luminance of the light-emitting body; for an object color, this value is related to the transmittance or reflectance of the object. Usually, the value range is from 0% (black) to 100% (white).

[0097] S206. For at least one pixel in the image to be rendered, obtain the chromaticity value in the HSV color space.

[0098] In one implementation, the entire image to be rendered is converted to the HSV color space to obtain the chromaticity values ​​in the HSV color space.

[0099] In another implementation, the image corresponding to the text area in the image to be rendered is converted to the HSV color space to obtain the chromaticity values ​​in the HSV color space.

[0100] S207. Determine the target color of the text based on the chromaticity values ​​of at least one or more pixels.

[0101] The target color of the text is determined based on the average value of the chroma component H_Avg, the average value of the saturation component S_Avg, and the average value of the luminance component V_Avg.

[0102] In this embodiment, the chromaticity values ​​are extracted from the image to be rendered, or from the image corresponding to the text region of the image to be rendered, and the average chromaticity value corresponding to multiple pixels is calculated to obtain the average chromaticity value H_Avg.

[0103] Find all colors in the entire color set S that have the smallest difference in the H-value dimension from the average chromaticity H_Avg, and use these colors as the candidate color set O for the text. For example... Figure 4 As shown, it is determined based on the average chromaticity H_Avg. Figure 4 Which column of colors in the table represents the candidate color set O? Minimizing the difference in H values ​​ensures that the text colors look harmonious and aesthetically pleasing.

[0104] Furthermore, any color can be selected from the color candidate set as the target text color; alternatively, the color with the highest saturation or brightness can be selected from the color candidate set as the target text color.

[0105] In one embodiment, determining the target text color based on the chromaticity values ​​of a plurality of pixels includes: calculating the chromaticity average of the chromaticity values ​​of the plurality of pixels; determining a color candidate set based on the chromaticity average; obtaining the saturation value and brightness value in the HSV color space for at least one pixel in the image to be rendered; and selecting the target text color from the color candidate set based on the saturation value and / or the brightness value of at least one or more pixels.

[0106] In one implementation, the chromaticity values ​​of the HSV color space are extracted from the image to be rendered, and the average chromaticity value H_Avg of the corresponding text region is calculated. The color value that differs the most from the average chromaticity value H_Avg is then used as the text color value.

[0107] In one embodiment, selecting a text target color from the color candidate set based on the saturation values ​​and / or brightness values ​​of multiple pixels includes: calculating the average saturation value and the average brightness value of the multiple pixels; for each color value in the color candidate set, calculating a first difference between the color value and the average saturation value, and / or calculating a second difference between the color value and the average brightness value; and determining the color corresponding to the maximum value of the first difference and / or the color corresponding to the maximum value of the second difference as the text target color.

[0108] Specifically, if the color value corresponding to the first maximum difference is the same as the color value corresponding to the second maximum difference, then the color corresponding to that value is determined as the target text color. If the color value corresponding to the first maximum difference is not the same as the color value corresponding to the second maximum difference, then either the color corresponding to the first maximum difference or the color corresponding to the second maximum difference is selected as the target text color.

[0109] In this embodiment, selecting the largest difference in saturation components and the largest difference in average brightness ensures strong contrast between text and background colors, which is beneficial for improving the reading experience.

[0110] S208. Render the image to be rendered based on the target text type and the target pattern type.

[0111] S209. Select the rendered image as the cover of the video to be processed.

[0112] In one embodiment, the image rendering method provided in this disclosure further includes: when the foreground region in the image mask is less than a first threshold, dividing the image to be rendered into a first region and a second region; and setting the text region in the first region or the second region.

[0113] The first threshold is used to determine whether the foreground region in the image mask is too small. If the foreground region in the image mask is smaller than the set first threshold, it indicates that the foreground region in the image mask is too small and unsuitable as a text region, requiring the selection of another region to be used as the text region. The first region and the second region can be understood as two regions with different main elements of the image. Optionally, the first region is the sky region and the second region is the ground region; alternatively, the first region is the beach region and the second region is the image area.

[0114] In this embodiment, the image to be rendered is divided into two different regions; the text region is set in either the first region or the second region. The method of dividing the image to be rendered into two regions will not be described in detail in this embodiment.

[0115] Furthermore, determine the size of the first and second regions, and place the text region in the larger region; if the two regions are not significantly different in size, select a region relatively close to the top edge or left side of the image to be rendered to place the text region; this ensures that the text is aesthetically pleasing and harmonious.

[0116] In one implementation, if the first region is smaller than a second threshold, or if the second region is smaller than a second threshold, then the text region is set at a preset position in the image to be rendered.

[0117] The second threshold is used to determine whether the first or second region is too small. If both the first and second regions are smaller than the set second threshold, it means that both regions are too small and unsuitable as text regions, requiring the selection of other regions for placement. In this case, the text region can be placed at any location in the image to be rendered.

[0118] Optionally, the preset position in the image to be rendered can be the center of the image, or the image can be divided according to a certain ratio, with the text area placed at the division point. This certain ratio can be a 4:6 ratio, a 3:7 ratio, or the golden ratio, etc. This ensures the text is aesthetically pleasing and harmonious.

[0119] like Figure 5 As shown, the image to be rendered is divided into a sky area and a ground area; the text area is placed on the sky area, that is, the text "On the Road to Dreams" is added to the sky area.

[0120] In one implementation, if the text information includes a main title and a subtitle, the text area can be divided into a main title area and a subtitle area. The text area can be divided into two equal areas, or it can be divided according to a certain ratio.

[0121] Furthermore, if the text area is too small to be divided into sections, the text area can be used as the main title area, and a section near the text area can be selected as the subtitle area.

[0122] Figure 6 This is a schematic diagram of the structure of an image rendering device according to an embodiment of the present disclosure. This embodiment can be applied to adding text effects to any image. The image rendering device can be implemented in software and / or hardware and can be configured in an electronic device.

[0123] like Figure 6The image rendering apparatus provided in this embodiment mainly includes the following modules: text region determination module 61, text type determination module 62, pattern type determination module 63, and rendering module 64.

[0124] The text region determination module 61 is used to process the image to be rendered to determine the text region; the target font size determination module 62 is used to determine the target text type based on the attribute information of the text region; the target color determination module 63 is used to determine the target text pattern type based on the image to be rendered; and the rendering module 64 is used to render the image to be rendered based on the target text type and the target text pattern type.

[0125] This disclosure provides an image rendering apparatus for performing the following steps: processing an image to be rendered to determine a text region; determining a target text type based on attribute information of the text region; determining a target text pattern type based on the image to be rendered; and rendering the image to be rendered based on the target text type and the target text pattern type. This disclosure achieves rapid image rendering by determining the text type based on the obtained text region, determining the text pattern type based on the image to be rendered, rendering the text according to the text type and text pattern type, and then adding the text to the text region on the image. This means that the given text is harmoniously and aesthetically placed in the image.

[0126] In one embodiment, the text region determination module includes: an image mask determination unit, configured to input the image to be rendered into a segmentation model to obtain an image mask; and a text region determination unit, configured to set the text region in the region corresponding to the foreground region in the image to be rendered when the foreground region in the image mask is greater than or equal to a first threshold.

[0127] In one embodiment, the text region determination module further includes: an image segmentation unit, configured to segment the image to be rendered into a first region and a second region when the foreground region in the image mask is less than a first threshold; the text region determination unit is further configured to use the sky region or the ground region as the text region.

[0128] In one embodiment, the text region determining unit is further configured to set the text region at a specified position in the image to be rendered if the first region is smaller than a second threshold, or if the second region is smaller than the second threshold.

[0129] In one embodiment, the text target pattern type includes: text target color; correspondingly, the pattern type determination module includes: an image conversion unit for converting the image to be rendered to the HSV color space; a chromaticity value acquisition unit for acquiring the chromaticity value in the HSV color space for at least one pixel in the image to be rendered; and a target color unit for determining the text target color based on the chromaticity values ​​of at least one or more pixels.

[0130] In one embodiment, the target color unit includes: a chromaticity average calculation subunit for calculating the chromaticity average of multiple pixels; a color candidate set determination subunit for determining a color candidate set based on the chromaticity average; a saturation and brightness value acquisition subunit for acquiring the saturation and brightness values ​​in the HSV color space for at least one pixel in the image to be rendered; and a target color determination subunit for selecting a text target color from the color candidate set based on the saturation and / or brightness values ​​of at least one or more pixels.

[0131] In one implementation, the target color determination subunit is specifically used to calculate the average saturation and average brightness of multiple pixels; for each color value in the color candidate set, calculate a first difference between the color value and the average saturation, and / or calculate a second difference between the color value and the average brightness; and determine the color corresponding to the maximum value of the first difference and / or the color corresponding to the maximum value of the second difference as the text target color.

[0132] In one implementation, the attribute information of the text region includes the width of the text region bounding box, and the target text type includes the target font size.

[0133] Correspondingly, the text type determination module is used to determine the target font size of the text based on the width of the outer frame and the number of text characters.

[0134] In one implementation, the text type determination module is specifically used to traverse each font size starting from the largest font size; determine the text width based on the current font size and the number of characters in the text; and when the text width is less than or equal to the width of the bounding box, determine the current font size as the target font size of the text.

[0135] In one embodiment, the apparatus further includes: a to-be-rendered image determination module, configured to select a video frame from the video to be processed as the to-be-rendered image; correspondingly, the apparatus further includes: a cover determination module, configured to render the to-be-rendered image based on the text target text type and the text target pattern type, and then determine the rendered image as the cover of the video to be processed.

[0136] The image rendering apparatus provided in this disclosure embodiment can execute the steps performed in the image rendering method provided in this disclosure method embodiment, and has the execution steps and beneficial effects, which will not be described in detail here.

[0137] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. See below for details. Figure 7 The diagram illustrates a structural schematic suitable for implementing the electronic device 700 in the embodiments of this disclosure. The electronic device 700 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable terminal devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 7 The terminal device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0138] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703 to implement the image rendering method as described in the embodiments of this disclosure. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing device 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0139] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0140] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the page navigation method as described above. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of embodiments of this disclosure.

[0141] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0142] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0143] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0144] The aforementioned computer-readable medium carries one or more programs, which, when executed by the terminal device, cause the terminal device to: process the image to be rendered to obtain a text region; determine the target font size of the text based on the attribute information of the text region; determine the target color of the text based on the background color information of the image to be rendered; and render the image to be rendered based on the target font size and the target color of the text.

[0145] Optionally, when one or more of the above-mentioned programs are executed by the terminal device, the terminal device may also execute other steps described in the above embodiments.

[0146] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0148] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0149] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0150] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0151] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method, including: processing an image to be rendered to determine a text region; determining a target text type based on attribute information of the text region; determining a target text pattern type based on the image to be rendered; and rendering the image to be rendered based on the target text type and the target text pattern type.

[0152] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method for processing an image to be rendered to determine a text region, including: inputting the image to be rendered into a segmentation model to obtain an image mask; and when the foreground region in the image mask is greater than or equal to a first threshold, setting the text region in the region corresponding to the foreground region in the image to be rendered.

[0153] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method, the method further comprising: when the foreground region in the image mask is less than a first threshold, dividing the image to be rendered into a first region and a second region; and setting the text region in the first region or the second region.

[0154] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method, the method comprising: if a first region is smaller than a second threshold, or if the second region is smaller than the second threshold, then setting the text region at a preset position in the image to be rendered.

[0155] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method, wherein the text target pattern type includes: text target color; correspondingly, determining the text target pattern type based on the image to be rendered includes: converting the image to be rendered to the HSV color space; obtaining the chromaticity value in the HSV color space for at least one pixel in the image to be rendered; and determining the text target color based on the chromaticity values ​​of at least one or more pixels.

[0156] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method for determining a target text color based on the chromaticity values ​​of a plurality of pixels, comprising: calculating a chromaticity average value of the chromaticity values ​​of the plurality of pixels; determining a color candidate set based on the chromaticity average value; obtaining a saturation value and a brightness value in the HSV color space for at least one pixel in the image to be rendered; and selecting a target text color from the color candidate set based on the saturation value and / or the brightness value of at least one or more pixels.

[0157] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method for selecting a text target color from a color candidate set based on the saturation values ​​and / or the brightness values ​​of a plurality of pixels, comprising: calculating an average saturation value and an average brightness value of a plurality of pixels; for each color value in the color candidate set, calculating a first difference between the color value and the average saturation value, and / or calculating a second difference between the color value and the average brightness value; and determining the color corresponding to the maximum value of the first difference and / or the color corresponding to the maximum value of the second difference as the text target color.

[0158] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method, wherein the attribute information of the text region includes the width of the text region's bounding box, and the text target text type includes the text target font size; correspondingly, determining the text target font size based on the attribute information of the text region includes: determining the text target font size based on the width of the bounding box and the number of text characters.

[0159] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method that determines the target font size of text based on the width of the bounding box and the number of text characters, including: traversing each font size starting from the largest font size; determining the text width based on the current font size traversed and the number of text characters; and determining the current font size traversed as the target font size when the text width is less than or equal to the width of the bounding box.

[0160] According to one or more embodiments of this disclosure, this disclosure provides an image rendering method, the method further comprising: selecting a video frame from the video to be processed as an image to be rendered; correspondingly, after rendering the image to be rendered based on the text target text type and the text target pattern type, the method further comprises: determining the rendered image as the cover of the video to be processed.

[0161] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus, including: a text region determination module, used to process an image to be rendered to determine a text region; a target font size determination module, used to determine a target text type based on attribute information of the text region; a target color determination module, used to determine a target text pattern type based on background color information of the image to be rendered; and a rendering module, used to render the image to be rendered based on the target text type and the target text pattern type.

[0162] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus and a text region determination module, comprising: an image mask determination unit, configured to input the image to be rendered into a segmentation model to obtain an image mask; and a text region determination unit, configured to set a text region in the region corresponding to the foreground region in the image to be rendered when the foreground region in the image mask is greater than or equal to a first threshold.

[0163] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus and a text region determination module, further comprising: an image segmentation unit, configured to segment the image to be rendered into a first region and a second region when the foreground region in the image mask is less than a first threshold; and a text region determination unit, further configured to use the sky region or the ground region as the text region.

[0164] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus, a text region determining unit, which is further configured to set the text region at a preset position in the image to be rendered if the first region is smaller than a second threshold, or if the second region is smaller than the second threshold.

[0165] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus, wherein the text target pattern type includes: a text target color; correspondingly, the pattern type determination module includes: an image conversion unit, configured to convert the image to be rendered to an HSV color space; a chromaticity value acquisition unit, configured to acquire a chromaticity value in the HSV color space for at least one pixel in the image to be rendered; and a target color unit, configured to determine the text target color based on the chromaticity values ​​of at least one or more pixels.

[0166] According to one or more embodiments of this disclosure, an image rendering apparatus is provided, including a target color unit, comprising: a chromaticity average calculation subunit for calculating the chromaticity average of multiple pixels; a color candidate set determination subunit for determining a color candidate set based on the chromaticity average; a saturation and brightness value acquisition subunit for acquiring a saturation value and a brightness value in an HSV color space for at least one pixel in the image to be rendered; and a target color determination subunit for selecting a text target color from the color candidate set based on the saturation value and / or the brightness value of at least one or more pixels.

[0167] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus, a target color determination subunit, specifically configured to calculate the average saturation and average brightness of a plurality of pixels; for each color value in a color candidate set, calculate a first difference between the color value and the average saturation, and / or calculate a second difference between the color value and the average brightness; and determine the color corresponding to the maximum value of the first difference and / or the color corresponding to the maximum value of the second difference as the text target color.

[0168] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus, wherein the attribute information of the text region includes the width of the text region bounding box, and the text target text type includes the text target font size; correspondingly, a text type determination module is used to determine the text target font size based on the width of the bounding box and the number of text characters.

[0169] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus, a text type determination module, specifically used to traverse each font size starting from the largest font size; determine the text width based on the current font size traversed and the number of text characters; and when the text width is less than or equal to the width of the outer frame, determine the current font size traversed as the target font size of the text.

[0170] According to one or more embodiments of this disclosure, this disclosure provides an image rendering apparatus, the apparatus further comprising: a to-be-rendered image determination module, configured to select a video frame from the video to be processed as the to-be-rendered image; correspondingly, the apparatus further comprises: a cover determination module, configured to render the to-be-rendered image based on the text target text type and the text target pattern type, and then determine the rendered image as the cover of the video to be processed.

[0171] According to one or more embodiments of this disclosure, this disclosure provides an electronic device, including:

[0172] One or more processors;

[0173] Memory, used to store one or more programs;

[0174] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the image rendering methods provided in this disclosure.

[0175] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements an image rendering method as described in any of the present disclosure.

[0176] This disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, implement the image rendering method described above.

[0177] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0178] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0179] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. An image rendering method, characterized in that, The method includes: Select video frames from the video to be processed as images to be rendered; The text region is determined by processing the image to be rendered. This process includes: inputting the image to be rendered into a segmentation model to obtain an image mask; setting the text region in the image to be rendered in the corresponding region of the foreground region when the foreground region in the image mask is greater than or equal to a first threshold; segmenting the image to be rendered into a first region and a second region when the foreground region in the image mask is less than the first threshold, and setting the text region in either the first region or the second region; and setting the text region at a preset position in the image to be rendered if the first region is less than a second threshold, or if the second region is less than the second threshold. The target text type is determined based on the attribute information of the text region; The text target pattern type is determined based on the image to be rendered. The text target pattern type includes: text target color. The text target color is the color in the color candidate set that has the largest difference between the average saturation value and / or the average brightness value of multiple pixels in the image to be rendered. The color candidate set is the color value in all color sets that has the smallest difference between the average chromaticity value and the image to be rendered. The image to be rendered is rendered based on the target text type and the target pattern type. The rendered image is selected as the cover image for the video to be processed.

2. The method according to claim 1, characterized in that, Determining the text target pattern type based on the image to be rendered includes: Convert the image to be rendered to the HSV color space; For at least one pixel in the image to be rendered, obtain the chromaticity value in the HSV color space; The target color of the text is determined based on the chromaticity values ​​of multiple pixels.

3. The method according to claim 2, characterized in that, Determining the target color of the text based on the chromaticity values ​​of multiple pixels includes: Calculate the chromaticity average of multiple pixels; A color candidate set is determined based on the chromaticity average value; For at least one pixel in the image to be rendered, obtain the saturation and brightness values ​​in the HSV color space; The text target color is selected from the color candidate set based on the saturation value and / or the brightness value of multiple pixels.

4. The method according to claim 3, characterized in that, Selecting a text target color from the color candidate set based on the saturation values ​​and / or brightness values ​​of multiple pixels includes: Calculate the average saturation and average brightness of multiple pixels; For each color value in the color candidate set, calculate a first difference between the color value and the average saturation value, and / or calculate a second difference between the color value and the average brightness value; The color corresponding to the maximum value of the first difference and / or the color corresponding to the maximum value of the second difference are determined as the target color for the text.

5. The method according to claim 1, characterized in that, The attribute information of the text region includes the width of the text region bounding box, and the target text type includes the target font size. Accordingly, determining the target font size of the text based on the attribute information of the text region includes: The target font size of the text is determined based on the width of the bounding box and the number of characters in the text.

6. The method according to claim 5, characterized in that, Determining the target font size of the text based on the width of the bounding box and the number of characters in the text includes: Iterate through each font size starting from the largest font size; The text width is determined based on the current font size and the number of characters in the text. When the width of the text is less than or equal to the width of the outer frame, the current font size encountered during the traversal is determined as the target font size of the text.

7. An image rendering apparatus, characterized in that, The device includes: The module for determining the image to be rendered is used to select video frames from the video to be processed as the images to be rendered. A text region determination module is used to process an image to be rendered to determine a text region. The process includes: inputting the image to be rendered into a segmentation model to obtain an image mask; when the foreground region in the image mask is greater than or equal to a first threshold, setting the text region in the region corresponding to the foreground region in the image to be rendered; when the foreground region in the image mask is less than the first threshold, segmenting the image to be rendered into a first region and a second region, and setting the text region in either the first region or the second region; if the first region is less than a second threshold, or the second region is less than the second threshold, then setting the text region at a preset position in the image to be rendered. The target font size determination module is used to determine the target text type based on the attribute information of the text region. The target color determination module is used to determine the text target pattern type based on the background color information of the image to be rendered. The text target pattern type includes: text target color, which is the color in the color candidate set that has the largest difference between the average saturation value and / or the average brightness value of multiple pixels in the image to be rendered. The color candidate set is the color value in all color sets that has the smallest difference between the average chromaticity value and the image to be rendered. The rendering module is used to render the image to be rendered based on the target text type and the target text pattern type. The cover determination module is used to determine the rendered image as the cover of the video to be processed.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.

10. A computer program product comprising a computer program or instructions that, when executed by a processor, implement the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method for adjusting color tone of text display area

    CN104076928A

  • Method, device for configuring text color in picture and electronic device

    CN109408177A

  • Method and device for adding characters into picture, electronic equipment and storage medium

    CN111161377A

  • OCR data synthesis method based on image structure information

    CN112949755A