A character rendering method, electronic device, medium and product

CN122886540APending Publication Date: 2026-10-09SHENZHEN TIANSHITONG INTELLIGENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611120116.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0003]但是,现有的OSD文字渲染方案存在背景适应性差、边缘抗干扰弱、画面遮挡严重、时间更新抖动及资源消耗高等问题,导致难以在在低功耗嵌入式设备上,同时兼顾OSD文字的全背景适应性、视觉稳定性与画面完整性

Benefits of technology

[0008]第五方面,本申请实施例提供一种计算机程序产品,包括计算机程序或计算机指令,所述计算机程序或所述计算机指令存储在计算机可读存储介质中,计算机设备的处理器从所述计算机可读存储介质读取所述计算机程序或所述计算机指令,所述处理器执行所述计算机程序或所述计算机指令,使得所述计算机设备执行如第一方面所述的文字渲染方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122886540A_ABST
    Figure CN122886540A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a character rendering method, relates to the technical field of video processing, and the method comprises the following steps: obtaining an original video frame and to-be-displayed text information; using a constant-width font to render the to-be-displayed text information, to generate a character bitmap; performing double-layered outline processing from outside to inside on the character bitmap, to form a double-color contour structure; through transparency mixing, superimposing the character bitmap after the double-layered outline processing to the original video frame, and outputting a rendered video frame. When OSD character rendering is performed, the embodiment of the application can take into account the full-background adaptability, visual stability and picture integrity of OSD characters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a text rendering method, electronic device, medium, and product. Background Technology

[0002] OSD (On-Screen Display) technology refers to overlaying character or graphic information onto a video screen, enabling monitoring personnel to obtain auxiliary information such as time, location, and channel number in real time. In fields such as video surveillance, security recording, and live streaming, on-screen display (OSD) components need to overlay information such as date, time, and channel name onto the video screen in real time.

[0003] However, existing OSD text rendering solutions suffer from poor background adaptability, weak edge interference resistance, severe screen occlusion, time update jitter, and high resource consumption, making it difficult to simultaneously ensure the full background adaptability, visual stability, and screen integrity of OSD text on low-power embedded devices. Summary of the Invention

[0004] This application provides a text rendering method, electronic device, medium, and product, which aims to balance the full background adaptability, visual stability, and image integrity of OSD text during OSD text rendering. In a first aspect, embodiments of this application provide a text rendering method, the method comprising: acquiring an original video frame and text information to be displayed; rendering the text information to be displayed using a fixed-width font to generate a text bitmap; performing a double-layer stroke processing on the text bitmap from the outside to the inside to form a two-color outline structure; and superimposing the text bitmap with the double-layer stroke processing onto the original video frame through transparency blending to output the rendered video frame.

[0005] Secondly, embodiments of this application provide a text rendering apparatus, the apparatus comprising: an acquisition module for acquiring an original video frame and text information to be displayed; a rendering module for rendering the text information to be displayed using a fixed-width font to generate a text bitmap; an outlining module for performing a double-layer outlining process on the text bitmap from the outside to the inside to form a two-color outline structure; and a blending module for overlaying the text bitmap with the double-layer outlining process onto the original video frame through transparency blending to output the rendered video frame.

[0006] Thirdly, embodiments of this application provide an electronic device, including: at least one processor; at least one memory for storing at least one program; and when at least one of the programs is executed by at least one of the processors, implementing the text rendering method as described in the first aspect.

[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions for performing the text rendering method as described in the first aspect.

[0008] Fifthly, embodiments of this application provide a computer program product, including a computer program or computer instructions, wherein the computer program or computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions, causing the computer device to perform the text rendering method as described in the first aspect.

[0009] In this embodiment, by using a monospaced font to render the text information to be displayed, text displacement and visual jitter caused by differences in character width during time updates can be effectively eliminated, thereby improving viewing comfort. By performing double-layer outlining processing on the text bitmap from the outside in to form a two-color outline structure, the width difference between the inner and outer outlining layers can be used to construct a clear boundary between the main body and the outline, significantly enhancing the edge recognizability of the text outline against a complex background. By using transparency blending, the text bitmap with the completed outlining processing is directly superimposed on the original video frame without introducing an additional background occlusion layer, which can avoid obscuring and destroying the original image content and preserve the complete information of the video image to the maximum extent. Overall, a clear and stable text visual effect is achieved while maintaining the simplicity of the rendering structure.

[0010] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0011] Figure 1 A flowchart illustrating the text rendering method provided in this application embodiment; Figure 2 A schematic diagram showing the OSD component provided in this application embodiment positioned in the upper left corner of the screen; Figure 3 A schematic diagram of the OSD component displayed against a dark background, provided for embodiments of this application; Figure 4 A schematic diagram of the OSD component displayed against a light-colored background, provided for an embodiment of this application; Figure 5 This is a schematic diagram of the structure of the text rendering device provided in the embodiments of this application; Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0013] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0014] In the description of the embodiments of this application, unless otherwise expressly limited, terms such as setting, installing, and connecting should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in the embodiments of this application in combination with the specific content of the technical solution.

[0015] In this application, the terms "furthermore," "exemplarily," or "optionally" are used as examples, illustrations, or descriptions and should not be construed as being more preferred or advantageous than other embodiments or designs. The use of terms such as "furthermore," "exemplarily," or "optionally" is intended to present the relevant concepts in a specific manner.

[0016] Figure 1 This is a flowchart of a text rendering method provided in an embodiment of this application. In this embodiment, the text rendering method includes, but is not limited to: Step S101: Obtain the original video frames and the text information to be displayed; Step S102: Render the text information to be displayed using a monospace font to generate a text bitmap; Step S103: Perform a double-layer stroke process from the outside to the inside on the text bitmap to form a two-color outline structure; Step S104: By blending transparency, the text bitmap with double-layer outline processing is superimposed onto the original video frame, and the rendered video frame is output.

[0017] In step S101, the original video frame is the basic image unit for overlaying text. Each frame is a two-dimensional image containing complete pixel color information, and its resolution and frame rate are determined by the attributes of the video source. The text information to be displayed is the character content that needs to be presented on the video screen, which can cover Chinese characters, letters, numbers, punctuation marks, and special symbols. Specific forms can include subtitles, timestamps, device identifiers, bullet comments, alarm prompts, etc.

[0018] The method for acquiring raw video frames can be selected according to the actual scenario: in real-time monitoring scenarios, raw YUV format video frames can be directly acquired from the camera through a video capture card; in video editing scenarios, they can be decoded and read frame by frame from locally stored video files; in live streaming scenarios, raw video frames can be obtained by decoding from network streaming.

[0019] The methods for obtaining the text information to be displayed are also flexible: for fixed subtitle content, it can be read from the subtitle file in advance and matched with the corresponding video frame according to the timestamp; for real-time bullet comments, the text content sent by the audience can be received through the network interface; for system information such as timestamps and channel numbers, it can be generated locally on the device in real time.

[0020] For example, in a security monitoring application scenario, the video input module acquires raw YUV video frames with a resolution of 1920×1080 from a high-definition camera, at a frame rate of 25 frames per second; the text management module generates the local system time "2026-07-21 14:30:00" and the channel identifier "CH-01" as the text information to be displayed, which is agreed to be superimposed on the upper left corner of the screen. Figure 2 As shown, Figure 2 An illustration of setting the OSD component in the upper left corner of the screen, with the OSD component maintaining a safe margin of A=25 pixels.

[0021] In step S102, a monospace font is a type of font in which all characters occupy the same horizontal width. Regardless of the complexity of the character's strokes, the number of horizontal pixels it occupies in the layout is exactly the same. Compared with proportional fonts, monospace fonts have the characteristics of neat layout and easy alignment, making them especially suitable for text scenarios that require vertical alignment, such as numbers and serial numbers. They also facilitate the unified calculation of the stroke extension range and character position coordinates.

[0022] A text bitmap is an image data representation of rasterized text content. Each pixel contains color and transparency information, with transparency represented by the alpha channel and used for subsequent transparency blending calculations. Pixels within the text area have higher transparency, while background pixels outside the text have zero transparency.

[0023] The specific process of rendering and generating a text bitmap is as follows: based on the configured font size and the fixed-width font file, the outline of the character is read character by character, rasterized according to the specified size, and then each character is arranged in turn according to the character spacing, and finally spliced ​​into a complete text bitmap.

[0024] In some implementations, the character spacing in the text information to be displayed is a preset multiple of the character width.

[0025] In this embodiment, character spacing refers to the blank space between two adjacent characters. Since each character in a monospaced font has the same width, setting the character spacing to a fixed multiple of the character width ensures that the spacing between all adjacent characters is uniform and consistent, avoiding the problem of spacing varying with the character width in proportional fonts, thus making the layout of the entire line of text neater and more aesthetically pleasing. The preset multiple can be adjusted within the range of 0.1 to 0.5 according to actual display requirements; a larger multiple results in sparser characters, while a smaller multiple results in more compact characters.

[0026] For example, if the "Source Han Monospace" font is selected, the font size is set to 20 pixels, and the standard width of each character is 14 pixels; if the preset multiplier is set to 0.2, then the character spacing is 14 × 0.2 = 2.8 pixels. When rendering the text "CH-01", there is a 2.8-pixel gap between every two adjacent characters, and the entire line of text is neatly arranged and visually uniform.

[0027] In step S103, the double-layer stroke processing is the core step in improving the text background adaptability. Specifically, two layers of strokes are drawn outward along the outline of the main text, from the outermost edge to the center of the text, consisting of an outer stroke, an inner stroke, and the main text, forming a three-layer progressive outline structure. The two-color outline structure, through different levels of color and width configuration, can adapt to dark and light backgrounds respectively, solving the deficiency of single-layer strokes in not being able to handle both light and dark scenes.

[0028] In some implementations, performing a double-layer stroke process from the outside in on the text bitmap to form a two-color outline structure specifically includes: An outer stroke layer is drawn around the text bitmap with a first width, an inner stroke layer is drawn inside the outer stroke layer with a second width, and the text body layer is drawn inside the inner stroke layer; the second width is smaller than the first width.

[0029] In this embodiment, the outer stroke layer is located at the outermost edge of the entire text outline, forming the outermost boundary where the text contacts the background. Its main function is to distinguish the text from the background against a light-colored background. The first width is the pixel width of the outer stroke, determining the thickness of the outermost outline. The inner stroke layer is located between the outer stroke layer and the main text layer. On one hand, it serves as a color transition, softening the color abruptness between the outer layer and the main text; on the other hand, it enhances the sense of depth at the edges of the text against a dark background. The second width is the pixel width of the inner stroke, and its value is smaller than the first width, ensuring a stepped shape where the outline gradually shrinks from the outside to the inside.

[0030] A morphological dilation algorithm can be used to draw a double-layered stroke: first, dilate the alpha channel of the main text by a first width to obtain the outer stroke area and fill it with the corresponding color; then, dilate it by a second width to obtain the inner stroke area and fill it with the corresponding color; the innermost part retains the original main text. A three-layered structure can be quickly obtained through two dilation operations, resulting in high computational efficiency.

[0031] For example, the first width is set to 2 pixels, and the second width is set to 1 pixel. When stroking the text "8", the text is first expanded outward by 2 pixels from the main body of the text to draw an outer black stroke; then expanded outward by 1 pixel to draw an inner black stroke; the innermost part is left as white text. Finally, from the outer edge of the text to the center, the text passes through a 2-pixel wide outer layer, a 1-pixel wide inner layer, and then back to the main body of the text, resulting in a total outline thickness of 3 pixels.

[0032] In this implementation, the outline of the text is expanded without significantly increasing the amount of computation by using a double-layered outline structure with a wider outer layer and a narrower inner layer. At the same time, the inner outline buffers the color difference between the outer layer and the main body, making the transition of the text edges more natural and reducing the visual abruptness.

[0033] In some implementations, the color of the text body layer is a first color, and the brightness contrast between the first color and the dark background is greater than a first preset value; the colors of the outer stroke layer and the inner stroke layer are both second colors, and the hue difference between the second color and the light background is greater than a second preset value.

[0034] In this embodiment, the first color is the fill color of the main text. The selection principle is to have sufficient brightness contrast against a dark background to ensure that the main text is clearly distinguishable against a dark background. The first preset value is the threshold of brightness contrast, which is usually set according to the standard settings of the web content accessibility guidelines to ensure that users with normal vision can clearly recognize the text.

[0035] The second color is used by both the outer and inner stroke layers. The selection principle is to ensure sufficient hue difference between the stroke color and the background color against a light background, guaranteeing that the stroke outline is clearly visible against the light background. The second preset value is the hue difference threshold, used to measure the degree of distinction between the stroke color and the light background; the larger the value, the higher the distinction.

[0036] "Dark background" refers to a color area with a luminance component Y ≤ 128, and "light background" refers to a color area with a luminance component Y > 128. The value range of the luminance component Y is 0~255.

[0037] The core logic of this color scheme is division of labor: the main text adapts to dark backgrounds, while the stroke layer adapts to light backgrounds. When text is overlaid on a dark screen, the bright text contrasts sharply with the dark background, making it clearly visible; when text is overlaid on a light screen, the dark stroke contrasts sharply with the light background, also clearly outlining the text. Therefore, it adapts to alternating light and dark video scenes without dynamically changing the text color based on the background.

[0038] For example, the first color is set to pure white, with a first preset value of 7:1. When the overlay area has a dark gray background, the brightness contrast between the white subject and the dark background is approximately 12:1, which is greater than the first preset value, making the text subject clear and prominent. The second color is set to pure black, with a second preset value of 4.5:1. When the overlay area has a light white background, the hue difference between the black outline and the light background is approximately 16:1, which is greater than the second preset value, making the text outline clearly distinguishable. Figure 3 and Figure 4 As shown, Figure 3 This is a schematic diagram showing the OSD components against a dark background; Figure 4 This is a schematic diagram of the OSD components displayed against a light-colored background.

[0039] In this embodiment, by dividing the color of the text body and the double-layer outline, respectively adapting to dark and light backgrounds, the text is ensured that no matter whether it is superimposed on a bright or dark video screen, there is at least one layer of structure that can form sufficient visual contrast with the background, fundamentally solving the problem of unstable text readability in dynamic video backgrounds.

[0040] In some implementations, when the background of the original video frame is a textured background, the width of the outer stroke layer is a third width, which is greater than the second width.

[0041] In this embodiment, a textured background refers to a background area with rich color details and complex patterns, such as tree branches and leaves, building facades, or dense crowds. Such backgrounds can easily cause visual interference with the edges of text, resulting in blurred text outlines. For textured backgrounds, appropriately increasing the width of the outer stroke can create a wider visual isolation band, effectively separating the main text from the complex background texture.

[0042] The third width is the outer stroke width specifically set for textured backgrounds. Its value is greater than the second width of the inner stroke and also greater than the first width of the outer stroke in normal scenes, thereby strengthening the isolation effect of the outline and resisting visual interference from the textured background.

[0043] Determining whether a background is textured can be achieved by calculating the pixel variance of the overlay area: statistically analyze the brightness variance of all pixels within the text overlay area. If the variance exceeds a preset threshold, it is determined to be a textured background; if it is below the threshold, it is determined to be a solid color or simple background.

[0044] For example, in a normal scene, the outer stroke width is 2 pixels and the inner stroke width is 1 pixel. When the text overlay area is detected to be a leaf texture background and the brightness variance exceeds the threshold, the outer stroke width is adjusted to a third width of 3 pixels, while the inner stroke remains at 1 pixel. The 3-pixel outer stroke can effectively cover the fine texture of the background, making the text still clear in the cluttered background.

[0045] In this embodiment, the outer stroke of the textured background is dynamically widened, which specifically solves the problem of reduced text recognition under complex textures and improves the display effect under special backgrounds without changing the overall rendering architecture.

[0046] In step S104, transparency blending, also known as alpha blending, is a crucial step in achieving a natural fusion of text and video frames. Its core idea is to perform a weighted calculation on the text pixels and background pixels based on the transparency of each pixel in the text bitmap to obtain the final output pixel. Through transparency blending, text edges can achieve a smooth transition from completely opaque to completely transparent, avoiding harsh, jagged edges and allowing the text to blend naturally into the video background.

[0047] In some implementations, the transparency blending is represented by a first formula, which is: Output_R=Source_R×Source_A+Background_R×(1-Source_A); Output_G=Source_G×Source_A+Background_G×(1-Source_A); Output_B=Source_B×Source_A+Background_B×(1-Source_A); Wherein, Output_R, Output_G, and Output_B are the R, G, and B channel values ​​of the mixed output pixel, respectively; Source_R, Source_G, and Source_B are the R, G, and B channel values ​​of the text bitmap pixel, respectively; Background_R, Background_G, and Background_B are the R, G, and B channel values ​​of the corresponding pixel in the original video frame, respectively; and Source_A is the opacity value of the text bitmap pixel.

[0048] In this embodiment, the formula is a standard per-channel alpha blending formula, performing the same blending calculation on the red, green, and blue color channels respectively. The opacity value Source_A ranges from 0 to 1. A value of 0 indicates that the pixel is completely transparent, and the background is fully displayed after blending; a value of 1 indicates that the pixel is completely opaque, and the text is fully displayed after blending; when the value is between 0 and 1, the colors of the text and the background are blended proportionally.

[0049] For the main text area, the pixel opacity is 1, ultimately displaying the full color of the text; for the stroke edge area, the opacity gradually transitions from 1 to 0, forming a smooth edge gradient; for the area outside the text, the opacity is 0, completely preserving the content of the original video frame.

[0050] For example, at a certain edge pixel location, the stroke pixel of the text bitmap is black, with R, G, and B values ​​all of 0 and an opacity of 0.6; the corresponding background pixel in the original video frame is light gray, with R, G, and B values ​​all of 200. Substituting these values ​​into the formula, we can obtain: Output_R=0×0.6+200×(1-0.6)=80; Output_G=0×0.6+200×(1-0.6)=80; Output_B=0×0.6+200×(1-0.6)=80; Ultimately, this position outputs a dark gray color, achieving a semi-transparent blend of the black outline and the light gray background, with a natural and smooth edge transition.

[0051] In this embodiment, by using a monospaced font to render the text information to be displayed, text displacement and visual jitter caused by differences in character width during time updates can be effectively eliminated, thereby improving viewing comfort. By performing double-layer outlining processing on the text bitmap from the outside in to form a two-color outline structure, the width difference between the inner and outer outlining layers can be used to construct a clear boundary between the main body and the outline, significantly enhancing the edge recognizability of the text outline against a complex background. By using transparency blending, the text bitmap with the completed outlining processing is directly superimposed on the original video frame without introducing an additional background occlusion layer, which can avoid obscuring and destroying the original image content and preserve the complete information of the video image to the maximum extent. Overall, a clear and stable text visual effect is achieved while maintaining the simplicity of the rendering structure.

[0052] In some embodiments, performing a double-layer stroke process from the outside in on the text bitmap to form a two-color outline structure includes: An outer stroke layer is drawn around the text bitmap with a first width, an inner stroke layer is drawn inside the outer stroke layer with a second width, and the text body layer is drawn inside the inner stroke layer; the second width is smaller than the first width.

[0053] To facilitate a comprehensive understanding of the execution process and effects of this solution, the complete processing will be explained using the rendering of bullet screen text in a live streaming scenario as an example: First, data acquisition. The live streaming device acquires raw video frames in real time, with a resolution of 1920×1080 and a frame rate of 30 frames per second; at the same time, it receives the bullet screen text "The picture quality is very clear" sent by the audience as the text information to be displayed, and it is agreed that the bullet screen text will be displayed from right to left.

[0054] Second, bitmap generation. Select the monospace font "Source Han Sans Monospace", set the font size to 24 pixels, and the width of each character to 18 pixels; set the character spacing to 0.15 times the character width, i.e., 2.7 pixels. Rasterize and render the text to be displayed character by character, generating a text bitmap with an alpha channel. The initial fill of the text is pure white.

[0055] Third, double-layered outlines. The standard configuration has an outer outline width of 2 pixels and an inner outline width of 1 pixel, both filled with pure black. The system detects background features in the overlay area of ​​the bullet comments in real time. When the bullet comments scroll to the audience texture area, if the background texture complexity exceeds the limit, the outer outline width is automatically adjusted to a third width of 3 pixels to enhance the outline isolation effect. This ultimately forms a three-layered structure with a wider outer and narrower inner black outline and a white main body.

[0056] Fourth, frame compositing output. Each pixel of the text bitmap is iterated through, and the text pixels are blended channel-by-channel with the corresponding pixels in the original video frame according to the alpha blending formula. The main text area is completely white, the outline area transitions gradually with transparency, and the transparent areas retain the original image. After compositing, a video frame with the superimposed subtitles is output and sent to subsequent encoding and streaming stages.

[0057] Figure 5 A schematic diagram of the structure of a text rendering apparatus provided in another embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0058] Reference Figure 5 The text rendering device may include: The acquisition module 501 is used to acquire the original video frames and the text information to be displayed; Rendering module 502 is used to render the text information to be displayed using a fixed-width font and generate a text bitmap; The stroke module 503 is used to perform a double-layer stroke process from the outside to the inside on the text bitmap to form a two-color outline structure; The blending module 504 is used to overlay the text bitmap with double-layer stroke processing onto the original video frame through transparency blending, and output the rendered video frame.

[0059] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application, and are devices corresponding to the above-mentioned methods. All implementation methods in the above-mentioned method embodiments are applicable to the embodiments of this device. For details on its specific functions and the technical effects it brings, please refer to the method embodiment section, which will not be repeated here.

[0060] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0061] Figure 6This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 6 As shown, the electronic device 600 includes a memory and a processor. The number of memories and processors can be one or more. Figure 6 Taking a memory 601 and a processor 602 as an example; the memory 601 and processor 602 in the network device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0062] The memory 601, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the methods provided in any embodiment of this application. The processor 602 implements the text rendering method provided in any of the above embodiments by running the software programs, instructions, and modules stored in the memory 601.

[0063] Memory 601 may primarily include a program storage area and a data storage area, wherein the program storage area may store the operating system and application programs required for at least one function. Furthermore, memory 601 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, memory 601 further includes memory remotely located relative to processor 602, and these remote memories can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0064] One embodiment of this application also provides a computer-readable storage medium storing computer-executable instructions for performing the text rendering method provided in any embodiment of this application.

[0065] An embodiment of this application also provides a computer program product, including a computer program or computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium and executes the computer program or computer instructions, causing the computer device to perform the text rendering method provided in any embodiment of this application.

[0066] The system architecture and application scenarios described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will understand that as system architectures evolve and new application scenarios emerge, the technical solutions provided in this application are also applicable to similar technical problems.

[0067] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0068] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0069] The terms “component,” “module,” “system,” etc., used in this specification are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process or execution thread, and components may be located on a single computer or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, or a network, such as the Internet interacting with other systems via signals).

[0070] The above description, with reference to the accompanying drawings, illustrates some embodiments of this application, but does not limit the scope of this application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of this application shall be within the scope of this application.

Claims

1. A text rendering method, characterized in that, The method includes: Obtain the original video frames and the text information to be displayed; The text information to be displayed is rendered using a monospace font to generate a text bitmap; Perform a double-layer stroke process from the outside to the inside on the text bitmap to form a two-color outline structure; By using transparency blending, the text bitmap with double-layered outline processing is overlaid onto the original video frame, and the rendered video frame is output.

2. The method as described in claim 1, characterized in that, The process of performing a double-layer stroke process from the outside to the inside on the text bitmap to form a two-color outline structure includes: An outer stroke layer is drawn around the text bitmap with a first width, an inner stroke layer is drawn inside the outer stroke layer with a second width, and the text body layer is drawn inside the inner stroke layer; the second width is smaller than the first width.

3. The method as described in claim 2, characterized in that, The text body layer is colored with a first color, and the brightness contrast between the first color and the dark background is greater than a first preset value. The outer stroke layer and the inner stroke layer are both colored with a second color, and the hue difference between the second color and the light background is greater than a second preset value.

4. The method as described in claim 2, characterized in that, When the background of the original video frame is a textured background, the width of the outer stroke layer is a third width, which is greater than the second width.

5. The method as described in claim 1, characterized in that, The transparency blending is represented by a first formula, which is: Output_R = Source_R × Source_A + Background_R × (1 - Source_A) Output_G = Source_G × Source_A + Background_G × (1 - Source_A) Output_B = Source_B × Source_A + Background_B × (1 - Source_A) Wherein, Output_R, Output_G, and Output_B are the R, G, and B channel values ​​of the mixed output pixel, respectively; Source_R, Source_G, and Source_B are the R, G, and B channel values ​​of the text bitmap pixel, respectively; Background_R, Background_G, and Background_B are the R, G, and B channel values ​​of the corresponding pixel in the original video frame, respectively; and Source_A is the opacity value of the text bitmap pixel.

6. The method as described in claim 1, characterized in that, The character spacing in the text information to be displayed is a preset multiple of the character width.

7. A text rendering device, characterized in that, The device includes: The acquisition module is used to acquire the original video frames and the text information to be displayed; The rendering module is used to render the text information to be displayed using a monospaced font and generate a text bitmap; The stroke module is used to perform a double-layer stroke process from the outside to the inside on the text bitmap to form a two-color outline structure; The blending module is used to overlay the text bitmap with double-layer outline processing onto the original video frame through transparency blending, and output the rendered video frame.

8. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; The method as described in any one of claims 1 to 6 is implemented when at least one of the programs is executed by at least one of the processors.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for performing the method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program or computer instructions, characterized in that, The computer program or the computer instructions are stored in a computer-readable storage medium, and the processor of the computer device reads the computer program or the computer instructions from the computer-readable storage medium. The processor executes the computer program or the computer instructions, causing the computer device to perform the method as described in any one of claims 1 to 6.