Methods, devices, electronic devices and storage media for replacing text in videos

By detecting and recognizing text regions and information in video frames, erasing the first text and inserting the second text, the problem of not being able to replace text at arbitrary positions in existing technologies is solved, enabling flexible replacement of text in videos.

CN115643464BActive Publication Date: 2026-07-17ALIBABA (CHINA) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA (CHINA) CO LTD
Filing Date
2022-09-26
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing methods for replacing text in videos cannot replace text at any location in the video, failing to meet the replacement needs in different scenarios and exhibiting poor applicability.

Method used

By detecting text regions in video frames, identifying text information, determining the first text region, erasing it, and inserting the corresponding second text information, text replacement at any position can be achieved.

Benefits of technology

It enables free replacement of text at any position in a video, meeting the text replacement needs in different scenarios and improving the applicability of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115643464B_ABST
    Figure CN115643464B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, electronic device, and storage medium for replacing text in a video. The method includes: detecting text regions containing text in video frames of video data to be processed; performing text recognition on the text regions to determine the text information included in the text regions; determining a first text region containing first text information from the text regions; erasing the first text information within the first text region; and inserting second text information corresponding to the first text information into the video frame after erasing the first text information. This solution can be used in fields such as film and television post-production and AR / VR, and has strong applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and storage medium for replacing text in a video. Background Technology

[0002] Text is one of the most common and important pieces of information in videos. Erasing and replacing specific text in videos has a wide range of applications. For example, in post-production of film and television, Chinese subtitles in videos can be erased and replaced with subtitles in other languages; the brand text of sponsor A in sports videos can be erased and replaced with the brand text of sponsor B; and sensitive text in videos can be erased and replaced with other text information.

[0003] Currently, when replacing text in a video, the text within a fixed area of ​​the video is erased, and other text information is inserted after the text is erased.

[0004] However, the text replacement method that erases text within a fixed area in a video and inserts other text information cannot replace text at any location in the video, nor can it replace specific text in the video. Therefore, the existing text replacement methods in videos have poor applicability and cannot meet the needs of replacing text in videos in different scenarios. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for replacing text in a video, in order to at least solve or alleviate the above-mentioned problems.

[0006] According to a first aspect of the present application, a method for text replacement in a video is provided, comprising: detecting a text region containing text in a video frame of video data to be processed; performing text recognition on the text region to determine the text information included in the text region; determining a first text region including first text information from the text region; erasing the first text information in the first text region; and inserting second text information corresponding to the first text information into the video frame after erasing the first text information.

[0007] According to a second aspect of the embodiments of this application, a video playback method is provided, applied to a video playback device, the video playback device including a virtual reality device or an augmented reality device, the method comprising: detecting a text region containing text in a video frame of video data to be processed; performing text recognition on the text region to determine the text information included in the text region; determining a first text region including first text information from the text region; erasing the first text information within the first text region; inserting second text information corresponding to the first text information into the video frame after erasing the first text information; and rendering the video frame after inserting the second text information onto the display of the video playback device.

[0008] According to a third aspect of the embodiments of this application, a video text replacement apparatus is provided, comprising: a text detection unit, configured to detect text regions containing text in video frames of video data to be processed; a text recognition unit, configured to perform text recognition on the text regions and determine the text information included in the text regions; a matching unit, configured to determine a first text region including first text information from the text regions; an erasure unit, configured to erase the first text information in the first text region; and a replacement unit, configured to insert second text information corresponding to the first text information into the video frame after erasing the first text information.

[0009] According to a fourth aspect of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the video text replacement method provided in the first aspect, or to perform the operation corresponding to the video playback method provided in the second aspect.

[0010] According to a fifth aspect of the embodiments of this application, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the video text replacement method described in the first aspect above, or implements the video playback method described in the second aspect above.

[0011] According to a sixth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that instruct a computing device to execute the video text replacement method described in the first aspect above, or to execute the video playback method provided in the second aspect above.

[0012] According to the scheme for replacing text in a video provided in this application, after detecting text regions in a video frame, the text information included in each text region is identified. Then, based on the text information included in each text region, a first text region containing first text information can be determined. The first text information in each first text region of the video frame is then erased, and second text information is inserted into the video frame after erasing the first text information, thus replacing the first text information in the video with the second text information. By detecting text regions and recognizing text information, the first text information at any position in the video frame can be identified. Therefore, text replacement in the video is no longer limited by the text position, and the first text information can be freely defined according to requirements. This allows for the replacement of any text in the video, meeting the needs of text replacement in different scenarios and making this video text replacement method highly applicable. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0014] Figure 1 This is a schematic diagram of an exemplary system applied in one embodiment of this application;

[0015] Figure 2 This is a flowchart of a video text replacement method according to an embodiment of this application;

[0016] Figure 3 This is a flowchart of a subtitle recognition method according to an embodiment of this application;

[0017] Figure 4 This is a flowchart of a video playback method according to an embodiment of this application;

[0018] Figure 5 This is a schematic diagram of a video text replacement device according to an embodiment of this application;

[0019] Figure 6 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0020] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the essence of the present application, well-known methods, processes, and flows are not described in detail. Furthermore, the accompanying drawings are not necessarily drawn to scale.

[0021] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows.

[0022] Subtitles: Subtitles (of motion picture) refer to non-visual content such as dialogue in television, film, and stage productions displayed in text form. They also broadly refer to text added during post-production of film and television works. In this application embodiment, subtitles refer to narration, lyrics, dialogue, explanatory text, etc., appearing in video frames.

[0023] Text detection: Using a computer to automatically detect regions containing text from an image; in this embodiment, it specifically refers to determining the regions where text is located in a video frame.

[0024] Text recognition: a technology that uses computers to automatically identify characters from images, specifically referring to the recognition of various texts in video frames in this application embodiment.

[0025] Page layout segmentation: Using computers to automatically classify image content at the pixel level.

[0026] Exemplary System

[0027] Figure 1 An exemplary system for a text replacement method in a video, applicable to embodiments of this application, is shown. For example... Figure 1 As shown, the system includes a cloud server 10, a communication network 20, and at least one user device 30. Figure 1 The example shown is for multiple user devices 30.

[0028] The cloud server 10 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 10 can perform any suitable function. For example, the cloud server 10 can be used to replace text in a video. In some embodiments, the cloud server 10 can receive instructions sent by the user device 30, replace the text in the target video according to the instructions, and send the video with the replaced text to the user device 30, so that the user device 30 can play the video with the replaced text.

[0029] Communication network 20 can be any suitable combination of one or more wired and / or wireless networks. For example, communication network 20 can include any one or more of the following: the Internet, intranet, wide area network (WAN), local area network (LAN), wireless network, digital subscriber line (DSL) network, frame relay network, asynchronous transfer mode (ATM) network, virtual private network (VPN), and / or any other suitable communication network. User equipment 30 can be connected to communication network 20 via one or more communication links (e.g., communication link 112), which can be connected to cloud server 10 via one or more communication links (e.g., communication link 114). Communication links can be any communication link suitable for transmitting data between cloud server 10 and user equipment 30, such as network links, dial-up links, wireless links, hardwired links, any other suitable communication links, or any suitable combination of such links.

[0030] User device 30 may include any one or more devices suitable for video playback, the text to be replaced, and the definition of the replacement text. User device 30 may include any suitable type of device; for example, user device 30 may be any suitable type of device such as a mobile device, tablet computer, laptop computer, desktop computer, wearable computer, game console, media player, or conferencing equipment.

[0031] It should be noted that replacing text in a video on the cloud server is only one application scenario of this application embodiment. The video text replacement method provided in this application embodiment can also be implemented by a local server, client, IoT device, etc., and this application embodiment does not limit this.

[0032] Methods for replacing text in videos

[0033] Figure 2 This is a flowchart of a video text replacement method according to an embodiment of this application. Figure 2 As shown, the text replacement method in this video includes the following steps:

[0034] Step 201: Detect text regions containing text in the video frames of the video data to be processed.

[0035] The video data refers to the video file that needs to be replaced with text. The video file contains one or more video frames. By decompiling the video data, the individual video frames that make up the video file can be obtained. A video frame is an image that constitutes a video; when these frames are displayed continuously on a monitor at a certain frame rate, a video visible to the human eye is formed.

[0036] By performing text detection on video frames, the text regions containing text within the video frames can be identified.

[0037] In one possible implementation, a text detection model can be used to detect text regions containing text within video frames. This model performs text detection on the input image and outputs location information identifying text regions within the image; these text regions are the areas in the input image that contain text. The text detection model is trained on a business dataset. The location information output by the model can be the position of a bounding box; the image within the bounding box represents the text region containing the text.

[0038] After decompiling the video data into frames, each frame can be sequentially input into a text detection model. The model can detect the text in the input frames and output the location information of the text regions within each frame. Depending on the video frame, the number of text regions detected by the model can be arbitrary. For example, the number of text regions is zero when the video frame contains no text, and two when the text in the frame is concentrated in two regions.

[0039] Step 202: Perform character recognition on the text region to determine the text information included in the text region.

[0040] After identifying the text region included in the video frame, text recognition can be performed on the text region to obtain the text information contained therein.

[0041] In one possible implementation, a pre-trained first character recognition model can be used to recognize text within a text region. This first character recognition model can recognize the text in the input image and output the recognized text information. The first character recognition model is trained on a relevant business dataset.

[0042] After identifying the text regions in a video frame, the positional information of the video frame and the text regions can be input into a first character recognition model. The first character recognition model then identifies the text within each text region on the video frame and outputs the text information included in each text region. Alternatively, after identifying the text regions in a video frame, images of each text region can be cropped from the video frame. Then, each image of a text region can be input into the first character recognition model, which then identifies the text in the input images and outputs the text information included in the corresponding text region.

[0043] Step 203: Determine the first text region containing the first text information from the text region.

[0044] The first text information is a pre-defined set of text to be replaced, such as sensitive words or sponsor names. After determining the text information included in each text region, the text information included in each text region is matched with the first text information. If a text region includes the first text information, then that text region is designated as the first text region.

[0045] For example, if there are 3 text regions in a video frame, and text regions 1 and 2 contain the first text information while text region 3 does not contain the first text information, then text regions 1 and 2 are determined as the first text regions.

[0046] Step 204: Erase the first text information in the first text area.

[0047] After identifying the first text region containing the first text information, the first text information within the first text region can be erased. When the first text region may contain multiple pieces of first text information, each piece of first text information within the first text region can be erased separately.

[0048] In one possible implementation, a text erasure model can be used to erase the first text information in the first text region. The text erasure model can erase the specified text in the input image and restore the background pixels, thus outputting an image with the corresponding text erased and the background pixels restored. The text erasure model is trained on a relevant business dataset.

[0049] After identifying the first text region containing the text to be replaced, the video frame and the position information of each first text region in the video frame are input into the text erasure model. The text erasure model then erases the first text information in each first text region of the input video frame and restores the background pixels of each first text region.

[0050] Step 205: Insert the second text information corresponding to the first text information into the video frame after erasing the first text information.

[0051] The second text information is pre-defined text used to replace the first text information. For example, when replacing Chinese subtitles in a video with subtitles in another language, the first text information is the Chinese subtitles in the video, and the second text information is the English subtitles used to replace them. Similarly, when replacing the brand text of sponsor A in a sports video with the brand text of sponsor B, depending on the viewing region, the first text information is the brand text of sponsor A in the sports video, and the second text information is the brand text of sponsor B. During video review, the first text information is text related to politics, pornography, or terrorism in the video, and the second text information is other text used to replace such text.

[0052] It should be understood that after the second text information is inserted into the video frame, the video frames after the insertion of the second text information are organized according to the order of each video frame in the video data to obtain the video data after text replacement. Then, after the video data after text replacement is sent to the user terminal, the user terminal can play the video based on the received video data. The video played does not include the first text information.

[0053] In this embodiment, after detecting text regions in a video frame, the text information included in each text region is identified. Then, based on the text information included in each text region, a first text region containing first text information can be determined. The first text information in each first text region of the video frame is then erased, and second text information is inserted into the video frame after erasing the first text information, thus replacing the first text information in the video with the second text information. By detecting text regions and recognizing text information, the first text information at any position in the video frame can be identified. Therefore, text replacement in the video is no longer limited by the text position, and the first text information can be freely defined according to requirements. This allows for the replacement of any text in the video, meeting the needs of text replacement in different scenarios and making this video text replacement method highly applicable.

[0054] In one possible implementation, after determining the text region containing text in the video frame, not only can text replacement be performed based on the text region, but also subtitle information recognition can be performed based on the text region. Figure 3 This is a flowchart of a subtitle recognition method according to an embodiment of this application. Figure 3 As shown, the subtitle recognition method includes the following steps:

[0055] Step 301: Detect text regions containing text in the video frames of the video data to be processed.

[0056] It should be noted that step 301 above can refer to step 201 in the aforementioned embodiment, and step 301 will not be described again here.

[0057] Step 302: Segment the video frame to determine at least one subtitle region in the video frame.

[0058] After obtaining video frames by deframing the video data, each video frame can be segmented to determine the subtitle area containing subtitle information in each video frame.

[0059] In one possible implementation, when performing page layout segmentation on video frames, the video frames can be input into a pre-trained page layout segmentation model. The model then segments the video frames to identify subtitle regions containing subtitle information. The page layout segmentation model can perform pixel-level classification of the input image to distinguish between subtitle and non-subtitle regions. The page layout segmentation model is trained on a relevant business dataset.

[0060] After inputting a video frame into the layout segmentation model, the model can output the position information of the subtitle region within the video frame. It should be understood that a video frame may or may not include a subtitle region; when a video frame includes a subtitle region, the number of subtitle regions can be one or more.

[0061] Step 303: Based on the positions of each text region and each subtitle region in the video frame, determine the second text region including the subtitle from each text region.

[0062] The text area is the area containing text information, and the subtitle area is the area containing subtitle information. However, the subtitle information included in the subtitle area may not be complete, while the text area includes complete text information within the corresponding area. Therefore, based on the text area and the subtitle area, the text area containing complete subtitle information can be determined as the second text area.

[0063] In one possible implementation, when determining text regions using a text detection model, the model divides the video frame into text regions and non-text regions. Text regions are areas within the video frame that contain text, while non-text regions are areas that do not contain text. Text regions can be multiple unconnected regions. Similarly, when determining subtitle regions using a layout segmentation model, the model distinguishes between subtitle regions and non-subtitle regions within the video frame through pixel-level classification. Subtitle regions are areas within the video frame that contain subtitles, while non-subtitle regions are areas that do not contain subtitles. Subtitle regions can be multiple unconnected regions.

[0064] For video frames containing subtitles, since a video frame may include non-subtitle text, not all text within a text region is subtitles; that is, only a portion of the text within each text region of the video frame is subtitle. The subtitle region is the area in the video frame that contains subtitles, but the subtitles within this region may be incomplete, meaning the subtitle region may only contain a portion of the subtitles. Based on the above explanation, a second text region containing complete subtitles can be determined from each text region based on the subtitle region.

[0065] Step 304: Perform text recognition on the second text region to obtain the subtitle information included in the video frame.

[0066] For a given video frame, after identifying one or more second text regions included in the video frame, text recognition can be performed on each second text region in the video frame to obtain the text information included in each second text region. Then, the text information included in each second text region in the video frame can be determined as the subtitle information included in the video frame.

[0067] In one possible implementation, a pre-trained second character recognition model can be used to recognize characters in the second character region. The second character recognition model can recognize characters in each of the second character regions separately and output the recognized character information. The second character recognition model can be trained on a corresponding business dataset. The second character recognition model can be the same model as the first character recognition model in the aforementioned embodiments, or it can be a different model. For example, the second character recognition model and the first character recognition model can be the same model; this application does not limit this aspect.

[0068] In this embodiment, text regions and subtitle regions in a video frame are detected. A text region is a region in a video frame that includes text, and a subtitle region is a region in a video frame that includes subtitles. A text region includes complete text information but also includes text information other than subtitles, while a subtitle region includes subtitles but the subtitle information is incomplete. Therefore, based on each text region and each subtitle region included in the same video frame, a second text region including subtitles can be determined from each text region, ensuring that the second text region includes complete subtitle information. Furthermore, by performing text recognition on the second text region, complete subtitle information in the video frame can be obtained, ensuring the accuracy of subtitle recognition in the video.

[0069] In one possible implementation, when determining the second text region from each text region based on the position of each text region and each subtitle region in the video frame, a target subtitle region with a number of pixels greater than a preset threshold can be determined from each subtitle region. Each subtitle region includes at least two adjacent pixels. Then, the minimum bounding rectangle of each target subtitle region is determined. Then, based on the position of the minimum bounding rectangle of each text region and each target subtitle region in the video frame, a second text region including the subtitle is determined from each text region. The text region is also a rectangular region.

[0070] In one possible implementation, when training the text detection model, by setting the parameters of the text detection model, the output of the text detection model can be made to be a rectangular region including the text, that is, the text region is a rectangular region.

[0071] The page segmentation model determines the subtitle region in a video frame through pixel-level classification. Therefore, the shape of the subtitle region is irregular; it can be a single pixel or composed of multiple adjacent pixels. Subtitle regions with a small number of pixels are considered interference regions, generated by the page text detection model to detect false non-subtitle text. A pre-set threshold is used to compare the number of pixels in each subtitle region with this threshold. If the number of pixels in a subtitle region is greater than the threshold, it is identified as the target subtitle region and used to determine the second text region containing the subtitle. If the number of pixels in a subtitle region is less than or equal to the threshold, it is identified as an interference region and is not included in the determination process of the second text region.

[0072] Since the subtitle area is an irregular image area, the target subtitle area is also an irregular image area, while the text area is a rectangular area. In order to facilitate the determination of the second text area including the subtitle based on the position of the subtitle area and the text area in the video frame, the minimum bounding rectangle of each target subtitle area is first determined. Then, based on the position of the minimum bounding rectangle of each text area and each target subtitle area in the video frame, the second text area including the subtitle is determined from each text area.

[0073] In this embodiment, subtitle regions with a pixel count greater than a threshold are selected as target subtitle regions. The minimum bounding rectangle of each target subtitle region is determined. Then, based on the positions of the minimum bounding rectangles of each text region and each target subtitle region in the video frame, a second text region containing the subtitle is determined from each text region. Filtering target subtitle regions by the number of pixels they contain eliminates interference from subtitle regions with fewer pixels, improving the accuracy of determining the second text region containing the subtitle from the video frame. Determining the second text region based on the positions of the minimum bounding rectangles of the text regions and target subtitle regions in the video frame, and using the minimum bounding rectangle of the target subtitle region to identify its position in the video frame, improves the efficiency of determining the second text region.

[0074] In one possible implementation, when determining the second text region including the subtitle from each text region based on the position of the minimum bounding rectangle of each text region and each target subtitle region in the video frame, the intersection-union ratio (IUU) of each text region with the minimum bounding rectangle of each target subtitle region in the same video frame can be calculated respectively. Then, the text region with an IUU greater than a preset IUU threshold with the minimum bounding rectangle of any target subtitle region is determined as the second text region.

[0075] The intersection-union ratio of the minimum bounding rectangles of the text region and the target subtitle region is the ratio of the area (or the number of pixels included) of the intersection of the minimum bounding rectangles of the text region and the target subtitle region to the area (or the number of pixels included) of the union of the minimum bounding rectangles of the text region and the target subtitle region.

[0076] For each text region in a video frame, calculate the intersection-union ratio (IUR) of the minimum bounding rectangle of the text region with each target subtitle region in the video frame. If the IUR of the text region with the minimum bounding rectangle of any one or more target subtitle regions is greater than a preset IUR threshold, then the text region is determined as the second text region including the subtitle.

[0077] In this embodiment, since the text region is a rectangular region in the video frame that includes text, and the minimum bounding rectangle of the target subtitle region is a rectangular region in the video frame that includes subtitles, if the intersection and union ratio of the minimum bounding rectangle of a text region and a target subtitle region is large, it indicates that the overlap area between the text region and the target subtitle region is large, and the probability that the text region includes subtitle information is high. Therefore, the text region is determined as the second text region that includes subtitles, thus ensuring the accuracy of subtitle recognition in the video.

[0078] In another possible implementation, when determining the second text region including the subtitle from each text region based on the position of the minimum bounding rectangle of each text region and each target subtitle region in the video frame, the center point of each text region and the target subtitle region can be determined separately. For any text region, when the distance between the center point of the text region and the center point of at least one target subtitle region is less than a preset distance threshold, the text region is determined as the second text region.

[0079] In this embodiment, the distance between the center point of the text region and the center point of the target subtitle region can characterize the degree of overlap between the text region and the target subtitle region. If the distance between the center point of a text region and the center point of a target subtitle region is small, the probability that the text region and the target subtitle region include the same subtitle information is high. Thus, the text region can be determined as the second text region that includes subtitle information, ensuring the accuracy of subtitle information recognition in the video.

[0080] In one possible implementation, when each video frame is input into a pre-trained page segmentation model to determine at least one subtitle region in the video frame, each video frame can be input into the pre-trained page segmentation model to obtain a binary image output by the page segmentation model. In the binary image, the pixel value of a pixel is either a first pixel value or a second pixel value, and the first pixel value and the second pixel value are different. Then, the region in the binary image where the pixel value is the first pixel value is determined as the subtitle region. Each subtitle region includes one pixel with the first pixel value, or includes multiple adjacent pixels with the first pixel value.

[0081] After inputting the video frame into the layout segmentation model, the model performs pixel-level classification of the video frame, distinguishing between subtitle regions and non-subtitle regions. The output of the layout segmentation model is a binary image corresponding to the input video frame. In the binary image, the pixel value of the pixel in the subtitle region is the first pixel value, and the pixel value of the pixel in the non-subtitle region is the second pixel value. For example, the first pixel value is 1 and the second pixel value is 0. Therefore, the subtitle region in the video frame that includes the subtitle can be determined based on the pixel values ​​of each pixel in the binary image.

[0082] In this embodiment, the page segmentation model can convert the input video frame into a binary image. Pixels in the subtitle area and non-subtitle area in the binary image have different pixel values. Then, based on the pixel values ​​of the pixels in the binary image, the subtitle area in the video frame that includes the subtitle can be determined, ensuring that the subtitle area in the video frame can be accurately located and improving the accuracy of text recognition in the video.

[0083] In one possible implementation, when performing text recognition on the second text region to obtain the subtitle information included in the video frame, an image of the second text region can be extracted from the video frame, and then the text in the image of the second text region can be recognized to obtain the subtitle information included in the video frame.

[0084] In one possible implementation, an image of each second text region can be extracted from the video frame, and then the image of each text region can be input into a second text recognition model. The second text recognition model can then recognize the text in the image of each second text region to obtain the subtitle information included in the video frame.

[0085] Since a video frame may contain multiple second text regions—for example, if the video frame includes vertically distributed Chinese and English subtitles, or if the subtitles are long and displayed in two lines—it will result in the video frame containing two second text regions. After identifying the second text regions included in the video frame, an image of each second text region is extracted from the video frame. This image is then input into a second text recognition model for text recognition to obtain the text information within each second text region. This text information is then used as the subtitle information included in the video frame.

[0086] In this embodiment, after determining the second text region in the video frame, images of each second text region are extracted from the video frame, and then each image is input into a second text recognition model for text recognition to obtain the text information in each second text region. The text information included in each second text region in the same video frame is determined as subtitle information. Extracting images of each second text region from the video frame and inputting them into the second text recognition model for text recognition allows for more accurate recognition of the text within each second text region, avoiding mutual interference between different second text regions and improving the accuracy of subtitle recognition in the video. Furthermore, inputting only images of the second text regions into the second text recognition model for text recognition reduces the amount of data that the second text recognition model needs to process during text recognition, thereby improving the efficiency of subtitle recognition in the video.

[0087] In one possible implementation, the second character recognition model includes a backbone network and an encoding / decoding network. The backbone network is a convolutional neural network (CNN), and the encoding / decoding network is a long short-term memory network (LSTM).

[0088] When identifying the text in each second text region image using the second text recognition model to obtain the subtitle information included in the video frame, the image of the second text region is first input into the backbone network of the second text recognition model to extract visual features. Then, the extracted visual features are input into the encoding and decoding network of the second text recognition model to encode and decode the visual features to obtain the subtitle information included in the video frame.

[0089] In this embodiment, when recognizing text in a second text region using a second text recognition model, visual features are first extracted from the image of the second text region via a backbone network. These visual features indicate the difference between the text image and the image itself. Then, the visual features are input into an encoding / decoding network to encode and decode them, yielding the text content. Extracting visual features using a CNN backbone network and then employing an LSTM encoder-decode structure to encode and decode these features ensures the accuracy of the recognized subtitle information.

[0090] In one possible implementation, there can be multiple first text messages, with different first text messages corresponding to the same or different second text messages.

[0091] The first text information is the text that needs to be replaced. Depending on the actual business requirements, there can be one or more texts to be replaced in the video. The second text information is the text used to replace the first text information. Depending on the actual business requirements, different first texts can be replaced with different second texts (different first texts correspond to different second texts), or different first texts can be replaced with the same second text (different first texts correspond to the same second text).

[0092] A mapping relationship between first and second text information can be pre-established. After recognizing the text information included in each text region, the recognized text information is sequentially matched with each first text information. If a recognized text information successfully matches a first text information, the first text information in that text region is erased, and the aforementioned mapping relationship is queried to determine the second text information mapped to the successfully matched first text information. The determined second text information then replaces the first text information in that text region. If no first text information matches the text information recognized from the text region, no processing is required for the text information in that text region.

[0093] In this embodiment of the application, the first text information is the text in the video that needs to be replaced, and the second text information is the text used to replace the first text information. The first text information can be one or more, and different text information can correspond to the same or different second text information. Thus, a certain text in the video can be replaced, or multiple texts in the video can be replaced. Moreover, different texts in the video can be replaced with a certain text, or different texts in the video can be replaced with different texts, which improves the applicability of the text replacement method in the video.

[0094] In one possible implementation, the text erasure model can be trained using Generative Adversarial Networks (GANs). Training the GAN on a business dataset yields a text erasure model. When video frames are input into the model, it effectively erases information from text regions and restores the background pixels.

[0095] In this embodiment, the generative adversarial network is used for image quality assessment. By training the generative adversarial network on a large-scale sample set, a text erasure model that can be used for text erasure is obtained. Compared with traditional rule-based or vision-based text erasure methods, it can more thoroughly erase the text in the image and make the difference between the restored area and the background area smaller, so that the text erasure has a better effect.

[0096] In one example, the text detection model can detect text in video frames based on the DBNet++ algorithm, while the page segmentation model can identify subtitle regions from video frames based on the BiSeNet-v2 algorithm. The first text recognition model can identify text within text regions based on the Master algorithm, employing a Transformer-encoder-decoder architecture to recognize the text, and its output is the text recognition result. The second text recognition model can be based on the SAR algorithm, first using a CNN backbone network to extract visual features from the input image, then using an LSTM encoder-decode structure to encode and decode the features to obtain the text content, thereby obtaining the subtitle content.

[0097] In one possible implementation, when inserting the second text information corresponding to the first text information into a video frame after erasing the first text information, the second text information can be inserted into the position in the video frame where the first text information was not erased.

[0098] In this embodiment of the application, after erasing the first text information in the video frame, the second text information used to replace the first text information can be inserted into the second text information in the video frame where the first text information was previously located, so that the second text information is in the same position as the first text information before replacement, ensuring that the video viewing effect is not affected after the text replacement, and improving the viewing effect.

[0099] Video playback method

[0100] Regarding the video text replacement scheme provided in this application embodiment for application scenarios such as Virtual Reality (VR) and Augmented Reality (AR), this application embodiment provides a video playback device, which includes a virtual reality device or an augmented reality device. For example... Figure 4 As shown, the video playback method includes the following steps:

[0101] Step 401: Detect text regions containing text in the video frames of the video data to be processed;

[0102] Step 402: Perform character recognition on the text region to determine the text information included in the text region;

[0103] Step 403: Determine the first text region containing the first text information from the text region;

[0104] Step 404: Erase the first text information within the first text area;

[0105] Step 405: Insert the second text information corresponding to the first text information into the video frame after erasing the first text information;

[0106] Step 406: Render the video frame after inserting the second text information onto the display of the video playback device.

[0107] It should be noted that, Figure 4 The embodiments shown are specific applications of the text replacement scheme in the embodiments of this application. For specific text replacement schemes, please refer to the description in the foregoing embodiments, which will not be repeated here.

[0108] Video text replacement device

[0109] Corresponding to the above method embodiments, Figure 5 A schematic diagram of a text replacement device in a video is shown. Figure 5 As shown, the text replacement device 500 in the video includes:

[0110] The text detection unit 501 is used to detect text regions containing text in the video frames of the processed video data;

[0111] The character recognition unit 502 is used to perform character recognition on the text region and determine the text information included in the text region;

[0112] Matching unit 503 is used to determine a first text region including first text information from the text region;

[0113] The erasing unit 504 is used to erase the first text information within the first text area;

[0114] Replacement unit 505 is used to insert second text information corresponding to the first text information into the video frame after erasing the first text information.

[0115] In this embodiment, the text detection unit 501 detects text regions in the video frame, the text recognition unit 502 recognizes the text information included in the text regions, the matching unit 503 determines the first text region including the first text information based on the text information included in each text region, the erasing unit 504 erases the first text information in each first text region of the video frame, and the replacement unit 506 inserts the second text information into the video frame after erasing the first text information, thus replacing the first text information in the video with the second text information. Through text region detection and text information recognition, the first text information at any position in the video frame can be identified. Therefore, text replacement in the video is no longer limited by the text position, and the first text information can be freely defined according to requirements, thereby allowing replacement of any text in the video. This meets the needs of text replacement in videos under different scenarios, making this video text replacement method highly applicable.

[0116] It should be noted that the video text replacement device in this embodiment is used to implement the corresponding video text replacement method in the foregoing method embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0117] electronic devices

[0118] Figure 6 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Specific embodiments of this application do not limit the specific implementation of the electronic device. Figure 6 As shown, the electronic device may include: a processor 602, a communications interface 604, a memory 606, and a communications bus 608. Wherein:

[0119] The processor 602, communication interface 604, and memory 606 communicate with each other via communication bus 608.

[0120] Communication interface 604 is used for communication with other electronic devices or servers.

[0121] The processor 602 is used to execute program 610, specifically to execute the relevant steps in any of the aforementioned video text replacement method embodiments.

[0122] Specifically, program 610 may include program code that includes computer operation instructions.

[0123] The processor 602 may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The smart device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0124] RISC-V is an open-source instruction set architecture based on the Reduced Instruction Set Computing (RISC) principle. It can be applied to various aspects of microcontrollers and FPGA chips, specifically in areas such as IoT security, industrial control, mobile phones, and personal computers. Because its design considers small size, speed, and low power consumption, it is particularly suitable for modern computing devices such as warehouse-scale cloud computers, high-end mobile phones, and tiny embedded systems. With the rise of AIoT (Artificial Intelligence of Things), the RISC-V instruction set architecture is receiving increasing attention and support and is expected to become the next generation of widely used CPU architecture.

[0125] The computer operation instructions in this embodiment can be computer operation instructions based on the RISC-V instruction set architecture. Correspondingly, the processor 602 can be designed based on the RISC-V instruction set. Specifically, the processor chip in the electronic device provided in this embodiment can be a chip designed using the RISC-V instruction set. This chip can execute executable code based on the configured instructions, thereby realizing the text replacement method in the video described above.

[0126] Memory 606 is used to store program 610. Memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0127] Specifically, program 610 can be used to cause processor 602 to execute the video text replacement method in any of the foregoing embodiments.

[0128] The specific implementation of each step in program 610 can be found in the corresponding steps and units described in any of the aforementioned video text replacement method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the aforementioned method embodiments, and will not be repeated here.

[0129] The electronic device in this application detects text regions in a video frame, identifies the text information included in each text region, and then determines a first text region containing first text information based on the text information included in each text region. The first text information in each first text region of the video frame is then erased, and second text information is inserted into the video frame after erasing the first text information, thus replacing the first text information in the video with the second text information. By detecting text regions and recognizing text information, first text information at any position in the video frame can be identified. Therefore, text replacement in the video is no longer limited by the text position, and the first text information can be freely defined according to requirements. This allows for the replacement of any text in the video, meeting the needs of text replacement in different scenarios and making this video text replacement method highly applicable.

[0130] Computer storage media

[0131] This application also provides a computer-readable storage medium storing instructions for causing a machine to perform the video text replacement method as described herein. Specifically, a system or apparatus equipped with a storage medium storing software program code that implements the functions of any of the embodiments described above, and enabling the computer (or CPU or MPU) of the system or apparatus to read and execute the program code stored in the storage medium.

[0132] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of this application.

[0133] Examples of storage media used to provide program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0134] Computer program products

[0135] This application also provides a computer program product, including computer instructions that instruct a computing device to perform any corresponding operation in the above-described plurality of method embodiments.

[0136] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0137] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0138] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0139] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. A method for replacing text in a video, comprising: The text regions in the video frames of the video data to be processed are detected. The text region is subjected to character recognition to determine the text information included in the text region; Determine a first text region that includes the first text information from the text region; Based on the position information of the first text region, the first text information within the first text region is erased and the background pixels of the first text region are restored. Insert the second text information corresponding to the first text information into the video frame after erasing the first text information; The method further includes: Based on the pixels of the video frame, the video frame is segmented to determine at least one subtitle region in the video frame, wherein the subtitle region is an irregular image region; Based on the position of the minimum bounding rectangle of the text region and the subtitle region, a second text region including the subtitle is determined from each of the text regions.

2. The method according to claim 1, wherein, After determining a second text region including the subtitle from each of the text regions based on the position of the minimum bounding rectangle of the text region and the subtitle region, the method further includes: The second text region is subjected to text recognition to obtain the subtitle information included in the video frame.

3. The method according to claim 1, wherein, Based on the position of the minimum bounding rectangle of the text region and the subtitle region, a second text region including the subtitle is determined from each of the text regions, including: From each of the subtitle regions, a target subtitle region is determined that includes more than a preset number threshold of pixels, wherein each subtitle region includes at least two adjacent pixels; Determine the minimum bounding rectangle for each of the target subtitle regions; Based on the position of the minimum bounding rectangle of each text region and each target subtitle region in the video frame, a second text region including the subtitle is determined from each text region, wherein the text region is a rectangular region.

4. The method according to claim 3, wherein, The step of determining a second text region including the subtitle from each of the text regions based on the position of the minimum bounding rectangle of each of the text regions and each of the target subtitle regions in the video frame includes: Calculate the intersection-union ratio (IUU) of the minimum bounding rectangle of each text region and each target subtitle region in the same video frame; The text region whose intersection-union ratio with the smallest bounding rectangle of any of the target subtitle regions is greater than a preset intersection-union ratio threshold is determined as the second text region.

5. The method according to claim 2, wherein, The step of performing text recognition on the second text region to obtain the subtitle information included in the video frame includes: Extract an image of the second text region from the video frame; The text in the image of the second text region is identified respectively to obtain the subtitle information included in the video frame.

6. The method according to claim 5, wherein, The step of identifying the text in the image of the second text region to obtain the subtitle information included in the video frame includes: The image of the second text region is input into the backbone network to extract visual features, wherein the backbone network is a convolutional neural network; The visual features are input into an encoding / decoding network, and the visual features are encoded / decoded to obtain the subtitle information included in the video frame. The encoding / decoding network is a long short-term memory network.

7. The method according to claim 1, wherein, There are multiple first text messages, and different first text messages correspond to the same or different second text messages.

8. The method according to any one of claims 1-7, wherein, The step of inserting the second text information corresponding to the first text information into the video frame after erasing the first text information includes: The second text information corresponding to the first text information is inserted into the position in the video frame where the first text information was not erased.

9. A video playback method applied to a video playback device, the video playback device including a virtual reality device or an augmented reality device, the method comprising: The text regions in the video frames of the video data to be processed are detected. The text region is subjected to character recognition to determine the text information included in the text region; Determine a first text region that includes the first text information from the text region; Based on the position information of the first text region, the first text information within the first text region is erased and the background pixels of the first text region are restored. Insert the second text information corresponding to the first text information into the video frame after erasing the first text information; The video frame after inserting the second text information is rendered onto the display of the video playback device; The method further includes: Based on the pixels of the video frame, the video frame is segmented to determine at least one subtitle region in the video frame, wherein the subtitle region is an irregular image region; Based on the position of the minimum bounding rectangle of the text region and the subtitle region, a second text region including the subtitle is determined from each of the text regions.

10. A video text replacement device, comprising: The text detection unit is used to detect text regions that contain text in the video frames of the video data to be processed; A character recognition unit is used to perform character recognition on the text region and determine the text information included in the text region; A matching unit is configured to determine a first text region including first text information from the text region; The erasing unit is used to erase the first text information within the first text area and restore the background pixels of the first text area based on the position information of the first text area. A replacement unit is used to insert second text information corresponding to the first text information into the video frame after erasing the first text information; The video text replacement device further includes: a subtitle recognition unit, used to perform layout segmentation on the video frame based on the pixels of the video frame, and determine at least one subtitle region in the video frame, wherein the subtitle region is an irregular image region; and to determine a second text region including the subtitle from each of the text regions based on the position of the minimum bounding rectangle of the text region and the subtitle region.

11. An electronic device, comprising: The processor, memory, communication interface, and communication bus communicate with each other through the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the video text replacement method as described in any one of claims 1-8, or to perform the operation corresponding to the video playback method as described in claim 9.

12. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the video text replacement method as claimed in any one of claims 1-8, or the video playback method as claimed in claim 9.

13. A computer program product comprising computer instructions that instruct a computing device to perform a video text replacement method as claimed in any one of claims 1-8, or to perform a video playback method as claimed in claim 9.