Video text line detection enhancement method and device, and electronic equipment

By injecting semantic features into the visual features of video images and performing feature enhancement, the problem that visual modalities cannot capture semantic characteristics in existing technologies is solved, thereby improving the accuracy and robustness of video text line detection.

CN120976903APending Publication Date: 2025-11-18WONDERSHARE TECH (HUNAN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510849583.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing technologies, video text detection methods rely on visual modalities, which cannot fully capture the semantic characteristics of text, resulting in insufficient detection accuracy and robustness.

Method used

By injecting semantic features into the visual features of video images, and using the semantic vectors corresponding to the semantic features for feature enhancement, including creating an initial feature map, filling semantic vectors, and performing self-distillation, combined with a semantic conditional batch normalization (BN) strategy, the detection performance of video text lines is improved.

Benefits of technology

It improves the accuracy and robustness of video text line detection, and is suitable for text detection tasks in complex video scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976903A_ABST
    Figure CN120976903A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video text line detection enhancement method. The method comprises the following steps: acquiring visual features of a video image; semantic features are injected into the visual features of the video image; and performing display feature enhancement on the video features by using the semantic vectors corresponding to the semantic features so as to enhance the detection of the video text lines. According to the video text line detection enhancement method provided by the embodiment of the invention, the visual features of the video image are enhanced by using the semantic features, so that the video text line detection effect is improved. The embodiment of the invention further provides a video text line detection enhancement device and electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of image processing, in particular, to a video text line detection enhancement method and device and electronic equipment. BACKGROUND

[0002] Scene text detection and recognition are always two core tasks in the field of computer vision. Automatically detecting and recognizing text information in images plays a key role in many practical applications, such as real-time translation, assisting blind people in traveling, personalized advertisement recommendation, automatic driving system optimization, intelligent robot interaction, online education content processing, and sentiment analysis. Text detection, as a pre-process of text recognition, its accuracy directly affects the accuracy of subsequent recognition results. In addition, text detection also plays an indispensable role in information extraction and document layout structuring.

[0003] Compared with general object detection tasks, text detection has significant differences. Text instances contain characters inside, which carry low-level but high-density semantic information. Current text detection methods mainly focus on visual modalities, and locate text regions by analyzing color, texture, shape and other visual features in images. However, relying only on visual modalities may not fully capture the semantic characteristics of text, thereby affecting the accuracy and robustness of detection. SUMMARY

[0004] To solve the problems in the prior art, embodiments of the present application provide a video text line detection enhancement method, device and electronic equipment, which enhances the visual features of video images using semantic features, thereby improving the effect of video text line detection.

[0005] In a first aspect, embodiments of the present application provide a video text line detection enhancement method, comprising:

[0006] obtaining visual features of a video image;

[0007] injecting semantic features into the visual features of the video image; and

[0008] using a semantic vector corresponding to the semantic features to perform feature enhancement on the video features to enhance detection of the video text line.

[0009] Further, the injecting semantic features into the visual features of the video image comprises:

[0010] creating an initialization feature map with a size consistent with that of a visual feature map corresponding to the visual features of the video image;

[0011] filling the semantic vector of the video image in the initialization feature map to obtain a semantic feature map corresponding to the video image; and

[0012] The visual feature corresponding to the visual feature of the video image and the semantic feature map are subjected to self-distillation processing to inject semantic features into the visual feature of the video image.

[0013] Further, the creating of the initialization feature map with a size consistent with the visual feature map corresponding to the visual feature of the video image comprises:

[0014] The multi-dimensional visual feature of the video image is obtained through a learnable embedding vector and a prompt fine-tuning method; and

[0015] The multi-dimensional visual feature of the video image is subjected to self-distillation processing to create the initialization feature map with a size consistent with the visual feature map corresponding to the visual feature of the video image.

[0016] Further, the multi-dimensional visual feature is a 768-dimensional visual feature.

[0017] Further, the filling of the semantic vector of the video image in the initialization feature map to obtain the semantic feature map corresponding to the video image comprises:

[0018] The semantic vector of the video image is filled in the initialization feature map to obtain the semantic feature map corresponding to the video image by using a GT text box.

[0019] Further, the feature enhancement of the video feature by using the semantic vector corresponding to the semantic feature comprises:

[0020] The feature enhancement of the video feature by using the semantic vector corresponding to the semantic feature is based on a semantic conditional BN strategy to enhance the detection of the video text line.

[0021] Further, the feature enhancement of the video feature by using the semantic vector corresponding to the semantic feature is based on a semantic conditional BN strategy.

[0022] The semantic vector corresponding to the semantic feature is subjected to MLP conversion of dimensions to obtain learnable parameter vectors respectively;

[0023] The video feature is subjected to display feature enhancement by using the learnable parameter vector.

[0024] In a second aspect, the embodiments of the present application further provide a video text line detection enhancement device, comprising:

[0025] A visual feature acquisition module is configured to acquire a visual feature of a video image.

[0026] a semantic feature injection module configured to inject semantic features into visual features of the video image; and

[0027] a video feature enhancement module configured to perform feature enhancement on the video features using semantic vectors corresponding to the semantic features, so as to enhance detection of the video text line.

[0028] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor is configured to implement the video text line detection enhancement method according to the first aspect when running the program.

[0029] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program for implementing the video text line detection enhancement method according to the first aspect.

[0030] In a fifth aspect, a computer program product is provided, which stores a computer program for implementing the video text line detection enhancement method according to the first aspect.

[0031] The embodiments of the present application have the following beneficial effects:

[0032] In the video text line detection enhancement method of the embodiments of the present application, first, visual features of a video image are obtained, then semantic features are injected into the visual features of the video image, and finally, feature enhancement on the video features is performed using semantic vectors corresponding to the semantic features, so as to enhance detection of the video text line. The video text line detection enhancement method of the embodiments of the present application enhances the visual features of the video image using semantic features, thereby improving the precision and robustness of video text line detection. BRIEF DESCRIPTION OF DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from the structures shown in the drawings without creative labor.

[0034] Figure 1 A flowchart of the video text line detection enhancement method provided by the embodiments of the present application is shown in the figure.

[0035] Figure 2A network architecture diagram of the video text line detection enhancement method provided by the embodiment of the present application is shown.

[0036] Figure 3 A specific details diagram of the SSCBN model used in the video text line detection enhancement method provided by the embodiment of the present application is shown.

[0037] Figure 4 A preliminary test effect diagram of the video text line detection enhancement method provided by the embodiment of the present application is shown.

[0038] Figure 5 A structure block diagram of the video text line detection enhancement device provided by the embodiment of the present application is shown.

[0039] Figure 6 A structure diagram of the electronic device provided by the embodiment of the present application is shown.

[0040] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0041] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the protection scope of the present application.

[0042] In the specification and claims of the present application and the above-mentioned drawings, the terms “first” and “second” are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by “first” and “second” can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of “multiple” is two or more. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0043] Figure 1 A flowchart of the video text line detection enhancement method according to an embodiment of the present application is shown.

[0044] As shown in Figure 1 The video text line detection enhancement method of the embodiment of the present application includes the following steps:

[0045] S101: acquiring visual features of a video image;

[0046] S102: injecting semantic features into the visual features of the video image; and

[0047] S103: Perform feature enhancement on the video features by using the semantic vector corresponding to the semantic features to enhance the detection of the video text line.

[0048] Specifically, in some embodiments of the present application, a pre-trained CNN (such as ResNet-50, HRNet) or a Transformer architecture (such as Swin Transformer) is used to extract the spatial features of the video frames, a time sequence modeling module (such as 3D CNN, LSTM) is used to capture dynamic time sequence information, and a visual feature map containing texture, edge, and motion trajectory is generated. Then, a pre-trained NLP model (such as BERT, CLIP) is used to generate a semantic vector of the text description, and a keyword vector is extracted by using ASR (Automatic Speech Recognition) or OCR (Optical Character Recognition) results. Cross-Attention or Concatenation is used to map the semantic vector to the visual feature space, and modality alignment is achieved. Finally, a spatial attention map is generated based on semantic similarity, dynamic weighting of visual features is performed, and an adversarial training or feature reconstruction (such as U-Net decoder) is introduced to generate a fine-grained text boundary heat map, thereby enhancing the detection capability of fuzzy and occluded text.

[0049] Therefore, in the video text line detection enhancement method of the embodiments of the present application, the visual features of the video image are first obtained, then the semantic features are injected into the visual features of the video image, and finally the video features are enhanced by using the semantic vector corresponding to the semantic features to enhance the detection of the video text line. The video text line detection enhancement method of the embodiments of the present application enhances the visual features of the video image by using semantic features, thereby improving the precision and robustness of video text line detection.

[0050] Further, the injecting of the semantic features into the visual features of the video image comprises:

[0051] creating an initialization feature map with a size consistent with the visual feature map corresponding to the visual features of the video image;

[0052] filling the initialization feature map with the semantic vector of the video image to obtain a semantic feature map corresponding to the video image; and

[0053] performing self-distillation processing on the visual feature map and the semantic feature map corresponding to the visual features of the video image to inject semantic features into the visual features of the video image.

[0054] Specifically, in order to inject semantic features into the visual features of a video image, an initialization feature map (for example, a full zero matrix) consistent with the spatial resolution of the visual feature map needs to be created first. The number of channels can be aligned with the visual feature map or adjusted through linear mapping to ensure the dimensional compatibility of subsequent fusion. Then the semantic vector (for example, the [CLS] vector of BERT or the ASR keyword) is filled into each spatial position of the initialization feature map to generate a semantic feature map. Finally, the teacher-student model, that is, the teacher network takes the semantic feature map as input to generate a high-dimensional semantic representation, and the student network takes the visual feature map as input to generate a visual representation. The similarity of the outputs of the teacher and the student is constrained by the KL divergence (KLD) or the mean square error (MSE), forcing the visual feature map to learn the key information in the semantic feature map. Through the distillation loss and the original detection loss (for example, cross-entropy, IoU loss), the weights can be adjusted through hyperparameters (such as λ = 0.1-0.5). Therefore, the detection enhancement method for video text lines according to the embodiments of the present application enhances the semantic relevance while retaining the spatial information of the visual features through the above implicit feature fusion, and is suitable for text detection tasks in complex video scenes.

[0055] Further, the creating an initialization feature map consistent with the size of the visual feature map corresponding to the visual features of the video image comprises:

[0056] obtaining the multi-dimensional visual features of the video image through the learnable embedding vector and the prompt fine-tuning method; and

[0057] performing self-distillation processing on the multi-dimensional visual features of the video image to create an initialization feature map consistent with the size of the visual feature map corresponding to the visual features of the video image.

[0058] Specifically, a learnable embedding vector (e.g., the [CLS] token in Transformer or an extra trainable parameter) is introduced to concatenate or interact with the visual features of the video image (e.g., the output of CNN / Transformer) to generate a multi-dimensional feature representation containing global semantics. During feature extraction, the visual features are fine-tuned by a lightweight prompt (e.g., a continuous prompt vector or a text-guided attention mask) to enhance sensitivity to the text region. Then, the channel number of the multi-dimensional visual features is mapped to the same dimension (e.g., 256 dimensions) as the visual feature map by a lightweight network (e.g., 1x1 convolution or linear layer). If the multi-dimensional feature is a global vector (e.g., the [CLS] output), it is duplicated to each spatial location; if it is a region feature (e.g., ROI feature), it is filled by position. Finally, through self-distillation processing, the KL divergence or feature similarity (e.g., cosine similarity) between the teacher and student outputs is calculated to force the initialized feature map to learn the spatial distribution of the original feature and semantic information. Therefore, through the above method, the initialized feature map not only has spatial consistency with the original feature map, but also incorporates learnable semantic information, laying a foundation for subsequent semantic feature injection and detection enhancement.

[0059] Further, the multi-dimensional visual feature is a 768-dimensional visual feature.

[0060] Specifically, the 768-dimensional visual feature is usually from the output of a pre-trained model such as ViT, CLIP's visual encoder or a custom Transformer, containing global semantic information. A learnable embedding vector matching the 768-dimensional is introduced, such as a 768-length parameter matrix, which is fused with the original feature through weighted summation or concatenation to enhance the representation ability of the text region. And use a lightweight prompt such as a 768-dimensional continuous vector of length 10 to perform cross-attention calculation with the visual feature, dynamically adjust the feature response, and focus on the text region.

[0061] Further, the filling of the semantic vector of the video image in the initialized feature map to obtain the semantic feature map corresponding to the video image comprises:

[0062] Using the GT text box, the semantic vector of the video image is filled in the initialized feature map to obtain the semantic feature map corresponding to the video image.

[0063] Specifically, the semantic vector can come from the [CLS] vector of a pre-trained language model such as BERT or a video text recognition result such as a word embedding of an ASR / OCR keyword, which needs to be aligned with the text content of the video image. The semantic vector is L2 normalized using the GT text box to ensure stable numerical range and avoid feature imbalance when padding. According to the GT text box (Ground Truth, labeled text box) of the video image, the semantic vector is accurately padded to the corresponding area of the initialized feature map to generate a semantic feature map. Therefore, through the above process, the semantic feature map not only retains semantic information but also aligns with the spatial structure of the visual feature map, providing high-quality input for subsequent cross-modal feature fusion and detection enhancement.

[0064] Further, the feature enhancement of the video feature using the semantic vector corresponding to the semantic feature includes:

[0065] Based on the semantic conditional BN strategy, the feature enhancement of the video feature using the semantic vector corresponding to the semantic feature is used to enhance the detection of the video text line.

[0066] Specifically, the semantic conditional BN (Batch Normalization) dynamically adjusts the BN parameters by introducing the semantic vector, so that the normalization process adapts to the text features. Based on the semantic vector, conditional BN parameters (γ, β) are generated to specifically enhance the text region in the video feature, which can effectively suppress background noise and further improve the detection accuracy and robustness of the video text line.

[0067] Further, the feature enhancement of the video feature using the semantic vector corresponding to the semantic feature based on the semantic conditional BN strategy includes:

[0068] The semantic vector corresponding to the semantic feature is converted in dimension by an MLP to obtain a learnable parameter vector respectively;

[0069] The video feature is enhanced using the learnable parameter vector.

[0070] Specifically, the semantic vector, such as an OCR text embedding or a pre-trained language model output, is usually of a fixed dimension (e.g., 768 dimensions) and needs to be aligned with the number of channels of the video features. The semantic vector is mapped into a learnable parameter vector through a lightweight multi-layer perceptron (MLP, e.g., 2-layer fully connected + ReLU), and the gamma and beta generated by the MLP are fused with the global parameters of batch normalization (BN) to generate conditional BN parameters, and then the video features are normalized, that is, the conditional BN parameters are applied to the video features. Therefore, through the above semantic conditional BN and MLP parameter conversion, the text region in the video features is specifically enhanced, and the background noise is effectively suppressed, which can further improve the robustness and accuracy of video text detection.

[0071] With reference to Figure 2 , Figure 2 The network architecture diagram of the video text line detection enhancement method provided by the embodiments of the present application is shown in FIG. 1. Figure 2 As shown in FIG. 1, the entire framework is improved based on DBnet and mainly divided into two stages:

[0072] Among them, the first stage uses the prompt tuning idea, freezes the clip text encoder and the image encoder, introduces a learnable embedding vector into the input part, and fine-tunes the prompt with a small learning rate; the second stage uses the clip model after fine-tuning to extract text lines, which is mainly divided into three parts:

[0073] 1) The picture is extracted into visual features through the clip image encoder, and the features are input into the DBhead to obtain the final segmentation map, that is, the loss1 part;

[0074] 2) The learnable embedding and the prompt are extracted into 768-dimensional features through the clip text encoder, and a feature map of all zeros is initialized using the self-distillation idea. The size and dimension of the feature map are the same as the feature map extracted by the visual encoder, and the input image is down-sampled in the H and W dimensions. Then, the GT text box is filled with a 768-dimensional semantic vector in the initialized feature map, and thus a semantic feature map is obtained. Then, the feature map and the original visual features are distilled to gradually inject semantic features into the visual features, that is, the loss2 part.

[0075] 3) Based on the semantic conditional BN strategy, the visual features are displayed and enhanced by the semantic vector. As shown in FIG. 3, first, the text vector is converted in dimension through two MLPs to obtain a gamma vector and a beta vector, and the vectors are used to enhance the intermediate features in the head part. Figure 3

[0076] With reference to Figure 4 ,​Figure 4 is a preliminary test effect schematic diagram of the video text line detection enhancement method provided by the embodiment of the present application. As shown in Figure 4 : before introducing the SSCBN (3D Semantic Scene Completion BN, three-dimensional semantic scene completion batch normalization) model proposed in the embodiment of the present application, the shallow channel features extracted by vision are relatively abstract and fuzzy, and the semantic information is rich after 100 channels, such as texture boundary contour and other semantic information. However, the features of the SSCBN branch have rich semantic information from the shallow channel to the high-level channel, which shows that compared with the traditional use of visual information, semantic information is helpful to the extraction of problem details contour and other features. At the same time, through preliminary experiments, the effect after SSCBN is better than that of using only visual features.

[0077] Figure 5 is a structural block diagram of the video text line detection enhancement device 200 of the embodiment of the present application. As shown in Figure 5 , the video text line detection enhancement device 200 of the embodiment of the present application comprises a visual feature acquisition module 210, a semantic feature injection module 220 and a video feature enhancement module 230, wherein:

[0078] The visual feature acquisition module 210 is configured to acquire visual features of a video image.

[0079] The semantic feature injection module 220 is configured to inject semantic features into the visual features of the video image; and

[0080] The video feature enhancement module 230 is configured to perform display feature enhancement on the video features by using a semantic vector corresponding to the semantic features, so as to enhance the detection of the video text line.

[0081] In the video text line detection enhancement device provided by the embodiment of the present application, the visual features of a video image are first acquired, then semantic features are injected into the visual features of the video image, and finally the display feature enhancement is performed on the video features by using a semantic vector corresponding to the semantic features, so as to enhance the detection of the video text line. The video text line detection enhancement device of the embodiment of the present application enhances the visual features of the video image by using semantic features, thereby improving the precision and robustness of the video text line detection.

[0082] It should be noted that the specific implementation mode of the video text line detection enhancement device of the embodiment of the present application is similar to that of the video text line detection enhancement method of the embodiment of the present application, and specific please refer to the description of the method part, this place does not make superfluous.

[0083] Figure 6 is a structural schematic diagram of the electronic device 300 of the embodiment of the present application.

[0084] like Figure 6 As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from the storage section 302 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0085] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. A removable medium 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 310 as needed so that computer programs read from it can be installed into storage section 308 as needed.

[0086] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the functions defined in the electronic device of this application.

[0087] It should be noted that the computer-readable medium in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor electronic device, device or apparatus, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0088] In the present application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution electronic device, device or apparatus. In the present application, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution electronic device, device or apparatus. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination thereof.

[0089] The flowcharts and block diagrams in the drawings illustrate the possible implementation architecture, function and operation of the processing receiving device, method and computer program product according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different order than that shown in the figure. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the function involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based electronic device that performs the specified function or operation, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0090] The units or modules described in the embodiments of the present application can be implemented in the form of software or in the form of hardware. The described units or modules can also be arranged in a processor for executing the program to implement the detection enhancement method of video text lines:

[0091] obtaining visual features of the video image;

[0092] injecting semantic features into the visual features of the video image; and

[0093] performing display feature enhancement on the video features by using semantic vectors corresponding to the semantic features, to enhance the detection of the video text lines.

[0094] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The computer readable storage medium stores one or more programs, and when the programs are executed by one or more processors, the detection enhancement method of video text lines described in the present application is implemented:

[0095] obtaining visual features of the video image;

[0096] injecting semantic features into the visual features of the video image; and

[0097] performing display feature enhancement on the video features by using semantic vectors corresponding to the semantic features, to enhance the detection of the video text lines.

[0098] As another aspect, the present application also provides a computer program product, which can be included in the electronic device described in the above embodiments, or can exist independently without being assembled into the electronic device. The computer program product stores one or more programs, and when the programs are executed by one or more processors, the detection enhancement method of video text lines described in the present application is implemented:

[0099] obtaining visual features of the video image;

[0100] injecting semantic features into the visual features of the video image; and

[0101] performing display feature enhancement on the video features by using semantic vectors corresponding to the semantic features, to enhance the detection of the video text lines.

[0102] The above merely describes the preferred embodiments of the present application, and is not intended to limit the patent scope of the present application. Any equivalent structural changes made according to the content of the present application specification and drawings, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for detection enhancement of video text lines, characterized in that, The method comprises: obtaining visual features of a video image; injecting semantic features into the visual features of the video image; and performing feature enhancement on the video features by using a semantic vector corresponding to the semantic features to enhance detection of the video text line. The injecting of the semantic features into the visual features of the video image comprises:

2. The method of claim 1, wherein, creating an initial feature map with a size consistent with the visual features of the video image; filling the initial feature map with a semantic vector of the video image to obtain a semantic feature map corresponding to the video image; and performing self-distillation processing on the visual feature map and the semantic feature map corresponding to the visual features of the video image to inject the semantic features into the visual features of the video image. The creating of the initial feature map with a size consistent with the visual features of the video image comprises:

3. The method of claim 2, wherein, obtaining multi-dimensional visual features of the video image by using a learnable embedding vector and a prompt fine-tuning method; and performing self-distillation processing on the multi-dimensional visual features of the video image to create the initial feature map with a size consistent with the visual features of the video image. The multi-dimensional visual features are 768-dimensional visual features.

4. The method of claim 3, wherein, The filling of the initial feature map with the semantic vector of the video image to obtain the semantic feature map corresponding to the video image comprises:

5. The method of claim 2, wherein, filling the initial feature map with the semantic vector of the video image by using a GT text box to obtain the semantic feature map corresponding to the video image. The performing of the feature enhancement on the video features by using the semantic vector corresponding to the semantic features comprises:

6. The method of claim 1, wherein, performing the feature enhancement on the video features by using the semantic vector corresponding to the semantic features based on a semantic conditional BN strategy to enhance the detection of the video text line. The performing of the feature enhancement on the video features by using the semantic vector corresponding to the semantic features based on the semantic conditional BN strategy comprises:

7. The method of claim 6, wherein, the semantic vector corresponding to the semantic features is converted in dimension by an MLP to obtain learnable parameter vectors respectively; and the feature enhancement is performed on the video features by using the learnable parameter vectors. The method comprises:

8. An apparatus for detecting enhancement of a video text line, characterized by, a visual feature acquisition module configured to obtain visual features of a video image; a semantic feature injection module configured to inject semantic features into the visual features of the video image; and a video feature enhancement module configured to perform feature enhancement on the video features by using a semantic vector corresponding to the semantic features to enhance detection of the video text line. The computer program product comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor is configured to implement the video text line detection enhancement method according to any one of claims 1-7 when executing the program. The computer readable storage medium stores a computer program, and the computer program is used to implement the video text line detection enhancement method according to any one of claims 1-7.

9. An electronic device, comprising: ​ 10. A computer-readable storage medium, characterized in that, ​

Citation Information

Patent Citations

  • Natural scene text recognition method based on characterization batch normalization

    CN114219990A

  • Text detection method and device, computing equipment and storage medium

    CN114758332A

  • Data processing method and device, equipment, storage medium and program product

    CN117194655A

  • Text detection method and device, text detection model optimization method and device and data annotation method and device

    CN117275005A

  • Training method of forged voice detection model, forged voice detection method and device

    CN118366433A