Text attribute recognition method and device, electronic equipment and storage medium

By utilizing text line information from consecutive frames in a video and a pre-trained model for feature encoding and inter-frame feature fusion, the problem of low accuracy in single-frame image recognition is solved, achieving high-accuracy recognition of video text attributes.

CN116311200BActive Publication Date: 2026-05-12BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2023-02-03
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing text attribute recognition methods based on single-frame images have low accuracy.

Method used

By acquiring text line-related information of text lines in each video frame based on at least two consecutive video frames in the video to be identified, and using a pre-trained video attribute recognition model for feature encoding, attribute recognition is performed by combining inter-frame features.

Benefits of technology

The accuracy of video text attribute recognition has been improved by combining the inter-frame features of text lines from multiple video frames for attribute recognition, thereby enhancing the accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311200B_ABST
    Figure CN116311200B_ABST
Patent Text Reader

Abstract

The method and device for text attribute recognition, the electronic device and the storage medium provided by the embodiments of the present disclosure relate to the technical field of computers. In the present disclosure, text line related information of each video frame is obtained based on text lines in at least two continuous video frames in a video to be recognized; the text line related information of each video frame is subjected to feature coding by a pre-trained video attribute recognition model to obtain text line features corresponding to each video text line; text line inter-frame features corresponding to each text line are obtained based on the text line features corresponding to at least two video frames by the video attribute recognition model; and attribute recognition is performed on the text line inter-frame features to determine text attributes corresponding to each text line in the video frame to be recognized. In this way, in the process of video text analysis, the text line inter-frame features in continuous multiple video frames are taken into account, and the accuracy of video text attribute recognition in each frame is improved to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a text attribute recognition method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of computer technology, short videos are widely welcomed by users due to their timeliness, speed, dynamic visuals, and enhanced entertainment value. As a result, there is a growing demand for applications based on the attributes of text content within videos, such as video search and video text extraction. In other words, there is an increasing number of applications that utilize the text attributes of text content within videos.

[0003] In related technologies, text attribute recognition is mainly performed on the text in a single frame image. However, this recognition method has low accuracy because it only uses the text information of a single frame image for analysis and recognition. Summary of the Invention

[0004] This disclosure provides a text attribute recognition method, apparatus, electronic device, and storage medium to at least address the problem of how to improve the accuracy of text attribute recognition. The technical solution of this disclosure is as follows:

[0005] According to a first aspect of the present disclosure, a text attribute recognition method is provided, comprising:

[0006] Based on the text lines in at least two consecutive video frames in the video to be identified, obtain the text line related information of the text lines in each video frame;

[0007] By using a pre-trained video attribute recognition model, the text line information of each text line in each video frame is encoded to obtain the text line features corresponding to each text line in each video frame.

[0008] Based on the text line features corresponding to at least two video frames, the video attribute recognition model obtains the inter-frame features of each text line.

[0009] The text line inter-frame features are analyzed to identify the text attributes corresponding to each text line in the video frame to be identified.

[0010] Optionally, the text line-related information includes text information, image information, and location information; the step of obtaining the text line-related information of the text lines in each of at least two consecutive video frames in the video to be identified includes:

[0011] For any one of the at least two video frames, local image features of the region where each text line is located in the video frame are obtained, and used as the image information;

[0012] The text content of each text line in the video frame is obtained as the text information;

[0013] Obtain the position coordinates of each text line in the video frame as the position information.

[0014] Optionally, the step of using a pre-trained video attribute recognition model to perform feature encoding on the text line-related information of each text line in each video frame to obtain the text line features corresponding to each text line in each video frame includes:

[0015] The first processing layer in the video attribute recognition model generates first splicing information for each text line based on the position information and text information of each text line, and generates second splicing information for each text line based on the position information and image information of each text line.

[0016] The first concatenation information and the second concatenation information of each text line are concatenated to obtain the concatenation information of each text line;

[0017] Based on the concatenation information of each of the text lines, the text line features of each of the text lines are generated.

[0018] Optionally, the method further includes:

[0019] For any of the video frames, the first processing layer determines the text sequence identifier of the text information and the image sequence identifier of the image information in the text line of the video frame; the text sequence identifier and the image sequence identifier of the same text line are the same;

[0020] A continuous first position identifier is set for the target information of all text lines in the video frame; the target information includes the text information and image information.

[0021] The step of generating first concatenation information for each text line based on the position information and text information of each text line includes: concatenating the position information, text information, first position identifier of the text information, and text sequence identifier of each text line to obtain the first concatenation information for each text line;

[0022] The step of generating second splicing information for each text line based on the position information and the image information of each text line includes: splicing the position information, the image information, the first position identifier of the image information, and the image sequence identifier of each text line to obtain the second splicing information for each text line.

[0023] Optionally, the step of obtaining the inter-frame features of each text line based on the text line features corresponding to the at least two video frames using the video attribute recognition model includes:

[0024] The first position identifier corresponding to the text information of the text lines in the at least two video frames is obtained through the second processing layer in the video attribute recognition model.

[0025] For any text line in any video frame, the inter-frame features of the text line are determined based on the text line features corresponding to the text line, the display duration, and the first position identifier; the display duration is used to characterize the duration of the text line in the at least two video frames.

[0026] Optionally, determining the inter-frame features of the text line based on the text line features corresponding to the text line, the display duration, and the first position identifier includes:

[0027] The display duration and the first location identifier are encoded respectively to obtain a display duration vector and a first location identifier vector;

[0028] The text line features, the display duration vector, and the first position identifier vector are concatenated to obtain the text line frame features.

[0029] Optionally, the step of performing attribute recognition based on the inter-frame features of the text lines to determine the text attributes corresponding to each text line in the video frame to be recognized includes:

[0030] For any given text line, attribute recognition is performed based on the inter-frame features of the text line to obtain the text type attribute, text source attribute, and text visual attribute corresponding to the text line.

[0031] Optionally, the video attribute recognition model is trained in the following manner:

[0032] The basic model is trained based on a first specified training task and a first sample video, and the basic model is trained based on a second specified training task and each video frame in the first sample video to obtain a recognition model to be trained; the first specified training task includes a text line feature consistency task; the second specified training task includes a text content recovery task, an image region recovery task and / or an image region and text content alignment task.

[0033] Use the text lines from at least two second sample video frames as input to the recognition model to be trained, and obtain the text attributes predicted by the recognition model to be trained;

[0034] Based on the text attributes and the text attribute labels of the text lines in the at least two second sample video frames, the parameters of the recognition model to be trained are adjusted; the text attribute labels are used to characterize the true text attributes of the text lines.

[0035] If the training recognition model reaches the stopping condition, the training recognition model that has reached the stopping condition is determined as the video attribute recognition model.

[0036] Optionally, the step of training the base model based on a first specified training task and a first sample video, and training the base model based on a second specified training task and each video frame in the first sample video to obtain a recognition model to be trained, includes:

[0037] The first sample video is used as the input to the base model to obtain the text line features of each text line in the first sample video output by the base model.

[0038] Furthermore, for any video frame in the first sample video, based on the text content, image region, and masked text content and image region in the video frame, the text content and image region output by the basic model are obtained;

[0039] Based on the text line features, text content, and image region output by the base model, the parameters of the base model are adjusted.

[0040] If the base model reaches the stopping condition, the base model that has reached the stopping condition is determined as the recognition model to be trained.

[0041] According to a second aspect of the present disclosure, a text attribute recognition device is provided, comprising:

[0042] The first acquisition module is used to acquire text line related information of the text lines in each of at least two consecutive video frames in the video to be identified.

[0043] The first encoding module is used to perform feature encoding on the text line information of the text lines in each video frame using a pre-trained video attribute recognition model, so as to obtain the text line features corresponding to each text line in each video frame.

[0044] The first recognition module is used to perform attribute recognition based on the text line features corresponding to the at least two video frames using the video attribute recognition model, so as to determine the text attributes corresponding to each text line in each video frame.

[0045] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0046] processor;

[0047] Memory used to store the processor's executable instructions;

[0048] The processor is configured to execute the instructions to implement the method as described in any one of the first aspects.

[0049] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, causes the electronic device to perform the method as described in any one of the first aspects.

[0050] According to a fifth aspect of the present disclosure, a computer program product is provided, the computer program product including readable program instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the method as described in any one of the first aspects.

[0051] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects: In the embodiments of this disclosure, a text attribute recognition method is provided by obtaining text line-related information of text lines in each video frame based on text lines in at least two consecutive video frames in a video to be recognized; using a pre-trained video attribute recognition model to perform feature encoding on the text line-related information of text lines in each video frame to obtain text line features corresponding to each text line in each video frame; using the video attribute recognition model based on the text line features corresponding to at least two video frames to obtain inter-frame features of text lines corresponding to each text line; and performing attribute recognition on the inter-frame features of text lines to determine the text attributes corresponding to each text line in the video frame to be recognized. Thus, by first using a video attribute recognition model to perform feature encoding on the text line-related information contained in a single video frame to obtain the text line features corresponding to the text lines in a single video frame, and then combining the inter-frame features of text lines in at least two consecutive video frames for attribute recognition, the text attributes corresponding to each text line can be obtained. Therefore, in the process of video text analysis, by taking into account the inter-frame features of text lines in multiple consecutive video frames and combining the inter-frame features of text lines in multiple video frames for attribute recognition, the accuracy of text attribute recognition in video can be improved to a certain extent.

[0052] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0053] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0054] Figure 1 This is a flowchart illustrating a text attribute recognition method according to an exemplary embodiment;

[0055] Figure 2 This is a schematic diagram of a processing flow according to an exemplary embodiment;

[0056] Figure 3 A block diagram of a text attribute recognition device according to an exemplary embodiment is shown;

[0057] Figure 4 This is a block diagram illustrating an apparatus for text attribute recognition according to an exemplary embodiment;

[0058] Figure 5 This is a block diagram illustrating another apparatus for text attribute recognition according to an exemplary embodiment. Detailed Implementation

[0059] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0060] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0061] Figure 1 This is a flowchart illustrating a text attribute recognition method according to an exemplary embodiment, such as... Figure 1 As shown, the following steps may be included:

[0062] Step 101: Based on the text lines in at least two consecutive video frames in the video to be identified, obtain the text line related information of the text lines in each video frame.

[0063] In this embodiment of the disclosure, the video to be identified can be a video stored locally on the user terminal, which the user terminal can directly obtain from its local storage. Alternatively, the target video can be a video stored on the servers of various website platforms on the Internet, which the user terminal can download via the network. A video to be identified can contain multiple video frames. At least two consecutive video frames refer to at least two video frames that are adjacent in time sequence. Text lines within video frames can be areas in the video where text appears. For example, the target video can be a short video from a short video platform, and the text lines can be text content contained within the short video.

[0064] Decoding the video to be recognized yields multiple consecutive video frames. Each video frame may contain one or more text lines. Text line-related information can be used to characterize the features of the text lines, including text and layout information. For example, text line-related information can be obtained from video frames through analysis and recognition processing using Optical Character Recognition (OCR) technology.

[0065] Step 102: Use a pre-trained video attribute recognition model to encode the text line information of each text line in each video frame to obtain the text line features corresponding to each text line in each video frame.

[0066] In this embodiment of the disclosure, the video attribute recognition model can be a multi-layer processing model obtained through pre-training. After recognizing the text line-related information corresponding to the text lines in each video frame, the video attribute recognition model encodes the text line-related information in each video frame to obtain the text line features corresponding to the text lines in each video frame. For example, assuming that the video file to be recognized contains m video frames, and each video frame contains n text lines, the video attribute recognition model encodes the n text line-related information corresponding to the n text lines in a single video frame to obtain the n text line features corresponding to the n text lines in that single video frame. Then, the text lines in the m video frames are encoded using the above method to obtain the n text line features corresponding to the n text lines in each video frame.

[0067] Step 103: Based on the text line features corresponding to the at least two video frames, obtain the inter-frame features of each text line using the video attribute recognition model.

[0068] In this embodiment of the disclosure, after obtaining the text line features corresponding to at least two consecutive video frames, the text line features corresponding to the at least two video frames are fused using a video attribute recognition model to obtain the inter-frame features of the text lines corresponding to each text line in each video frame. For example, assuming the video file to be identified contains m video frames, the text line features corresponding to n text lines in the m video frames are obtained using the method in step 102. Then, the video attribute recognition model is used to perform attribute recognition on the n text line features corresponding to the m video frames, finally obtaining the text attributes corresponding to each of the n text lines in the m video frames.

[0069] Step 104: Perform attribute recognition on the inter-frame features of the text lines to determine the text attributes corresponding to each text line in the video frame to be recognized.

[0070] In this embodiment of the disclosure, attribute recognition is performed based on the inter-frame features of each text line. The text attributes corresponding to each text line can be determined according to the inter-frame features of the text lines. Since the inter-frame features of the text lines are obtained by combining the text line features of each text line in at least two video frames, the inter-frame features of the text lines can better reflect the true features of the text lines, and thus obtain more realistic and accurate text attributes.

[0071] In summary, the embodiments of this disclosure obtain text line-related information of text lines in each video frame based on text lines in at least two consecutive video frames in the video to be identified; perform feature encoding on the text line-related information of text lines in each video frame using a pre-trained video attribute recognition model to obtain text line features corresponding to each text line in each video frame; obtain inter-frame features of text lines corresponding to each text line based on the text line features corresponding to at least two video frames using the video attribute recognition model; and perform attribute recognition on the inter-frame features of text lines to determine the text attributes corresponding to each text line in the video frame to be identified. Thus, by first using the video attribute recognition model to perform feature encoding on the text line-related information contained in a single video frame to obtain the text line features corresponding to the text lines in that single video frame, and then combining the inter-frame features of text lines in at least two consecutive video frames for attribute recognition, the text attributes corresponding to each text line can be obtained. Therefore, in the process of video text analysis, by taking into account the inter-frame features of text lines in multiple consecutive video frames and combining the inter-frame features of text lines in multiple video frames for attribute recognition, the accuracy of text attribute recognition in video is improved to a certain extent.

[0072] Optionally, the text line-related information includes text information, image information, and location information.

[0073] Text lines in video frames possess multi-dimensional features, including textual, image, and positional information. Since the semantics, images, and positions of different text lines may vary, these differences will all affect the text attribute recognition results. Incorporating this information as text line-related data enriches the content represented by the text line features, resulting in more comprehensive text line features and facilitating more accurate text attribute identification based on these features. Specifically, textual information can be the text content corresponding to the text line, image information can be the text region features corresponding to the text line, and positional information can be the positional features of the text line within the video frame, which can be represented using coordinates.

[0074] Accordingly, step 101 may include the following steps:

[0075] Step 1011: For any one of the at least two video frames, obtain the local image features of the region where each text line in the video frame is located, as the image information.

[0076] In this embodiment of the disclosure, for a single video frame, image information corresponding to any text line in the video frame can be obtained based on the text information and position information corresponding to that text line. For example, the text information and image information corresponding to any text line can be input into a deep neural network, and the image information corresponding to that text line can be extracted using the ROI Align method. It can be understood that for a video frame containing n text lines, the obtained n image information can be represented as {image1, image2, ... image...} n The image information is used to characterize the local image features of the region where the text line is located.

[0077] Step 1012: Obtain the text content of each text line in the video frame as the text information.

[0078] In this embodiment of the disclosure, OCR technology is used to recognize the text content of each text line in a video frame, extracting the text content from the image and converting it into text information. For example, the text content can be segmented into individual characters to obtain the text information corresponding to the text content in the text line.

[0079] Step 1013: Obtain the position coordinates of each text line in the video frame as the position information.

[0080] In this embodiment, OCR technology can be used to detect the location, range, and layout of text lines, typically including layout analysis and text line detection. The obtained location information is used to characterize which area in the video frame contains text and the size of that text area. The location information can be represented by coordinate pairs, and based on these coordinate pairs, the X and Y coordinates of the text line in the video frame can be obtained, thereby determining the specific location range of the text line in the video frame.

[0081] In this embodiment of the disclosure, by acquiring text information, location information and image information, the text-related information corresponding to the text line can be obtained by combining these three types of information. In the process of obtaining text line features based on text-related information, the content represented by the text line features is enriched, thereby improving the accuracy of text attribute recognition to a certain extent.

[0082] Optionally, step 102 may include the following steps:

[0083] Step 1021: Through the first processing layer in the video attribute recognition model, generate first splicing information for each text line based on the position information and text information of each text line, and generate second splicing information for each text line based on the position information and image information of each text line.

[0084] In this embodiment of the disclosure, the first processing layer in the video attribute recognition model concatenates the text information and position information corresponding to each text line to obtain the first concatenation information, which serves as the Input1 of the transformer model. The image information and position information corresponding to each text line are concatenated to obtain the second concatenation information, which serves as the Input2 of the transformer model.

[0085] Step 1022: Combine the first concatenation information and the second concatenation information of each text line to obtain the concatenation information of each text line.

[0086] In this embodiment of the disclosure, the concatenation information includes first concatenation information and second concatenation information, and one text line corresponds to a set of first concatenation information and second concatenation information. The first concatenation information and the second concatenation information are concatenated using the concat function to obtain the concatenation information of the text line: Input = concat(Input1, Input2).

[0087] Step 1023: Based on the splicing information of each of the text lines, generate the text line features of each of the text lines.

[0088] In this embodiment of the disclosure, the concatenated information Input is used as the input to the first processing layer, so as to generate text line features text1, text2, ... text through the first processing layer. n =transformer(Input). The text line feature can be represented by `text`. For example, if a video frame contains n text lines, the text line feature corresponding to that video frame can be represented as (text... t1 text t2 , ...text tn ).

[0089] In this embodiment of the disclosure, by splicing the text information, image information and position information of a text line in a single video frame, and inputting the spliced ​​information into a video attribute recognition model, text line features are generated. This allows for feature recognition of text lines by combining multi-dimensional information, thereby improving the accuracy of feature recognition and thus improving the data reliability of text line features.

[0090] Optionally, step 102 may also include the following steps:

[0091] Step 102a: For any of the video frames, the text sequence identifier of the text information and the image sequence identifier of the image information in the video frame are determined by the first processing layer in the video attribute recognition model; the text sequence identifier and the image sequence identifier of the same text line are the same.

[0092] In this embodiment of the disclosure, for any one of at least two video frames, the first processing layer of the video attribute recognition model determines the text sequence identifier (ID) of the text information and the image sequence identifier of the image information in the text line of the video frame. The text sequence identifier and the image sequence identifier are used to align the image information and text information of the same text line. That is, since the text sequence identifier and the image sequence identifier of the same text line are the same, the corresponding text information and image information can be aligned according to the text sequence identifier and the image sequence identifier. One text line corresponds to one text sequence identifier and one image sequence identifier. For example, assuming that a video frame contains n text lines, there are corresponding n text information and n image information. The first processing layer can determine n text sequence identifiers and n image sequence identifiers. Since the text sequence identifier and the image sequence identifier of the same text line are the same, the text sequence identifier and the image sequence identifier can both be represented as (1, 2, ..., n). The first processing layer can be a transformer model based on a multi-head attention mechanism. Multi-head attention mechanisms use multiple attention mechanisms to perform independent computations to obtain more layers of semantic information, and then combine the results obtained by each attention mechanism to obtain the final result.

[0093] Step 102b: Set a continuous first position identifier for the target information of all text lines in the video frame; the target information includes the text information and image information.

[0094] In this embodiment of the disclosure, the target information may include text information and image information of the text lines. A position identifier is randomly initialized for each text line and image information in the video frame, serving as a first position identifier corresponding to each text line and image information, respectively. This first position identifier may be continuously set. The first position identifier can be represented as a 1d position. For example, assuming a video frame contains n text lines, the value range of the first position identifier can be (1, 2, ..., 2n). Specifically, the value range of the first position identifier corresponding to n text lines can be (1, 2, ..., n), and the value range of the first position identifier corresponding to n image lines can be (n+1, n+2, ..., 2n).

[0095] Step 102c: Concatenate the position information, text information, first position identifier of the text information, and text sequence identifier of each text line to obtain the first concatenation information of each text line.

[0096] In this embodiment of the disclosure, the position information, the text sequence identifier corresponding to the text line, and the first position identifier in the text line-related information can be mapped into vectors respectively to obtain a position information vector, a text sequence identifier vector, and a first position identifier vector. The position information vector, the text sequence identifier vector, and the first position identifier vector are then concatenated to obtain position concatenation information. For example, if the position information of a text line is {x} ti :y ti If the text sequence identifier is i and the first position identifier is n+i, then a matrix Emb can be used for mapping to obtain Emb(x). ti Emb(y) ti Emb(i) and Emb(n+i) are used to concatenate the above vectors to obtain the position concatenation information. The obtained location information is then combined with the text information to obtain the first combined information.

[0097] Step 102d: Concatenate the position information, image information, first position identifier of the image information, and image sequence identifier of each text line to obtain the second concatenation information of each text line.

[0098] In this embodiment of the disclosure, the position information, the image sequence identifier corresponding to the text line, and the first position identifier in the text line-related information can be mapped into vectors respectively to obtain a position information vector, an image sequence identifier vector, and a first position identifier vector. The position information vector, the image sequence identifier vector, and the first position identifier vector are then concatenated to obtain position concatenation information. For example, if the position information of a text line is {x} ti :y ti If the image sequence is labeled i and the first position is labeled i, then a matrix Emb can be used for mapping to obtain Emb(x). ti Emb(y) ti The above vectors, Emb(i) and Emb(i), are concatenated to obtain the position concatenation information. The obtained position information is then combined with the text information to obtain the second combined information.

[0099] In this embodiment of the disclosure, by splicing the text information, image information and position information of a text line in a single video frame, and combining the text sequence identifier, image sequence identifier and first position identifier, a text line feature is generated. This allows for feature recognition of the text line by combining multi-dimensional information, improving the accuracy of feature recognition and thus enhancing the data reliability of the text line feature.

[0100] Optionally, step 103 may include the following steps:

[0101] Step 1031: Obtain the first position identifier corresponding to the text information of the text lines in the at least two video frames through the second processing layer in the video attribute recognition model.

[0102] In this embodiment of the disclosure, the first position identifier is the first position identifier set in step 1022, wherein the second processing layer can be a multi-layer transformer model based on a multi-head attention mechanism, and the architecture of the second processing layer is similar to that of the first processing layer.

[0103] Step 1032: For any text line in any video frame, determine the inter-frame features of the text line based on the text line features corresponding to the text line, the display duration, and the first position identifier corresponding to the text information; the display duration is used to characterize the duration of the text line in the at least two video frames.

[0104] In this embodiment of the disclosure, the dwell time information of each text line in at least two video frames can be obtained first to obtain the display duration corresponding to the text line. For example, for any text line, the number of video frames containing the text line in at least two video frames can be determined, and the display duration can be calculated based on the number of video frames. For instance, if a video frame equals 1 / 12 of a second, and a text line is contained in 24 video frames in at least two video frames, then the display duration of the text line is determined to be 2 seconds. This embodiment of the invention does not limit the method for determining the display duration of text lines.

[0105] In one possible implementation, when at least two video frames contain text lines with identical text and location information, the text line features corresponding to these text lines can be obtained, and their average value can be calculated to obtain the mean text line feature. This mean text line feature is then used as a new text line feature to deduplicate overlapping text lines within at least two consecutive video frames. Based on the mean text line feature, display duration, and first location identifier, the inter-frame feature of the corresponding text line is determined. In other words, an inter-frame feature of a text line can be determined based on the information of each text line containing the same text and location information. Thus, by deduplicating overlapping text lines in at least two video frames, it is equivalent to combining the text line features of the same text lines from multiple video frames, and then performing subsequent processing based on the obtained mean text line feature. Therefore, by referencing the temporal information in multiple consecutive video frames, the obtained inter-frame feature of the text line is closer to the true features of the text line, resulting in more accurate text attributes. Meanwhile, by transforming the text line features of overlapping text lines in at least two video frames into the average text line feature of a single text line, the number of calculation steps is reduced, which to some extent lowers the complexity of recognition calculation.

[0106] Step 1033: Based on the classifier in the second processing layer, perform attribute recognition on the target text features of each text line in the at least two video frames to obtain the text attributes corresponding to each text line in the at least two video frames.

[0107] In this embodiment, the second processing layer includes a classifier, which can be a classification model. The classifier maps target text features to a given category, thereby enabling text attribute recognition. The classifier maps the target text features of each text line in at least two video frames to obtain the text attributes corresponding to each text line in at least two video frames.

[0108] In this embodiment of the disclosure, by further combining the text line features in a single video frame with the display duration of the text line and the first position identifier, a text line frame-to-frame feature that integrates multiple modalities such as time, space, image, and text is obtained. Then, the inter-frame feature of the text line is used to perform attribute recognition by a classifier. The multimodal inter-frame feature of the text line in the preceding and following video frames is used as reference information to obtain a more accurate text attribute recognition result.

[0109] Optionally, step 1032 may include the following steps:

[0110] Step 1032a: Encode the display duration and the first location identifier respectively to obtain the display duration vector and the first location identifier vector.

[0111] In this embodiment of the disclosure, for any text line, the display duration corresponding to the text line is mapped to a display duration vector using the matrix Emb, and the first position identifier corresponding to the text line is mapped to a first position identifier vector using the matrix Emb.

[0112] Step 1032b: Concatenate the text line features, the display duration vector, and the first position identifier vector to obtain the text line inter-frame features.

[0113] In this embodiment of the disclosure, the text line feature corresponding to the text line, the display duration vector, and the first position vector are concatenated end to end to obtain the text line inter-frame feature corresponding to the text line. The obtained text line inter-frame feature can be represented as {obj1,…obj…} n}express.

[0114] In this embodiment of the disclosure, by concatenating the display duration vector and the position vector with the text line features, that is, by combining the semantic and position information of the text lines in the preceding and following frames to obtain the inter-frame features of the text lines, the subsequent text attribute recognition results can be made more accurate to a certain extent, and at the same time, the misjudgment of text attributes caused by the ambiguity of the text can be eliminated to a certain extent.

[0115] Optionally, the classifier includes a type attribute classifier, a source attribute classifier, and a visual attribute classifier.

[0116] Among them, the type attribute classifier can be represented as a classifier. func The source attribute classifier can be represented as a classifier. source And visual attribute classifiers can be represented as classifiers. vision For example, the loss can be calculated using the cross-entropy loss function, and the three classifiers mentioned above can be obtained by training with gradient backpropagation.

[0117] Step 1033 may include the following steps:

[0118] Step 1033a: For any given text line, perform attribute recognition based on the inter-frame features of the text line to obtain the text type attribute, text source attribute, and text visual attribute corresponding to the text line.

[0119] In this embodiment of the disclosure, the target text features of the text line can be input into the type attribute classifier, the source attribute classifier, and the visual attribute classifier respectively to obtain the text type attribute output by the type attribute classifier, the text source attribute output by the source attribute classifier, and the text visual attribute output by the visual attribute classifier.

[0120] The text type attributes can include titles, main content, and auxiliary text. Main content can include subtitles, chat logs, comic dialogues, body text, complex layouts, cast and crew lists, etc. Auxiliary text can include logos (e.g., station logos, signs, product brands), annotations, guiding text, bullet comments, special effects text (e.g., magic emoji effects, variety show effects, comic effects), template text, watermarks, comment text, score-related text (e.g., score, team name, time), emoticon text, and advertising sticky notes, etc. Text source attributes can include background text (e.g., scene text, electronic screen text), foreground text (e.g., post-edited text), etc. Text visual attributes can include font size (e.g., below size 1, size 1, 2, 3, 4, 5 and above), font (e.g., printed fonts (such as Song, Kai, Hei), artistic fonts, handwritten fonts), and clarity (e.g., fully clear text, more than half clear text, blurry text), etc. For example, the text attributes output by the video attribute recognition model can be titles, post-edited text, or fully clear text in KaiTi font (size 4).

[0121] For example, the results of the text type attribute, the text source attribute, and the text visual attribute can be obtained in the following ways:

[0122]

[0123]

[0124]

[0125] in, This represents the text type attribute results corresponding to n lines of text. This represents the visual attribute results of n lines of text. This represents the text source attribute results corresponding to n lines of text.

[0126] In this embodiment of the disclosure, three types of attribute classifiers can be used to obtain three dimensions of text attributes corresponding to text lines. Based on the text attributes, the text lines in the video are judged to obtain more comprehensive multi-dimensional text attributes, so that the identified text attributes can be provided to the corresponding business according to business needs.

[0127] Optionally, the video attribute recognition model in this embodiment of the disclosure is trained in the following manner:

[0128] Step 201: Train the basic model based on the first specified training task and the first sample video, and train the basic model based on the second specified training task and each video frame in the first sample video to obtain the recognition model to be trained; the first specified training task includes a text line feature consistency task; the second specified training task includes a text content recovery task, an image region recovery task and / or an image region and text content alignment task.

[0129] In this embodiment, text lines from each video frame of the first sample video are obtained, and a base model is pre-trained based on these text lines. The text content recovery task and the image region recovery task can be performed by masking the text content or image information and then recovering it based on the unmasked text information. The image region and text content alignment task can be performed by masking the image information and then training the image-text alignment based on the correspondence between text and image information. For example, the text content recovery task can specifically involve masking 15% of the text information with a certain probability, inputting the masked text into the base model, and having the base model recover the text characters based on the remaining text information, image information, and position information. The obtained text content is then compared with the original text content, and parameters are adjusted based on the comparison results. Similarly, the image region recovery task can specifically involve masking 20% ​​of the image information of an image region with a certain probability, inputting the masked image into the base model, and having the base model recover the image information of the masked image based on the remaining image information, text information, and position information. The obtained image information is then compared with the original image information, and parameters are adjusted based on the comparison results. The image region-to-text content alignment task can be specifically implemented by blackening 20% ​​of the image information in a text line according to a certain probability, and then inputting the masked text line into a base model. The base model outputs the text content. Since there is a one-to-one correspondence between the text content and the image region, the modal alignment between the text and the image is trained based on whether the output text content and the corresponding image block are masked. The text line feature consistency task can be specifically implemented by training to ensure that the similarity of text line features of the same text line appearing in different frames is higher than a first preset threshold (i.e., to maximize the similarity of text line features of the same text line appearing in different frames), and that the similarity of features of different text lines is lower than a second preset threshold (i.e., to minimize the similarity of features of different text lines).

[0130] In this embodiment, the base model can be a transformer model, such as a multi-layer transformer model based on a multi-head attention mechanism. The network layer corresponding to the text content restoration task can be a layer used only when performing the text content restoration task; the network layer corresponding to the image region restoration task can be a layer used only when performing the image region restoration task; the network layer corresponding to the image region and text content alignment task can be a layer used only when performing the image region and text content alignment task; and the network layer corresponding to the text line feature consistency task can be a layer used only when performing the text line feature consistency task. Each network layer corresponding to each training task can be considered a task-independent network, specifically an independent fully connected layer.

[0131] Optionally, step 201 may include the following steps:

[0132] Step 2011: Use the first sample video as input to the base model, and obtain the text line features of each text line in the first sample video output by the base model.

[0133] In this embodiment of the disclosure, a first sample video is input into the base model. The first sample video contains multiple video frames, and the base model can output the text line features corresponding to each text line in the multiple video frames.

[0134] Step 2012, and, for any video frame in the first sample video, based on the text content, image region, and masked text content and image region in the video frame, obtain the text content and image region output by the base model.

[0135] In this embodiment of the disclosure, for any video frame, a portion of the text content in the video frame is masked, so that the base model outputs the corresponding text content based on the remaining unmasked text content in the video frame; a portion of the image region in the video frame is masked, so that the base model outputs the corresponding image region based on the remaining unmasked image region in the video frame; and a portion of the image region in the video frame is masked, so that the base model outputs the corresponding text content based on the remaining unmasked image region in the video frame.

[0136] Step 2013: Based on the text line features, text content, and image region output by the basic model, adjust the parameters of the basic model.

[0137] In this embodiment of the disclosure, in order to maximize the similarity of text line features of the same text line appearing in different frames and minimize the similarity of different text line features (i.e., to minimize the similarity of different text line features), the parameters of the base model are adjusted; in order to make the output text content as identical as possible to the masked text content, the parameters of the base model are adjusted; in order to make the output image region as identical as possible to the masked image region, the parameters of the base model are adjusted; in order to make the output text content as aligned as possible with the masked image region, the text content corresponding to the masked image region output by the base model is made empty.

[0138] Step 2014: If the basic model reaches the stopping condition, the basic model that has reached the stopping condition is determined as the recognition model to be trained.

[0139] In this embodiment of the disclosure, the stopping conditions may include conditions such as the loss value of the base model reaching a preset threshold or the number of training rounds of the base model reaching a preset number of rounds threshold.

[0140] In this embodiment of the disclosure, by training the base model on a specified task, the base model can learn general video representation capabilities during the training process, better initialize the model training, and train the model's text attribute recognition capabilities in the subsequent fine-tuning process.

[0141] Step 202: Use the text lines in at least two second sample video frames as input to the recognition model to be trained, and obtain the text attributes predicted by the recognition model to be trained.

[0142] In this embodiment of the disclosure, each second sample video frame is pre-annotated with a text line containing the actual text attributes of the text line, and each sample video frame contains text-related information, text sequence identifiers, image sequence identifiers, first position identifiers, and display durations corresponding to each text line.

[0143] Step 203: Based on the text attributes and the text attribute labels of the text lines in the at least two second sample video frames, adjust the parameters of the recognition model to be trained; the text attribute labels are used to characterize the true text attributes of the text lines.

[0144] In this embodiment, the real text attributes include the text type attribute, text source attribute, and text visual attribute corresponding to the text line. The three text attributes output by the recognition model to be trained are compared with the three real text attributes of the text line, and the parameters of the recognition model to be trained are adjusted based on the comparison results. For example, three types of loss functions can be obtained based on the differences between the three text attributes output by the recognition model to be trained and the three real text attributes of the text line. The derivative of this loss function is calculated to obtain the gradient. The gradient backpropagation method is used to optimize the loss function, thereby adjusting the parameters of the recognition model to be trained. The recognition model to be trained with adjusted parameters continues to be trained until the recognition model reaches the stopping condition.

[0145] Step 204: If the recognition model to be trained reaches the stopping condition, the recognition model to be trained that has reached the stopping condition is determined as the video attribute recognition model.

[0146] In this embodiment, the stopping condition may include conditions such as the loss value of the recognition model to be trained reaching a preset threshold, or the number of training epochs of the recognition model to be trained reaching a preset epoch threshold. When the three text attributes output by the recognition model to be trained are consistent with the three real text attributes of the text line, the recognition model to be trained can be considered to have been trained and identified as a video attribute recognition model, so as to realize the fine-tuning process of the recognition model to be trained.

[0147] For example, the optimization of the loss value can be achieved in the following way:

[0148]

[0149]

[0150] Where, loss func Loss represents the loss of text type attributes; vision The loss represents the visual attributes of the text; source This represents the loss of text source attributes; This represents the actual text type attribute corresponding to n lines of text. Represents the visual attributes of the actual text corresponding to n lines of text; This represents the actual text source attribute corresponding to n lines of text.

[0151] In this embodiment, the base model is first unsupervised pre-trained based on multiple specified training tasks and a first sample video. This allows the model to learn general video text object representations, such as text information, image information, and image-text alignment, which are valuable for downstream tasks, ensuring the effectiveness of pre-training. Simultaneously, supervised learning is performed on the model based on a labeled dataset with text attribute labels (at least two second sample video frames), resulting in a superior video attribute recognition model. Since the model has been pre-trained, the parameter tuning calculations during supervised learning converge more easily, i.e., compared to the unpre-trained case, training time is shortened, thus improving the training efficiency of the video attribute recognition model to some extent.

[0152] Optionally, embodiments of this disclosure may further include the following steps:

[0153] Step 301: Obtain the text content of the text line with the text type attribute "title" in the video to be identified; use the text content as the search information for the video to be identified.

[0154] In this embodiment, the text content of text lines with the text type attribute "title" in the video to be identified is obtained. Since the title generally summarizes the main content of the video, the text content can be used as search information for the video to be identified, specifically applicable to video search scenarios. In one possible implementation, if the video to be identified contains the text content "bursting red bean paste rice cake," and the video attribute recognition model determines that the text type attribute corresponding to this text content is "title," then when a user enters "bursting red bean paste rice cake" in the search box, the video to be identified will appear in the search results.

[0155] Step 302: Obtain the text content in the text lines of the video to be identified where the text type attribute is auxiliary text; obtain the virtual item information included in the video to be identified from the text content; associate the virtual item information with the video to be identified.

[0156] In this embodiment, virtual item information is obtained from text content with the text type attribute of auxiliary text, and the video to be identified is associated with the virtual item information. This can be specifically applied to e-commerce product identification scenarios. In one possible implementation, the video to be identified contains the text content "XX brand TV". The video attribute recognition model determines that the text type attribute corresponding to this text content is auxiliary text, specifically a product brand. The virtual item information "TV" is extracted from the text content. Then, when a user enters "TV" in the search box, the video to be identified will appear in the search results.

[0157] Step 303: Determine the distribution degree of the video to be identified based on the visual attributes of the text lines in the video to be identified; the distribution degree is positively correlated with the prominence of the visual attributes of the text.

[0158] In this embodiment, based on the visual attributes of the text lines in the video to be identified, the distribution level of the video to be identified can be determined according to the prominence of the visual attributes. That is, the higher the prominence of the visual attributes, the higher the distribution level of the video to be identified. This can be specifically applied to video distribution and recommendation scenarios. Furthermore, the distribution level of the video to be identified can also be determined by combining the visual attributes of the text lines in the video to be identified with the text type attribute. In one possible implementation, if the video to be identified contains text content in size 1 that is a title, then the distribution level of this video to be identified is higher than that of a video containing text content in size 4 that is a subtitle.

[0159] In this embodiment of the disclosure, based on the different text attributes of the text content in the video to be identified, the video to be identified can be provided to different business scenarios to achieve more accurate recommendations or searches. At the same time, the business scenarios can be more accurately located based on the text attributes, which to a certain extent ensures the application value of the video to be identified.

[0160] For example, Figure 2 A schematic diagram of the processing flow of a video attribute recognition model is shown, such as... Figure 2 As described above, Images 1 and 2 represent two consecutive video frames in a video. The first processing layer represents the first processing layer (stageone) in the video attribute recognition model, and the second processing layer represents the second processing layer (stagetwo) in the same model. First, the text line features in the video frames of Images 1 and 2 are obtained respectively. Figure 2Taking the acquisition of text line features from image 1 as an example, the video frame is first input into a deep neural network, such as a Feature Pyramid Network (FPN), to generate an image feature map and obtain image information. Then, OCR technology is used to recognize the text for text processing to obtain text information and location information, such as... Figure 2 The positional information in the first processing layer is shown as coordinates 1, 2, 3, and 4. For example, coordinates 1, 2, 3, and 4 can be respectively (x... t1 y t1 ), (x t2 y t2 ), (x t3 y t3 ), (x t4 y t4 The text information and image information correspond to the same location information. Determine the sequence identifiers corresponding to the image information and text information, such as... Figure 2 The sequence identifiers shown in the first processing layer are 1-4 for the four text lines. Text and image information within the same text line correspond to the same sequence identifier, and consecutive first position identifiers are set for the text and image information within each text line. Figure 2 The first position identifiers shown in the first processing layer are 1-4 for image information and 5-8 for text information. Based on text line information, text sequence identifiers, image sequence identifiers, and the first position identifiers, splicing information is generated, thereby obtaining the text line features corresponding to each text line: such as... Figure 2 Features 1, 2, 3, and 4 shown in the first processing layer can, for example, be represented as text. t1 text t2 text t3 text t4 After obtaining the text line features of each text line in the video frames of Image 1 and Image 2, the text line features of the video frames of Image 1 and Image 2 (such as...) are then obtained. Figure 2 The second processing layer contains features 1, 3, 4, and 5, which are used to display the duration vector: (e.g., ...) Figure 2 As shown in the second processing layer, vectors 1, 3, 4, and 5 can be exemplarily represented as Emb(t1), Emb(t3), Emb(t4), and Emb(t5), respectively, along with the first position identifier vector (e.g., ...). Figure 2The second processing layer (shown as 1, 3, 4, and 5) removes duplicate text lines from the video frames of images 1 and 2, and calculates the average value of the text line features corresponding to the duplicate text lines, such as... Figure 2 As shown, since "The Yang family is preparing to go to the suburbs," "outing" (not shown in the image), and "the thirteenth paragraph" appear in both Image 1 and Image 2 video frames, the text embeddings corresponding to these text lines are obtained by averaging the text line features in Image 1 and Image 2 video frames. Based on the processed text line features (textembedding), the display duration vector (time embedding), and the first position identifier vector (PositionEmbedding), the inter-frame features of the text lines are obtained. The classifier in the second processing layer then identifies the text attributes corresponding to each text line. For example... Figure 2 As shown, the text attributes of “Yang Family is going to the suburbs”, “outing” (not shown in the figure) and “the thirteenth paragraph” are identified as titles, while the text attributes of “The weather is nice today” and “We can go out and play” are identified as subtitles.

[0161] Figure 3 This is a block diagram illustrating a text attribute recognition device according to an exemplary embodiment, such as... Figure 3 As shown, the device 40 may include:

[0162] The first acquisition module 401 is used to acquire text line related information of the text lines in each of the video frames based on the text lines in at least two consecutive video frames in the video to be identified.

[0163] The first encoding module 402 is used to perform feature encoding on the text line related information of the text lines in each video frame through a pre-trained video attribute recognition model, so as to obtain the text line features corresponding to each text line in each video frame.

[0164] The second acquisition module 403 is used to acquire the text line frame features corresponding to each text line based on the text line features corresponding to the at least two video frames through the video attribute recognition model.

[0165] The first recognition module 404 is used to perform attribute recognition on the inter-frame features of the text lines and determine the text attributes corresponding to each text line in the video frame to be recognized.

[0166] In one optional embodiment, the text line-related information includes text information, image information, and location information; the first acquisition module 401 is specifically configured to:

[0167] The first acquisition submodule is used to acquire local image features of the region where each text line in any of the at least two video frames is located, as the image information.

[0168] The second acquisition submodule is used to acquire the text content of each text line in the video frame as the text information.

[0169] The third acquisition submodule is used to acquire the position coordinates of each text line in the video frame as the position information.

[0170] In one alternative embodiment, the first encoding module 402 is specifically configured as follows:

[0171] The first generation module is configured to generate first splicing information for each text line based on the position information and text information of each text line through the first processing layer in the video attribute recognition model, and to generate second splicing information for each text line based on the position information and image information of each text line.

[0172] The first splicing module is used to splice the first splicing information and the second splicing information of each text line to obtain the splicing information of each text line.

[0173] The second generation module is used to generate text line features for each text line based on the splicing information of each text line.

[0174] In one alternative embodiment, the first encoding module 402 is further configured to:

[0175] The first determining module is used to determine, for any video frame, the text sequence identifier of the text information and the image sequence identifier of the image information in the text line of the video frame through the first processing layer in the video attribute recognition model; the text sequence identifier and the image sequence identifier of the same text line are the same.

[0176] The first setting module is used to set a continuous first position identifier for the target information of all text lines in the video frame; the target information includes the text information and image information.

[0177] The first generation module is specifically configured as follows:

[0178] The first splicing submodule is used to splice the position information, the text information, the first position identifier of the text information, and the text sequence identifier of each text line to obtain the first splicing information of each text line.

[0179] The second splicing submodule is used to splice the position information, image information, first position identifier of the image information and image sequence identifier of each text line to obtain the second splicing information of each text line.

[0180] In one alternative embodiment, the second acquisition module 403 is specifically configured as follows:

[0181] The fourth acquisition submodule is used to acquire the first position identifier of the text line in the at least two video frames through the second processing layer in the video attribute recognition model.

[0182] The second determining module is used to determine the inter-frame features of a text line for any text line in any video frame based on the text line features corresponding to the text line, the display duration, and the first position identifier; the display duration is used to characterize the duration of the text line in the video frame.

[0183] In one alternative embodiment, the second determining module is specifically configured as follows:

[0184] The first encoding module is used to encode the display duration and the first location identifier respectively to obtain the display duration vector and the first location identifier vector.

[0185] The first splicing module is used to splice the text line features, the display duration vector, and the first position identifier vector to obtain the text line frame features.

[0186] In one alternative embodiment, the first identification submodule may specifically be used for:

[0187] For any given text line, attribute recognition is performed based on the inter-frame features of the text line to obtain the text type attribute, text source attribute, and text visual attribute corresponding to the text line.

[0188] In one alternative embodiment, the device 40 further includes:

[0189] The second acquisition module is used to acquire the text content in the text line with the text type attribute of "title" in the video to be identified; and to use the text content as the search information for the video to be identified.

[0190] The third acquisition module is used to acquire the text content in the text lines of the video to be identified where the text type attribute is auxiliary text; acquire the virtual item information included in the video to be identified from the text content; and associate the virtual item information with the video to be identified.

[0191] The third determining module is used to determine the distribution degree of the video to be identified based on the visual attributes of the text lines in the video to be identified; the distribution degree is positively correlated with the prominence of the visual attributes of the text.

[0192] According to one embodiment of this disclosure, an electronic device is provided, including: a processor and a memory for storing processor-executable instructions, wherein the processor is configured to implement, when executed, the steps of the text attribute recognition method as described in any of the above embodiments.

[0193] According to one embodiment of the present disclosure, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the steps of the text attribute recognition method as described in any of the above embodiments.

[0194] According to one embodiment of this disclosure, a computer program product is also provided, which includes readable program instructions that, when executed by a processor of an electronic device, enable the electronic device to perform the steps in the text attribute recognition method as described in any of the above embodiments.

[0195] Figure 4 This is a block diagram illustrating an apparatus for text attribute recognition according to an exemplary embodiment.

[0196] The device 500 may include a processing component 502, a memory 504, a power supply component 506, a multimedia component 508, an audio component 510, an input / output interface 512, a sensor component 514, a communication component 516, and a processor 520. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the text attribute recognition method described above. In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as the memory 504 including instructions, which can be executed by the processor 520 of the device 500 to complete the method described above. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0197] Figure 5This is a block diagram illustrating another apparatus for text attribute recognition according to an exemplary embodiment. The apparatus 600 may include a processing component 622, a memory 632, an input / output interface 658, a network interface 650, and a power supply component 626. The apparatus 600 may be provided as a server. The application program stored in the memory 632 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 622 is configured to execute instructions to perform the aforementioned text attribute recognition method.

[0198] All user information (including but not limited to user device information, user personal information, etc.) and related data involved in this disclosure are information authorized by the user or by the parties involved.

[0199] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0200] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A text attribute recognition method, characterized in that, The method includes: Based on the text lines in at least two consecutive video frames in the video to be identified, obtain the text line related information of the text lines in each video frame; the text line related information includes text information; By using a pre-trained video attribute recognition model, the text line information of each text line in each video frame is encoded to obtain the text line features corresponding to each text line in each video frame. Based on the text line features corresponding to at least two video frames, the video attribute recognition model obtains the inter-frame features of each text line. Attribute recognition is performed on the inter-frame features of the text lines to determine the text attributes corresponding to each text line in the video frame to be identified. The method further includes: A continuous first position identifier is set for the target information of all text lines in the video frame; the target information includes the text information. The step of obtaining inter-frame features of each text line based on the text line features corresponding to at least two video frames using the video attribute recognition model includes: The first position identifier corresponding to the text information of the text lines in the at least two video frames is obtained through the second processing layer in the video attribute recognition model. For any text line in any video frame, the inter-frame features of the text line are determined based on the text line features corresponding to the text line, the display duration, and the first position identifier corresponding to the text information; the display duration is used to characterize the duration of the text line in the at least two video frames.

2. The method according to claim 1, characterized in that, The text line information also includes image information and location information; the step of obtaining text line information for each video frame based on text lines in at least two consecutive video frames in the video to be identified includes: For any one of the at least two video frames, local image features of the region where each text line is located in the video frame are obtained, and used as the image information; The text content of each text line in the video frame is obtained as the text information; Obtain the position coordinates of each text line in the video frame as the position information.

3. The method according to claim 2, characterized in that, The step of using a pre-trained video attribute recognition model to encode the text line information of each text line in each video frame to obtain the text line features corresponding to each text line in each video frame includes: The first processing layer in the video attribute recognition model generates first splicing information for each text line based on the position information and text information of each text line, and generates second splicing information for each text line based on the position information and image information of each text line. The first concatenation information and the second concatenation information of each text line are concatenated to obtain the concatenation information of each text line; Based on the concatenation information of each of the text lines, the text line features of each of the text lines are generated.

4. The method according to claim 3, characterized in that, The method further includes: For any of the video frames, the first processing layer determines the text sequence identifier of the text information and the image sequence identifier of the image information in the text line of the video frame; the text sequence identifier and the image sequence identifier of the same text line are the same; The target information also includes image information; The step of generating first concatenation information for each text line based on the position information and text information of each text line includes: concatenating the position information, text information, first position identifier of the text information, and text sequence identifier of each text line to obtain the first concatenation information for each text line; The step of generating second splicing information for each text line based on the position information and the image information of each text line includes: splicing the position information, the image information, the first position identifier of the image information, and the image sequence identifier of each text line to obtain the second splicing information for each text line.

5. The method according to claim 1, characterized in that, The step of determining the inter-frame features of the text line based on the text line features corresponding to the text line, the display duration, and the first position identifier corresponding to the text information includes: The display duration and the first position identifier corresponding to the text information are encoded to obtain a display duration vector and a first position identifier vector, respectively. The text line features, the display duration vector, and the first position identifier vector are concatenated to obtain the text line frame features.

6. The method according to claim 1, characterized in that, The video attribute recognition model is trained in the following manner: The basic model is trained based on a first specified training task and a first sample video, and the basic model is trained based on a second specified training task and each video frame in the first sample video to obtain a recognition model to be trained; the first specified training task includes a text line feature consistency task. The second designated training task includes a text content recovery task, an image region recovery task, and / or an image region and text content alignment task; Use the text lines from at least two second sample video frames as input to the recognition model to be trained, and obtain the text attributes predicted by the recognition model to be trained; Based on the text attributes and the text attribute labels of the text lines in the at least two second sample video frames, the parameters of the recognition model to be trained are adjusted; the text attribute labels are used to characterize the true text attributes of the text lines. If the training recognition model reaches the stopping condition, the training recognition model that has reached the stopping condition is determined as the video attribute recognition model.

7. A text attribute recognition device, characterized in that, The device includes: The first acquisition module is used to acquire text line related information of the text lines in each of at least two consecutive video frames in the video to be identified; the text line related information includes text information. The first encoding module is used to perform feature encoding on the text line information of the text lines in each video frame using a pre-trained video attribute recognition model, so as to obtain the text line features corresponding to each text line in each video frame. The second acquisition module is used to acquire the inter-frame features of each text line based on the text line features corresponding to the at least two video frames through the video attribute recognition model. The first recognition module is used to perform attribute recognition on the inter-frame features of the text lines and determine the text attributes corresponding to each text line in the video frame to be recognized. The device further includes: The first setting module is used to set a continuous first position identifier for the target information of all text lines in the video frame; the target information includes the text information. The second acquisition module is specifically configured as follows: The fourth acquisition submodule is used to acquire the first position identifier of the text line in the at least two video frames through the second processing layer in the video attribute recognition model; The second determining module is used to determine the inter-frame features of a text line for any text line in any video frame based on the text line features corresponding to the text line, the display duration, and the first position identifier; the display duration is used to characterize the duration of the text line in the video frame.

8. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device performs the method as described in any one of claims 1 to 6.