Video character processing method and device, equipment, storage medium and program product
By performing frame sampling, character recognition, text tracking, and type identification on the video, and integrating the spatiotemporal attribute information of the text trajectory, the problems of repetitive output and scene noise interference in video text processing are solved, thereby improving the accuracy and semantic coherence of the video text structuring results.
Patent Information
- Application Number
- CN202511072931.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies for video text processing suffer from issues such as repetitive text output, scene noise interference, and a lack of structured information leading to semantic incoherence.
By performing frame sampling, character recognition, text tracking, and type identification on the video, and integrating the spatiotemporal attribute information of the text trajectory, a structured video text result is obtained.
It improves the accuracy of video text structuring results, reduces repetitive text output, enhances semantic coherence, and facilitates downstream applications in accurately obtaining content of interest.
Smart Images

Figure CN120976906A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video processing, in particular to a video text processing method and device, equipment, storage medium and program product. BACKGROUND
[0002] The text appearing in a video usually contains rich semantic information and is an indispensable feature of video content understanding. Accordingly, the accuracy of video text processing will affect downstream businesses. Therefore, it is necessary to provide a video text processing method to ensure the accuracy of video text processing. SUMMARY
[0003] Therefore, the present application provides a video text processing method, device, equipment, storage medium and program product to solve the problem of video text processing.
[0004] In a first aspect, the present application provides a video text processing method, comprising:
[0005] obtaining a first video and frame sampling the first video to obtain a first image frame;
[0006] performing text recognition on the first image frame to obtain a text recognition result of the first image frame;
[0007] performing text tracking based on the text recognition result to obtain a first text track of the first video, the first text track comprising text content and spatio-temporal attribute information of the text content;
[0008] performing type recognition based on the first image frame and the first text track to obtain a type of the first text track, the type of the first text track representing a function of the text content in the first image frame;
[0009] integrating the first text track based on the type of the first text track and the spatio-temporal attribute information of the text content in the first text track to obtain a text structured result of the first video.
[0010] In a second aspect, the present application provides a video text processing device, comprising:
[0011] an acquisition module configured to obtain a first video and frame sample the first video to obtain a first image frame;
[0012] a text recognition module configured to perform text recognition on the first image frame to obtain a text recognition result of the first image frame;
[0013] a text tracking module, configured to perform text tracking based on the character recognition result, to obtain a first text track of the first video, the first text track comprising text content and spatio-temporal attribute information of the text content;
[0014] a type identification module, configured to perform type identification based on the first image frame and the first text track, to obtain a type of the first text track, the type of the first text track being used to represent a function of the text content in the first image frame;
[0015] a track integration module, configured to integrate the first text track based on the type of the first text track and the spatio-temporal attribute information of the text content in the first text track, to obtain a character structured result of the first video.
[0016] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, the memory and the processor are communicatively connected with each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the video character processing method in the first aspect or any of the corresponding implementation manners thereof.
[0017] In a fourth aspect, the present application provides a computer readable storage medium, which stores computer instructions, and the computer instructions are used to make a computer execute the video character processing method in the first aspect or any of the corresponding implementation manners thereof.
[0018] In a fifth aspect, the present application provides a computer program product, which comprises computer instructions, and the computer instructions are used to make a computer execute the video character processing method in the first aspect or any of the corresponding implementation manners thereof.
[0019] The video text processing method provided by the embodiment of the application comprises the following steps: acquiring a first video, performing frame sampling on the first video to obtain a first image frame; performing text recognition on the first image frame to obtain a text recognition result of the first image frame; performing text tracking based on the text recognition result to obtain a first text track of the first video, the first text track comprising text content and space-time attribute information of the text content; performing type recognition based on the first image frame and the first text track to obtain a type of the first text track, the type of the first text track being used to represent a function of the text content in the first image frame; and integrating the first text track based on the type of the first text track and the space-time attribute information of the text content of the first text track to obtain a text structured result of the first video. The method performs frame sampling, text recognition, text tracking and type recognition on the first video, and then performs track integration according to the type of the first text track and the space-time attribute information of the text content of the first text track, so as to obtain a structured video text result, provide a text result with richer attributes from multiple angles, improve the accuracy of the text structured result of the first video, and enable downstream applications to accurately obtain interested content. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the specific embodiments or the related art, the drawings needed to be used in the specific embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0021] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present application;
[0022] Figure 2 is a first principle schematic diagram of a video text processing method according to an embodiment of the present application;
[0023] Figure 3 is a first flow schematic diagram of a video text processing method according to an embodiment of the present application;
[0024] Figures 4a-4b is a schematic diagram of a text recognition result of a first image frame according to an embodiment of the present application;
[0025] Figure 5 is a second flow schematic diagram of a video text processing method according to an embodiment of the present application;
[0026] Figure 6 is a second principle schematic diagram of a video text processing method according to an embodiment of the present application;
[0027] Figure 7 is a second flow diagram of a video text processing method according to an embodiment of the present application;
[0028] Figure 8 is a display page diagram of text structuring according to an embodiment of the present application;
[0029] Figure 9 is a structural block diagram of a video text processing apparatus according to an embodiment of the present application;
[0030] Figure 10 is a hardware structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0032] It can be understood that, before using the technical solutions disclosed in the embodiments of the present application, the type, use range, use scenario and the like of personal information involved in the present application should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0033] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will need to obtain and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the electronic device, application program, server or storage medium and the like software or hardware performing the operation of the technical solutions of the present application according to the prompt information.
[0034] As an optional but not limited implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be a pop-up window manner, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0035] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present application, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present application.
[0036] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws and regulations and relevant provisions.
[0037] In the related art, for the processing of video text, the text in the video is generally recognized first, and then the text recognition result is output in sequence according to the order of the image frame where the text is located. However, this processing method of video text may have the problem of repeated text in similar frames. That is, for the same subtitle that may exist in multiple consecutive image frames, the same subtitle will be output multiple times when the text is output.
[0038] In addition, due to the existence of scene text in the image frame, for example, the text on a certain object in the image frame. The output of such text will cause the output result to contain more scene noise text interference. Further, since the text in the video is output in sequence according to the order of the image frame where the text is located, due to the lack of structured information, it may cause semantic incoherence and the like.
[0039] Based on this, the embodiment of the present application provides a video text processing method, which comprises the following steps: acquiring a first video, and performing frame sampling on the first video to obtain a first image frame; performing text recognition on the first image frame to obtain a text recognition result of the first image frame; performing text tracking based on the text recognition result to obtain a first text track of the first video, the first text track comprising text content and spatio-temporal attribute information of the text content; performing type identification based on the first image frame and the first text track to obtain a type of the first text track, the type of the first text track being used to represent the function of the text content in the first image frame; and integrating the first text track based on the type of the first text track and the spatio-temporal attribute information of the text content of the first text track to obtain a text structured result of the first video.
[0040] As shown in Figure 2 , the method performs frame extraction, text recognition, text tracking and video text structuring (i.e., type identification) on the obtained video Li Lu, and then performs track integration according to the type of the obtained first text track and the spatio-temporal attribute information of the text content of the first text track to obtain a structured video text result and output. The method provides a text result with more rich attributes from multiple angles, improves the accuracy of the text structured result of the first video, so that the downstream application can accurately obtain the content of interest according to the text structured result.
[0041] As an optional application scenario of the embodiment of the present application, as shown in Figure 1 , as an optional application scenario of the embodiment of the present application, as shown in Figure 1As shown, the terminal device 110 is installed with an application 101, and the user 130 can interact with the application 101 through the terminal device 110 and / or the access device of the terminal device 110.
[0042] Exemplarily, the application 101 can be any application, and can provide a video text processing application. Figure 1 In the application scenario as shown, if the application 101 is in an active state, the terminal device 110 can present an interface 102 of the application 101. The interface 102 can include various pages that can be provided by the application 101, such as an interactive page, a setting page, a query page, and the like.
[0043] In some embodiments, the terminal device 110 is in communication connection with the server 120 to implement the provision of services of the application 101. The terminal device 110 can be a mobile terminal, a fixed terminal, or a portable terminal, and the like, including but not limited to a mobile phone, a desktop computer, a notebook computer, a multimedia tablet, an electronic book device, a game device, or any combination of the above, including accessories and peripherals of the devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface, and the server 120 can be various types of computing systems, servers, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like, which can provide computing capabilities.
[0044] It should be noted that, Figure 1 This is only an example of an application scenario, and does not limit the protection scope of the present disclosure.
[0045] The embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the pages shown in the drawings are only examples, and various page designs can actually exist. Various graphical elements in the page can have different arrangements and different visual representations, one or more elements of which can be omitted or replaced, and one or more other elements can also exist, which are not limited in the embodiments of the present disclosure. In addition, the embodiments are mainly described below with respect to the server.
[0046] According to the embodiments of the present application, a video text processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0047] In the present embodiment, a video text processing method is provided, which can be used for the server as described above, Figure 3 is a flowchart of the video text processing method according to the embodiments of the present application, as Figure 3 shown, the flow includes the following steps:
[0048] In step S301, a first video is acquired, and frame sampling is performed on the first video to obtain first image frames.
[0049] The first video is a video that needs to be subjected to video-text processing, and includes a plurality of continuous image frames. The first video can be a locally stored video, a video acquired through a communication connection with another device, or a locally generated video, etc. The source of the first video is not limited herein.
[0050] Frame sampling means that the first video is sampled according to a preset sampling frequency. The preset sampling frequency is set according to actual needs, for example, once every 1 second, once every 2 seconds, etc.
[0051] After the preset sampling frequency is obtained, frame sampling is performed on the first video to obtain first image frames. It should be understood that the first image frames are a general term for a plurality of image frames, and do not specifically refer to one image frame. Without loss of generality, the first frame of the first video belongs to the first image frames, and subsequent frame sampling is performed according to the preset sampling frequency to obtain the first image frames. The first image frames are used for subsequent text recognition and other processing.
[0052] In step S302, text recognition is performed on the first image frames to obtain text recognition results of the first image frames.
[0053] The subsequent processing described for the first image frames is performed on each first image frame, and does not specifically refer to a certain first image frame. The text in the first image frames includes but is not limited to titles, backgrounds, subtitles, etc. in the image frames, for example, there is an article in the first image frames, and there is text on the article, so the text recognition result will include the text on the article. Further, if there is a subtitle in the first image frames, the text recognition result will include the subtitle.
[0054] The manner of performing text recognition (Optical Character Recognition, OCR) on the first image frames can be a non-deep learning OCR method and a deep learning-based method, etc. Taking the non-deep learning OCR method as an example, first, the first image frames are preprocessed, including but not limited to image noise reduction, binarization, tilt correction, and text positioning and segmentation; then the preprocessed image frames are extracted for geometric features, statistical features, and structural features, etc.; and the extracted features are classified and recognized to obtain the text recognition result.
[0055] The text recognition result of the first image frame includes the text content, the spatial information of the text content, and the image frame to which the text content belongs. For example, if the text content in the first image frame appears continuously, a line of text can be recognized as a text block and output as one piece of text content. The spatial information of the text content is used to characterize the position of the text content in the first image frame, such as the size of the text box corresponding to the text content and the coordinates of the text box, etc. The image frame to which the text content belongs, that is, the correspondence between the text content and the image frame, can also be referred to as the temporal information of the text content.
[0056] Since text content may exist in multiple locations within the first image frame, multiple text contents can be generated from the same first image frame. The text recognition result records the spatial information of these multiple text contents and the image frame to which the text contents belong. Without loss of generality, the image frame to which the text contents belong can be represented by the frame number of the image frame.
[0057] For example, such as Figure 4a As shown, the text content identified in the first image frame includes: text content 401, text content 402, text content 403 and text content 404. These text contents are located in different positions in the first image frame. When outputting the text recognition result, in addition to outputting the specific text content, the position information of the text content is also included.
[0058] Figure 4a The frame number of the first image frame shown is 10. Figure 4b The first image frame shown has a frame number of 15, and its corresponding text content includes: text content 401, text content 403, text content 404 and text content 405. Accordingly, the position information of the text content is output, etc.
[0059] Step S303: Based on the text recognition results, perform text tracking to obtain the first text trajectory of the first video.
[0060] The first text trajectory includes the text content and the spatiotemporal attribute information of the text content.
[0061] Because text content in the same position may appear in multiple consecutive frames, for example, Figure 4a The text content 401 in the middle, not only in Figure 4a The text 401 appears in frame 10 and also in frame 15 of Figure 4. It should be noted that the same position in frames 11 through 14 also contains text content 401, and text content 401 first appears in frame 10 and last appears in frame 15. Therefore, the duration of text content 401's appearance is from frame 10 to frame 15.
[0062] In other words, if the same text content appears consecutively for 6 frames, then by performing text tracking on the text recognition results, the same text content can be represented in the form of text trajectories, which can reduce the output of repeated text content.
[0063] Text tracking can determine whether text in adjacent frames belongs to the same text by calculating the feature similarity of text content. For example, for multiple first image frames, text detection and recognition are performed on each frame. Feature extraction, recognition, and fusion are performed through an end-to-end text tracking network to obtain a fused feature descriptor. Based on this, the preset tracking trajectory is updated to obtain the text tracking trajectory.
[0064] Text trajectories are primarily used to represent the location information of text content, including which image frames the text content is distributed in and its specific location within each image frame. For example, the target tracking trajectory of text content in a video might indicate that the text content exists continuously from frame 10 to frame 15 of a video, and its specific location information in each frame can also be determined through the text trajectory.
[0065] Specifically, the first text trajectory includes text content and its spatiotemporal attribute information. The text content refers to the text itself, such as all the text in a line of text. The spatiotemporal attribute information includes temporal and spatial attributes. The temporal attribute information represents the image frame to which the text content belongs, such as frames 10 to 15. The spatial attribute information represents the position of the text content within the image frame, such as the position of the text box corresponding to the text content and the size of the text box, etc.
[0066] Continuing with the example above, such as Figure 4a as well as Figure 4b As shown, Figure 4a The image shown is frame 10, and the text content in the image frames from frame 11 to frame 14 is the same (not only the text content itself, but also the position information of the text content, etc.). Starting from frame 15, the text content of the image frames changes.
[0067] Therefore, the text track of the text content 401 can be obtained by text tracking: the text content 401, the position information of the text content 401, and the image frames to which the text content 401 belongs are the 10th frame to the 15th frame; the text track of the text content 402: the text content 402, the position information of the text content 402, and the image frames to which the text content 402 belongs are the 10th frame to the 14th frame; the text track of the text content 403: the text content 403, the position information of the text content 403, and the image frames to which the text content 403 belongs are the 10th frame to the 15th frame; the text track of the text content 404: the text content 404, the position information of the text content 404, and the image frames to which the text content 404 belongs are the 10th frame to the 15th frame; and the text track of the text content 405: the text content 405, the position information of the text content 405, and the image frames to which the text content 405 belongs are the 15th frame.
[0068] In step S304, type recognition is performed on the first image frames and the first text tracks to obtain the types of the first text tracks.
[0069] The type of the first text track is used to represent the function of the text content in the first image frame.
[0070] The function of the text content in the first image frame includes title, background, and subtitles, etc. The specific categories are set according to actual needs. The types of the first text tracks are obtained by sequentially performing type recognition on the first text tracks.
[0071] For example, the title generally appears in the upper middle position of the image frame, and the subtitles appear in the lower middle position of the image frame. Therefore, the position information of the text content of the first text track can be used to identify the type of the first text track. Alternatively, a text track classification model can also be used, the input of the text track classification model includes the first text track and the image frame involved by the first text track, and the output of the text track classification model includes the type corresponding to the first text track.
[0072] The types of the first text tracks are obtained by performing type recognition on all the first text tracks obtained in step S303.
[0073] In step S305, the first text tracks are integrated based on the types of the first text tracks and the spatio-temporal attribute information of the text content of the first text tracks to obtain the text structured result of the first video.
[0074] As described above, the spatio-temporal attribute information of the text content can include temporal attribute information and spatial attribute information. In the integration of the first text track, the first text track can be divided according to the type of the text track, and the first text track of the same type can be integrated according to the spatio-temporal attribute information of the text content to obtain the text structured result of the first video.
[0075] When the text structured result is output, the type of the text track can be taken as a first major category, and the text content can be sorted according to the time sequence and the position relationship under the text track of the same type.
[0076] Of course, other ways can also be used to integrate the first text track, which is not limited herein and can be set according to actual needs.
[0077] The video text processing method provided in this embodiment frames the first video, recognizes the text, tracks the text, and identifies the type, and then integrates the track according to the type of the obtained first text track and the spatio-temporal attribute information of the text content of the first text track to obtain the structured video text result. The method provides a text result with richer attributes from multiple perspectives, improves the accuracy of the text structured result of the first video, and enables downstream applications to accurately obtain the content of interest.
[0078] In this embodiment, a video text processing method is provided, which can be used in the server, Figure 5 is a flowchart of the video text processing method according to an embodiment of the application, as Figure 5 shown, the flowchart includes the following steps:
[0079] In step S501, a first video is obtained, and the first video is frame sampled to obtain a first image frame. For details, refer to step S301 of the embodiment shown in Figure 3 , which will not be repeated here.
[0080] In step S502, text recognition is performed on the first image frame to obtain a text recognition result of the first image frame. For details, refer to step S302 of the embodiment shown in Figure 3 , which will not be repeated here.
[0081] In step S503, text tracking is performed based on the text recognition result to obtain a first text track of the first video.
[0082] The first text track includes text content and spatio-temporal attribute information of the text content. For details, refer to step S303 of the embodiment shown in Figure 3 , which will not be repeated here.
[0083] In step S504, type recognition is performed based on the first image frame and the first text track, to obtain a type of the first text track.
[0084] The type of the first text track is used to represent a function of the text content in the first image frame.
[0085] Specifically, the step S504 includes:
[0086] In step S5041, text encoding is performed on the first text track, to obtain a text vector.
[0087] In addition to the text content, the first text track also includes spatio-temporal attribute information of the text content. Accordingly, the text content and the spatio-temporal attribute information are encoded together to obtain the text vector. For example, the text information (the text content and the spatio-temporal attribute information) is first represented by token, and then vector mapping is performed to obtain the text vector.
[0088] In step S5042, visual encoding is performed on the first image frame, to obtain a visual vector.
[0089] Each first image frame can be encoded by a visual encoder to obtain a corresponding visual sub-vector. The visual sub-vectors are spliced to obtain the visual vector. The specific implementation of the visual encoder is not limited herein, and can be set according to actual requirements.
[0090] In some optional embodiments, the step S5042 includes:
[0091] In step a1, the first image frame is sampled based on a duration of the first text track, to obtain a second image frame. The text content in all the second image frames covers the text content of the first text track.
[0092] In step a2, visual encoding is performed on the second image frame, to obtain the visual vector.
[0093] As described above, the same text content can appear in multiple continuous image frames. In addition, the time information of each text content is represented in the first text track, i.e., the corresponding start frame and end frame of each text content are included. Accordingly, the start frame and the end frame of the text content correspond to the duration of the first text track. For example, if the start frame of the text content of the first text track is the 10th frame and the end frame is the 15th frame, then the duration of the first text track is from the 10th frame to the 15th frame.
[0094] Since the text content can exist in multiple consecutive image frames, the image frames are repetitive information for the text content. Therefore, the first image frames are sampled according to the duration of the first text track to obtain second image frames. That is, the text content in all the obtained second image frames can cover the text content of all the first text tracks.
[0095] Exemplarily, the time information of each first text track can be compared to determine the repetitive image frames. One frame in the repetitive image frames is taken as a second image frame, and the non-repetitive image frames are sampled to ensure that the text content of the obtained second image frames can cover the text content in all the first text tracks.
[0096] Through the sampling of the first image frames, the number of image frames for subsequent visual coding can be reduced. That is, the number of second image frames is less than the number of first image frames. The second image frames are then visually coded to obtain visual vectors.
[0097] In the process of visually coding the first image frames, the first image frames are sampled according to the duration of the first text track to ensure that the text content in all the obtained second image frames can cover the text content of all the first text tracks, which on the one hand removes the image frames with the same text content and on the other hand improves the processing efficiency.
[0098] In some optional embodiments, the above step a2 includes:
[0099] Step a21, aggregating the first text track based on the spatiotemporal attribute information of the first text track to obtain a second text track.
[0100] Step a22, visually coding the second image frames to obtain visual features.
[0101] Step a23, vectorizing based on the second text track and the visual features to obtain a visual vector.
[0102] In the process of visually coding the second image frames, the first text track can be introduced. Since the text content of the first text track comes from the first image frames, introducing the first text track into the visual coding can further improve the reliability of the obtained visual vector.
[0103] Specifically, the first text track is aggregated based on the spatiotemporal attribute information of the first text track. Here, since the type of the first text track has not been obtained, the aggregation here is aggregation using time information and spatial information. That is, through feature fusion in the spatial and temporal dimensions, trajectory optimization, and multi-modal collaboration, discrete text content is converted into coherent structured information.
[0104] Visual features are obtained by visual encoding of the second image frame, and then vectorized by combining the aggregated second text trajectory to obtain a visual vector.
[0105] By aggregating the first text trajectory using spatiotemporal attribute information, scattered text regions can be associated and merged in time and space dimensions to form a coherent text trajectory, thereby improving recognition accuracy and processing efficiency, and further enhancing the reliability of the obtained visual vectors.
[0106] Step S5043: Perform feature encoding based on text vectors and visual vectors to obtain the first feature.
[0107] Feature encoding is performed on the text vector and the visual vector to obtain the first feature. Furthermore, the first image frame can also be incorporated during feature encoding. In other words, the input to feature encoding includes the text vector, the visual vector, and the first image frame, and the output includes the first feature.
[0108] Step S5044: Based on the first feature, perform type recognition to obtain the type of the first text trajectory.
[0109] Type recognition is implemented based on a type classification model. Its input includes a first feature, and its output includes the type of the first text trajectory. The specific structure of the type classification model is not limited here; it can be set according to actual needs.
[0110] For example, such as Figure 6 As shown, text recognition is performed on the first image frame to obtain the text recognition result; then, text tracking is performed on the text recognition result to obtain the text trajectory. Frame sampling is performed based on the text trajectory to obtain the second image frame. The second image frame is visually encoded, and then combined with the text trajectory aggregation result (representing refined text box-level visual features) for visual feature compression to obtain fixed-dimensional visual features; the visual features are then vectorized to obtain visual vectors. In addition, word vectorization is performed on the text trajectory to obtain text vectors. Feature encoding is performed on the text vectors and visual vectors to obtain the first feature, and a multilayer perceptron is used to process the first feature to obtain the type of text trajectory.
[0111] Step S505: Based on the type of the first text trajectory and the spatiotemporal attribute information of the text content in the first text trajectory, the first text trajectory is integrated to obtain the text structured result of the first video. See details... Figure 3 Step S505 of the illustrated embodiment will not be described again here.
[0112] The video text processing method provided in the embodiment can provide coding information from multiple angles by performing text and visual coding on the first text track and the first image frame when identifying the type of the first text track, wherein the spatio-temporal attribute information in the first text track provides rich information for type identification judgment, and the accuracy of type identification is further improved.
[0113] A video text processing method is provided in the embodiment, which can be used for the server, Figure 7 is a flowchart of the video text processing method according to the embodiment of the application, as Figure 7 shown, the flowchart includes the following steps:
[0114] In step S701, a first video is acquired, and a first image frame is obtained by frame sampling the first video. For details, refer to step S301 of the embodiment shown in Figure 3 , which will not be repeated here.
[0115] In step S702, text recognition is performed on the first image frame to obtain a text recognition result of the first image frame. For details, refer to step S302 of the embodiment shown in Figure 3 , which will not be repeated here.
[0116] In step S703, text tracking is performed based on the text recognition result to obtain a first text track of the first video.
[0117] The first text track includes text content and spatio-temporal attribute information of the text content.
[0118] Specifically, step S703 includes the following steps:
[0119] In step S7031, similarity calculation is performed on the text recognition results of two adjacent first image frames based on the text recognition results to obtain a similarity result.
[0120] According to the order of the first image frames, similarity calculation is performed on the text recognition results of two adjacent first image frames to obtain a similarity result of the text content between the two adjacent first image frames. The purpose of the similarity calculation is to merge the text content with a similarity higher than a threshold value and filter redundant and repetitive text content.
[0121] In some optional embodiments, the text recognition result includes text content, time attribute information of the text content, and spatial attribute information of the text content. The time attribute information includes a third image frame corresponding to the text content, and the spatial attribute information includes a text box size and a text box position of the text content.
[0122] Based on this, step S7031 includes the following steps:
[0123] Step b1, calculate the similarity of the text content of the two adjacent first image frames to obtain a first similarity value.
[0124] Step b2, calculate the similarity of the text box size of the two adjacent first image frames to obtain a second similarity value.
[0125] Step b3, calculate the similarity of the text box position of the two adjacent first image frames to obtain a third similarity value.
[0126] Step b4, based on the distance between the third image frames corresponding to the text content of the two adjacent first image frames, obtain a fourth similarity value, the size of the fourth similarity value is negatively related to the size of the distance.
[0127] Step b5, fuse the first similarity value, the second similarity value, the third similarity value and the fourth similarity value to obtain a similarity result.
[0128] For any two text contents, the similarity result includes four parts of the similarity value, which are the first similarity value representing the similarity of the text content, the second similarity value representing the similarity of the text box size, the third similarity value representing the similarity of the text box position, and the fourth similarity value obtained by the distance between the third image frames corresponding to the text content. The distance between the third image frames represents the interval frame number of the two third image frames. If the first third image frame is the 10th frame and the second third image frame is the 15th frame, the interval frame number between them is 5 frames; if the first third image frame is the 10th frame and the second third image frame is the 11th frame, the interval frame number between them is 1 frame. Therefore, the fourth similarity value of the interval frame number of 1 frame is greater than the fourth similarity value of the interval frame number of 5 frames.
[0129] It should be understood that since the first image frame is obtained by frame sampling the first video, even if the first image frames are adjacent, the frame numbers are not necessarily continuous. It is possible that the previous first image frame is the 10th frame and the adjacent next first image frame is the 15th frame.
[0130] After obtaining the four similarity values, the weighted calculation can be performed by combining the respective weights of the respective similarity values to obtain the similarity result. The respective weights of the respective similarity values can be the same or can be set according to actual requirements. In any case, the sum of the four weights is 1.
[0131] The similarity calculation is performed from multiple dimensions, which combines the text content, spatial information and time information, and improves the accuracy of the similarity calculation result.
[0132] Step S7032, grouping the text recognition result based on the similarity result to obtain an optional text track.
[0133] After obtaining the similarity results of the two text contents, the similarity results are compared with a preset threshold value. If the similarity results are higher than the preset threshold value, it is considered that the two text contents are consistent and belong to the same text track. Therefore, after the similarity calculation and track grouping of all text contents, a plurality of optional text tracks can be obtained.
[0134] In step S7033, the optional text tracks are screened to obtain a first text track.
[0135] The screening of the optional text tracks can be calculating an evaluation index of the optional text track, comparing the evaluation index with an index threshold value, and retaining the optional text track as the first text track if the evaluation index is higher than the index threshold value, otherwise discarding the optional text track. The evaluation index can be a score average. For example, in the optional text track, there can be text contents corresponding to a plurality of image frames. The similarity results of the text contents can be averaged to obtain the score average.
[0136] Of course, the evaluation index and the index threshold value can be set according to actual needs, and here they are not limited.
[0137] In step S704, a type of the first text track is identified based on the first image frame and the first text track to obtain the type of the first text track.
[0138] The type of the first text track is used to represent the function of the text content in the first image frame. For details, see Figure 5 The step S504 of the embodiment shown in the figure will not be described here again.
[0139] In step S705, the first text track is integrated based on the type of the first text track and the spatio-temporal attribute information of the text content of the first text track to obtain a text structured result of the first video.
[0140] Specifically, the spatio-temporal attribute information includes a start frame and an end frame corresponding to the text content of the first text track, and spatial information of the text content of the first text track. Based on this, the above step S705 includes:
[0141] In step S7051, the first text track is grouped according to the type of the first text track.
[0142] As described above, the first text track is grouped according to the type of each first text track to obtain the first text track under the same type.
[0143] In step S7052, the text content of the first text track of the same type is sorted based on the order of the start frame corresponding to the text content of the first text track.
[0144] After obtaining the first text tracks of the same type, there can be multiple first text tracks under the same type. For these first text tracks, the text content is sorted in combination with the order of the starting frames corresponding to the text content of the first text tracks.
[0145] For example, there are 3 first text tracks under type 1, the starting frame of the first text track 1 is the 10th frame, the starting frame of the first text track 2 is the 12th frame, and the starting frame of the first text track 3 is the 14th frame. The text content of the first text tracks is sorted according to the order of the starting frames, obtaining the text content of the first text track 1, the text content of the first text track 2, and the text content of the first text track 3.
[0146] In step S7053, if there are multiple third text tracks with the same starting frame and ending frame in the first text tracks of the same type, the text content of the third text tracks is sorted based on the spatial information of the text content, obtaining the sorting result of the text content of the third text tracks.
[0147] The text structured result of the first video includes the grouping result of the first text tracks and the sorting result of the text content in the same group.
[0148] It should be understood that in the process of sorting according to the starting frame, there can be multiple third text tracks with the same starting frame and ending frame in the first text tracks of the same type. For these third text tracks, the text content in the third text tracks can be sorted according to the spatial information of the text content, for example, the text content in the third text tracks is sorted in the order from top to bottom and from left to right, obtaining the sorting result.
[0149] In some optional embodiments, the text structured result of the first video further includes: a result of labeling the text content in the corresponding first image frame based on the labeling style corresponding to the type of the first text track.
[0150] Further, in the text structured result of the first video, in addition to the above-mentioned text structured content, there is also a labeling result of labeling the text content in the first image frame. Different types of text tracks have different labeling styles of the corresponding text content, for example, the text content can be distinguished by the color of the text box, or distinguished by the line type of the text box, and the like.
[0151] In the text structured result, there is also a labeling result in the image frame, and the labeling style corresponds to the type of the first text track, so that the text content in the image frame can be intuitively represented.
[0152] The video text processing method provided in this embodiment groups text content by trajectory through similarity calculation, which can filter redundant and repetitive text content. Furthermore, filtering based on trajectory grouping can remove false positives and other interference, ensuring the accuracy of the obtained first text trajectory. The first text trajectory is first grouped according to its type to ensure the categorization of text content of the same type; then, it is further sorted by combining the temporal and spatial information of the text content within each group, ensuring the accuracy of the structured text display.
[0153] As a specific application embodiment of this application, combined with Figure 1 The application scenario is illustrated. The terminal device uploads a first video to the server for video text processing. The server obtains the video text structured result of the first video by executing the video text processing method described in this embodiment, and then sends the text structured result to the terminal device. Accordingly, the text structured result is displayed on the interface of the terminal device.
[0154] For example, such as Figure 8 As shown, for each image frame, text boxes are used to annotate the text content based on the type of text trajectory to which the text content belongs. Figure 8 In the text box marked with a solid line 801, the text trajectory type is title; the text content marked with a wide dashed line 802, the text trajectory type is background; and the text content marked with a thin dashed line 803, the text trajectory type is subtitle.
[0155] For text structuring, the text trajectory is grouped and displayed, for example... Figure 8 The title 804 is shown. Under each type, the start and end image frame numbers to which the text content belongs are displayed. For example... Figure 8 In the text content 1, there are start frame 1 and end frame 1.
[0156] The display of video text processing results facilitates an intuitive understanding of the text content and its structured information. Specifically, text trajectory generation removes repetitive text, and the text structuring ensures the text content is displayed in the correct order.
[0157] This embodiment also provides a video text processing apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0158] This embodiment provides a video text processing device, such as... Figure 9As shown, comprising:
[0159] The acquisition module 901 is configured to acquire a first video and perform frame sampling on the first video to obtain a first image frame.
[0160] The text recognition module 902 is configured to perform text recognition on the first image frame to obtain a text recognition result of the first image frame.
[0161] The text tracking module 903 is configured to perform text tracking based on the text recognition result to obtain a first text track of the first video, the first text track comprising text content and spatiotemporal attribute information of the text content.
[0162] The type identification module 904 is configured to perform type identification based on the first image frame and the first text track to obtain a type of the first text track, the type of the first text track being used to represent a function of the text content in the first image frame.
[0163] The track integration module 905 is configured to integrate the first text track based on the type of the first text track and the spatiotemporal attribute information of the text content of the first text track to obtain a text structured result of the first video.
[0164] In some optional embodiments, the type identification module 904 comprises:
[0165] The text encoding unit is configured to perform text encoding on the first text track to obtain a text vector.
[0166] The visual encoding unit is configured to perform visual encoding on the first image frame to obtain a visual vector.
[0167] The feature encoding unit is configured to perform feature encoding based on the text vector and the visual vector to obtain a first feature.
[0168] The type identification unit is configured to perform type identification based on the first feature to obtain the type of the first text track.
[0169] In some optional embodiments, the visual encoding unit comprises:
[0170] The sampling subunit is configured to sample the first image frame based on a duration of the first text track to obtain a second image frame, and all text content in the second image frame covers the text content of the first text track.
[0171] The visual encoding subunit is configured to perform visual encoding on the second image frame to obtain the visual vector.
[0172] In some optional embodiments, the visual encoding subunit comprises:
[0173] aggregating the first text track based on the spatio-temporal attribute information of the first text track to obtain a second text track.
[0174] encoding the second image frame to obtain a visual feature.
[0175] vectorizing based on the second text track and the visual feature to obtain a visual vector.
[0176] In some optional embodiments, the text tracking module 903 includes:
[0177] calculating a similarity between the text recognition results of the adjacent two first image frames based on the text recognition results to obtain a similarity result.
[0178] grouping the text recognition results into text tracks based on the similarity result to obtain optional text tracks.
[0179] screening the optional text tracks to obtain the first text track.
[0180] In some optional embodiments, the text recognition result includes text content, time attribute information of the text content, and spatial attribute information of the text content, the time attribute information includes a third image frame corresponding to the text content, and the spatial attribute information includes a text box size and a text box position of the text content; the similarity calculation unit includes:
[0181] calculating a similarity between the text contents of the adjacent two first image frames to obtain a first similarity value.
[0182] calculating a similarity between the text box sizes of the adjacent two first image frames to obtain a second similarity value.
[0183] calculating a similarity between the text box positions of the adjacent two first image frames to obtain a third similarity value.
[0184] obtaining a fourth similarity value based on a distance between the third image frames corresponding to the text contents of the adjacent two first image frames, the fourth similarity value being negatively related to the distance.
[0185] fusing the first similarity value, the second similarity value, the third similarity value, and the fourth similarity value to obtain the similarity result.
[0186] In some optional embodiments, the spatio-temporal attribute information comprises a start frame and an end frame corresponding to the text content of the first text track, and spatial information of the text content of the first text track; the track integration module 905 comprises:
[0187] a grouping unit configured to group the first text tracks according to types of the first text tracks.
[0188] a first sorting unit configured to sort the text content of the first text tracks of the same type based on an order of the start frames corresponding to the text content of the first text tracks.
[0189] a second sorting unit configured to, if there are multiple third text tracks with the same start frame and end frame in the first text tracks of the same type, sort the third text tracks based on the spatial information of the text content of the third text tracks to obtain a text content sorting result of the third text tracks, and the text structure result of the first video comprises the grouping result of the first text tracks and the text content sorting result in the same group.
[0190] In some optional embodiments, the text structure result of the first video further comprises a result of labeling the text content in the corresponding first image frame based on a labeling style corresponding to the type of the first text track.
[0191] The video text processing apparatus provided by the embodiments of the present disclosure can execute the video text processing method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of executing the method. The further function description of each module and unit is the same as that of the corresponding embodiments, which will not be repeated here.
[0192] Figure 10 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure.
[0193] The following will be specifically described with reference to Figure 10 which shows a structural schematic diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present disclosure. The electronic device can include a processor (such as a central processor, a graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or loaded from a storage 1008 to a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0194] In general, the following devices can be connected to the I / O interface 1005: input devices 1006, including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 1007, including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 1008, including, for example, a magnetic tape, a hard disk, and the like; and communication devices 1009. The communication devices 1009 can allow the electronic device to communicate wirelessly or via a wire with other devices to exchange data. Although Figure 10 An electronic device having various devices is shown, but it is understood that all of the shown devices are not required and more or fewer devices can be implemented or provided instead.
[0195] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 1009, or installed from the storage devices 1008, or installed from the ROM 1002. When the computer program is executed by the processor 1001, the above-described functions defined in the video text processing method of the embodiments of the present disclosure are performed.
[0196] Figure 10 The electronic device shown is merely an example and should not impose any limitation on the functions and scope of use of embodiments of the present disclosure.
[0197] The embodiments of the present application also provide a computer-readable storage medium, and the above-mentioned method according to the embodiments of the present application can be implemented in hardware, firmware, or as computer code recorded on a storage medium, or as computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded from a network and stored in a local storage medium, so that the method described herein can be processed by such software on a storage medium using a general-purpose computer, a special-purpose processor, or programmable or special-purpose hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk or a solid state disk, etc.; further, the storage medium can also include a combination of the above-mentioned types of storage. It can be understood that the computer, processor, microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code, which, when accessed and executed by the computer, processor or hardware, implements the video text processing method shown in the above embodiments.
[0198] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, can invoke or provide methods and / or technical solutions according to the present application. Those skilled in the art should understand that the form of computer program instructions in computer readable medium includes but is not limited to source files, executable files, installation package files and the like, and accordingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.
[0199] Although the embodiments of the present application are described in conjunction with the drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope defined by the appended claims.
Claims
1. A method for processing video text, characterized in that, The method includes: Acquire a first video and perform frame sampling on the first video to obtain a first image frame; Perform text recognition on the first image frame to obtain the text recognition result of the first image frame; Based on the text recognition results, text tracking is performed to obtain the first text trajectory of the first video. The first text trajectory includes text content and spatiotemporal attribute information of the text content. Type recognition is performed based on the first image frame and the first text trajectory to obtain the type of the first text trajectory. The type of the first text trajectory is used to characterize the function of the text content in the first image frame. Based on the type of the first text trajectory and the spatiotemporal attribute information of the text content in the first text trajectory, the first text trajectory is integrated to obtain the text structured result of the first video.
2. The method according to claim 1, characterized in that, The step of performing type recognition based on the first image frame and the first text trajectory to obtain the type of the first text trajectory includes: The first text trajectory is text-encoded to obtain a text vector; The first image frame is visually encoded to obtain a visual vector; Based on the text vector and the visual vector, feature encoding is performed to obtain the first feature; Based on the first feature, type recognition is performed to obtain the type of the first text trajectory.
3. The method according to claim 2, characterized in that, The step of visually encoding the first image frame to obtain a visual vector includes: Based on the duration of the first text trajectory, the first image frame is sampled to obtain a second image frame, and the text content in all the second image frames covers the text content of the first text trajectory. The visual vector is obtained by performing visual encoding based on the second image frame.
4. The method according to claim 3, characterized in that, The step of visual encoding based on the second image frame to obtain the visual vector includes: The first text trajectory is aggregated based on its spatiotemporal attribute information to obtain the second text trajectory; Visual features are obtained by visually encoding the second image frame. The visual vector is obtained by vectorizing the second text trajectory and the visual features.
5. The method according to claim 1, characterized in that, The step of performing text tracking based on the text recognition result to obtain the first text trajectory of the first video includes: For the text recognition results of two adjacent first image frames, a similarity calculation is performed based on the text recognition results to obtain a similarity result; Based on the similarity results, the text recognition results are grouped into trajectories to obtain selectable text trajectories; The first text trajectory is obtained by filtering the optional text trajectories.
6. The method according to claim 5, characterized in that, The text recognition result includes the text content, the time attribute information of the text content, and the spatial attribute information of the text content. The time attribute information includes the third image frame corresponding to the text content, and the spatial attribute information includes the text box size and text box position of the text content. The similarity calculation based on the text recognition results to obtain similarity results includes: Calculate the similarity of the text content of two adjacent first image frames to obtain the first similarity value; Calculate the similarity of the text box sizes of two adjacent first image frames to obtain the second similarity value; Calculate the similarity between the text box positions of two adjacent first image frames to obtain the third similarity value; A fourth similarity value is obtained based on the distance between the third image frames corresponding to the text content of two adjacent first image frames. The magnitude of the fourth similarity value is negatively correlated with the magnitude of the distance. The first similarity value, the second similarity value, the third similarity value, and the fourth similarity value are fused to obtain the similarity result.
7. The method according to claim 1, characterized in that, The spatiotemporal attribute information includes the start and end frames corresponding to the text content in the first text trajectory, as well as the spatial information of the text content in the first text trajectory; the integration of the first text trajectory based on its type and the spatiotemporal attribute information of the text content in the first text trajectory to obtain the text structured result of the first video includes: The first text trajectory is grouped according to its type; For the same type of first text trajectory, the text content of the same type of first text trajectory is sorted based on the order of the starting frames corresponding to the text content in the first text trajectory; If there are multiple third text trajectories of the same type with the same start and end frames in the first text trajectory, then the text content of the third text trajectory is sorted based on the spatial information of the text content in the third text trajectory to obtain the text content sorting result of the third text trajectory. The text structuring result of the first video includes the grouping result of the first text trajectory and the text content sorting result within the same group.
8. The method according to claim 1, characterized in that, The text structuring results of the first video also include: The result of annotating the text content in the corresponding first image frame based on the annotation style corresponding to the type of the first text trajectory.
9. A video text processing device, characterized in that, The device includes: An acquisition module is used to acquire a first video and perform frame sampling on the first video to obtain a first image frame; The text recognition module is used to perform text recognition on the first image frame and obtain the text recognition result of the first image frame; The text tracking module is used to perform text tracking based on the text recognition results to obtain the first text trajectory of the first video. The first text trajectory includes text content and spatiotemporal attribute information of the text content. A type recognition module is used to perform type recognition based on the first image frame and the first text trajectory to obtain the type of the first text trajectory. The type of the first text trajectory is used to characterize the function of the text content in the first image frame. The trajectory integration module is used to integrate the first text trajectory based on the type of the first text trajectory and the spatiotemporal attribute information of the text content in the first text trajectory to obtain the text structured result of the first video.
10. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the video text processing method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the video text processing method according to any one of claims 1 to 8.
12. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the video text processing method according to any one of claims 1 to 8.