Video text task processing method and device, equipment, medium and program product

By performing text recognition and tracking on video frames and combining this with a large language model to process video text tasks, the problem of low accuracy in existing technologies has been solved, achieving efficient and accurate video text task processing.

CN120976905APending Publication Date: 2025-11-18BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511072926.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

The accuracy of video text processing in existing technologies is low, which affects downstream businesses.

Method used

By acquiring video frames and performing text recognition, combined with text tracking and spatiotemporal attribute information, a large language model is used to process video text tasks, reducing the acquisition of repetitive content and improving processing efficiency and accuracy.

Benefits of technology

It enables end-to-end video text task processing, improving processing efficiency and accuracy while reducing redundant output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976905A_ABST
    Figure CN120976905A_ABST
Patent Text Reader

Abstract

The invention discloses a video character task processing method and device, equipment, a medium and a program product, and relates to the field of video processing technologies, artificial intelligence technologies, large model technologies and large language models.The method comprises the steps that a first video is obtained, and frame sampling is conducted on the first video to obtain a first image frame; performing character recognition on the first image frame to obtain a character recognition result of the first image frame; performing text tracking based on the character recognition result to obtain a first text track of the first video, the first text track including text content and time-space attribute information of the text content; and processing the first image frame, the first text track and first prompt information based on the first text task processing model to obtain a processing result of the target video text task, the first prompt information being used for indicating the first text task processing model to process the target video text task. According to the method, the processing efficiency and accuracy of the video text task are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, artificial intelligence, large model and large language model, in particular to a processing method and device for a video text task, equipment, medium and program product. BACKGROUND

[0002] Texts appearing in a video usually contain rich semantic information and are indispensable features for understanding the content of the video. A video text task is used to extract different types of texts in a video and output, for example, to extract titles, subtitles and backgrounds in the video, etc. The processing result of the video text task is used for consumption by downstream businesses, and accordingly, the accuracy of the video text task processing will affect the downstream business. SUMMARY

[0003] Therefore, the present application provides a processing method and device for a video text task, equipment, medium and program product to solve the problem of the accuracy of the video text task processing.

[0004] In a first aspect, the present application provides a processing method for a video text task, comprising:

[0005] obtaining a first video and frame sampling the first video to obtain a first image frame;

[0006] performing text recognition on the first image frame to obtain a text recognition result of the first image frame;

[0007] performing text tracking based on the text recognition result to obtain a first text track of the first video, the first text track comprising text content and spatio-temporal attribute information of the text content;

[0008] processing the first image frame, the first text track and first prompt information based on a first text task processing model to obtain a processing result of a target video text task, the first prompt information being used to instruct the first text task processing model to perform processing of the target video text task.

[0009] In a second aspect, the present application provides a processing device for a video text task, comprising:

[0010] a video obtaining module configured to obtain a first video and frame sample the first video to obtain a first image frame;

[0011] a text recognition module configured to perform text recognition on the first image frame to obtain a text recognition result of the first image frame;

[0012] a text tracking module, configured to perform text tracking based on the character recognition result to obtain a first text track of the first video, the first text track including text content and spatio-temporal attribute information of the text content;

[0013] a task processing module, configured to process the first image frame, the first text track and first prompt information based on a first character task processing model to obtain a processing result of a target video character task, the first prompt information being used to instruct the first character task processing model to perform processing of the target video character task.

[0014] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, which are communicatively connected with each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the video character task processing method in the first aspect or any of the corresponding implementation manners thereof.

[0015] In a fourth aspect, the present application provides a computer readable storage medium, which stores computer instructions, and the computer instructions are used to make a computer execute the video character task processing method in the first aspect or any of the corresponding implementation manners thereof.

[0016] In a fifth aspect, the present application provides a computer program product, which comprises computer instructions, and the computer instructions are used to make a computer execute the video character task processing method in the first aspect or any of the corresponding implementation manners thereof.

[0017] The video character task processing method provided by the embodiments of the present application comprises the following steps: obtaining a first video, and performing frame sampling on the first video to obtain a first image frame; performing character recognition on the first image frame to obtain a character recognition result of the first image frame; performing text tracking based on the character recognition result to obtain a first text track of the first video, the first text track including text content and spatio-temporal attribute information of the text content; and processing the first image frame, the first text track and first prompt information based on a first character task processing model to obtain a processing result of a target video character task, the first prompt information being used to instruct the first character task processing model to perform processing of the target video character task. The method performs character recognition and text tracking on the first image frame obtained by frame sampling on the first video to obtain the first text track. Since the first text track includes not only the text content but also the spatio-temporal attribute information of the text content, the text content is represented in the time and space dimensions, which can reduce the acquisition of repeated text content on one hand, and can process the target video character task in combination with the first prompt information on the other hand, so as to obtain an end-to-end processing result of the video character task, and improve the processing efficiency and accuracy of the video character task. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the specific embodiments or the related art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the related art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0019] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present application;

[0020] Figure 2 is a first flowchart of a processing method of a video text task according to an embodiment of the present application;

[0021] Figure 3 is a second flowchart of a processing method of a video text task according to an embodiment of the present application;

[0022] Figures 4a-4b is a schematic diagram of a text recognition result of a first image frame according to an embodiment of the present application;

[0023] Figure 5 is a schematic diagram of the working principle of a first text task processing model according to an embodiment of the present application;

[0024] Figure 6 is a structural block diagram of a processing device of a video text task according to an embodiment of the present application;

[0025] Figure 7 is a hardware structure schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0027] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the type, use range, use scenario and the like of the personal information involved in the present application should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.

[0028] For example, in response to receiving an active request of a user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide personal information to the software or hardware, such as an electronic device, an application program, a server, or a storage medium, performing the operation of the technical solution of the present disclosure according to the prompt information.

[0029] As an optional but non-limiting implementation, in response to receiving an active request of a user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in the form of text. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0030] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0031] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0032] In the related art, for processing of a video text task, a video text task to be processed is generally given, for example, subtitle recognition, background recognition, title recognition, and the like, and then text recognition is performed on image frames in the video, Key Information Extraction (KIE) is performed based on key information of the image, and a certain label value is assigned to each text in the image frame to determine and extract text of a corresponding category. The category corresponds to the type of the video text task.

[0033] This processing manner uses multi-modal information of the image to further filter the text in the image and extract text content corresponding to the video text task. However, in this manner, the text is filtered using the category label, and the accuracy of text filtering may be low due to incorrect recognition of the category label.

[0034] Based on this, the embodiment of the application provides a processing method of a video text task, the method comprising: acquiring a first video, and performing frame sampling on the first video to obtain a first image frame; performing text recognition on the first image frame to obtain a text recognition result of the first image frame; performing text tracking based on the text recognition result to obtain a first text track of the first video, the first text track comprising text content and spatio-temporal attribute information of the text content; and processing the first image frame, the first text track and first prompt information based on a first text task processing model to obtain a processing result of a target video text task, the first prompt information being used to instruct the first text task processing model to perform processing of the target video text task.

[0035] The method performs text recognition and text tracking on the first image frame obtained by frame sampling on the first video to obtain the first text track, wherein the first text track comprises not only text content but also spatio-temporal attribute information of the text content, and the text content is represented from the time and space dimensions, which can reduce the acquisition of repeated text content on one hand, and can process the target video text task in combination with the first prompt information on the other hand, so as to obtain an end-to-end processing result of the video text task and improve the processing efficiency and accuracy of the video text task.

[0036] As an optional application scenario of the embodiment of the application, as shown in Figure 1 As an optional application scenario of the embodiment of the application, as shown in Figure 1 The terminal device 110 is installed with the application 101, and the user 130 can interact with the application 101 through the terminal device 110 and / or an access device of the terminal device 110.

[0037] Exemplarily, the application 101 can be any application, which can provide a video text processing application. In Figure 1 As shown in the application scenario, if the application 101 is in an active state, the terminal device 110 can present an interface 102 of the application 101. The interface 102 can comprise various pages that can be provided by the application 101, such as an interactive page, a setting page, a query page and the like.

[0038] In some embodiments, the terminal device 110 is in communication connection with the server 120 to implement the provision of services of the application 101. The terminal device 110 can be a mobile terminal, a fixed terminal or a portable terminal, etc., including but not limited to a mobile phone, a desktop computer, a notebook computer, a multimedia tablet, an electronic book device, a game device or any combination of the above, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface, and the server 120 can be various types of computing systems, servers that can provide computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc.

[0039] It should be noted that, Figure 1 is only an example of an application scenario, and does not limit the protection scope of the present disclosure.

[0040] The embodiments of the present disclosure will be described below with reference to the drawings. It should be understood that the pages shown in the drawings are only examples, and various page designs can actually exist. Various graphical elements in the page can have different arrangements and different visual representations, one or more elements of which can be omitted or replaced, and one or more other elements can also exist, which are not limited in the embodiments of the present disclosure. In addition, the embodiments are mainly described below for the server.

[0041] According to the embodiments of the present application, a video text task processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from here.

[0042] A video text task processing method is provided in the present embodiment, which can be used in the server described above, Figure 2 is a flowchart of the video text task processing method according to the embodiments of the present application, as Figure 2 shown, the flow includes the following steps:

[0043] Step S201, obtaining a first video and performing frame sampling on the first video to obtain a first image frame.

[0044] The first video is a video that needs to be processed by video text, which includes multiple consecutive image frames. The first video can be a locally stored video, or a video obtained through a communication connection with other devices, or a locally generated video, etc. The source of the first video is not limited here.

[0045] Frame sampling means that the first video is sampled according to a preset sampling frequency. The preset sampling frequency is set according to actual needs, for example, sampling once every 1S, or sampling once every 2S, etc.

[0046] After obtaining the preset sampling frequency, the first video is frame sampled to obtain the first image frame. It should be understood that the first image frame is a general term for multiple image frames, and does not specifically refer to one image frame. Without loss of generality, the first frame of the first video belongs to the first image frame, and subsequent frame sampling is performed according to the preset sampling frequency to obtain the first image frame. The first image frame is used for subsequent text recognition and other processing.

[0047] In step S202, text recognition is performed on the first image frame to obtain a text recognition result of the first image frame.

[0048] The subsequent processing described for the first image frame is performed for each first image frame, rather than a specific first image frame. The text in the first image frame includes, but is not limited to, a title, a background, and subtitles in the image frame, and the like. For example, there is an article in the first image frame, and there is text on the article. Therefore, the text on the article is included in the text recognition result. Further, there is a subtitle in the first image frame. Therefore, the subtitle is included in the text recognition result.

[0049] The text recognition (Optical Character Recognition, OCR) performed on the first image frame can be a non-deep learning OCR method, a deep learning-based method, or the like. Taking the non-deep learning OCR method as an example, the first image frame is preprocessed, including but not limited to image noise reduction, binarization, tilt correction, and text positioning and segmentation. The preprocessed image frame is extracted for geometric features, statistical features, and structural features, and the like. The extracted features are classified and recognized to obtain the text recognition result.

[0050] The text recognition result of the first image frame includes text content, spatial information of the text content, and an image frame to which the text content belongs. For example, if the text content in the first image frame is continuously present, a line of text can be recognized as a text block and output as a text content. The spatial information of the text content is used to represent the position of the text content in the first image frame, for example, the size of the text box corresponding to the text content and the coordinates of the text box. The image frame to which the text content belongs, i.e., the correspondence between the text content and the image frame, can also be referred to as the time information of the text content.

[0051] Since there can be text content in multiple positions in the first image frame, multiple text contents can be generated in the same first image frame. The spatial information of the multiple text contents and the image frame to which the text content belongs are recorded in the text recognition result. Without loss of generality, the image frame to which the text content belongs can be represented by the frame number of the image frame.

[0052] For example, as shown in FIG. 4, the text content recognized in the first image frame includes text content 401, text content 402, text content 403, and text content 404. These text contents are located at different positions in the first image frame. In addition to outputting the specific text content, the position information of the text content is also outputted. Figure 4a

[0053] Figure 4a ​The frame number of the first image frame shown is 10. Figure 4b The first image frame shown has a frame number of 15, and its corresponding text content includes: text content 401, text content 403, text content 404 and text content 405. Accordingly, the position information of the text content is output, etc.

[0054] Step S203: Based on the text recognition results, perform text tracking to obtain the first text trajectory of the first video.

[0055] The first text trajectory includes the text content and the spatiotemporal attribute information of the text content.

[0056] Because text content in the same position may appear in multiple consecutive frames, for example, Figure 4a The text content 401 in the middle, not only in Figure 4a The text 401 appears in frame 10 and also in frame 15 of Figure 4. It should be noted that the same position in frames 11 through 14 also contains text content 401, and text content 401 first appears in frame 10 and last appears in frame 15. Therefore, the duration of text content 401's appearance is from frame 10 to frame 15.

[0057] In other words, if the same text content appears consecutively for 6 frames, then by performing text tracking on the text recognition results, the same text content can be represented in the form of text trajectories, which can reduce the output of repeated text content.

[0058] Text tracking can determine whether text in adjacent frames belongs to the same text by calculating the feature similarity of text content. For example, for multiple first image frames, text detection and recognition are performed on each frame. Feature extraction, recognition, and fusion are performed through an end-to-end text tracking network to obtain a fused feature descriptor. Based on this, the preset tracking trajectory is updated to obtain the text tracking trajectory.

[0059] Text trajectories are primarily used to represent the location information of text content, including which image frames the text content is distributed in and its specific location within each image frame. For example, the target tracking trajectory of text content in a video might indicate that the text content exists continuously from frame 10 to frame 15 of a video, and its specific location information in each frame can also be determined through the text trajectory.

[0060] Specifically, the first text track includes text content and spatio-temporal attribute information of the text content. The text content refers to the text itself, for example, all the text of a line of text. The spatio-temporal attribute information includes temporal attribute information and spatial attribute information. The temporal attribute information represents the image frame to which the text content belongs, for example, the 10th frame to the 15th frame. The spatial attribute information represents the position information of the text content in the image frame, for example, the position information of the text box corresponding to the text content and the size of the text box, and the like.

[0061] Continuing with the above example, as shown in Figure 4a and Figure 4b , it is shown that the 10th frame, and the text content in the image frames from the 11th frame to the 14th frame is consistent (not only the text content itself, but also the position information of the text content, and the like), and the text content of the image frames changes from the 15th frame. Figure 4a

[0062] Therefore, through text tracking, the text track of the text content 401 can be obtained: the text content 401, the position information of the text content 401, and the image frame to which the text content 401 belongs is the 10th frame to the 15th frame; the text track of the text content 402: the text content 402, the position information of the text content 402, and the image frame to which the text content 402 belongs is the 10th frame to the 14th frame; the text track of the text content 403: the text content 403, the position information of the text content 403, and the image frame to which the text content 403 belongs is the 10th frame to the 15th frame; the text track of the text content 404: the text content 404, the position information of the text content 404, and the image frame to which the text content 404 belongs is the 10th frame to the 15th frame; and the text track of the text content 405: the text content 405, the position information of the text content 405, and the image frame to which the text content 405 belongs is the 15th frame.

[0063] In step S204, the first image frame, the first text track, and the first prompt information are processed based on the first text task processing model to obtain a processing result of the target video text task.

[0064] The first prompt information is used to instruct the first text task processing model to process the target video text task.

[0065] The first text task processing model is constructed based on a large language model. The input of the first text task processing model includes the first image frame, the first text track, and the first prompt information, and the output of the first text task processing model includes the processing result of the target video text task. Specifically, in the first text task processing module, since the processing object includes images and text, a text encoding module and a visual encoding module are required. The encoding results and the first prompt information are input into the large language module of the first text task processing model to obtain the corresponding processing result.​

[0066] The first prompt information is used to instruct the first text task processing model to process the target video text task, that is, the description of the target video text task needs to be included in the first prompt information. For example, extracting subtitles in a video, extracting titles in a video, and the like, which are specifically set according to actual needs.

[0067] Exemplarily, in addition to the description of the target video text task, the token of the first image frame and the description of the first text track, and the like can also be included in the first prompt information. Including multi-dimensional information in the first prompt information can improve the processing accuracy of the first text task processing model.

[0068] The video text task processing method provided in this embodiment can obtain the first text track by performing text recognition and text tracking on the first image frame obtained by frame extraction on the first video. Since the first text track not only includes text content but also includes the spatio-temporal attribute information of the text content, the text content is characterized from the time and space dimensions, which can reduce the acquisition of repeated text content on the one hand, and can process the target video text task in combination with the first prompt information, thereby obtaining an end-to-end video text task processing result and improving the processing efficiency and accuracy of the video text task.

[0069] In this embodiment, a video text task processing method is provided, which can be used in the server, Figure 3 is a flowchart of the video text task processing method according to an embodiment of the application, as Figure 3 shown, the flow includes the following steps:

[0070] In step S301, a first video is obtained, and frame sampling is performed on the first video to obtain a first image frame. For details, refer to step S201 of the embodiment shown in Figure 2 , which will not be repeated here.

[0071] In step S302, text recognition is performed on the first image frame to obtain a text recognition result of the first image frame. For details, refer to step S202 of the embodiment shown in Figure 2 , which will not be repeated here.

[0072] In step S303, text tracking is performed based on the text recognition result to obtain a first text track of the first video.

[0073] The first text track includes text content and spatio-temporal attribute information of the text content.

[0074] Specifically, the above step S303 includes:

[0075] In step S3031, the similarity between the text recognition results of the two adjacent first image frames is calculated based on the text recognition results, to obtain a similarity result.

[0076] According to the order of the first image frames, the similarity between the text recognition results of the two adjacent first image frames is calculated, to obtain a similarity result of the text content between the two adjacent first image frames. The purpose of the similarity calculation is to merge the text content with a similarity higher than a threshold value, and filter the redundant and repetitive text content.

[0077] In some optional embodiments, the text recognition result includes the text content, time attribute information of the text content, and spatial attribute information of the text content. The time attribute information includes the third image frame corresponding to the text content, and the spatial attribute information includes the text box size and the text box position of the text content.

[0078] Based on this, the above step S7031 includes:

[0079] In step a1, the similarity of the text content of the two adjacent first image frames is calculated, to obtain a first similarity value.

[0080] In step a2, the similarity of the text box size of the two adjacent first image frames is calculated, to obtain a second similarity value.

[0081] In step a3, the similarity of the text box position of the two adjacent first image frames is calculated, to obtain a third similarity value.

[0082] In step a4, the fourth similarity value is obtained based on the distance between the third image frames corresponding to the text content of the two adjacent first image frames. The size of the fourth similarity value is negatively related to the size of the distance.

[0083] In step a5, the first similarity value, the second similarity value, the third similarity value, and the fourth similarity value are fused to obtain the similarity result.

[0084] For any two text contents, the similarity result includes four parts of similarity values, which are the first similarity value representing the similarity of the text contents, the second similarity value representing the similarity of the text box size, the third similarity value representing the similarity of the text box position, and the fourth similarity value obtained based on the distance between the third image frames corresponding to the text contents. The distance between the third image frames represents the interval frame number of the two third image frames. If the first third image frame is the 10th frame and the second third image frame is the 15th frame, the interval frame number between them is 5 frames. If the first third image frame is the 10th frame and the second third image frame is the 11th frame, the interval frame number between them is 1 frame. Therefore, the fourth similarity value when the interval frame number is 1 frame is greater than the fourth similarity value when the interval frame number is 5 frames.

[0085] It should be understood that since the first image frames are obtained by frame sampling the first video, even if adjacent first image frames, the frame sequence number is not necessarily continuous. It is possible that the previous first image frame is the 10th frame, and the adjacent next first image frame is the 15th frame.

[0086] After obtaining the four similarity values, the similarity results can be obtained by weighted calculation combined with the respective weights of the respective similarity values. The respective weights of the respective similarity values can be the same or can be set according to actual needs. In any case, the sum of the four weights is 1.

[0087] The similarity calculation from multiple dimensions combines text content, spatial information, and time information, improving the accuracy of the similarity calculation result.

[0088] Step S3032, grouping the text recognition result based on the similarity result to obtain an optional text trajectory.

[0089] After obtaining the similarity results of the two text contents, the similarity results are compared with the preset threshold value. If it is higher than the preset threshold value, it is considered that the two text contents are consistent and belong to the same text trajectory. Therefore, after the above similarity calculation and trajectory grouping are performed on all text contents, a plurality of optional text trajectories can be obtained.

[0090] Step S3033, screening the optional text trajectory to obtain a first text trajectory.

[0091] The screening of the optional text trajectory can be to calculate an evaluation index of the optional text trajectory, compare the evaluation index with an index threshold value, and if it is higher than the index threshold value, it is retained as the first text trajectory, otherwise discarded. The evaluation index can be a score average. For example, in the optional text trajectory, there can be text content corresponding to a plurality of image frames. The similarity results of these text contents can be averaged to obtain the score average.

[0092] Of course, the evaluation index and the index threshold value can be set according to actual needs, and here they are not limited in any way.

[0093] Step S304, processing the first image frame, the first text trajectory, and the first prompt information based on the first text task processing model to obtain a processing result of the target video text task.

[0094] The first prompt information is used to instruct the first text task processing model to process the target video text task.

[0095] Specifically, the above step S304 includes:

[0096] At step S3041, the text encoding module of the first text task processing model is used to perform text encoding on the first text track to obtain a text vector.

[0097] The first text task processing model includes a text encoding module, a visual encoding module, and a large language module. Specifically, the text encoding module is configured to perform text encoding on the first text track, the visual encoding module is configured to perform visual encoding on the first image frame, and the large language module is configured to obtain the processing result of the target video text task based on the text encoding result, the visual encoding result, and the first prompt information.

[0098] In some optional embodiments, the step S3041 includes:

[0099] At step b1, the text encoding unit of the text encoding module is used to encode the text content in the first text track to obtain a first vector.

[0100] At step b2, the space-time encoding unit of the text encoding module is used to encode the space-time attribute information of the text content in the first text track to obtain a second vector.

[0101] At step b3, the first vector and the second vector are spliced to obtain a text vector.

[0102] Since the first text track includes text content and space-time attribute information of the text content, different types of information are processed by corresponding encoding units. Specifically, the text content in the first text track is encoded by the text encoding unit to obtain a first vector. The space-time attribute information in the first text track is encoded by the space-time encoding unit to obtain a second vector.

[0103] As described above, the space-time attribute information includes time attribute information and space attribute information. The time attribute information represents the duration of the first text track, i.e., the start frame and the end frame. The space attribute information represents the position information of the text content in the corresponding image frame, such as the coordinates of the text box corresponding to the text content, the size of the text box corresponding to the text content, and the like.

[0104] For example, after the space-time attribute information of the first text track is processed by word segmentation, fixed-dimensional feature vectors are obtained by layout, space-time encoders (Layout Embedder, Spatial Embedder, Temporal Embedder) and layout, space-time projectors (Layout Projector, Spatial Projector, Temporal Projector). That is, the first vector obtained by the text content and the second vector obtained by the space-time attribute information are spliced to obtain a text vector.

[0105] For the text vector, not only the text content in the first text track is included, but also the spatio-temporal attribute information of the text content is included. The two are encoded by corresponding encoding units to obtain the text vector, thereby ensuring the content richness of the text vector and providing a basic condition for subsequent target video text task processing.

[0106] In step S3042, the visual encoding module of the first text task processing model is used to perform visual encoding on the first image frame to obtain a visual vector.

[0107] The visual encoding of the first image frame can be that all the first image frames need to be encoded, or that the first image frames are frame-sampled and the frame-sampled results are encoded, and the like, which can be set according to actual needs.

[0108] In some optional embodiments, the above step S3042 includes:

[0109] In step c1, the first image frame is sampled based on the duration of the first text track to obtain a second image frame, and the text content in all the second image frames covers the text content of the first text track.

[0110] In step c2, the visual encoding module is used to perform visual encoding on the second image frame to obtain a visual vector.

[0111] As described above, the same text content can appear in multiple consecutive image frames. In addition, the time information of each text content is represented in the first text track, that is, the starting frame and the ending frame corresponding to each text content are included. Accordingly, the starting frame and the ending frame of the text content correspond to the duration of the first text track. For example, if the starting frame of the text content of the first text track is the 10th frame and the ending frame is the 15th frame, then the duration of the first text track is from the 10th frame to the 15th frame.

[0112] Since the text content can exist in multiple consecutive image frames, for the text content, these image frames belong to repeated information. Therefore, the first image frame is sampled according to the duration of the first text track to obtain a second image frame. That is, the text content in all the obtained second image frames can cover the text content of all the first text tracks.

[0113] For example, the time information of each first text track can be compared to determine the repeated image frames. One frame in the repeated image frames is taken as a second image frame, and the non-repeated image frames are sampled to ensure that the text content of the obtained second image frames can cover the text content in all the first text tracks.

[0114] By sampling the first image frames, the number of image frames for subsequent visual coding can be reduced. That is, the number of second image frames is less than the number of first image frames. Visual coding is then performed on the second image frames to obtain visual vectors.

[0115] In the process of visual coding on the first image frames, the first image frames are sampled according to the duration of the first text track, so that the text content in all the obtained second image frames can cover the text content of all the first text tracks, on the one hand, the image frames with the same text content are removed, and on the other hand, the processing efficiency is improved.

[0116] In step S3043, the first prompt information is obtained by filling the prompt template of the target video text task based on the first image frames and the first text track.

[0117] The prompt template can correspond to the video text task, or can be applicable to all video text tasks. In an exemplary first, the prompt template is provided with fixed text description, and variables corresponding to image frames, text tracks and video text tasks, after obtaining the first image frames, the first text track and the target video text task, the obtained information is used to fill the prompt template to obtain the first prompt information.

[0118] It should be understood that the first prompt information can also include other content, which is specifically set according to actual needs, and here it is not limited in any way.

[0119] In some optional embodiments, the arrangement order of the first image frames and the first text track in the first prompt information includes: arranging all the first image frames at a first position of the first prompt information, and interlacing the text content of the first text track and the spatio-temporal attribute information of the text content at a second position of the first prompt information.

[0120] The order of the first position and the second position is not limited in any way, and can be set according to actual needs. In the first position, all the first image frames are arranged, and in the second position, the interlacing arrangement is performed, that is, the text content of the first text track and the spatio-temporal attribute information are interlaced.

[0121] Exemplarily, the arrangement information at the second position can be represented as: text content of text track 1, spatio-temporal attribute information of text track 1; text content of text track 2, spatio-temporal attribute information of text track 2; …; text content of text track N, spatio-temporal attribute information of text track N.

[0122] In some other optional embodiments, the arrangement order of the first image frames and the first text track in the first prompt information includes: interlacing the first image frames, the text content of the first text track and the spatio-temporal attribute information of the text content.

[0123] The staggered arrangement of the first image frame and the first text trajectory can be characterized as follows: image frame 1, text content corresponding to image frame 1, and spatiotemporal attribute information of the text content; image frame 2, text content corresponding to image frame 2, and spatiotemporal attribute information of the text content; ...; text content of image frame N, text content corresponding to image frame N, and spatiotemporal attribute information of the text content.

[0124] When arranging the first image frame and the first text trajectory in the first prompt information, the text content can be arranged in an alternating manner with its spatiotemporal attributes, or the text content and its spatiotemporal attribute information in the first image frame and the first text trajectory can be arranged in an alternating manner. That is, by using an alternating arrangement, the model's understanding efficiency and generation accuracy can be improved.

[0125] Step S3044: Input the text vector, visual vector, and first prompt information into the large language module of the first text task processing model to obtain the processing result of the target video text task.

[0126] The input to the large language module in the first text task processing model includes text vectors, visual vectors, and the first prompt information, and the output is the processing result.

[0127] It should be noted that the specific module structures of the text encoding module, visual encoding module, and large language module described in the embodiments of this application are set according to actual needs, and no limitations are made here.

[0128] For example, such as Figure 5 As shown, a first text trajectory is obtained by performing text recognition and text tracking on the first image frame. The first image frame is then sampled to obtain a second image frame, which is encoded using a visual encoder to obtain a visual vector. The text content in the first text trajectory is encoded using a text encoding unit to obtain a first vector, and the first text trajectory itself is encoded using a spatiotemporal encoding unit to obtain a second vector. The first and second vectors are concatenated to obtain a text vector. The visual vector, text vector, and first prompt information are then input into the large language module of the first text task processing model to obtain the processing result of the target video text task.

[0129] The video text task processing method provided in this embodiment groups text content into trajectories through similarity calculation, which can filter out redundant and repetitive text content. Furthermore, filtering based on trajectory grouping can remove false positives and other interference, ensuring the accuracy of the obtained first text trajectory. The first text trajectory is included in the first prompt information, thereby enriching the content of the prompt information and further improving the accuracy of the target video text task processing results.

[0130] In some optional embodiments, the training process of the first text processing task model comprises:

[0131] Step d1, training the initial text processing task model by using the first sample set to obtain a second text processing task model, the first sample set comprising sample pictures and text labels in the pictures, and the training task of the first sample set comprising recognizing the text in the sample pictures.

[0132] Step d2, training the second text processing task model by using the second sample set to obtain a third text processing task model, the second sample set comprising the first frame images of the first sample video, the text tracks of the first frame images, and the text labels of the first sample video, and the training task of the second sample set comprising processing of a single type of first video text task.

[0133] Step d3, training the third text processing task model by using the third sample set to obtain the first text processing task model, the third sample set comprising the second frame images of the second sample video, the text tracks of the second frame images, and the text labels of the second sample video, and the training task of the second sample set comprising processing of multiple types of second video text tasks.

[0134] The training process of the first text processing task model is a progressive multi-stage learning mode, the first stage training is to train the initial text processing task model to recognize the text in the sample pictures to obtain the second text processing task model, the second stage training is to train the second text processing task model to process a single type of first video text task, for example, only processing the caption recognition task or only processing the title recognition task, etc.

[0135] After the second stage training is completed, the third text processing task model is obtained. The third stage training is to train the third text processing task model, and in the training, a second video text task is randomly specified for processing to obtain the first text processing model.

[0136] Exemplarily, the input of the first stage training is a single sample picture, the training task is text recognition, and the purpose of this stage training is to improve the ability of the model in recognizing the text labels and text detection in the video service scenario.

[0137] The input of the second stage training is the first frame images of the first sample video, the text tracks, and the text labels, and the task of this stage training is to extract one of the title text, the caption text, and the background text of the first sample video in an end-to-end manner.

[0138] The input of the third stage training is the second frame images of the second sample video, the text tracks, and the text labels, and in the training of this stage, the training task is randomly specified, and the purpose is to let the model learn multiple tasks at the same time.

[0139] In the process of training the first text processing task model, a progressive multi-stage learning manner is adopted to process training tasks from easy to difficult and from few to many, so as to improve the training efficiency of the first text processing task model and the processing accuracy of the model.

[0140] As a specific application embodiment of the present application, in combination with the application scenario shown in Figure 1 , the user performs the target video text task and the given first video through the interaction with the interface 102 provided by the application 101 in the terminal device 110. The terminal device 110 sends the first video and the target video text task to the server 120, and the server 120 is deployed with the video text task processing method described in the present application. The processing result corresponding to the target video text task is obtained by executing the method, and the processing result is fed back to the terminal device 110. Accordingly, the processing result of the target video text task of the first video can be obtained in the terminal device 110.

[0141] In the present embodiment, a video text task processing apparatus is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments, and will not be described again. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware is also possible and is contemplated.

[0142] The present embodiment provides a video text task processing apparatus, as shown in Figure 6 , comprising:

[0143] The video acquisition module 601 is configured to acquire a first video and perform frame sampling on the first video to obtain a first image frame.

[0144] The text recognition module 602 is configured to perform text recognition on the first image frame to obtain a text recognition result of the first image frame.

[0145] The text tracking module 603 is configured to perform text tracking based on the text recognition result to obtain a first text track of the first video, the first text track comprising text content and spatiotemporal attribute information of the text content.

[0146] The task processing module 604 is configured to process the first image frame, the first text track, and first prompt information based on a first text processing task model to obtain a processing result of a target video text task, the first prompt information being used to instruct the first text processing task model to process the target video text task.

[0147] In some optional embodiments, the task processing module 604 comprises:

[0148] a text encoding unit, configured to perform text encoding on the first text track by using a text encoding module of the first text task processing model to obtain a text vector.

[0149] a visual encoding unit, configured to perform visual encoding on the first image frame by using a visual encoding module of the first text task processing model to obtain a visual vector.

[0150] a template filling unit, configured to fill a prompt template of the target video text task based on the first image frame and the first text track to obtain first prompt information.

[0151] a text processing unit, configured to input the text vector, the visual vector, and the first prompt information into a large language module of the first text task processing model to obtain a processing result of the target video text task.

[0152] In some optional embodiments, the text encoding unit comprises:

[0153] a first encoding subunit, configured to encode text content in the first text track by using the text encoding module in the text encoding unit to obtain a first vector.

[0154] a second encoding subunit, configured to encode spatial and temporal attribute information of the text content in the first text track by using a spatial and temporal encoding unit in the text encoding module to obtain a second vector.

[0155] a vector splicing subunit, configured to splice the first vector and the second vector to obtain the text vector.

[0156] In some optional embodiments, the visual encoding unit comprises:

[0157] a frame sampling subunit, configured to sample the first image frame based on a duration of the first text track to obtain a second image frame, and all text content in the second image frame covers the text content of the first text track.

[0158] a visual encoding subunit, configured to perform visual encoding on the second image frame by using the visual encoding module to obtain the visual vector.

[0159] In some optional embodiments, the arrangement order of the first image frame and the first text track in the first prompt information comprises:

[0160] arranging all the first image frames at a first position of the first prompt information, and arranging the text content of the first text track and the spatial and temporal attribute information of the text content at a second position of the first prompt information in an interleaved manner;

[0161] or,

[0162] The text content of the first image frame, the first text track, and the spatio-temporal attribute information of the text content are interleaved.

[0163] In some optional embodiments, the text tracking module 603 comprises:

[0164] The similarity calculation unit is configured to perform similarity calculation on the character recognition results of the adjacent two first image frames based on the character recognition results to obtain similarity results.

[0165] The track grouping unit is configured to perform track grouping on the character recognition results based on the similarity results to obtain optional text tracks.

[0166] The track screening unit is configured to screen the optional text tracks to obtain the first text track.

[0167] In some optional embodiments, the method further comprises:

[0168] The first training module is configured to train the initial character processing task model by using the first sample set to obtain a second character processing task model, the first sample set comprises sample pictures and character labels in the pictures, and a training task of the first sample set comprises recognizing characters in the sample pictures.

[0169] The second training module is configured to train the second character processing task model by using a second sample set to obtain a third character processing task model, the second sample set comprises first frame extraction images of a first sample video, text tracks of the first frame extraction images, and character labels of the first sample video, and a training task of the second sample set comprises processing of a single type of first video character task.

[0170] The third training module is configured to train the third character processing task model by using a third sample set to obtain the first character processing task model, the third sample set comprises second frame extraction images of a second sample video, text tracks of the second frame extraction images, and character labels of the second sample video, and a training task of the third sample set comprises processing of multiple types of second video character tasks.

[0171] The video character task processing apparatus provided by the embodiments of the present disclosure can perform the video character task processing method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method. The further function description of each of the above modules and units is the same as that of the corresponding embodiments, which will not be described here.

[0172] Figure 7 A structural schematic diagram of an electronic device provided by the embodiments of the present disclosure.

[0173] The following will be specifically described with reference to Figure 7which shows a structural schematic diagram suitable for being used to implement an electronic device in the embodiments of the present disclosure. The electronic device can include a processor (such as a central processor, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a memory 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for operation of the electronic device are also stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0174] Generally, the following devices can be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device with various devices is shown, but it should be understood that all the shown devices are not required to be implemented or possessed, and more or less devices can be alternatively implemented or possessed.

[0175] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product including a computer program carried on a non-transitory computer readable medium, the computer program containing program codes for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device 709, or installed from the memory 708, or installed from the ROM 702. When the computer program is executed by the processor 701, the above-mentioned functions defined in the processing method of the video text task of the embodiments of the present disclosure are performed.

[0176] Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0177] The embodiments of the present application further provide a computer readable storage medium, and the method according to the embodiments of the present application can be implemented in hardware, firmware, or recorded in a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine readable storage medium and downloaded through a network and stored in a local storage medium, so that the method described herein can be processed by such software on a storage medium using a general purpose computer, a special purpose processor, or programmable or special hardware. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state disk, etc. Further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that the computer, processor, microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, when the software or computer code is accessed and executed by the computer, processor, or hardware, the processing method of the video text task shown in the above embodiments is implemented.

[0178] Part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be called or provided. Those skilled in the art should understand that the form of computer program instructions in a computer readable medium includes but is not limited to source files, executable files, installation package files, etc. Correspondingly, the way of executing computer program instructions by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.

[0179] Although the embodiments of the present application are described in conjunction with the accompanying drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application, and such modifications and changes fall within the scope defined by the appended claims.

Claims

1. A method for processing video text tasks, characterized in that, include: Acquire a first video, and perform frame sampling on the first video to obtain a first image frame; Perform text recognition on the first image frame to obtain the text recognition result of the first image frame; Based on the text recognition results, text tracking is performed to obtain the first text trajectory of the first video. The first text trajectory includes text content and spatiotemporal attribute information of the text content. The first image frame, the first text trajectory, and the first prompt information are processed based on the first text task processing model to obtain the processing result of the target video text task. The first prompt information is used to instruct the first text task processing model to process the target video text task.

2. The method according to claim 1, characterized in that, The processing of the first image frame, the first text trajectory, and the first prompt information based on the first text task processing model to obtain the processing result of the target video text task includes: The first text trajectory is encoded using the text encoding module of the first text task processing model to obtain a text vector; The first image frame is visually encoded using the visual encoding module of the first text task processing model to obtain a visual vector; Based on the first image frame and the first text trajectory, the prompt template of the target video text task is filled to obtain the first prompt information; The text vector, the visual vector, and the first prompt information are input into the large language module of the first text task processing model to obtain the processing result of the target video text task.

3. The method according to claim 2, characterized in that, The step of encoding the first text trajectory using the text encoding module of the first text task processing model to obtain a text vector includes: The text content in the first text trajectory is encoded using the text encoding unit in the text encoding module to obtain a first vector; Using the spatiotemporal encoding unit in the text encoding module, the spatiotemporal attribute information of the text content in the first text trajectory is encoded to obtain the second vector; The first vector and the second vector are concatenated to obtain the text vector.

4. The method according to claim 2, characterized in that, The step of visually encoding the first image frame using the visual encoding module of the first text task processing model to obtain a visual vector includes: Based on the duration of the first text trajectory, the first image frame is sampled to obtain a second image frame, and the text content in all the second image frames covers the text content of the first text trajectory. The visual encoding module is used to perform visual encoding on the second image frame to obtain the visual vector.

5. The method according to claim 2, characterized in that, The arrangement order of the first image frame and the first text trajectory in the first prompt information includes: All the first image frames are arranged at the first position of the first prompt information, and the text content of the first text trajectory and the spatiotemporal attribute information of the text content are arranged alternately at the second position of the first prompt information. or, The first image frame, the text content of the first text trajectory, and the spatiotemporal attribute information of the text content are arranged in an alternating manner.

6. The method according to claim 1, characterized in that, The step of performing text tracking based on the text recognition result to obtain the first text trajectory of the first video includes: For the text recognition results of two adjacent first image frames, a similarity calculation is performed based on the text recognition results to obtain a similarity result; Based on the similarity results, the text recognition results are grouped into trajectories to obtain selectable text trajectories; The first text trajectory is obtained by filtering the optional text trajectories.

7. The method according to claim 1, characterized in that, Also includes: An initial text processing model is trained using a first sample set to obtain a second text processing model. The first sample set includes sample images and text labels in the images. The training task of the first sample set includes recognizing the text in the sample images. The second text processing task model is trained using the second sample set to obtain the third text processing task model. The second sample set includes the first frame image of the first sample video, the text trajectory of the first frame image, and the text label of the first sample video. The training task of the second sample set includes processing of a single type of first video text task. The third text processing task model is trained using the third sample set to obtain the first text processing task model. The third sample set includes the second frame image of the second sample video, the text trajectory of the second frame image, and the text label of the second sample video. The training task of the second sample set includes the processing of various types of second video text tasks.

8. A processing apparatus for video text tasks, characterized in that, include: The video acquisition module is used to acquire a first video and perform frame sampling on the first video to obtain a first image frame; The text recognition module is used to perform text recognition on the first image frame and obtain the text recognition result of the first image frame; The text tracking module is used to perform text tracking based on the text recognition results to obtain the first text trajectory of the first video. The first text trajectory includes text content and spatiotemporal attribute information of the text content. The task processing module is used to process the first image frame, the first text trajectory, and the first prompt information based on the first text task processing model to obtain the processing result of the target video text task. The first prompt information is used to instruct the first text task processing model to process the target video text task.

9. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the video text task processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the video text task processing method according to any one of claims 1 to 7.

11. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the video text task processing method according to any one of claims 1 to 7.