Video text tracking method, video processing method, device, equipment and medium
By using particle generation and similarity calculation methods in video text tracking, the problems of long video text tracking processing time and high computing resources in the prior art are solved, and the speed of video text tracking is improved and the computing resources are saved.
Patent Information
- Application Number
- CN202011565988.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-12-25
AI Technical Summary
The prior art requires detection and matching of each video frame in video text tracking, resulting in a long processing time and high computing resources consumption, making it difficult to meet the practical application needs.
By determining the text box in the first video frame of the video, and generating a plurality of particles in the second video frame adjacent to the first video frame, determining the second text box according to the position of the particles, calculating the similarity between the first text box and the second text box, selecting the second text box with the highest similarity as the third text box, and then determining the target tracking track of the video text.
It effectively reduces the processing time required during video text tracking, improves the speed of video text tracking, saves computing resources, and simplifies the analysis and review process of video content.
Smart Images

Figure CN113392689B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a video text tracking method, a video processing method, a device, a device and a medium. Background Art
[0002] In recent years, the artificial intelligence technology has developed rapidly and achieved good application effects in the fields of image classification, face recognition, autonomous driving, etc. For example, the artificial intelligence technology can be used to track the trajectory of text in a video to determine which video frames contain the same continuously displayed text information, which is convenient for video content analysis, review, etc.
[0003] In the related art, generally when tracking the trajectory of text, it is necessary to detect and match each video frame. This implementation method takes a long time, consumes a large amount of computing resources, has a high cost, and is difficult to meet the actual application requirements. In summary, the problems existing in the prior art need to be solved urgently. Summary of the Invention
[0004] The purpose of this application is to solve at least to some extent one of the technical problems existing in the prior art.
[0005] To this end, an object of an embodiment of this application is to provide a video text tracking method, a video processing method, a device, a device and a medium. The video text tracking method can effectively improve the processing speed of video text tracking and reduce the consumption of computing resources.
[0006] One aspect of this application provides a video text tracking method, including the following steps:
[0007] Determine a first text box from the first video frame of the video;
[0008] Generate a plurality of particles at a position corresponding to the first text box in the second video frame of the video; the first video frame and the second video frame are adjacent;
[0009] Determine a plurality of second text boxes in the second video frame according to the positions of the respective particles;
[0010] Determine a first similarity between the first text box and each of the second text boxes, and use the second text box with the highest first similarity as the third text box;
[0011] Determine a target tracking trajectory of the video text according to the first text box and the third text box; the target tracking trajectory is used to represent the position information of the video text.
[0012] Another aspect of this application provides a video processing method, including:
[0013] Obtain multiple consecutive video frames of a video;
[0014] By using the video text tracking method described above, obtain the target tracking trajectories of multiple video texts in the video;
[0015] According to each of the target tracking trajectories, extract the video frames to obtain a set of key frames of the video.
[0016] Another aspect of the present application provides a video text tracking device, including:
[0017] A first processing module, configured to determine a first text box from a first video frame of a video;
[0018] A particle generation module, configured to generate a plurality of particles at a position corresponding to the first text box in a second video frame of the video; the first video frame and the second video frame are adjacent;
[0019] A second processing module, configured to determine a plurality of second text boxes in the second video frame according to the positions of the respective particles;
[0020] A similarity determination module, configured to determine a first similarity between the first text box and each of the second text boxes, and use the second text box with the highest first similarity as a third text box;
[0021] A trajectory determination module, configured to determine a target tracking trajectory of the video text according to the first text box and the third text box; the target tracking trajectory is used to characterize the position information of the video text.
[0022] Another aspect of the present application provides an electronic device, including:
[0023] At least one processor;
[0024] At least one memory, configured to store at least one program;
[0025] When the at least one program is executed by the at least one processor, the at least one processor implements the video text tracking method or the video processing method described above.
[0026] Another aspect of the present application provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the video text tracking method or the video processing method described above when executed by the processor.
[0027] Another aspect of the present application also provides a computer program product or a computer program, which includes computer instructions stored in the aforementioned computer-readable storage medium; the processor of the aforementioned electronic device can read the computer instructions from the aforementioned computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the aforementioned video text tracking method or video processing method.
[0028] The advantages and beneficial effects of the present application will be partially given in the following description, partially will become obvious from the following description, or be understood through the practice of the present application:
[0029] In the video text tracking method in the embodiments of the present application, when tracking and recognizing video text, after determining the first text box from the first video frame, a plurality of particles are generated at the position corresponding to the first text box in the second video frame adjacent to the first video frame, the second text box is determined according to the positions of the respective particles, then the similarity between the first text box and each second text box is determined, the second text box with the highest similarity is used as the third text box, and the target tracking trajectory of the video text is determined according to the third text box and the first text box. In the video text tracking method in the embodiments of the present application, when tracking video text, the possible positions of the same text in adjacent video frames are determined by the position of the text in the current video frame, without having to detect the adjacent video frames from the beginning, which can effectively reduce the processing time required in the video text tracking process, improve the speed of video text tracking, and save computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the accompanying drawings of the related technical solutions in the embodiments of the present application or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present application, and those skilled in the art can also obtain other accompanying drawings according to these drawings without creative efforts.
[0031] Figure 1 It is a schematic diagram of the implementation environment of a video text tracking method provided by an embodiment of the present application;
[0032] Figure 2 It is a schematic flow chart of a video text tracking method provided by an embodiment of the present application;
[0033] Figure 3 It is a schematic diagram of a first text tracking network adopted in the video text tracking method provided by an embodiment of the present application;
[0034] Figure 4Another schematic diagram of the first text tracking network adopted in the video text tracking method provided by the embodiment of the present application;
[0035] Figure 5 A schematic diagram of a second text tracking network adopted in the video text tracking method provided by the embodiment of the present application;
[0036] Figure 6 A schematic diagram of particle generation in the video text tracking method provided by the embodiment of the present application;
[0037] Figure 7 The first schematic diagram of a text tracking network based on the Yolo-v3 network adopted in the video text tracking method provided by the embodiment of the present application;
[0038] Figure 8 The second schematic diagram of a text tracking network based on the Yolo-v3 network adopted in the video text tracking method provided by the embodiment of the present application;
[0039] Figure 9 A schematic diagram of the process of a video processing method provided by the embodiment of the present application;
[0040] Figure 10 A schematic diagram of obtaining a key frame set in a video processing method provided by the embodiment of the present application;
[0041] Figure 11 A schematic diagram of the structure of a video text tracking device provided by the embodiment of the present application;
[0042] Figure 12 A schematic diagram of the structure of an electronic device provided by the embodiment of the present application. Detailed implementation manners
[0043] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not constitute a specific limitation on the present application.
[0044] Next, the technical field related to the present application will be introduced first:
[0045] Artificial Intelligence (AI): This technology uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also involves researching the design principles and implementation methods of various intelligent machines to enable them to have functions of perception, reasoning, and decision-making.
[0046] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0047] Computer Vision (CV): Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing image processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0048] The video text tracking method provided in the embodiments of the present application can be used in various application scenarios of video processing. For example, these application scenarios include but are not limited to video content analysis, video promotion, and video review, etc. Specifically, taking the review of the text content in a video as an example, through this video text tracking method, a certain video frame in the video can be detected first to determine the text area therein, obtaining a text box. According to the position of the text box in this video frame, at the corresponding position in the video frame adjacent to this video frame (which can be a video frame before this video frame or a video frame after this video frame), a text box with the highest similarity to this text box is determined as the position of the text area in the adjacent video frame, thereby determining the target tracking trajectory of the video text. When it is necessary to extract the content of the video text, it can be extracted from the corresponding text box through the position information in the target tracking trajectory.
[0049] When extracting or reviewing video text, since the object to be recognized is the text in the video frame, and the refresh rate of the video frame is generally relatively fast. Therefore, if the text information in each video frame is directly detected and extracted, a large amount of redundant data will appear, which is not conducive to subsequent analysis and processing. In the related art, generally, a trajectory tracking method is adopted, that is, it is determined which video frames contain the same text information that is continuously displayed. For this part of the repeated text information, extraction or review is only performed once. However, in this implementation method, when tracking video text, it is necessary to detect the position of the video text in each frame, and then determine the target tracking trajectory of the video text through matching. It requires a long processing time, consumes a large amount of computing resources, has a high cost, and the application benefit is not high.
[0050] In view of this, in the embodiments of the present application, a text box is determined in the first video frame of a video, and based on the positions of particles around the text box, the possible positions where the text box may appear in the second video frame adjacent to the first video frame are determined, thereby achieving the matching of video text trajectory tracking. When processing adjacent video frames, the video text tracking method does not need to perform the step of detecting the first text box in the first video frame every time, that is, it does not need to determine the position of the text box in each video frame from the beginning. On the one hand, it can effectively reduce the time required in the video text trajectory tracking process, save computing resources, and improve the processing speed. On the other hand, after determining the target tracking trajectory of the video text, when it is necessary to extract or review the text content of the video, any video frame covered by the target tracking trajectory of the video text can be processed, which facilitates the analysis and review of the video content. It should be noted that the video in the embodiments of the present application may refer to an aggregate composed of multiple consecutive pictures, and a video frame refers to one of the pictures in the aggregate. Therefore, it can be understood that the aggregate includes but is not limited to the content that can be played on a multimedia platform, files in formats such as MPEG (Moving Picture Experts Group), AVI (Audio Video Interleaved), nAVI (new AVI), ASF (Advanced Streaming Format), MOV (the movie format of software QuickTime), WMV (Windows Media Video), 3GP (3rd Generation Partnership Project), RM (RealMedia), RMVB (RealMedia Variable Bitrate), FLV (FLASH VIDEO), MP4 (Moving Picture Experts Group 4), etc., or dynamic graphics, multiple pictures during the change of lyrics in music playback, and so on.
[0051] Figure 1 is an optional application environment schematic diagram of the video text tracking method provided by the embodiments of the present application. Refer to Figure 1, the video text tracking method provided by the embodiments of the present application can be applied to a video text tracking system 100. The video text tracking system 100 may include a terminal 110 and a server 120, and the specific number of the terminal 110 and the server 120 can be set arbitrarily. The terminal 110 and the server 120 can establish a communication connection through a wireless network or a wired network. The wireless network or the wired network uses standard communication technologies and / or protocols. The network can be set as the Internet, or any other network, such as including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. The terminal 110 can send the video that needs text tracking to the server 120 based on the established communication connection. The server 120 performs corresponding processing by executing the video text tracking method provided by the embodiments of the present application to obtain the target tracking trajectory of the video text in the video, and then returns the processing result to the terminal 110.
[0052] In some embodiments, the above terminal 110 can be any kind of electronic product that can perform human-computer interaction with a user in one or more ways such as a keyboard, a touchpad, a touch screen, a remote control, a mouse, voice interaction or a handwriting device. For example, the electronic product can include but not limited to a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a pocket PC (PPC), a tablet computer, etc. The server 120 can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server that provides services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0053] It should be understood that Figure 1 What is shown is only an optional implementation environment of the video text tracking method of the embodiments of the present application. In actual applications, it is not fixedly implemented through Figure 1 the video text tracking system 100 therein. For example, in some embodiments, the video text tracking method can be independently implemented by the terminal 110 locally. For example, the video text tracking method can be executed by some application programs installed on the terminal 110. The application program can be a video playback software, a web platform, etc. Similarly, the video text tracking method can also be independently implemented by the server 120.
[0054] Reference Figure 2 , Figure 2 FIG. 1 is an optional flowchart of a video text tracking method provided by an embodiment of the present application. This method can be applied to the above-mentioned video text tracking system 100, Figure 2 and the method in FIG. 1 includes steps S201-S205.
[0055] Step S201: Determine a first text box from the first video frame of the video.
[0056] In the embodiment of the present application, a video can be divided into a plurality of consecutive video frames in advance through video stream decoding technology, and then any one of these video frames is selected as the first video frame. Then, the text area in the first video frame is determined, and a text box is delimited according to the text area, denoted as the first text box. Specifically, the size and shape of the first text box can be determined according to the text area in the video frame. For example, if the subtitles in the video are to be detected, the first text box can be set as a rectangular box. It can be understood that the first text box is mainly used to represent the position information of the text in the first video frame. The video text here can be the text in the subtitles or the text at any position in the video screen. The text box in the embodiment of the present application includes the bounding box line and the text content of the text therein or the picture content presented based on the text.
[0057] Specifically, step S201 in the embodiment of the present application can be implemented through step S210, or through steps S220-S230, where:
[0058] S210: Detect the text area in the first video frame to obtain the first text box.
[0059] In the embodiments of the present application, the text area refers to the area containing text. As described above, this area can be any area in the first video frame. Here, the detection of the text area can be completed by using a machine learning model in artificial intelligence technology. For the machine learning model, the detection task of the text area can be regarded as a segmentation task, that is, to segment the part of the video frame containing text from the complete video frame, and then the first text box can be determined according to the outer shape of this part of the picture. Specifically, in some embodiments, the shape of the first text box can be preset in advance. For example, it can be set as a rectangle. The first text box can be as small as possible on the premise of including the segmented text picture. For example, the minimum circumscribed rectangle of the text area detected in the first video frame can be taken as the first text box. In some embodiments, the shape and size of the first text box can be preset in advance, and then the center position of the text area detected in the first video frame can be used as the center of the first text box to determine the first text box. Here, it should be noted that the text area is an area where some video texts gather. When there are texts in multiple places in the first video frame, multiple text areas can be determined in the first video frame at the same time, and a corresponding first text box can be determined for each text area.
[0060] In the embodiments of the present application, an initial tracking trajectory can be established in advance. For example, obtain an adjacent video frame of the first video frame, detect the text area in the first video frame to obtain a text box, detect the text area in the adjacent video frame to obtain another text box, determine the similarity between the two text boxes, and determine whether the two text boxes match according to the similarity. If they match, an initial tracking trajectory can be established according to the position information of the two text boxes. It should be noted that in the embodiments of the present application, an adjacent video frame adjacent to the above adjacent video frame can also be continuously obtained, the similarity between the text boxes can be identified and determined through the above steps, and the similarity between the text boxes can be continuously determined according to the similarity to determine whether to add it to the initial tracking trajectory, that is, the initial tracking trajectory can include the position information of the text boxes in two or more video frames. In the embodiments of the present application, the position information includes but is not limited to the coordinates of the text box, and the coordinates can be the coordinates of the center of the text box or the coordinates of each vertex. Optionally, in the embodiments of the present application, a preset threshold can also be set. When it is detected that there is the same matching text box in the video frames of the continuous preset threshold number of frames, that is, an initial tracking trajectory is established; when it is detected that there is no same matching text box in the video frames of the continuous preset threshold number of frames, no initial tracking trajectory is established.
[0061] Step S220: Obtain a third video frame adjacent to the first video frame from the video.
[0062] In the embodiments of the present application, adjacent video frames refer to video frames adjacent in chronological order. For example, among multiple video frames of a video, there are video frame A, video frame B, and video frame C that are consecutive in chronological order. Then both video frame A and video frame C are adjacent to video frame B. It can be understood that the third video frame can be located before the first video frame or after the first video frame in chronological order, and the first video frame is located between the second video frame and the third video frame. Among them, the third video frame can be the adjacent video frame adjacent to the first video frame in the above-mentioned initial tracking trajectory.
[0063] Step S230: Use the text box in the first video frame that matches the third video frame as the first text box.
[0064] In the embodiments of the present application, the first text box is determined by matching the text box in the first video frame with the text box in the third video frame. Here, the purpose of matching the first video frame and the third video frame is to determine the initial tracking trajectory of video text. Specifically, step S230 can be implemented through the following steps S240 - S280.
[0065] Step S240: Input the first video frame and the third video frame into the first text tracking network; the first text tracking network includes a first detection branch network and a second detection branch network.
[0066] As Figure 3 shown, taking the first video frame 101 and the third video frame 102 as an example for illustration, both the first video frame 101 and the third video frame 102 have text regions, and the text "XXXXX" is in the text regions. Input the first video frame 101 and the third video frame 102 into the first text tracking network 310. In the embodiments of the present application, the first text tracking network 310 can be pre-trained, and the first detection branch network and the second detection branch network are used to detect the text box and output the detection result. In the embodiments of the present application, the structures of the first detection branch network and the second detection branch network can be the same and can share weights. Specifically, in the embodiments of the present application, the detection branch network is a network capable of performing text detection, which can be, but is not limited to, the Yolo network (You Only Look Once), CNN (Convolutional Neural Networks), LSTM (Long-Short Term Memory artificial neural network), etc.
[0067] Step S250: Detect the first video frame through the first detection branch network to obtain the fourth text box.
[0068] Specifically, the first video frame 101 is input into the first detection branch network, and through the processing of the first detection branch network, a fourth text box is obtained, and the fourth text box is used to represent the position information of the text in the first video frame.
[0069] Step S260: Detect the third video frame through the second detection branch network to obtain a fifth text box.
[0070] Specifically, the third video frame 102 is input into the second detection branch network, and through the processing of the second detection branch network, a fifth text box is obtained, and the fifth text box is used to represent the position information of the text in the third video frame.
[0071] Step S270: Determine the second similarity between the fourth text box and the fifth text box.
[0072] Refer to Figure 4 , in the embodiment of the present application, the first text tracking network 310 may further include a first tracking branch network and a second tracking branch network, and step S270 may include steps S301 - S303:
[0073] Step S301: Extract the fourth text box through the first tracking branch network to obtain a first feature vector.
[0074] Specifically, in the embodiment of the present application, the first tracking branch network is connected to the first detection branch network, and the first tracking branch network receives the output of the first detection branch network as input, so as to perform feature extraction on the fourth text box to obtain a first feature vector.
[0075] Step S302: Extract the fifth text box through the second tracking branch network to obtain a second feature vector.
[0076] Specifically, the second tracking branch network is connected to the second detection branch network, and the second tracking branch network receives the output of the second detection branch network as input, so as to perform feature extraction on the fifth text box to obtain a second feature vector. In the embodiment of the present application, the structures of the first tracking branch network and the second tracking branch network may be the same and may share weights.
[0077] Step S303: Determine the second similarity according to the first feature vector and the second feature vector.
[0078] Specifically, the second similarity may be determined by the Euclidean distance, Manhattan distance, Minkowski distance, or cosine similarity, etc. between the first feature vector and the second feature vector.
[0079] Step S280: When the second similarity is greater than the first threshold, determine the fourth text box as the first text box.
[0080] It can be understood that the first threshold can be adjusted as needed. When the second similarity is greater than the first threshold, the fourth text box is determined as the first text box. As Figure 3 shown, after obtaining the second similarity by performing similarity matching between the fourth text box and the fifth text box, the second similarity is compared with the first threshold. When the second similarity is greater than the first threshold, the recognition result is obtained, that is, the fourth text box is determined as the first text box.
[0081] Referring to Figure 5 , specifically, step S230 can also be implemented through steps S310 - S360.
[0082] Step S310: Input the first video frame and the third video frame into the second text tracking network; the second text tracking network includes a third detection branch network and a fourth detection branch network; the third detection branch network includes a first sub - network and a second sub - network, and the fourth detection branch network includes a third sub - network and a fourth sub - network.
[0083] Specifically, Figure 5 in, still taking the aforementioned first video frame 101 and third video frame 102 as examples for illustration, input the first video frame 101 and the third video frame 102 into the second text tracking network 320. Similarly, the second text tracking network can also be pre - trained. The first sub - network and the second sub - network in the third detection branch network, and the third sub - network and the fourth sub - network in the fourth detection branch network are all used to detect text boxes and output detection results. Among them, the first sub - network and the second sub - network can receive the same input information, and the third sub - network and the fourth sub - network can receive the same input information. In the embodiments of the present application, the structures and weight parameters of the third detection branch network and the fourth detection branch network can also be set to be the same. It should be noted that in the embodiments of the present application, the third detection branch network and the fourth detection branch network can also respectively include more than two sub - networks, and the number of sub - networks in the third detection branch network and the fourth detection branch network can be the same or different.
[0084] Step S320: Detect the first video frame through the first sub - network to obtain a sixth text box, and detect the first video frame through the second sub - network to obtain a seventh text box;
[0085] Specifically, step S320 can be implemented through steps S401 - S402:
[0086] Step S401: Downsample the first video frame by a first multiple through the first sub - network to perform feature extraction on the first video frame and detect and obtain a sixth text box;
[0087] Step S402: Downsample the second multiple through the second sub-network to perform feature extraction on the first video frame and detect the seventh text box.
[0088] In the embodiments of the present application, downsampling is a processing method for image compression. After the downsampling operation, the image size will be reduced, and the degree of reduction is related to the sampling period of downsampling. In the embodiments of the present application, the first multiple and the second multiple of downsampling are used for feature extraction. The purpose is to extract image features of different depths at different image scales for text region detection. Specifically, the difference between the first multiple and the second multiple can be adjusted according to actual needs and is not limited herein. It can be understood that when the number of sub-networks is more than two, the first video frame can also be subjected to feature extraction by setting a third multiple different from both the first multiple and the second multiple, etc., to obtain more different text box detection results for improving the matching accuracy.
[0089] Step S330: Detect the third video frame through the third sub-network to obtain the eighth text box, and detect the third video frame through the fourth sub-network to obtain the ninth text box;
[0090] Specifically, step S330 can be implemented through steps S403 - S404:
[0091] Step S403: Downsample the first multiple through the third sub-network to perform feature extraction on the third video frame and detect the eighth text box;
[0092] Step S404: Downsample the second multiple through the fourth sub-network to perform feature extraction on the third video frame and detect the ninth text box.
[0093] In the embodiments of the present application, when extracting features from the third video frame, different sampling multiples of the third sub-network and the fourth sub-network can also be used. Moreover, the sampling multiple of the third sub-network can be the same as that of the aforementioned first sub-network, and the sampling multiple of the fourth sub-network can also be the same as that of the aforementioned second sub-network for facilitating subsequent matching.
[0094] Step S340: Determine the third similarity between the sixth text box and the eighth text box, and determine the fourth similarity between the seventh text box and the ninth text box;
[0095] In the embodiments of the present application, when determining the third similarity between the sixth text box and the eighth text box and determining the fourth similarity between the seventh text box and the ninth text box, it can be implemented in the manner of step S270.
[0096] Step S350: Determine the fifth similarity according to the third similarity and the fourth similarity;
[0097] Specifically, step S350 can be implemented through steps S501 - S504:
[0098] Step S501, obtain the confidence levels of the sixth text box, the seventh text box, the eighth text box, and the ninth text box;
[0099] Step S502, determine the first weight according to the average confidence level of the sixth text box and the eighth text box;
[0100] Specifically, according to the confidence level of the sixth text box and the confidence level of the eighth text box, calculate the average value to obtain the average confidence level of the sixth text box and the eighth text box as the first weight.
[0101] Step S503, determine the second weight according to the average confidence level of the seventh text box and the ninth text box;
[0102] Specifically, according to the confidence level of the seventh text box and the confidence level of the ninth text box, calculate the average value to obtain the average confidence level of the seventh text box and the ninth text box as the second weight.
[0103] Step S504, perform weighted summation on the third similarity and the fourth similarity according to the first weight and the second weight to obtain the fifth similarity.
[0104] Specifically, the fifth similarity can be calculated through the following formula:
[0105]
[0106] is the confidence level of the text box b1 of the i-th sub-network of the third detection branch network, is the confidence level of the text box b2 of the i-th sub-network of the fourth detection branch network, is the similarity of b1 and b2 in the corresponding i-th sub-network, is the similarity result of b1 and b2.
[0107] For example, when i = 1, is the confidence level of the text box b1 of the first sub-network of the third detection branch network (i.e., the confidence level of the sixth text box in the first sub-network), is the confidence level of the text box b1 of the first sub-network of the fourth detection branch network (i.e., the confidence level of the eighth text box in the third sub-network), is the similarity of the sixth text box and the eighth text box (the third similarity). Similarly, when i = 2, it will not be elaborated here.
[0108] It can be understood that when the third detection branch network and the fourth detection branch network have more than two sub-networks, the confidence levels of different text boxes can also be determined according to steps S501 - S504, and the average confidence level and weight between different text boxes can be determined, and weighted summation can be performed using the above formula to determine the fifth similarity.
[0109] Among them, when determining the above initial tracking trajectory, when there are more than two text boxes in two consecutive video frames of video text, all text boxes in the two video frames can be detected first, and the similarity between each text box in one video frame and each text box in the other video frame can be calculated, and combined with the intersection over union between the text boxes to form a similarity matrix, and the bipartite graph maximum weight matching method is used to pair the text box combinations so that the pairing result satisfies the maximum sum of the similarity and the intersection over union, thereby completing the pairing of each text box. It should be noted that the pairing can also be achieved by setting a pairing threshold. When the sum of the similarity and the intersection over union of the paired text boxes is greater than or equal to the pairing threshold, the pairing is considered successful, that is, the matching is successful. Among them, the similarity between two text boxes refers to the similarity result determined by the calculation formula in step S503.
[0110] Step S360: When the fifth similarity is greater than the second threshold, determine the sixth text box or the seventh text box as the first text box.
[0111] It can be understood that the second threshold can be adjusted according to the actual situation. When the fifth similarity is greater than the second threshold, one of the sixth text box or the seventh text box can be randomly selected as the first text box, or the text box with a higher confidence level can be further used as the first text box according to the confidence levels of the sixth text box and the seventh text box.
[0112] Step S202: Generate multiple particles at the position corresponding to the first text box in the second video frame of the video; the first video frame and the second video frame are adjacent.
[0113] It can be understood that when the first text box is determined through step S220, the order of the first video frame, the second video frame, and the third video frame can be the third video frame, the first video frame, the second video frame, or the second video frame, the first video frame, the third video frame.
[0114] In the embodiments of the present application, a plurality of particles are generated at corresponding positions, including but not limited to being generated within the corresponding positions, or around the corresponding positions, or being generated centered on the corners of the corresponding positions. For example, when the first text box is rectangular, a plurality of particles can be generated within the position of the rectangle corresponding to the first text box in the second video frame, or around the position of the rectangle, or centered on one of the four corner positions of the rectangle. It can be understood that the number of generated particles can be adjusted, and the particles can be used to represent the position information of the text box. For example, the particles include but are not limited to representing the coordinates of one of the vertices of the text box or the center coordinates of the text box; the size, shape, etc. of the particles can be adjusted as needed.
[0115] Step S203: Determine a plurality of second text boxes in the second video frame according to the positions of the respective particles.
[0116] Specifically, step S203 can be determined through the following steps:
[0117] Use the positions of the respective particles as the midpoints or any vertex of the text box to determine a plurality of second text boxes in the second video frame.
[0118] Specifically, take the positions of the respective particles as the midpoint of the text box or any vertex of the text box, and combine the size information of the first text box. The size information includes but is not limited to length and width, so as to determine a plurality of second text boxes in the second video frame. It should be noted that the size of the second text box is the same as the size of the first text box.
[0119] As Figure 6 shown, Figure 6 shows a schematic diagram of generating particles 1031 in the second video frame 103 through the first text box 1011 of the first video frame 101. Taking the first text box 1011 in the first video frame 101 as a rectangle as an example, a plurality of particles 1031 are generated at the position of the rectangle corresponding to the first text box 1011 in the second video frame 103. Specifically, for example, a plurality of particles 1031 can be generated around the vertex in the upper left corner of the rectangle. At this time, each particle 1031 can represent the coordinates of the vertex in the upper left corner of the rectangular text box. Combining the size information of the first text box 1011, the second text box corresponding to each particle 1031 can be determined. It should be noted that when generating the particles 1031, they can be generated according to a preset rule or randomly. The preset rule includes but is not limited to being generated within the figure formed with the upper left corner as the center. It can be understood that the particles 1031 can also be generated at the center position of the rectangle or at the other vertex positions of the rectangle. Correspondingly, the particles 1031 correspondingly represent the center coordinates of the rectangular text box or the coordinates of other vertices of the rectangular text box.
[0120] Step S204: Determine the first similarity between the first text box and each second text box, and use the second text box with the highest first similarity as the third text box;
[0121] In the embodiment of the present application, the first similarity is determined by the calculation result of the formula in step S504. It can be understood that the first similarity can also be determined by the method in step S270.
[0122] Step S205: Determine the target tracking trajectory of the video text according to the first text box and the third text box. The target tracking trajectory is used to represent the position information of the video text.
[0123] Specifically, step S205 may include step S601 or step S602.
[0124] Step S601: When the first similarity between the first text box and the third text box is greater than the third threshold, add the position information of the third text box to the target tracking trajectory.
[0125] Step S602: When the first similarity between the first text box and the third text box is less than the fourth threshold, end the trajectory tracking of the video text to obtain the target tracking trajectory.
[0126] In the embodiment of the present application, for the first similarity between the first text box and the third text box, two thresholds can be set, denoted as the third threshold and the fourth threshold. The third threshold and the fourth threshold can be set simultaneously, and the value of the third threshold should be greater than or equal to the fourth threshold. For example, taking the percentage as the measurement method of similarity, when the first similarity between the first text box and the third text box is 100%, it means that the first text box and the third text box are exactly the same. The third threshold can be set to 80%, and the fourth threshold can be set to 50%. Of course, the above values are only for convenient illustration, and the actual threshold values can be adjusted flexibly according to needs.
[0127] When the first similarity between the first text box and the third text box is greater than the third threshold, for example, the first similarity between the first text box and the third text box is 90%, it means that there is a text box in the second text box of the second video frame that is very similar to the first text box, that is, the content in the third text box is very likely to be the same as the content in the first text box. Therefore, it can be considered that these text contents exist in both the first video frame and the second video frame, so the position information of the third text box can be added to the target tracking trajectory of this text. Specifically, in the embodiment of the present application, the target tracking trajectory refers to the position information of the video text in a series of consecutive video frames, which includes two aspects. The first aspect is which video frames the video text is distributed in; the second aspect is the specific position of the video text in each video frame.
[0128] Conversely, when the first similarity between the first text box and the third text box is less than the fourth threshold, for example, the first similarity between the first text box and the third text box is 30%, it indicates that among the text boxes in the second video frame, even the text box (i.e., the third text box) that is most similar to the first text box does not have a very high degree of substantial similarity with the first text box. Therefore, it can be considered that at this time, the text content in the first text box in the first video frame no longer exists at the corresponding position in the second video frame, that is, the last frame where the video text exists is the first video frame. At this time, the trajectory tracking of the video text in the first text box ends, and it is considered that the trajectory tracking of the video text in the first text box is completed, and the target tracking trajectory can be obtained.
[0129] It should be noted that the processing objects of step S601 and step S602 in the embodiments of the present application can be understood as any pair of video frames in the target tracking trajectory processing process. For example, for the target tracking trajectory of a certain video text, it is characterized that the video text continuously exists in the 15th video frame to the 25th video frame of a video. When determining the target tracking trajectory by the video text tracking method in the embodiments of the present application, assuming that the processing starts from the 15th frame according to the video frame number, for a pair of video frames composed of the 17th frame and the 18th frame, the first text box including the video text can be determined from the 17th frame, and the third text box can be determined from the 18th frame. By comparison, it can be known that the first similarity between the third text box in the 18th frame and the first text box in the 17th frame is greater than the third threshold. Therefore, the position information of the third text box in the 18th frame can be added to the target tracking trajectory of the video text.
[0130] For a pair of video frames composed of the 25th frame and the 26th frame, the first text box including the video text can be determined from the 25th frame, and the third text box can be determined from the 26th frame. By comparison, it can be known that the first similarity between the third text box in the 25th frame and the first text box in the 26th frame is less than the fourth threshold. Therefore, at this time, it can be considered that the trajectory tracking of the video text is completed. By backward deduction from the first video frame (i.e., the 25th frame) at the completion point to the starting frame of recognition (i.e., the 15th frame), the target tracking trajectory of the video text can be obtained. Moreover, it should be supplemented that except for the starting frame of recognition, for each of the remaining frames, when determining the first text box, the third text box in the previous recognition can be used as the first text box for the next recognition. For example, for a pair of video frames composed of the 18th frame and the 19th frame, the third text box in the 18th frame determined during the previous recognition of the 17th frame and the 18th frame can be used as the first text box in the 18th frame for this recognition.
[0131] Next, in combination with specific application embodiments, the technical solutions of the present application will be described in detail. It should be understood that the types of models and model structures adopted below do not constitute a limitation to the actual application of the present application.
[0132] In the embodiments of the present application, the Yolo-v3 (You Only Look Once-v3) network can be used as the detection branch network to build the text tracking network. Specifically, the Yolo-v3 network is a classic neural network in the field of object detection. This network is a fully convolutional network, and skip connections with a large number of residual mechanisms are used in the network. Downsampling is performed on the feature maps through convolutions with a stride of 2. In terms of image feature extraction, the Yolo-v3 network adopts a partial network structure of Darknet-53 (containing 53 convolutional layers). It is worth noting that the Yolo-v3 network can detect objects at 32x downsampling, 16x downsampling, and 8x downsampling, that is, identify the positions of objects on feature maps of three scales of 52*52, 26*26, and 13*13, and generate the feature representations of the target boxes, that is, the feature maps of the target box part. For the embodiments of the present application, the area where the text is located is the part that the Yolo-v3 network needs to detect and frame, that is, the target text box. The prediction results of three feature maps will be generated for this target text box under the three-scale prediction. The prediction process of each feature map can represent a process of detecting the text box.
[0133] Refer to Figure 7 , Figure 7 shows a partial schematic diagram of the text tracking network built with the Yolo-v3 network detection branch network when processing the text tracking task. In Figure 7 , video frame 401 and video frame 402 are two adjacent video frames, which are respectively input into the text tracking network. Taking the processing flow of video frame 401 as an example, by processing video frame 401 through the first Yolo-v3 network, detection results at three scales can be obtained. For example, in the embodiments of the present application, the target text box feature map generated by detection at 8x downsampling is denoted as feature map A1, the target text box feature map generated by detection at 16x downsampling is denoted as feature map A2, and the target text box feature map generated by detection at 32x downsampling is denoted as feature map A3. Here, these three feature maps can be aligned through ROI Align (Region of Interest Align layer) to make the sizes of the feature maps the same, for example, all aligned to 14*14. Then, feature vectors are extracted from feature map A1, that is, the feature map is mapped to a vector space, and the obtained feature vector is denoted as feature vector C1. At the same time, the same processing is performed on feature map A2 and feature map A3, and the obtained feature vectors are denoted as feature vector C2 and feature vector C3 respectively.
[0134] The processing of video frame 402 is quite similar to that of video frame 401, except that it is detected by another Yolo-v3 network, which is denoted as the second Yolo-v3 network. Here, it should be noted that the same Yolo-v3 network can also be used to process video frame 402. The purpose of using another Yolo-v3 network to process video frame 402 is to achieve synchronous processing of video frame 401 and video frame 402, that is, there is no need to wait for video frame 401 to be processed before processing video frame 402, thus greatly shortening the time required for processing.
[0135] Weight sharing can be performed between the first Yolo-v3 network and the second Yolo-v3 network, that is, the network parameters in the first Yolo-v3 network and the second Yolo-v3 network can be set to be the same to reduce the interference of network parameter differences on the obtained recognition results. Similarly, after the second Yolo-v3 network detects video frame 402, detection results of three feature maps are generated. The target text box feature map generated during 8-fold downsampling is denoted as feature map B1, the target text box feature map generated during 16-fold downsampling is denoted as feature map B2, and the target text box feature map generated during 32-fold downsampling is denoted as feature map B3. Then, the feature vectors of feature map B1, feature map B2, and feature map B3 are extracted respectively to obtain feature vectors D1, D2, and D3. Here, the network used to extract feature map B1, feature map B2, and feature map B3 can have the same network structure and parameter settings as the network used to extract feature map A1, feature map A2, and feature map A3 before. The purpose is also to reduce the interference of network structure and parameter differences on the obtained recognition results.
[0136] After obtaining the feature vectors C1, C2, C3, D1, D2, and D3, similarity matching is performed on the feature vectors C1 and D1, C2 and D2, and C3 and D3, respectively. The similarity between the feature vectors C1 and D1 is denoted as similarity S1, the similarity between the feature vectors C2 and D2 is denoted as similarity S2, and the similarity between the feature vectors C3 and D3 is denoted as similarity S3. Here, since the structures and network parameters of the first Yolo-v3 network and the second Yolo-v3 network are the same, and the network structures and parameters for extracting the feature vectors C1 and D1 are also the same, the magnitude of the similarity S1 between the feature vectors C1 and D1 can effectively reflect the similarity of the text boxes in video frame 401 and video frame 402. Similarly, the similarities S2 and S3 can also effectively reflect the similarity of the text boxes in video frame 401 and video frame 402. Therefore, in the embodiments of the present application, the similarity of the text boxes in video frame 401 and video frame 402 can be comprehensively judged based on the similarities S1, S2, and S3. This similarity is denoted as X. In some embodiments, the similarity X can be determined according to the average value of the similarities S1, S2, and S3. In this way, the detection results of the neural network for the text boxes at different scales are considered in an equilibrium manner, and the obtained similarity X can reduce the negative impact caused by the inaccurate single prediction of the neural network. In some embodiments, the confidence information of the detection results generated by the Yolo-v3 network at each prediction scale can also be obtained, and based on these confidences, the reliability of the similarities S1, S2, and S3 can be determined. For example, at the prediction scale for generating the feature map A1, the confidence of the first Yolo-v3 network is 0.8; at the prediction scale for generating the feature map A2, the confidence of the first Yolo-v3 network is 0.9; at the prediction scale for generating the feature map A3, the confidence of the first Yolo-v3 network is 0.85. Since the structure and parameters of the second Yolo-v3 network and the first Yolo-v3 network in the embodiments of the present application are exactly the same, it can be considered that the confidence of the second Yolo-v3 network is the same as that of the first Yolo-v3 network. Since the confidences of the feature map A1 and the feature map B1 are 0.85, it means that the reliability of the similarity S1 can be characterized by the confidence 0.8. Similarly, the reliabilities of the similarities S2 and S3 can be characterized by the confidences 0.9 and 0.85, respectively. Therefore, the weights of the similarities S1, S2, and S3 can be determined based on the magnitudes of these three confidences. Taking the sum of the weights of the similarities S1, S2, and S3 as 1 as an example, the calculation formula for the similarity X can be expressed as:
[0137] X = 0.314 * S1 + 0.353 * S2 + 0.333 * S3
[0138] S1, S2, and S3 in the above formula represent the magnitudes of similarity S1, similarity S2, and similarity S3 respectively; X represents the magnitude of similarity X.
[0139] Of course, it can be understood that when the network structures or parameters of the two detection branch networks of the text tracking network are different, a similar method can also be used to determine the weights of each similarity, and only the mean of the corresponding confidence levels needs to be calculated. For example, when the structures or parameters of the first Yolo-v3 network and the second Yolo-v3 network are different, the mean of the confidence level of feature map A1 and the confidence level of feature map B1 can be used as the judgment basis for measuring the reliability of similarity S1.
[0140] It should be noted that in the embodiments of the present application, for two consecutive video frames, the recognized text boxes can be either the text boxes of subtitles or the text boxes of the text in the picture content. When there are multiple groups of text boxes that may contain the same text content to be recognized in two consecutive video frames, either the method in the foregoing embodiments can be used to sequentially recognize each group of text boxes, or multiple groups of text boxes can be recognized simultaneously.
[0141] Refer to Figure 8 , Figure 8 The text tracking network built with the Yolo-v3 network detection branch network, as well as one specific structure of the first tracking branch network and the second tracking branch network, are shown in . Specifically, the first tracking branch network includes ROI Align layer 501 (target region alignment layer), ROI Align layer 502, ROI Align layer 503, layer 601, layer 602, layer 603, connection layer 701, connection layer 702, connection layer 703, where layer 601, layer 602, and layer 603 each include a convolutional layer and an average pooling layer. Similarly, the second tracking branch network includes ROI Align layer 504, ROI Align layer 505, ROI Align layer 506, layer 604, layer 605, layer 606, connection layer 704, connection layer 705, connection layer 706, where layer 604, layer 605, and layer 606 each include a convolutional layer and an average pooling layer.
[0142] In an embodiment of the present application, similarly, video frame 401 and video frame 402 are two adjacent video frames, which are respectively input into the text tracking network. Through the first-scale detection, second-scale detection, and third-scale detection, feature maps A1, A2, A3, B1, B2, and B3 can also be obtained. The feature map A1 is input into the ROI Align layer 501, the output result of the ROI Align layer 501 is input into the layer 601, and then the output result of the layer 601 is input into the connection layer 701, and thus the above-mentioned feature vector C1 can be obtained; similarly, the feature map A2 is input into the ROI Align layer 502, the output result of the ROI Align layer 502 is input into the layer 602, and then the output result of the layer 602 is input into the connection layer 702, and thus the above-mentioned feature vector C2 can be obtained; the feature map A3 is input into the ROI Align layer 503, the output result of the ROI Align layer 503 is input into the layer 603, and then the output result of the layer 603 is input into the connection layer 703, and thus the above-mentioned feature vector C3 can be obtained; the feature map A4 is input into the ROI Align layer 504, the output result of the ROI Align layer 504 is input into the layer 604, and then the output result of the layer 604 is input into the connection layer 704, and thus the above-mentioned feature vector D1 can be obtained; the feature map A5 is input into the ROI Align layer 505, the output result of the ROI Align layer 505 is input into the layer 605, and then the output result of the layer 605 is input into the connection layer 705, and thus the above-mentioned feature vector D2 can be obtained; the feature map A6 is input into the ROI Align layer 606, the output result of the ROI Align layer 606 is input into the layer 606, and then the output result of the layer 606 is input into the connection layer 706, and thus the above-mentioned feature vector D3 can be obtained. Then, the feature vector C1 and the feature vector D1 are input into the connection layer 801 to obtain the similarity S1, the feature vector C2 and the feature vector D2 are input into the connection layer 802 to obtain the similarity S2, and the feature vector C3 and the feature vector D3 are input into the connection layer 803 to obtain the similarity S3.
[0143] Refer to Figure 9, in the embodiments of the present application, a video processing method is further provided. This video processing method can be applied to a terminal, or to a server, or to software in a terminal or a server, and is used to implement part of the software functions. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or as a server cluster or a distributed system composed of multiple physical servers, or as a cloud server providing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, and big data and artificial intelligence platforms; the software can be an application program for video playback, etc., but is not limited to the above forms. Figure 9 Shown in it is an optional flowchart of the video processing method provided in the embodiments of the present application. This method mainly includes steps 701 to 703:
[0144] Step 701, obtain multiple consecutive video frames of the video;
[0145] Step 702, obtain the tracking trajectories of multiple video texts in the video through the aforementioned video text tracking method;
[0146] Step 703, extract video frames according to the tracking trajectories to obtain a set of key frames of the video.
[0147] In the embodiments of the present application, a video processing method is provided. Through this method, a set of key frames of the video can be effectively extracted. Here, the set of key frames refers to a set of video frames that can reflect and describe the video content. For example, the lines of the video are very helpful for understanding the video content. Taking each piece of text presented in the subtitle as a line, video frames covering each line can be selected as the set of key frames, which is convenient for video content review or recommendation. For example, referring to Figure 10, for example, a certain video clip includes 50 consecutive video frames, and there are a total of five lines of dialogue displayed in these video frames. The first line of dialogue T1 is distributed from frame 1 to frame 15, the second line of dialogue T2 is distributed from frame 16 to frame 21, the third line of dialogue T3 is distributed from frame 22 to frame 37, the fourth line of dialogue T4 is distributed from frame 37 to frame 42, and the fifth line of dialogue is distributed from frame 43 to frame 50. Then, through the aforementioned video text tracking method, here, the line of dialogue is the tracking target of the video text, and five tracking trajectories of the five lines of dialogue can be obtained. The first tracking trajectory records the position information of the first line of dialogue T1, and this position information characterizes which video frames the first line of dialogue T1 is distributed in, that is, from the first frame to the fifteenth frame. Therefore, according to the tracking trajectory of the first line of dialogue T1, a frame can be extracted from the video frames covered by this tracking trajectory to reflect the text content of the first line of dialogue T1. Similarly, for the second line of dialogue T2 to the fifth line of dialogue T5, a frame is extracted from the video frames covered by their corresponding tracking trajectories to reflect the text content of the corresponding line of dialogue, and thus these extracted video frames are used as the key frame set. For example, for the aforementioned 50 consecutive video frames, the key frame set of this video segment can be obtained by extracting the 10th frame, the 18th frame, the 29th frame, the 41st frame, and the 44th frame. It can be understood that in the embodiments of the present application, the number of video frames and the selection of key frames are only for convenient illustration, and can be flexibly adjusted according to needs during actual implementation.
[0148] In the above video processing method, the target tracking trajectory of video text is mainly used to determine in which video frames the video text is distributed, so that a video frame can be selected from them for analyzing and auditing the video content. In some other embodiments, the target tracking trajectory of video text can also be used to determine the specific position of the video text in each video frame. For example, when it is found that there is text in a certain video that does not meet the relevant specifications and needs to be blocked, the specific position of the video text in each video frame can be quickly determined according to the target tracking trajectory of the video text, which is convenient for the staff to perform censor processing in a timely manner.
[0149] Referring to Figure 11 , an embodiment of the present application also discloses a video text tracking device, including:
[0150] The first processing module 910 is used to determine the first text box from the first video frame of the video;
[0151] The particle generation module 920 is used to generate a plurality of particles at the position corresponding to the first text box in the second video frame of the video; the first video frame and the second video frame are adjacent;
[0152] The second processing module 930 is used to determine a plurality of second text boxes in the second video frame according to the positions of the respective particles;
[0153] A similarity determination module 940 is configured to determine a first similarity between the first text box and each second text box, and use the second text box with the highest first similarity as the third text box;
[0154] A trajectory determination module 950 is configured to determine a target tracking trajectory of the video text according to the first text box and the third text box; the target tracking trajectory is used to represent the position information of the video text.
[0155] It can be understood that Figure 2 The content in the video text tracking method embodiment shown is applicable to the video text tracking device embodiment of this application. The functions specifically implemented by the video text tracking device embodiment of this application are the same as those Figure 2 shown in the video text tracking method embodiment, and the beneficial effects achieved are the same as those Figure 2 shown in the video text tracking method embodiment.
[0156] Referring to Figure 12 , an electronic device is further disclosed in an embodiment of this application, including:
[0157] At least one processor 1010;
[0158] At least one memory 1020, configured to store at least one program;
[0159] When at least one program is executed by at least one processor 1010, at least one processor 1010 is caused to implement the video text tracking method embodiment as Figure 2 shown or Figure 7 the video processing method embodiment shown.
[0160] It can be understood that, as Figure 2 shown in the video text tracking method embodiment or Figure 9 the video processing method embodiment shown, the content is applicable to the electronic device embodiment of this application. The functions specifically implemented by the electronic device embodiment of this application are the same as those in the video text tracking method embodiment as Figure 2 shown or Figure 9 the video processing method embodiment shown, and the beneficial effects achieved are the same as those in the video text tracking method embodiment as Figure 2 shown or Figure 9 the video processing method embodiment shown.
[0161] An embodiment of this application further discloses a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the video text tracking method embodiment as Figure 2 shown or Figure 9An embodiment of the video processing method shown.
[0162] It can be understood that Figure 2 The content in the embodiment of the video text tracking method shown or Figure 9 The content in the embodiment of the video processing method shown is applicable to this embodiment of the computer-readable storage medium. The functions specifically implemented in this embodiment of the computer-readable storage medium are the same as those in Figure 2 The embodiment of the video text tracking method shown or Figure 9 The embodiment of the video processing method shown, and the beneficial effects achieved are the same as those in Figure 2 The embodiment of the video text tracking method shown or Figure 9 The beneficial effects achieved by the embodiment of the video processing method shown.
[0163] This embodiment of the application also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in the above-mentioned computer-readable storage medium; Figure 12 The processor of the electronic device shown can read the computer instructions from the above-mentioned computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes Figure 2 The embodiment of the video text tracking method shown or Figure 9 The embodiment of the video processing method shown.
[0164] It can be understood that Figure 2 The content in the embodiment of the video text tracking method shown or Figure 9 The content in the embodiment of the video processing method shown is applicable to this embodiment of the computer program product or the computer program. The functions specifically implemented in this embodiment of the computer program product or the computer program are the same as those in Figure 2 The embodiment of the video text tracking method shown or Figure 9 The embodiment of the video processing method shown, and the beneficial effects achieved are the same as those in Figure 2 The embodiment of the video text tracking method shown or Figure 9 The beneficial effects achieved by the embodiment of the video processing method shown.
[0165] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. Further, the embodiments presented and described in the flowcharts of the present application are provided by way of example in order to provide a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are executed independently.
[0166] In addition, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present application. Rather, given the attributes, functions and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those skilled in the art can implement the present application as set forth in the claims without undue experimentation. It should also be understood that the particular concepts disclosed are illustrative only and are not intended to limit the scope of the present application, the scope of which is determined by the full scope of the appended claims and their equivalents.
[0167] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the relevant technology, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0168] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0169] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0170] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0171] In the above description of this specification, the description referring to terms such as "one embodiment / Example", "another embodiment / Example", or "certain embodiments / Examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0172] Although embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
[0173] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A video text tracking method, characterized in that, Including the following steps: Determine a first text box from the first video frame of the video; Generate a plurality of particles at a position corresponding to the first text box in the second video frame of the video; The first video frame and the second video frame are adjacent; Determine a plurality of second text boxes in the second video frame according to the positions of the respective particles; Determine a first similarity between the first text box and each of the second text boxes, and use the second text box with the highest first similarity as the third text box; Determine a target tracking trajectory of the video text according to the first text box and the third text box; The target tracking trajectory is used to characterize the position information of the video text; The determining the first text box from the first video frame of the video includes: Obtain a third video frame adjacent to the first video frame from the video, and the first video frame is located between the second video frame and the third video frame; Determine the text box that matches the first video frame and the third video frame as the first text box; The determining the text box that matches the first video frame and the third video frame as the first text box includes: Input the first video frame and the third video frame into a second text tracking network; the second text tracking network includes a third detection branch network and a fourth detection branch network; the third detection branch network includes a first sub-network and a second sub-network, and the fourth detection branch network includes a third sub-network and a fourth sub-network; Detect the first video frame through the first sub-network to obtain a sixth text box, and detect the first video frame through the second sub-network to obtain a seventh text box; Detect the third video frame through the third sub-network to obtain an eighth text box, and detect the third video frame through the fourth sub-network to obtain a ninth text box; Determine a third similarity between the sixth text box and the eighth text box, and determine a fourth similarity between the seventh text box and the ninth text box; Determine a fifth similarity according to the third similarity and the fourth similarity; When the fifth similarity is greater than a second threshold, determine the sixth text box or the seventh text box as the first text box.
2. The method according to claim 1, characterized in that, The determining the first text box from the first video frame of the video includes: Detect the text area in the first video frame to obtain the first text box.
3. The method according to claim 1, characterized in that, The determining the text box that matches the first video frame and the third video frame as the first text box includes: Input the first video frame and the third video frame into a first text tracking network; the first text tracking network includes a first detection branch network and a second detection branch network; Detect the first video frame through the first detection branch network to obtain a fourth text box; Detect the third video frame through the second detection branch network to obtain a fifth text box; Determine a second similarity between the fourth text box and the fifth text box; When the second similarity is greater than a first threshold, determine the fourth text box as the first text box.
4. The method according to claim 3, characterized in that, The first text tracking network further includes a first tracking branch network and a second tracking branch network; Determining the second similarity between the fourth text box and the fifth text box includes: Extracting the fourth text box through the first tracking branch network to obtain a first feature vector; Extracting the fifth text box through the second tracking branch network to obtain a second feature vector; Determining the second similarity according to the first feature vector and the second feature vector.
5. The method according to claim 1, characterized in that, Detecting a sixth text box for the first video frame through the first sub-network and detecting a seventh text box for the first video frame through the second sub-network includes: Downsampling the first video frame by a first multiple through the first sub-network to perform feature extraction on the first video frame and detecting the sixth text box; Downsampling the first video frame by a second multiple through the second sub-network to perform feature extraction on the first video frame and detecting the seventh text box.
6. The method according to claim 5, characterized in that, Detecting an eighth text box for the third video frame through the third sub-network and detecting a ninth text box for the third video frame through the fourth sub-network includes: Downsampling the third video frame by the first multiple through the third sub-network to perform feature extraction on the third video frame and detecting the eighth text box; Downsampling the third video frame by the second multiple through the fourth sub-network to perform feature extraction on the third video frame and detecting the ninth text box.
7. The method according to any one of claims 5 - 6, characterized in that, Determining the fifth similarity according to the third similarity and the fourth similarity includes: Obtaining the confidence levels of the sixth text box, the seventh text box, the eighth text box, and the ninth text box; Determining a first weight according to the average confidence level of the sixth text box and the eighth text box; Determining a second weight according to the average confidence level of the seventh text box and the ninth text box; Performing weighted summation on the third similarity and the fourth similarity according to the first weight and the second weight to obtain the fifth similarity.
8. The method according to claim 1, characterized in that, Determining a plurality of second text boxes in the second video frame according to the positions of the respective particles includes: Taking the positions of the respective particles as the midpoints or any vertices of the text boxes to determine a plurality of the second text boxes in the second video frame.
9. The method according to claim 1, characterized in that, Determining the target tracking trajectory of the video text according to the first text box and the third text box includes: When the first similarity between the first text box and the third text box is greater than a third threshold, adding the position information of the third text box to the target tracking trajectory; Or, When the first similarity between the first text box and the third text box is less than a fourth threshold, ending the trajectory tracking of the video text to obtain the target tracking trajectory.
10. A video processing method, characterized in that, Including the following steps: Obtaining a plurality of consecutive video frames of the video; Obtaining the target tracking trajectories of a plurality of video texts in the video through the video text tracking method according to any one of claims 1-9; Extracting the video frames according to the respective target tracking trajectories to obtain a key frame set of the video.
11. A video text tracking device, characterized in that, Including: A first processing module, configured to determine a first text box from a first video frame of the video; A particle generation module, configured to generate a plurality of particles at a position corresponding to the first text box in a second video frame of the video; The first video frame and the second video frame are adjacent; A second processing module, configured to determine a plurality of second text boxes in the second video frame according to the positions of the respective particles; A similarity determination module, configured to determine a first similarity between the first text box and each of the second text boxes, and use the second text box with the highest first similarity as a third text box; A trajectory determination module, configured to determine a target tracking trajectory of the video text according to the first text box and the third text box; The target tracking trajectory is used to characterize the position information of the video text; The determining the first text box from the first video frame of the video includes: Obtaining a third video frame adjacent to the first video frame from the video, where the first video frame is located between the second video frame and the third video frame; Determining the text box that matches the first video frame and the third video frame as the first text box; The determining the text box that matches the first video frame and the third video frame as the first text box includes: Inputting the first video frame and the third video frame into a second text tracking network; the second text tracking network includes a third detection branch network and a fourth detection branch network; the third detection branch network includes a first sub-network and a second sub-network, and the fourth detection branch network includes a third sub-network and a fourth sub-network; Detecting the first video frame through the first sub-network to obtain a sixth text box, and detecting the first video frame through the second sub-network to obtain a seventh text box; Detecting the third video frame through the third sub-network to obtain an eighth text box, and detecting the third video frame through the fourth sub-network to obtain a ninth text box; Determining a third similarity between the sixth text box and the eighth text box, and determining a fourth similarity between the seventh text box and the ninth text box; Determining a fifth similarity according to the third similarity and the fourth similarity; When the fifth similarity is greater than a second threshold, determining the sixth text box or the seventh text box as the first text box.
12. An electronic device, characterized in that, Includes: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-10.
13. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the method according to any one of claims 1-10 when executed by the processor.
Citation Information
Patent Citations
Video text tracking method and device
CN112101344A