Method and system for intelligently and accurately extracting video axis file based on OCR (Optical Character Recognition)

Through the intelligent and accurate extraction method of video axis files based on OCR, using preprocessing, multimodal feature fusion and deep learning technology, the accuracy and completeness problems in video axis file extraction are solved, and efficient and accurate understanding and extraction of video content is achieved.

CN120298951AActive Publication Date: 2025-07-11XIAN LINGXIANG BIRD CULTURE COMM CO LTD

Patent Information

Application Number
CN202510443080.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art has problems with poor accuracy and completeness in video axis file extraction, especially the lack of comprehensive understanding of video content and the inefficient automatic extraction.

Method used

An OCR-based method is used to generate accurate and complete video frame axis files through preprocessing, multimodal feature extraction, scene analysis, adaptive attention model and long-term and short-term memory networks.

Benefits of technology

It improves the accuracy and completeness of the extraction of video axis files, and can automatically adjust the extraction strategy according to different video types and contents, solving the problem of insufficient accuracy and completeness in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298951A_ABST
    Figure CN120298951A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of optical character recognition, and discloses an intelligent and accurate video axis file extraction method and system based on OCR, and the method comprises the steps: extracting OCR text features, image features and audio features of a target video frame; analyzing the target video frame to obtain a scene type, fusing the feature vector, and obtaining a weighted feature vector based on a pre-trained adaptive attention model and the fused feature vector; modeling the weighted feature vector to obtain hidden state sequence information; and based on the hidden state sequence information and the long short-term memory network model, extracting task information, a scene type, a pre-trained deep network method and a target video frame, and generating a target video frame axis file. Through the implementation of the method, the information and the method are comprehensively utilized, the extraction strategy can be automatically adjusted according to different video types and contents, the accurate and complete video frame axis file is generated, and the problem that the video axis file extracted by a traditional method is poor in accuracy and integrity is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optical character recognition, and particularly to an intelligent and accurate extraction method and system for video axis files based on OCR. Background Art

[0002] In the current era of digital information explosion, video, as an important information carrier, is widely used in various fields. Video axis files record the correspondence between video content and time, which are crucial for applications such as video editing, content retrieval, and video analysis. In video editing, accurate video axis files can help editors quickly locate and clip video segments; in video content retrieval, users can quickly find the required video content based on the text information in the video axis files.

[0003] However, for the text information in videos, traditional methods often rely on manual annotation, which is inefficient and error-prone. There are also some automatic extraction methods currently, but due to the lack of comprehensive understanding and analysis of video content, the accuracy and integrity of the extracted video axis files are poor. Therefore, how to efficiently, accurately, and completely extract video axis files is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the present invention provides an intelligent and accurate extraction method and system for video axis files based on OCR, a computer device, and a storage medium to solve the problem of how to efficiently, accurately, and completely extract video axis files.

[0005] In a first aspect, the present invention provides an intelligent and accurate extraction method for video axis files based on OCR, the method comprising: obtaining a video to be extracted and extraction task information, preprocessing the video to be extracted to obtain target video frames; extracting OCR text features, image features, and audio features of the target video frames; performing scene analysis on the target video frames to obtain scene types; obtaining a fusion feature vector based on the scene types, OCR text features, image features, and audio features; obtaining a weighted feature vector based on a pre-trained adaptive attention model and the fusion feature vector; modeling the weighted feature vector using a long short-term memory network model to obtain hidden state sequence information; generating a target video frame axis file based on the hidden state sequence information, the long short-term memory network model, the extraction task information, the scene types, a pre-trained deep network method, and the target video frames.

[0006] The intelligent and accurate extraction method of video axis files based on OCR provided in this embodiment first performs preprocessing operations such as decoding, sampling, and image enhancement on the video to reduce the computational amount, remove noise, and enhance text edges, laying a foundation for accurately extracting various features subsequently and improving the overall processing efficiency. Secondly, by comprehensively extracting multi-modal features of video frames, it changes the traditional way of relying on manual annotation of text information, avoids the problems of low efficiency and easy errors in manual annotation, and obtains video content information from multiple dimensions. Then, by analyzing the scene type, it provides a basis for subsequent feature fusion and extraction strategy formulation, enabling the system to perform differential processing for different scenes and improving the comprehensiveness of understanding video content. Furthermore, by fusing multi-modal features, it makes up for the defect that existing automatic extraction methods lack a comprehensive understanding of video content, comprehensively considers various modal information, enhances the expression ability of video content, and helps to improve the accuracy and integrity of extraction. Then, through a pre-trained adaptive attention model, attention weights are dynamically allocated according to the multi-modal fusion features to highlight key features, further enhancing the ability to capture key content of the video, thereby improving the accuracy of extraction. Next, a long short-term memory network is used to model the weighted feature vector sequence, effectively capturing the temporal dependence relationship of video content, enabling the system to better process the information of the video in the temporal dimension, and improving the accuracy and integrity of understanding video content. Finally, by comprehensively using various information and pre-trained deep network methods, the system can automatically adjust the extraction strategy according to different video types and content, generate accurate and complete video frame axis files, and solve the problem of poor accuracy and integrity of the video axis files extracted by traditional methods. In summary, the intelligent and accurate extraction method of video axis files based on OCR provided in this embodiment effectively solves the problem of how to efficiently, accurately, and completely extract video axis files.

[0007] In an alternative embodiment, the video to be extracted and extraction task information are obtained, and the video to be extracted is preprocessed to obtain target video frames, including: using a video decoding library to convert the video to be extracted into a video frame sequence; sampling the video frame sequence based on a preset sampling interval to obtain sampled video frames; and performing histogram equalization processing on the sampled video frames to obtain target video frames.

[0008] In an alternative embodiment, the OCR text features of the target video frames are extracted, including: using an OCR engine to recognize the characters in the target video frames to obtain text information; performing word segmentation processing, part-of-speech tagging processing, and keyword extraction processing on the text information to obtain the word frequency data and inverse document frequency data of each keyword; and constructing OCR text features based on the word frequency data and inverse document frequency data of each keyword.

[0009] In an alternative embodiment, extracting the image features of the target video frame includes: inputting the target video frame into the convolutional layer of a pre-trained convolutional neural network, and using the convolutional layer to generate a preliminary feature map based on the target video frame, where the preliminary feature map includes edge features and texture features; inputting the preliminary feature map into the pooling layer, using the pooling layer to divide the preliminary feature map into several small regions, performing max pooling processing on each small region, and taking the maximum value within each small region as the output to obtain a reduced feature map; flattening the reduced feature map into a one-dimensional vector, and using the fully connected layer to map the one-dimensional vector into a fixed-length vector to obtain the image features.

[0010] In an alternative embodiment, the pre-trained convolutional neural network includes at least three combinations of convolutional layers and pooling layers.

[0011] In an alternative embodiment, obtaining a fused feature vector based on the scene type, OCR text features, image features, and audio features includes: determining a default weight based on the correspondence between the default weight and the scene type and the scene type; performing feature fusion on the OCR text features, image features, and audio features based on the default weight to obtain the fused feature vector.

[0012] In an alternative embodiment, obtaining a weighted feature vector based on the pre-trained adaptive attention model and the fused feature vector includes: calculating an attention weight vector based on the pre-trained adaptive attention model and the fused feature vector; using the attention weight vector to perform weighted processing on the fused feature vector to obtain the weighted feature vector.

[0013] In an alternative embodiment, generating a target video frame axis file based on the hidden state sequence information, long short-term memory network model, extraction task information, scene type, pre-trained deep network method, and target video frame includes: defining environmental state information based on the hidden state sequence information; defining an extraction strategy set based on the scene type; defining a reward strategy set based on the extraction task information; using the pre-trained deep network method to obtain a target extraction strategy set based on the environmental state information, extraction strategy set, and reward strategy set; processing the target video frame based on the target extraction strategy set to generate the target video frame axis file.

[0014] In an alternative embodiment, the method further includes: extracting target OCR text features, target image features, and target audio features of a target video frame axis file; analyzing the target OCR text features, target image features, and target audio features by using a fusion analysis technique based on the Transformer architecture to obtain consistency data of the target OCR text features, target image features, and target audio features on the time axis, and complementary data of the target OCR text features, target image features, and target audio features on the time axis; obtaining collaborative score data based on the consistency data and the complementary data; and performing an optimization process on the target video frame axis file based on a reinforcement learning model, a meta-learning model, and the collaborative score data to obtain an optimized video frame axis file.

[0015] In a second aspect, the present invention provides an intelligent and accurate extraction system for a video axis file based on OCR. The system includes: a preprocessing module, configured to obtain a video to be extracted and extraction task information, and perform preprocessing on the video to be extracted to obtain target video frames; an extraction module, configured to extract OCR text features, image features, and audio features of the target video frames; an analysis module, configured to perform scene analysis on the target video frames to obtain a scene type; a fusion module, configured to obtain a fusion feature vector based on the scene type, the OCR text features, the image features, and the audio features; a calculation module, configured to obtain a weighted feature vector based on a pre-trained adaptive attention model and the fusion feature vector; a modeling module, configured to model the weighted feature vector by using a long short-term memory network model to obtain hidden state sequence information; and a generation module, configured to generate a target video frame axis file based on the hidden state sequence information, the long short-term memory network model, the extraction task information, the scene type, a pre-trained deep network method, and the target video frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 is a flowchart of an intelligent and accurate extraction method for a video axis file based on OCR according to an embodiment of the present invention;

[0018] Figure 2 is a structural block diagram of an intelligent and accurate extraction system for a video axis file based on OCR according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0020] In the current era of digital information explosion, video, as an important information carrier, is widely used in various fields. The video axis file records the correspondence between video content and time, which is crucial for applications such as video editing, content retrieval, and video analysis. In video editing, an accurate video axis file can help editors quickly locate and clip video segments; in video content retrieval, users can quickly find the required video content based on the text information in the video axis file.

[0021] However, for the text information in videos, traditional methods often rely on manual annotation, which is inefficient and error-prone. There are also some existing automatic extraction methods. Due to the lack of a comprehensive understanding and analysis of video content, the accuracy and integrity of the extracted video axis files are relatively poor. Therefore, how to efficiently, accurately, and completely extract video axis files is an urgent problem to be solved.

[0022] The intelligent and precise extraction method of video axis files based on OCR provided by this embodiment first performs preprocessing operations such as decoding, sampling, and image enhancement on the video to reduce the computational load, remove noise, and enhance text edges, laying a foundation for accurately extracting various features subsequently and improving the overall processing efficiency. Secondly, by comprehensively extracting multi-modal features of video frames, it changes the traditional way of relying on manual annotation of text information, avoids the problems of low efficiency and easy errors in manual annotation, and obtains video content information from multiple dimensions. After that, by analyzing the scene type, it provides a basis for subsequent feature fusion and extraction strategy formulation, enabling the system to perform differential processing for different scenes and improving the comprehensiveness of understanding video content. Furthermore, by fusing multi-modal features, it makes up for the defect that existing automatic extraction methods lack a comprehensive understanding of video content, comprehensively considers various modal information, enhances the expression ability of video content, and helps improve the accuracy and integrity of extraction. Then, through the pre-trained adaptive attention model, it dynamically assigns attention weights according to the multi-modal fusion features, highlights key features, and further enhances the ability to capture key content in the video, thereby improving the accuracy of extraction. Next, it models the weighted feature vector sequence through a long short-term memory network, effectively captures the temporal dependence relationship of video content, enables the system to better process the information of the video in the time dimension, and improves the accuracy and integrity of understanding video content. Finally, by comprehensively using various information and the pre-trained deep network method, the system can automatically adjust the extraction strategy according to different video types and content, generate accurate and complete video frame axis files, and solve the problem of poor accuracy and integrity of the video axis files extracted by traditional methods. In summary, the intelligent and precise extraction method of video axis files based on OCR provided by this embodiment effectively solves the problem of how to efficiently, accurately, and completely extract video axis files.

[0023] According to an embodiment of the present invention, an embodiment of an intelligent and precise extraction method of video axis files based on OCR is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0024] In this embodiment, an intelligent and precise extraction method of video axis files based on OCR is provided. Figure 1 It is a flowchart of the intelligent and precise extraction method of video axis files based on an embodiment of the present invention. As Figure 1 shown, the process includes the following steps:

[0025] Step S101, obtain the video to be extracted and extraction task information, and perform preprocessing on the video to be extracted to obtain target video frames.

[0026] The videos to be extracted include video files of various storage formats, such as the common MP4, AVI, MKV, etc. These videos come from a wide range of sources, such as surveillance videos, film and television materials, teaching videos, etc. The extraction task information includes extraction targets, format requirements and accuracy requirements. The extraction target represents the specific information to be obtained from the video, such as whether to extract all dialogue texts in the video, or only extract the speech of a specific person; whether to locate the time node of a key event, or identify the scene where a specific object appears in the video. For example, in the film and television material extraction task, it may be required to extract the protagonist's lines and corresponding timestamps. The format requirements specify the format of the final generated video axis file, such as CSV, JSON and other formats. Different formats are suitable for different application scenarios. The CSV format is often used for data table display, and the JSON format is convenient for data transmission on the network and program parsing. Accuracy requirements involve time accuracy and content accuracy. In terms of time accuracy, it is specified whether the extracted timestamp is accurate to seconds, milliseconds or finer granularity; in terms of content accuracy, it is clear that the requirements for the accuracy of text extraction and the level of detail of image feature recognition are required. For example, in surveillance video analysis, the timestamp may be required to be accurate to milliseconds to accurately capture the moment when the event occurs. In addition, there may be special extraction conditions set for specific video content. For example, when extracting a teaching video reel file, you need to skip the opening and ending parts and only extract the part containing the blackboard content; or when processing multi-language videos, you need to specify the extraction of text information in a certain language.

[0027] Specifically, the above step S101 includes:

[0028] Step a1: using the video decoding library, convert the video to be extracted into a video frame sequence.

[0029] Furthermore, videos are usually stored in specific encoding formats, such as MP4, AVI, etc. These encoding formats use compression technology to reduce the size of video files for easy storage and transmission. However, when extracting video axis files, the video needs to be decoded into a series of image frames for subsequent processing. The video decoding library is a tool to achieve this function. For example, the widely used FFmpeg is a powerful open source multimedia framework that can support the decoding of multiple video encoding formats. The specific operation process is to input the file path of the video to be extracted as a parameter into the video decoding library. The decoding library will parse the video content frame by frame according to the encoding format of the video and the frame rate of the video to generate a continuous sequence of video frames. The video frame rate indicates the number of frames displayed in a video in one second. Common frame rates include 24fps, 30fps, 60fps, etc. For example, a video with a duration of 10 seconds and a frame rate of 30fps will get 300 video frames after decoding.

[0030] Step a2: sampling the video frame sequence based on a preset sampling interval to obtain sampled video frames.

[0031] Furthermore, not all frames in the video frame sequence are of equal importance for the video axis file extraction task, and processing all frames will greatly increase the computational amount and processing time. Therefore, it is necessary to sample the video frame sequence and only select a part of the key frames for subsequent processing. The preset sampling interval is a parameter set before sampling, indicating how many frames to select as a sampling frame at intervals. For example, if the preset sampling interval is 5, then in the video frame sequence, the 1st frame, the 6th frame, the 11th frame, etc. will be selected as sampling frames in sequence. By reasonably setting the sampling interval, the number of video frames to be processed can be effectively reduced on the premise of ensuring the accuracy of the extraction result, and the processing efficiency can be improved. The setting of the sampling interval needs to be determined according to the complexity of the video content and the accuracy requirements of the extraction task. For videos with relatively slow content changes, the sampling interval can be appropriately increased; while for videos with frequent content changes, the sampling interval needs to be reduced to ensure that key information is not missed.

[0032] Step a3, perform histogram equalization processing on the sampled video frames to obtain target video frames.

[0033] Furthermore, for the sampled video frames obtained through sampling, there are differences in their image quality, such as problems like uneven brightness and low contrast, which will affect the accuracy of subsequent OCR text feature extraction and image feature extraction. Histogram equalization is a commonly used image enhancement technique for improving the contrast and brightness distribution of images. In digital images, each pixel has a corresponding gray value. For color images, they will be converted into grayscale images before processing. The principle of histogram equalization is to transform the gray histogram of the image, redistribute the gray values of the image, make the gray distribution of the image more uniform, and thus enhance the contrast of the image. The specific calculation process is as follows: First, count the frequency of each gray value in the image to obtain the gray histogram. Then, according to a certain formula, map the original gray values to a new gray value range, thereby realizing the redistribution of gray values. After the sampled video frames are processed by histogram equalization, the image quality is improved, which is more conducive to subsequent feature extraction operations. At this time, the video frames are the target video frames.

[0034] Step S102, extract the OCR text features, image features, and audio features of the target video frames.

[0035] Specifically, extracting the OCR text features of the target video frames includes:

[0036] Step b1, use the OCR engine to recognize the characters in the target video frames to obtain text information.

[0037] Furthermore, the Optical Character Recognition engine is a software tool that can convert text in an image into computer-editable text. Its core principle is based on pattern recognition and machine learning technology. When processing the target video frame, the image area in the video frame is first divided into candidate character areas, and then the image features of each candidate area are extracted, such as the outline of the character, the stroke structure, etc. Next, these features are matched with the pre-trained character models in the OCR engine. These models contain character sample features of various fonts, font sizes, and styles. By calculating the similarity between the features, the characters corresponding to the character candidate area are determined, and finally all the recognized characters are combined into text information. For example, when recognizing a video frame containing Chinese and English, the OCR engine can accurately recognize and convert the "Hello, hello!" in the image into the corresponding text string.

[0038] Step b2, performing word segmentation, part-of-speech tagging and keyword extraction on the text information to obtain word frequency data and inverse document frequency data of each keyword.

[0039] Furthermore, word segmentation refers to dividing continuous text information into independent words according to certain rules. In Chinese, word segmentation is particularly important because there is no obvious space between words. Common word segmentation methods include dictionary matching-based methods, statistical model-based methods, etc. For example, using a dictionary matching-based method to segment "I like to eat apples", we will get the words "I", "like", "eat", and "apple". Part-of-speech tagging refers to marking the part of speech of each word after segmentation, such as noun, verb, adjective, etc. Part-of-speech tagging helps to understand the grammatical structure and semantic information of the text. Commonly used part-of-speech tagging tools are based on hidden Markov models, conditional random fields, and other methods. For example, when the sentence "Apple is a kind of fruit" is tagged with part of speech, "apple" will be marked as a noun, "is" will be marked as a verb, "a kind of" will be marked as a quantifier, and "fruit" will be marked as a noun. Keyword extraction refers to extracting words that can represent the core content of the text from the text. Common methods are based on TF-IDF (Term Frequency-Inverse Document Frequency) methods, TextRank methods, etc. After keyword extraction, each keyword will be obtained. At the same time, the number of occurrences of each keyword in the current text is calculated to obtain word frequency data; the inverse document frequency is calculated, which indicates the general importance of a word in the entire document collection. The higher the inverse document frequency, the rarer the word is in the document collection, and the higher its discrimination. For example, in an article about apple planting, words such as "apple" and "planting" may be keywords, and their word frequency and inverse document frequency data are obtained by calculation.

[0040] Step b3: Based on the word frequency data and inverse document frequency data of each keyword, construct OCR text features.

[0041] Furthermore, combine the word frequency data and inverse document frequency data of each keyword to form a feature vector. For example, assume there are three keywords "apple", "fruit", and "planting", with their word frequencies being 5, 3, and 2 respectively, and their inverse document frequencies being 0.5, 0.3, and 0.8 respectively. Then a three-dimensional feature vector [5×0.5, 3×0.3, 2×0.8] = [2.5, 0.9, 1.6] can be constructed. This feature vector can quantitatively describe the text content in the target video frame from the semantic level of the text, and can be used for subsequent tasks such as multi-modal feature fusion and video content analysis, providing an important basis of text information for the accurate extraction of video axis files.

[0042] Specifically, extract the image features of the target video frame, including:

[0043] Step c1: Input the target video frame into the convolutional layer of a pre-trained convolutional neural network, and use the convolutional layer to generate a preliminary feature map based on the target video frame. The preliminary feature map contains edge features and texture features.

[0044] Furthermore, a convolutional neural network is a deep learning model specifically designed for processing data with a grid structure. The pre-trained CNN is trained on a large number of image datasets and has learned many general image feature representations. The convolutional layer is the core component of the CNN and consists of multiple convolutional kernels. A convolutional kernel is a small matrix, and common sizes are 3×3, 5×5, etc. When processing the target video frame, the convolutional kernel slides on the video frame with a certain stride. For each sliding position, the convolutional kernel performs a weighted sum of the pixel values in the corresponding area of the video frame to obtain a new pixel value. For example, for a 3×3 convolutional kernel and a video frame area containing three RGB channels, the convolutional kernel will multiply and sum the corresponding pixel values of the R, G, and B channels respectively, and finally obtain a new pixel value. This process is performed for each position of the video frame, thus generating a new feature map. Since different convolutional kernels can capture different image features, such as horizontal edges, vertical edges, textures, etc., through the parallel operation of multiple convolutional kernels, a preliminary feature map containing various edge features and texture features can be generated. For example, a convolutional layer containing 16 3×3 convolutional kernels will generate 16 different preliminary feature maps, and each feature map highlights different types of local features in the video frame.

[0045] Step c2: Input the preliminary feature map into the pooling layer, and use the pooling layer to divide the preliminary feature map into several small regions, perform max-pooling processing on each small region, and take the maximum value in each small region as the output to obtain a reduced feature map.

[0046] Furthermore, the main function of the pooling layer is to reduce the dimension of the feature map, reduce the computational amount, while retaining important image features and improving the robustness of the model. Common pooling methods include max pooling and average pooling, and max pooling is adopted here. During the max pooling operation, the preliminary feature map is divided into non-overlapping small regions, and the common region sizes are 2×2 or 3×3. For each small region, the maximum value of the pixel values within the region is taken as the output after pooling of this region. For example, for a 2×2 small region, where the four pixel values are 10, 20, 15, and 30 respectively, after max pooling, the output value of this region is 30. In this way, each small region only retains the most significant features, and at the same time, the size of the feature map will be reduced. For example, a preliminary feature map with a size of 100×100, after 2×2 max pooling, the size becomes 50×50, and the generated new feature map is the reduced feature map. The reduced feature map reduces the data volume while retaining the key features of the image, and alleviates the burden for subsequent calculations and processing.

[0047] Step c3, flatten the reduced feature map into a one-dimensional vector, and use the fully connected layer to map the one-dimensional vector into a fixed-length vector to obtain the image features.

[0048] Furthermore, after being processed by the convolutional layer and the pooling layer, the reduced feature map still maintains a certain spatial structure (such as a two-dimensional matrix form), but subsequent neural network layers (such as classifiers, etc.) usually require a one-dimensional vector as the input. Therefore, it is necessary to flatten the reduced feature map into a one-dimensional vector. For example, a reduced feature map with a size of 50×50×16 (height is 50, width is 50, and the number of channels is 16), after flattening, will obtain a one-dimensional vector with a length of 50×50×16 = 40000. The fully connected layer is a common layer in the neural network, and each neuron of it is connected to all neurons of the previous layer. Input the flattened one-dimensional vector into the fully connected layer, and the fully connected layer performs a linear transformation on the input vector through a series of weight matrices and bias vectors, and then performs a non-linear transformation through an activation function (such as the ReLU function), and finally maps the input vector into a fixed-length vector. This fixed-length vector is the extracted image feature, which contains the comprehensive feature information of the image in the video frame, and can be used for subsequent multi-modal feature fusion, video content analysis and other tasks, providing an important image information basis for the accurate extraction of the video axis file. For example, map the flattened 40000-dimensional vector through the fully connected layer into a 256-dimensional fixed-length vector, and this 256-dimensional vector represents the image feature of the target video frame.

[0049] In an alternative of some embodiments, the pre-trained convolutional neural network includes at least three groups of combinations of convolutional layers and pooling layers.

[0050] Specifically, a single convolutional layer and pooling layer can only extract simple and low-level image features, such as edges and textures. As the network depth increases, each combination of convolutional and pooling layers can extract higher-level and more abstract features. The first combination of convolutional and pooling layers extracts basic features such as edges and corners; on this basis, the second combination can identify more complex shapes or local structures; the third combination can capture features with semantic information, such as parts of an object or complete object category features. Through multiple combinations, different levels of features can be comprehensively extracted from video frames, providing rich information for subsequent analysis. As convolutional and pooling layers are stacked, the receptive field gradually increases. The receptive field refers to the mapped area of neurons in a convolutional neural network on the original image. In a shallow network, the receptive field of neurons is small, and they can only focus on local details of the image; while in a deep network, the receptive field of neurons is large, and they can obtain a larger range of image information. At least three combinations can ensure that the network is sensitive to local details and can grasp the overall image structure, which is crucial for accurately understanding the content of video frames. The pooling layer reduces the dimensionality of the feature map while retaining key information and reducing the computational amount. If the number of layers in a convolutional neural network is too small, in order to achieve sufficient feature extraction ability, it may be necessary to use larger convolutional kernels and more parameters, which will not only increase the computational burden but also easily lead to overfitting. Through multiple combinations of convolutional and pooling layers, the size of the feature map is gradually reduced, the number of parameters is reduced, the generalization ability of the model is improved, and the risk of overfitting is reduced, enabling the model to have a stable performance on different video frames.

[0051] Specifically, extracting the audio features of the target video frame includes:

[0052] Step d1, using an audio processing library to separate the audio track from the target video.

[0053] Furthermore, audio is usually stored as continuous time-series data, and its sampling rate determines the number of audio samples collected per second. Common sampling rates are 44100Hz, 48000Hz, etc. The separated audio data is segmented into audio segments of a fixed length, and each segment can be regarded as an independent analysis unit. For example, the audio is segmented into audio segments of 1 second in length for subsequent fine processing of each segment. In this step, the audio data is also normalized, adjusting the amplitude of the audio signal to a fixed range, usually [-1, 1], to eliminate the differences in amplitude between different audio signals and provide a unified input format for subsequent feature extraction.

[0054] Step d2, performing a short-time Fourier transform on each audio segment to convert the audio signal in the time domain into a frequency-domain representation.

[0055] Furthermore, STFT obtains the energy distribution of the audio signal at different times and frequencies by performing Fourier transform on the audio signal within a short time window, generating a spectrogram. For example, setting a window of 2048 sample points and a step size of 512 sample points for STFT, the resulting spectrogram can clearly show the changes in the audio in the time and frequency dimensions. To more effectively extract audio features, the spectrogram is usually subjected to Mel-frequency transformation to convert linear frequencies to Mel frequencies, because Mel frequencies are more in line with the frequency perception characteristics of the human auditory system. Under the Mel-frequency scale, Mel-Frequency Cepstral Coefficients (MFCCs) are calculated. MFCCs can well reflect features such as the timbre and formants of the audio and are commonly used feature parameters in audio analysis. In addition, other time-frequency features such as spectral centroid and spectral bandwidth can also be calculated. These features describe the frequency characteristics of the audio signal from different angles, further enriching the audio feature information.

[0056] Step d3, perform dimensionality reduction processing on the extracted audio features using principal component analysis.

[0057] Furthermore, after time-frequency feature extraction, the resulting audio features have a high dimension and may contain redundant information. To reduce the computational amount and improve the model training efficiency, principal component analysis projects the high-dimensional data into a low-dimensional space through linear transformation, removing noise and redundant information while retaining the main features of the data. For example, a high-dimensional audio feature vector containing multiple features such as 13 MFCCs, spectral centroid, and spectral bandwidth is projected into a 10-dimensional low-dimensional space through PCA. Finally, the dimensionality-reduced audio features are fused with the features of other modalities to form a comprehensive multi-modal feature vector, providing comprehensive data support for subsequent video content analysis and video axis file extraction based on multi-modal information.

[0058] Step S103, perform scene analysis on the target video frame to obtain the scene type.

[0059] Specifically, the above step S103 includes:

[0060] Step e1: Clean the extracted OCR text by removing special characters, garbled characters, and stop words to purify the text content. Utilize pre-trained language models with the Transformer architecture, such as GPT-Neo, etc., to convert the cleaned text into high-dimensional semantic vectors. In this way, rich semantic information contained in the text can be mined. For example, in the sports event scenario, keywords such as "goal", "game", "referee", etc. can be accurately captured and converted into semantic vectors, providing a basis for subsequent analysis. From the already extracted image features, screen out key information such as scene layout, object contours, and color distribution. Use a method that combines principal component analysis and local linear embedding for dimensionality reduction. PCA is used to initially reduce the feature dimension and remove redundant information; LLE further mines the local geometric structure of the data, retains the key features of the scene in the image, and enhances the distinguishability of image features in different scenes. For example, in the urban street view and natural scenery scenarios, it can effectively highlight the feature differences between buildings and natural elements. For audio features, such as MFCCs, zero-crossing rate, etc., perform normalization and standardization processing to ensure that different audio features are analyzed on the same scale. Use wavelet transform to enhance the features of key sound elements in the audio, such as the car horn sound and siren sound in the traffic scene, and the instrument performance sound in the concert scene, etc., making the audio features more distinguishable.

[0061] Step e2: Select a multi-layer perceptron (MLP) as the text analysis model. Train it with text data from large-scale and diverse scenarios, which include but are not limited to traffic, meetings, family gatherings, etc. The model learns the co-occurrence patterns and semantic associations of text vocabulary in different scenarios. For example, in the meeting scenario, words such as "agenda", "report", "discussion", etc. frequently co-occur, and based on this, it judges the scene type to which the text belongs. Adopt a convolutional neural network improved based on the DenseNet architecture. Input the preprocessed image features into the model, and through densely connected convolutional layers and pooling layers, fully extract the scene features in the image. During the training process, use a large number of image data sets containing various scenarios, allowing the model to learn the image features of different scenarios, such as roads and vehicles in the traffic scene, and human interactions and indoor layouts in the family gathering scene, etc., to achieve accurate prediction of the scene type. Use a model that combines a recurrent neural network and an attention mechanism, such as a model that combines long short-term memory network and attention mechanism, and train it with audio samples from different scenarios, allowing the model to learn the change patterns of audio features in the time series, such as the differences between natural sound feature sequences such as rain sounds and bird calls and the feature sequence of factory machine roars, so as to judge the scene type to which the audio belongs.

[0062] Step e3: Input the preprocessed OCR text features, image features, and audio features into their respective trained models to obtain the predicted probability distributions of each model for different scene types. For example, the text analysis model outputs a traffic scene probability of 0.6, a meeting scene probability of 0.2, and other scene probabilities of 0.2; the image analysis model outputs a traffic scene probability of 0.5, a meeting scene probability of 0.3, and other scene probabilities of 0.2; the audio analysis model outputs a traffic scene probability of 0.4, a meeting scene probability of 0.3, and other scene probabilities of 0.3. The dynamic weight fusion method is used to comprehensively judge the prediction results of the three models. The weights are dynamically adjusted according to the reliability and importance of each modality feature in different scenarios. For example, in the traffic scene, image and audio features may be more important, with the text analysis weight set to 0.3, the image analysis weight set to 0.4, and the audio analysis weight set to 0.3; while in the meeting scene, text features may be more critical, and the weights are adjusted accordingly to text analysis weight 0.4, image analysis weight 0.3, and audio analysis weight 0.3. Calculate the weighted average probability, and take the scene type with the highest probability as the final result. Finally, record the scene type result to provide a basis for formulating subsequent video axis file extraction strategies, improving the accuracy and pertinence of extraction.

[0063] Step S4, obtain the fused feature vector based on the scene type, OCR text features, image features, and audio features.

[0064] Specifically, the above Step S4 includes:

[0065] Step f1, determine the default weight based on the correspondence between the default weight and the scene type, and the scene type.

[0066] Furthermore, in different scene types, the importance of OCR text features, image features, and audio features for describing the scene and extracting key information varies. For example, in a speech scene, OCR text features (such as the speaker's lines) may be the most crucial for understanding the video content; while in a natural scenery video, image features (such as landscapes, the forms of animals and plants) may be dominant; in a music performance scene, audio features (such as melody, rhythm) are more important. By establishing the correspondence between default weights and scene types, it is possible to allocate appropriate weights to each modality feature according to the scene type of the video, highlight key information, and improve the accuracy of subsequent analysis and processing. A weight mapping table is established in advance, which records the correspondence between various common scene types and the default weights of each modality. For example, for the "speech scene", the weight of OCR text features is set to 0.5, the weight of image features is 0.3, and the weight of audio features is 0.2; for the "sports event scene", the weight of OCR text features is 0.3, the weight of image features is 0.5, and the weight of audio features is 0.2. After determining the scene type of the current target video frame through step S3, the corresponding default weights can be quickly found in this weight mapping table. For example, if the current scene type is determined to be the "speech scene", the OCR text feature weight of 0.5, the image feature weight of 0.3, and the audio feature weight of 0.2 are obtained from the table. The method of determining default weights based on scene types enables the system to perform differential processing for different scenes and effectively utilize the advantages of each modality feature. It avoids using fixed weights in all scenes, resulting in some key information being ignored or some redundant information being overemphasized, thereby enhancing the adaptability and accuracy of the entire video axis file extraction process for different scenes.

[0067] Step f2, perform feature fusion on the OCR text features, image features, and audio features based on the default weights to obtain a fused feature vector.

[0068] Furthermore, feature fusion combines features from different modalities together according to certain rules to form a comprehensive feature vector. Here, weighted summation is used for fusion. The principle is to weight the feature vectors of each modality according to the default weights of each modality feature, and then add the weighted vectors to obtain the fused feature vector. This method can comprehensively consider the information of different modality features, give full play to the complementarity of each modality, and make the fused feature vector more comprehensively and accurately describe the content of the video frame. The fused feature vector synthesizes multi-modal information and can provide a more rich and comprehensive description of the video frame content compared to the feature vector of a single modality. This is of great significance for subsequent video content analysis, target extraction, and other tasks based on this feature vector, which can significantly improve the accuracy and reliability of the tasks and lay a solid foundation for generating high-quality video axis files.

[0069] Step S5: Obtain a weighted feature vector based on the pre-trained adaptive attention model and the fused feature vector.

[0070] Specifically, the pre-trained adaptive attention model is trained using the backpropagation method, including but not limited to news reports, film and television programs, teaching videos, sports event videos, etc. Fused feature vectors are extracted from these videos as training data. Each fused feature vector corresponds to a corresponding video frame, and at the same time, labels such as the scene type and key information position to which the video frame belongs are marked. For example, in a news report video, the time stamps and text content corresponding to the host's lines are marked; in a sports event video, the times and relevant descriptions of key events such as goals and fouls are marked. These labels will serve as the supervision information for model training to help the model learn how to accurately allocate attention weights according to the fused feature vectors.

[0071] The above step S5 includes:

[0072] Step g1: Calculate an attention weight vector based on the pre-trained adaptive attention model and the fused feature vector.

[0073] Furthermore, the pre-trained adaptive attention model is based on the attention mechanism, which draws on the characteristic of humans automatically focusing on key content when processing information. In multi-modal information processing, such as in the video axis file extraction task, the fused feature vector contains features from multiple aspects such as OCR text, images, and audio. These features have complex dimensions and different degrees of importance. The adaptive attention model can accurately calculate the importance of each feature dimension through a series of complex operations, and then generate an attention weight vector. Its core implementation method is to perform a linear transformation on the fused feature vector by means of multiple learnable weight matrices. First, the fused feature vector is multiplied by the first weight matrix W1 for matrix multiplication to map the fused feature vector from the original space to a new intermediate feature space. This process can be understood as a preliminary screening and recombination of the features. Then, the ReLU activation function is applied to filter out negative eigenvalue features and only retain and enhance valuable feature signals to obtain a more discriminative intermediate feature representation M2. Then, M2 is multiplied by the second weight matrix W2 again for matrix multiplication to further explore the potential relationships between features and map the intermediate features to the attention weight space. Finally, the result is normalized to the 0-1 interval through the Softmax activation function so that the sum of the weights of all dimensions is 1, and thus the attention weights corresponding to each feature dimension are obtained. These weights together constitute the attention weight vector. The role of the Softmax function is like a "referee", presenting the importance of features in different dimensions in the form of a probability distribution, enabling the model to clearly distinguish which feature dimensions are more critical for the analysis of the current video frame. Suppose the fused feature vector F is a d-dimensional vector, the weight matrix W1 in the adaptive attention model has a dimension of d×k, and the weight matrix W2 has a dimension of k×1. Here, k is an intermediate dimension determined according to the model design, which plays a bridging role connecting the original feature space and the attention weight space. During specific operations, first multiply the fused feature vector F by W1, that is, M1 = F×W1, to obtain a k-dimensional intermediate vector M1. Apply the ReLU function to M1, M2 = ReLU(M1). The ReLU function will set the negative values in M1 to 0 and keep the positive values unchanged to enhance the effective features. Then, multiply M2 by W2, M3 = M2×W2, to obtain a one-dimensional result M3. Finally, apply the Softmax function to M3, α = Softmax(M3), to obtain an attention weight vector α with a dimension of d, where α iRepresents the attention weight of the i-th dimension in the fused feature vector. For example, when d = 5 and k = 3, F = [1, 2, 3, 4, 5], W1 is a 5×3 matrix. After matrix multiplication, M1 is obtained, and then after passing through the ReLU function to get M2, multiplying with W2 and then passing through the Softmax function to obtain the attention weight vector α = [0.1, 0.2, 0.3, 0.2, 0.2]. In the complex multi-modal information processing process, not all feature dimensions are equally important for video axis file extraction. By calculating the attention weight vector, the adaptive attention model acts like an intelligent "filter" that can automatically identify the feature dimensions in the fused feature vector that are more crucial for the current task, highlighting the key information. For example, in a speech video, it can accurately identify the text feature dimensions closely related to the speech content, the image feature dimensions of the person's facial expressions, and the audio feature dimensions of the speaker's intonation, while suppressing those redundant or secondary information, such as unimportant decorative elements in the video background. This not only helps improve the accuracy of the subsequent model in analyzing video content but also greatly reduces the waste of computing resources, enabling the entire video axis file extraction system to efficiently focus on the most valuable parts when processing complex multi-modal information and enhancing the system performance.

[0074] Step g2, use the attention weight vector to perform weighted processing on the fused feature vector to obtain a weighted feature vector.

[0075] Furthermore, the principle of weighted processing is based on the importance of each feature dimension reflected by the attention weight vector. Multiplying the attention weight vector with the corresponding elements of the fused feature vector essentially makes a weighted adjustment to each dimension of the fused feature vector according to the magnitude of the attention weight. The larger the attention weight, the more important the feature represented by that dimension is for describing the key information of the video frame, and the greater the proportion it occupies in the weighted feature vector; conversely, the smaller the weight, the smaller its influence in the weighted feature vector. This way can re-encode the fused feature vector according to the importance calculated by the adaptive attention model, enabling the weighted feature vector to more accurately reflect the feature representation of the key information in the video frame. For example, in a sports event video, if the attention weight vector indicates that the image feature dimension of the athlete's actions has a higher weight, then in the weighted processing, this part of the feature dimension will be enhanced in the weighted feature vector, more prominently reflecting the key action information of the event.

[0076] Given the attention weight vector α = [α1, α2,..., α d ;

[0077] The fused feature vector F = [F1, F2,..., F d ;

[0078] The calculation method of the weighted feature vector G is as follows: G i = α i × F i , where i = 1, 2,..., d.

[0079] For example, when the attention weight vector α = [0.2, 0.3, 0.5] and the fused feature vector F = [10, 20, 30], the weighted feature vector G = [0.2 × 10, 0.3 × 20, 0.5 × 30] = [2, 6, 15]. By multiplying the corresponding elements in this way, the attention weights are incorporated into the fused feature vector to achieve weighted adjustment of the features. The weighted feature vector obtained through the weighted processing is an optimization and refinement of the fused feature vector. It highlights the feature dimensions determined to be important by the adaptive attention model, enabling the subsequent long short-term memory network to more effectively capture the key information in the video content when processing this feature vector. The long short-term memory network can utilize these weighted features to better understand the changes in the video over time series and provide a more representative feature representation for generating accurate video axis files. In practical applications, the weighted feature vector can help the system more accurately locate the key time points, key events, etc. in the video, further improving the accuracy and reliability of video axis file extraction and providing high-quality data support for downstream tasks such as video editing and content retrieval.

[0080] Step S6: Use the long short-term memory network model to model the weighted feature vector to obtain hidden state sequence information.

[0081] Specifically, the above step S6 includes:

[0082] Step h1: Initialize the long short-term memory network.

[0083] Furthermore, the long short-term memory network (LSTM) is a special type of recurrent neural network that is good at processing time series data and can effectively solve the problems of vanishing gradients and exploding gradients in traditional RNNs. Before using LSTM to model the weighted feature vector, it needs to be initialized. The LSTM cell consists of an input gate, a forget gate, an output gate, and a memory cell. The input gate controls the input of new information, the forget gate determines whether to retain or discard the old information in the memory cell, and the output gate determines the output information. Multiple LSTM cells are connected in sequence to form an LSTM network. For example, to construct a network with 3 layers of LSTM, the number of LSTM cells in each layer can be set according to the task requirements and data characteristics. For example, 64 cells are set in the first layer, 32 cells in the second layer, and 16 cells in the third layer. The weight matrices in the LSTM network include input weights, forget weights, output weights, and memory cell weights, etc., which need to be initialized. Common initialization methods include random initialization, Xavier initialization, etc. Taking Xavier initialization as an example, the initial value of the weight is determined according to the dimensions of the input and output, which can make the network converge better in the initial stage of training. For example, for the input weight matrix, according to the dimension of the weighted feature vector and the number of LSTM cells, the initial weight value is calculated and set according to the Xavier initialization formula.

[0084] Step h2: Input the weighted feature vector.

[0085] Furthermore, the weighted feature vectors obtained by processing through the adaptive attention model are sequentially input into the initialized LSTM network in the order of time series. Suppose there are T time steps in the video frame sequence, and each time step corresponds to a weighted feature vector. For example, in a video with a duration of 10 seconds and a frame rate of 30 fps, there are 300 video frames in total. After feature extraction and fusion processing, 300 weighted feature vectors are obtained, and each time step is 1 / 30 second. At the first time step t = 1, the first weighted feature vector G1 is input into the first layer of the LSTM network. The LSTM network will update the memory cell and the hidden state through the collaborative operation of the input gate, the forget gate, the output gate, and the memory cell according to the input feature vector and the hidden state of the previous time step. For example, the input gate calculates the weight of the input signal based on G1 and the hidden state of the previous time step, and determines which new information can enter the memory cell; the forget gate calculates the retention weight of the old information and determines which old information in the memory cell needs to be retained; the output gate calculates the output hidden state based on the updated memory cell and the current input.

[0086] Step h3: LSTM network operation.

[0087] Furthermore, at each time step, the LSTM network performs complex operations on the input weighted feature vectors to update the memory cells and hidden states. Input gate operation: The input gate calculates the weights of the input signals through a Sigmoid activation function. The Sigmoid function maps the output value to between 0 and 1. The closer the value is to 1, the more important the new information of the input is and the easier it is to enter the memory cell. Forget gate operation: The forget gate also calculates the retention weights of the old information through the Sigmoid function. The closer the forget gate value is to 1, the more old information in the memory cell is retained. Memory cell update: Update the memory cell according to the calculation results of the input gate and the forget gate. In this way, the memory cell can not only retain important old information but also receive new information. Output gate operation: The output gate calculates the output hidden state based on the updated memory cell and the current input. The hidden state contains information about the current time step and previous time steps and is the output result of the LSTM network.

[0088] Step h4: Obtain the hidden state sequence information.

[0089] Furthermore, as the time steps progress, the LSTM network processes the weighted feature vectors of each time step in sequence, and finally obtains the hidden state sequence information of the entire time series. After processing the weighted feature vectors of all T time steps, a hidden state sequence of length T will be obtained. Each hidden state in this sequence contains video frame information from the first time step to the current time step, and through the operations of the LSTM network, it effectively captures the long-term dependencies in the time series. For example, when analyzing a continuous sports event video, the hidden state sequence can record information such as the action changes of athletes at different time points and the development of the game situation. The hidden state sequence information is crucial for generating the subsequent video axis file. It can be used as the input of a policy generation model based on reinforcement learning to help the model understand the changes in video content in the time dimension, thereby accurately determining the time points of key events and the occurrence positions of key information in the video, providing a key basis for generating an accurate and complete video axis file.

[0090] Step S7, generate the target video frame axis file based on the hidden state sequence information, long short-term memory network model, extraction task information, scene type, pre-trained deep network method, and target video frame.

[0091] Specifically, the above step S7 includes:

[0092] Step i1, define the environmental state information based on the hidden state sequence information.

[0093] Furthermore, the hidden state sequence information is obtained by the long short-term memory network modeling the weighted feature vectors, which contains the key information of video frames in the time series. The environmental state information is a concept in reinforcement learning that describes the environmental state in which the method model for generating the video axis file is located. Defining the environmental state information based on the hidden state sequence information is to transform the hidden state output by the LSTM into an environmental description meaningful for the current video analysis task. For example, the hidden state may contain information such as the actions of people in the video, the language content, and the scene changes. Integrating these information into the environmental state, such as the key scenes of the current video, the activities of the main characters, etc., so that subsequent methods can make decisions based on the environmental state. First, feature extraction and screening are performed on the hidden state sequence. Dimensionality reduction techniques such as principal component analysis can be used to extract the main components from the high-dimensional hidden state sequence and remove redundant information. Then, according to the characteristics of the video analysis task, semantic mapping is performed on the feature vectors. For example, if specific action or scene keywords appear in the video, they are mapped to corresponding environmental state labels, such as "intense competition scene", "speech scene", etc. Finally, the processed information is combined into environmental state information for subsequent policy formulation. Accurate environmental state information enables the method to understand the key features and states of the current video, providing a basis for subsequent definition of the extraction policy and the reward policy. It is like providing a "map" for the method, allowing the method to clearly know its location and the situation it faces, so as to formulate the generation policy of the video axis file targeted.

[0094] Step i2, define the extraction policy set based on the scene type.

[0095] Furthermore, different scene types have different characteristics and rules. Defining an extraction strategy set based on scene types is to utilize the association between scene types and extraction strategies to formulate a series of possible extraction strategies for each scene type. For example, in a meeting scene, the extraction strategy may focus on extracting the speech content of the speaker and the text information on the PPT; in a sports event scene, the extraction strategy may focus on the key actions of the athletes and the game scores. By pre-analyzing the characteristics of different scene types, a mapping relationship between scene types and extraction strategies is established. First, a scene type library is established, which includes various common scene types, such as meetings, sports events, movies, teaching, etc. Then, for each scene type, the characteristics and key information of its video content are analyzed. For example, for the teaching scene, the key information may include the knowledge points explained by the teacher, the blackboard writing content, etc. According to these characteristics, corresponding extraction strategies are formulated. The extraction strategies include the selection of feature extraction methods. For example, for the blackboard writing content, specific OCR recognition parameters are selected, and the determination method of time stamps, such as dividing time segments according to the pauses in the teacher's explanation. Finally, the extraction strategies formulated for each scene type are summarized into an extraction strategy set. The extraction strategy set provides a variety of feasible operation plans for the generation of video axis files. Selecting appropriate extraction strategies according to different scene types can improve the accuracy and efficiency of extraction. It is like a "strategy toolbox", and methods can select the most suitable tool from the toolbox according to the scene type of the current video to complete the task of generating the video axis file.

[0096] Step i3, define a reward strategy set based on the extraction task information.

[0097] Furthermore, the reward strategy set is a mechanism in reinforcement learning used to motivate the agent to take optimal actions. According to the goals and requirements of the extraction task, a series of reward rules are formulated. For example, if the extraction task is to accurately extract the key person conversations in a video, then the reward strategy can be to give a positive reward when the method accurately extracts the key person conversations, and a negative reward when the extraction is incorrect or omitted. Through the reward strategy, the method is guided to develop in the direction of meeting the requirements of the extraction task. First, clarify the specific goals and requirements of the extraction task, such as extraction accuracy, integrity, time precision, etc. Then, according to these goals and requirements, formulate corresponding reward rules. For example, for extraction accuracy, it can be set that when the similarity between the extracted text content and the actual video content reaches 90%, a positive reward value of 10 is given; when the similarity is below a certain threshold of 70%, a negative reward value of -5 is given. For time precision, if the timestamp error of the extracted key event and the actual timestamp is within the range of ±1 second, a reward is given, and a penalty is given if it exceeds the range. Finally, these reward rules are summarized into a reward strategy set. The reward strategy set provides a clear goal orientation for the method, enabling the method to continuously adjust its behavior according to the reward feedback during the process of generating the video axis file, improving the quality of extraction, and finally generating a video axis file that meets the requirements of the extraction task.

[0098] Step i4, using the pre-trained deep network method, obtain the target extraction strategy set based on the environmental state information, the extraction strategy set, and the reward strategy set.

[0099] Furthermore, the pre-trained deep network method is used to learn the optimal policy in reinforcement learning. By continuously interacting with the environment, in this task, the environmental state information is equivalent to the environment, the extraction policy set is equivalent to the action space of the agent, and the reward policy set is equivalent to the reward feedback. It learns which extraction policy to adopt under different environmental states to obtain the maximum reward. The deep network method utilizes the powerful fitting ability of the neural network to learn and analyze the environmental state information, the extraction policy set, and the reward policy set, and finds the optimal extraction policy combination. First, the environmental state information, the extraction policy set, and the reward policy set are used as inputs and input into the pre-trained deep network method. The neural network part in the deep network method extracts and processes the input information. It extracts the features of the environmental state information through multiple convolutional layers and fully connected layers, and encodes the extraction policy set into a processable vector form. Then, according to the output of the neural network, the value of each extraction policy in the current environmental state is calculated, that is, the reward that may be obtained by adopting this policy. Through continuous iterative training, the deep network method gradually learns the optimal extraction policies under different environmental states and forms the target extraction policy set. The target extraction policy set is the optimal policy combination learned by the deep network method based on the environmental state, the extraction policy, and the reward policy. It comprehensively considers various information of the video and the requirements of the extraction task, and can guide the method to efficiently and accurately generate the video axis file.

[0100] Step i5, process the target video frames based on the target extraction policy set to generate the target video frame axis file.

[0101] Furthermore, the set of target extraction strategies includes a series of optimal extraction strategies for the target video frames. Processing the target video frames based on these strategies means extracting key information from the target video frames according to the methods and steps specified in the strategies and organizing it in the format of a video axis file. For example, according to the timestamp determination method specified in the strategy, the time points of key events are extracted from the video frame sequence; according to the feature extraction method specified in the strategy, key features such as text, images, and audio in the video are extracted, and this information is organized and stored according to the format requirements of the video axis file. First, according to the timestamp determination strategy in the set of target extraction strategies, the target video frame sequence is divided into time points to determine the start and end times of each key event. For example, in a sports event video, the timestamps of key events such as goals and fouls are determined according to the strategy. Then, corresponding text, image, and audio features are extracted according to the feature extraction method in the strategy. For example, specific OCR methods are used to extract text information in the video, and the image features of key scenes are extracted according to the image feature extraction strategy. Finally, the extracted information such as timestamps, text, images, and audio is organized and stored in the format of a video axis file to generate the target video frame axis file. This step transforms the previous analysis, strategy formulation, and learning results into an actual video axis file, which is the ultimate goal and key link in the entire video axis file generation process. By processing the target video frames according to the set of target extraction strategies, high-quality video axis files that meet the requirements of the extraction task can be generated, providing strong support for subsequent video editing, content retrieval, and other applications.

[0102] In an alternative embodiment of some embodiments, the above method further includes:

[0103] Step j1, extracting the target OCR text feature, target image feature, and target audio feature of the target video frame axis file.

[0104] Furthermore, the target video frame axis file has integrated the key information in the video. Extracting its OCR text features, image features, and audio features again is to conduct in-depth analysis of the file content from a multi-modal perspective. For OCR text features, by re-identifying the text content in the axis file and using techniques such as lexical and syntactic analysis, the semantic information and potential logical relationships in the text are mined. Image feature extraction is targeted at the key images associated in the axis file. Based on visual elements such as the color, texture, and shape of the images, computer vision methods are used to obtain feature vectors that can represent the image content. Audio feature extraction is from the corresponding audio part of the axis file. By analyzing the time domain and frequency domain of the audio signal, such as calculating Mel Frequency Cepstral Coefficients (MFCCs), the key features of the audio are obtained. For OCR text features, a professional OCR recognition tool is used to scan and recognize the text areas in the axis file. The recognized text is cleaned to remove noise and invalid characters, and then natural language processing tools are used for word segmentation, part-of-speech tagging, and keyword extraction to form OCR text features. For image features, the image data in the axis file is read out and input into a pre-trained convolutional neural network model. After multiple convolutional and pooling operations, the high-level features of the image are extracted, and finally, the features are mapped to vectors of a fixed length through a fully connected layer. In terms of audio feature extraction, with the help of an audio processing library, the audio signal is separated from the axis file. After frame division and windowing processing of the audio signal, the audio feature parameters are calculated to form an audio feature vector. Extracting these multi-modal features is the basis for subsequent fusion analysis and optimization processing. They provide comprehensive data support for in-depth understanding of the content of the video frame axis file and help discover possible information inconsistencies or incompleteness in the file.

[0105] Step j2, using the fusion analysis technology based on the Transformer architecture to analyze the target OCR text features, target image features, and target audio features, to obtain the consistency data of the target OCR text features, target image features, and target audio features on the time axis, and the complementary data of the target OCR text features, target image features, and target audio features on the time axis.

[0106] Furthermore, the Transformer architecture, based on the self-attention mechanism, can effectively capture the long-range dependencies between different modal features and achieve in-depth fusion analysis of multimodal information. When processing the target OCR text features, target image features, and target audio features, Transformer takes the feature vectors of different modalities as inputs and calculates the degree of association of different modal features on the time axis through the self-attention mechanism. For consistency analysis, it is judged whether the information expressed by different modalities is consistent at the same time point or time period; for complementary analysis, it is determined whether different modalities can complement each other on the time axis to provide a more comprehensive description of the video content. The extracted target OCR text features, target image features, and target audio features are encoded and converted into a format suitable for the input of Transformer. For example, the text features are converted into vector representations through a word embedding layer, and the image features and audio features are normalized and then concatenated with the text feature vectors. The concatenated feature vectors are input into the Transformer model, and the multi-head attention mechanism in the model calculates the relationships between different modal features in the time dimension. Through multiple iterative calculations, the association matrix of different modal features on the time axis is obtained. Based on the association matrix, consistency data is analyzed and calculated, such as calculating the similarity of different modal feature vectors at the same time point; by analyzing the distribution differences of different modal features on the time axis, complementary data is obtained to determine which time periods different modalities can provide unique information. The consistency data and complementary data obtained through this fusion analysis technology can evaluate the fusion quality of multimodal information in the video frame axis file, discover problems of inconsistent information or insufficient complementarity, and provide directions for subsequent optimization.

[0107] Step j3, obtaining collaborative score data based on the consistency data and the complementary data.

[0108] Furthermore, the collaborative score data is a quantitative evaluation of the effect of the collaborative expression of the target OCR text features, target image features, and target audio features on the time axis for video content. Based on the consistency data and complementarity data, a comprehensive score is calculated through a certain method to reflect the degree of collaboration of multimodal information in the video frame axis file. For example, the weighted summation method can be used to combine the consistency score and the complementarity score according to certain weights to obtain the final collaborative score. Operation process: First, according to the actual needs and experience, weights are assigned to the consistency data and complementarity data. For example, the consistency weight is set to 0.6, and the complementarity weight is set to 0.4. Then, the consistency data and complementarity data are normalized to make them in the same numerical range for easy calculation. For example, both the consistency score and the complementarity score are normalized to the interval [0, 1]. Finally, weighted summation is performed according to the set weights to calculate the collaborative score data, such as: collaborative score = consistency score × 0.6 + complementarity score × 0.4. The collaborative score data provides an intuitive quantitative index for evaluating the quality of the video frame axis file, can quickly judge the fusion effect of multimodal information in the file, and provides a quantitative basis for subsequent optimization processing.

[0109] Step j4, optimize the target video frame axis file based on the reinforcement learning model, meta-learning model, and collaborative score data to obtain an optimized video frame axis file.

[0110] Furthermore, the reinforcement learning model interacts with the target video frame axis file and related data, and learns the optimal optimization strategy based on the collaborative score data. The meta-learning model is "learning how to learn". It uses past learning experiences to quickly adjust the parameters and learning methods of the reinforcement learning model, enabling the reinforcement learning model to more efficiently optimize the current video frame axis file. Based on the collaborative score data, the reinforcement learning model continuously tries different optimization operations, such as adjusting the keywords for text extraction, correcting the parameters for image feature extraction, etc. It judges the effect of the optimization operation according to the change in the collaborative score, and gradually finds the optimal optimization plan, thereby obtaining the optimized video frame axis file. Initialize the reinforcement learning model and the meta-learning model, and set the action space and state space of the reinforcement learning model. The action space represents various possible optimization operations, and the state space represents the features of the target video frame axis file and the collaborative score data. Input the relevant information of the target video frame axis file and the collaborative score data into the reinforcement learning model. The reinforcement learning model selects an optimization action according to the current state, such as adjusting the threshold for OCR text feature extraction. After executing the optimization action, recalculate the collaborative score data of the target video frame axis file, and feedback the new state and reward, that is, the change value of the collaborative score, to the reinforcement learning model. The reinforcement learning model updates its policy according to this information. The meta-learning model adjusts parameters such as the learning rate and exploration rate of the reinforcement learning model according to the learning situation of the reinforcement learning model in multiple tasks, accelerating the convergence speed of the reinforcement learning model. After multiple iterations of optimization, when the reinforcement learning model converges to a certain extent, the optimized target video frame axis file obtained is the optimized video frame axis file. Through the collaborative action of the reinforcement learning model and the meta-learning model, it is possible to optimize the target video frame axis file targeted according to the collaborative score data of multi-modal information, improve the quality and accuracy of the file, and make it more meet the needs of practical applications.

[0111] The intelligent and accurate extraction method of video axis files based on OCR provided in this embodiment first performs preprocessing operations such as decoding, sampling, and image enhancement on the video to reduce the computational load, remove noise, and enhance the text edges, laying a foundation for accurately extracting various features subsequently and improving the overall processing efficiency. Secondly, by comprehensively extracting the multi-modal features of video frames, the traditional way of relying on manual annotation of text information is changed, avoiding the problems of low efficiency and easy errors in manual annotation, and obtaining video content information from multiple dimensions. After that, by analyzing the scene type, it provides a basis for formulating subsequent feature fusion and extraction strategies, enabling the system to perform differential processing for different scenes and improving the comprehensiveness of understanding video content. Furthermore, by fusing multi-modal features, the defect that existing automatic extraction methods lack a comprehensive understanding of video content is made up for, considering various modal information comprehensively, enhancing the expression ability of video content, and helping to improve the accuracy and integrity of extraction. Then, through the pre-trained adaptive attention model, attention weights are dynamically assigned according to the multi-modal fusion features to highlight key features, further enhancing the ability to capture key content in the video, thereby improving the accuracy of extraction. Next, the weighted feature vector sequence is modeled through a long short-term memory network to effectively capture the temporal dependence relationship of video content, enabling the system to better process the information of the video in the temporal dimension and improving the accuracy and integrity of understanding video content. Finally, by comprehensively using various information and the pre-trained deep network method, the system can automatically adjust the extraction strategy according to different video types and contents, generate accurate and complete video frame axis files, and solve the problem of poor accuracy and integrity of the video axis files extracted by traditional methods. In summary, the intelligent and accurate extraction method of video axis files based on OCR provided in this embodiment effectively solves the problem of how to efficiently, accurately, and completely extract video axis files.

[0112] The above is the embodiment of the intelligent and accurate extraction system of video axis files based on OCR provided by this application. Other embodiments of the intelligent and accurate extraction of video axis files based on OCR provided by this application will be described below. See the following for details.

[0113] In this embodiment, an intelligent and accurate extraction system of video axis files based on OCR is also provided. This system is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0114] This embodiment provides an intelligent and accurate extraction system of video axis files based on OCR, as Figure 2 shown, including:

[0115] The preprocessing module 201 is used to obtain the video to be extracted and extraction task information, preprocess the video to be extracted, and obtain target video frames.

[0116] The extraction module 202 is used to extract the OCR text features, image features, and audio features of the target video frames.

[0117] The analysis module 203 is used to perform scene analysis on the target video frames to obtain the scene type.

[0118] The fusion module 204 is used to obtain a fused feature vector based on the scene type, OCR text features, image features, and audio features.

[0119] The calculation module 205 is used to obtain a weighted feature vector based on a pre-trained adaptive attention model and the fused feature vector.

[0120] The modeling module 206 is used to model the weighted feature vector using a long short-term memory network model to obtain hidden state sequence information.

[0121] The generation module 207 is used to generate a target video frame axis file based on the hidden state sequence information, long short-term memory network model, extraction task information, scene type, pre-trained deep network method, and target video frames.

[0122] The intelligent and accurate extraction system of video axis files based on OCR provided in this embodiment first performs preprocessing operations such as decoding, sampling, and image enhancement on the video to reduce the computational load, remove noise, and enhance text edges, laying a foundation for accurately extracting various features in the subsequent steps and improving the overall processing efficiency. Secondly, by comprehensively extracting multi-modal features of video frames, the traditional way of relying on manual annotation of text information is changed, avoiding the problems of low efficiency and easy errors in manual annotation, and obtaining video content information from multiple dimensions. After that, by analyzing the scene type, it provides a basis for subsequent feature fusion and extraction strategy formulation, enabling the system to perform differential processing for different scenes and improving the comprehensiveness of understanding video content. Furthermore, by fusing multi-modal features, the defect that existing automatic extraction methods lack a comprehensive understanding of video content is made up for, considering various modal information comprehensively, enhancing the expression ability of video content, and helping to improve the accuracy and integrity of extraction. Then, through the pre-trained adaptive attention model, attention weights are dynamically allocated according to the multi-modal fusion features to highlight key features, further enhancing the ability to capture key content in the video, thereby improving the accuracy of extraction. Next, the weighted feature vector sequence is modeled through a long short-term memory network to effectively capture the temporal dependence relationship of video content, enabling the system to better process the information in the time dimension of the video and improving the accuracy and integrity of understanding video content. Finally, by comprehensively using various information and pre-trained deep network methods, the system can automatically adjust the extraction strategy according to different video types and contents, generate accurate and complete video frame axis files, and solve the problem of poor accuracy and integrity of the video axis files extracted by traditional methods. In summary, the intelligent and accurate extraction system of video axis files based on OCR provided in this embodiment effectively solves the problem of how to efficiently, accurately, and completely extract video axis files.

[0123] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

Claims

1. An intelligent and accurate extraction method for video axis files based on OCR, characterized in that, The method includes: Obtaining the video to be extracted and extraction task information, preprocessing the video to be extracted to obtain target video frames; Extracting the OCR text features, image features, and audio features of the target video frames; Performing scene analysis on the target video frames to obtain the scene type; Obtaining a fused feature vector based on the scene type, the OCR text features, image features, and audio features; Obtaining a weighted feature vector based on a pre-trained adaptive attention model and the fused feature vector; Using a long short-term memory network model to model the weighted feature vector to obtain hidden state sequence information; Generating a target video frame axis file based on the hidden state sequence information, the long short-term memory network model, the extraction task information, the scene type, a pre-trained deep network method, and the target video frames.

2. The method according to claim 1, wherein The obtaining the video to be extracted and extraction task information, preprocessing the video to be extracted to obtain target video frames includes: Using a video decoding library to convert the video to be extracted into a video frame sequence; Sampling the video frame sequence based on a preset sampling interval to obtain sampled video frames; Performing histogram equalization processing on the sampled video frames to obtain target video frames.

3. The method according to claim 1, characterized in that Extracting the OCR text features of the target video frames includes: Using an OCR engine to recognize the characters in the target video frames to obtain text information; Performing word segmentation processing, part-of-speech tagging processing, and keyword extraction processing on the text information to obtain the word frequency data and inverse document frequency data of each keyword; Constructing OCR text features based on the word frequency data and inverse document frequency data of each keyword.

4. The method according to claim 3, characterized in that, Extracting the image features of the target video frames includes: Inputting the target video frames into the convolutional layer of a pre-trained convolutional neural network, and using the convolutional layer to generate a preliminary feature map based on the target video frames, where the preliminary feature map includes edge features and texture features; Inputting the preliminary feature map into a pooling layer, using the pooling layer to divide the preliminary feature map into several small regions, performing max pooling processing on each small region, and taking the maximum value in each small region as the output to obtain a reduced feature map; Flattening the reduced feature map into a one-dimensional vector, and using a fully connected layer to map the one-dimensional vector into a fixed-length vector to obtain image features.

5. The method according to claim 4, characterized in that The pre-trained convolutional neural network includes at least three groups of combinations of convolutional layers and pooling layers.

6. The method according to claim 5, characterized in that, The obtaining a fused feature vector based on the scene type, the OCR text features, image features, and audio features includes: Determining a default weight based on the corresponding relationship between the default weight and the scene type and the scene type; Performing feature fusion on the OCR text features, image features, and audio features based on the default weight to obtain a fused feature vector.

7. The method according to claim 6, wherein The obtaining a weighted feature vector based on a pre-trained adaptive attention model and the fused feature vector includes: Calculating an attention weight vector based on a pre-trained adaptive attention model and the fused feature vector; Using the attention weight vector to perform weighted processing on the fused feature vector to obtain a weighted feature vector.

8. The method according to claim 7, wherein Generating a target video frame axis file based on the hidden state sequence information, long short-term memory network model, extraction task information, scene type, pre-trained deep network method, and target video frame, includes: Defining environmental state information based on the hidden state sequence information; Defining an extraction policy set based on the scene type; Defining a reward policy set based on the extraction task information; Using the pre-trained deep network method to obtain a target extraction policy set based on the environmental state information, extraction policy set, and reward policy set; Processing the target video frame based on the target extraction policy set to generate a target video frame axis file.

9. The method according to claim 8, wherein The method further includes: Extracting the target OCR text feature, target image feature, and target audio feature of the target video frame axis file; Using a fusion analysis technique based on the Transformer architecture to analyze the target OCR text feature, target image feature, and target audio feature to obtain the consistency data of the target OCR text feature, target image feature, and target audio feature on the time axis, and the complementary data of the target OCR text feature, target image feature, and target audio feature on the time axis; Obtaining collaborative score data based on the consistency data and complementary data; Performing an optimization process on the target video frame axis file based on the reinforcement learning model, meta-learning model, and the collaborative score data to obtain an optimized video frame axis file.

10. An intelligent and accurate extraction system for video axis files based on OCR, characterized in that, The system includes: A preprocessing module, configured to obtain a video to be extracted and extraction task information, and preprocess the video to be extracted to obtain a target video frame; An extraction module, configured to extract the OCR text feature, image feature, and audio feature of the target video frame; An analysis module, configured to perform a scene analysis on the target video frame to obtain a scene type; A fusion module, configured to obtain a fusion feature vector based on the scene type, the OCR text feature, image feature, and audio feature; A calculation module, configured to obtain a weighted feature vector based on a pre-trained adaptive attention model and the fusion feature vector; A modeling module, configured to model the weighted feature vector using a long short-term memory network model to obtain hidden state sequence information; A generation module, configured to generate a target video frame axis file based on the hidden state sequence information, long short-term memory network model, extraction task information, scene type, pre-trained deep network method, and target video frame.

Citation Information

Patent Citations

  • Video behavior automatic description method based on deep reinforcement learning

    CN111460883A

  • Method and device for generating description information of multimedia data, equipment and medium

    CN111723937A

  • Text information extraction method, system and equipment and medium

    CN113094509A

  • Video production system based on AI technology

    CN117376502A

  • Manipulator visual strategy model training method, manipulator visual strategy model control method and manipulator visual strategy model training system

    CN118493396A

Cited By

  • Family safety AI monitoring method and system based on multi-modal algorithm fusion

    CN120976853A