Video content description method and device, electronic equipment and storage medium
A video content description method based on multimodal data extraction and cross-modal feature integration solves the problem of difficult language template construction and improves the accuracy of video content description.
Patent Information
- Application Number
- CN202511113563.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-14
AI Technical Summary
In existing video content description methods, language templates are difficult to construct, resulting in insufficient accuracy in description, inability to cover all situations, and a tendency to make mistakes.
By acquiring target video data, multimodal data extraction and spiking neural coding are performed to generate spatiotemporal feature maps. Hierarchical temporal transformation and cross-modal feature integration are then carried out to generate cross-modal context representations and perform text decoding, ensuring accurate alignment between semantic features and auditory features.
It improves the accuracy of video content description, retains key action and audio information in the target video data, and enhances the accuracy of the description.
Smart Images

Figure CN120956982A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, applicable to the fields of financial technology and medical technology, and particularly to a video content description method and apparatus, electronic device and storage medium. Background Technology
[0002] Video content description is used to convert visual and auditory information from videos into textual descriptions. For example, in the fintech field, video content description of bank surveillance footage helps regulators quickly understand the bank's basic information; in the medical technology field, video content description of surgical videos helps doctors or patients quickly understand their condition.
[0003] Currently, video content description methods typically describe the content of a video based on pre-defined language templates. However, in this method, language templates are difficult to construct and cannot cover all situations, which can easily lead to errors in video content description. Therefore, improving the accuracy of video content description has become an urgent technical problem to be solved. Summary of the Invention
[0004] The main objective of this application is to provide a video content description method, apparatus, electronic device, and storage medium, which aims to improve the accuracy of video content description.
[0005] To achieve the above objectives, a first aspect of this application proposes a video content description method, the method comprising:
[0006] Acquire target video data;
[0007] Multimodal data extraction is performed on the target video data to obtain a multimodal data stream;
[0008] The multimodal data stream is subjected to pulse neural coding to obtain a spatiotemporal feature map;
[0009] The spatiotemporal feature map is subjected to hierarchical temporal transformation to obtain semantic units and Mel frequency cepstral coefficients;
[0010] Cross-modal feature integration is performed on the semantic units and the Mel frequency cepstral coefficients to obtain a cross-modal context representation;
[0011] Text decoding is performed on the cross-modal context representation to obtain the target natural language information.
[0012] In some embodiments, the hierarchical temporal transformation of the spatiotemporal feature map to obtain semantic units and Mel frequency cepstral coefficients includes:
[0013] The spatiotemporal feature map is dynamically adjusted in terms of temporal granularity to obtain a time window feature.
[0014] The time window features are subjected to hierarchical feature weighting to obtain time window weighted features;
[0015] The time-window weighted features are aggregated to obtain the semantic unit and the Mel frequency cepstral coefficients.
[0016] In some embodiments, the step of dynamically adjusting the temporal granularity of the spatiotemporal feature map to obtain a time window feature includes:
[0017] The temporal complexity of the spatiotemporal feature map is evaluated to obtain a temporal complexity score;
[0018] Based on the time complexity score, a window size decision is made to obtain a time window size scheme;
[0019] Based on the time window size scheme, the spatiotemporal feature map is windowed to obtain the time window feature.
[0020] In some embodiments, the step of performing hierarchical feature weighting on the time window features to obtain time window weighted features includes:
[0021] The time window features are segmented to obtain segmented features;
[0022] The hierarchical features are weighted to obtain the hierarchical attention weights;
[0023] Based on the hierarchical attention weights, the hierarchical features are weighted to obtain the time window weighted features.
[0024] In some embodiments, the step of performing feature aggregation on the time-window weighted features to obtain the semantic unit and the Mel frequency cepstral coefficients includes:
[0025] The time-window weighted features are integrated to obtain aggregated features;
[0026] The aggregated features are semantically annotated to obtain the semantic units;
[0027] Audio features are extracted from the aggregated features to obtain the Mel frequency cepstral coefficients.
[0028] In some embodiments, the cross-modal feature integration of the semantic unit and the Mel frequency cepstral coefficients to obtain a cross-modal context representation includes:
[0029] Timestamp matching is performed on the semantic units and the Mel frequency cepstral coefficients to obtain multimodal alignment features;
[0030] The multimodal alignment features are parsed to obtain interactive features;
[0031] The interactive features are merged to obtain fused features;
[0032] The fused features are used to construct a context to obtain the cross-modal context representation.
[0033] In some embodiments, the step of performing multimodal data extraction on the target video data to obtain a multimodal data stream includes:
[0034] The target video data is sampled frame by frame to obtain an image data stream;
[0035] The target video data is segmented into audio frames to obtain an audio data stream;
[0036] The image data stream and the audio data stream are merged to obtain the multimodal data stream.
[0037] To achieve the above objectives, a second aspect of this application provides a video content description apparatus, the apparatus comprising:
[0038] The data acquisition module is used to acquire target video data;
[0039] The data extraction module is used to extract multimodal data from the target video data to obtain a multimodal data stream;
[0040] The pulse coding module is used to perform pulse neural coding on the multimodal data stream to obtain a spatiotemporal feature map;
[0041] The temporal transformation module is used to perform hierarchical temporal transformation on the spatiotemporal feature map to obtain semantic units and Mel frequency cepstral coefficients.
[0042] The feature integration module is used to perform cross-modal feature integration on the semantic unit and the Mel frequency cepstral coefficients to obtain a cross-modal context representation;
[0043] The text decoding module is used to decode the cross-modal context representation to obtain the target natural language information.
[0044] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.
[0045] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0046] The video content description method, apparatus, electronic device, and storage medium proposed in this application acquire target video data and perform multimodal data extraction to obtain a multimodal data stream, which provides a synchronized audio-visual foundation for video content description. Subsequently, pulse neural coding is performed on the multimodal data stream to convert pixels and sound waves in the multimodal data stream into high-resolution spatiotemporal feature maps, preserving millisecond-level key events in the target video data, thereby reducing redundancy in the multimodal data stream. Next, hierarchical temporal transformation is performed on the spatiotemporal feature maps to obtain semantic units and Mel-frequency cepstral coefficients, enabling semantic capture from action to dialogue-level events, thereby ensuring accurate alignment of semantic features and auditory features. Furthermore, cross-modal feature integration is performed on the semantic units and Mel-frequency cepstral coefficients to obtain a cross-modal context representation, and text decoding is performed on the cross-modal context representation, so that the output target natural language information retains key action information and key audio information in the target video data, thereby improving the accuracy of video content description. Attached Figure Description
[0047] Figure 1 This is a flowchart of the video content description method provided in the embodiments of this application;
[0048] Figure 2 yes Figure 1 The flowchart of step S102 in the document;
[0049] Figure 3 yes Figure 1 The flowchart of step S104 in the process;
[0050] Figure 4 yes Figure 3 The flowchart of step S301 in the process;
[0051] Figure 5 yes Figure 3 The flowchart of step S302 in the text;
[0052] Figure 6 yes Figure 3 The flowchart of step S303 in the process;
[0053] Figure 7 yes Figure 1 The flowchart of step S105 in the process;
[0054] Figure 8This is a schematic diagram of the structure of the video content description device provided in the embodiments of this application;
[0055] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0057] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0059] First, let's analyze some of the terms used in this application:
[0060] Video Content Description System: The video content description system is an end-to-end multimodal semantic mapping framework. The input of the video content description system is the original video stream, and the output is a natural language description of the corresponding scene. The core process of the video content description system includes cross-modal feature extraction, temporal context modeling, and symbolic text generation. The video content description system aims to convert visual and auditory signals into structured and interpretable semantic representations.
[0061] Autoregressive techniques are a sequence modeling method. The core idea of autoregressive techniques is to transform the sequence generation task into a chain of conditional probabilities. Specifically, the output of each step depends on all previously generated elements, and the overall sequence is constructed through progressive sampling. In video content description systems, autoregressive techniques use cross-modal context representations as conditions and leverage neural networks to iteratively predict the probability distribution of the next word until a complete natural language description is generated. Autoregressive techniques use causal masks to ensure that predictions are based solely on historical information, thereby guaranteeing the rationality and coherence of the generation process.
[0062] Video content description is used to convert visual and auditory information from videos into textual descriptions. For example, in the fintech field, video content description of bank surveillance footage helps regulators quickly understand the bank's basic information; in the medical technology field, video content description of surgical videos helps doctors or patients quickly understand their condition.
[0063] Currently, video content description methods typically describe the content of a video based on pre-defined language templates. However, in this method, language templates are difficult to construct and cannot cover all situations, which can easily lead to errors in video content description. Therefore, improving the accuracy of video content description has become an urgent technical problem to be solved.
[0064] Based on this, embodiments of this application provide a video content description method and apparatus, electronic device and storage medium, aiming to improve the accuracy of video content description.
[0065] The video content description method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the video content description method in this application is described.
[0066] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0067] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0068] The video content description method provided in this application relates to the field of video processing technology and is applicable to the fintech and medical fields. The video content description method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the video content description method, but is not limited to the above forms.
[0069] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0070] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0071] Figure 1 This is an optional flowchart of a video content description method provided in this application embodiment. This method can be used in a video content description system. Figure 1The method may include, but is not limited to, steps S101 to S106.
[0072] Step S101: Obtain target video data;
[0073] Step S102: Perform multimodal data extraction on the target video data to obtain a multimodal data stream;
[0074] Step S103: Perform spiking neural coding on the multimodal data stream to obtain a spatiotemporal feature map;
[0075] Step S104: Perform hierarchical temporal transformation on the spatiotemporal feature map to obtain semantic units and Mel frequency cepstral coefficients;
[0076] Step S105: Perform cross-modal feature integration on semantic units and Mel frequency cepstral coefficients to obtain cross-modal context representation;
[0077] Step S106: Decode the cross-modal context representation to obtain the target natural language information.
[0078] Steps S101 to S106 of this embodiment involve acquiring target video data and performing multimodal data extraction to obtain a multimodal data stream. This provides a synchronized audio-visual foundation for describing video content. Subsequently, pulse neural coding is performed on the multimodal data stream to convert pixels and sound waves in the multimodal data stream into a high-resolution spatiotemporal feature map, preserving millisecond-level key events in the target video data and reducing redundancy in the multimodal data stream. Next, hierarchical temporal transformation is performed on the spatiotemporal feature map to obtain semantic units and Mel-frequency cepstral coefficients, enabling semantic capture from action to dialogue-level events, thereby ensuring accurate alignment of semantic features and auditory features. Furthermore, cross-modal feature integration is performed on the semantic units and Mel-frequency cepstral coefficients to obtain a cross-modal context representation, and text decoding is performed on the cross-modal context representation. This ensures that the output target natural language information retains key action information and key audio information in the target video data, thereby improving the accuracy of video content description.
[0079] In step S101 of some embodiments, the target video data refers to the video content that needs to be described. For example, in the financial field, the target video data may be bank surveillance video, and in the medical field, the target video data may be video recorded during surgery.
[0080] This application embodiment can use a camera or other video capture device to record video information of a target scene. For example, in the financial field, cameras are installed near bank ATMs to capture video data of the transaction process; in the medical field, cameras in operating rooms can capture video data of the surgical procedure. Further, storing this video information yields video data. Finally, by extracting the video data, the target video data can be obtained. It is important to note that to ensure the target video data content can be accurately described, attention must be paid to information such as the resolution, frame rate, and compression format of the target video data to ensure that the target video data meets the requirements for video content description.
[0081] In step S102 of some embodiments, a multimodal data stream refers to a data set that includes image data and audio data.
[0082] This application embodiment obtains an image data stream by sampling video frames of the target video data, and obtains an audio data stream by performing audio frame segmentation on the target video data. Then, the image data stream and the audio data stream are merged to form a multimodal data stream containing visual and auditory information.
[0083] For details, please refer to Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S203:
[0084] Step S201: Sample video frames from the target video data to obtain an image data stream;
[0085] Step S202: Perform audio frame segmentation on the target video data to obtain an audio data stream;
[0086] Step S203: Merge the image data stream and the audio data stream to obtain a multimodal data stream.
[0087] In step S201 of some embodiments, the image data stream refers to a data sequence consisting of consecutive image frames.
[0088] This application embodiment captures image frames periodically from target video data by setting a specific time interval. Furthermore, the captured image frames are sorted in chronological order to obtain an image data stream. For example, in a financial transaction scenario, image frame data from the monitoring video of the trading hall can be captured by setting a sampling frequency of 30 frames per second to form an image data stream. In a medical surgery scenario, an image data stream can be generated by sampling the surgical video at 30 frames per second.
[0089] In step S202 of some embodiments, the audio data stream refers to a data sequence consisting of consecutive audio frames.
[0090] This application embodiment extracts audio signals from target video data to obtain video audio signals. Furthermore, the video audio signals are divided into multiple short frames by a pre-set audio frame time length. The audio frame time length refers to the time length of the audio data in the audio frame. For example, an audio frame time length of 20-40 milliseconds means that the audio frame contains 20-40 milliseconds of audio data.
[0091] In step S203 of some embodiments, the image data stream and audio data stream sampled from the target video data are synchronized according to the same timestamp to obtain a timestamp-aligned image data stream and audio data stream. Furthermore, the timestamp-aligned image data stream and audio data stream are merged according to the timestamp to form a multimodal data stream containing visual and auditory information.
[0092] Steps S201 to S203, as illustrated in this embodiment, involve sampling video frames and segmenting audio frames from the target video data. This ensures the synchronization and accuracy of the image and audio data streams extracted from the target video data, laying the foundation for video content parsing. Furthermore, by merging the aforementioned image and audio data streams into a multimodal data stream, comprehensive capture of the video content of the target video data is achieved, enhancing the richness and completeness of the multimodal data stream.
[0093] In step S103 of some embodiments, the spatiotemporal feature map refers to the feature representation of a multimodal data stream in time and space.
[0094] This application embodiment achieves pulse neural encoding of the multimodal data stream by performing spatiotemporal convolution on the multimodal data stream, obtaining a spatiotemporal feature map. Specifically, in the spatial dimension, dynamically transformed weight distribution is used to detect features such as edges and corners in the image data stream of the multimodal data stream, and local comparisons are performed on these features. In the temporal dimension, first-order or second-order difference is used to detect abrupt changes in pixel values between two frames in the image data stream, thereby transforming the "bright-dark-bright" or "dark-bright-dark" transition representation of pixel values in the multimodal data stream into the timing and intensity of pulse firing. Simultaneously, the audio data stream in the multimodal data stream is decomposed to obtain energy trajectories of several Mel filter banks. Differentiating each energy trajectory yields pulse sequences corresponding to "energy rising edge" and "energy falling edge". Furthermore, the pulse sequences corresponding to "energy rising edge" and "energy falling edge" are input into a two-dimensional grid composed of leakage integral-firing neurons to obtain the audio features of the audio data stream in the temporal and spatial dimensions.
[0095] In step S104 of some embodiments, a semantic unit refers to a basic unit with a clear semantic meaning extracted from the spatiotemporal feature map of the image data stream. Mel frequency cepstral coefficients refer to features in the spatiotemporal feature map of the audio data stream.
[0096] This application embodiment dynamically adjusts the temporal granularity of the spatiotemporal feature map to generate time window features adapted to different time scales. Subsequently, the time window features are hierarchically weighted to highlight key information in the time window features, thereby forming time window weighted features. Finally, by aggregating the various time window weighted features, semantic units and Mel frequency cepstral coefficients that can represent the semantic content in the target video data can be extracted from the spatiotemporal feature map.
[0097] For details, please refer to Figure 3 In some embodiments, step S104 may include, but is not limited to, steps S301 to S303:
[0098] Step S301: Dynamically adjust the temporal granularity of the spatiotemporal feature map to obtain the time window feature;
[0099] Step S302: Perform hierarchical feature weighting on the time window features to obtain time window weighted features;
[0100] Step S303: Perform feature aggregation on the time window weighted features to obtain semantic units and Mel frequency cepstral coefficients.
[0101] In step S301 of some embodiments, the time window feature refers to the feature in the spatiotemporal feature map extracted within a specific time window. The time window feature reflects the video information of the target video data within a specific time period.
[0102] This application embodiment evaluates the temporal complexity of the spatiotemporal feature map to determine the complexity of the video content of the target video data as it changes over time and space. Then, based on the evaluated temporal complexity score, it determines the optimal time window size and generates a time window size scheme to adapt to the dynamic characteristics of the video content of the target video data. Finally, it applies the time window size scheme to perform feature windowing processing on the spatiotemporal feature map, thereby accurately extracting the time window features that reflect the temporal characteristics of the target video data.
[0103] For details, please refer to Figure 4 In some embodiments, step S301 may include, but is not limited to, steps S401 to S403:
[0104] Step S401: Evaluate the temporal complexity of the spatiotemporal feature map and obtain a temporal complexity score;
[0105] Step S402: Based on the time complexity score, make a window size decision to obtain a time window size scheme;
[0106] Step S403: Based on the time window size scheme, the spatiotemporal feature map is windowed to obtain the time window features.
[0107] In step S401 of some embodiments, the temporal complexity score refers to a score that measures the temporal variation characteristics of video content.
[0108] This application embodiment can perform differential or optical flow estimation on each adjacent frame in a multimodal data stream based on the spatiotemporal feature map to obtain the differences between each adjacent frame in the multimodal data stream, such as the magnitude and direction of motion vectors and changes in audio signals. Furthermore, the temporal complexity of the video content in the target video data is evaluated based on the differences between each adjacent frame in the multimodal data stream, thereby generating a corresponding temporal complexity score. Specifically, the greater the difference between each adjacent frame in the multimodal data stream, the greater the temporal complexity score; the smaller the difference between each adjacent frame in the multimodal data stream, the smaller the temporal complexity score.
[0109] In step S402 of some embodiments, the time window size scheme refers to a strategy or plan for determining the specific time window size for the video content description task.
[0110] In this embodiment, the above-mentioned time complexity score is input into a constrained optimization function or optimization model to output the optimal time window. It should be noted that the constraint of the optimization function or optimization model can be expressed as the time window needing to be able to include semantically complete actions in the multimodal data stream and also needing to retain time precision. Therefore, the constraint can be a restriction on parameters such as the minimum number of frames, the maximum number of frames, and the overlap ratio of adjacent windows.
[0111] In step S403 of some embodiments, by cutting the four-dimensional tensor features in the spatiotemporal feature map into several feature representations along the time axis, and further performing pooling or attention weighting on each feature representation, a fixed-length time window feature can be obtained.
[0112] Steps S401 to S403, as illustrated in this embodiment, involve evaluating the temporal complexity of the spatiotemporal feature map to obtain a temporal complexity score. This enables the quantification of the dynamic temporal changes in the target video data content, providing temporal information for video content description. Secondly, based on the temporal complexity score, a window size decision is made to obtain a time window size scheme, allowing the selected time window to adaptively match the video content complexity of the target video data. Finally, based on the time window size scheme, the spatiotemporal feature map is windowed to obtain time window features, making the analysis of the target video data content more detailed and thus improving the accuracy of video content description.
[0113] In step S302 of some embodiments, the time window weighted feature refers to the time window feature after weighting. It should be noted that the weight of the time window feature can reflect whether the video content of the target video data is key information.
[0114] In this embodiment, the time window features are decomposed into multiple levels through feature layering to form multiple layered features. Next, the weight of each layered feature is calculated to obtain the layered attention weight. Finally, the layered attention weight is used to weight the above layered features to obtain the time window weighted features.
[0115] For details, please refer to Figure 5 In some embodiments, step S302 may include, but is not limited to, steps S501 to S503:
[0116] Step S501: Perform feature stratification on the time window features to obtain stratified features;
[0117] Step S502: Calculate the weights of the hierarchical features to obtain the hierarchical attention weights;
[0118] Step S503: Based on the hierarchical attention weights, the hierarchical features are weighted to obtain time window weighted features.
[0119] In step S501 of some embodiments, the hierarchical feature refers to the feature of the time window feature at each level.
[0120] This application embodiment can identify the interrelationships between features of various time windows to obtain feature associations. Furthermore, it can identify the degree of influence of each time window feature on the description of video content. Finally, based on the above feature associations and degree of influence, each time window feature is divided into different levels, thereby forming a hierarchical feature.
[0121] In step S502 of some embodiments, the hierarchical attention weight is a weight used to indicate the importance of each hierarchical feature.
[0122] In this embodiment, the hierarchical features can be globally pooled to obtain a 1×d-dimensional low-dimensional vector. This low-dimensional vector can then be mapped to a score space to obtain a scalar score, which represents the importance of the hierarchical features in the video content description task. Finally, the scalar score can be normalized to obtain the hierarchical attention weights.
[0123] In step S503 of some embodiments, the hierarchical features can be transformed into time window weighted features by calculating the product of each hierarchical feature and its corresponding hierarchical attention weight.
[0124] Steps S501 to S503 as shown in the embodiments of this application, by performing feature layering on the time window features, can more finely distinguish information at different levels in the multimodal data stream, thereby laying the foundation for video content description. Furthermore, weight calculation is performed on the layered features to obtain layered attention weights, and then feature weighting is performed on the layered features based on the layered attention weights to obtain time window weighted features, which can strengthen the influence of important features in the layered features, thereby improving the accuracy of video content description.
[0125] In step S303 of some embodiments, the above-mentioned time window weighted features are integrated to form aggregated features, and then the aggregated features are semantically annotated to generate semantic units representing specific semantic content in the target video data. At the same time, audio features are extracted from the aggregated features to obtain Mel frequency cepstral coefficients.
[0126] For details, please refer to Figure 6 In some embodiments, step S303 may include, but is not limited to, steps S601 to S603:
[0127] Step S601: Integrate the time-window weighted features to obtain aggregated features;
[0128] Step S602: Semantic annotation is performed on the aggregated features to obtain semantic units;
[0129] Step S603: Extract audio features from the aggregated features to obtain Mel frequency cepstral coefficients.
[0130] In step S601 of some embodiments, aggregated features refer to feature vectors that contain information of multiple time window weighted features.
[0131] In this embodiment, the time window weighted features are concatenated into a high-dimensional sequence according to the time steps of the target video data. Then, the high-dimensional sequences of adjacent frames in the multimodal data stream are recursively fused in the temporal dimension using a pre-set gated memory network to obtain the aggregated features.
[0132] In step S602 of some embodiments, the semantic representation of the aggregated features, i.e., semantic features, can be obtained by mapping the above-mentioned aggregated features to a predefined semantic prototype space. For example, in the financial field, when analyzing transaction behavior, the transaction behavior information in the aggregated features can be mapped to specific transaction types, such as "buy" or "sell" through semantic annotation. In the medical field, the physiological monitoring behavior of patients can be mapped to specific health conditions, such as "normal" or "abnormal," through semantic annotation.
[0133] In step S603 of some embodiments, an audio feature vector can be obtained by locating the audio-related feature vector in the aggregated features. Then, the audio feature vector is mapped back to the time-frequency domain to obtain the energy representation of the approximate Mel filter bank. Finally, the logarithm of the energy representation is taken and a discrete cosine transform is performed to obtain the Mel frequency cepstral coefficients in the aggregated features.
[0134] Steps S601 to S603, as illustrated in this embodiment, integrate the time-window weighted features to obtain aggregated features, providing a multi-dimensional data foundation for describing the video content of the target video data. Subsequently, semantic annotation is performed on the aggregated features to obtain semantic units, making the understanding of the video content of the target video data more intuitive and accurate. At the same time, audio features are extracted from the aggregated features to obtain Mel-frequency cepstral coefficients, enhancing the quantification and recognition capabilities of audio information in the target video data and providing auditory dimension information for multimodal analysis of video content description, thereby improving the accuracy of video content description.
[0135] Steps S301 to S303, as illustrated in this embodiment, involve dynamically adjusting the temporal granularity of the spatiotemporal feature map to obtain time window features, making the captured events at different time scales more accurate. Next, the time window features are subjected to hierarchical feature weighting to obtain time window weighted features, ensuring that key information in the target video data receives sufficient attention, thereby improving the accuracy of video content description. Finally, the time window weighted features are aggregated to obtain semantic units and Mel-frequency cepstral coefficients, which helps to deepen the semantic understanding of the video content of the target video data, thereby improving the accuracy of video content description.
[0136] In step S105 of some embodiments, cross-modal context representation refers to a context representation that can reflect the multimodal information of the video content of the target video data.
[0137] This application embodiment aligns the timestamped semantic units with Mel frequency cepstral coefficients to form multimodal alignment features, then analyzes the interaction between the multimodal alignment features, and extracts interaction features based on the interaction. Next, the above interaction features are merged to obtain fused features. Finally, based on the context features of the fused features, context is constructed to obtain cross-modal context representation, thereby realizing the integrated expression of image and audio semantics.
[0138] For details, please refer to Figure 7 In some embodiments, step S105 may include, but is not limited to, steps S701 to S704:
[0139] Step S701: Perform timestamp matching on semantic units and Mel frequency cepstral coefficients to obtain multimodal alignment features;
[0140] Step S702: Perform feature interaction parsing on the multimodal alignment features to obtain the interaction features;
[0141] Step S703: Merge the interactive features to obtain fused features;
[0142] Step S704: Construct context for the fused features to obtain cross-modal context representation.
[0143] In step S701 of some embodiments, multimodal alignment features refer to a set of features of different modalities having the same time stamp.
[0144] In this embodiment of the application, a time axis is formed with the timestamp as the axis, and then the semantic units and Mel frequency cepstral coefficients are aligned on the time axis to obtain aligned semantic units and Mel frequency cepstral coefficients, i.e., multimodal alignment features.
[0145] In step S702 of some embodiments, the interaction feature can characterize the interaction relationship between features of different modalities in the multimodal alignment feature.
[0146] This application embodiment obtains the interaction matrix by performing outer product calculation on the features of different modalities in the multimodal alignment feature. Then, by performing low-rank decomposition on this interaction matrix, the higher-order covariance of the interaction matrix can be obtained, which represents the interaction features between the features of different modalities in the multimodal alignment feature.
[0147] In step S703 of some embodiments, the fused feature refers to a comprehensive feature that includes information from multiple modalities.
[0148] In this embodiment of the application, fused features can be obtained by fusing interactive features with corresponding semantic units and Mel frequency cepstral coefficients.
[0149] In step S704 of some embodiments, the above-mentioned fused features can be understood in context by a self-attention network with causal masking, so as to establish long-range dependencies on the fused features and thus obtain cross-modal context representation.
[0150] Steps S701 to S704, as illustrated in this embodiment, obtain multimodal alignment features by performing timestamp matching on semantic units and Mel-frequency cepstral coefficients. This ensures the temporal consistency between semantic units and Mel-frequency cepstral coefficients, laying the foundation for cross-modal feature integration. Next, feature interaction parsing is performed on the multimodal alignment features to obtain interactive features, which clarifies the potential connections and influences between semantic units and Mel-frequency cepstral coefficients. Then, feature merging is performed on the interactive features to obtain fused features, improving the expressive power of the fused features. Finally, context construction is performed on the fused features to obtain a cross-modal context representation, making the description of video content more accurate.
[0151] In step S106 of some embodiments, by inputting the cross-modal context representation into a pre-trained text model, the text model can output a complete sentence based on the cross-modal context representation in an autoregressive manner. Further, key numbers or entities are detected in the complete sentence, and when the detection results ensure that the complete sentence is consistent with the cross-modal context representation, the complete sentence, i.e., the target natural language information, is output.
[0152] This application acquires target video data and performs multimodal data extraction to obtain a multimodal data stream, providing a synchronized audio-visual foundation for video content description. Subsequently, pulse neural coding is applied to the multimodal data stream to transform pixels and sound waves in the multimodal data stream into high-resolution spatiotemporal feature maps, preserving millisecond-level key events in the target video data and reducing redundancy in the multimodal data stream. Next, hierarchical temporal transformation is performed on the spatiotemporal feature maps to obtain semantic units and Mel-frequency cepstral coefficients, enabling semantic capture from action to dialogue-level events, thus ensuring accurate alignment of semantic features and auditory features. Furthermore, cross-modal feature integration is performed on the semantic units and Mel-frequency cepstral coefficients to obtain a cross-modal context representation, and text decoding is performed on the cross-modal context representation, so that the output target natural language information retains key action information and key audio information in the target video data, thereby improving the accuracy of video content description.
[0153] Please see Figure 8 This application also provides a video content description apparatus that can implement the above-described video content description method. The apparatus includes:
[0154] Data acquisition module 801 is used to acquire target video data;
[0155] The data extraction module 802 is used to extract multimodal data from the target video data to obtain a multimodal data stream;
[0156] The pulse coding module 803 is used to perform pulse neural coding on the multimodal data stream to obtain a spatiotemporal feature map;
[0157] The timing transformation module 804 is used to perform hierarchical timing transformation on the spatiotemporal feature map to obtain semantic units and Mel frequency cepstral coefficients.
[0158] The feature integration module 805 is used to perform cross-modal feature integration on semantic units and Mel frequency cepstral coefficients to obtain cross-modal context representation;
[0159] The text decoding module 806 is used to decode the cross-modal context representation to obtain the target natural language information.
[0160] The specific implementation of the video content description device is basically the same as the specific embodiment of the video content description method described above, and will not be repeated here.
[0161] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described video content description method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0162] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0163] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0164] The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the video content description method of the embodiments of this application.
[0165] The input / output interface 903 is used to implement information input and output;
[0166] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0167] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);
[0168] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0169] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video content description method.
[0170] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0171] The video content description method, video content description device, electronic device, and storage medium provided in this application embodiment acquire target video data, perform multimodal data extraction on the target video data to obtain a multimodal data stream, perform pulse neural coding on the multimodal data stream to obtain a spatiotemporal feature map, perform hierarchical temporal transformation on the spatiotemporal feature map to obtain semantic units and Mel-frequency cepstral coefficients, perform cross-modal feature integration on the semantic units and Mel-frequency cepstral coefficients to obtain a cross-modal context representation, and perform text decoding on the cross-modal context representation to obtain target natural language information.
[0172] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0173] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0174] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0175] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0176] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0177] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0178] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.
[0179] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0180] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0181] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0182] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for describing video content, characterized in that, The method includes: Acquire target video data; Multimodal data extraction is performed on the target video data to obtain a multimodal data stream; The multimodal data stream is subjected to pulse neural coding to obtain a spatiotemporal feature map; The spatiotemporal feature map is subjected to hierarchical temporal transformation to obtain semantic units and Mel frequency cepstral coefficients; Cross-modal feature integration is performed on the semantic units and the Mel frequency cepstral coefficients to obtain a cross-modal context representation; Text decoding is performed on the cross-modal context representation to obtain the target natural language information.
2. The method according to claim 1, characterized in that, The hierarchical temporal transformation of the spatiotemporal feature map to obtain semantic units and Mel-frequency cepstral coefficients includes: The spatiotemporal feature map is dynamically adjusted in terms of temporal granularity to obtain a time window feature. The time window features are subjected to hierarchical feature weighting to obtain time window weighted features; The time-window weighted features are aggregated to obtain the semantic unit and the Mel frequency cepstral coefficients.
3. The method according to claim 2, characterized in that, The step of dynamically adjusting the temporal granularity of the spatiotemporal feature map to obtain time window features includes: The temporal complexity of the spatiotemporal feature map is evaluated to obtain a temporal complexity score; Based on the time complexity score, a window size decision is made to obtain a time window size scheme; Based on the time window size scheme, the spatiotemporal feature map is windowed to obtain the time window feature.
4. The method according to claim 2, characterized in that, The step of performing hierarchical feature weighting on the time window features to obtain time window weighted features includes: The time window features are segmented to obtain segmented features; The hierarchical features are weighted to obtain the hierarchical attention weights; Based on the hierarchical attention weights, the hierarchical features are weighted to obtain the time window weighted features.
5. The method according to claim 2, characterized in that, The step of performing feature aggregation on the time-window weighted features to obtain the semantic unit and the Mel frequency cepstral coefficients includes: The time-window weighted features are integrated to obtain aggregated features; The aggregated features are semantically annotated to obtain the semantic units; Audio features are extracted from the aggregated features to obtain the Mel frequency cepstral coefficients.
6. The method according to any one of claims 1-5, characterized in that, The step of integrating the semantic units and the Mel frequency cepstral coefficients across modal features to obtain a cross-modal context representation includes: Timestamp matching is performed on the semantic units and the Mel frequency cepstral coefficients to obtain multimodal alignment features; The multimodal alignment features are parsed to obtain interactive features; The interactive features are merged to obtain fused features; The fused features are used to construct a context to obtain the cross-modal context representation.
7. The method according to any one of claims 1-5, characterized in that, The step of extracting multimodal data from the target video data to obtain a multimodal data stream includes: The target video data is sampled frame by frame to obtain an image data stream; The target video data is segmented into audio frames to obtain an audio data stream; The image data stream and the audio data stream are merged to obtain the multimodal data stream.
8. A video content description device, characterized in that, The device includes: The data acquisition module is used to acquire target video data; The data extraction module is used to extract multimodal data from the target video data to obtain a multimodal data stream; The pulse coding module is used to perform pulse neural coding on the multimodal data stream to obtain a spatiotemporal feature map; The temporal transformation module is used to perform hierarchical temporal transformation on the spatiotemporal feature map to obtain semantic units and Mel frequency cepstral coefficients. The feature integration module is used to perform cross-modal feature integration on the semantic unit and the Mel frequency cepstral coefficients to obtain a cross-modal context representation; The text decoding module is used to decode the cross-modal context representation to obtain the target natural language information.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the video content description method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the video content description method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Pulse neural network multi-mode lip reading method and system based on attention mechanism
CN115482582A
Video content description method based on combination of multi-modal fusion and multi-layer attention
CN115661697A
Multi-modal video description generation method and device, and computer readable medium
CN119136020A
Video content understanding method of multi-mode advertisement inventory intelligent matching system
CN120388324A
Video content structured disassembly analysis method and system based on time sequence segmentation
CN120411861A
Cited By
Time sequence feature calculation method and system, computer equipment and readable storage medium
CN121188077A