A video image data processing method and system based on a large-scale temporal model

Through the video image data processing method based on the timing model, the problems of dynamic spatiotemporal feature extraction and key information capture in the video data are solved, and a comprehensive understanding and efficient prediction of video content are achieved.

CN119723427BActive Publication Date: 2025-07-11SITENG HELI TIANJIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510220089.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-11
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively extract dynamic spatiotemporal features in video data, the key information is not captured accurately, the historical information is lacking in utilization, and the feature fusion and transformation are unreasonable, resulting in the inaccurate and comprehensive understanding of the video content.

Method used

Using a method based on a time series model, video image data is collected for preprocessing, video frames at key time points are extracted and interpolated, historical information and interpolated features are fused, feature transformation is performed, and finally input the video prediction model to predict the center position of the target area.

Benefits of technology

It realizes effective extraction of dynamic spatio-temporal features of video data and accurate capture of key information, improves the prediction accuracy and efficiency of the video prediction model, and can better understand the contextual relationship of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723427B_ABST
    Figure CN119723427B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for processing video image data based on a temporal large model, which relates to the technical field of video image data processing. Video image data is collected and preprocessed; key information of the preprocessed video data is captured to obtain video frames at key time points of the video data, interpolation operations are performed on the video frames at key time points to obtain interpolated eigenvalue; historical information is retrieved from the global memory pool and fused with the interpolated eigenvalue, and feature transformation operations are performed on the fused features to obtain transformed features; the transformed features are input into a video prediction model to predict the central position of the target area at future moments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a method and system for processing video image data based on a large temporal model, which relates to the technical field of video image data processing. Background Art

[0002] In today's digital age, video data has been widely used in many fields such as security monitoring, autonomous driving, and intelligent media. These scenarios have put forward higher and higher requirements for the analysis and understanding of video data, and it is necessary to accurately extract key information from videos, such as the positions and behaviors of objects.

[0003] Video data has obvious temporal characteristics, and there is temporal continuity and relevance between each frame of image. Traditional image processing methods often only focus on the information of a single image frame and cannot make full use of the temporal features of video data. Therefore, a technology that can effectively process temporal data is needed to mine the dynamic information that changes over time in videos.

[0004] In recent years, large models have achieved great success in the fields of natural language processing, computer vision, etc. Large models have powerful feature representation learning capabilities and can automatically learn complex patterns and features from massive data. Introducing large models into the field of video image data processing can leverage their powerful learning capabilities to handle the complexity and diversity of video data.

[0005] However, in the existing technology, there are still the following technical problems:

[0006] 1. Difficulty in extracting dynamic spatio-temporal features: The objects in video data move continuously in space and change continuously in time. How to effectively extract these dynamic spatio-temporal features is a key issue. Traditional methods cannot well capture the dynamic changes of objects in videos and the spatio-temporal dependence relationships between different frames, resulting in inaccurate and incomplete understanding of video content.

[0007] 2. Inaccurate capture of key information: Video data often contains a large amount of redundant information, and it is not easy to accurately capture key information from it. If key information cannot be accurately captured, it will lead to deviations in subsequent analysis and prediction tasks. For example, important targets may be missed in object detection tasks or wrong judgments may be made in behavior analysis.

[0008] 3. Lack of utilization of historical information: The existing processing methods do not make full use of the historical information of video data, making the analysis of the current frame by the model lack the support of context. In many actual application scenarios, historical information is crucial for understanding current events and predicting future trends. For example, historical information such as the previous driving trajectory of a vehicle is very important for predicting its future position and driving direction.

[0009] 4. Unreasonable feature fusion and transformation: Even if features from different sources can be obtained (such as current frame features and historical information features), how to perform effective fusion and subsequent reasonable transformation on the fused features to meet the requirements of the prediction model is still an unsolved problem. Unreasonable feature fusion and transformation may lead to information loss or introduce noise, affecting the performance of the final prediction model. Summary of the Invention

[0010] To solve the above technical problems, the present invention proposes a method for processing video image data based on a large temporal model, including the following steps:

[0011] Step S1: Collect video image data and preprocess the video image data;

[0012] Step S2: Capture key information from the preprocessed video data to obtain video frames at key time points of the video data, and perform interpolation operations on the video frames at key time points to obtain interpolated eigenvalue;

[0013] Step S3: Retrieve historical information from the global memory pool and fuse the interpolated eigenvalue, and perform feature transformation operations on the fused features to obtain transformed features;

[0014] Step S4: Input the transformed features into the video prediction model to predict the central position of the target area at future moments.

[0015] In a preferred embodiment, the step S2 includes:

[0016] S21: Extract video frames at key time points of the video data;

[0017] S22: Calculate the interpolated eigenvalue of the video frames at key time points.

[0018] In a preferred embodiment, let the video frame X at the key time point S The eigenvalue of the pixel at each coordinate (i, j) in is X ZS (i, j), and the interpolated eigenvalue is calculated from adjacent frame images using the interpolation method :

[0019] ;

[0020] f def Is the interpolation function, let , , for the video frame X at the key time point S , sample at non-integer coordinates (u, v), then the interpolation function is:

[0021] ;

[0022] where w mn is a weight coefficient calculated based on the fractional parts of u and v, and are the floor functions of u and v respectively, and m and n are index variables for summation.

[0023] In a preferred embodiment, in step S3, the interpolated eigenvalue is concatenated with the retrieved historical information M:

[0024] ,

[0025] where Concat represents the concatenation operation;

[0026] The fused feature F is subjected to a feature transformation operation to obtain the transformed feature F t :

[0027] .

[0028] In a preferred embodiment, in step S4, the transformed feature F t is input into the trained object detection model, and the object detection model outputs the bounding box information of the detected target region. For a two-dimensional image, the bounding box information is represented by the upper left coordinates (X min , Y min ) and the lower right coordinates (X max , Y max ). Then the center position of the target region is obtained by the following formula as the current position of the target region :

[0029] ;

[0030] .

[0031] In a preferred embodiment, the transformed feature F t is represented as a column vector , and after inputting the transformed feature F t into the neural network, the adjustment amounts and are obtained:

[0032] ;

[0033] where B a is the bias vector and Q a is the weight matrix;

[0034] is represented as:

[0035] ;

[0036] , obtain ;

[0037] Let: ;

[0038] ;

[0039] Among them, is the j-th matrix element of the first row of the weight matrix Q a of, is the j-th matrix element of the second row of the weight matrix Q a of, is the transformed feature F t expressed as the j-th eigenvector after being a column vector.

[0040] In a preferred embodiment, the central position of the target region predicted at the next moment is:

[0041] ;

[0042] .

[0043] In a preferred embodiment, in step S21, the video frame X with dimensions of T×H×W×C is input into the dimension expansion layer to obtain a feature map F with dimensions of T×H×W×C′ Z , perform dimension averaging on the feature map F Z to obtain an averaged result F with dimensions of T×1×1×C′ avg , where T is the number of time frames, H is the height of the video frame, W is the width of the video frame, C is the number of channels, and C′ is the number of intermediate channels;

[0044] Process the averaged result F avg through the data sequence recognition mechanism to obtain the time focus weight S, and multiply the time focus weight S by the video frame X in an element-wise multiplication manner to obtain the video frame at the key time point:

[0045] ;

[0046] represents element-wise multiplication.

[0047] The present invention also proposes a video image data processing system based on a temporal large model for implementing the above-mentioned video image data processing method based on a temporal large model, including: a data acquisition and preprocessing module, a key information capture module, an information fusion and transformation module, and a video prediction module;

[0048] The data acquisition and preprocessing module is used to acquire video image data and preprocess the video image data;

[0049] The key information capture module is used to capture key information from the preprocessed video data, obtain the video frames at the key time points of the video data, perform interpolation operations on the video frames at the key time points, and obtain the interpolated eigenvalue;

[0050] The information fusion and transformation module is used to retrieve historical information from the global memory pool and fuse the interpolated eigenvalue, perform feature transformation operations on the fused features, and obtain the transformed features;

[0051] The video prediction module is used to input the transformed features into the video prediction model to predict the central position of the target area at future moments.

[0052] Compared with the prior art, the present invention has the following beneficial technical effects:

[0053] 1. Effectively extract dynamic spatio-temporal features and key information: Extract dynamic time features and capture key information from the preprocessed video data, which can fully explore the dynamic change features of the video data in the time dimension, accurately capture key information, and make the finally obtained feature representation more comprehensively and accurately reflect the video content, providing a high-quality data basis for subsequent processing and analysis. In the video surveillance scenario, key information such as the movement trajectory and behavior changes of the target object can be effectively captured.

[0054] 2. Fusion of historical information to enhance feature representation ability: By performing a splicing operation on the interpolated eigenvalue and the retrieved historical information, combining the features of the current video frame with the historical information, making full use of the information correlation in the time series. This fusion method can enable the model to better understand the context relationship of the video content, further enrich the feature representation, and improve the model's processing ability for complex scenarios and long-term dependency relationships. In video behavior analysis, combining historical information can more accurately judge whether the current behavior is abnormal.

[0055] 3. Optimize the input of the prediction model to improve the prediction accuracy: Performing a linear transformation operation on the spliced new feature vector can adjust the dimension and distribution of the features, making it more in line with the input requirements of the video prediction model, improving the accuracy and efficiency of the model's prediction of the target position. Highlight the information useful for the prediction task and reduce the interference of redundant information, thereby improving the performance of the entire video prediction model. Description of the Drawings

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a schematic flowchart of the method for processing video image data based on the time-series large model of the present invention.

[0058] Figure 2 It is a schematic flowchart of obtaining the interpolated eigenvalue of the present invention.

[0059] Figure 3 It is a structural diagram of the video image data processing system based on the time-series large model of the present invention. Detailed implementation manners

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0061] In the drawings of the specific embodiments of the present invention, in order to better and more clearly describe the working principles of the components in the system and show the connection relationships of the various parts of the device, only the relative positional relationships between the components are clearly distinguished, and it does not constitute a limitation on the signal transmission direction, connection sequence, and the sizes, dimensions, and shapes of the various parts of the structure within the component or structure.

[0062] Embodiment 1

[0063] As Figure 1 shown, it is a flowchart of the method for processing video image data based on the time-series large model of the present invention. The method for processing video image data based on the time-series large model includes the following steps:

[0064] S1. Collect video image data and preprocess the video image data.

[0065] S11. Video image data collection.

[0066] Use a high-resolution camera (such as 4K / 8K), industrial camera, or smartphone to ensure stable frame rate and support dynamic scene capture; for special scenes (such as infrared, depth perception), multi-modal sensors (such as Kinect, LiDAR) need to be equipped.

[0067] Cover diverse scenarios according to task requirements (such as indoor / outdoor, lighting changes, occluded environments); collect different perspectives (top-down, eye-level, multi-angle) and motion patterns (uniform motion, accelerated motion, random motion).

[0068] The multi-camera system needs to perform time synchronization, spatially calibrate the devices, and eliminate lens distortion.

[0069] Extract key frames from the video or sample evenly to avoid redundant calculations. Long videos can be segmented into equal-length segments (such as 16 frames / segment) to adapt to the input length of the model.

[0070] S12. Preprocess the video image data.

[0071] Normalize the pixel values (such as [-1, 1]), adjust the resolution (such as 224x224), apply spatio-temporal augmentation to improve generalization, and preferably use augmentation methods such as random cropping, temporal flipping, and color perturbation.

[0072] S13. Temporal annotation of data and metadata recording.

[0073] Timestamp the key events in the video. Key events in the video are, for example, the start / end frames of actions. Use dense annotation for continuous actions, for example, annotate once per second.

[0074] Save the acquisition parameters: resolution, frame rate, light intensity, device model, etc.; record the scene semantic labels such as "driving in the rain" and "crowded people".

[0075] In the preferred embodiment, data quality control is required.

[0076] Anomaly detection: Automatically filter black frames and blurred frames, and remove video segments with stuttering or dropped frames; manually sample and check the annotation consistency, for example, the overlap rate IoU > 0.8.

[0077] Data balance: Statistically analyze the class distribution, and perform oversampling or synthesis on the long-tailed data.

[0078] S2. Capture the key information from the preprocessed video data to obtain the video frames at the key time points of the video data, and perform interpolation operations on the video frames at the key time points to obtain the interpolated eigenvalue.

[0079] As Figure 2 shown, it specifically includes the following steps:

[0080] S21. Extract the video frames X at the key time points of the video data S :

[0081] The video frames of the input video data, where the dimension of video frame X is T×H×W×C, where T is the number of time frames, H is the height of the video frame, W is the width of the video frame, and C is the number of channels. Each pixel value in the video frame is a real number and is stored in the real number set R.

[0082] S211. Generate a feature map.

[0083] Input the video frame X through a dimension expansion layer to obtain a feature map F with a dimension of T×H×W×C′ Z , where C′ is the number of intermediate channels.

[0084] F Z =Conv2D(X;W conv )

[0085] W conv is the dimension expansion kernel parameter, and the feature map F Z retains the original spatial dimension H×W, but the number of channels becomes C′.

[0086] The role of the dimension expansion layer is to extract features from the video frames of the video data to obtain a more representative feature map F Z .

[0087] S212. Dimension averaging.

[0088] Perform dimension averaging on the feature map F Z to obtain an averaged result F with a dimension of T×1×1×C′ avg :

[0089] For each time step t and channel C′ of the feature map F Z , calculate the averaged result F avg :

[0090] ;

[0091] The purpose of dimension averaging is to compress the information in the time dimension and only retain the global information in the channel dimension, which can reduce the computational amount and highlight the importance of each channel.

[0092] S213. Generate time-focusing weights through a data sequence recognition mechanism.

[0093] Map the C′-dimensional vector to the H×W dimension through a neural network and adjust the shape through the Reshape function operation:

[0094] ;

[0095] where W1 and W2 are weights, is the activation function. The role of the Reshape function operation is to readjust the output flattened vector into a shape that matches the spatial dimensions of the original video frame. b1 and b2 are constant terms.

[0096] The result after dimension averaging is processed through the data sequence recognition mechanism to obtain the time focus weight S, and its calculation formula is:

[0097] ;

[0098] where is the sigmoid activation function.

[0099] The sigmoid function maps the output value to the interval [0, 1], making the weight at each position represent the importance degree of that position in space.

[0100] S214, focus on the video frames at key time points.

[0101] Multiply the time focus weight S and the original video frame X element-wise to obtain the video frames at key time points:

[0102] .

[0103] represents element-wise multiplication.

[0104] In this way, the video frames at key time points are enhanced, while the video frames at non-key time points are relatively weakened.

[0105] S22, calculate the interpolated eigenvalue of the video frame X S at key time points.

[0106] To solve the problem of feature blur caused by violent movement. The displacement information of the object is obtained through the fuzzy value estimation method, and then the attention area is dynamically adjusted according to the displacement information, so that the model can better focus on the important parts in the video and maintain the accuracy of features even when the object is moving violently. This method can effectively improve the performance of video analysis tasks, especially when dealing with moving objects.

[0107] S221, obtain the central displacement information of the target area using the fuzzy value estimation method.

[0108] The fuzzy value estimation method is a method for calculating the movement of pixel points in an image, and the movement of the target area is described by calculating the displacement of pixels between adjacent frames. In the embodiment, the displacement information of the center of the target area in the video frame is obtained using the fuzzy value estimation method.

[0109] Use (△p x (i, j), △p Y(i, j)) represents the displacement vector of the pixel at the central position coordinates (i, j) of the image frame from the (t - 1)-th frame to the t-th frame, and these displacement vectors will be used to determine the blur value subsequently.

[0110] S222. Determine the blur value based on the displacement vectors obtained by the blur value estimation method.

[0111] For the feature of the pixel at each coordinate (i, j), its blur value Δp(i, j) can be adjusted according to the displacement vector:

[0112] Δp(i, j)=(Δp x (i, j), Δp Y (i, j)).

[0113] S223. Perform interpolation operation based on the blur value to obtain the interpolated feature value.

[0114] Let the video frame X at the key time point S be a two-dimensional feature map with a size of H × W (H is the height and W is the width). The feature value of the pixel at each coordinate (i, j) (0 ≤ i < H, 0 ≤ j < W) in the video frame X at the key time point S is X ZS (i, j). Calculate the interpolated feature value from the adjacent frame images through the interpolation method .

[0115] Calculate the interpolated feature value from the adjacent frame images using the interpolation method.

[0116] Let the interpolated feature value be , then:

[0117] ;

[0118] f def is the interpolation function. Let , , that is, for the video frame X at the key time point S , sample at the non-integer coordinates (u, v), then the interpolation function is:

[0119] ;

[0120] where w mn is the weight coefficient calculated according to the fractional parts of u and v, and are the integer parts of u and v respectively. This formula calculates f ZS by performing weighted summation within a certain neighborhood centered on and according to the weight coefficient w mn . defThe value of .

[0121] m and n are index variables for summation, and their value ranges are both from 0 to 1. When calculating the interpolation at non-integer coordinates (u, v), double summation of m and n in the range of 0 to 1 is performed, combined with the weight coefficient w mn And after the coordinates are rounded down The video frame value at is used to obtain the final interpolation result. m and n are used to traverse the items related to the surrounding pixels involved in the interpolation calculation.

[0122] After the above interpolation operation, the interpolated eigenvalue is obtained. , 0≤i <H,0≤j<W。

[0123] Through the above steps, the interpolated eigenvalues ​​are calculated. ,Therefore the feature blur problem caused by the violent movement of the target can be solved.

[0124] S3. Retrieve historical information from the global memory pool and fuse the interpolated feature values, perform feature transformation operations on the fused features, and obtain transformed features.

[0125] S31. Obtain historical information M from a global memory pool storing historical information.

[0126] The global memory pool here can be understood as a storage structure that stores information such as features related to past video frames for subsequent retrieval.

[0127] The global memory pool is a storage structure that stores historical information such as features related to past video frames. The historical information M is obtained through the following steps:

[0128] Locate the storage location and clarify the storage location of the global memory pool in the system, such as a specific area in the computer memory, a table in the database, or a specific storage module, or in the deep learning video processing framework, a cache area specially opened in the memory for storing historical features.

[0129] Set the search conditions and search rules according to current needs. When you want to obtain historical information within a specific time period, you need to determine it based on the timestamp of the video frame; when you want to obtain historical information with similar scenes to the current frame, you need to search based on the similarity of scene features.

[0130] Execute the search operation, and execute the search instruction in the storage location of the global memory pool according to the set search conditions. Preferably, execute the SELECT statement in the database to filter out the historical information M that meets the conditions from the storage structure. In the deep learning framework, call the corresponding function or method to implement the search.

[0131] Extract the information and retrieve the historical information M from the storage location for subsequent fusion operations with the final features.

[0132] S32. Obtain the interpolated eigenvalue at the current moment , and the interpolated eigenvalue is concatenated with the retrieved historical information M. It is expressed by the formula:

[0133] ,

[0134] where Concat represents the concatenation operation, which connects two features in a certain dimension to form a new feature vector.

[0135] Among them, assume that the interpolated eigenvalue is a feature vector with dimension d1, and the retrieved historical information M is a feature vector with dimension d2. After the fusion concatenation operation, the dimension of the fused feature F is d3 = d1 + d2.

[0136] Perform a feature transformation operation on the fused feature F, and the transformed feature F t is:

[0137] ;

[0138] The dimension of the transformed feature F t is also d3.

[0139] S4. Input the transformed feature into the video prediction model to predict the center position of the target area at the future moment.

[0140] First, calculate the center position of the current target area of the object from the transformed feature F t as :

[0141] Input the transformed feature F t into the trained object detection model. The object detection model outputs the bounding box information of the detected target area. For a two-dimensional image, the bounding box is represented by the upper-left coordinates (X min , Y min ) and the lower-right coordinates (X max , Y max ). Then, the center position of the target area is obtained through the following formula as the current position of the target area :

[0142] ;

[0143] ;

[0144] Then, the transformed feature Ft Expressed as a column vector ,

[0145] After the transformed feature F t is input into the neural network, an adjustment amount is obtained and :

[0146]

[0147] where B a is a bias vector and Q a is a weight matrix;

[0148] Expressed as:

[0149] ;

[0150] , obtain ;

[0151] Let: ;

[0152] ;

[0153] where is the j-th matrix element of the first row of the weight matrix Q a , is the j-th matrix element of the second row of the weight matrix Q a , is the j-th eigenvector after the transformed feature F t is expressed as a column vector

[0154] Then DX and DY represent the movement values of the predicted center position of the output of the fully connected neural network layer

[0155] The predicted value of the center position of the target area at the next moment is:

[0156] ;

[0157] ;

[0158] Through the above steps, the center movement value of the target area predicted by the fully connected neural network is combined with the center position of the target area at the previous moment to predict the center position of the target area at the next moment

[0159] Embodiment 2

[0160] Such as Figure 3As shown in the figure, it is the structural diagram of the video data processing system for implementing the video image data processing method based on the sequential large model of the present invention. The video data processing system includes: a data acquisition and preprocessing module, a key information capture module, an information fusion and transformation module, and a video prediction module.

[0161] The data acquisition and preprocessing module is used to acquire video image data and preprocess the video image data.

[0162] The key information capture module is used to capture key information from the preprocessed video data, obtain the video frames at the key time points of the video data, and perform interpolation operations on the video frames at the key time points to obtain the interpolated eigenvalues.

[0163] The information fusion and transformation module is used to retrieve historical information from the global memory pool and fuse the interpolated eigenvalues, and perform feature transformation operations on the fused features to obtain the transformed features.

[0164] The video prediction module is used to input the transformed features into the video prediction model to predict the central position of the target area at future moments.

[0165] In a preferred embodiment, the data acquisition and preprocessing module includes: a video image acquisition unit and a preprocessing unit. The video image acquisition unit is responsible for acquiring video image data from various video sources (such as cameras, video files, etc.). The video image acquisition unit includes a hardware interface adaptation module for connecting different types of video input devices to ensure stable data acquisition. The preprocessing unit preprocesses the acquired video image data, including denoising the image to remove noise interference introduced during the acquisition process; size normalization to uniformly adjust images of different resolutions to a size suitable for subsequent processing; and color space conversion to convert the image data into a suitable color space representation.

[0166] The key information capture module captures key information from the preprocessed video data, obtains the video frames at the key time points of the video data, and performs interpolation operations on the video frames at the key time points to obtain the interpolated eigenvalues.

[0167] The information fusion and transformation module includes a storage unit, a fusion unit, and a feature transformation unit. The storage unit is used to store historical information, which can be the key features of previously processed video data, etc. The storage unit uses an efficient retrieval mechanism to quickly retrieve relevant historical information according to current requirements. The fusion unit fuses the historical information retrieved from the storage unit with the interpolated eigenvalues obtained by the key information capture module and performs interpolation operations. The fusion method can be simple feature splicing or weighted fusion to highlight the importance of different features. The feature transformation unit performs feature transformation operations on the fused features to make them more suitable for subsequent prediction tasks.

[0168] The video prediction module predicts the position information of the center of the target area in the video by learning and analyzing the input features. The prediction result can be the coordinates of the center of the target area or other representation forms related to the center of the target area, such as the position and size of the bounding box, etc.

[0169] In a preferred embodiment, the video data processing system of the video image data processing method based on the temporal large model of the present invention further includes a system control and coordination module, which is responsible for the process control between the various modules of the entire system to ensure that data is sequentially transmitted and processed between the modules in order. For example, after the data acquisition is completed, the preprocessing module is triggered to start working; after the preprocessing is completed, the key information capture module is started, etc.; at the same time, the parameters involved in each module in the system are managed, such as the data acquisition frequency, the parameters of the preprocessing operation, the hyperparameters of the feature extraction and prediction models, etc., and these parameters are adjusted and optimized according to the actual application scenario and requirements.

[0170] In one embodiment, a computer device is also provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0171] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0172] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0173] Those skilled in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., and are not limited thereto.

[0174] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0175] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for processing video image data based on a large temporal model, characterized in that, It includes the following steps: Step S1: Collect video image data and preprocess the video image data; Step S2: Capture key information from the preprocessed video data to obtain video frames at key time points of the video data, and perform interpolation operations on the video frames at key time points to obtain interpolated eigenvalues; Input the original video frame X of the video data. The dimension of the original video frame X is T×H×W×C, where T is the number of time frames, H is the height of the video frame, W is the width of the video frame, and C is the number of channels; Multiply the time focusing weight S and the original video frame X element by element to obtain the video frame at the key time point : ; Among them, represents element-wise multiplication; Step S3: Retrieve historical information from the global memory pool and fuse the interpolated eigenvalues, and perform feature transformation operations on the fused features to obtain transformed features; Interpolated eigenvalue Perform a concatenation operation with the retrieved historical information M: , Among them, Concat represents the concatenation operation; Interpolated eigenvalue is an eigenvector with dimension d1, and the retrieved historical information M is an eigenvector with dimension d2. After the fusion and concatenation operation, the dimension of the fused feature F is d3 = d1 + d2; perform a feature transformation operation on the fused feature F to obtain the transformed feature F t : ; Step S4: Input the transformed features into the video prediction model to predict the center position of the target area at future times.

2. The method for processing video image data based on a time series large model according to claim 1, wherein In the step S2, assume that the video frame X at the key time point S The eigenvalue of the pixel at each coordinate (i, j) in is X ZS (i, j), and the interpolated eigenvalue is calculated from adjacent frame images using the interpolation method : ; where, f def is an interpolation function, and represents the displacement vector of the pixel at the coordinates (i, j) of the image frame from the (t - 1)-th frame to the t-th frame; Let , , for the video frame X at the key time point S , sampling at non-integer coordinates (u, v), the interpolation function is: ; where, w mn is a weight coefficient calculated according to the fractional parts of u and v, and are the floor functions of u and v respectively, and m and n are index variables for summation.

3. The method for processing video image data based on a time series large model according to claim 1, wherein In the step S4, the transformed feature F t is input into the trained object detection model, and the object detection model outputs the bounding box information of the detected target region. For a two-dimensional image, the bounding box information is represented by the upper left coordinates (X min , Y min ) and the lower right coordinates (X max , Y max ), and the current position of the center position of the target region is obtained through the following formula : ; 。 4. The method for processing video image data based on a time series large model according to claim 3, wherein Represent the transformed feature F t as a column vector , where the dimension of the column vector is d 3, After inputting the transformed feature F t into the neural network, an adjustment amount and are obtained ; Among them, B a is the bias vector, and Q a is the weight matrix; Expressed as: ; Expressed as: ; Let: ; ; Among them, is the j-th matrix element of the first row of the weight matrix Q a ; is the j-th matrix element of the second row of the weight matrix Q a ; is the transformed feature F t which is the j-th eigenvector after being represented as a column vector, B a,x is the x-component of the bias vector B a,y is the y-component of the bias vector B.

5. The method for processing video image data based on a time series large model according to claim 4, wherein The central position of the predicted target area at the next moment t is as follows: ; 。 6. The method for processing video image data based on a time-series large model according to claim 1, wherein In the step S2, the video frame X with the dimension of T×H×W×C is input into the dimension expansion layer to obtain the feature map F with the dimension of T×H×W×C'. Z , for the feature map F Z perform dimension averaging to obtain the averaged result F with the dimension of T×1×1×C'. avg , where T is the number of time frames, H is the height of the video frame, W is the width of the video frame, C is the number of channels, and C' is the number of intermediate channels; Process the mean value result F through the data sequence recognition mechanism avg to obtain the time focusing weight S.

7. A video image data processing system based on a temporal large model, which is used to implement the video image data processing method based on the temporal large model according to any one of claims 1-6, and is characterized in that, It includes: Data acquisition and preprocessing module, key information capture module, information fusion and transformation module, video prediction module; The data acquisition and preprocessing module is used to collect video image data and preprocess the video image data; The key information capture module is used to capture key information from the preprocessed video data to obtain video frames at key time points of the video data, and perform interpolation operations on the video frames at key time points to obtain interpolated eigenvalues; The information fusion and transformation module is used to retrieve historical information from the global memory pool and fuse the interpolated eigenvalues, and perform feature transformation operations on the fused features to obtain transformed features; The video prediction module is used to input the transformed features into the video prediction model to predict the center position of the target area at future times.

Citation Information

Patent Citations

  • Target detection method and device and electronic equipment

    CN116665181A

  • DBT reconstruction method based on three-dimensional coordinate displacement, medium and electronic equipment

    CN119206097A