Video data processing method and device

Through machine learning technology, the automatic identification and extraction of keyframes in videos has been solved, and the problem that the existing technology cannot effectively adapt to the diversity and complexity of video content is achieved, and more efficient and accurate video analysis is achieved.

CN119946326APending Publication Date: 2025-05-06JIASHAN LITONG INFORMATION & TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411969335.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing video frame extraction methods cannot effectively adapt to the diversity and complexity of video content, resulting in the loss of key information or excessive redundant information.

Method used

By obtaining the pending video data for pre-processing, the features in the image frame data are extracted, the object recognition model is established, and the keyframes are automatically recognized and extracted using machine learning technology.

Benefits of technology

It realizes more precise extraction of keyframes, improves the speed and accuracy of video analysis, reduces unnecessary frame processing, and can handle more video streams under the same hardware conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946326A_ABST
    Figure CN119946326A_ABST
Patent Text Reader

Abstract

The invention discloses a video data processing method and device. The method comprises the following steps: acquiring to-be-processed video data, and preprocessing the to-be-processed video data to generate image frame data; performing feature extraction on the image frame data, and identifying an object in the image frame data based on image features; establishing an association relationship between the image frame data and the identified object, and generating training data; performing machine learning model training based on the training data to obtain a machine learning model for object recognition; and performing object identification on the video data by using the trained machine learning model. The video analysis speed is improved, the accuracy of video monitoring and analysis is improved, and unnecessary frame processing is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a key frame extraction technology in surveillance videos, and more particularly to a method and device for processing video data. Background Art

[0002] There are two main existing methods for video frame extraction: one is based on a fixed time interval, and the other is based on scene change detection. Among them, the method based on a fixed time interval is the simplest and most intuitive frame extraction method, that is, extracting a frame at a fixed time interval. For example, extracting one frame per second, or one frame per minute. The method based on scene change detection extracts frames based on changes in video content. For example, when the scene in the video changes significantly, such as the appearance or disappearance of an object, or the brightness of the scene changes, a frame is extracted.

[0003] The main disadvantage of the method based on fixed time intervals is that it cannot adapt to the diversity and complexity of video content. For some videos with fast or slow content changes, the fixed time interval frame extraction method may lead to the loss of key information or excessive redundant information. The main disadvantage of the method based on scene change detection is that it is often difficult to accurately define and detect scene changes, and it is easily affected by noise. For example, due to changes in lighting, shadows and other factors, it may cause false detection or missed detection. Summary of the invention

[0004] The present application provides a method and device for processing video data to at least solve the above technical problems existing in the prior art.

[0005] According to a first aspect of the present application, a method for processing video data is provided, comprising:

[0006] Acquire video data to be processed, preprocess the video data to be processed, and generate image frame data;

[0007] Extracting features from the image frame data, and identifying objects in the image frame data based on image features;

[0008] Establishing an association relationship between the image frame data and the identified object to generate training data;

[0009] Performing machine learning model training based on the training data to obtain a machine learning model for object recognition;

[0010] Use the trained machine learning model to perform object recognition on video data.

[0011] In some executable embodiments, the preprocessing of the to-be-processed video data includes:

[0012] The size and format of the video frames in the video data to be processed are adjusted to adapt to feature detection of the object.

[0013] In some executable embodiments, the preprocessing of the to-be-processed video data includes:

[0014] A full video frame extraction method is performed on the video data to be processed, a timestamp is generated for each video frame, and a corresponding relationship between the timestamp and the video frame is established.

[0015] In some executable embodiments, the performing machine learning model training based on the training data to obtain a machine learning model for object recognition includes:

[0016] A decision tree is generated using the random forest algorithm. Each decision tree is trained on a random subset of the training data. Based on the object information, it is predicted whether each video frame is a key frame.

[0017] Use the test data in the training data to evaluate the precision, recall, and F1 score of the machine learning model and obtain the evaluation results;

[0018] Based on the evaluation results, the test data is adjusted, the parameters of the machine learning model are adjusted, and the model is retrained to obtain the adjusted machine learning model.

[0019] In some executable embodiments, extracting features from the image frame data and identifying objects in the image frame data based on image features includes:

[0020] The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame;

[0021] Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information;

[0022] The normalized local feature information is input into the activation function to determine the object in the video frame.

[0023] According to a second aspect of the present application, a video data processing device is provided, comprising:

[0024] A first generating unit is used to obtain the video data to be processed, pre-process the video data to be processed, and generate image frame data;

[0025] A first recognition unit, configured to extract features from the image frame data and recognize objects in the image frame data based on image features;

[0026] A second generating unit, used to establish an association relationship between the image frame data and the identified object to generate training data;

[0027] A training unit, used to perform machine learning model training based on the training data to obtain a machine learning model for object recognition;

[0028] The second recognition unit is used to perform object recognition on the video data using the trained machine learning model.

[0029] In some executable embodiments, the first generating unit is further configured to:

[0030] The size and format of the video frames in the video data to be processed are adjusted to adapt to feature detection of the object.

[0031] In some executable embodiments, the first generating unit is further configured to:

[0032] A full video frame extraction method is performed on the video data to be processed, a timestamp is generated for each video frame, and a corresponding relationship between the timestamp and the video frame is established.

[0033] In some executable embodiments, the training unit is further used for:

[0034] A decision tree is generated using the random forest algorithm. Each decision tree is trained on a random subset of the training data. Based on the object information, it is predicted whether each video frame is a key frame.

[0035] Use the test data in the training data to evaluate the precision, recall, and F1 score of the machine learning model and obtain the evaluation results;

[0036] Based on the evaluation results, the test data is adjusted, the parameters of the machine learning model are adjusted, and the model is retrained to obtain the adjusted machine learning model.

[0037] In some executable embodiments, the first identification unit is further configured to:

[0038] The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame;

[0039] Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information;

[0040] The normalized local feature information is input into the activation function to determine the object in the video frame.

[0041] According to a third aspect of the present application, an electronic device is provided, including:

[0042] at least one processor; and

[0043] a memory communicatively connected to the at least one processor; wherein,

[0044] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the video data processing method.

[0045] According to a fourth aspect of the present application, a non-temporary computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of the method for processing video data.

[0046] The video data processing method and device, electronic device, and storage medium of the present application can better adapt to video content of different types and styles by automatically learning and generating frame extraction rules. By optimizing the time interval and frequency of frame extraction, key frames can be extracted more accurately, thereby improving the speed of video analysis. The embodiments of the present application can more accurately identify and extract key frames, thereby improving the accuracy of video monitoring and analysis, reducing unnecessary frame processing, and can process more video streams under the same hardware conditions, or when processing the same number of video streams, lower computing resources can be used, thereby improving the performance and efficiency of the entire system. Based on machine learning technology, the embodiments of the present application can easily introduce new features and rules to adapt to new video content and application scenarios, and improve the applicability and scalability of the technical solution.

[0047] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become readily understood. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, wherein:

[0049] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0050] Figure 1 A schematic diagram showing a flow chart of a method for processing video data according to an embodiment of the present application;

[0051] Figure 2 A schematic diagram showing the composition structure of a video data processing device according to an embodiment of the present application is shown;

[0052] Figure 3 A schematic diagram of the structure of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0053] In order to make the purpose, features, and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0054] Figure 1 A schematic diagram showing a process flow of a method for processing video data according to an embodiment of the present application is shown. Figure 1 As shown, the method for processing video data in the embodiment of the present application includes the following processing steps:

[0055] Step 101: Obtain video data to be processed, pre-process the video data to be processed, and generate image frame data.

[0056] The technical solution of the embodiment of the present application is applied to various surveillance video analysis scenarios, for example, it can be applied to the monitoring of personnel identification or specific scenes, for example, it can monitor whether there are personnel in a specific area, or monitor the environmental conditions of a specific area such as fire conditions, public security incidents, etc. The embodiment of the present application uses a video key frame extraction solution based on machine learning, and generates certain rules according to the frequency, time and duration of appearance of objects and characters in the video through machine learning technology, and then uses these rules to optimize the time interval and frequency of frame extraction, thereby improving the recognition efficiency while ensuring the accuracy of the recognition of related objects in the video.

[0057] In an embodiment of the present application, the size and format of the video frame in the video data to be processed are adjusted to adapt to the feature detection of the object. Specifically, the input of the surveillance video data to be processed is received, and the input video data is preprocessed, including adjusting the size and format of the video to adapt to the subsequent object detection.

[0058] In the embodiment of the present application, the video data to be processed is subjected to full video frame extraction, a timestamp is generated for each video frame, and a correspondence between the timestamp and the video frame is established. Specifically, a full frame extraction scheme is adopted for the video data to be processed, that is, each frame is preprocessed and related objects are identified, and a timestamp is generated for each frame, forming a correspondence between the timestamp and the video frame.

[0059] Step 102: extract features from the image frame data, and identify objects in the image frame data based on image features.

[0060] In an embodiment of the present application, the object to be identified in the video image frame is a person or other object. Specifically, the embodiment of the present application uses yolo to detect the sampled video data. The Yolo network structure can use yolov8 to identify the objects and their number in the video image frame. Specifically, the image information features of each video frame of the sampled video data are extracted through the backbone network (Backbone), and the image information features are convoluted to obtain the local feature information of each video frame; the local feature information of each video frame is normalized to reduce the covariate shift of the local feature information; the normalized local feature information is input into the activation function to determine the object in the video frame. Among them, Backbone uses a series of convolution and deconvolution layers to extract features, and uses residual connections and bottleneck structures to reduce the size of the network and improve performance. The C2f module is used to extract feature information, which improves the ability to enhance feature information. And the multi-scale feature fusion technology is used to fuse the feature maps from different stages of Backbone to enhance the feature representation capability.

[0061] Step 103: establishing an association relationship between the image frame data and the identified object to generate training data.

[0062] In the embodiment of the present application, the characters in the video frames and the detection results are summarized to form data in a unified format, including the timestamp of each frame and the information of the objects and characters appearing in the frame. These training data will be used as input for the machine learning model.

[0063] Specifically, the data collected after video detection can be divided into training data and test data. The data is divided according to the ratio of 70% training data and 30% test data, that is, 70% of the collected data is used as training data, and the remaining data is used as test data. The collected data can be set in proportion as needed, and the above example is only for reference.

[0064] Step 104: Perform machine learning model training based on the training data to obtain a machine learning model for object recognition.

[0065] In the embodiment of the present application, a decision tree is generated using a random forest algorithm, and each decision tree is trained on a random subset of the training data; based on the information of the object, it is predicted whether each video frame is a key frame; the precision, recall rate and F1 score of the machine learning model are evaluated using the test data in the training data to obtain an evaluation result; based on the evaluation result, the test data is adjusted, the parameters of the machine learning model are adjusted, and the model is retrained to obtain an adjusted machine learning model. The evaluation parameters of the machine learning model are not limited to precision, recall and F1 score, but can also be other evaluation parameters.

[0066] In the embodiment of the present application, a machine learning model is trained using the aggregated data, and the trained machine learning model is applied to video detection, and the time interval and frequency of frame extraction can be dynamically adjusted. At the same time, the machine learning model is optimized and adjusted according to the effect of the machine learning model in actual application.

[0067] Step 105: Use the trained machine learning model to perform object recognition on the video data.

[0068] In the embodiment of the present application, after the model training and optimization are completed, the machine learning model can be applied to actual video detection, the time interval and frequency of frame extraction can be dynamically adjusted, and the model optimization and adjustment can continue based on the effect of the model in actual application.

[0069] The embodiment of the present application uses a machine learning model to automatically identify and extract key frames, which greatly reduces the number of video frames that need to be processed, thereby improving the efficiency of video processing. Compared with the traditional frame extraction scheme, the present invention can process a large amount of video data more efficiently. Through intelligent key frame extraction, the embodiment of the present application can greatly reduce the load of the hardware, so that the hardware can focus more on processing more important tasks. At the same time, by reducing the amount of data to be processed, the pressure of video processing on the hardware can also be reduced, so that low-performance devices can process more videos. By automatically extracting key frames, a more accurate and representative video summary can be generated to help users quickly understand the content of the video. At the same time, by only saving key frames, the storage space requirement for video data can be greatly reduced. Through model verification and optimization, the embodiment of the present application can continuously improve the performance of the model to better adapt it to actual application needs. In addition, as more data is processed, the performance of the model will continue to improve over time. The embodiment of the present application can be applied to various scenarios that require video analysis, such as security monitoring, behavior recognition, event detection, etc. The embodiment of the present application improves the efficiency of video processing, improves the processing power of hardware, provides better video summaries, saves storage space, and shows great commercial value and social value.

[0070] Figure 2 FIG. 4 shows a schematic diagram of the structure of a video data processing device according to an embodiment of the present application. Figure 2 As shown, the video data processing device of the embodiment of the present application includes:

[0071] The first generating unit 20 is used to obtain the video data to be processed, pre-process the video data to be processed, and generate image frame data;

[0072] A first recognition unit 21 is used to extract features from the image frame data and recognize objects in the image frame data based on image features;

[0073] A second generating unit 22, used to establish an association relationship between the image frame data and the identified object to generate training data;

[0074] A training unit 23, configured to perform machine learning model training based on the training data to obtain a machine learning model for object recognition;

[0075] The second recognition unit 24 is used to perform object recognition on the video data using the trained machine learning model.

[0076] In some executable embodiments, the first generating unit 20 is further configured to:

[0077] The size and format of the video frames in the video data to be processed are adjusted to adapt to feature detection of the object.

[0078] In some executable embodiments, the first generating unit 20 is further configured to:

[0079] A full video frame extraction method is performed on the video data to be processed, a timestamp is generated for each video frame, and a corresponding relationship between the timestamp and the video frame is established.

[0080] In some executable embodiments, the training unit 23 is further configured to:

[0081] A decision tree is generated using the random forest algorithm. Each decision tree is trained on a random subset of the training data. Based on the object information, it is predicted whether each video frame is a key frame.

[0082] Use the test data in the training data to evaluate the precision, recall, and F1 score of the machine learning model and obtain the evaluation results;

[0083] Based on the evaluation results, the test data is adjusted, the parameters of the machine learning model are adjusted, and the model is retrained to obtain the adjusted machine learning model.

[0084] In some executable embodiments, the first identification unit 21 is further configured to:

[0085] The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame;

[0086] Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information;

[0087] The normalized local feature information is input into the activation function to determine the object in the video frame.

[0088] In an exemplary embodiment, the aforementioned units and the like may be implemented by one or more central processing units (CPU), graphics processing units (GPU), application specific integrated circuits (ASIC), DSPs, programmable logic devices (PLD), complex programmable logic devices (CPLD), field programmable gate arrays (FPGA), general-purpose processors, controllers, microcontrollers (MCU), microprocessors, or other electronic components.

[0089] Regarding the device in the above embodiment, the specific manner in which each module and unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0090] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.

[0091] Figure 3 8 is a schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present application. Figure 3As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0092] A number of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a data processing transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0093] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a processing method for video data. For example, in some embodiments, the processing method for video data may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the processing method for video data described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the processing method for video data in any other appropriate manner (e.g., by means of firmware).

[0094] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0095] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0096] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0097] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0098] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0099] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0100] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in this application can be achieved, and this document is not limited here.

[0101] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0102] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for processing video data, characterized in that: The method comprises: Acquire video data to be processed, preprocess the video data to be processed, and generate image frame data; Extracting features from the image frame data, and identifying objects in the image frame data based on image features; Establishing an association relationship between the image frame data and the identified object to generate training data; Performing machine learning model training based on the training data to obtain a machine learning model for object recognition; Use the trained machine learning model to perform object recognition on video data.

2. The method for processing video data according to claim 1, characterized in that: The preprocessing of the video data to be processed includes: The size and format of the video frames in the video data to be processed are adjusted to adapt to feature detection of the object.

3. The method for processing video data according to claim 1, characterized in that: The preprocessing of the video data to be processed includes: A full video frame extraction method is performed on the video data to be processed, a timestamp is generated for each video frame, and a corresponding relationship between the timestamp and the video frame is established.

4. The method for processing video data according to claim 1, characterized in that: The performing of machine learning model training based on the training data to obtain a machine learning model for object recognition includes: A decision tree is generated using the random forest algorithm. Each decision tree is trained on a random subset of the training data. Based on the object information, it is predicted whether each video frame is a key frame. Use the test data in the training data to evaluate the precision, recall, and F1 score of the machine learning model and obtain the evaluation results; Based on the evaluation results, the test data is adjusted, the parameters of the machine learning model are adjusted, and the model is retrained to obtain the adjusted machine learning model.

5. The method for processing video data according to claim 1, characterized in that: The extracting features from the image frame data and identifying objects in the image frame data based on image features includes: The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame; Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information; The normalized local feature information is input into the activation function to determine the object in the video frame.

6. A video data processing device, characterized in that: The device comprises: A first generating unit is used to obtain the video data to be processed, pre-process the video data to be processed, and generate image frame data; A first recognition unit, configured to extract features from the image frame data and recognize objects in the image frame data based on image features; A second generating unit, used to establish an association relationship between the image frame data and the identified object to generate training data; A training unit, used to perform machine learning model training based on the training data to obtain a machine learning model for object recognition; The second recognition unit is used to perform object recognition on the video data using the trained machine learning model.

7. The video data processing device according to claim 6, characterized in that: The first generating unit is further configured to: The size and format of the video frames in the video data to be processed are adjusted to adapt to feature detection of the object.

8. The video data processing device according to claim 6, characterized in that: The first generating unit is further configured to: A full video frame extraction method is performed on the video data to be processed, a timestamp is generated for each video frame, and a corresponding relationship between the timestamp and the video frame is established.

9. The video data processing device according to claim 6, characterized in that: The training unit is further used for: A decision tree is generated using the random forest algorithm. Each decision tree is trained on a random subset of the training data. Based on the object information, it is predicted whether each video frame is a key frame. Use the test data in the training data to evaluate the precision, recall, and F1 score of the machine learning model and obtain the evaluation results; Based on the evaluation results, the test data is adjusted, the parameters of the machine learning model are adjusted, and the model is retrained to obtain the adjusted machine learning model.

10. The video data processing device according to claim 6, characterized in that: The first identification unit is further used for: The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame; Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information; The normalized local feature information is input into the activation function to determine the object in the video frame.