Video frame extraction intelligent identification acceleration method

By decoding and editing image frames in the GPU video memory and optimizing the video frame extraction process, the problems of high bandwidth usage and CPU limitations in video frame extraction intelligent recognition are solved, and efficient and real-time video recognition acceleration is achieved.

CN120751132APending Publication Date: 2025-10-03ZAOZHUANG POWER SUPPLY COMPANY OF STATE GRID SHANDONG ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510832188.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing technology has problems in the process of video frame extraction and intelligent recognition, such as high bandwidth resource usage, CPU architecture limitations, and underutilized GPU parallel processing capabilities, resulting in low efficiency and insufficient real-time performance.

Method used

The GPU hardware decoder is used to decode the video stream, and the image frames are edited and encoded directly in the GPU memory to reduce data transmission. The video stream is divided into four steps: image frame extraction, editing, forwarding and analysis through a structured OD/OA separation mode to optimize the processing flow.

Benefits of technology

It significantly improves the efficiency of video frame extraction, reduces dependence on bandwidth resources, and improves system operating costs and the real-time and accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751132A_ABST
    Figure CN120751132A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent recognition acceleration method for video frame extraction. The method comprises the following steps: step 1, image frame extraction; step 2, image frame editing; step 3, image frame forwarding; according to the method, the video frame extraction efficiency is remarkably improved, and the intelligent recognition process is accelerated. The dependence on bandwidth resources is reduced, and the system operation cost is reduced. By optimizing the processing flow, the accuracy and real-time performance of video intelligent identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video recognition technology, and in particular to a video frame extraction intelligent recognition acceleration method. Background Art

[0002] There are currently two main ways to implement video frame extraction and intelligent recognition. One is to directly pull the video stream from the intelligent recognition model to the model for decoding and recognition based on the SDK. The other is to extract video images at a slower speed of more than 6 seconds based on the CPU and send them to the intelligent recognition model for analysis and recognition. The first method will occupy more bandwidth resources due to the large volume of video stream data. The second method is extremely cost-effective for image decoding and frame extraction due to CPU architecture and resource usage issues.

[0003] Based on the physical structure of the integrated graphics, including the CPU core, the last-level cache LLC, the Slice and Un-Slice structure of the GPU, and the MFF module for codecs, it can accelerate video decoding and image processing.

[0004] The optimal solution to slowing down AI video analysis is to structure the video stream before matching it with various algorithm models. Enhanced streaming, based on unified video, innovatively introduces a structured OD / OA separation mode. This leverages powerful GPU resources to perform video frame extraction (OD), converting the video stream into images. After user-defined video thinning, these images are transmitted to the AI ​​platform via subscription notifications. The AI ​​platform then performs model matching and analysis on the images, accelerating AI video recognition efficiency by approximately five times without increasing hardware investment.

[0005] The existing technology has the following major defects in video frame extraction intelligent recognition: 1. SDK-based direct video stream recognition requires processing large amounts of video stream data, resulting in high bandwidth usage and increased costs.

[0006] 2. The CPU-based video frame extraction method has low frame extraction efficiency due to CPU architecture limitations and resource usage issues, and cannot meet real-time requirements.

[0007] 3. Existing technologies fail to fully utilize the parallel processing capabilities of GPUs, resulting in limited video decoding and image processing speeds. Summary of the Invention

[0008] The purpose of the present invention is to provide a video frame extraction intelligent recognition acceleration method to solve the above technical problems.

[0009] To achieve the above-mentioned purpose, the present invention adopts the following technical solutions: A video frame extraction intelligent recognition acceleration method comprises the following steps: Step 1: Image frame extraction: Obtain the video stream by pushing or pulling the stream. According to the encapsulation protocol of the video stream, demultiplex the socket data of the video stream, obtain the nalu data in the network data packet, input the nalu data into the GPU hardware decoder, and obtain the YUV image frame. Step 2: Image frame editing: Directly edit the YUV image frame in the GPU memory, align the GPU memory data according to the API input parameters, scale the image, and input the scaled YUV image into the GPU hardware encoder to obtain a JPG image. Step 3: Image frame forwarding: copy the JPG image from the GPU memory to the host memory, and use the HTTP protocol to send the JPG image to the designated artificial intelligence platform; Step 4: Image frame analysis: The AI ​​platform receives the JPG image through the HTTP protocol and performs AI analysis on the image using the corresponding model.

[0010] Compared with the prior art, the present invention has the following advantages: the present invention significantly improves the efficiency of video frame extraction and accelerates the intelligent recognition process.

[0011] It reduces the dependence on bandwidth resources and lowers the system operation cost.

[0012] By optimizing the processing flow, the accuracy and real-time performance of video intelligent recognition are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flow chart of the steps of the present invention. DETAILED DESCRIPTION

[0014] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0015] A video frame extraction intelligent recognition acceleration method comprises the following steps: Step 1: Image frame extraction: Obtain the video stream by pushing or pulling the stream. According to the encapsulation protocol of the video stream, demultiplex the socket data of the video stream, obtain the nalu data in the network data packet, input the nalu data into the GPU hardware decoder, and obtain the YUV image frame. Step 2: Image frame editing: Directly edit the YUV image frame in the GPU video memory, align the GPU memory data according to the API input parameters, scale the image, and input the scaled YUV image into the GPU hardware encoder to obtain a JPG image. This process greatly improves image processing efficiency because it avoids copying between the GPU video memory and the host memory.

[0016] Step 3: Image frame forwarding: copy the JPG image from the GPU memory to the host memory, and use the HTTP protocol to send the JPG image to the designated artificial intelligence platform; Step 4: Image frame analysis: The AI ​​platform receives the JPG image through the HTTP protocol and performs AI analysis on the image using the corresponding model.

[0017] This invention uses a GPU hardware decoder to decode video streams, improving decoding efficiency. Image frame editing and encoding are performed directly in the GPU video memory, reducing data transmission and increasing processing speed. The introduction of a structured video OD / OA separation mode divides video stream analysis into steps such as image frame extraction, editing, forwarding, and analysis, optimizing the processing flow.

[0018] Edge IoT agents and image recognition algorithms are deployed at charging stations and connected to the streaming media platform. The platform leverages capabilities such as visual recognition services, device catalog subscription and notification management, alarm subscription and notification services, video recognition result overlay services, video confluence services, tagging / query services, and device pre-position management services to comprehensively coordinate charging station video image recognition, identify event alerts, overlay events on video screens, and quickly search and access historical videos.

[0019] (1) Visual recognition service: The streaming media middleware is linked to the edge IoT agent to start and access the video image recognition model, including vehicle entry and exit recognition, charging station environmental sanitation recognition, charging vehicle queue recognition, etc.

[0020] (2) Device catalog subscription and notification management service: The streaming media middle platform is connected to the edge IoT agent device of the charging station to subscribe to and display the accessed video channels, scale, additions, deletions and other information.

[0021] (3) Alarm subscription and notification service: The streaming media middle platform is connected to the edge IoT agent device of the charging station to subscribe to messages and provide notification reminders for identified abnormal events at the charging station.

[0022] (4) Video recognition result overlay service: The streaming media center will overlay and display abnormal events at the charging station in the video stream.

[0023] (5) Tag labeling / query service: The streaming media platform labels the historical videos and event videos stored in the charging station and performs quick query and display based on the tags.

[0024] (6) Equipment preset position management service: remotely manage and update the preset positions of charging station video terminals.

[0025] The above is a preferred embodiment of the present invention. For ordinary technicians in this field, based on the teachings of the present invention, without departing from the principles and spirit of the present invention, changes, modifications, substitutions and variations made to the implementation methods are still within the scope of protection of the present invention.

Claims

1. A video frame extraction intelligent recognition acceleration method, characterized in that: The steps include: Step 1 Image frame extraction: Obtain through push or pull streaming Video stream, according to the video stream encapsulation protocol, demultiplex the socket data of the video stream, obtain the nalu data in the network data packet, input the nalu data into the GPU hardware decoder, and obtain the YUV image frame; Step 2: Image frame editing: Directly edit the YUV image frame in the GPU memory, align the GPU memory data according to the API input parameters, scale the image, and input the scaled YUV image into the GPU hardware encoder to obtain a JPG image. Step 3: Image frame forwarding: copy the JPG image from the GPU memory to the host memory, and use the HTTP protocol to send the JPG image to the designated artificial intelligence platform; Step 4: Image frame analysis: The AI ​​platform receives the JPG image through the HTTP protocol and performs AI analysis on the image using the corresponding model.