Heterogeneous equipment-oriented low-coupling intelligent video analysis framework
By designing a low-coupled intelligent video analysis framework for heterogeneous devices, efficient adaptation and decoupling of video analysis systems in heterogeneous hardware environments are achieved, compatibility and scalability problems are solved, multi-scenario application needs are met, and development and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510234609.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-18
AI Technical Summary
The existing video analysis systems have problems such as poor compatibility, high development and maintenance costs, and difficulty in flexibly coping with the needs of different scenarios in heterogeneous hardware environments. The coupling degree of traditional frameworks is too high, resulting in low system scalability and efficiency.
Design a low-coupled intelligent video analysis framework for heterogeneous devices, including video preprocessing module, video inference analysis module, post-processing module and OSD module. Through modular design and dynamic adaptation mechanism, it supports the unified processing of multiple hardware-accelerated devices and model formats, and realizes the decoupling of preprocessing, inference and post-processing.
It improves the compatibility and efficiency of video analysis systems in heterogeneous equipment environments, reduces development and maintenance costs, supports multi-scenario application requirements, and provides efficient and flexible video analysis solutions.
Smart Images

Figure CN120339897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly relates to a low-coupling intelligent video analysis framework for heterogeneous devices. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technologies, intelligent video analysis has been increasingly widely used in fields such as smart cities, industrial quality inspection, and autonomous driving. However, in actual deployment, video analysis systems face huge challenges brought by heterogeneous hardware environments. Different devices such as CPUs, GPUs, and NPUs have significant differences in computing power, memory architecture, instruction sets, etc., resulting in complex adaptation and optimization required for video analysis tasks during cross-platform deployment. In addition, the diversity of video data in terms of resolution, encoding format, frame rate, etc., and the complexity of analysis tasks (such as object detection, behavior recognition, etc.) further increase the difficulty of system design.
[0003] Traditional video analysis frameworks are usually optimized for specific hardware platforms or task scenarios and lack general support for heterogeneous devices, resulting in poor system scalability and high development and maintenance costs. At the same time, the coupling degree of existing frameworks in preprocessing, inference, postprocessing, etc. is too high, making it difficult to flexibly respond to the changing needs of different scenarios. These problems seriously restrict the large-scale application and industrial development of intelligent video analysis technologies.
[0004] Therefore, there is an urgent need for a low-coupling intelligent video analysis framework that can efficiently adapt to heterogeneous devices and support modular expansion to solve the compatibility and efficiency problems of video analysis systems in heterogeneous device environments and meet the application requirements of multiple scenarios. Summary of the Invention
[0005] The purpose of the present invention is to provide a low-coupling intelligent video analysis framework for heterogeneous devices to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A low-coupling intelligent video analysis framework for heterogeneous devices, comprising:
[0007] A video preprocessing module, which dynamically adjusts the video stream format according to the input requirements of the inference model for efficient adaptation of video input;
[0008] A video inference analysis module, which performs unified processing of multiple model formats and heterogeneous devices through inference combining model conversion and hardware accelerator adaptation;
[0009] A postprocessing module, which performs result parsing and structured processing according to the postprocessing logic template, and the postprocessing logic template is matched with the dynamic loading of inference results;
[0010] An OSD module, which is used to superimpose the inference result on the video stream and generate a visual output.
[0011] Preferably, the video preprocessing module adopts a real-time parameter adjustment technology based on model requirements, integrates the multi-format decoding capabilities of OpenCV and FFmpeg, combines an adaptive interpolation algorithm for dynamic resolution matching, establishes a device capability matrix, and automatically selects normalization parameters and data enhancement strategies;
[0012] The video preprocessing module supports the reading of multiple video stream files and picture formats, and dynamically adjusts parameters including but not limited to resolution, image cropping, image normalization, and data enhancement of video frames according to the input requirements of the inference model and heterogeneous devices. The video preprocessing module is used for efficient adaptation of video input in multiple hardware environments.
[0013] Preferably, the video inference and analysis module includes:
[0014] A model conversion sub-module, which is used to support the mutual conversion of models in different framework formats, and the model conversion sub-module is adapted to a variety of heterogeneous hardware acceleration devices. The model conversion sub-module has an ONNX intermediate representation layer built-in, supports the automatic topology conversion of TensorFlow / PyTorch / MXNet, and uses a layer fusion optimization technology to improve the conversion efficiency;
[0015] A hardware adaptation sub-module, which imports the converted model into the corresponding inference framework according to the device characteristics of the running environment, and dynamically selects the optimal inference acceleration device to generate an inference engine for real-time video processing and analysis. The hardware adaptation sub-module integrates NVIDIA TensorRT, Intel OpenVINO, and ARM NN multi-platform SDKs, and constructs an inference engine selection decision tree through device feature fingerprint recognition.
[0016] Preferably, the post-processing module includes:
[0017] A post-processing logic component library, which contains the post-processing logic of common video analysis tasks and supports user-defined post-processing processes;
[0018] A registration and management post-processing module, which is used to register and manage all post-processing modules in the post-processing logic component library. The registration and management module performs a dependency injection mechanism, supports the definition of a processing pipeline through a YAML configuration file, and the custom interface provides Python / C++ dual-language bindings, not limited to users registering processing logic based on Lambda expressions;
[0019] A dynamic loading module, which selects a qualified module from all registered post - processing modules according to the output format of the inference result, executes the corresponding processing logic, and loads the processing module in the form of a dynamic link library.
[0020] Preferably, the OSD module is used for the result overlay of multiple video analysis tasks, supports the visual display of detection frames, recognition results, segmentation masks, and text labels, and provides configuration interfaces for colors, fonts, and sizes.
[0021] The present invention also provides a low - coupling intelligent video analysis task construction method for heterogeneous devices, which is applied to the low - coupling intelligent video analysis framework for heterogeneous devices described in any one of the above, and includes the following steps:
[0022] S1: Video pre - processing: According to the requirements of the input model of the heterogeneous device, dynamically adjust the video stream format to adapt to the model input, and output it to the next module;
[0023] S2: Video inference analysis: Receive the pre - processed video frames and the model, perform model conversion and adaptation according to the current heterogeneous device environment, import the converted model into the corresponding inference framework, select a suitable hardware acceleration device to generate an inference engine, process and analyze the video frames, and output the inference result to the next module;
[0024] S3: Post - processing: Perform result parsing and structuring processing on the input inference result, dynamically load qualified post - processing logic according to the inference result, and perform corresponding post - processing operations. The processed inference result is output to the next module;
[0025] S4: OSD processing: Receive the processed inference result and the video frame, overlay the inference result on the video frame, and finally output and display it to the user along with the video stream. The visual display of the inference result supports the overlay of detection frames, recognition results, segmentation masks, and text labels, and provides configurable interfaces for colors, fonts, sizes, and positions.
[0026] Preferably, in the video pre - processing step S1, dynamically adjusting the video stream format includes resolution adjustment, image cropping, image normalization, and data augmentation, and the specific parameters are automatically configured according to the inference model and the input requirements of the heterogeneous device.
[0027] Preferably, in the video inference analysis step S2, the model conversion and adaptation support the ONNX intermediate representation layer, perform automatic topology conversion of multiple framework models such as TensorFlow, PyTorch, and MXNet, and improve the conversion efficiency through layer fusion optimization technology.
[0028] Preferably, in the post - processing step S3, the dynamically loaded post - processing logic selects a matching module from the post - processing logic component library based on the output format of the inference result, and supports users to register custom post - processing logic through Python / C++ interfaces.
[0029] The present invention also provides a hardware device, including a video acquisition device, a main control device, a hardware - accelerated inference device, and a display device. It is characterized in that the hardware device supports the collaborative work of heterogeneous hardware devices through modular design. The video acquisition device is used for video stream acquisition; the hardware - accelerated inference device supports multiple inference frameworks and device characteristics to perform inference acceleration during video processing and analysis. The display device is used for displaying the video stream with the superimposed inference result at the end. The main control device is used for unified management and task scheduling of heterogeneous hardware and executes the computer program to implement a method for constructing a low - coupling video analysis task for heterogeneous devices as described in claim 6.
[0030] The technical effects and advantages of the present invention:
[0031] (1) Through the cooperation of the video pre - processing module, the video inference analysis module, the post - processing module, and the OSD module, the present invention supports the unified processing of multiple hardware - accelerated devices and model formats, provides an efficient and flexible video analysis solution, and solves the compatibility and efficiency problems of the video analysis system in a heterogeneous device environment through modular design and dynamic adaptation mechanism, and is applicable to the application requirements of different scenarios such as smart cities, industrial quality inspection, and autonomous driving.
[0032] (2) The framework of the present invention supports video analysis tasks in multiple hardware environments through modular design, provides efficient adaptation and decoupling for heterogeneous devices, and the framework supports multiple hardware - accelerated devices, and monitors the running status and resource usage of each module in real - time. The framework supports a multi - user collaboration mode to ensure the efficient use of system resources, realizes the decoupling of links such as pre - processing, inference, and post - processing, and improves the flexibility and scalability of the system.
[0033] (3) Through the multi - scenario support of the framework, the present invention reduces the development and maintenance costs. Through the visualization and configurability of OSD processing, it provides an intuitive visual output and rich configuration interfaces to meet the personalized needs of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic structural diagram of the low - coupling intelligent video analysis framework provided by the present invention.
[0035] Figure 2 It is a schematic working diagram of the video pre - processing module in the low - coupling intelligent video analysis framework provided by the present invention.
[0036] Figure 3Schematic diagram of the working process of the video inference and analysis module in the low-coupling intelligent video analysis framework provided by the present invention.
[0037] Figure 4 Schematic diagram of the working process of the post-processing module in the low-coupling intelligent video analysis framework provided by the present invention.
[0038] Figure 5 Schematic diagram of the working process of the OSD module in the low-coupling intelligent video analysis framework provided by the present invention.
[0039] Figure 6 Schematic flowchart of the method for constructing a low-coupling video analysis task provided by the present invention.
[0040] Figure 7 Schematic diagram of the structure of the hardware device provided by the present invention. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] The present invention provides a Figure 1-5 low-coupling intelligent video analysis framework for heterogeneous devices as shown in the figure, including: a video preprocessing module, a video inference and analysis module, a post-processing module, and an OSD module. The video preprocessing module, the video inference and analysis module, the post-processing module, and the OSD module form a stable framework. The framework supports federated learning technology to achieve collaborative inference between edge devices and automatically optimize the model structure through neural architecture search technology; the framework supports native adaptation of photon computing chips and further improves the video analysis efficiency through optical computing acceleration technology; the framework provides end-to-end performance optimization tools, including functions such as model quantization, pruning, and distillation, and supports efficient operation on resource-constrained edge devices; the framework provides open API interfaces, supports integration with third-party systems, and simplifies the development process through the SDK toolkit; the framework supports a multi-user collaboration mode, provides permission management and task scheduling functions to ensure the efficient use of system resources; the framework supports real-time performance monitoring and fault diagnosis functions, and helps users quickly locate and solve problems through log analysis and performance metric visualization; the framework supports a multi-language development environment, including Python, C++, Java, etc., and provides detailed development documents and example codes to lower the usage threshold for developers; the framework supports multi-scenario applications, including fields such as smart city video surveillance, industrial quality inspection, autonomous driving, and medical image analysis.
[0043] Among them, the video preprocessing module dynamically adjusts the video stream format according to the input requirements of the inference model to achieve efficient adaptation of video input, realizing the decoupling of the preprocessing process in the construction of video analysis tasks. The video preprocessing module adopts real-time parameter adjustment technology based on model requirements, integrates the multi-format decoding capabilities of OpenCV and FFmpeg, combines adaptive interpolation algorithms (such as bilinear / Bicubic) for dynamic resolution matching, establishes a device capability matrix (such as the quantization support of the NPU), and automatically selects normalization parameters and data augmentation strategies. The video preprocessing module supports the reading of various video stream files and picture formats, and dynamically adjusts the parameters of video frames including but not limited to resolution, image cropping, image normalization, and data augmentation according to the input requirements of the inference model and heterogeneous devices. The video preprocessing module is used for the efficient adaptation of video input in multiple hardware environments, for example Figure 2 In the embodiments of Figure 2 , three inference frameworks, namely TensorFlow, OnnxRuntime, and OpenVINO, are designed to implement the structure parsing of corresponding models. In addition, the video stream is extracted into frames of images after entering the video preprocessing module, and preprocessing operations such as resolution adjustment, image cropping, normalization, data augmentation, etc. are performed on each frame of the image according to the parsed model input information. Finally, the processed video frames and the model are output to the next module together.
[0044] In addition, the video inference analysis module performs unified processing of multiple model formats and heterogeneous devices through inference combining model conversion and hardware accelerator adaptation, efficiently completing video inference in different device environments, and realizing the decoupling of the inference process in the construction of video analysis tasks at both the software and hardware levels;
[0045] The video inference analysis module receives the processed video frames and the model output by the previous module, the video preprocessing module. The video inference analysis module includes a model conversion sub-module and a hardware adaptation sub-module. The model conversion sub-module is used to support the mutual conversion of models in different framework formats, and the model conversion sub-module is adapted to a variety of heterogeneous hardware acceleration devices. The model conversion sub-module has an ONNX intermediate representation layer built-in, supports the automatic topology conversion of TensorFlow / PyTorch / MXNet, and uses layer fusion optimization technology to improve the conversion efficiency; such as Figure 3In the illustrated embodiment, the model import module is used to create an inference engine. Currently, the framework supports three types of hardware acceleration devices, namely GPU, TPU, and VPU. The model conversion sub-module supports several types of models that can be converted to each other, including but not limited to the model formats supported by these three types of hardware acceleration devices for accelerated inference, namely the onnx model, the tflite model, and the OpenVINO model. This sub-module converts the model according to the current hardware environment used. For example, if the currently used hardware acceleration device is GPU, the model will be converted into an onnx file; if it is a TPU hardware acceleration device, the model will be converted into a tflite file; and if it is a VPU hardware acceleration device, the model will be converted into an OpenVINO model.
[0046] The hardware adaptation sub-module imports the converted model into the corresponding inference framework according to the device characteristics of the running environment, and dynamically selects the optimal inference acceleration device to generate an inference engine. The inference framework includes but not limited to TensorFlow, OnnxRuntime, and OpenVINO, and selects the current best running environment. Finally, an inference engine adapted to the current video analysis task and the hardware running environment is generated to achieve the efficiency and adaptability of real-time video processing and analysis for video frame inference analysis. The inference results generated by the inference analysis will be output to the next module together with the video frames. The hardware adaptation sub-module integrates multi-platform SDKs such as NVIDIA TensorRT, Intel OpenVINO, and ARM NN, and constructs an inference engine selection decision tree through device feature fingerprint recognition (CUDA core count / VPU computing power).
[0047] At the same time, the post-processing module performs result parsing and structured processing according to the post-processing logic template. The post-processing logic template is dynamically loaded and matched with the relevant information of the inference results, and at the same time supports users to register custom post-processing logics, realizing the decoupling of the post-processing process in the construction of video analysis tasks. The post-processing module includes a post-processing logic component library, a registration and management post-processing module, and a dynamic loading module. The post-processing logic component library contains the post-processing logics for common video analysis tasks, supports users to customize the post-processing process, and at the same time the dynamic library supports users to customize post-processing logics. Users can refer to the implementation of existing templates to quickly create post-processing logics that meet the requirements of specific video analysis tasks, improving the scalability of the framework, reducing the difficulty of user customization, and enhancing the development efficiency.
[0048] Specifically, the registration and management post - processing module is used to register and manage all post - processing modules in the post - processing logic component library. The registration and management module implements a dependency injection mechanism, supports defining the processing pipeline through a YAML configuration file, provides Python / C++ dual - language binding through a custom interface, and does not limit users to registering processing logic based on lambda expressions. The dynamic loading module selects eligible modules from all registered post - processing modules and executes the corresponding processing logic according to the output format of the inference result, and loads the processing module in the form of a dynamic link library.
[0049] As Figure 4 In the embodiment shown, during the operation of the video analysis task, the post - processing module dynamically registers each post - processing module into a global post - processing manager through the registration mechanism. When the inference result is input into the post - processing module, it will dynamically load and call the post - processing logic in the manager according to its specific information. In the embodiment of the present invention, the framework selects the output size of the inference result to correspond to the appropriate post - processing logic because the output sizes of different models may be different. At the same time, a judgment function is designed in each post - processing logic. When the inference result meets the input of a specific post - processing function, the post - processing logic will be called and the corresponding post - processing operation will be performed, and finally the parsed inference result will be output to the next module.
[0050] Furthermore, the OSD module is used to overlay the inference result on the video stream and generate a visual output, realizing the decoupling of the OSD process in the construction of the video analysis task. The OSD module is used for the result overlay of multiple video analysis tasks, supports the visual display of detection boxes, recognition results, segmentation masks, and text labels, and provides a configuration interface for colors, fonts, and sizes. The OSD module draws different visual contents for different video analysis tasks. For example, in the object detection task, the detected object needs to be framed with a detection box; in the object recognition task, the category of the detected object needs to be represented by a text label; in the semantic segmentation task, a segmentation mask needs to be drawn to mark and distinguish different regions; in the action recognition task, the corresponding key points need to be drawn and connected to mark the detected object, and after being overlaid on the video frame, it will be output and displayed to the user along with the video stream.
[0051] A low - coupling intelligent video analysis task construction method for heterogeneous devices, which is applied to the above - mentioned low - coupling intelligent video analysis framework for heterogeneous devices, as Figure 6 shown, includes the following steps:
[0052] S1: Video preprocessing: Dynamically adjust the video stream format according to the requirements of the heterogeneous device input model to adapt to the model input and output to the next module. Among them, dynamically adjusting the video stream format includes resolution adjustment, image cropping, image normalization, and data augmentation, and the specific parameters are automatically configured according to the input requirements of the inference model and heterogeneous devices.
[0053] S2: Video inference analysis: Receive the preprocessed video frames and the model, perform model conversion and adaptation according to the current heterogeneous device environment, import the converted model into the corresponding inference framework and select a suitable hardware acceleration device to generate an inference engine, process and analyze the video frames, and output the inference results to the next module. Model conversion and adaptation support the ONNX intermediate representation layer, perform automatic topology conversion of multiple framework models such as TensorFlow, PyTorch, and MXNet, and improve the conversion efficiency through layer fusion optimization technology.
[0054] S3: Post-processing: Parse and structure the input inference results, dynamically load eligible post-processing logic according to the inference results and perform corresponding post-processing operations, and output the processed inference results to the next module. Dynamically loading post-processing logic is based on the output format of the inference results, selects a matching module from the post-processing logic component library, and supports users to register custom post-processing logic through Python / C++ interfaces.
[0055] S4: OSD processing: Receive the processed inference results and video frames, overlay the inference results on the video frames, and finally output and display them to the user along with the video stream. Among them, the visual display of the inference results supports the overlay of detection boxes, recognition results, segmentation masks, and text labels, and provides configurable interfaces for color, font, size, and position.
[0056] A hardware device, such as Figure 7As shown, it includes a video acquisition device, a main control device, a hardware acceleration inference device, and a display device. The hardware device supports the collaborative work of heterogeneous hardware devices through modular design. The video acquisition device is used for video stream acquisition; the hardware acceleration inference device supports multiple inference frameworks and device characteristics to accelerate inference during video processing and analysis. The display device is used for displaying the video stream with the superimposed inference result. The main control device is a computer program for unified management and task scheduling of heterogeneous hardware and execution to implement a low-coupling video analysis task construction method for heterogeneous devices as claimed in claim 6. Among them, the video stream acquisition device completes the acquisition of real-time video and transmits the video stream to the main control device. The main control device receives the video stream and executes the computer program to implement the above-mentioned low-coupling video analysis task construction method for heterogeneous devices, realizing unified management and task scheduling of heterogeneous hardware. The hardware acceleration inference device supports multiple inference frameworks and device characteristics, and is used to accelerate the video inference process in the computer program. In an embodiment of a hardware device according to the present invention, Raspberry Pi's Pi4 and NVIDIA's Jetson TX2 are used as the main control device, and the above-mentioned low-coupling video analysis task construction method can be implemented on any device. In addition, the hardware acceleration inference device adopts the GPU modules of intel NCS2, Coral USB Accelerator, and Jetson TX2 corresponding to three different architectures of hardware acceleration devices, namely VPU, TPU, and GPU, to realize video analysis acceleration inference under heterogeneous devices and solve the decoupling problem of the video analysis task construction method at the hardware level. The display device, i.e., the display, is used to display the inference result to the user.
[0057] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A low-coupling intelligent video analysis framework for heterogeneous devices, characterized in that Including: A video preprocessing module, which dynamically adjusts the video stream format according to the input requirements of the inference model to achieve efficient adaptation of video input; A video inference analysis module, which performs unified processing of multiple model formats and heterogeneous devices through inference combining model conversion and hardware accelerator adaptation; A post-processing module, which performs result parsing and structured processing according to the post-processing logic template, and the post-processing logic template matches the dynamic loading of inference results; An OSD module, which is used to overlay the inference results on the video stream and generate a visual output.
2. The low-coupling intelligent video analysis framework for heterogeneous devices according to claim 1, wherein The video preprocessing module adopts a real-time parameter adjustment technology based on model requirements, integrates the multi-format decoding capabilities of OpenCV and FFmpeg, combines an adaptive interpolation algorithm for dynamic resolution matching, establishes a device capability matrix, and automatically selects normalization parameters and data augmentation strategies; The video preprocessing module supports the reading of multiple video stream files and picture formats, and dynamically adjusts the parameters of the video frame including but not limited to resolution, image cropping, image normalization, and data augmentation according to the input requirements of the inference model and heterogeneous devices. The video preprocessing module is used for efficient adaptation of video input in multiple hardware environments.
3. A low-coupling intelligent video analysis framework for heterogeneous devices according to claim 1, characterized in that, The video inference analysis module includes: A model conversion sub-module, which is used to support the mutual conversion of models in different framework formats, and the model conversion sub-module adapts to a variety of heterogeneous hardware acceleration devices. The model conversion sub-module has an ONNX intermediate representation layer, supports automatic topology conversion of TensorFlow / PyTorch / MXNet, and uses layer fusion optimization technology to improve conversion efficiency; A hardware adaptation sub-module, which imports the converted model into the corresponding inference framework according to the device characteristics of the running environment, and dynamically selects the optimal inference acceleration device to generate an inference engine for real-time video processing analysis. The hardware adaptation sub-module integrates NVIDIA TensorRT, Intel OpenVINO, and ARM NN multi-platform SDKs, and constructs an inference engine selection decision tree through device feature fingerprint recognition.
4. A low-coupling intelligent video analysis framework for heterogeneous devices according to claim 1, characterized in that The post-processing module includes: A post-processing logic component library, which contains the post-processing logics of common video analysis tasks and supports user-defined post-processing processes; A registration and management post-processing module, which is used to register and manage all post-processing modules in the post-processing logic component library. The registration and management module performs a dependency injection mechanism, supports defining a processing pipeline in a YAML configuration file, and the custom interface provides Python / C++ dual-language bindings, not limited to users registering processing logics based on Lambda expressions; A dynamic loading module, which selects eligible modules from all registered post-processing modules according to the output format of the inference results and executes the corresponding processing logics, and loads the processing modules in the form of dynamic link libraries.
5. The low-coupling intelligent video analysis framework for heterogeneous devices according to claim 1, characterized in that The OSD module is used for the overlay of the results of various video analysis tasks, supports the visual display of detection boxes, recognition results, segmentation masks, and text labels, and provides configuration interfaces for colors, fonts, and sizes.
6. A low-coupling intelligent video analysis task construction method for heterogeneous devices, which is applied to the low-coupling intelligent video analysis framework for heterogeneous devices described in any one of claims 1-5, and is characterized in that It includes the following steps: S1: Video preprocessing: According to the requirements of the heterogeneous device input model, dynamically adjust the video stream format to adapt to the model input and output it to the next module; S2: Video inference analysis: Receive the preprocessed video frames and the model, perform model conversion and adaptation according to the current heterogeneous device environment, import the converted model into the corresponding inference framework and select a suitable hardware acceleration device to generate an inference engine, process and analyze the video frames, and output the inference results to the next module; S3: Post-processing: Perform result parsing and structured processing on the input inference results, dynamically load the eligible post-processing logic according to the inference results and perform corresponding post-processing operations, and output the processed inference results to the next module; S4: OSD processing: Receive the processed inference results and video frames, overlay the inference results onto the video frames, and finally output and display them to the user along with the video stream. The visual display of the inference results supports the overlay of detection boxes, recognition results, segmentation masks, and text labels, and provides configurable interfaces for colors, fonts, sizes, and positions.
7. A method for constructing a low-coupling intelligent video analysis task for heterogeneous devices according to claim 6, characterized in that, In the video preprocessing step S1, dynamically adjusting the video stream format includes resolution adjustment, image cropping, image normalization, and data augmentation, and the specific parameters are automatically configured according to the input requirements of the inference model and heterogeneous devices.
8. A method for constructing a low-coupling intelligent video analysis task for heterogeneous devices according to claim 6, characterized in that In the video inference analysis step S2, the model conversion and adaptation support the ONNX intermediate representation layer, perform automatic topology conversion of multiple framework models such as TensorFlow, PyTorch, and MXNet, and improve the conversion efficiency through layer fusion optimization technology.
9. A method for constructing a low-coupling intelligent video analysis task for heterogeneous devices according to claim 6, wherein, In the post-processing step S3, dynamically loading the post-processing logic is based on the output format of the inference results, selects the matching module from the post-processing logic component library, and supports users to register custom post-processing logic through Python / C++ interfaces.
10. A hardware device, comprising a video acquisition device, a main control device, a hardware acceleration inference device, and a display device, characterized in that, The hardware device supports the collaborative work of heterogeneous hardware devices through modular design. The video acquisition device is used for the acquisition of video streams; the hardware acceleration inference device supports multiple inference frameworks and device characteristics, performs inference acceleration during video processing and analysis, the display device is used for the display of the video stream with the finally overlaid inference results, and the main control device is used for the unified management and task scheduling of heterogeneous hardware and executes the computer program to implement a low-coupling video analysis task construction method for heterogeneous devices as described in claim 6.