Video fusion method and device based on edge computing platform

By employing a video fusion method on an edge computing platform, this approach utilizes OpenCV operators and MPI multithreading to pull video streams in parallel, perform decoding and affine transformation, and combine this with an NPU fusion matrix for edge inference of video frames. This solves the problems of interoperability barriers and high costs in video fusion, achieving efficient and stable video fusion processing.

CN120980269APending Publication Date: 2025-11-18ARMOR ACADEMY OF CHINESE PEOPLES LIBERATION ARMY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511185999.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing video fusion methods suffer from problems such as incompatible video protocols, inconsistent video encoding formats, and limited processing performance, which lead to obstacles in the interoperability of video resources during the fusion process, as well as high costs and poor stability.

Method used

A video fusion method based on an edge computing platform is adopted. The IP protocol address is determined by OpenCV operators, the video stream is pulled in parallel using MPI multi-threading, and decoding, affine transformation and energy correlation matrix calculation are performed. The NPU fusion matrix of the edge computing platform is used to perform edge inference of video frames, so as to achieve efficient fusion of video frames.

Benefits of technology

It achieves efficient, stable, and low-cost video fusion processing, improves streaming speed by 3 to 4 times, solves the problems of video protocol incompatibility and inconsistent encoding formats, and is suitable for edge deployment in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980269A_ABST
    Figure CN120980269A_ABST
Patent Text Reader

Abstract

The invention discloses a video fusion method and device based on an edge computing platform. The method comprises the following steps: parallelly pulling each path of to-be-fused video stream by using MPI (Message Passing Interface) multithreading; decoding all the video streams to be fused; performing affine transformation on each video frame in each decoded video stream, and determining an energy correlation matrix corresponding to the video frame based on an affine transformation result corresponding to each video frame; storing the feature points and the edge feature angular points corresponding to the video frames into the same Jason file; replacing the data of the pixel points in the video frame corresponding to the feature points with the feature point data corresponding to the feature points in the Jason data file; and filling the NPU fusion matrix corresponding to each video frame based on each replaced video frame data to determine an edge pixel range of each video frame, and performing edge reasoning to obtain a fused video. The method solves the problems of incompatibility of video protocols, non-uniform video coding formats, limited processing performance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, in particular to a video fusion method and device based on an edge computing platform. BACKGROUND

[0002] In the existing field of video processing, video fusion is a key technology and is widely used in commercial display, security monitoring, intelligent transportation and other scenes. However, the existing video fusion method has problems such as video protocol incompatibility, video coding format inconsistency, and limited processing performance.

[0003] The video data fusion method with application number CN118784893A is based on FPGA multi-layer image processing: cropping and reducing, DMA read and write operations of video data, etc., to realize video data splicing and fusion. This method can effectively process video data, but the FPGA chip is expensive and has poor versatility, which is not convenient for edge deployment.

[0004] The Internet of Things video fusion method with application number CN119380516A constructs a 5G Internet of Things video monitoring area, sets a regional risk coefficient and a judgment threshold, intelligently warns the video building, identifies and locates the personnel position coordinates in the video through the YOLO algorithm. This method can effectively realize video fusion, but it depends on high-performance servers to complete, has low cost performance, and is not suitable for special scenes.

[0005] Different communication systems and video systems use different communication and transmission protocols, resulting in interoperability problems in video resource fusion; differences in video coding formats, such as H.264 and H.265, make it impossible to achieve video fusion even if the protocols support interoperability; traditional video fusion systems mostly use independent graphics servers or FPGA array cards for processing, which have problems such as low cost performance, poor stability, and high cost. In addition, the conventional multi-threaded parallel network camera streaming only establishes multi-threaded tasks and then performs parallel streaming, which is not very fast, and when the number of cameras is large, the delay will be large, affecting target capture and panoramic monitoring. SUMMARY

[0006] The present application proposes a video fusion method and device based on an edge computing platform, which can solve the above technical problems.

[0007] In the method embodiments of the present application, a video fusion method based on an edge computing platform includes:

[0008] Step S1: inputting multiple videos captured by multiple cameras into a server device; each camera corresponds to an IP protocol address;

[0009] Step S2: establishing a first OpenCV operator on the server device; using the first OpenCV operator to determine IP protocol addresses corresponding to each camera that needs to be video fused, and to determine video streams to be fused corresponding to the IP protocol addresses corresponding to each camera that needs to be video fused, using MPI multi-threading to pull each video stream to be fused in parallel;

[0010] Step S3: establishing a second OpenCV operator on the server device; using the second OpenCV operator to decode all video streams to be fused;

[0011] Step S4: performing affine transformation on each video frame in each decoded video stream respectively, determining an energy-related matrix corresponding to each video frame based on the affine transformation result corresponding to the video frame, determining a feature point and an edge feature corner point corresponding to each video frame based on the energy-related matrix corresponding to the video frame, and obtaining feature point data of the feature point corresponding to each video frame and edge feature corner point data of the edge feature corner point corresponding to each video frame based on the affine transformation result corresponding to the video frame;

[0012] Step S5: storing the feature point and the edge feature corner point corresponding to each video frame into the same Jason file;

[0013] Step S6: connecting multiple cameras to an edge computing platform, importing the Jason data file into the edge computing platform, replacing data of a pixel point in a video frame corresponding to a feature point with feature point data corresponding to the feature point in the Jason data file, filling an NPU fusion matrix corresponding to each video frame based on the replaced video frame data, determining an edge pixel range of each video frame based on each NPU fusion matrix, and performing edge inference based on the edge pixel range of each video frame to obtain a fused video.

[0014] Optionally, in the step S4, the affine transformation on each video frame in each decoded video stream comprises:

[0015] The following operations are performed on each video frame in each decoded video stream:

[0016] Step S41: obtaining three first vertices at fixed positions in a video frame src, and determining second vertices corresponding to the three first vertices respectively in a target video frame dst through affine transformation;

[0017] The affine transformation formula is:

[0018] =

[0019] wherein is a parameter matrix of the affine transformation, ( , ) represents the pixel data in the src of the video frame. () represents the pixel data in the target video frame dst after affine transformation;

[0020] Step S42: Based on the first vertex and the second vertex, perform an affine transformation on the video frame src to obtain the target video frame dst.

[0021] Optionally, in step S4, determining the energy correlation matrix corresponding to each video frame based on the affine transformation result includes:

[0022] For each target video frame dst after affine transformation, perform the following operations:

[0023] Obtain each pixel p in the src of the video frame corresponding to the target video frame dst. ij Determine p ij The pixel p* corresponding to the target video frame dst ij The correlation;

[0024] The energy correlation matrix Emat_P corresponding to this video frame ij for:

[0025] Ematrix_P ij = Corr((H_frameA[p ij ]) 2 , (H_frameB[p ij ]) 2 )

[0026] Where i = 0, 1, 2, ..., witdth; j = 0, 1, 2, 3, ..., height, witdth and height are the width and height of the video frame src, respectively; Corr(·) is the correlation calculation, H_frameA[p ij ]= src ( , ), H_frameB[p ij ])=dst( ).

[0027] Optionally, in step S4, determining the feature points and edge feature corner points corresponding to each video frame based on the energy correlation matrix corresponding to each video frame includes:

[0028] For each video frame, perform the following operations:

[0029] Determine the energy value threshold corresponding to the video frame, and use the pixels in the target video frame corresponding to the matrix elements whose matrix element values ​​exceed the energy value threshold in the energy correlation matrix of the video frame as feature points.

[0030] constructing a feature operator fusion list, the feature operator fusion list including all feature points;

[0031] writing all edge feature corner points into a corner point feature list, wherein the preset condition is an edge condition.

[0032] Optionally, in the step S1, a plurality of same type cameras are selected, and the cameras are connected to a server device through an Ethernet switch.

[0033] In the above-mentioned method embodiments of the application, a video fusion device based on an edge computing platform comprises:

[0034] An initialization module is configured to connect a plurality of videos captured by a plurality of cameras to a server device; each camera corresponds to an IP protocol address;

[0035] A data acquisition module is configured to establish a first OpenCV operator on the server device; determine IP protocol addresses corresponding to each camera that needs to perform video fusion using the first OpenCV operator, and determine video streams to be fused corresponding to the IP protocol addresses of each camera that needs to perform video fusion, and use MPI multi-threading to pull each video stream to be fused in parallel;

[0036] A decoding module is configured to establish a second OpenCV operator on the server device; and decode all video streams to be fused using the second OpenCV operator;

[0037] An affine module is configured to perform affine transformation on each video frame in each decoded video stream respectively, determine an energy-related matrix corresponding to each video frame based on the affine transformation result corresponding to the video frame, determine feature points and edge feature corner points corresponding to each video frame based on the energy-related matrix corresponding to the video frame, and obtain feature point data of the feature points corresponding to each video frame and edge feature corner point data of the edge feature corner points corresponding to each video frame based on the affine transformation result corresponding to each video frame;

[0038] A storage module is configured to store the feature points and edge feature corner points corresponding to each video frame into a same Jason file;

[0039] A fusion module is configured to connect the plurality of cameras to an edge computing platform, import the Jason data file into the edge computing platform, replace data of pixel points in a video frame corresponding to a feature point with feature point data of the feature point in the Jason data file, fill an NPU fusion matrix corresponding to each video frame based on the replaced video frame data, determine an edge pixel range of each video frame based on the NPU fusion matrix, and perform edge inference based on the edge pixel range of each video frame to obtain a fused video.

[0040] In the method embodiments of the present application, a computer readable storage medium, the storage medium stores a plurality of instructions, the plurality of instructions are used to be loaded and executed by a processor as the method described above.

[0041] In the method embodiments of the present application, an electronic device, the electronic device comprises: a processor, used to execute a plurality of instructions; a memory, used to store a plurality of instructions; wherein the plurality of instructions are used to be stored by the memory, and loaded and executed by the processor as the method described above.

[0042] The present application can realize efficient, stable and low-cost video fusion processing, effectively solve the problems of video protocol incompatibility, video encoding format non-uniformity and processing performance limitation in the existing video fusion method; in addition, compared with the conventional multi-thread parallel pull stream, the MPI multi-thread parallel pull stream can increase the pull stream speed by 3-4 times on the same hardware platform through the shared memory mode.

[0043] The technical solutions of the present application will be described in further detail below by means of the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0044] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of embodiments of the present application taken in conjunction with the accompanying drawings. The drawings provided in the specification and the embodiments of the present application together serve to provide a further understanding of the present application, and constitute a part of the specification, explain the present application together with the embodiments of the present application, and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally designate the same components or steps throughout the specification.

[0045] Figure 1 The present application is a flowchart of the video fusion method based on the edge computing platform;

[0046] Figure 2 The present application is a flowchart of pulling each road to be fused video stream;

[0047] Figure 3 The present application is a flowchart of decoding the to-be-fused video stream;

[0048] Figure 4 The present application is a flowchart of pre-processing the decoded video stream;

[0049] Figure 5 The present application is a flowchart of realizing the edge platform multi-channel video fusion;

[0050] Figure 6 The present application is a structure diagram of the video fusion device based on the edge computing platform;

[0051] Figure 7 FIG. 1 is a structural schematic diagram of a video fusion electronic device based on an edge computing platform according to an embodiment of the present application. DETAILED DESCRIPTION

[0052] Hereinafter, example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of embodiments of the present application, and thus are not to limit the whole embodiments of the present application. It should be understood that the present application is not limited to the example embodiments described herein. It should be noted that the relative arrangement of the components and steps, numerical expressions, and numerical values set forth in these embodiments are not to limit the scope of the present application unless otherwise specifically stated.

[0053] Those skilled in the art can understand that the terms "first", "second", S1, S2 and the like in the embodiments of the present application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they represent the inevitable logical sequence between them. It should also be understood that in the embodiments of the present application, "a plurality of" can mean two or more, and "at least one" can mean one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the present application, it can be understood as one or more in general without explicit limitation or in the context of the preceding and following text. In addition, the term "and / or" in the present application is only to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects. It should also be understood that the description of each embodiment of the present application emphasizes the differences between each embodiment, and the same or similar parts can be referred to each other, and for the sake of brevity, they will not be repeated. At the same time, it should be understood that the size of each part shown in the drawings is not drawn in accordance with the actual proportional relationship. The following description of at least one example embodiment is actually only illustrative, but not as any limitation on the present application and its application or use. The techniques, methods and devices known to those skilled in the relevant art can not be discussed in detail, but in appropriate cases, the techniques, methods and devices should be considered as part of the specification. It should be noted that similar reference numbers and letters represent similar items in the following drawings, and thus once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0054] The embodiments of the present application can be applied to terminal devices, computer systems, servers and other electronic devices, which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and the like. Terminal devices, computer systems, servers and other electronic devices can be described in the general context of computer system executable instructions, such as program modules, executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, in which tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.

[0055] Exemplary method

[0056] Figure 1 is a flowchart of a video fusion method based on an edge computing platform provided by an exemplary embodiment of the present application. As shown in Figure 1 , the method comprises the following steps:

[0057] Step S1: inputting multiple videos captured by multiple cameras into a server device; each camera corresponds to an IP protocol address;

[0058] Step S2: establishing a first OpenCV operator on the server device; using the first OpenCV operator to determine the IP protocol addresses of the cameras that need to be fused, and determining the video streams to be fused corresponding to the IP protocol addresses of the cameras that need to be fused, and using MPI multi-threading to pull the video streams to be fused in parallel;

[0059] Step S3: establishing a second OpenCV operator on the server device; using the second OpenCV operator to decode all the video streams to be fused;

[0060] Step S4: affine transformation is performed on each video frame in the decoded video stream respectively, the energy-related matrix corresponding to each video frame is determined based on the affine transformation result corresponding to the video frame, the feature points and edge feature corner points corresponding to each video frame are determined based on the energy-related matrix corresponding to the video frame, and the feature point data of the feature points corresponding to each video frame and the edge feature corner point data of the edge feature corner points are obtained based on the affine transformation result corresponding to the video frame;

[0061] Step S5: the feature points and edge feature corner points corresponding to each video frame are stored in the same Jason file.

[0062] Step S6: the multi-camera is connected to the edge computing platform, the Jason feature data file is imported into the edge computing platform, the feature point data corresponding to the feature points in the Jason feature data file is used to replace the data of the pixel points in the video frame corresponding to the feature points, the NPU fusion matrix corresponding to each video frame is filled based on the replaced video frame data, the edge pixel range of each video frame is determined based on the NPU fusion matrix, and the edge inference is performed based on the edge pixel range of each video frame to obtain the fusion video.

[0063] The NPU fusion matrix is constructed through the edge feature corner points, the feature point data corresponding to the Jason file coordinates in the image matrix after affine transformation is used to replace the image pixel point data before affine transformation, the data processed above is filled into the NPU fusion matrix according to the edge pixel range of the NPU fusion matrix to perform edge inference to obtain the fusion video.

[0064] The MPI multi-thread parallel extraction and NPU acceleration technology are adopted, the fusion processing of the video stream can be efficiently promoted, the edge computing platform deployed in the NPU has low cost and can be deployed in multiple scenes and in all directions, the video protocol decoding is automatically realized through the OpenCV operator specification, and the non-uniformity of video frame encoding and decoding is avoided.

[0065] The edge device related to the edge computing platform has a hardware environment: 8-core NPU, 8G memory, 2.4G main frequency, Arm architecture and four-channel LPDDR4, and has a software environment: Python 3.8, pytorch 1.10, ffmpeg-Rockchip, MPP, RGA, OpenCV 4.10.1 and Ubuntu 2204 system.

[0066] In the step S1, N cameras of the same type are selected and numbered as M1, M2, M3, …, M N The N cameras and the server are connected to the switch in an Ethernet mode, and the import of the multi-channel video stream is realized.

[0067] As Figure 2As shown, IP protocol addresses are set for N cameras. Based on the video resolution, a shared memory address `video_arr` is allocated. The first OpenCV operator `_OpenAI_Video` is created to access and acquire data from the shared memory address of the N cameras, achieving multi-threaded parallel streaming.

[0068] video_arr={ IP-dev_arr1, IP-dev_arr2, IP-dev_arr3,…, IP-dev_arrN}

[0069] Stream1 = _OpenAI_Video(IP-dev_arr1) -- MPI thread 1

[0070] Stream2 = _OpenAI_Video(IP-dev_arr2) -- MPI thread 2

[0071] Stream3 = _OpenAI_Video(IP-dev_arr3) -- MPI thread 3

[0072] ...

[0073] StreamN = _OpenAI_Video(IP-dev_arrN) -- MPI thread N

[0074] Among them, _OpenAI_Video is the OpenCV built-in streaming module. The streaming module takes the IP shared memory address of the network camera (IP-dev_arr1, IP-dev_arr2, IP-dev_arr3,…, IP-dev_arrN) as input. By using shared memory, it uses frombuffer to obtain shared memory video data frames for parallel streaming via MPI multi-threading, thereby improving the streaming speed of conventional cameras.

[0075] like Figure 3 As shown, import the OpenCV operator decoding module _OpenCV_Decode to achieve parallel automatic decoding of multiple video streams (the built-in operators can automatically recognize video encoding protocols such as H.264 and H.265):

[0076] Frame1 = _OpenCV_Decode(Stream1) -- MPI thread 1

[0077] Frame2 = _OpenCV_Decode(Stream2) -- MPI thread 2

[0078] Frame3 = _OpenCV_Decode(Stream3) -- MPI thread 3

[0079] …

[0080] FrameN = _OpenCV_Decode(StreamN) -- MPI thread N

[0081] Wherein, _OpenCV_Decode is an OpenCV built-in video protocol decoding module, according to different video encoding protocols, OpenCV can automatically identify the protocol frame header of the video image, confirm the decoding mode of the video image, so as to realize the automatic decoding of the video stream image.

[0082] Further, as shown in Figure 4 , the step S4, the decoded video stream in each video frame is respectively subjected to affine transformation, including:

[0083] The decoded video stream in each video frame is subjected to the following operation:

[0084] Step S41: obtaining three fixed positions of the first vertex in the video frame src, determining the second vertex corresponding to the three first vertices in the target video frame dst through affine transformation;

[0085] The affine transformation formula is:

[0086] =

[0087] Wherein is the parameter matrix of affine transformation, ( , ) is the pixel point data in the video frame src, ( ) is the pixel point data in the target video frame dst after affine transformation;

[0088] Step S42: based on the first vertex and the second vertex, affine transformation is carried out on the video frame src to obtain the target video frame dst.

[0089] In the step S4, the energy related matrix corresponding to each video frame is determined based on the affine transformation result corresponding to the video frame, including:

[0090] The following operation is performed on each target video frame dst after affine transformation:

[0091] Obtaining each pixel p ij in the video frame src corresponding to the target video frame dst, determining p ij and the corresponding pixel p* ijcorrelation of the video frame;

[0092] the energy correlation matrix Ematrix_P corresponding to the video frame ij is:

[0093] Ematrix_P ij = Corr((H_frameA[p ij ]) 2 , (H_frameB[p ij ]) 2 ).

[0094] wherein, i=0,1,2,…,witdth; j=0,1,2,3,…,height, witdth and height are width and height of the video frame src respectively; Corr(·) is correlation calculation, H_frameA[p ij ]= src( , ), H_frameB[p ij ])=dst( )。

[0095] In the step S4, the feature points and edge feature corner points corresponding to each video frame are determined based on the energy correlation matrix corresponding to each video frame, comprising:

[0096] For each video frame, the following operations are performed:

[0097] determine the energy value threshold corresponding to the video frame, and the pixel points in the target video frame corresponding to the matrix elements in the energy correlation matrix of the video frame whose element values exceed the energy value threshold are taken as the feature points;

[0098] a feature operator fusion list is constructed, and the feature operator fusion list includes all feature points;

[0099] the feature points in the feature operator fusion list meeting the preset condition are taken as the edge feature corner points, and all edge feature corner points are written into a corner feature list; wherein, the preset condition is an edge condition.

[0100] In the present application, the affine matrix H_trans is constructed by the three-point method, the affine transformation of the access video frame is realized to obtain H_frame, the energy of the contrast video frame is compared, the energy correlation matrix Ematrix is constructed, and the energy correlation feature operator fusion CombineList is performed.

[0101] The three-point method constructs the affine mapping matrix mainly through the getAffineTransform module of OpenCV:

[0102] H_trans = _OpenCV_getAffineTransform(const Point2f* src, const Point2f*dst)

[0103] Where the parameter const Point2f* src is the three fixed vertices of the original image, and the parameter const Point2f*dst: the three fixed vertices of the target image. The target image and the original image are fixed through the initial image (xth frame image) pulled by the video stream, where x represents the video frame number corresponding to the clear image quality. Then the video frame is subjected to affine transformation to obtain H_frame:

[0104] H_frame1 = Frame1 *H_trans -- MPI thread 1

[0105] H_frame2 = Frame2 *H_trans -- MPI thread 2

[0106] H_frame3 = Frame3 *H_trans -- MPI thread 3

[0107] …

[0108] H_frameN = FrameN *H_trans -- MPI thread N

[0109] Where * represents matrix multiplication, and the dimension of the affine transformed matrix is consistent with that of Frame.

[0110] Energy-related matrix Ematrix_P ij = Corr((H_frameA[pij]) 2 , (H_frameB[pij]) 2 ).

[0111] Where, p ij is the pixel value of the image, i=0,1,2,…,witdth; j=0,1,2,3,…,height, witdth and height correspond to the resolution of the video frame image; Corr(·) is the correlation calculation. Set the energy correlation threshold r, and use the numpy acceleration library to find the feature points that satisfy > r to obtain the feature operator fusion CombineList:

[0112] CombineList=[{x1 ij ,y1 ij}, {x2 ij ,y2 ij},{x3 ij ,y3ij},…{xL ij ,yL ij}],

[0113] wherein L is the number of feature points meeting the requirement, {xk ij ,yk ij} is the pixel coordinate meeting the requirement, k = 1, 2, 3, … L.

[0114] Construct the corner feature table TrEdgeList, and find the edge feature corner points in CombineList:

[0115] TrEdgeList = [{x* m1n1 , y* m1n1},{x* m2n2 , y* m2n2}, {x* m3n3 , y* m3n3},…,{x* mInI , y* mInI}],

[0116] wherein I is the total number of corners, and mInI is the corresponding pixel coordinate.

[0117] In step S5, the feature points and edge feature corner points corresponding to each video frame are stored in the Jason file, wherein:

[0118] Construct the Jason file FeatureSet.jason to access the feature points and corner points:

[0119] “Bine_1”:{

[0120] CombineList1: {x1 ij ,y1 ij}, {x2 ij ,y2 ij},{x3 ij ,y3 ij},…,{xL ij ,yL ij}

[0121] TrEdgeList1: {x* m1n1 , y* m1n1},{x* m2n2 , y* m2n2}, {x* m3n3 , y* m3n3},…, {x* mInI ,y* mInI}

[0122] }

[0123] "Bine_2": {

[0124] CombineList2: {x1 ij ,y1 ij}, {x2 ij ,y2 ij},{x3 ij ,y3 ij},…,{xL ij ,yL ij}

[0125] TrEdgeList2: {x* m1n1 , y* m1n1},{x* m2n2 , y* m2n2}, {x* m3n3 , y* m3n3},…, {x* mInI ,y* mInI}

[0126] }

[0127] …

[0128] "Bine_N": {

[0129] CombineListN: {x1 ij ,y1 ij}, {x2 ij ,y2 ij},{x3 ij ,y3 ij},…,{xL ij ,yL ij}

[0130] TrEdgeListN: {x* m1n1 , y* m1n1},{x* m2n2 , y* m2n2}, {x* m3n3 , y* m3n3},…, {x* mInI ,y* mInI}

[0131] }

[0132] As Figure 5As shown, FeatureSet.jason is imported into the edge computing platform to construct an NPU fusion matrix MatrixNPU. The imported JSON data is extracted into the NPU fusion matrix, and the pixel coordinates in the NPU fusion matrix are applied to the video stream deployed on the edge computing platform to accelerate video fusion (CombineVideoStream). The specific steps are as follows:

[0133] MatrixNPU is a one-dimensional pointer sequence that uses the feature point and corner point values ​​from the FeatureSet.jason data to convert them into a one-dimensional array using the NumPy third-party data processing library in Python and then stores them in MatrixNPU.

[0134] MatrixNPU={ x1 ij ,y1 ij , x2 ij ,y2 ij x3 ij y3 ij ,…,xL ij ,yL ij, x* m1n1 , y* m1n1 ,x* m2n2 ,y* m2n2 , x* m3n3 , y* m3n3 ,…, x* mInI , y* mInI Since the network camera is fixed in its installation, the values ​​in the MatrixNPU fusion matrix do not need to be changed and are treated as fixed parameters in the hardware device.

[0135] Therefore, for edge devices, the input is MPI multi-threaded shared memory accelerated video frames, which are then fused by MatrixNPU matrix inference output to obtain the multi-channel fused video frame stream CombineVideoStream.

[0136] In step S6, the edge computing device still uses MPI multi-threaded parallel method to pull video stream, establishes OpenCV operator to decode video to obtain video frames, performs affine transformation, reads JSON data obtained by server device, constructs NPU fusion matrix through edge feature corner points, replaces image pixel data before affine transformation with feature point data corresponding to JSON file coordinates in the image matrix after affine transformation, and fills the NPU fusion matrix with the above-processed data according to the edge pixel range of NPU fusion matrix to perform edge inference and obtain fused video.

[0137] The video frame of the multi-path video stream fusion adopts ffmpeg-rockip to perform fast H265 / H264 hardware compression and encoding, and performs efficient streaming output through rtsp / rtmp.

[0138] As described above, the application can greatly improve the pull streaming speed of a conventional camera by implementing memory sharing of video frame images of the pull stream through MPI multi-thread parallel computing technology; meanwhile, the application solves the problems of incompatible video protocols and non-uniform video encoding formats through OpenCV library; finally, the application performs edge deployment, efficiently uses NPU hardware inference module, constructs NPU fusion matrix (Neural Processing Unit Fusion Matrix), and accelerates the video fusion process. The application has high performance-price ratio, and can deploy the video fusion method on a large number of edge computing platforms to realize efficient, stable and low-cost video fusion processing.

[0139] Exemplary apparatus

[0140] Figure 6 is a structural schematic diagram of a video fusion device based on an edge computing platform provided by an exemplary embodiment of the application. As shown in Figure 6 , the embodiment includes:

[0141] The initialization module is configured to access a server device by a plurality of cameras; each camera corresponds to an IP protocol address;

[0142] The data acquisition module is configured to establish a first OpenCV operator on the server device; determine the IP protocol addresses of the cameras that need to perform video fusion using the first OpenCV operator, and determine the video streams to be fused corresponding to the IP protocol addresses of the cameras that need to perform video fusion, and use MPI multi-thread to pull each video stream to be fused;

[0143] The decoding module is configured to establish a second OpenCV operator on the server device; and use the second OpenCV operator to decode all the video streams to be fused;

[0144] The affine module is configured to perform affine transformation on each video frame in each decoded video stream respectively, determine the energy-related matrix corresponding to each video frame based on the affine transformation result corresponding to each video frame, determine the feature points and edge feature corner points corresponding to each video frame based on the energy-related matrix corresponding to each video frame, and obtain the feature point data of the feature points corresponding to each video frame and the edge feature corner point data corresponding to the edge feature corner points based on the affine transformation result corresponding to each video frame.

[0145] Storage module: Configured to store the feature points and edge feature points corresponding to each video frame into the same Jason file;

[0146] The fusion module is configured to connect multiple cameras to the edge computing platform. The edge computing platform imports the Jason data file and replaces the pixel data in the video frame corresponding to the feature point with the feature point data corresponding to the feature point in the Jason data file. Based on the replaced video frame data, the NPU fusion matrix corresponding to each video frame is filled. Based on each NPU fusion matrix, the edge pixel range of each video frame is determined. Based on the edge pixel range of each video frame, edge inference is performed to obtain the fused video.

[0147] Exemplary electronic device

[0148] Figure 7 This is the structure of an electronic device 70 provided in an exemplary embodiment of the present invention. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them. Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Figure 7 As shown, the electronic device includes one or more processors 71 and memory 72.

[0149] The processor 71 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0150] The memory 72 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 71 may execute the program instructions to implement the methods of the software programs of the various embodiments of this disclosure described above, and / or other desired functions. In one example, the electronic device may also include an input device 73 and an output device 74, these components being interconnected via a bus system and / or other forms of connection mechanisms (not shown). Furthermore, the input device 73 may include, for example, a keyboard, a mouse, etc. The output device 74 may output various information to the outside. The output device 74 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0151] Of course, in order to simplify, Figure 7 Only some of the components of the electronic device related to the present disclosure are shown in FIG. 1, and components such as a bus, an input / output interface, and the like are omitted. In addition, the electronic device can further include any other appropriate components according to a specific application.

[0152] Exemplary computer program product and computer readable storage medium

[0153] In addition to the above-mentioned method and device, embodiments of the present disclosure can also be a computer program product including computer program instructions, which, when executed by a processor, cause the processor to perform steps of the methods according to various embodiments of the present disclosure described in the above “Exemplary Method” section of the specification.

[0154] The computer program product can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0155] In addition, embodiments of the present disclosure can also be a computer readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform steps of the methods according to various embodiments of the present disclosure described in the above “Exemplary Method” section of the specification.

[0156] The computer readable storage medium can be any combination of one or more non-transitory media. The non-transitory medium can be a non-transitory signal medium or a non-transitory storage medium. The non-transitory storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the non-transitory storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0157] The above describes the basic principles of the present disclosure in conjunction with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present disclosure are merely examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as necessary for each embodiment of the present disclosure. In addition, the above specific details of the disclosure are only for the purpose of example and for the purpose of understanding, and are not limiting, and the above details do not limit the present disclosure to be necessarily implemented with the above specific details.

[0158] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between each embodiment can be understood by referring to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts can be understood by referring to the part of the method embodiment.

[0159] The block diagrams of the devices, apparatuses, equipment, systems involved in the present disclosure are only exemplary examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.

[0160] The methods and devices of the present disclosure can be implemented in many ways. For example, the methods and devices of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present disclosure are not limited to the above specific description, unless otherwise specifically described. In addition, in some embodiments, the present disclosure can also be implemented as programs recorded in recording media, which include machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers the recording media storing the programs for executing the method according to the present disclosure.

[0161] It is also important to note that the devices, apparatuses and methods described in the disclosure can be embodied in a variety of other forms, including but not limited to oral, written, and / or visual forms. It is also noted that scripting languages can be used in the devices, apparatuses and methods described in the disclosure. It is also noted that the devices, apparatuses and methods described in the disclosure can be implemented in software and / or hardware. It is also noted that the devices, apparatuses and methods described in the disclosure can be implemented in a variety of ways, including as a computer program product, as a computer-implemented process, and / or as an apparatus. Furthermore, the described aspects can be implemented by hardware, software, firmware or any combination thereof. If implemented in software, the described aspects can be stored in or implemented with one or more computer-readable storage mediums. Any of the described aspects can also be embodied as computer-readable storage mediums containing instructions that, when executed by a processor, can cause the processor to carry out the described aspects. The computer-readable storage medium can be a storage medium such as any of the storage media described above. The computer-readable storage medium can be tangible and / or non-transitory. The computer-readable storage medium can be a recording medium. The computer-readable storage medium can also be articles of manufacture.

[0162] The above description has been presented for the purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the disclosure to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those of ordinary skill in the art will appreciate a variety of modifications, alternatives, permutations, additions, and sub-combinations.

Claims

1. A video fusion method based on an edge computing platform, characterized in that, The method includes: Step S1: Connect the multiple video streams captured by the multiple cameras to the server device; each camera corresponds to an IP address; Step S2: Establish the first OpenCV operator on the server device; use the first OpenCV operator to determine the IP protocol address corresponding to each camera that needs to be fused, and determine the video stream to be fused corresponding to the IP protocol address of each camera that needs to be fused; use MPI multi-threading to pull each video stream to be fused in parallel. Step S3: Establish a second OpenCV operator on the server device; use the second OpenCV operator to decode all video streams to be merged; Step S4: Perform affine transformation on each video frame in each decoded video stream, and determine the energy correlation matrix corresponding to each video frame based on the affine transformation result; determine the feature points and edge feature corner points corresponding to each video frame based on the energy correlation matrix; the feature point data of the feature points and the edge feature corner point data corresponding to each video frame are obtained based on the affine transformation result of each video frame. Step S5: Store the feature points and edge feature points corresponding to each video frame into the same Jason file; Step S6: Multiple cameras are connected to the edge computing platform. The edge computing platform imports the Jason data file and replaces the pixel data in the video frame corresponding to the feature point with the feature point data corresponding to the feature point in the Jason data file. The NPU fusion matrix corresponding to each video frame is filled based on the replaced video frame data. The edge pixel range of each video frame is determined based on each NPU fusion matrix. Edge inference is performed based on the edge pixel range of each video frame to obtain the fused video.

2. The method as described in claim 1, characterized in that, In step S4, affine transformation is performed on each video frame in each decoded video stream, including: For each video frame in each decoded video stream, perform the following operations: Step S41: Obtain the first vertices at three fixed positions in the video frame src, and determine the second vertices in the target video frame dst that correspond to the three first vertices respectively through affine transformation; The formula for affine transformation is: = , in Let be the parameter matrix of the affine transformation, ( , ) represents the pixel data in the src of the video frame. () represents the pixel data in the target video frame dst after affine transformation; Step S42: Based on the first vertex and the second vertex, perform an affine transformation on the video frame src to obtain the target video frame dst.

3. The method as described in claim 2, characterized in that, In step S4, the energy correlation matrix corresponding to each video frame is determined based on the affine transformation results, including: For each target video frame dst after affine transformation, perform the following operations: Obtain each pixel p in the src of the video frame corresponding to the target video frame dst. ij Determine p ij The pixel p* corresponding to the target video frame dst ij The correlation; The energy correlation matrix Emat_P corresponding to this video frame ij for: Ematrix_P ij = Corr((H_frameA[p ij ]) 2 , (H_frameB[p ij ]) 2 ), Where i = 0, 1, 2, ..., witdth; j = 0, 1, 2, 3, ..., height, witdth and height are the width and height of the video frame src, respectively; Corr(·) is the correlation calculation, H_frameA[p ij ]= src ( , ), H_frameB[p ij ])=dst( ).

4. The method as described in claim 3, characterized in that, In step S4, the feature points and edge feature corner points corresponding to each video frame are determined based on the energy correlation matrix corresponding to each video frame, including: For each video frame, perform the following operations: Determine the energy value threshold corresponding to the video frame, and use the pixels in the target video frame corresponding to the matrix elements whose matrix element values ​​exceed the energy value threshold in the energy correlation matrix of the video frame as feature points. Construct a feature operator fusion list that includes all feature points; Feature points that meet the preset conditions in the feature operator fusion list are taken as edge feature corner points, and all edge feature corner points are written into the corner feature list; where the preset conditions are edge conditions.

5. The method according to any one of claims 1-4, characterized in that, In step S1, multiple cameras of the same model are selected, and the cameras and server devices are connected to a switch via Ethernet.

6. A video fusion device based on an edge computing platform, characterized in that, The device includes: Initialization module: Configured to connect multiple video streams captured by multiple cameras to the server device; each camera corresponds to an IP address; Data acquisition module: configured to establish the first OpenCV operator on the server device; use the first OpenCV operator to determine the IP protocol address corresponding to each camera that needs to be fused, and determine the video stream to be fused corresponding to the IP protocol address of each camera that needs to be fused, and use MPI multi-threading to pull each video stream to be fused in parallel; Decoding module: Configured to establish a second OpenCV operator on the server device; use the second OpenCV operator to decode all video streams to be merged; Affine module: configured to perform affine transformation on each video frame in each decoded video stream, determine the energy correlation matrix corresponding to each video frame based on the affine transformation result, determine the feature points and edge feature corner points corresponding to each video frame based on the energy correlation matrix, and obtain the feature point data and edge feature corner point data corresponding to each video frame based on the affine transformation result of each video frame; Storage module: Configured to store the feature points and edge feature points corresponding to each video frame into the same Jason file; The fusion module is configured to connect multiple cameras to the edge computing platform. The edge computing platform imports the Jason data file and replaces the pixel data in the video frame corresponding to the feature point with the feature point data corresponding to the feature point in the Jason data file. Based on the replaced video frame data, the NPU fusion matrix corresponding to each video frame is filled. Based on each NPU fusion matrix, the edge pixel range of each video frame is determined. Based on the edge pixel range of each video frame, edge inference is performed to obtain the fused video.

7. A computer-readable storage medium, characterized in that, The storage medium stores a plurality of instructions; the plurality of instructions are loaded by a processor and executed as described in any one of claims 1-5.

8. An electronic device, characterized in that, The electronic device includes: A processor is used to execute multiple instructions; Memory, used to store multiple instructions; The plurality of instructions are to be stored in the memory and loaded by the processor and executed as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-layer video data fusion method and system based on FPGA

    CN118784893A

  • Building intelligent early warning method based on 5G Internet of Things integrated video monitoring

    CN119380516A