A semantic correlation multi-source monitoring video fusion analysis method and system

By constructing a multi-source fusion monitoring network system and an image similarity measurement model, the problem of fusion of monitoring videos from different sources was solved, enabling real-time tracking and intelligent monitoring of targets and avoiding monitoring blind spots.

CN116109965BActive Publication Date: 2026-01-02JINAN RAILWAY INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211594911.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2026-01-02
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

Existing video surveillance systems cannot effectively integrate surveillance videos from different sources, resulting in inaccurate target positioning and blind spots, and failing to achieve intelligent real-time target monitoring.

Method used

A multi-source fusion monitoring network system is constructed, which uploads and integrates monitoring images via a wireless network. Image registration and fusion are performed using an image similarity measurement model and a registration model, and registration parameters are optimized to achieve registration and integration of multi-source monitoring videos and target tracking.

Benefits of technology

It effectively integrates surveillance videos from different sources, avoids blind spots in monitoring, improves the accuracy and intelligence of target monitoring, and enables real-time tracking of the targets to be monitored.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116109965B_ABST
    Figure CN116109965B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of monitoring video fusion, and discloses a semantic correlation multi-source monitoring video fusion analysis method and system, which comprises the following steps: a multi-source mixed target global similarity measurement model is constructed, and monitoring images captured by monitoring video devices and target images to be monitored are input into the model; image similarity measurement values of different monitoring images and the target images to be monitored are taken as probability characteristics of the target monitored by the different monitoring images, and the probability characteristics of the monitoring images are correlated to determine whether the target to be monitored is monitored; a registration model is used to perform initial matching on the target images to be monitored and the monitoring images of the target monitored; an improved quantum particle swarm optimization algorithm is used to optimize parameters of the registration model; the monitoring images are input into an optimal registration model for registration fusion, and the position of the target to be monitored at different moments is tracked. The application realizes fusion of monitoring videos from different sources and real-time monitoring and tracking of the target to be monitored.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of monitoring video fusion, and particularly relates to a semantic correlation multi-source monitoring video fusion analysis method and system. BACKGROUND

[0002] In recent years, the protection of life and property safety gradually enters the vision of everyone, and the traditional video monitoring system has been unable to meet the demand of people for real-time and intelligence. The existing video monitoring analysis mainly uses single-source monitoring video analysis, and cannot solve the joint analysis of different source monitoring videos, that is, quickly positioning the time position of the same target subject in different source monitoring videos and intelligently extracting and analyzing. At the same time, the single source monitoring video has a monitoring dead angle, and there may be a target missing situation. In view of the problem, the present application provides a semantic correlation multi-source monitoring video fusion analysis method to improve the intelligent level of video monitoring and the target monitoring accuracy. SUMMARY

[0003] Therefore, the application provides a semantic correlation multi-source monitoring video fusion analysis method, which aims to 1) realize the fusion of different source monitoring videos, prestore a target image to be monitored in a network layer by constructing a multi-source fusion monitoring network system, wherein a physical layer is composed of a plurality of monitoring video devices, the monitoring video devices can capture monitoring video images of a monitoring area as monitoring images in real time, the captured monitoring images are uploaded to the network layer through a wireless network forwarding device, the network layer is used for associating and integrating all the captured monitoring images, and whether the target to be monitored is monitored is determined based on context probability features of the associated monitoring images at the same time, if the target to be monitored is monitored, then adjacent monitoring images corresponding to the same time are input into a registration model for registration fusion, registration integration of the multi-source monitoring videos is realized, three-dimensional positions of the target to be monitored at different times are obtained, and real-time monitoring and tracking of the target to be monitored are realized; 2) similarity calculation functions in image gray domain and frequency domain are respectively constructed, image similarity measurement values of different monitoring images and the target image to be monitored are determined based on the similarity in the image gray domain and the frequency domain, the image similarity measurement values are used as probability features of the target monitored by the monitoring images, and the associated probability of the target monitored in the monitoring images is calculated by associating the probability features of the adjacent monitoring images and the adjacent time monitoring images, if the associated probability reaches a threshold value, then it is indicated that the target is monitored in the current monitoring image, and probability integration of the monitoring target of the multi-source monitoring images is realized; 3) initial matching is performed on the target image to be monitored and the monitoring images of the target monitored, initial registration parameters are obtained, a registration model containing a plurality of groups of registration parameters is constructed, the monitoring images are registration fused by using the registration parameters, in the registration parameter optimization process, the initial registration parameters based on random interference are used as initial values of algorithm iteration, and then the registration parameters of the monitoring images captured by the adjacent monitoring video devices are iteratively obtained, the different monitoring images are registration fused by using the obtained registration parameters, multi-source fused monitoring videos are obtained, and then the positions of the target to be monitored at different times are obtained.

[0004] To achieve the above object, in one aspect, the application provides a semantic correlation multi-source monitoring video fusion analysis method, which comprises the following steps:

[0005] S1: constructing a multi-source fusion video monitoring network system, networking and integrating a plurality of monitoring video devices and pre-storing a target image to be monitored;

[0006] S2: constructing a multi-source mixed target global similarity measurement model, and inputting the monitoring images captured by the monitoring video devices and the target image to be monitored into the model, wherein the model takes the target image to be monitored and different monitoring images as inputs,

[0007] and takes image similarity measurement values as outputs;

[0008] S3: taking the image similarity measurement value of the different monitoring image and the image of the target to be monitored as the probability feature of the target monitored by the different monitoring image, and judging whether the target to be monitored is monitored or not by the probability feature of the monitoring image;

[0009] S4: if the target to be monitored is monitored, performing initial matching on the target to be monitored image and the monitoring image of the monitored target by using the registration model;

[0010] S5: optimizing the registration model parameters by using the improved quantum particle swarm optimization algorithm to obtain the optimal registration model;

[0011] S6: selecting the monitoring image captured by the nearest monitoring video device of the monitoring video device of the monitored target, inputting the selected monitoring image into the optimal registration model for registration fusion to obtain the three-dimensional position of the target to be monitored at different time.

[0012] As a further improved method of the application:

[0013] Optionally, the S1 step of constructing the multi-source fusion video monitoring network system comprises:

[0014] The multi-source fusion monitoring network system comprises a network layer and a physical layer, the target to be monitored image is preset in the network layer, the physical layer is composed of a plurality of monitoring video devices, the monitoring video devices can capture video images of the monitoring area as monitoring images in real time, the captured monitoring images are uploaded to the network layer through a wireless network forwarding device, the network layer is used for associating all the captured monitoring images, and whether the target to be monitored is monitored or not is judged based on the context probability feature of the associated monitoring images; if the target to be monitored is monitored, the corresponding monitoring image is input into the registration model for registration fusion to obtain the three-dimensional position of the target to be monitored at different time, and real-time monitoring and tracking of the target to be monitored are realized.

[0015] Optionally, the S2 step of constructing the multi-source mixed target global similarity measurement model comprises:

[0016] The multi-source mixed target global similarity measurement model takes the target to be monitored image and different monitoring images as input and takes the image similarity measurement value as output, and the image similarity measurement process based on the target global similarity measurement model comprises:

[0017] S21: the target global similarity measurement model receives the preset target to be monitored image Q0 and the monitoring image Q k (t), wherein the monitoring image Q k(t) represents a monitoring image taken by the kth monitoring video device in the multi-source fusion video monitoring network system at time t, k∈[1, K], K represents the total number of monitoring video devices in the multi-source fusion video monitoring network system, and the kth monitoring video device and the k+1th monitoring video device are adjacent monitoring video devices;

[0018] S22: Fourier transform processing is used to obtain the to-be-monitored target image Q0and the monitoring image Q k (t) and F k (t) and extract the frequency domain phase information of the frequency domain representation results, respectively:

[0019]

[0020]

[0021] wherein:

[0022] represents the amplitude information of the frequency domain representation F0, represents the phase information of the frequency domain representation F0;

[0023] represents the amplitude information of the frequency domain representation F k (t), represents the phase information of the frequency domain representation F k (t).

[0024] wherein the phase information of any ith pixel in the to-be-monitored target image Q0is All the pixel numbers of the images are N, and the frequency domain phase average of the to-be-monitored target image Q0is the frequency domain phase average of the monitoring image Q k (t) is represents the phase information of any ith pixel in the monitoring image Q k (t).

[0025] S23: the frequency domain similarity of the to-be-monitored target image Q0and the monitoring image Q k (t) is calculated

[0026]

[0027]

[0028]

[0029]

[0030] wherein:

[0031] N represents the total number of pixels of the image;

[0032] respectively represent the to-be-monitored target image Q0and the monitoring image Q k the phase information standard deviation of (t), represent the phase information covariance of both;

[0033] S24: convert the to-be-monitored target image Q0and the monitoring image Q k (t) into gray-scale images G0and G k (t) respectively, and calculate the gray-scale domain similarity sim_b(Q0, Q k (t)) of the to-be-monitored target gray-scale image G0and the monitoring gray-scale image G k (t):

[0034]

[0035] wherein:

[0036] μ0 represents the average pixel gray-scale value of the to-be-monitored target gray-scale image G0, and σ0 represents the gray-scale value standard deviation of the to-be-monitored target gray-scale image G0;

[0037] μ kt represents the average pixel gray-scale value of the monitoring image Q k (t), and σ kt represents the gray-scale value standard deviation of the monitoring image Q k (t);

[0038] S25: take the average value of the frequency domain similarity and the gray-scale domain similarity sim_b(Q0, Q k (t)) as the similarity measurement value sim(Q0, Q k (t)) of the to-be-monitored target image Q0and the monitoring image Q k (t), and perform normalization processing on the similarity measurement value to obtain the normalized similarity measurement value sim'(Q0, Q k (t), and the formula of the normalization processing is:

[0039]

[0040] wherein:

[0041] max t represents the maximum similarity measurement value of the monitoring images taken by the K monitoring video devices and the to-be-monitored target image at the t time;

[0042] min t represents the minimum similarity measurement value of the monitoring images taken by the K monitoring video devices and the to-be-monitored target image at the t time.

[0043] Optionally, the step S2 of inputting the monitoring image captured by the monitoring video device and the target image to be monitored into the target global similarity measure model comprises:

[0044] inputting the monitoring images captured by different monitoring video devices at different time and the target image to be monitored into the target global similarity measure model to obtain a set of similarity measure values of the captured monitoring images and the target image to be monitored:

[0045] {sim'(Q0, Q k (t))|k∈[1,K]},t∈[t0,t L ]

[0046] wherein:

[0047] {sim'(Q0, Q k (t))|k∈[1,K]} represents a set of similarity measure values of the captured monitoring images and the target image to be monitored at time t;

[0048] t0 represents an initial time for monitoring the target image to be monitored, and t L represents a cut-off time for monitoring the target image to be monitored.

[0049] Optionally, the step S3 of judging whether the target to be monitored is monitored according to the probability feature of the monitoring image comprises:

[0050] taking the image similarity measure value of different monitoring images and the target image to be monitored as the probability feature of the target monitored by the different monitoring images, then for the similarity measure value sim'(Q0, Q k (t)) of any monitoring image and the target image to be monitored calculated at time t, the corresponding probability feature is p k (t) = sim'(Q0, Q k (t));

[0051] the process of judging whether the target to be monitored is monitored according to the probability feature of the monitoring image comprises:

[0052] S31: traversing the set of probability features {p k (t)|k∈[1,K]} at time t to obtain the probability feature p k′ (t) with the largest probability feature and the k'th monitoring video device corresponding thereto;

[0053] S32: calculating the position distance between the k'th monitoring video device and the adjacent monitoring video device, and selecting the adjacent monitoring video device with a position distance less than a position threshold as the nearest monitoring video device of the k'th monitoring video device;

[0054] S33: constructing a monitoring video device set with the k'th monitoring video device and its nearest monitoring video devices, and calculating a probability feature product of monitoring images taken by any two monitoring video devices in the monitoring video device set at time t, wherein the weight of the probability feature product is the distance between the two monitoring video devices, and the sum of the weighted probability feature products is taken as the monitoring conflict θ;

[0055] S34: calculating the probability P that the k'th monitoring video device takes the target to be monitored at time t k′ (t):

[0056]

[0057] wherein:

[0058] Ω(k') represents the monitoring video device set corresponding to the k'th monitoring video device, and c is a monitoring video device therein;

[0059] p c (t) represents the probability feature of the image taken by the monitoring video device c at time t, and Δt represents the time interval;

[0060] S35: if P k′ (t) ≥ α, it indicates that the k'th monitoring video device monitors the target to be monitored at time t, wherein α represents a time threshold; otherwise, returning to step S31, selecting the probability feature next to the largest probability feature and the corresponding monitoring video device, and if both are less than the time threshold, it indicates that the target to be monitored is not monitored at time t.

[0061] Optionally, the initial matching of the target image to be monitored and the monitoring image of the monitored target in the S4 step utilizes a registration model, comprising:

[0062] If the k"th monitoring video device monitors the target at time t, the corresponding monitoring image is Q k″ (t), and the initial matching of the target image to be monitored Q0 and the monitoring image Q k″ (t) of the monitored target utilizes a registration model, wherein the registration model contains a plurality of sets of registration parameters, and the registration parameters are used for registration and fusion processing of the monitoring image to obtain three-dimensional feature information of the target to be monitored in the monitoring scene;

[0063] The initial matching process is:

[0064] S41: extracting SIFT feature points of the target image to be monitored Q0 and the monitoring image Q k″ (t) respectively, and constructing a 64-dimensional SIFT feature descriptor for each feature point;

[0065] S42: calculate the cosine similarity between the SIFT feature descriptors in the target image Q0 to be monitored and the monitoring image Q k″ (t), and the matching descriptors of the SIFT feature descriptors in the target image Q0 to be monitored are the SIFT feature descriptors in the monitoring image Q k″ (t) with the minimum cosine similarity to the SIFT feature descriptors in the target image Q0 to be monitored;

[0066] S43: obtain the initial registration parameters s0 of the SIFT feature descriptors in the target image Q0 to be monitored and the matching descriptors by using the least square method, and register the monitoring image Q k″ (t) using the initial registration parameters s0: s0SIFT k″ (t), wherein SIFT k″ (t) represents the SIFT feature descriptors of the monitoring image Q k″ (t).

[0067] Optionally, the registration model parameters are optimized in the S5 step using the improved quantum particle swarm optimization algorithm, comprising:

[0068] traverse the monitoring image {Q k″,1 (t), Q k″,2 (t),..., Q k″,u (t),..., Q k″,U (t)} taken by the nearest neighboring monitoring video device of the k" monitoring video device at time t, wherein Q k″,u (t) represents the monitoring image taken by the nearest neighboring u-th monitoring video device corresponding to the k" monitoring video device at time t, and the SIFT feature descriptors of the traversed monitoring images are extracted;

[0069] The registration parameters of different monitoring images are optimized using the improved quantum particle swarm optimization algorithm, and all the optimized registration parameters constitute an optimal registration model, and the optimization process of the quantum particle swarm optimization algorithm is as follows:

[0070] S51: initialize to generate a plurality of particles, and each particle z constitutes a set of registration parameters of different monitoring images wherein z 0 represents the initial particle z, represents the registration parameters corresponding to the monitoring image Q k″,u (t), and each registration parameter generated by the initialization is the generated result of adding random disturbance to the initial registration parameter s0;

[0071] S52: let the current iteration number of the algorithm be r, the initial value of r be 1, and the maximum value be MAX;

[0072] S53: construct a fitness function of quantum particle swarm, the fitness function is:

[0073]

[0074] Wherein:

[0075] SIFT k″,u (t) represents the SIFT feature descriptor of the monitoring image Q k″,u (t), SIFT0 represents the SIFT feature descriptor in the target image Q0 to be monitored;

[0076] The registration parameters of the particle z obtained by the L-1th iteration are represented by f(z L-1 ) represents the fitness function value of the particle z obtained by the L-1th iteration;

[0077] C(·) represents a cosine similarity calculation function;

[0078] The particle with the maximum fitness function value of the L-1th iteration is taken as the optimal particle of the L-1th iteration

[0079] S54: iterate any particle z to obtain the Lth iteration particle z L :

[0080] z L =z L-1 +β|z L-1,* -z L-1 |×rand(0,1)

[0081]

[0082] Wherein:

[0083] z L-1 represents the particle of the L-1th iteration, and rand(0,1) represents a random number between 0 and 1;

[0084] ω represents a learning rate, which is set to 0.8;

[0085] S55: calculate the fitness function value of each particle after the Lth iteration, if L=MAX, the registration parameters corresponding to the particle with the maximum fitness function are taken as the optimized registration parameters of different monitoring images, otherwise L=L+1, return to step S54.

[0086] Optionally, in the S6 step, the monitoring image taken by the nearest monitoring video device of the monitoring video device monitoring the target is selected, and the selected monitoring image is input into the optimal registration model for registration fusion, comprising:

[0087] The monitoring image photographed by the most adjacent monitoring video device of the monitoring video device monitoring the target at any moment is selected, the selected monitoring image is input into the optimal registration model, each monitoring image is registered and fused by using the registration parameter corresponding to the monitoring image, the monitoring image monitoring the target is registered and fused by using the initial registration parameter, the position of the target to be monitored after registration and fusion is taken as the three-dimensional position of the target to be monitored in the monitoring scene, and the three-dimensional feature information and three-dimensional position of the target to be monitored at any moment in the monitoring scene are obtained.

[0088] To solve the above problems, in another aspect, the application further provides a semantic correlation multi-source monitoring video fusion analysis system capable of realizing the above method, the system comprising:

[0089] The monitoring judgment device is used for constructing a multi-source mixed target global similarity measurement model, inputting the monitoring image photographed by the monitoring video device and the target image to be monitored into the model, taking the image similarity measurement value of different monitoring images and the target image to be monitored as the probability feature of the target monitored by different monitoring images, and judging whether the target to be monitored is monitored according to the probability feature of the monitoring image;

[0090] The initial registration device is used for performing initial matching on the target image to be monitored and the monitoring image monitoring the target by using the registration model;

[0091] The registration fusion module is used for optimizing the registration model parameter by using the improved quantum particle swarm optimization algorithm, obtaining the optimal registration model, selecting the monitoring image photographed by the most adjacent monitoring video device of the monitoring video device monitoring the target, inputting the selected monitoring image into the optimal registration model for registration fusion, and obtaining the three-dimensional position of the target to be monitored at different moments.

[0092] To solve the above problems, the application further provides an electronic device, comprising:

[0093] The memory stores at least one instruction; and

[0094] The processor executes the instruction stored in the memory to realize the semantic correlation multi-source monitoring video fusion analysis method described above.

[0095] To solve the above problems, the application further provides a computer readable storage medium, the computer readable storage medium stores at least one instruction, and the at least one instruction is executed by the processor in the electronic device to realize the semantic correlation multi-source monitoring video fusion analysis method described above.

[0096] Compared with the prior art, the application provides a semantic correlation multi-source monitoring video fusion analysis method and system, which has the following advantages:

[0097] Firstly, the application provides a fusion method for different source monitoring videos, a multi-source fusion monitoring network system is constructed, and a target image to be monitored is preset in the network layer, wherein the physical layer is composed of a plurality of monitoring video devices, the monitoring video devices can capture monitoring video images of a monitoring area as monitoring images in real time, the captured monitoring images are uploaded to the network layer through a wireless network forwarding device, the network layer is used for associating and integrating all the captured monitoring images, and whether the target to be monitored is monitored is determined based on context probability features of the associated monitoring images at the same time, so as to avoid the monitoring dead angle problem caused by a single video source, if the target to be monitored is monitored, then the adjacent monitoring images corresponding to the same time are input into a registration model for registration fusion, registration and integration of the multi-source monitoring videos are realized, three-dimensional positions of the target to be monitored at different times are obtained, and real-time monitoring and tracking of the target to be monitored are realized.

[0098] Meanwhile, the application provides an image similarity measurement method, and the image similarity measurement process based on the target global similarity measurement model is as follows: the target global similarity measurement model receives a preset target image to be monitored Q0 and a monitoring image Q k (t), wherein the monitoring image Q k (t) represents a monitoring image captured by the kth monitoring video device at the t time in the multi-source fusion video monitoring network system, k [1, K], K represents the total number of monitoring video devices in the multi-source fusion video monitoring network system, and the kth monitoring video device and the k+1th monitoring video device are adjacent monitoring video devices; Fourier transform processing is used to obtain frequency domain representations F0 and F k (t) of the target image to be monitored Q0 and the monitoring image Q k (t) respectively, and frequency domain phase information of the frequency domain representation results is extracted respectively.

[0099]

[0100]

[0101] Wherein: F0 represents amplitude information of the frequency domain representation F0, F0 represents phase information of the frequency domain representation F0; F represents amplitude information of the frequency domain representation F k (t), F represents phase information of the frequency domain representation F k (t); wherein the phase information of any ith pixel in the target image to be monitored Q0 is The frequency domain phase mean value of the target image Q0 to be monitored is The monitoring image Q k The frequency domain phase mean value of the target image Q0 to be monitored is The phase information of any i-th pixel in the monitoring image Q k (t) is represented; the frequency domain similarity between the target image Q0 to be monitored and the monitoring image Q k (t) is calculated

[0102]

[0103]

[0104]

[0105]

[0106] Wherein: N represents the total number of pixels of the image; The phase information standard deviation of the target image Q0 to be monitored and the monitoring image Q k (t) is represented respectively, The phase information covariance of the target image Q0 to be monitored and the monitoring image Q k (t) is represented respectively; the target image Q0 to be monitored and the monitoring image Q k (t) are converted into gray images G0 and G k (t) respectively, and the gray domain similarity sim_b(Q0,Q k (t)) between the target gray image G0 and the monitoring gray image G

[0107]

[0108] Wherein: μ0 represents the average pixel gray value of the target gray image G0 to be monitored, σ0 represents the gray value standard deviation of the target gray image G0 to be monitored; μ k,t The average pixel gray value of the monitoring image Q k (t) is represented, and σ k,t The gray value standard deviation of the monitoring image Q k (t) is represented; the average value of the frequency domain similarity And the gray domain similarity sim_b(Q0,Q k (t)) is taken as the similarity measurement value sim(Q0,Q k (t)) of the target image Q0 to be monitored and the monitoring image Q k (t), and the similarity measurement value is normalized to obtain the normalized similarity measurement value sim'(Q0,Q k (t)), and the formula of the normalization processing is:

[0109]

[0110] max t represents the maximum similarity measure value of the K monitoring images and the target image at time t; min t represents the minimum similarity measure value of the K monitoring images and the target image at time t. The scheme constructs similarity calculation functions in the image gray domain and the frequency domain, respectively, judges the image similarity measure value of different monitoring images and the target image based on the similarity in the image gray domain and the frequency domain, takes the image similarity measure value as the probability feature of the target monitored by the monitoring image, and calculates the correlation probability of the target monitored in the monitoring image by correlating the probability features of the adjacent monitoring images and the adjacent time monitoring images. If the correlation probability reaches a threshold, it indicates that the current monitoring image monitors the target. Since all monitoring videos have probability features, the monitoring dead angle of a single source monitoring video is avoided, and the target missed detection may exist. The monitoring target probability of the multi-source monitoring image is integrated.

[0111] Finally, the scheme proposes a target video tracking method. The monitoring image captured by the most adjacent monitoring video device corresponding to the monitoring video device that monitors the target at any time is selected. The selected monitoring image is input into the optimal registration model, the registration parameters corresponding to the monitoring image are used to perform registration fusion on each monitoring image, and the initial registration parameters are used to perform registration fusion on the monitoring image that monitors the target. The position of the target after registration fusion is taken as the three-dimensional position of the target in the monitoring scene, and the three-dimensional feature information and the three-dimensional position of the target at any time in the monitoring scene are obtained. The scheme performs initial matching on the target image and the monitoring image that monitors the target, obtains the initial registration parameters, constructs a registration model containing several groups of registration parameters, performs registration fusion on the monitoring image by using the registration parameters, and uses the initial registration parameters based on random interference as the initial value of algorithm iteration in the registration parameter optimization process. Then, the registration parameters of the monitoring images captured by the adjacent monitoring video devices are obtained by iteration. The registration fusion of different monitoring images is performed by using the optimized registration parameters, the multi-source fused monitoring video is obtained, and the position of the target at different times is obtained. The target video tracking is realized. BRIEF DESCRIPTION OF DRAWINGS

[0112] Figure 1 A flowchart of a semantic correlation multi-source monitoring video fusion analysis method provided by an embodiment of the application is shown in the figure.

[0113] Figure 2 A functional module diagram of a semantic correlation multi-source monitoring video fusion analysis system provided by an embodiment of the application is shown in the figure.

[0114] Figure 3 The structural schematic diagram of an electronic device for implementing the semantic correlation multi-source monitoring video fusion analysis method provided by an embodiment of the present application is shown in the figure.

[0115] In the figure: 100, a semantic correlation multi-source monitoring video fusion analysis system; 101, a monitoring and judging device; 102, an initial registration device; 103, a registration fusion module; 1, an electronic device; 10, a processor; 11, a memory; 12, a program; and 13, a communication interface.

[0116] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0117] It should be understood that the specific embodiments described herein are merely intended to explain the present application and not to limit the present application.

[0118] An embodiment of the present application provides a semantic correlation multi-source monitoring video fusion analysis method. The execution subject of the semantic correlation multi-source monitoring video fusion analysis method includes but is not limited to at least one of electronic devices such as a server and a terminal which can be configured to execute the method provided by the embodiment of the present application. In other words, the semantic correlation multi-source monitoring video fusion analysis method can be executed by software or hardware installed in a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to a single server, a server cluster, a cloud server or a cloud server cluster, etc.

[0119] Embodiment 1

[0120] S1: Construct a multi-source fusion video monitoring network system, network and integrate various monitoring video devices and preset target images to be monitored.

[0121] The S1 step of constructing a multi-source fusion video monitoring network system includes:

[0122] The multi-source fusion monitoring network system includes a network layer and a physical layer. The network layer is used to preset target images to be monitored. The physical layer is composed of a plurality of monitoring video devices. The monitoring video devices can capture video images of a monitoring area as monitoring images in real time. The captured monitoring images are uploaded to the network layer through a wireless network forwarding device. The network layer is used to associate all the captured monitoring images and determine whether the target to be monitored is monitored based on the context probability features of the associated monitoring images. If the target to be monitored is monitored, the corresponding monitoring image is input into a registration model for registration fusion to obtain the three-dimensional position of the target to be monitored at different time, thereby realizing real-time monitoring and tracking of the target to be monitored.

[0123] S2: a multi-source mixed target global similarity measure model is constructed, and a monitoring image captured by a monitoring video device and a target image to be monitored are input into the model, the model taking the target image to be monitored and different monitoring images as inputs and taking an image similarity measure value as an output.

[0124] The S2 step of constructing the multi-source mixed target global similarity measure model comprises:

[0125] The multi-source mixed target global similarity measure model is constructed, the model taking the target image to be monitored and different monitoring images as inputs and taking an image similarity measure value as an output, and the image similarity measure process based on the target global similarity measure model comprises:

[0126] S21: the target global similarity measure model receives a preset target image to be monitored Q0 and monitoring images Q k (t), wherein the monitoring images Q k (t) represent monitoring images captured by a kth monitoring video device in a multi-source fusion video monitoring network system at a time t, k [1, K], K represents a total number of monitoring video devices in the multi-source fusion video monitoring network system, and the kth monitoring video device and a (k+1)th monitoring video device are adjacent monitoring video devices.

[0127] S22: Fourier transform processing is used to obtain frequency domain representations F0 and F k (t) of the target image to be monitored Q0 and the monitoring images Q k (t), respectively, and frequency domain phase information of the frequency domain representation results is extracted, respectively.

[0128]

[0129] Wherein:

[0130] represents amplitude information of the frequency domain representation F0, represents phase information of the frequency domain representation F0;

[0131] represents amplitude information of the frequency domain representation F k (t), represents phase information of the frequency domain representation F k (t).

[0132] Wherein, phase information of any ith pixel in the target image to be monitored Q0 is All pixel numbers of the images are N, and the frequency domain phase average of the target image to be monitored Q0 is The frequency domain phase average of the monitoring images Q k (t) is represents the monitoring images Qk (t) the phase information of any i-th pixel in (t);

[0133] S23: Calculate the target image Q0to be monitored and the monitoring image Q k (t) in the frequency domain

[0134]

[0135]

[0136]

[0137]

[0138] wherein:

[0139] N represents the total number of pixels of the image;

[0140] respectively represent the target image Q0to be monitored and the monitoring image Q k (t) the standard deviation of the phase information of (t), represent the covariance of the phase information of both;

[0141] S24: Convert the target image Q0to be monitored and the monitoring image Q k (t) into gray-scale images G0and G k (t), respectively, and calculate the target gray-scale image G0and the monitoring gray-scale image G k (t) in the gray-scale domain; k (t):

[0142]

[0143] wherein:

[0144] μ0represents the average pixel gray-scale value of the target gray-scale image G0to be monitored, and σ0represents the standard deviation of the gray-scale value of the target gray-scale image G0to be monitored;

[0145] μ k,t represents the average pixel gray-scale value of the monitoring image Q k (t), and σ k,t represents the standard deviation of the gray-scale value of the monitoring image Q k (t);

[0146] S25: Calculate the average value of the frequency domain similarity sim_a(Q0,Q k (t)) and the gray-scale domain similarity sim_b(Q0,Q k (t) as the similarity measure value sim(Q0,Qk (t)) and normalizing the similarity measure value to obtain a normalized similarity measure value sim'(Q0, Q k (t)) and normalizing the similarity measure value to obtain a normalized similarity measure value sim'(Q0, Q

[0147]

[0148] wherein:

[0149] max t represents the maximum similarity measure value of the monitoring image captured by the K monitoring video devices and the target image at time t;

[0150] min t represents the minimum similarity measure value of the monitoring image captured by the K monitoring video devices and the target image at time t.

[0151] The step S2 of inputting the monitoring image captured by the monitoring video device and the target image to be monitored into the target global similarity measure model comprises:

[0152] inputting the monitoring images captured by different monitoring video devices at different times and the target image to be monitored into the target global similarity measure model to obtain a set of similarity measure values of the captured monitoring images and the target image to be monitored:

[0153] {sim'(Q0, Q k (t))|k∈[1,K]},t∈[t0,t L ]

[0154] wherein:

[0155] {sim'(Q0, Q k (t))|k∈[1,K]} represents a set of similarity measure values of the captured monitoring image and the target image to be monitored at time t;

[0156] t0 represents the initial time of monitoring the target image to be monitored, and t L represents the cutoff time of monitoring the target image to be monitored.

[0157] S3: taking the image similarity measure value of different monitoring images and the target image to be monitored as the probability feature of the target monitored by different monitoring images, and associating the probability feature of the monitoring image to determine whether the target to be monitored is monitored.

[0158] The step S3 of associating the probability feature of the monitoring image to determine whether the target to be monitored is monitored comprises:

[0159] The image similarity measure value of different monitoring images and the image of the target to be monitored is taken as the probability feature of the target monitored by different monitoring images. The similarity measure value of any monitoring image and the image of the target to be monitored at time t is sim'(Q0, Q k The corresponding probability feature is p k (t) = sim'(Q0, Q k (t));

[0160] The flow of judging whether the target to be monitored is monitored by the probability feature of the associated monitoring image is as follows:

[0161] S31: traverse the probability feature set {p k (t) | k ∈ [1, K]} at time t to obtain the probability feature p k′ (t) with the maximum probability and the corresponding k'th monitoring video device;

[0162] S32: calculate the position distance of the k'th monitoring video device and the adjacent monitoring video device, and select the adjacent monitoring video device with a position distance less than a position threshold as the nearest adjacent monitoring video device of the k'th monitoring video device;

[0163] S33: construct a monitoring video device set with the k'th monitoring video device and its nearest adjacent monitoring video device, and calculate the probability feature product of the monitoring images captured by any two monitoring video devices in the monitoring video device set at time t, wherein the weight of the probability feature product is the distance of the two monitoring video devices, and the sum of the weighted probability feature products is taken as the monitoring conflict θ;

[0164] S34: calculate the probability P k′ (t) that the k'th monitoring video device captures the target to be monitored at time t:

[0165]

[0166] Wherein:

[0167] Ω(k') represents the monitoring video device set corresponding to the k'th monitoring video device, and c is a monitoring video device therein.

[0168] p c (t) represents the probability feature of the image captured by the monitoring video device c at time t, and Δt represents the time interval.

[0169] S35: if P k′(t)≥α, it indicates that the k'th monitoring video device monitors the target to be monitored at time t, wherein α represents a time threshold; otherwise, return to step S31, select the probability feature next to the smallest probability feature and the corresponding monitoring video device, and if they are all less than the time threshold, it indicates that the target to be monitored is not monitored at time t.

[0170] S4: If the target to be monitored is monitored, an initial matching is performed on the target to be monitored image and the monitoring image of the monitored target by using a registration model.

[0171] The initial matching of the target to be monitored image and the monitoring image of the monitored target by using the registration model in the S4 step comprises:

[0172] If the k''th monitoring video device monitors the target at time t, the corresponding monitoring image is Q k″ (t), an initial matching is performed on the target to be monitored image Q0 and the monitoring image Q k″ (t) of the monitored target by using a registration model, wherein the registration model comprises a plurality of sets of registration parameters, and the monitoring image is fused by using the registration parameters to obtain three-dimensional feature information of the target to be monitored in the monitoring scene;

[0173] The initial matching process is as follows:

[0174] S41: SIFT feature points of the target to be monitored image Q0 and the monitoring image Q k″ (t) are extracted respectively, and a 64-dimensional SIFT feature descriptor is constructed for each feature point;

[0175] S42: Cosine similarity of the SIFT feature descriptors in the target to be monitored image Q0 and the monitoring image Q k″ (t) is calculated, and the SIFT feature descriptor with the minimum cosine similarity between the SIFT feature descriptors in the target to be monitored image Q0 and the monitoring image Q k″ (t) is taken as the matching descriptor of the SIFT feature descriptor in the target to be monitored image Q0.

[0176] S43: An initial registration parameter s0 of the SIFT feature descriptor in the target to be monitored image Q0 and the matching descriptor is solved by using a least square method, and the monitoring image Q k″ (t) is registered by using the initial registration parameter s0: s0SIFT k″ (t), wherein SIFT k″ (t) represents the SIFT feature descriptor of the monitoring image Q k″ (t).

[0177] S5: An improved quantum particle swarm optimization algorithm is used to optimize the registration model parameters to obtain an optimal registration model.

[0178] The S5 step optimizes the registration model parameters using an improved quantum particle swarm optimization algorithm, including:

[0179] The most adjacent monitoring image of the kth monitoring video device at time t is traversed to obtain a monitoring image {Q k″,1 (t),Q k″,2 (t),...,Q k″,u (t),...,Q k″,U (t)} wherein Q k″,u (t) represents the monitoring image of the u-th monitoring video device corresponding to the most adjacent kth monitoring video device at time t, and the SIFT feature descriptor of the traversed monitoring image is extracted;

[0180] The registration parameters of different monitoring images are optimized using the improved quantum particle swarm optimization algorithm, and all the optimized registration parameters constitute an optimal registration model, and the optimization process of the quantum particle swarm optimization algorithm is as follows:

[0181] S51: initialize to generate a plurality of particles, each particle z constitutes a set of registration parameters of different monitoring images wherein z 0 represents the initial particle z, represents the registration parameter corresponding to the monitoring image Q k″,u (t), and each registration parameter generated by the initialization is the generated result of adding random disturbance to the initial registration parameter s0;

[0182] S52: let the current iteration number of the algorithm be r, the initial value of r be 1, and the maximum value be MAX;

[0183] S53: construct the fitness function of the quantum particle swarm, and the fitness function is:

[0184]

[0185] wherein:

[0186] SIFT k″,u (t) represents the SIFT feature descriptor of the monitoring image Q k″,u (t), and SIFT0 represents the SIFT feature descriptor in the monitoring target image Q0;

[0187] represents the registration parameter of the particle z obtained by the L-1th iteration, and f(z L-1 ) represents the fitness function value of the particle z obtained by the L-1th iteration;

[0188] C(·) represents the cosine similarity calculation function;

[0189] The particle with the maximum fitness function value of the L-1th iteration is taken as the optimal particle of the L-1th iteration

[0190] S54: Iterating any particle z to obtain the particle z of the Lth iteration L :

[0191] z L = z L-1 + β|z L-1,* -z L-1 | × rand (0, 1)

[0192]

[0193] Wherein:

[0194] z L-1 represents the particle of the L-1th iteration, and rand (0, 1) represents a random number between 0 and 1;

[0195] ω represents a learning rate, which is set to 0.8;

[0196] S55: Calculate the fitness function value of each particle after the Lth iteration, if L = MAX, the particles corresponding to the maximum fitness function are taken as the optimal registration parameters of the different monitoring images, otherwise L = L + 1, return to step S54.

[0197] S6: Selecting the monitoring image captured by the nearest monitoring video device of the monitoring video device monitoring the target, inputting the selected monitoring image into the optimal registration model for registration fusion to obtain the three-dimensional position of the target to be monitored at different times.

[0198] The S6 step of selecting the monitoring image captured by the nearest monitoring video device of the monitoring video device monitoring the target, inputting the selected monitoring image into the optimal registration model for registration fusion, comprises:

[0199] Selecting the monitoring image captured by the nearest monitoring video device corresponding to the monitoring video device monitoring the target at any time, inputting the selected monitoring image into the optimal registration model, using the registration parameters corresponding to the monitoring image to perform registration fusion on each monitoring image, and using the initial registration parameters to perform registration fusion on the monitoring image monitoring the target, taking the position of the target to be monitored after registration fusion as its three-dimensional position in the monitoring scene, and obtaining the three-dimensional feature information and three-dimensional position of the target to be monitored at any time in the monitoring scene.

[0200] Embodiment 2:

[0201] AsFigure 2 Fig. 1 is a functional module diagram of a semantic correlation multi-source monitoring video fusion analysis system according to an embodiment of the present application, which can realize the semantic correlation multi-source monitoring video fusion analysis method in Embodiment 1.

[0202] The semantic correlation multi-source monitoring video fusion analysis system 100 according to the present application can be installed in an electronic device. According to the functions to be realized, the semantic correlation multi-source monitoring video fusion analysis system can include a monitoring judgment device 101, an initial registration device 102 and a registration fusion module 103. The modules according to the present application can also be referred to as units, which refer to a series of computer program segments capable of being executed by an electronic device processor and capable of completing fixed functions, and are stored in the memory of the electronic device.

[0203] The monitoring judgment device 101 is configured to construct a multi-source mixed target global similarity measurement model, input the monitoring images captured by the monitoring video devices and the target images to be monitored into the model, take the image similarity measurement values of different monitoring images and the target images to be monitored as the probability features of the target monitoring of different monitoring images, and judge whether the target to be monitored is monitored according to the probability features of the monitoring images.

[0204] The initial registration device 102 is configured to perform initial matching on the target images to be monitored and the monitoring images of the monitored target by using a registration model.

[0205] The registration fusion module 103 is configured to optimize the registration model parameters by using an improved quantum particle swarm optimization algorithm, obtain an optimal registration model, select the monitoring images captured by the nearest monitoring video devices of the monitoring video devices of the monitored target, input the selected monitoring images into the optimal registration model for registration fusion, and obtain the three-dimensional positions of the target to be monitored at different time points.

[0206] In detail, the modules in the semantic correlation multi-source monitoring video fusion analysis system 100 according to the present application use the same technical means as the semantic correlation multi-source monitoring video fusion analysis method in Embodiment 1 and can produce the same technical effects, which will not be described herein again. Figure 1

[0207] Embodiment 3

[0208] As shown in Fig. 1, an electronic device for realizing the semantic correlation multi-source monitoring video fusion analysis method according to an embodiment of the present application is provided. Figure 3

[0209] The electronic device 1 can include a processor 10, a memory 11, a communication interface 13 and a bus, and can further include a computer program, such as a program 12, stored in the memory 11 and executable on the processor 10.​​

[0210] The memory 11 includes at least one type of readable storage medium, such as a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 11 can include both an internal storage unit and an external storage device of the electronic device 1. The memory 11 can be used not only to store application software and various data installed in the electronic device 1, such as the code of the program 12, etc., but also to temporarily store data that has been output or will be output.

[0211] The processor 10 can be composed of an integrated circuit in some embodiments, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more combinations of a central processing unit (CPU), a microprocessor, a digital processing chip, a graphics processor, and various control chips, etc. The processor 10 is the control unit of the electronic device, which connects various components of the entire electronic device through various interfaces and lines, executes programs or modules stored in the memory 11 (such as the program 12 for implementing semantic association-based multi-source monitoring video fusion analysis, etc.), and calls data stored in the memory 11, to perform various functions and process data of the electronic device 1.

[0212] The communication interface 13 can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the electronic device 1 and other electronic devices, and to realize connection and communication between internal components of the electronic device.

[0213] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection and communication between the memory 11, the at least one processor 10, etc.

[0214] Figure 3 Only the electronic device with components is shown, and those skilled in the art can understand that, Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and can include fewer or more components than shown, or combine certain components, or different component arrangements.

[0215] For example, although not shown, the electronic device 1 can also include a power supply (such as a battery) to power each component. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, so that the power management device can implement functions such as charge management, discharge management, and power consumption management. The power supply can also include one or more direct current or alternating current power supplies, recharging devices, power supply fault detection circuits, power supply converters or inverters, power supply status indicators, etc. The electronic device 1 can also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which are not described here.

[0216] Optionally, the electronic device 1 can also include a user interface, which can be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch, etc. The display can also be appropriately referred to as a display screen or a display unit, for displaying information processed in the electronic device 1 and for displaying a visualized user interface.

[0217] It should be understood that the embodiments are only for illustration and are not limited in the scope of the patent application by this structure.

[0218] The program 12 stored in the memory 11 in the electronic device 1 is a combination of multiple instructions, which, when executed in the processor 10, can implement:

[0219] A multi-source fusion video monitoring network system is constructed, and multiple monitoring video devices are networked and integrated, and a target image to be monitored is preset.

[0220] A target global similarity measure model of multi-source mixing is constructed, and the monitoring image captured by the monitoring video device and the target image to be monitored are input into the model;

[0221] The image similarity measure value of different monitoring images and the target image to be monitored is taken as the probability feature of the target monitored by the different monitoring images, and the probability feature of the monitoring image is associated to determine whether the target to be monitored is monitored;

[0222] If the target to be monitored is monitored, the target image to be monitored and the monitoring image of the target monitored are initially matched by using the registration model;

[0223] The registration model parameters are optimized by using the improved quantum particle swarm optimization algorithm to obtain the optimal registration model;

[0224] The monitoring image captured by the nearest monitoring video device of the monitoring video device of the target monitored is selected, and the selected monitoring image is input into the optimal registration model for registration fusion to obtain the three-dimensional position of the target to be monitored at different time.

[0225] Specifically, the specific implementation method of the processor 10 to the above instructions can refer to Figures 1 to 3 The description of the related steps in the corresponding embodiments will not be repeated here.

[0226] It should be noted that the above sequence numbers of the embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments. Moreover, the terms "include", "contain" or any other variants thereof in this paper are intended to cover non-exclusive inclusion, so that the process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, device, article or method. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, device, article or method including the element.

[0227] From the above description of the embodiments, those skilled in the art can clearly understand that the above embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the method described in each embodiment of the present application.

[0228] The above merely preferred embodiments of the present application and are not intended to limit the patent scope of the present application, any equivalent structure or equivalent process transformation made by using the content of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A semantic correlation multi-source monitoring video fusion analysis method, characterized in that, The method comprises: S1: constructing a multi-source fusion video monitoring network system, networking and integrating multiple monitoring video devices and presetting a target image to be monitored; S2: constructing a multi-source mixed target global similarity measurement model, and inputting the monitoring images captured by the monitoring video devices and the target image to be monitored into the model; S3: taking the image similarity measurement values of different monitoring images and the target image to be monitored as the probability characteristics of the target monitored by the different monitoring images, and associating the probability characteristics of the monitoring images to determine whether the target to be monitored is monitored; S4: if the target to be monitored is monitored, performing initial matching on the target image to be monitored and the monitoring image of the monitored target by using a registration model; S5: optimizing the parameters of the registration model by using an improved quantum particle swarm optimization algorithm to obtain an optimal registration model; S6: selecting a monitoring image captured by the nearest monitoring video device of the monitoring video device monitoring the target, inputting the selected monitoring image into the optimal registration model for registration fusion, and obtaining the three-dimensional position of the target to be monitored at different time points; In the S2 step, the multi-source mixed target global similarity measurement model is constructed, comprising: The multi-source mixed target global similarity measurement model is constructed, the model takes the target image to be monitored and different monitoring images as inputs, and takes image similarity measurement values as outputs, and the image similarity measurement process based on the target global similarity measurement model is: S21: the target global similarity measure model receives the preset target image to be monitored and the monitoring image , wherein the monitoring image represents a monitoring image captured by the kth monitoring video device in the multi-source fusion video monitoring network system at time t, K represents the total number of monitoring video devices in the multi-source fusion video monitoring network system, and the kth monitoring video device and the k+1th monitoring video device are adjacent monitoring video devices; S22: Fourier transform processing is used to obtain the frequency domain representations of the monitoring target image and the monitoring image respectively and the frequency domain phase information of the frequency domain representation results is extracted respectively and ​​ ; ; Wherein: amplitude information of the frequency domain representation amplitude information of the frequency domain representation phase information of the frequency domain representation phase information of the frequency domain representation amplitude information of the frequency domain representation amplitude information of the frequency domain representation phase information of the frequency domain representation phase information of the frequency domain representation Among them, the target image to be monitored The phase information of any i-th pixel is If all images have N pixels, then the target image to be monitored is... The frequency domain phase mean is Surveillance images The frequency domain phase mean is , Represents surveillance images Phase information of any i-th pixel; S23: Calculate the target image to be monitored and the monitoring image the frequency domain similarity : ; ; ; ; Wherein: N represents the total number of pixels of the image; respectively denote the phase information standard deviation of the target image and the monitoring image respectively denote the phase information covariance of the target image and the monitoring image S24: Transfer the images of the target to be monitored to... and surveillance images Convert to grayscale and And calculate the grayscale image of the target to be monitored. and monitoring grayscale images grayscale similarity : ; Wherein: the average pixel gray value of the gray image of the target to be monitored, the standard deviation of the gray value of the gray image of the target to be monitored, the standard deviation of the gray value of the gray image of the target to be monitored,​ the average pixel intensity value of the monitoring image the average pixel intensity value of the monitoring image the standard deviation of the intensity values of the monitoring image the standard deviation of the intensity values of the monitoring image S25: Frequency domain similarity Similarity with grayscale The mean value of the target image to be monitored is used as the mean value. and surveillance images Similarity metric The similarity metrics were then normalized to obtain normalized similarity metrics. The normalization formula is as follows: ; Wherein: represents the maximum similarity measure value of the monitoring image captured by the K monitoring video devices and the target image to be monitored at time t; represents the minimum similarity measure value of the monitoring images captured by the K monitoring video devices and the target image to be monitored at time t.

2. The semantic correlation multi-source monitoring video fusion analysis method of claim 1, wherein, In the S1 step, the multi-source fusion video monitoring network system is constructed, comprising: The multi-source fusion monitoring network system comprises a network layer and a physical layer, the target image to be monitored is preset in the network layer, the physical layer is composed of a plurality of monitoring video devices, the monitoring video devices can capture video images of the monitoring area as monitoring images in real time, the captured monitoring images are uploaded to the network layer through a wireless network forwarding device, the network layer is used for associating all the captured monitoring images, and whether the target to be monitored is monitored is determined based on the context probability characteristics of the associated monitoring images, if the target to be monitored is monitored, the corresponding monitoring image is input into the registration model for registration fusion, and the three-dimensional position of the target to be monitored at different time points is obtained.

3. The semantic associated multi-source monitoring video fusion analysis method of claim 1, wherein, In the S2 step, the monitoring images captured by the monitoring video devices and the target image to be monitored are input into the target global similarity measurement model, comprising: The different monitoring images captured by the different monitoring video devices at different time points and the target image to be monitored are input into the target global similarity measurement model to obtain a set of similarity measurement values of the captured monitoring images and the target image to be monitored: ; Wherein: represent a set of similarity measure values of the monitoring image and the target image to be monitored at time t; represents an initial time instant at which monitoring of the target image to be monitored starts, represents a cut-off time instant at which monitoring of the target image to be monitored ends.

4. The semantic associated multi-source monitoring video fusion analysis method of claim 3, wherein, In the S3 step, the probability characteristics of the monitoring images are associated to determine whether the target to be monitored is monitored, comprising: The image similarity measure value of different monitoring images and the image of the target to be monitored is taken as the probability feature of the target monitored by the different monitoring images, and the similarity measure value of any monitoring image and the image of the target to be monitored is calculated at time t The corresponding probability feature is ; The process of determining whether the target to be monitored is monitored by associating the probability characteristics of the monitoring images is: S31: traverse the probability feature set at time t ; S32: calculate the position distance between the first monitoring video device and the adjacent monitoring video devices, and select the adjacent monitoring video device with a position distance less than a position threshold as the nearest monitoring video device of the first monitoring video device; S32: calculate the position distance between the first monitoring video device and the adjacent monitoring video devices, and select the adjacent monitoring video device with a position distance less than a position threshold as the nearest monitoring video device of the first monitoring video device;​ S33: The first A set of surveillance video devices is constructed by considering each surveillance video device and its nearest neighbor. At time t, the product of the probability features of the images captured by any two surveillance video devices in the set is calculated, where the weight of the probability feature product is the distance between the two surveillance video devices. The sum of the weighted probability feature products is used as the monitoring conflict rate. ; S34: calculate the probability that the i-th monitoring video device captures the target to be monitored at time t :​ ; Wherein: Indicates the first The set of surveillance video devices corresponding to each surveillance video device, where c is the surveillance video device in the set. representing a probability feature of an image taken by the monitoring video device c at time t, representing a time interval; S35: If , it indicates that the th monitoring video device monitors the target to be monitored at time t, wherein represents a time threshold; otherwise, return to step S31 to select the probability feature with the second smallest probability and the corresponding monitoring video device, and if they are both smaller than the time threshold, it indicates that the target to be monitored is not monitored at time t.

5. The semantic associated multi-source monitoring video fusion analysis method of claim 1, wherein, In the S4 step, the initial matching of the target image to be monitored and the monitoring image of the monitored target is performed by using the registration model, comprising: If the first If a surveillance video device detects a target at time t, then the corresponding surveillance image is: Using a registration model to monitor the target image and surveillance images of the detected target Initial matching is performed. The registration model contains several sets of registration parameters. The monitoring images are registered and fused using the registration parameters to obtain the three-dimensional feature information of the target under the monitoring scene. The initial matching process is: S41: Extract the target image to be monitored respectively SIFT feature points of the monitoring image and construct a 64-dimensional SIFT feature descriptor for each feature point; S42: calculating the target image to be monitored SIFT feature descriptors in the monitoring image SIFT feature descriptors in the monitoring image SIFT feature descriptors in the monitoring image SIFT feature descriptors in the monitoring image S43: obtaining the target image to be monitored by using a least square method SIFT feature descriptors and the initial registration parameters of the matching descriptors , using the initial registration parameters to the monitoring image registration: , wherein SIFT feature descriptors of the monitoring image .

6. The semantic associated multi-source monitoring video fusion analysis method of claim 5, wherein, In the S5 step, the parameters of the registration model are optimized by using the improved quantum particle swarm optimization algorithm, comprising: The monitoring image captured by the nearest monitoring video device of the i-th monitoring video device at time t is denoted as The monitoring image captured by the nearest monitoring video device of the i-th monitoring video device at time t is denoted as wherein The monitoring image captured by the nearest monitoring video device of the i-th monitoring video device at time t is denoted as The monitoring image captured by the nearest monitoring video device of the i-th monitoring video device at time t is denoted as The improved quantum particle swarm optimization algorithm is used to optimize the registration parameters of different monitoring images, and all the optimized registration parameters are used to form an optimal registration model, and the optimization process of the quantum particle swarm optimization algorithm is as follows: S51: initialize generating a number of particles, each particle z constituting a set of registration parameters of different monitoring images wherein denotes an initial particle z, a monitoring image a corresponding registration parameter, each registration parameter generated in the initialization being an initial registration parameter adding a random perturbation to the generated result; S52: Let the current iteration number of the algorithm be r, the initial value of r be 1, and the maximum value be MAX; S53: Construct the fitness function of the quantum particle swarm, and the fitness function is: ; Wherein: representing a monitoring image SIFT feature descriptors of the monitoring image, representing SIFT feature descriptors in a target image to be monitored SIFT feature descriptors in the target image to be monitored represents the registration parameters of the particle z obtained in the L-1th iteration, represents the fitness function value of the particle z obtained in the L-1th iteration; denotes a cosine similarity computation function; the particle with the maximum fitness function value of the L-1th iteration is taken as the optimal particle of the L-1th iteration ; S54: iterate for any particle z, get the Lth iteration particle : ; ; Wherein: represents the particle of the L-1 iteration, represents a random number between 0-1; learning rate, set to 0.8; S55: Calculate the fitness function value of each particle after the Lth iteration, if L=MAX, then the registration parameters corresponding to the particle with the maximum fitness function value are used as the optimized registration parameters of different monitoring images, otherwise L=L+1, and return to step S54.

7. The semantic associated multi-source monitoring video fusion analysis method of claim 6, wherein, In the S6 step, the monitoring image captured by the nearest monitoring video device of the monitoring video device monitoring the target is selected, and the selected monitoring image is input into the optimal registration model for registration fusion, including: The monitoring image captured by the nearest monitoring video device of the monitoring video device monitoring the target is selected, and the selected monitoring image is input into the optimal registration model, the registration parameters corresponding to the monitoring image are used to perform registration fusion on each monitoring image, and the initial registration parameters are used to perform registration fusion on the monitoring image monitoring the target, the position of the monitoring target after registration fusion is used as the three-dimensional position of the monitoring target in the monitoring scene, and the three-dimensional feature information and three-dimensional position of the monitoring target at any time in the monitoring scene are obtained.

8. A semantic associated multi-source surveillance video fusion analysis system for implementing the semantic associated multi-source surveillance video fusion analysis method of any one of claims 1-7, characterized in that, The system comprises: A monitoring judgment device is used to construct a multi-source mixed target global similarity measurement model, and the monitoring image captured by the monitoring video device and the monitoring target image are input into the model, the image similarity measurement value of different monitoring images and the monitoring target image is used as the probability feature of the monitoring target in different monitoring images, and the probability feature of the monitoring image is used to determine whether the monitoring target is monitored; An initial registration device is used to use the registration model to perform initial matching on the monitoring target image and the monitoring image monitoring the target; A registration fusion module is used to optimize the registration model parameters by using the improved quantum particle swarm optimization algorithm, obtain the optimal registration model, select the monitoring image captured by the nearest monitoring video device of the monitoring video device monitoring the target, input the selected monitoring image into the optimal registration model for registration fusion, and obtain the three-dimensional position of the monitoring target at different times, so as to realize the semantic association multi-source monitoring video fusion analysis method as claimed in claims 1-7.

Citation Information

Patent Citations

  • Video multi-target fuzzy data correlation method and device

    CN107423686A

  • Video target multi-target tracking method and device and storage medium

    CN109859245A