Unmanned aerial vehicle formation flight state diagnosis system and method based on audio and video monitoring

CN116977938BActive Publication Date: 2026-09-11ZHEJIANG UNIV OF SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311009616.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2026-09-11
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

多无人机飞行过程中,每个无人机都有传感器系统来监测无人机集群是否在正常协同编队飞行,但是单一的主动感知监测精确度欠佳,且多无人机系统协同飞行时,可能会出现部分无人机损毁、通信失联等问题,主动感知的传感器无法正常工作,故需要被动感知监测,对失去联系的无人机进行定位追踪,使得操纵台及时调整,恢复正常工作

Benefits of technology

[0069] (1) Passive monitoring of the flight status of multiple UAVs in formation: This invention provides a diagnostic system and method for the flight status of UAVs in formation based on audio and video monitoring. By generating and collecting audio and video information of multiple UAVs, it can determine whether the flight status of multiple UAVs is normal, thereby improving the safety of UAV flight.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977938B_ABST
    Figure CN116977938B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle teaming flight state diagnostic system and method based on audio and video monitoring.The flight audio and video generation component is used to generate three-dimensional simulation video image and simulation audio information in flight process after receiving the flight plan information of multiple unmanned aerial vehicles;Flight audio and video monitoring component is used to monitor audio and video information in flight process in real time and send to environmental variable-oriented feature fusion removal component;Feature fusion removal component is used to generate real-time simulation audio and video information and real-time real audio and video information, realize environmental occlusion information fusion and environmental noise information removal;Flight state judging component is used to construct time series atlas and input into graph neural network for state judgment, obtain the flight state result of unmanned aerial vehicle.The application has the advantages of real-time, accuracy and visualization, applied to the passive monitoring of the flight of multiple unmanned aerial vehicles, can track the abnormal state of the flight of multiple unmanned aerial vehicles and report, improve the coordination and safety of teaming flight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for identifying abnormal flight states of multiple drones in the field of drone supervision, specifically involving a drone fleet flight state diagnosis system and method based on audio and video monitoring. Background Technology

[0002] The coordinated reconnaissance and combat operations of multiple drones flying in formation can, to some extent, increase the success rate of a single drone mission. In military reconnaissance, target strikes, communications relay, electronic warfare, battlefield assessment, and harassment / deception, drone formation flying can improve the efficiency of completing a single mission. During multi-drone flight, each drone has a sensor system to monitor whether the drone swarm is flying in normal coordinated formation. However, the accuracy of single active sensing and monitoring is insufficient, and when multiple drone systems are flying in coordination, problems such as damage to some drones or loss of communication may occur, rendering the active sensing sensors inoperable. Therefore, passive sensing and monitoring are needed to locate and track drones that have lost contact, allowing the control center to make timely adjustments and restore normal operation. Thus, a ground-based passive sensing monitoring system is needed to assist in monitoring whether the drone swarm is flying in normal coordinated formation. Summary of the Invention

[0003] In order to address the problems and needs existing in the background technology, the purpose of this invention is to propose a diagnostic system and method for the flight status of unmanned aerial vehicles (UAVs) fleets based on audio and video monitoring.

[0004] This invention utilizes multi-UAV flight plans to generate 3D simulated video images and simulated audio information during flight. An audio-visual acquisition module monitors the audio-visual information during multi-UAV flight in real time. An audio track separation method is used to separate environmental noise from the mixed audio information monitored in real time, obtaining real-time flight audio information. Real-time monitoring of field-of-view occlusion information during flight is added to the simulated 3D video images of multi-UAV flight, resulting in real-time simulated 3D video images. A real-time simulation time-series graph is constructed based on the real-time simulated 3D video images and simulated audio information. A real-time real-world time-series graph is also constructed based on the real-time monitored 3D video images and real-time flight audio information. Both the real-time simulation time-series graph and the real-time real-world time-series graph are input into a graph neural network for state judgment, obtaining the UAV flight state results and promptly sending commands to the control console for timely error correction.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] I. A UAV Flying Status Diagnostic System Based on Audio and Video Monitoring

[0007] include:

[0008] The flight audio and video generation component is used to receive flight plan information from multiple UAVs and perform simulation analysis, generate three-dimensional simulation video images and simulation audio information during the flight process, and send them to the feature fusion and removal component;

[0009] The flight audio and video monitoring component is used to monitor the audio and video information of multiple UAVs in flight in real time using the audio and video acquisition module and send it to the feature fusion and removal component;

[0010] The feature fusion and removal component is used to generate real-time simulated audio and video information and real-time real audio and video information based on the three-dimensional simulated video images and simulated audio information during flight and the real-time monitored audio and video information, and send the real-time simulated audio and video information and real-time real audio and video information to the flight status judgment component.

[0011] The flight status determination component is used to construct a time-series graph based on real-time simulated audio and video information and real-time real audio and video information, input the time-series graph into a graph neural network for status determination, and obtain the flight status result of the UAV.

[0012] The audio and video acquisition module consists of a microphone array and a binocular camera sensor. Both the microphone array and the binocular camera sensor are fixedly mounted on the platform. The microphone array is used to acquire audio information from multiple drones, and the binocular camera sensor is used to acquire three-dimensional video image information from multiple drones.

[0013] The observation parameters of the microphone array and the binocular camera satisfy the following formula:

[0014] r′ o =Preset fixed value

[0015] θ′ o _min=0

[0016]

[0017]

[0018]

[0019] Where, r o θ represents the distance between the microphone array, the binocular camera, and the drone's takeoff point. o ′_min represents the minimum polar angle, θ o ′_max represents the maximum polar angle. Indicates the minimum azimuth angle. This indicates the maximum azimuth angle.

[0020] II. A method for diagnosing the flight status of drone squadrons based on audio and video monitoring

[0021] Step 1: Generate 3D simulation video images and simulation audio information of multiple UAVs during flight based on the flight plan information of multiple UAVs;

[0022] Step 2: Use a microphone array to monitor the mixed audio information of multiple drones in flight in real time, and use a binocular camera to monitor the three-dimensional video image information of multiple drones in flight in real time;

[0023] Step 3: Use the audio track separation method to separate the environmental noise in the real-time monitored mixed audio information to obtain real-time flight audio information;

[0024] Step 4: Extract the field-of-view occlusion information from the real-time monitored 3D video image information and add it to the 3D simulation video image to obtain the real-time simulation 3D video image;

[0025] Step 5: Construct a real-time simulation time series graph based on real-time simulated 3D video images and simulated audio information; construct a real-time real time series graph based on real-time monitored 3D video image information and real-time flight audio information.

[0026] Step 6: Input the real-time simulation time series graph and the real-time actual time series graph into the graph neural network model to make state judgments and obtain the flight state results of the UAV.

[0027] The observation parameters of the microphone array and the binocular camera satisfy the following formula:

[0028] r′ o =Preset fixed value

[0029] θ′ o _min=0

[0030] θ′ o _max=180°

[0031]

[0032]

[0033] Where, r o θ represents the distance between the microphone array, the binocular camera, and the drone's takeoff point. o ′_min represents the minimum polar angle, θ o ′_max represents the maximum polar angle. Indicates the minimum azimuth angle. This indicates the maximum azimuth angle.

[0034] Step 3 specifically involves:

[0035] First, the real-time monitored mixed audio information Represented as a linear combination, the formula is as follows:

[0036]

[0037] Where A is the mixing matrix, s(t1, t2, ..., t i () is a mixed source signal;

[0038] Next, the mixing matrix A is decomposed to obtain the separating matrix W;

[0039] Then, the separation matrix W is used to analyze the mixed audio signals. Reconstruction is performed to obtain the independent source signal s′ real (t), satisfying

[0040]

[0041] Finally, the ICA signal separation algorithm is used to separate the mixed audio signals using the following formula. The signals are separated to obtain the environmental audio signal e(t) and the multi-UAV sound signals s. real (t):

[0042]

[0043]

[0044] in, The environmental noise signal at time t. Let a be the audio signal of the nth drone at time t. n This is the audio information of the nth drone.

[0045] In step 4, at time t, the real-time monitored video image information z collected by the binocular camera real (t) includes the left camera image l real (t) and the right camera image r real First, the disparity image d at time t is calculated using the following formula. real (t):

[0046] d real (x, y) t =(x left -x right ) t

[0047] Where, d real (x, y) t The image d at time t represents the disparity image. real The pixel coordinates (x, y) in (t) t The disparity value, x left and xright These are images from the left camera. real (t) and the right camera image r real The horizontal pixel coordinates of the feature points in (t);

[0048] Next, based on the disparity image d at time t... real (t), and then the depth image z is calculated using the following formula. real (t):

[0049]

[0050] Where B is the baseline length of the binocular camera, f is the focal length, and z is the focal length. real (x, y) t Represents the depth image z at time t real The pixel coordinates (x, y) in (t) t The depth value;

[0051] Then, the depth image z at time t real In F(t), depth values ​​less than the threshold T are marked as foreground, otherwise as background, thus obtaining the foreground image F(t) at time t;

[0052] Then, based on the foreground image F(t) at time t, the depth values ​​of the left camera and the right camera at time t, the field-of-view occlusion information W(t) at time t is determined using the following formula:

[0053]

[0054] Where W(x, y) t Let W(t) represent the pixel coordinates (x, y) in the field-of-view occlusion information at time t. t The occlusion value, z left (x, y) t Z represents the depth value of the left camera at time t. right (x, y) t Let F(x, y) represent the depth value of the right camera at time t. t Let F(t) be the pixel coordinates (x, y) in the foreground image F(t) at time t. t The pixel value.

[0055] Finally, the field-of-view occlusion information W(t) at time t is compared with the simulated 3D video images of the multiple UAVs at time t. After combining, the multi-UAV real-time simulation 3D video image Z at time t, with added field-of-view occlusion information, is obtained. sim (t), the formula is as follows:

[0056]

[0057] Step 5 specifically involves:

[0058] Using multiple UAVs as entity nodes and the flight time series of multiple UAVs as the main line of the time series graph, the time series graph is divided into t1~t2, t2~t3, ..., t n-1 ~t n The time series is the edge of the time series graph. The real-time simulation time series graph is obtained by using real-time simulated 3D video images and simulated audio information as attributes of the time series graph. The real-time real time series graph is obtained by using real-time monitored 3D video image information and real-time flight audio information as attributes of the time series graph.

[0059] Step 6 specifically involves:

[0060] 6.1) Using real-time simulation time series graph G sim (t n ) and real-time time series graph g real (t n Using the graph data as the object, feature extraction is performed on each entity node to form a real-time simulation time-series feature matrix G. A1sim (t n ) and real-time time series feature matrix g A2real (t n Based on the two temporal feature matrices, the corresponding adjacency matrix is ​​generated and used as the input to the graph convolutional neural network model;

[0061] 6.2) Use a graph convolutional neural network (GCN) to transform the input adjacency matrix into a real-time simulation time-series graph feature matrix. and real-time time series graph feature matrix The formula is as follows:

[0062]

[0063]

[0064] Among them, G G1sim (t n ) and g G2real (t n ) are the nth real-time simulation adjacency matrix and the real-time actual adjacency matrix, respectively, and T represents the matrix transpose;

[0065] 6.3) After classifying the flight state of the output of the Graph Convolutional Neural Network (GCN) using a Support Vector Machine (SVM), the flight state results of multiple UAVs are obtained, as shown in the following formula:

[0066]

[0067] Where y represents the flight status result of multiple UAVs, f SVM() indicates that the binary classification method using support vector machines is used for classification.

[0068] The beneficial effects of this invention are:

[0069] (1) Passive monitoring of the flight status of multiple UAVs in formation: This invention provides a diagnostic system and method for the flight status of UAVs in formation based on audio and video monitoring. By generating and collecting audio and video information of multiple UAVs, it can determine whether the flight status of multiple UAVs is normal, thereby improving the safety of UAV flight.

[0070] (2) Efficient and accurate monitoring of UAV flight status: The system uses multiple sensors for monitoring, including binocular cameras and microphone arrays, which can efficiently and accurately acquire audio and video information of UAV flight, and perform feature extraction and analysis to effectively monitor the flight status of multiple UAVs.

[0071] (3) Automated and intelligent monitoring of UAV flight status: The system adopts the method of constructing time series graphs and analysis and comparison based on graph neural network models, which can automatically and intelligently monitor the flight status of UAVs, reduce the burden on operators, and improve monitoring efficiency and accuracy.

[0072] (4) Improved efficiency of flight data acquisition and analysis: The system can acquire audio and video information of multiple UAVs in real time. By controlling environmental factor variables, adding real-time field-of-view obstruction information and removing environmental noise during flight, the efficiency of flight data acquisition and analysis is improved, which helps to better understand the flight status of UAVs and guide flight control and decision-making. Attached Figure Description

[0073] Figure 1 This is a system structure diagram of the system of the present invention.

[0074] Figure 2 This is a flowchart of the implementation steps of the present invention.

[0075] Figure 3 Examples of multi-UAV flight monitored in real time by this invention.

[0076] Figure 4 This is the time series map constructed by the present invention. Detailed Implementation

[0077] The technical solutions of the present invention will be further described below with reference to the accompanying drawings of the embodiments of the present invention, but the scope of protection of the present invention is not limited thereto.

[0078] like Figure 1 A drone platoon flight status diagnostic system based on audio and video monitoring includes:

[0079] The flight audio and video generation component is used to receive flight plan information from multiple UAVs and perform simulation analysis, generate three-dimensional simulation video images and simulation audio information during the flight process, and send them to the feature fusion and removal component oriented towards environmental variables.

[0080] The flight audio and video monitoring component is used to monitor the audio and video information of multiple UAVs in real time during flight using the audio and video acquisition module and send it to the feature fusion and removal component oriented towards environmental variables;

[0081] The feature fusion and removal component generates real-time simulated audio-visual information and real-time realistic audio-visual information based on 3D simulated video images and simulated audio information during flight, as well as real-time monitored audio-visual information. This real-time simulated audio-visual information and real-time realistic audio-visual information are then sent to the multi-UAV flight status judgment component. The real-time simulated audio-visual information consists of simulated audio information and 3D simulated video images with added real-time occlusion information (i.e., real-time simulated 3D video images). The real-time realistic audio-visual information consists of real-time monitored audio-visual information (i.e., real-time flight audio information) after removing real-time environmental noise and real-time monitored 3D video image information.

[0082] The flight status judgment component is used to construct a time series graph based on real-time simulated audio and video information and real-time real audio and video information. The time series graph is then input into a graph neural network for status judgment to obtain the flight status results of the UAV. Based on the flight status results of the UAV, if an abnormal situation occurs at a certain moment during the flight of one of the UAVs, it will promptly report to the control console and make adjustments.

[0083] The audio and video acquisition module consists of a microphone array and a binocular camera sensor. Both the microphone array and the binocular camera sensor are fixedly mounted on the platform. The microphone array is used to acquire audio information from multiple drones, and the binocular camera sensor is used to acquire three-dimensional video image information from multiple drones.

[0084] like Figure 2 As shown, a method for diagnosing the flight status of drone fleets based on audio and video monitoring includes the following steps:

[0085] Step 1: Generate 3D simulation video images and simulation audio information of multiple UAVs during flight based on the flight plan information of multiple UAVs;

[0086] In step 1, the multi-drone flight plan information includes the drone model, quantity (Num), and route information (Route), as shown in the following formula:

[0087] Route = {D i,t A i,t H i,t P i,t V i,t}

[0088] Among them, D i,t Let A be the flight direction of the i-th UAV at time t. i,t Let H be the flight attitude of the i-th UAV at time t. i,t Let P be the flight altitude of the i-th UAV at time t. i,t Let V be the flight position of the i-th UAV at time t. i,t Let be the flight speed of the i-th UAV at time t.

[0089] The formula for generating 3D simulation video images and simulation audio information of multiple UAVs during flight is as follows:

[0090]

[0091] in, S provides 3D simulation video images of multiple drones in flight. sim (t) represents the simulated audio information of multiple drone flights.

[0092] Suppose there are n drones, n = (1, 2, ..., i), and they are all the same model. Let the model of each drone be X, and let the position P of each drone in the global coordinate system be... i,t Represented using a three-dimensional vector, the formula is as follows:

[0093] P i,t =(X i,t Y i,t Z i,t )

[0094] Among them, X i,t Y i,t Z i,t These are the X, Y, and Z coordinates of the UAV in the global coordinate system;

[0095] Its attitude angle A i,t This can be expressed by the following formula:

[0096] A i,t =(φ i,t θ i,t , ψ i,t )

[0097] Where, φ i,t θ i,t , ψ i,t These are the roll angle, pitch angle, and yaw angle of the drone, respectively.

[0098] Suppose that the position of a ground observer in the global coordinate system is represented by the following formula:

[0099] P0 = (X0, Y0, Z0)

[0100] Where X0, Y0, and Z0 are the X, Y, and Z coordinates of the ground observer in the global coordinate system, respectively;

[0101] The attitude angle is expressed by the following formula:

[0102] A0 = (φ0, θ0, ψ0)

[0103] Where φ0, θ0, and ψ0 are the roll angle, pitch angle, and yaw angle observed by the ground observer, respectively.

[0104] If the distance from the starting point of the multiple UAVs is r0, then the position and attitude of the multiple UAVs from the observer's perspective can be represented by (P). i,0 ) t The formula is as follows:

[0105] (P i,0 ) t =R(A0)*(P i,t -P0)

[0106] Where R(A0) is the rotation matrix in the ground observer coordinate system, representing the rotation of the vector in the global coordinate system to the vector in the ground observer coordinate system, P0 is the position of the ground observer in the global coordinate system, and R(A0) is calculated from the attitude angle of the ground observer according to the following formula:

[0107] R(A0)=R Z (ψ0)*R Y (θ0)*R X (φ0)*R X (-φ i )*R Y (-θ i )*R Z (-ψ i )

[0108] Among them, R X (φ0) represents the rotation matrix for rotating about the x-axis by an angle of φ0, R z (ψ0) represents the rotation matrix for rotating about the z-axis by an angle ψ0, R Y (θ0) represents the rotation matrix θ0 around the y-axis, R X (-φ i R represents the rotation matrix for rotating about the x-axis by an angle of -φ0. Y (-θ i R represents the rotation matrix for rotating about the z-axis by an angle -ψ0. Z (-ψ i ) represents the rotation matrix that rotates around the y-axis by an angle of -θ0.

[0109] From the above, we can obtain the position and attitude (P) of multiple UAVs from the perspective of a ground observer. i,0 ) t Then, simulated 3D video images of multiple drones in flight can be generated and represented as follows:

[0110]

[0111] Where X, Y, and Z represent the three-dimensional spatial coordinates of the image, and this simulates a three-dimensional video image. It provides 3D images and videos in all directions.

[0112] The flight plan information for multiple drones can be decomposed into the flight plan information for each individual drone. The flight plan information for multiple drones can then be represented by the following formula:

[0113] B={B1, B2, B3,…,B i}

[0114] Among them, B i This represents the flight plan information for the i-th single UAV. A single UAV flight audio library is introduced, expressed by the following formula:

[0115] s={s1, s2, s3,..., s n}

[0116] Where n represents the nth audio signal, which corresponds one-to-one with the flight plan information of a single drone. For audio information from multiple drones, the audio signal s of each single drone can be used as the basis for the data. n Use the following formula to mix the audio and generate a mixed signal S from multiple signals. sim (t).

[0117]

[0118] Step 2: Use a microphone array to monitor the mixed audio information of multiple drones in flight in real time, and use a binocular camera to monitor the three-dimensional video image information of multiple drones in flight in real time, thereby obtaining real-time monitored audio and video information;

[0119] In step 2, such as Figure 3 As shown, taking the flight of 5 drones as an example, first determine the overall monitoring position of the microphone array and binocular camera. Let the microphone array and binocular camera be represented by MK&CAM. The position of MK&CAM is expressed by the following formula:

[0120]

[0121] Where, r o ′ represents the distance from MK&CAM to the takeoff point, θ o′ represents the angle between MK&CAM and the vertical axis. This indicates the rotation angle of MK&CAM on the horizontal plane.

[0122] To facilitate the description of the observation angle, the solid angle is decomposed into two parts: the horizontal angular range (azimuth). ) and the range of angles in the vertical direction (polar angle θ) o ′). Horizontal angular range (azimuth). ) can be described as in Indicates the minimum azimuth angle. Indicates the maximum azimuth angle. The range of angles in the vertical direction (polar angle θ) o ′) can be described as θ o ′_min,θ o [′_max], where θ o ′_min represents the minimum polar angle, θ o ′_max represents the maximum polar angle.

[0123] Therefore, the range of observation angles can be expressed by the following formula:

[0124] θ o ′_min≤θ o ′≤θ o ′_max

[0125]

[0126] When the distance r between the MK&CAM and the takeoff point is a fixed distance. o ′, and hope to observe the azimuth angle when observing the omnidirectional angular range of multiple drones. and polar angle θ o The following conditions must be met:

[0127] r′ o =Preset fixed value

[0128] θ′ o _min=0

[0129]

[0130]

[0131]

[0132] This allows for observation of the entire angular range, including 180 degrees vertically and 360 degrees horizontally.

[0133] Once the location of MK&CAM is determined, that is, based on the multi-UAV flight plan, the distance r between MK&CAM and the takeoff point is fixed. o After that, θ can be determined at a fixed time during the flight of multiple drones. o 'and The angle, that is, the observation angle of MK&CAM at time t1 for multiple UAVs, is At time t2, the observation angle of MK&CAM for multiple drones is... ...multiple drones in t n At that moment, the observation angle of MK&CAM was

[0134] Real-time mixed audio information from multiple drones was collected using a microphone array. The real-time 3D video image information of multiple drones acquired using a binocular camera is expressed by the following formula:

[0135] z real (t) = (x, y, z, t)

[0136] Where (x, y, z) represent the three-dimensional coordinates of the image. This is expressed by the following formula:

[0137]

[0138] Step 3: Use the audio track separation method to separate the environmental noise in the real-time monitored mixed audio information to obtain real-time flight audio information;

[0139] Step 3 specifically involves:

[0140] First, the real-time monitored mixed audio information Represented as a linear combination, the formula is as follows:

[0141]

[0142] Where A is the mixing matrix, s(t1, t2, ..., t i () is a mixed source signal;

[0143] Next, the mixture matrix A is decomposed by taking the inverse of the mixture matrix A to obtain the separation matrix W;

[0144] Then, the separation matrix W is used to analyze the mixed audio signals. Reconstruction is performed to obtain the independent source signal s′ real (t), satisfying

[0145]

[0146] Finally, the ICA signal separation algorithm is used to separate the mixed audio signals using the following formula. The signals are separated to obtain the environmental audio signal e(t) and the multi-UAV sound signals s. real (t):

[0147]

[0148]

[0149] in, The environmental noise signal at time t. Let a be the audio signal of the nth drone at time t. n This is the audio information of the nth drone.

[0150] Step 4: Extract the field-of-view occlusion information from the real-time monitored 3D video image information and add it to the 3D simulation video image to obtain the real-time simulation 3D video image;

[0151] At time t, the real-time monitored video image information z collected by the binocular cameras real (t) includes the left camera image l real (t) and the right camera image r real First, the disparity image d at time t is calculated using the following formula. real (t):

[0152] d real (x, y) t =(x left -X right ) t

[0153] Where, d real (x, y) t The image d at time t represents the disparity image. real The pixel coordinates (x, y) in (t) t The disparity value, x left and x right These are images from the left camera. real (t) and the right camera image r real The horizontal pixel coordinates of the feature points in (t);

[0154] Next, based on the disparity image d at time t... real (t), and then the depth image z is calculated using the following formula. real (t):

[0155]

[0156] Where B is the baseline length of the binocular camera, f is the focal length, and z is the focal length.real (x, y) t Represents the depth image z at time t real The pixel coordinates (x, y) in (t) t The depth value;

[0157] Then, the depth image z at time t real In F(t), depth values ​​less than the threshold T are marked as foreground, i.e., F(x, y). t =1 indicates foreground; otherwise, it represents background, i.e., F(x, y). t =0, thus obtaining the foreground image F(t) at time t; the formula is as follows:

[0158]

[0159] Then, based on the foreground image F(t) at time t, the depth values ​​of the left camera and the right camera at time t, the field-of-view occlusion information W(t) at time t is determined using the following formula:

[0160]

[0161] Where W(x, y) t Let W(t) represent the pixel coordinates (x, y) in the field-of-view occlusion information at time t. t The occlusion value, z left (x, y) t Z represents the depth value of the left camera at time t. right (x, y) t Let F(x, y) represent the depth value of the right camera at time t. t Let F(t) be the pixel coordinates (x, y) in the foreground image F(t) at time t. t The pixel value.

[0162] Specifically:

[0163] The depth value of the left camera at time t is z. left (x, y) t The depth value of the right camera at time t is z. right (x, y) t A binary mask is used to represent whether occlusion exists at a certain pixel. If, at time t, the depth value of the left camera at pixel position (x, y) is less than the depth value of the right camera at time t, and the corresponding pixel in the foreground image is also in the foreground, then occlusion exists, i.e., W(x, y). t =1; otherwise, there is no occlusion, i.e., W(x, y). t =0.

[0164] Finally, the field-of-view occlusion information W(t) at time t is compared with the simulated 3D video images of the multiple UAVs at time t. After combining, specifically, the pixels of the occluded parts are set to 0, while the unoccluded parts are retained, to obtain the multi-UAV real-time simulation 3D video image Z at time t after adding the field-of-view occlusion information. sim (t), the formula is as follows:

[0165]

[0166] Step 5: Construct a real-time simulation time series graph based on real-time simulated 3D video images and simulated audio information. Construct a real-time real-time time series graph based on real-time monitored 3D video image information and real-time flight audio information, such as... Figure 4 (a) and Figure 4 As shown in (b);

[0167] Step 5 specifically involves:

[0168] (1) Determining entities: This invention relates to a time series map, the entities of which are multiple unmanned aerial vehicles;

[0169] (2) Constructing the main thread: Using the flight time series of multiple UAVs as the main thread;

[0170] (3) Add edges: based on t1~t2, t2~t3, ..., t n-1 ~t n The time series is used as the edges of the time series graph to express the correlation between the flight states of multiple UAVs in different time periods;

[0171] (4) Add attributes: Use real-time simulation 3D video images and simulation audio information as attributes of each entity node in the real-time simulation time series graph. At the same time, use real-time monitored 3D video images and real-time flight audio information as attributes of each entity node in the real-time real time series graph.

[0172] (5) Constructing the time series graph: Integrate entity nodes, time series main lines, edges, and attributes to construct a real-time simulation time series graph G. sim (t n ) and real-time time series graph g real (t n );

[0173] Step 6: Input the real-time simulation time series graph and the real-time actual time series graph into the graph neural network model to make state judgments and obtain the flight state results of the UAV.

[0174] 6.1) Using real-time simulation time series graph G sim (t n ) and real-time time series graph g real (tn Using the graph data as the object, feature extraction is performed on each entity node to form a real-time simulation time-series feature matrix G. A1sim (t n ) and real-time time series feature matrix g A2real (t n Based on the two temporal feature matrices, the corresponding adjacency matrix is ​​generated and used as the input to the graph convolutional neural network model;

[0175] 6.2) Use a graph convolutional neural network (GCN) to transform the input adjacency matrix into a real-time simulation time-series graph feature matrix. and real-time time series graph feature matrix The formula is as follows:

[0176]

[0177]

[0178] Among them, G G1sim (t n ) and g G2real (t n ) are the nth real-time simulation adjacency matrix and the real-time actual adjacency matrix, respectively, and T represents the matrix transpose;

[0179] In specific implementation, the graph convolutional neural network function f is used. GCN ()accomplish:

[0180]

[0181]

[0182] 6.3) Support Vector Machines are used to classify the processed data to monitor anomalies. Specifically, the flight status of the output of the Graph Convolutional Neural Network (GCN) is classified to obtain the flight status results of multiple UAVs. The flight status results are categorized as normal or abnormal. In practice, normal is represented by C and abnormal by A. The formula is as follows:

[0183]

[0184] Where y represents the flight status result of multiple UAVs, f SVM () indicates that the binary classification method using support vector machines is used for classification.

[0185] Finally, it should be noted that the above embodiments and descriptions are only used to illustrate the technical solutions of the present invention and not to limit it. Those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the protection scope of the claims of the present invention.

Claims

1. A drone platoon flight status diagnostic system based on audio and video monitoring, characterized in that, include: The flight audio and video generation component is used to receive flight plan information from multiple UAVs and perform simulation analysis, generate three-dimensional simulation video images and simulation audio information during the flight process, and send them to the feature fusion and removal component; The flight audio and video monitoring component is used to monitor the audio and video information of multiple UAVs in real time during flight using the audio and video acquisition module and send it to the feature fusion and removal component; the audio and video acquisition module consists of a microphone array and a binocular camera sensor, both of which are fixedly mounted on the platform. The microphone array is used to acquire the audio information of multiple UAVs, and the binocular camera sensor is used to acquire the three-dimensional video image information of multiple UAVs. The feature fusion and removal component is used to generate real-time simulated audio and video information and real-time real audio and video information based on the three-dimensional simulated video images and simulated audio information during flight and the real-time monitored audio and video information, and send the real-time simulated audio and video information and real-time real audio and video information to the flight status judgment component. The flight status determination component is used to construct a time series graph based on real-time simulated audio and video information and real-time real audio and video information, input the time series graph into a graph neural network for status determination, and obtain the flight status result of the UAV. Specifically, it includes: Using multiple drones as entity nodes and the flight time series of multiple drones as the main line of the time series graph, , ... The time series is the edge of the time series graph. The real-time simulation time series graph is obtained by using real-time simulated 3D video images and simulated audio information as attributes of the time series graph. The real-time real time series graph is obtained by using real-time monitored 3D video image information and real-time flight audio information as attributes of the time series graph. Real-time simulation time series graph and real-time time series graph Using the graph data as the object, feature extraction is performed on each entity node to form a real-time simulation time-series feature matrix. and real-time time series feature matrix Based on the two temporal feature matrices, the corresponding adjacency matrix is ​​generated and used as the input of the graph convolutional neural network model; The graph convolutional neural network (GCN) is used to transform the input adjacency matrix into a real-time simulation time-series graph feature matrix. and real-time time series graph feature matrix The formula is as follows: in, and These are the nth real-time simulation adjacency matrix and the real-time actual adjacency matrix, respectively, and T represents the matrix transpose; After classifying the flight state of the graph convolutional neural network (GCN) using a support vector machine, the flight state results of multiple UAVs are obtained, as shown in the following formula: Where y represents the flight status result of multiple drones, This indicates that the binary classification method using support vector machines is employed.

2. The UAV platoon flight status diagnosis system based on audio and video monitoring according to claim 1, characterized in that, The observation parameters of the microphone array and the binocular camera satisfy the following formula: in, This indicates the distance between the microphone array, the binocular camera, and the drone's takeoff point. _min represents the minimum polar angle. _max represents the maximum polar angle. _min represents the minimum azimuth angle. _max represents the maximum azimuth angle.

3. A method for diagnosing the flight status of unmanned aerial vehicles (UAVs) fleets based on audio and video monitoring, characterized in that, Includes the following steps: Step 1: Generate 3D simulation video images and simulation audio information of multiple UAVs during flight based on the flight plan information of multiple UAVs; Step 2: Use a microphone array to monitor the mixed audio information of multiple drones in flight in real time, and use a binocular camera to monitor the three-dimensional video image information of multiple drones in flight in real time; Step 3: Use the audio track separation method to separate the environmental noise in the real-time monitored mixed audio information to obtain real-time flight audio information; Step 4: Extract the field-of-view occlusion information from the real-time monitored 3D video image information and add it to the 3D simulation video image to obtain the real-time simulation 3D video image; Step 5: Construct a real-time simulation time series graph based on real-time simulated 3D video images and simulated audio information; construct a real-time real time series graph based on real-time monitored 3D video image information and real-time flight audio information. Step 5 specifically involves: Using multiple drones as entity nodes and the flight time series of multiple drones as the main line of the time series graph, , ... The time series is the edge of the time series graph. The real-time simulation time series graph is obtained by using real-time simulated 3D video images and simulated audio information as attributes of the time series graph. The real-time real time series graph is obtained by using real-time monitored 3D video image information and real-time flight audio information as attributes of the time series graph. Step 6: Input the real-time simulation time series graph and the real-time actual time series graph into the graph neural network model for state determination to obtain the flight state results of the UAV; Step 6 specifically involves: 6.1) Using real-time simulation time series plots and real-time time series graph Using the graph data as the object, feature extraction is performed on each entity node to form a real-time simulation time-series feature matrix. and real-time time series feature matrix Based on the two temporal feature matrices, the corresponding adjacency matrix is ​​generated and used as the input of the graph convolutional neural network model; 6.2) Use a graph convolutional neural network (GCN) to transform the input adjacency matrix into a real-time simulation time-series graph feature matrix. and real-time time series graph feature matrix The formula is as follows: in, and These are the nth real-time simulation adjacency matrix and the real-time actual adjacency matrix, respectively, and T represents the matrix transpose; 6.3) After classifying the flight state of the output of the Graph Convolutional Neural Network (GCN) using a Support Vector Machine (SVM), the flight state results of multiple UAVs are obtained, as shown in the following formula: Where y represents the flight status result of multiple drones, This indicates that the binary classification method using support vector machines is employed.

4. The method for diagnosing the flight status of unmanned aerial vehicles (UAVs) fleets based on audio and video monitoring according to claim 3, characterized in that, The observation parameters of the microphone array and the binocular camera satisfy the following formula: in, This indicates the distance between the microphone array, the binocular camera, and the drone's takeoff point. _min represents the minimum polar angle. _max represents the maximum polar angle. _min represents the minimum azimuth angle. _max represents the maximum azimuth angle.

5. The method for diagnosing the flight status of unmanned aerial vehicles (UAVs) fleets based on audio and video monitoring according to claim 3, characterized in that, Step 3 specifically involves: First, the real-time monitored mixed audio information Represented as a linear combination, the formula is as follows: in, It is a mixed matrix. It is a mixed source signal; Next, for the mixing matrix Decompose to obtain the separation matrix ; Then, use the separation matrix. For mixed audio signals Reconstruction is performed to obtain independent source signals. ,satisfy : Finally, the ICA signal separation algorithm is used to separate the mixed audio signals using the following formula. Separate the signals to obtain the ambient audio signal. and multiple drone sound signals : in, For time Environmental noise signals, For time No. The audio signal of a drone. For the first Audio information from a drone.

6. The method for diagnosing the flight status of unmanned aerial vehicles (UAVs) fleets based on audio and video monitoring according to claim 3, characterized in that, In step 4, At any given moment, the binocular cameras capture real-time monitoring video image information. Including images from the left camera and right camera image First, calculate using the following formula. Parallax image at time : in, express Time parallax image pixel coordinates The disparity value, and These are images from the left camera. and right camera image Horizontal pixel coordinates of the feature points in the middle; Next, based on Parallax image at time The depth image is then calculated using the following formula. : in, The baseline length of the binocular camera. Focal length express Time-deep image pixel coordinates The depth value; Then, Depth image at time Medium less than the threshold The depth value is marked as foreground, otherwise as background, thus obtaining... Foreground image at time ; Then according to Foreground image at time Left camera Depth values ​​at any given time and the right camera at [time]. The depth value at time step is obtained using the following formula. Information obstructed from view at any time The formula is as follows: in, express Information obstructed from view at any time Mid-pixel coordinates The occlusion value, Indicates the left camera is in Depth value at time, Indicates the right camera is in Depth value at time, for Foreground image at any moment pixel coordinates Pixel values; Finally, Information obstructed from view at any time and Simulated 3D video images of multiple drones in real time After combining, obtain Real-time simulation of 3D video images of multiple drones after adding field-of-view occlusion information. The formula is as follows: 。