Unmanned aerial vehicle front end violation behavior identification method, system, device and medium
By combining multimodal fusion technology with visual images and inertial measurement unit data, the feature extraction and timing modeling of generative adversarial networks and hybrid convolutional neural long short-term memory networks are used to perform feature extraction and timing modeling, the problem of limitations of single modal information is solved, and high-precision and real-time identification of violations is achieved.
Patent Information
- Application Number
- CN202510273507.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
Single-modal image information has limitations in the identification of violations at the front end of the drone. The visual image is affected by noise and low contrast, and the inertial measurement unit data lacks spatial details, resulting in a reduced recognition accuracy.
By combining multimodal fusion technology with visual images and inertial measurement unit data, the image is pre-processed with the generative adversarial network defog module, visual features and inertial features are extracted, and a unified multimodal feature representation is generated through the fusion module. The mixed convolutional neural long short-term memory network is used for timing characteristics modeling to realize the classification and positioning of violations.
It improves the accuracy and real-timeness of the identification of violations, and the generated results have high accuracy, strong real-timeness and scene adaptability, which improves the accuracy and efficiency of violation detection.
Smart Images

Figure CN120220088A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image enhancement, and particularly relates to a method, system, device and medium for identifying illegal behaviors at the front end of an unmanned aerial vehicle (UAV). Background Art
[0002] Identifying illegal behaviors at the front end of a UAV is an important task in the fields of intelligent transportation and public safety, and is widely used in scenarios such as traffic management and environmental monitoring. In the identification of illegal behaviors at the front end of a UAV, conventional image acquisition includes visual images and inertial measurement unit (IMU) data. These two types of modal information can provide different perception capabilities. Visual images can clearly show the object features in the scene, especially for the identification of traffic violations; while IMU data can provide the motion trajectory and attitude information of the UAV.
[0003] However, single-modal image information has limitations. Visual images are affected by factors such as noise and low contrast, which may lead to difficulties in identifying some behaviors; although IMU data can provide flight path information, it lacks spatial details when identifying specific illegal behaviors. Therefore, relying solely on a certain type of modal data for illegal behavior identification may reduce the accuracy. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, system, device and medium for identifying illegal behaviors at the front end of a UAV, which can improve the accuracy and real-time performance of illegal behavior identification through a multi-modal fusion technology that combines visual images and IMU data.
[0005] To achieve the above purpose, the technical solution provided by the present invention is as follows:
[0006] The first aspect of the present application provides a method for identifying illegal behaviors at the front end of a UAV, including the following steps:
[0007] Collect video frame images through a camera carried by the UAV, and at the same time obtain IMU data through an IMU carried by the UAV to form a dataset for identifying illegal behaviors at the front end of the UAV; divide the dataset for identifying illegal behaviors at the front end of the UAV into a training set and a test set;
[0008] Construct a model for identifying illegal behaviors at the front end of the UAV, and use the training set and the test set to train and evaluate the model respectively; wherein, the model for identifying illegal behaviors at the front end of the UAV includes:
[0009] The video frame images are pre-processed with the generative adversarial network defogging module; the video frame images after defogging pre-processing are feature extracted to obtain visual features; the inertial measurement unit data is feature extracted to obtain inertial features; the visual features and inertial features are combined through the fusion module to generate a unified multi-modal feature representation; the hybrid convolutional neural long short-term memory network is used to obtain the modeling time series characteristics of the multi-modal feature representation to achieve the classification and positioning of violations;
[0010] The video frame images collected by the drone camera in real time and the inertial measurement unit data collected by the inertial measurement unit in real time are input into the trained and evaluated drone front-end violation behavior recognition model to generate the recognition results of the violation behavior.
[0011] To optimize the above technical solutions, the specific measures taken also include:
[0012] The generative adversarial network defogging module consists of two parts: a generator and a discriminator. The generator is responsible for generating a defogged image, and the discriminator is used to determine whether the generated image is a real defogged image or a generated fake image.
[0013] The identification result includes the category of the illegal behavior, as well as the corresponding time and corresponding spatial position of the illegal behavior.
[0014] The hybrid convolutional neural long short-term memory network is used to obtain the modeling time series characteristics of multimodal feature representation, realize the classification and positioning of illegal behaviors, and generate the identification results of illegal behaviors, specifically:
[0015] A hybrid convolutional neural long short-term memory network is used to process multimodal feature representation. The hybrid convolutional neural long short-term memory network consists of a convolutional neural network and a long short-term memory network. The convolutional neural network is responsible for extracting spatial features in the multimodal feature representation, and the long short-term memory network is responsible for capturing the temporal relationship in the multimodal feature representation. The recognition results are shown as follows:
[0016] (h t , v t ,φ t )=RNN(z t ,h t-1 )
[0017] where h t is the hidden state at the current moment, h t-1 is the hidden state of the previous moment. The long short-term memory network transfers memory through recurrent connections. t is the visual features and inertial features fused at the current moment, v t and φ t Represent the predicted spatial location information and behavior category respectively.
[0018] The described generative adversarial network dehazing module is improved using DenseNet-121. DenseNet-121 is introduced into the generator of the generative adversarial network dehazing module to enhance the complexity and realism of the generated image. For the input low-quality video frame image I raw is processed, and an enhanced image is generated by the generator, which is expressed as follows:
[0019] I enhanced = G dehaze (I raw )
[0020] where I raw is the original image that may contain haze interference, and G dehaze is the improved generator of the generative adversarial network.
[0021] The described video frame image after dehazing preprocessing is subjected to feature extraction to obtain visual features; the inertial measurement unit data is subjected to feature extraction to obtain inertial features, including:
[0022] Using a convolutional neural network to extract spatial features of the enhanced image, and the convolutional neural network extracts local and global features in the image through convolutional layers:
[0023] x v = E visual (I enhanced )
[0024] where x v is the image feature vector, I enhanced represents the spatial features extracted from the enhanced image, and E visual is the visual feature extraction network;
[0025] Using a one-dimensional convolutional network or a fully connected network to extract the motion features of the inertial measurement unit data:
[0026] x i = E inertial (M imu )
[0027] where x i is the inertial feature vector, representing the motion features extracted from the inertial measurement unit data M imu , and E inertial is the inertial feature extraction network.
[0028] The described visual features and inertial features are combined through a fusion module to generate a unified multi-modal feature representation, specifically:
[0029] Combining the image feature vector x v and the inertial feature xi Perform feature splicing and weighted fusion:
[0030] z t = g concat (x v ,x i )
[0031] where z t is the fused multi-modal feature representation and contains the joint information of images and motion, and g concat is the feature splicing and fusion function.
[0032] The second aspect of the present application provides a UAV front-end illegal behavior recognition system, including:
[0033] A UAV data acquisition module for real-time collecting video frame images through a camera carried by the UAV, and at the same time obtaining inertial measurement unit data through an inertial measurement unit carried by the UAV;
[0034] A UAV front-end illegal behavior recognition model construction module for constructing a UAV front-end illegal behavior recognition model for detecting and recognizing illegal behaviors;
[0035] An illegal behavior detection module for inputting the video frame images real-time collected by the UAV camera and the inertial measurement unit data real-time collected by the inertial measurement unit into the UAV front-end illegal behavior recognition model to generate the recognition result of the illegal behavior;
[0036] Among them, the UAV front-end illegal behavior recognition model construction module includes:
[0037] A UAV image preprocessing module for performing defogging preprocessing on the video frame images by using a generative adversarial network defogging module;
[0038] A feature extraction module for extracting features from the defogging preprocessed video frame images to obtain visual features; extracting features from the inertial measurement unit data to obtain inertial features;
[0039] A multi-modal feature fusion module for combining the visual features and the inertial features through a fusion module to generate a unified multi-modal feature representation;
[0040] A time series modeling behavior recognition module for obtaining the modeling time series characteristics of the multi-modal feature representation by using a hybrid convolutional neural long short-term memory network to realize the classification and positioning of illegal behaviors.
[0041] The third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned UAV front-end illegal behavior recognition method is implemented.
[0042] The fourth aspect of the present application provides a computer-readable storage medium storing a computer program, which causes a computer to execute the method for identifying illegal front-end behaviors of an unmanned aerial vehicle as described above.
[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0044] The present invention uses a multi-modal feature extraction module to extract features from images and inertial measurement unit data respectively, captures the object information in the images and the movement trajectories of the aircraft, combines the characteristics of the camera video frame images and the inertial measurement unit data, effectively retains the spatial information and movement trajectory features of the images through the feature extraction module, and captures the key behavior features in the multi-modal mode, providing high-quality input for subsequent fusion and enhancing the accuracy of behavior recognition.
[0045] The present invention ensures the clarity and quality of the images by de-hazing and enhancing the collected images.
[0046] The present invention effectively integrates the information of images and motion data through a multi-modal feature fusion module based on multi-modal fusion, generates a multi-modal feature representation, integrates spatial information and motion information, and improves the recognition accuracy and positioning effect of illegal behaviors globally.
[0047] Based on the fused features, the present invention captures the temporal changes of behaviors, classifies and identifies illegal behaviors, and the output module generates the recognition results of illegal behaviors, including behavior categories, time, and spatial positions.
[0048] Through the collaborative work of each module, the present invention improves the real-time performance of illegal behavior recognition. The generated results have the characteristics of high precision, strong real-time performance, and scene adaptability, improving the accuracy and efficiency of illegal detection, and having significant application value in intelligent transportation and urban management. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a structural diagram of the system for identifying illegal front-end behaviors of an unmanned aerial vehicle according to the present invention.
[0050] Figure 2 is a schematic diagram of the multi-modal feature fusion technology in the present invention.
[0051] Figure 3 The left half is the low-resolution front-end image of an unmanned aerial vehicle provided by the present invention, Figure 3 and the right half is the effect diagram after image enhancement processing. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The above content of the present invention will be further described in detail below in the form of specific embodiments. However, it should not be understood that the scope of the above subject matter of the present invention is limited to the following embodiments. All technologies implemented based on the above content of the present invention belong to the scope of the present invention.
[0053] In one embodiment, the present invention provides a method for identifying illegal behaviors at the front end of an unmanned aerial vehicle, including the following steps:
[0054] S1. Collect video frame images through the camera carried by the unmanned aerial vehicle, and at the same time obtain inertial measurement unit data through the inertial measurement unit carried by the unmanned aerial vehicle to form an illegal behavior identification data set at the front end of the unmanned aerial vehicle; divide the illegal behavior identification data set at the front end of the unmanned aerial vehicle into a training set and a test set for training and evaluating the illegal behavior identification model at the front end of the unmanned aerial vehicle;
[0055] S2. Construct an illegal behavior identification model at the front end of the unmanned aerial vehicle, including:
[0056] S2.1 Use a generative adversarial network defogging module to perform defogging preprocessing on the video frame images to remove visual interference and enhance image clarity;
[0057] S2.2 Extract features from the defogging preprocessed video frame images to obtain visual features; extract features from the inertial measurement unit data to obtain inertial features;
[0058] S2.3 Combine the visual features and the inertial features through a fusion module to generate a unified multi-modal feature representation, as Figure 2 shown;
[0059] S2.4 Use a hybrid convolutional neural long short-term memory network to obtain the modeling time series characteristics of the multi-modal feature representation, realize the classification and positioning of illegal behaviors, and generate the identification result of illegal behaviors. The identification result includes the category of illegal behaviors, as well as the corresponding time and corresponding spatial position of the illegal behaviors.
[0060] S3. Input the video frame images collected in real time by the unmanned aerial vehicle camera and the inertial measurement unit data collected in real time by the inertial measurement unit into the trained and evaluated illegal behavior identification model at the front end of the unmanned aerial vehicle to generate the identification result of illegal behaviors
[0061] The generative adversarial network defogging module in step S2.1 consists of a generator and a discriminator. The generator is responsible for generating the defogged image, and the discriminator is used to judge whether the generated image is a real defogged image or a generated fake image.
[0062] The dehazing module of the generative adversarial network is improved using DenseNet-121. DenseNet-121 is introduced into the generator of the dehazing module of the generative adversarial network to enhance the complexity and realism of the generated images. For the input low-quality video frame image I raw it is processed, and an enhanced image is generated by the generator, which is expressed as follows:
[0063] I enhanced = G dehaze (I raw )
[0064] where I raw is the original image that may contain haze interference, and G dehaze is the improved generator of the generative adversarial network, which converts the input drone image into an enhanced image.
[0065] In step S2.2, feature extraction is performed on the dehazed preprocessed video frame image to obtain visual features; feature extraction is performed on the inertial measurement unit data to obtain inertial features, including:
[0066] Use a convolutional neural network to extract spatial features of the enhanced image. The convolutional neural network extracts local and global features in the image through convolutional layers:
[0067] x v = E visual (I enhanced )
[0068] where x v is the image feature vector, and I enhanced represents the spatial features extracted from the enhanced image, and E visual is the visual feature extraction network;
[0069] Use a one-dimensional convolutional network or a fully connected network to extract the motion features of the inertial measurement unit data:
[0070] x i = E inertial (M imu )
[0071] where x i is the inertial feature vector, representing the motion features extracted from the inertial measurement unit data M imu , and E inertial is the inertial feature extraction network.
[0072] In step S2.3, the visual features and inertial features are combined through a fusion module to generate a unified multi-modal feature representation, specifically:
[0073] Combine the image feature vector x v and the inertial feature xi Perform feature splicing and weighted fusion:
[0074] z t = g concat (x v , x i )
[0075] where z t is the fused multi-modal feature representation and contains the joint information of images and motion, and g concat is the feature splicing and fusion function.
[0076] Step S2.4 uses a hybrid convolutional neural long short-term memory network to model the temporal characteristics of the multi-modal feature representation, realize the classification and localization of illegal behaviors, and generate the recognition results of illegal behaviors. Specifically:
[0077] Use a hybrid convolutional neural long short-term memory network to process the multi-modal feature representation. The hybrid convolutional neural long short-term memory network consists of a convolutional neural network and a long short-term memory network. The convolutional neural network part is responsible for extracting the spatial features in the multi-modal feature representation, and the long short-term memory network part is responsible for capturing the temporal relationship in the multi-modal feature representation. The recognition result is expressed as follows:
[0078] (h t , v t , φ t ) = RNN(z t , h t-1 )
[0079] where h t is the hidden state at the current moment, h t-1 is the hidden state at the previous moment. The long short-term memory network passes the memory through cyclic connections. z t is the fused visual feature and inertial feature at the current moment, v t and φ t represent the predicted spatial position information and behavior category respectively.
[0080] In a specific application embodiment, 4000 drone video frame images are manually annotated, and at the same time, the inertial measurement unit data is obtained through the inertial measurement unit carried by the drone to form a drone front-end illegal behavior recognition dataset. The drone front-end illegal behavior recognition dataset is divided into a training set and a validation set according to a ratio of 7:3.
[0081] Train and evaluate the drone front-end illegal behavior according to the above method, and use the trained and evaluated drone front-end illegal behavior recognition model to recognize the illegal behavior. The comparison graph is as Figure 3 shown, Figure 3 The left half shows the low-resolution drone images,Figure 3 The right half shows the effect diagram after image enhancement processing. It can be seen that both the image clarity and detail performance have been improved, the details are presented more clearly, making the recognition of illegal behaviors more accurate.
[0082] In another embodiment of the present invention, a front-end illegal behavior recognition system for drones is proposed, as Figure 1 shown, including:
[0083] A drone data acquisition module, which is used to collect video frame images in real time through a camera carried by the drone, and at the same time obtain inertial measurement unit data through an inertial measurement unit carried by the drone;
[0084] A front-end illegal behavior recognition model construction module for drones, which is used to construct a front-end illegal behavior recognition model for drones to detect and recognize illegal behaviors;
[0085] An illegal behavior detection module, which is used to input the video frame images collected in real time by the drone camera and the inertial measurement unit data collected in real time by the inertial measurement unit into the front-end illegal behavior recognition model for drones to generate the recognition result of illegal behaviors;
[0086] Among them, the front-end illegal behavior recognition model construction module for drones includes:
[0087] A drone image preprocessing module, which is used to perform defogging preprocessing on the video frame images by using a generative adversarial network defogging module;
[0088] A feature extraction module, which is used to extract features from the defogging preprocessed video frame images to obtain visual features; extract features from the inertial measurement unit data to obtain inertial features;
[0089] A multi-modal feature fusion module, which is used to combine the visual features and inertial features through a fusion module to generate a unified multi-modal feature representation;
[0090] A time series modeling behavior recognition module, which is used to adopt a hybrid convolutional neural long short-term memory network to obtain the modeling time series characteristics of the multi-modal feature representation, and realize the classification and positioning of illegal behaviors.
[0091] In another embodiment of the present invention, an electronic device is proposed, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned front-end illegal behavior recognition method for drones is realized.
[0092] In another embodiment of the present invention, a computer-readable storage medium is proposed, storing a computer program, and the computer program enables a computer to execute the above-mentioned front-end illegal behavior recognition method for drones.
[0093] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0094] The foregoing are only preferred embodiments of the present invention and do not impose any formal limitations on the present invention. Any person skilled in the art, without departing from the scope of the technical solution of the present invention and based on the technical essence of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for identifying illegal behaviors of a drone front end, characterized in that: The following steps are involved: The video frame images are collected through the camera carried by the drone, and the inertial measurement unit data is obtained through the inertial measurement unit carried by the drone to form the drone front-end violation behavior recognition data set; the drone front-end violation behavior recognition data set is divided into a training set and a test set; Construct a UAV front-end violation behavior recognition model, and use the training set and test set to train and evaluate the model respectively; wherein, the UAV front-end violation behavior recognition model includes: The video frame images are pre-processed with the generative adversarial network defogging module; the video frame images after defogging pre-processing are feature extracted to obtain visual features; the inertial measurement unit data is feature extracted to obtain inertial features; the visual features and inertial features are combined through the fusion module to generate a unified multi-modal feature representation; the hybrid convolutional neural long short-term memory network is used to obtain the modeling time series characteristics of the multi-modal feature representation to achieve the classification and positioning of violations; The video frame images collected by the drone camera in real time and the inertial measurement unit data collected by the inertial measurement unit in real time are input into the trained and evaluated drone front-end violation behavior recognition model to generate the recognition results of the violation behavior.
2. The method for identifying illegal behaviors of a drone front end according to claim 1, characterized in that: The generative adversarial network defogging module consists of two parts: a generator and a discriminator. The generator is responsible for generating a defogged image, and the discriminator is used to determine whether the generated image is a real defogged image or a generated fake image.
3. The method for identifying illegal behaviors of a drone front end according to claim 1, characterized in that: The identification result includes the category of the illegal behavior, as well as the corresponding time and corresponding spatial position of the illegal behavior.
4. The method for identifying illegal behaviors of a drone front end according to claim 1, characterized in that: The hybrid convolutional neural long short-term memory network is used to obtain the modeling time series characteristics of multimodal feature representation, realize the classification and positioning of illegal behaviors, and generate the identification results of illegal behaviors, specifically: A hybrid convolutional neural long short-term memory network is used to process multimodal feature representation. The hybrid convolutional neural long short-term memory network consists of a convolutional neural network and a long short-term memory network. The convolutional neural network is responsible for extracting spatial features in the multimodal feature representation, and the long short-term memory network is responsible for capturing the temporal relationship in the multimodal feature representation. The recognition results are shown as follows: (h t ,in t ,φ t )=RNN(z t ,h t-1 ) where h t is the hidden state at the current moment, h t-1 is the hidden state of the previous moment. The long short-term memory network transfers memory through recurrent connections. t is the visual features and inertial features fused at the current moment, v t and φ t Represent the predicted spatial location information and behavior category respectively.
5. The method for identifying illegal behaviors of a drone front end according to claim 2, characterized in that: The generative adversarial network dehazing module is improved using DenseNet-121, specifically: DenseNet-121 is introduced into the generator of the generative adversarial network defogging module to enhance the image and input low-quality video frame image I raw Processed, the enhanced image I generated by the improved generator enhanced It is expressed as follows: I enhanced =G dehaze (I raw ) Among them I raw is the original image that may contain haze interference, G dehaze Improved generators for generative adversarial networks.
6. The method for identifying illegal behaviors of a drone front end according to claim 5, characterized in that: The above-mentioned extracting features from the video frame images after defogging preprocessing to obtain visual features; Extract features from the inertial measurement unit data to obtain inertial features, including: Use convolutional neural network to extract spatial features of enhanced images. Convolutional neural network extracts local and global features in images through convolutional layers: x v =E visual (I enhanced ) where x v is the image feature vector, I enhanced represents the spatial features extracted from the enhanced image, E visual It is a visual feature extraction network; Extract motion features from IMU data using a 1D convolutional network or a fully connected network: x i =E inertial (M imu ) where x i is the inertial eigenvector, representing the inertial measurement unit data M imu The motion features extracted from E inertial For inertial feature extraction network.
7. The method for identifying illegal behaviors of a drone front end according to claim 1, characterized in that: The visual features and inertial features are combined through a fusion module to generate a unified multimodal feature representation, specifically: The image feature vector x v and the inertial characteristic x i Perform feature splicing and weighted fusion: z t =g concat (x v ,x i ) where z t is the fused multimodal feature representation and contains the joint information of image and motion, g concat is the feature concatenation and fusion function.
8. A drone front-end violation behavior recognition system, characterized in that: include: The UAV data acquisition module is used to collect video frame images in real time through the camera carried by the UAV, and to obtain inertial measurement unit data through the inertial measurement unit carried by the UAV; A drone front-end illegal behavior recognition model building module is used to build a drone front-end illegal behavior recognition model for illegal behavior detection and recognition; The violation behavior detection module is used to input the video frame images collected by the drone camera in real time and the inertial measurement unit data collected by the inertial measurement unit in real time into the drone front-end violation behavior recognition model to generate the recognition result of the violation behavior; Among them, the drone front-end violation behavior recognition model construction module includes: The drone image preprocessing module is used to perform defogging preprocessing on the video frame images using the generative adversarial network defogging module; The feature extraction module is used to extract features from the video frame images after defogging preprocessing to obtain visual features; and to extract features from the inertial measurement unit data to obtain inertial features; Multimodal feature fusion module, used to combine visual features and inertial features through the fusion module to generate a unified multimodal feature representation; The temporal modeling behavior recognition module is used to obtain the modeling temporal characteristics of multimodal feature representation using a hybrid convolutional neural long short-term memory network to achieve the classification and positioning of illegal behaviors.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for identifying illegal behaviors of the front end of a drone as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the method for identifying illegal behaviors of a front-end drone as described in any one of claims 1 to 7.