Target Behavior Analysis Method Based on Image Detection and Recognition

By adopting an algorithm collaborative analysis system based on edge computing devices in open scenarios, combined with the Inception V3 convolutional neural network optimization model, the problems of limited training samples and environmental interference in aircraft image detection and recognition and target behavior analysis are solved, and more efficient and accurate aircraft target detection and behavior analysis are achieved.

CN117197528BActive Publication Date: 2025-06-17UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310933939.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-06-17
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

In open scenarios, aircraft image detection and recognition and target behavior analysis face problems such as limited training samples and susceptibility to environmental interference. Traditional methods have limitations such as low accuracy, low efficiency and susceptibility to image background noise.

Method used

An algorithm collaborative analysis system based on edge computing devices is adopted to perform image decoding, scale transformation and feature extraction through the collaborative work of terminal devices, edge computing devices and servers. Combined with the Inception V3 convolutional neural network optimization model, aircraft classification and behavior analysis are realized.

Benefits of technology

It improves the accuracy and efficiency of aircraft target detection, enhances the analysis ability of multi-view scenarios, reduces the sensitivity to environmental interference, and improves the accuracy of model identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197528B_ABST
    Figure CN117197528B_ABST
Patent Text Reader

Abstract

The present invention discloses a target behavior analysis method based on image detection and recognition, belonging to the technical field of image recognition. The present invention includes: constructing an algorithm collaborative analysis system based on an edge computing device, including a terminal device, an edge computing device, and a server; a user sends an operation instruction to the edge computing device to determine the selected target object, so that the server can obtain the first task data and the second task data of the target object to complete the aircraft classification recognition result and the behavior type recognition result. The present invention has broken through key technologies such as representation learning based on small sample training data, target behavior transfer learning for open scenarios, and algorithm collaborative analysis in multi-perspective scenarios. Based on the constructed algorithm collaborative analysis system based on the edge computing device, an algorithm collaborative target behavior analysis method based on the edge computing device is realized, and the Inception V3 convolutional neural network is optimized, thereby improving the detection accuracy of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a target behavior analysis method based on image detection and recognition. Background Art

[0002] Images are a carrier for people to store, transmit, and obtain information in daily life. In recent years, the way people obtain information has gradually shifted from text to image and video. With the booming development of artificial intelligence in all walks of life, image recognition and behavior analysis are important fields of artificial intelligence. Image recognition refers to using a computer to process images instead of humans, analyze the content of images, and understand them. Traditional image recognition methods mainly consist of two steps: feature extraction and classification recognition. The overall contour shape features of image targets are obtained through image preprocessing, and then methods such as machine learning or template matching are used to complete the classification and recognition of images. However, traditional image recognition methods have problems such as low accuracy, low efficiency, and being easily affected by image background noise, resulting in unsatisfactory overall recognition effects. With the rise of deep learning, it provides a more accurate and efficient image detection and recognition method. Moreover, when it comes to objects like airplanes whose behaviors are difficult for humans to judge through images, deep learning has the ability to autonomously extract image features and perform classification and behavior analysis.

[0003] In the field of aircraft image recognition and detection, traditional methods include aircraft recognition algorithms based on image matching and principal component analysis (PCA), schemes that use a mixture of various features such as geometric features, principal component analysis features, and HU invariant moment features of aircraft to automatically identify aircraft types, and an aircraft target detection algorithm based on the combination of DSmT theory and SVM for multi-feature fusion proposed for aircraft targets in different flight postures. The extracted features are used with specific calculation methods to determine the category of aircraft in the image. Currently, automatically detecting and recognizing aircraft images and quickly judging their target behaviors has always been an urgent problem to be solved. Traditional image recognition and behavior analysis methods have certain limitations and are easily affected by open conditions. Specifically, there are the following aspects:

[0004] (1) Research on aircraft image detection and recognition and target behavior analysis, aiming at problems such as limited training samples for target behavior analysis in open scenarios and being easily affected by environmental interference.

[0005] (2) Aiming at problems such as difficulties in aircraft target behavior analysis and limited aircraft datasets.

[0006] (3) Aiming at technical problems such as algorithm collaborative analysis in multi-view scenarios and edge computing devices. Summary of the Invention

[0007] In view of the technical problems such as limited training samples for target behavior analysis in open scenarios and susceptibility to environmental interference, the present invention discloses a target behavior analysis method based on image detection and recognition.

[0008] The technical solution adopted by the present invention is as follows:

[0009] A target behavior analysis method based on image detection and recognition, the method comprising the following steps:

[0010] Step 1: Construct an algorithm collaborative analysis system based on edge computing devices, the system comprising a terminal device, an edge computing device (also referred to as an edge device) and a server;

[0011] The terminal device includes a first terminal device and a second terminal device. Among them, the first terminal device is used to collect aircraft images from the first perspective of the target aircraft based on the flight attitude and store them in the aircraft image database at its own end; the second terminal device is used to collect aircraft images from the first perspective of the target aircraft based on the flight attitude and store them in the aircraft image database at its own end; the first perspective and the second perspective are preset angles;

[0012] An image database at the terminal device layer is obtained based on the aircraft image databases at the respective terminal devices.

[0013] The edge computing device includes a first edge computing device and a second edge computing device;

[0014] The first edge computing device selects aircraft images from the first perspective of a specified target from the image database at the terminal device layer, decodes the selected images to obtain image data in the form of a two-dimensional matrix of image channels, and then performs scale transformation processing on the image data to obtain first task data, so that the obtained first task data matches the input of the aircraft classification model preset on the server;

[0015] The second edge computing device obtains aircraft images from the second perspective of a specified target from the image database at the terminal device layer, decodes the obtained images to obtain image data in the form of a two-dimensional matrix, and then performs scale transformation processing on the image data to obtain second task data, so that the obtained second task data matches the input of the behavior analysis model preset on the server;

[0016] An image database at the edge computing device layer is obtained based on the first task data and the second task data and transmitted to the server; it can be that each edge computing device transmits it to the server based on its corresponding image transmission module, or the transmission of this image database can be completed based on a common image transmission module;

[0017] The server is pre - installed with an aircraft classification model and a behavior analysis model. The server receives the image database from the edge computing device layer and stores it in the server database, and inputs the first task data therein into the aircraft classification model. The aircraft classification model outputs the aircraft classification and recognition result of the current image based on the image detection and recognition algorithm; the aircraft classification model is used to output the classification probability of each aircraft category, and the maximum probability is the aircraft classification and recognition result of the current image; the behavior analysis model outputs the behavior type recognition result of the aircraft of the specified category based on the image detection and recognition algorithm; the behavior analysis model is used to output the classification probability of each specified behavior type, and the maximum probability is the behavior type recognition result of the current object.

[0018] The server performs collaborative analysis on the aircraft classification and recognition result and the behavior type recognition result based on the algorithm collaborative analysis algorithm to obtain the final detection and recognition result and return it to the terminal device and / or the edge computing device.

[0019] Step 2: The user issues an operation instruction to the edge computing device to determine the selected target object, so that the server can obtain the first task data and the second task data of the target object.

[0020] Step 3: When the server receives the first task data and the second task data of the target object, it inputs the first task data into the aircraft classification model to obtain the aircraft classification and recognition result.

[0021] If the current aircraft classification and recognition result is of the specified category and the corresponding classification probability is greater than or equal to the preset first threshold, then input the corresponding second task data into the behavior analysis model to obtain the behavior type recognition result; if the classification probability corresponding to the current behavior type recognition result is greater than or equal to the preset second threshold, the server returns the determined aircraft classification and recognition result and behavior type recognition result to the terminal device and / or the edge computing device.

[0022] Furthermore, the server also includes incrementally learning and training the aircraft classification model and the behavior analysis model based on the accumulated first task data and second task data respectively to further improve the recognition accuracy of the models.

[0023] Furthermore, the aircraft classification model is an optimized Inception V3 convolutional neural network model. In this optimized model, the Relu activation function in the Inception V3 model is replaced with the LeakyRelu function, and the dropout layer in the Inception V3 model is replaced with the dropblock layer. The dropblock layer discards regions of 2×2 and above to better improve the detection effect; the aircraft classification model of the present invention is obtained by transfer learning for the optimized Inception V3 convolutional neural network model.

[0024] Further, the recognition categories of the aircraft classification model include: fighter jets, transport aircraft, and general aircraft.

[0025] Further, the behavior analysis model of the present invention is obtained through transfer learning of the optimized model of the Inception V3 convolutional neural network. Preferably, the behavior analysis model is mainly used for classifying and identifying the behavior types of fighter jets.

[0026] Further, when the aircraft classification model and the behavior analysis model are trained by transfer learning, the acquisition and preprocessing methods of the image data set related to training are specifically as follows:

[0027] Set the type of the aircraft classification model and the labels corresponding to each aircraft type;

[0028] In the public data set of aircraft images, randomly extract images of each label type, obtain a specified number of sample images for each aircraft type, and obtain the initial training set for each aircraft type; for a certain aircraft type, if the specified number of sample images cannot be extracted, obtain the side images of the aircraft corresponding to the current aircraft type through web crawling to meet the requirements of the number of sample images;

[0029] Randomly extract a specified number of images (such as 20 images) from each initial training set to form a small sample training set for each aircraft type; then perform data augmentation processing on each small sample training set, including: operations such as image flipping, rotation, and taking grayscale images; preferably, perform rotation by 90 degrees three times, mirror processing once, and grayscale processing once;

[0030] Classify the labels of the sample images of the initial training set of the target aircraft category (the aircraft category for behavior classification) to obtain the behavior category labels of the target aircraft category; for example, classify the behavior into landing, takeoff, reconnaissance (such as terrain reconnaissance), exercise, penetration behavior, etc.;

[0031] Extract a specified number of images from each behavior label to obtain a behavior analysis training set, and perform data augmentation processing on it. Preferably, the data augmentation processing of the behavior analysis training set is specifically: rotate by 180 degrees once, mirror processing once, and grayscale processing once.

[0032] The technical solution provided by the present invention at least brings the following beneficial effects:

[0033] The present invention has broken through key technologies such as representation learning based on small-sample training data, target behavior transfer learning for open scenarios, and algorithm collaborative analysis in multi-view scenarios. Based on the established algorithm collaborative analysis system based on edge computing devices, an algorithm collaborative target behavior analysis method based on edge computing devices is realized, and the Inception V3 convolutional neural network is optimized, thereby improving the detection accuracy of target objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0035] Figure 1 It is a schematic diagram of the system architecture of the algorithm collaborative analysis system based on edge computing devices provided by the embodiments of the present invention;

[0036] Figure 2 It is a schematic diagram of the processing process of a target behavior analysis method based on image detection and recognition provided by the embodiments of the present invention;

[0037] Figure 3 It is a schematic diagram of the processing process of the image acquisition module adopted by the embodiments of the present invention;

[0038] Figure 4 It is a schematic diagram of the processing process of the functional modules of the edge computing device provided by the embodiments of the present invention;

[0039] Figure 5 It is a schematic diagram of the processing process of the functional modules of the server provided by the embodiments of the present invention.

[0040] Figure 6 Schematic diagram of the system simulation interface adopted by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.

[0042] Currently, there are few publicly available aircraft image datasets on the Internet, and most of the aircraft images studied in many papers are remote sensing images rather than aircraft images obtained on the ground. In view of this, the embodiments of the present invention propose a target behavior analysis method based on image detection and recognition, which is mainly implemented in the following aspects:

[0043] (1) Research on aircraft image detection and recognition and target behavior analysis. Aiming at problems such as limited training samples for target behavior analysis and susceptibility to environmental interference in open scenarios, data augmentation and data processing for noise reduction are adopted.

[0044] First, collect and organize the aircraft image dataset. Independently collect and organize the aircraft side image dataset to obtain the initial training dataset. Currently, the aircraft side image dataset is relatively small. In the embodiment of the present invention, based on the subset CALTECH-Airplane of the CALTECH-101 dataset of the California Institute of Technology, aircraft side images in the network are crawled to expand the dataset, and then the aircraft side image dataset Airplane side image and the fighter behavior image dataset fighterside image are obtained through manual annotation.

[0045] Then, preprocess the aircraft images under small sample and open conditions, such as data augmentation and noise reduction. Combining the image enhancement method of deep learning, perform operations such as rotation, mirroring, and grayscale processing on the small sample dataset to expand the dataset, which helps the feature learning of the Inception V3 model later and solves the small sample problem in the field of aircraft images.

[0046] (2) Aiming at problems such as difficulties in aircraft target behavior analysis and limited aircraft datasets, convolutional neural networks and transfer learning are used to solve them.

[0047] Select the optimal network model for transfer learning through experimental data, retrain some layers of the model to classify aircraft images and perform behavior analysis, and further optimize the model according to experimental data and theoretical knowledge to obtain the final model.

[0048] First, aiming at problems such as small sample training sets and the particularity of aircraft behavior analysis, use deep learning and transfer learning to propose an optimized Inception V3 neural network model to achieve target behavior analysis based on image classification.

[0049] Secondly, aiming at problems such as aircraft image target detection, recognition, and behavior analysis in small sample training sets, through comparative experiments, the Inception V3 model is selected to be transferred to this task. The aircraft icon behavior analysis is divided into two types of tasks: aircraft type classification and aircraft behavior classification according to the edge computing device architecture. Optimize the Inception V3 model by replacing the dropout layer with the dropblock layer and replacing the Relu function with the Leaky Relu function, and verify through experiments that the two optimization schemes improve the detection accuracy of the two types of tasks under the initial training set and the small sample training set respectively.

[0050] Then, explore the effectiveness of Inception V3 and other neural networks in the detection, recognition, and classification of aircraft side images, and propose optimization schemes such as replacing dropblock to further improve the detection accuracy of the Inception V3 model in aircraft side image classification.

[0051] Again, conduct model control experiments and result analysis. First, conduct comparative experiments on traditional CNN, Inception V1 and Inception V3, the original data training set, and the small sample training set. According to the experimental results, it is found that Inception V3 has better performance. Subsequently, give network optimization operations and experimental data to obtain the best network model and parameters.

[0052] (3) For technical issues such as algorithm collaborative analysis and edge computing devices in multi-view scenarios, construct a simulation system composed of multiple edge devices and servers for experimental verification to meet the requirements of fast and real-time data exchange between edge devices and servers.

[0053] First, divide the target behavior analysis task based on aircraft side images into two categories: classification task and behavior analysis task, and propose an aircraft target behavior analysis framework based on the collaboration of edge computing devices by matching the tasks with the corresponding edge computing devices.

[0054] Secondly, conduct algorithm collaborative analysis based on edge computing devices. Give the key data for transmission and the collaborative algorithm between devices. Two edge devices respectively process different image recognition tasks and transmit the results to the server. The server uses the optimized Inception V3 model for calculation, and finally comprehensively judges the target behavior through the detection results.

[0055] In view of the problem that deep learning requires a certain amount of data set training and optimization to achieve better results, this invention constructs a corresponding aircraft data set and labels the types to prepare for subsequent network training. At the same time, for the problem of the small sample training set, the initial data set is expanded in a specific way to meet the requirements of the convolutional neural network for the number of images.

[0056] In the embodiments of the present invention, the dataset mainly comes from the CALTECH-Airplane subset of the CALTECH-101 dataset of the California Institute of Technology. At the same time, according to the dataset and label requirements, some side images of airplanes are crawled from the Internet to expand the dataset. The entire dataset contains more than 2,000 airplane images. According to the characteristics of different airplanes, the images are labeled into three categories, namely "fighter", "transport plane", and "aircraft", representing fighter planes, transport planes, and ordinary airplanes respectively. The fighter planes are further classified into "detect", "penetrate", "exercise", "land", and "launch", representing reconnaissance behavior, penetration behavior, exercise behavior, landing behavior, and takeoff behavior respectively.

[0057] To ensure the accuracy of the processing results, images of each label type are randomly selected. Each initial training set consists of 20 airplane images of this type, and 50 images of each type are selected to form the test set.

[0058] Before image processing, it is necessary to obtain the features of the data according to the distribution of the data in the image. The airplane image dataset used in the present invention has only 20 images in each initial training set. To meet the dependence of deep learning on the data scale, a batch data processing script is now written to use OpenCV functions to perform operations such as image flipping, rotation, and grayscale extraction on the initial images to expand the dataset.

[0059] Among them, image rotation rotates the entire image clockwise or counterclockwise by a certain angle about a certain point in the image (usually the center point) to obtain the rotated image, so as to obtain airplane images in different poses. The principle is to select the rotation center, and then transform the coordinate system to take the rotation center as the coordinate origin. The transformation process is shown in formula (1), where (x, y) is the coordinate of the image coordinate system, and (x', y') is the coordinate of the transformed coordinate system.

[0060]

[0061] After transforming to the coordinate system with the rotation center as the origin, the transformed coordinates are rotated. The transformation process is shown in formula (2), where (x'', y'') is the transformed coordinate, and α is the clockwise rotation angle.

[0062]

[0063] After the rotation operation, another transformation is required to change the coordinate origin of the coordinate system from the rotation center back to the coordinate origin of the image. The transformation operation is shown in formula (3), (x''', y''') is the transformed coordinate, left is the abscissa of the leftmost point after rotation, and top is the ordinate of the uppermost point after rotation.

[0064]

[0065] Image flipping is a way to augment the dataset through geometric coordinate transformation. The present invention adopts the horizontal image flipping method, and the principle is to make the abscissa of the image pixel points symmetric about the midline of the image. The transformation process is as shown in formula (4).

[0066]

[0067] A grayscale image is a special color image with the same three components of B, R, and G. The three components of B, R, and G together determine the color of a pixel point. The reasons for taking grayscale images are as follows: firstly, grayscale images can retain the changes in the brightness and chromaticity of the entire image, while reducing the calculation data and thus the calculation amount; secondly, since they are different from some data of color images, it can be explored whether color will affect the result of the neural network's recognition of aircraft images. There are many methods for taking grayscale images. The method adopted in the present invention is maximum grayscale processing, as shown in formula (5):

[0068] B = G = R = max([B, G, R]) (5)

[0069] In the embodiment of the present invention, more than 1000 aircraft images are used as the initial dataset. After aircraft type annotation, more than 200 "aircraft" labeled images, more than 100 "fighter" labeled images, and more than 600 "transport plane" labeled images are obtained. The number of "aircraft" and "fighter" labeled images is small. Some side images of aircraft are crawled from the Internet to augment the dataset. After augmentation, a total of 400 "aircraft" labeled images and more than 400 "fighter" labeled images are obtained. After removing the test set images, the experimental initial dataset origin classification is obtained.

[0070] Randomly select 20 images from each label in the origin classification to form the initial images of the small sample training set. Subsequently, write a script using OpenCV functions to rotate each group of images 90 degrees three times, perform mirroring once, and perform grayscale processing once. Each label's images are expanded from 20 to 320 to obtain the small sample classification training set small simpleclassification. Extract 100 images from each label in the origin classification as the test set testclassification. Reclassify the "fighter" label in the origin classification to obtain more than 100 "land" label images, more than 60 "launch" label images, more than 70 "detect" label images, more than 80 "exercise" label images, and more than 80 "penetrate" label images, representing landing, takeoff, reconnaissance, exercise, and penetration behaviors respectively.

[0071] Remove 20 images from each label to obtain the behavior analysis training set test behavior and the initial behavior analysis training set origin behavior. Extract 20 images from each label in the origin behavior. Since the aircraft inclination angle may affect the experimental results through experiments, perform rotation by 180 degrees once, mirroring once, grayscale processing once, and expand each group of 20 images to 160 images to obtain the small sample behavior analysis training set small simplebehavior.

[0072] In the embodiment of the present invention, the Inception V3 model uses the Relu function as the activation function. Since the Relu function may cause some neurons to be unable to extract features normally. The LeakyRelu function is an optimized version of the Relu function, and its calculation formula is shown in formula (6):

[0073]

[0074] When x <= 0, its derivative is a, rather than 0 of the Relu function, solving the problem that some neurons cannot work properly. LeakyRelu can solve the problem of neuron "death". LeakyRelu is very similar to Relu, only different in the part where the input is less than 0. For the part where the Relu input is less than 0, the value is 0, while for the part where the LeakyRelu input is less than 0, the value is negative and has a small gradient.

[0075] Compared with the Relu function, the LeakyRelu function has the following advantages in the task of analyzing the target behavior of aircraft side images:

[0076] (1) Regarding the Dead Relu problem existing in the Relu function, when the input of the Leaky Relu function is negative, a very small slope is given to the input value. On the basis of solving the zero-gradient problem in the case of negative input, the Dead Relu problem is also well alleviated, that is, the problem that some neurons cannot work properly is solved, thereby improving the model detection accuracy.

[0077] (2) The output of this function ranges from negative infinity to positive infinity, that is, leaky expands the range of the Relu function. The value of α is generally set to a small value, such as 0.01. Expanding the range of the Relu function makes it more adaptable to the detection of certain special images or special data.

[0078] In view of the small training samples and the easy overfitting problem of the present invention, it is proposed to replace the dropout layer in the neural network with a dropblock layer on the basis of Inception V3. The important parameters in dropout are blocksize and γ. Blocksize represents the size of the discarded area, and γ represents the probability of discarding. The inspiration for the dropout (random inactivation) layer comes from genetic inheritance. Some connections between two layers of neurons are randomly blocked to reduce the dependence between nodes, thereby realizing the regularization of the neural network. Reflected in the image, after obtaining the feature map, some pixels are randomly discarded. The feature distribution of aircraft images is relatively concentrated. Using the dropblock method to discard areas of 2×2 and above can better improve the detection effect. In addition, a dropblock layer can also be added to the input layer and a larger blocksize can be set to randomly dig out the original image to achieve the purpose of data augmentation.

[0079] On the basis of optimizing the model parameters, the present invention proposes two model optimization schemes. Optimization one on the basis of the original model is to replace the activation function Relu function with the LeakyRelu function, and optimization two is the optimization algorithm of replacing the dropout layer with the dropblock layer. Transfer learning is used for aircraft image classification and behavior analysis, and three versions of neural network models are obtained. After optimization, the data indicators of the optimized models are analyzed.

[0080] The network model of the Inception V3 model algorithm has a total of 46 layers, consisting of 10 Inception modules, convolutional layers, pooling layers, input layers, output layers, etc.

[0081] Based on the Inception V1 model (as shown in Table 1 specifically) and the Inception V3 model, corresponding control experiments are designed, namely the model control experiment and the Inception V3 model control experiment. The model control experiment selects a suitable image detection and recognition algorithm, and the Inception V3 model control experiment further verifies whether the optimization of the corresponding network model can achieve the expected effect. The Inception V1 and Inception V3 models have relatively excellent performance in image classification. The original models are trained with image training sets in the order of hundreds of thousands or millions. To solve their detection and recognition effects in the field of aircraft side images and under the condition of small sample training sets, experiments need to be conducted for testing.

[0082] Currently, there are already many relatively mature neural network models. Through the model control experiment, the present invention selects the model with the best performance in aircraft images as the model for transfer learning. The models used in the model control experiment are the traditional CNN model, the Inception V1 model, and the Inception V3 model. The CNN model is a simple convolutional neural network model built independently, consisting of an input layer, two convolutional layers, two max-pooling layers, and a fully connected layer, a total of six layers. The filter size in the convolutional layer is 3×3, Relu is the activation function, and softmax is the loss function. The Inception V1 model is the first-generation GoogleNet model of Google, which consists of nine Inception modules, two convolutional layers, four max-pooling layers, an average pooling layer, etc. The specific parameters are shown in Table 1, and the output layer is similar to that of Inception V3.

[0083] Table 1 Composition and Parameters of Inception V1 Model

[0084] Type Convolution Kernel Size / Stride Output Size Convolutional Layer 7×7 / 2 112×112×64 Max Pooling Layer 3×3 / 2 56×56×64 Convolutional Layer 3×3 / 1 56×56×192 Max Pooling Layer 3×3 / 2 28×28×192 Inception Module (3a) 28×28×256 Inception Module (3b) 28×28×480 Max Pooling Layer 3×3 / 2 14×14×480 Inception Module (4a) 14×14×512 Inception Module (4b) 14×14×512 Inception Module (4c) 14×14×512 Inception Module (4d) 14×14×528 Inception Module (4e) 14×14×832 Max Pooling Layer 3×3 / 2 7×7×832 Inception Module (5a) 7×7×832 Inception Module (5b) 7×7×1024 Average Pooling Layer 7×7 / 1 1×1×1024 Dropout Layer 1×1×1024 Fully Connected Layer Logits 1×1×1000 Softmax Layer Classification 1×1×1000

[0085] Control Experiment:

[0086] The first group of controls uses the training set origin classification to train the three models. After training is completed, the test set test classification is used to test the models, and the experimental results are obtained for analysis.

[0087] The second group of controls uses the training set small simple classification to train the three models. After training is completed, the test set test classification is used to test the models, and the experimental results are obtained for analysis.

[0088] The control experiments of Inception V3 are mainly divided into two parts, corresponding to the functions of two edge computing devices: classification and behavior analysis. According to the proposed optimization scheme, the Inception V3 model is optimized. The first optimization scheme is to replace the Relu activation function in the Inception V3 model with the optimized version LeakyRelu function to obtain InceptionV3 Optimization 1, which solves the problem that some neurons cannot extract features normally due to the derivative of the Relu function being 0 in the negative region. The second optimization scheme is to replace the dropout layer in the Inception V3 model with the dropblock layer. By increasing the block_size and randomly discarding the n×n area in the feature map, where n is the size of the block_size and the block_size is set to 2, Inception V3 Optimization 2 is obtained. Since the feature distribution of aircraft images is relatively concentrated, this optimization alleviates the overfitting problem and improves the training accuracy.

[0089] Control experiment:

[0090] For the first group of controls, after training the three versions of the model using the training set origin classification, the model is tested using the test set test classification, and the experimental results are analyzed. Then, the three versions of the model are trained using small simple classification, and after training is completed, the model is tested using the test set test classification, and the experimental results are analyzed.

[0091] For the second group of controls, after training the three versions of the model using the training set origin behavior, the model is tested using the test set test behavior, and the experimental results are analyzed. Then, the three versions of the model are trained using small simple behavior, and after training is completed, the model is tested using the test set test behavior, and the experimental results are analyzed.

[0092] The purpose of designing the first group of control experiments is to explore the detection accuracy of each model in detecting and classifying aircraft image targets without the limitation of few-shot learning. The purpose of designing the second group of control experiments is to explore the comparison of the detection accuracy of each model under the condition of few-shot learning and whether the model will be affected by few-shot learning compared with the first group of control experiments.

[0093] In the first group of control experiments, both the Inception V1 and Inception V3 models had good training accuracies, which were 94.2% and 96.8% respectively, while the traditional CNN model had a poor training accuracy of 75.1%. The test accuracy of the Inception V3 model was 91.67%, which was approximately 4% higher than that of the Inception V1 model (87.78%), while the influence of the traditional CNN model on the test accuracy training effect was still relatively low, at 67.10%. In the second group of control experiments, the limitation condition of few-shot learning was added. The result was that the training accuracy of the Inception V3 model was hardly affected, but the Inception V1 and traditional CNN models were affected. The Inception V1 model was less affected, and its training accuracy decreased from 94.2% to 92.3%, while the traditional CNN model was more affected, and its training accuracy decreased by 18.2%. The test accuracies of the Inception V1 and Inception V3 models were affected to a certain extent, decreasing by 4.11% and 2% respectively. Due to the poor training effect of the traditional CNN model, the test accuracy was only 52.22%. Through the two groups of control experiments, it can be considered that the few-shot learning limitation has an impact on the test accuracies of the models, and the impact sizes are different according to different models. In completing the task of object detection, recognition, and classification of aircraft images, both the Inception V1 and Inception V3 models have good effects and are less restricted by the few-shot training set, but the Inception V3 model has higher training accuracy and test accuracy. The reason why the CNN model has a poor effect is that in addition to the reasons of the network structure itself, the aircraft images in the dataset were obtained under open conditions, and the images more or less have certain noise, which affects the models with weak detection capabilities.

[0094] In summary, it can be considered that the Inception V3 model has good detection capabilities in aircraft image classification and can also have good adaptability under limited conditions such as few-shot training sets and open conditions. Therefore, the present invention selects Inception V3 as the source model for transfer learning.

[0095] The first group of control experiments was for the first type of task. Among the three versions, all had good training accuracies and test accuracies for the initial classification training set. The training accuracies were 96.8%, 95.6%, and 96.9% respectively, and the test accuracies were 91.67%, 91.67%, and 92.3% respectively.

[0096] Under the condition of the initial classification training set, the training accuracy of the optimized model with the Relu function replaced by the Leaky Relu function decreased, but the test accuracy remained unchanged, indicating that the impact of the optimization on the detection accuracy of a pair of models was very small under this condition. The training accuracy of the optimized model with the dropout layer replaced by the dropblock layer remained almost unchanged, but the test accuracy increased slightly. Whether it has an effect under the condition of the initial classification training set requires verification with larger experimental data. Since the research content of this invention is few-shot learning, no further experimental verification has been carried out. Under the condition of the few-shot classification training set, all three versions had good training accuracy and test accuracy for the few-shot training set, with training accuracies of 96.9%, 95.4%, and 96.2% respectively, and test accuracies of 89.67%, 90.77%, and 91.53% respectively.

[0097] Under the condition of the few-shot classification training set, the training accuracy of the optimized model 1 decreased by 1.5%, and the test accuracy increased by approximately 1%, indicating that the optimization had a certain impact on the training accuracy and a certain improvement on the test accuracy. The training accuracy of the optimized model 2 decreased by 0.7%, and the test accuracy increased by 1.86%, indicating that the optimized model 2 had a slight improvement on the training accuracy and a certain improvement on the test accuracy. The second group of control experiments was for the second type of task. For all three versions, the training accuracies and test accuracies for the initial behavior analysis training set were relatively close, with training accuracies of 87.5%, 86.67%, and 87.2% respectively, and test accuracies of 92.2%, 91.53%, and 92.3% respectively. Under the condition of the initial behavior analysis training set, both the training accuracy and test accuracy of the optimized model 1 decreased slightly, indicating that the optimization had a certain impact on both the training accuracy and test accuracy. The training accuracy and test accuracy of the optimized model 2 remained almost unchanged, indicating that the optimization 2 might have a slight impact on the training accuracy and test accuracy. Under the condition of the few-shot behavior analysis training set, all three versions had good training accuracy and test accuracy for the few-shot training set, with training accuracies of 92.1%, 89.73%, and 93.5% respectively, and test accuracies of 86.15%, 86.67%, and 89.23% respectively. Under the condition of the few-shot behavior analysis training set, the training accuracy of the optimized model 1 decreased by 2.37%, and the test accuracy increased by 0.52%, indicating that the optimization had a certain impact on the training accuracy and a slightly improved impact on the test accuracy under this condition. The training accuracy of the optimized model 2 increased by 1.4%, and the test accuracy increased by 3.38%, indicating that the optimization 2 had a certain improvement on both the training accuracy and test accuracy.

[0098] In addition, when processing the experimental result data, it is found that the misjudged aircraft images have different labels but relatively similar features. This shows that in terms of aircraft behavior analysis, simply analyzing side images has certain difficulties in correctly judging what behavior the aircraft is performing, or there is a certain upper limit of accuracy. Due to the particularity of such aircraft images and the existence of certain noise in aircraft images under open conditions, the detection accuracy of the model is affected. Therefore, a method can be proposed to add features other than the output image for feature fusion in the optimized network to further improve the detection accuracy.

[0099] In the embodiments of the present invention, the target behavior analysis based on the detection and recognition of aircraft side images is divided into two types of tasks, namely the aircraft classification task and the fighter target behavior analysis task. Different terminal devices need to be set to obtain aircraft side images from different perspectives. After the terminal device acquires the image, the corresponding edge device acquires the image for data processing to obtain a two-dimensional matrix and transmits it to the server. The neural network model corresponding to the task is used for calculation to obtain the detection result, and the judgment result is obtained through collaborative analysis of the algorithm and returned to the edge device.

[0100] That is, in the embodiments of the present invention, the algorithm collaborative analysis and processing process based on edge computing devices is as follows:

[0101] (1) The terminal device acquires the image and transmits it to the edge device for storage. The edge device selects the corresponding image for image processing work to obtain the corresponding two-dimensional matrix storing the information of each channel of the stored image.

[0102] (2) The edge device processes to obtain the two-dimensional matrix and transmits the two-dimensional matrix to the server. The server calls different neural network models according to the edge devices corresponding to different tasks for calculation and detection to obtain the result.

[0103] (3) Obtain the corresponding results from different perspectives, conduct collaborative analysis to obtain the final judgment result, and then return the result to all edge devices, with high detection accuracy.

[0104] With the rapid development of the Internet of Things, a large amount of information needs to be sent to the data center in real time, which easily leads to problems such as high latency, unstable network, and low bandwidth in the traditional cloud computing mode, and the requirements in fields such as vehicle networking and intelligent monitoring that require high bandwidth and low latency cannot be met. In order to meet the accuracy and response speed of aircraft image target detection and behavior analysis and simulate the actual application scenario, that is, the combination of edge computing and cloud computing. In the embodiments of the present invention, an algorithm collaborative analysis system based on edge computing devices with an edge computing + cloud computing architecture is proposed, as Figure 1As shown in the figure, it can be divided into three layers: the data access layer, the business logic layer, and the application layer. The database in the data access layer is used to store images of different perspectives of the corresponding target obtained by the terminal device according to the target's flight attitude. The devices involved in the business logic layer include the terminal device, the edge device, and the server. Among them, the terminal device is used to obtain images of different perspectives of the aircraft and store them locally. The edge device selects the images obtained by the terminal device, performs image decoding, scale transformation, etc., and then transmits them to the server. The server implements two types of image task detections (including aircraft classification tasks and fighter target behavior analysis tasks) based on a convolutional neural network. The application layer is used to implement the detection instructions input by the control personnel for the target aircraft, mainly to control the edge device to select images that meet the current requirements. That is, the algorithm collaborative analysis system based on edge computing devices constructed in the embodiments of the present invention includes two terminal devices, two edge devices (edge device A and B), and a server. Among them, an image acquisition module is provided on the terminal device, and an image transmission, image selection, image decoding, and scale transformation module is provided on the edge device. An image detection, recognition, and target behavior analysis module (algorithm collaborative analysis module) is provided on the server. See Figure 2 , the specific functions of each functional module are as follows:

[0105] (1) Image acquisition function module: The terminal device obtains images of different perspectives of the corresponding target according to the target's flight attitude, and stores the images in its own local memory to construct a local aircraft image database. As Figure 3 shown, terminal device A is used to obtain the aircraft image of perspective one (perspective one image), and terminal device B is used to obtain the aircraft image of perspective two (perspective two image). Based on the aircraft images of the two perspectives, a terminal database can be constructed.

[0106] (2) Image selection module: According to the content of the database (the aircraft image database of the terminal device), the corresponding images in the terminal device database are obtained for image detection, recognition, and target behavior analysis.

[0107] (3) Image decoding and scale transformation function module: The obtained image is decoded to obtain a two-dimensional matrix, and the scale of the input image is transformed according to the requirements of the behavior analysis network model in the server to obtain the processed image data, and the data is stored in the corresponding local database (edge database).

[0108] (4) Image transmission module: Selects the processed image data and transmits it to the server for reception.

[0109] As Figure 4 shown, the processing flow of the image selection module, the image decoding and scale transformation function module, and the image transmission module of the edge device is as follows:

[0110] The edge device A is used to process the image of view one acquired by the terminal device A, and the edge device B is used to process the image of view two acquired by the terminal device B. After image scale transformation and image decoding respectively, an edge database is obtained, and then it is uploaded to the server through the image transmission module, and the server stores it in the service database for subsequent processing.

[0111] (5) Image detection, recognition and target behavior analysis module. After the server receives the image data from the edge computing device, according to the tasks corresponding to the corresponding edge devices, it calls the models of the corresponding tasks for behavior analysis. The task corresponding to the edge device A is aircraft classification, and the task corresponding to the edge device B is aircraft target behavior analysis. After obtaining the experimental results, algorithm collaborative analysis is carried out according to the experimental results to obtain the final judgment result, and then the judgment result is stored in the server storage and transmitted back to the edge device.

[0112] (6) Algorithm collaborative analysis module. The role of the edge device A is to obtain aircraft images and transmit the two-dimensional matrix data of the extracted image channels to the server. The server receives the data and then uses the aircraft classification model obtained by transfer learning from the Inception V3 model to calculate the data, and obtains the weight mi that may be a certain type of aircraft, where i represents the i-th type of aircraft, and mi represents the possibility that the target image is the i-th type of aircraft. The role of the edge device B is to obtain aircraft images and transmit the two-dimensional matrix data of the extracted image channels to the server. The server receives the data and then uses the fighter behavior analysis model obtained by transfer learning from the Inception V3 model to calculate the data, and obtains the weight nj that may be a certain behavior, where j represents the j-th behavior, and nj represents the possibility that the target aircraft is performing the j-th behavior. After the server obtains the results of the two models, it conducts analysis. In the case where i is a fighter and mi is greater than a certain value (a preset first threshold), it can be considered that the target in the image is a fighter. At the same time, in the case where nj is greater than a certain value (a preset second threshold), it can be judged what threatening behavior the fighter in the image is performing, realizing algorithm collaborative analysis. After making a judgment, the judgment result is returned to the terminal device, and the edge device stores the acquired image and continues to acquire the image of the next target aircraft.

[0113] Such as Figure 5As shown in the figure, the processing flow of the image detection and recognition module and the algorithm collaborative analysis module of the server is as follows: Extract the data of Task 1 (image data after image scale transformation and image decoding of the image from Perspective 1) and the data of Task 2 (image data after image scale transformation and image decoding of the image from Perspective 2) from the server database. For the data of Task 1, call the model (aircraft classification model) for detection and recognition. For the data of Task 2, call the model (fighter behavior analysis model) for detection and recognition. Then, the algorithm collaborative analysis module performs algorithm collaborative analysis on the two detection and recognition results, stores the obtained algorithm collaborative analysis results in the server database, and feeds back to the terminal device and / or the edge device.

[0114] It should be noted that when the algorithm collaborative analysis system based on the edge computing device described in the above embodiments realizes its functions, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0115] As a possible implementation manner, in the embodiment of the present invention, the control interface of the algorithm collaborative analysis system based on the edge computing device is as Figure 6 shown, including: two terminal devices, two edge devices and a server. The terminal device acquires images, and the edge device processes and transmits the images to the server for calculation. After the server obtains the calculation result, it makes a comprehensive judgment and returns the judgment result. Set the select file and transfer data file controls, select the image file in the corresponding folder, return the image address in the set label, and then click the transfer data control to transmit the selected image to the server for calculation, simulating the process of the edge device acquiring images and uploading the processed data. The server receives the data, calculates the result, and returns it to the set label, simulating the process of the server receiving the data and processing it to obtain the result.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0117] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the inventive concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for target behavior analysis based on image detection and recognition, characterized in that, It includes the following steps: Step 1: Construct an algorithm collaborative analysis system based on edge computing devices, which includes terminal devices, edge computing devices, and servers; The terminal devices include a first terminal device and a second terminal device. Among them, the first terminal device is used to collect aircraft images from the first perspective of the target aircraft based on the flight attitude and store them in the local aircraft image database; the second terminal device is used to collect aircraft images from the first perspective of the target aircraft based on the flight attitude and store them in the local aircraft image database; the first perspective and the second perspective are preset angles; Based on the aircraft image databases of each terminal device, an image database at the terminal device layer is obtained; The edge computing devices include a first edge computing device and a second edge computing device; The first edge computing device selects the aircraft images from the first perspective of the specified target from the image database at the terminal device layer, decodes the selected images to obtain image data in the form of a two-dimensional matrix of image channels, and then performs scale transformation processing on the image data to obtain first task data, so that the obtained first task data matches the input of the aircraft classification model preset on the server; The second edge computing device obtains the aircraft images from the second perspective of the specified target from the image database at the terminal device layer, decodes the obtained images to obtain image data in the form of a two-dimensional matrix, and then performs scale transformation processing on the image data to obtain second task data, so that the obtained second task data matches the input of the behavior analysis model preset on the server; Based on the first task data and the second task data, an image database at the edge computing device layer is obtained and transmitted to the server; The server is preset with an aircraft classification model and a behavior analysis model. The server receives the image database at the edge computing device layer and stores it in the server database, and inputs the first task data into the aircraft classification model. The aircraft classification model outputs the aircraft classification recognition result of the current image based on the image detection and recognition algorithm; the aircraft classification model is used to output the classification probabilities of each aircraft category, and the maximum probability is the aircraft classification recognition result of the current image; the behavior analysis model outputs the behavior type recognition result of the aircraft of the specified category based on the image detection and recognition algorithm; the behavior analysis model is used to output the classification probabilities of each specified behavior type, and the maximum probability is the behavior type recognition result of the current object; The server performs collaborative analysis on the aircraft classification recognition result and the behavior type recognition result based on the algorithm collaborative analysis algorithm to obtain the final detection and recognition result and return it to the terminal device and / or the edge computing device; Step 2, the user sends an operation instruction to the edge computing device to determine the selected target object, so that the server obtains the first task data and the second task data of the target object; Step 3, when the server receives the first task data and the second task data of the target object, it inputs the first task data into the aircraft classification model to obtain the aircraft classification recognition result; When the current aircraft classification recognition result is the specified category and the corresponding classification probability is greater than or equal to the preset first threshold, then input the corresponding second task data into the behavior analysis model to obtain the behavior type recognition result; when the classification probability corresponding to the current behavior type recognition result is greater than or equal to the preset second threshold, the server returns the determined aircraft classification recognition result and behavior type recognition result to the terminal device and / or the edge computing device.

2. The method according to claim 1, characterized in that, The server also includes incrementally learning and training the aircraft classification model and the behavior analysis model based on the accumulated first task data and second task data respectively.

3. The method according to claim 1, characterized in that, The aircraft classification model is an optimized Inception V3 convolutional neural network model. In this optimized model, the Relu activation function in the Inception V3 model is replaced with the LeakyRelu function, and the dropout layer in the Inception V3 model is replaced with the dropblock layer, and the 2×2 and above regions are discarded through the dropblock layer; the aircraft classification model is obtained by transfer learning of the optimized Inception V3 convolutional neural network model.

4. The method according to claim 1, characterized in that, The recognition categories of the aircraft classification model include: fighter aircraft, transport aircraft, and general aircraft.

5. The method according to claim 3, characterized in that, The behavior analysis model is obtained by transfer learning of the optimized Inception V3 convolutional neural network model.

6. The method according to any one of claims 1 to 5, characterized in that, When the aircraft classification model and the behavior analysis model are undergoing transfer learning training, the specific methods for obtaining and preprocessing the image data sets related to training are as follows: Set the type of the aircraft classification model and the corresponding labels for each type; Randomly extract images of each label type in the public aircraft image data set to obtain a specified number of sample images for each aircraft type, and obtain the initial training set for each aircraft type; For a certain aircraft type, if the specified number of sample images cannot be extracted, obtain the side images of the aircraft corresponding to the current aircraft type through web crawling to meet the requirements for the number of sample images; Randomly extract a specified number of images from each of the initial training sets to form a small sample training set for each aircraft type; Then perform data augmentation processing on each small sample training set, including: image flipping, rotation, and grayscale conversion; Classify the labels of the sample images in the initial training set of the target aircraft category to obtain the behavior category labels of the target aircraft category; Extract a specified number of images from each of the behavior labels to obtain the behavior analysis training set, and perform data augmentation processing on it.

Citation Information

Patent Citations

  • Insulator state edge recognition method based on edge calculation and deep learning

    CN113780371A

  • Lightweight multi-task classification model and method for image classification, and edge device

    CN115546552A