Firearm type identification method and system based on deep learning
By using the ResNet-50 deep learning network and feature database, the problems of low efficiency and poor accuracy in firearm identification are solved, and high-precision identification and information output are achieved in complex environments.
Patent Information
- Application Number
- CN202511587356.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies for firearm identification suffer from low efficiency and poor accuracy, especially in complex environments where it is difficult to distinguish between firearms with similar appearances, and they also lack robustness and information support.
A deep learning-based approach is adopted, using ResNet-50 deep convolutional neural network for transfer learning, combined with global average pooling and L2 normalization to extract high-dimensional feature vectors, and constructing a firearm feature database. Recognition is performed through cosine similarity calculation, and detailed information data is output.
It improves the accuracy and robustness of firearm type recognition, enabling accurate identification of firearm types in complex environments and providing structured detailed information to meet real-time or near-real-time recognition requirements.
Smart Images

Figure CN121033818A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing and recognition, in particular, to a firearm type recognition method and system based on deep learning. BACKGROUND
[0002] There are many types of imitation firearms, and the appearance similarity is high. Artificial comparison not only consumes a long time, but also is affected by experience and subjective factors, and there is a situation of recognition lag and high error rate. At present, there is an obvious demand for fast and accurate image recognition technology in the domestic scenes of firearm management, identification and display.
[0003] In order to improve the degree of automation, an automatic recognition method based on traditional image processing and shallow machine learning appears, and common technologies include edge detection, texture analysis, local feature descriptor (such as SIFT, SURF) and classifier based on these features (such as SVM, KNN). However, these methods have obvious limitations in the task of firearm recognition: limited by artificially designed feature expression, it is difficult to describe the subtle structural differences of firearms; poor robustness to complex shooting conditions (scale, angle, light, noise, partial occlusion); when the appearance of firearms is highly similar or in the case of insufficient training samples, the recognition performance decreases significantly.
[0004] Therefore, it is necessary to design a firearm type recognition method and system based on deep learning to solve the problems existing in the current technology, which can not only improve the recognition efficiency and reduce the labor cost, but also establish a standardized recognition process and provide technical support for related units. SUMMARY
[0005] In view of this, the present application provides a firearm type recognition method and system based on deep learning, aiming at solving the problem of poor performance of firearm type recognition.
[0006] In one aspect, the present application provides a firearm type recognition method based on deep learning, comprising: extracting a plurality of feature vector sets based on a deep convolutional neural network model and according to standard firearm images, each of the feature vector sets comprising a corresponding firearm type identifier and detailed information data; and constructing a firearm feature database according to the plurality of feature vector sets; collecting a to-be-recognized firearm image, and performing a preprocessing operation on the to-be-recognized firearm image; input the preprocessed to-be-identified firearm image into a deep convolutional neural network model, the deep convolutional neural network model adopts a ResNet-50 architecture and performs transfer learning on an ImageNet dataset, retains all convolutional layer structures of an ImageNet pre-trained model, and freezes weight parameters of the convolutional layers; removes a fully connected layer and a classification layer of the pre-trained deep convolutional neural network model, and replaces them with an output layer adapted to a firearm identification task to output a high-dimensional feature vector of a fixed dimension; calculate cosine similarities of the high-dimensional feature vector with all feature vector sets in the firearm feature database, generate a similarity value list, sort the similarity value list in descending order, and select the top four firearm type identifiers with the highest similarity values; based on the top four firearm type identifiers, extract the detailed information data corresponding to the top four firearm type identifiers from the firearm feature database.
[0007] Further, when extracting a plurality of feature vector sets based on the deep convolutional neural network model and according to standard firearm images, the method comprises: obtain a plurality of standard firearm images, each standard firearm image corresponding to a known firearm type; perform forward propagation calculation on each standard firearm image, extract a multi-level feature response map through a convolutional layer of the deep convolutional neural network model; input the multi-level feature response map into a global average pooling layer, compress the spatial dimension, and generate a high-dimensional feature vector of a fixed dimension; generate a unique reference image file identifier for each standard firearm image, the reference image file identifier being composed of a firearm type identifier and an image serial number; bind the high-dimensional feature vector and the corresponding firearm type identifier; take the firearm name, the specification parameter, the manufacturer information, the application scenario classification label, and the reference image file identifier as detailed information data, and form a structured feature vector set together with the high-dimensional feature vector.
[0008] Further, when constructing a firearm feature database according to a plurality of feature vector sets, the method comprises: create a feature vector data table in the firearm feature database, the feature vector data table being provided with a feature vector field, a firearm type identifier primary key field, and a detailed information field; store the high-dimensional feature vector in each feature vector set in the feature vector field in the form of a floating-point number array; store the firearm type identifier in the feature vector set as a unique identifier value in the firearm type identifier primary key field; and store the detailed information data in the feature vector set in the detailed information field in a structured data format.
[0009] Further, when performing a preprocessing operation on the to-be-identified firearm image, the method comprises: The preprocessing includes input size adjustment, Gaussian filtering, contrast adjustment, data format normalization, and channel order adjustment.
[0010] Further, when the deep convolutional neural network model adopts a ResNet-50 architecture and performs transfer learning on an ImageNet dataset, the following steps are included: During the transfer learning process, the learnable parameters of the batch normalization layer in the deep convolutional neural network model are fine-tuned; the preprocessed image to be recognized is input into the adjusted deep convolutional neural network model, and a final feature map is generated through the convolutional layer; a global average pooling operation is performed on the final feature map to compress the spatial dimension to 1x1, obtaining a high-dimensional feature vector; and L2 norm normalization processing is performed on the high-dimensional feature vector.
[0011] Further, when performing L2 norm normalization processing on the high-dimensional feature vector, the following steps are included: A square operation is performed on each element in the high-dimensional feature vector, the results of all square operations are accumulated and summed, and a square root operation is performed on the accumulated sum result to obtain an L2 norm value; each element in the high-dimensional feature vector is divided by the L2 norm value to obtain a normalized high-dimensional feature vector.
[0012] Further, when calculating the cosine similarity between the high-dimensional feature vector and all feature vector sets in the firearm feature database to generate a similarity value list, the following steps are included: The feature vector field is read from the feature vector data table of the firearm feature database; the high-dimensional feature vector in each feature vector record field is extracted; the vector inner product of the high-dimensional feature vector of the image to be recognized and the high-dimensional feature vector in the firearm feature database is calculated, and the Euclidean norm of the high-dimensional feature vector of the image to be recognized and the high-dimensional feature vector in the firearm feature database is calculated; and the similarity value is calculated according to the vector inner product and the Euclidean norm; the similarity value and the corresponding firearm type identifier are mapped; all mapping relationships are constructed into a similarity value list containing the corresponding relationship between the firearm type identifier and the similarity value according to the calculation order.
[0013] Further, according to the similarity value list, the following steps are included when sorting in descending order and selecting the top four firearm type identifiers with the highest similarity values: A quicksort algorithm is performed on the similarity values in the similarity value list, and an ordered list is generated according to the similarity values from high to low; when the number of firearm type identifiers in the ordered list is greater than or equal to four, the top four firearm type identifiers are selected; when the number of firearm type identifiers in the ordered list is less than four, all valid firearm type identifiers are selected.
[0014] Further, when extracting the detailed information data corresponding to the first four firearm type identifiers from the firearm feature database, comprising: Performing a database query operation with the selected firearm type identifier as a query condition; extracting the corresponding detailed information data from the firearm feature database.
[0015] Compared with the prior art, the beneficial effects of the present application are that: based on the deep convolutional neural network (ResNet-50 transfer learning) to extract multi-level high-dimensional features and combined with global average pooling and L2 normalization, a feature vector with discriminative ability and fixed length can be generated, thereby improving the consistency and discrimination of firearm type representation, and improving the recognition accuracy and robustness. By using the transfer learning and freezing the convolutional layer, and fine-tuning the batch normalization parameter strategy, the general visual features of the large-scale ImageNet pre-trained model can be used, which reduces the dependence on a large amount of labeled firearm data, and the firearm recognition task can be efficiently fine-tuned, thereby shortening the training time and reducing the labeling cost. By packing the high-dimensional features, type identifiers and structured detailed information of standard firearm images into a feature vector set and establishing a feature vector data table, synchronous management of features and metadata is realized, and the query efficiency and scalability of the database are improved. Using the L2 normalization based on cosine similarity retrieval method, the similarity calculation can be simplified to vector inner product and norm operation, and the calculation overhead is controllable, which is convenient for efficient similarity retrieval (including combination with approximate nearest neighbor index) in actual system, thereby meeting the real-time or near real-time recognition demand. Outputting the first four candidate firearm types with the highest similarity and returning the corresponding structured detailed information (name, specification, manufacturer, application scenario and reference image identifier) not only improves the explainability of the recognition result, but also provides rich context information for manual review, evidence chain preservation or subsequent decision-making, and reduces the risk of misjudgment. The image preprocessing process (size adjustment, Gaussian filtering, contrast adjustment, format normalization and channel order adjustment) enhances the adaptability of the model to input image illumination, noise and scale changes, and improves the stability and generalization ability of the system in complex field environment.
[0016] On the other hand, the present application also provides a firearm type identification system based on deep learning, which is used to apply the above-mentioned firearm type identification method based on deep learning, comprising: A database construction unit configured to extract a plurality of feature vector sets based on the deep convolutional neural network model and according to standard firearm images, each feature vector set comprising a corresponding firearm type identifier and detailed information data; and construct a firearm feature database according to the plurality of feature vector sets; An image acquisition unit configured to acquire a to-be-identified firearm image and perform a preprocessing operation on the to-be-identified firearm image; The first processing unit is configured to input the preprocessed image of the to-be-identified firearm into a deep convolutional neural network model, the deep convolutional neural network model adopts a ResNet-50 architecture and is migrated on an ImageNet dataset, all convolutional layer structures of an ImageNet pre-trained model are retained, and weight parameters of the convolutional layers are frozen; a fully connected layer and a classification layer of the pre-trained deep convolutional neural network model are removed and are replaced with an output layer adapted to a firearm identification task to output a high-dimensional feature vector of a fixed dimension; The second processing unit is configured to calculate cosine similarity of the high-dimensional feature vector with all feature vector sets in the firearm feature database to generate a similarity value list; the similarity value list is sorted in descending order, and the top four firearm type identifiers with the highest similarity values are selected; The data extraction unit is configured to extract the detailed information data corresponding to the top four firearm type identifiers from the firearm feature database based on the top four firearm type identifiers.
[0017] It can be understood that the above-mentioned deep learning-based firearm type identification method and system have the same beneficial effects, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0018] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of preferred embodiments, and are not meant to limit the present application. Furthermore, the same reference numerals are used throughout the several drawings to refer to same or like parts. In the drawings: Figure 1 A flowchart of a deep learning-based firearm type identification method provided by an embodiment of the present application; Figure 2 A functional block diagram of a deep learning-based firearm type identification system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0019] Exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be accurately conveyed to those skilled in the art. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0020] The prior art has significant deficiencies in firearm identification. Traditional manual identification relies on experience and judgment, is easily affected by light, angle, occlusion and subjective factors of the operator, has low identification efficiency and is difficult to guarantee accuracy. The feature extraction capability of methods based on traditional image processing and shallow machine learning is limited, and edge detection, texture analysis or local description is often used. These low-level features are difficult to fully capture the subtle structural differences of firearms, making it difficult to distinguish between models with similar appearances. Existing identification methods generally lack robustness and perform poorly on blurred, low-resolution or noisy images, leading to misjudgments in actual applications. Most existing methods can only output firearm category labels and cannot be associated with metadata such as firearm specifications, manufacturer information and use scenarios, which is not conducive to subsequent evidence management or multi-dimensional analysis. In the application scenario of large-scale firearm databases, traditional matching and retrieval methods are inefficient and cannot meet the real-time and scalability requirements.
[0021] For example, during security checks at a border crossing, a surveillance device captured an image of a replica gun that closely resembled a standard-issue handgun. Due to the skewed shooting angle and insufficient light, traditional methods based on edge features and texture analysis failed to effectively distinguish between the two, misjudging the replica gun as a standard-issue handgun. This error not only delayed the security process but also caused controversy among law enforcement agencies regarding subsequent handling, affecting work efficiency and accuracy. This case highlights the problems of existing technology, such as insufficient recognition accuracy in complex environments, lack of robustness, and insufficient information support.
[0022] To address this, refer to Figure 1 a deep learning-based firearm type identification method, comprising: S100: Based on a deep convolutional neural network model and according to standard firearm images, extract a plurality of feature vector sets, each feature vector set including a corresponding firearm type identifier and detailed information data; and construct a firearm feature database based on the plurality of feature vector sets; S200: Collect the image of the firearm to be identified and perform preprocessing operations on the image of the firearm to be identified; S300: Input the preprocessed image of the firearm to be identified into the deep convolutional neural network model, the deep convolutional neural network model uses the ResNet-50 architecture and performs transfer learning on the ImageNet dataset, retains all the convolutional layer structures of the ImageNet pre-trained model, and freezes the weight parameters of the convolutional layers; remove the fully connected layer and classification layer of the pre-trained deep convolutional neural network model, and replace it with an output layer adapted to the firearm identification task, outputting a fixed-dimensional high-dimensional feature vector; S400: Calculate the cosine similarity of the high-dimensional feature vector with all feature vector sets in the firearm feature database to generate a list of similarity values; sort the list of similarity values in descending order and select the top four firearm type identifiers with the highest similarity values; S500: Extract detailed information data corresponding to the first four gun type identifiers from the gun feature database based on the first four gun type identifiers.
[0023] In particular, the implementation process of the present application includes two parts of offline construction and online identification: in S100, a plurality of standard firearm images covering multiple perspectives, multiple illuminations and different resolutions are collected, and a unique firearm type identifier (type_id) and structured detailed information data (such as firearm name, model, caliber, length, manufacturer, applicable scene label, reference image file identification, collection time and quality score) are labeled for each image. These images are batched into the ResNet-50 convolutional backbone network pre-trained by ImageNet (remove the original classification head) for forward inference, the featuremap of the last convolutional block is obtained by global average pooling (GAP) to obtain a fixed dimension vector (typically 2048 dimensions or 512 dimensions after projection dimension reduction), L2 normalization is performed on the vector for subsequent cosine similarity calculation, and the normalized vector, type_id and detailed metadata are packaged together as a structured "feature vector set" entry, and written into the feature vector data table according to the version number; the storage layer is recommended to support vector type or vector index (such as pgvector / FAISS / Milvus), and ANN index (HNSW, IVF+PQ, etc.) is established for the vector, and the index construction parameters and entry QA scores are saved for subsequent management and rollback; in S200, the image to be identified is subjected to strict preprocessing: after recording the acquisition device and time information, short side equal scaling and necessary padding or center cropping are performed to adapt to the model input size (such as 224x224 or higher), optional Gaussian denoising, CLAHE contrast enhancement and gamma correction are applied, and the image is converted to the channel order and floating point format normalized according to the ImageNet mean / variance expected by the model; in S300, the preprocessed image is input into the homologous ResNet-50 model (the same backbone and normalization process as offline construction is retained during deployment, and only forward propagation is performed online), the GAP output is taken and L2 normalized to obtain the query vector, and the projection head can be selected for dimension matching or quantization during deployment to optimize the delay and storage; in S400, vector retrieval is performed from the firearm feature database or vector index, if vector DB is used, top-k query is called to directly obtain the candidate and its similarity score, the retrieval results are sorted in descending order of similarity and aggregated and de-duplicated by type_id (if there are multiple references for the same type_id, the highest score entry is retained), when the number of candidates ≥4, the four type_ids with the highest similarity are returned, otherwise all valid candidates are returned and the candidate shortage is marked; a similarity threshold should be set to identify low confidence situations (for example, the verified experience threshold interval needs to be determined before deployment), and when approaching the boundary or the scores are extremely close, the heat map / Grad-CAM visualization is also returned to assist manual review;In S500, the corresponding detailed information data, reference image identifier, QA analysis index / model version number are extracted from the database with the selected four type_ids as the query key, and the results are returned to the caller in a structured format (such as containing query_id, top_candidates array, score and metadata of each candidate, and low confidence or manual review label), while recording the unforgeable audit log (including input image hash, timestamp, model and index version, returned candidate and score) to ensure traceability.
[0024] The working principle and process of the present application are as follows: in S100 stage, standard firearm images covering different angles, lighting and resolutions are collected, and a unique firearm type identifier and detailed information data (such as firearm name, model, caliber, length, manufacturer and application scenario, etc.) are labeled for each image. These images are input into the ResNet-50 convolutional neural network pre-trained by ImageNet for forward propagation, multi-level feature response maps are extracted by convolutional layers, fixed-dimensional high-dimensional feature vectors are generated by global average pooling (GAP), and L2 normalization is performed to ensure the stability of subsequent cosine similarity calculation. The normalized feature vectors, corresponding firearm type identifiers and detailed information data are packaged into a structured feature vector set and stored in a feature database, and a vector index can be established to speed up retrieval. In S200 stage, the pre-processed image is pre-processed, including size adjustment, necessary cropping, denoising, high-contrast enhancement and normalization operations to ensure that the data distribution of the input model is consistent with the training set. In S300 stage, the pre-processed image is input into the ResNet-50 model with the same architecture, features are extracted by convolution, GAP pooling and L2 normalization are performed, and the high-dimensional feature vector of the image to be identified is obtained. In S400 stage, the cosine similarity between the feature vector and all stored feature vector sets in the database is calculated, the similarity score is obtained by vector inner product divided by Euclidean norm of each vector, and a descending list is generated, and the top four firearm type identifiers with the highest similarity are selected as candidates. In S500 stage, the corresponding detailed information data including name, model, parameters and reference image identifier are extracted from the database with the selected four candidate identifiers as the query condition, and the results are returned in a structured form, which can be combined with a visual method to assist manual review. The whole process realizes efficient and accurate firearm type identification through offline training and vector retrieval.
[0025] As a preferred embodiment, the scheme of the present application is implemented as follows: for example, the law enforcement department needs to identify the type of a handgun found on site, a firearm feature database is established in S100 stage: standard images of different models of Glock, Beretta, SIG Sauer and Colt etc. are collected, each image is labeled with a unique firearm type identifier and detailed information data such as model, caliber, barrel length, manufacturer, application scenario (civilian, police), reference image file identifier and collection time, and a high-dimensional feature vector is extracted through ResNet-50 convolutional neural network, and after global average pooling and L2 normalization, it is stored in the database, and an ANN vector index is established to speed up the retrieval. In S200 stage, law enforcement personnel use a portable camera to take pictures of the on-site handgun, adjust the image size to 224x224, center crop, Gaussian denoising and contrast enhancement, and normalize the channel order to match the model input. In S300 stage, the processed image is input into the pre-trained and fine-tuned ResNet-50 model, the convolutional layer extracts spatial features, outputs a 2048-dimensional high-dimensional feature vector through global average pooling, and performs L2 normalization. In S400 stage, the feature vector is calculated with all feature vectors in the database to calculate the cosine similarity, for example, the similarity scores with Glock 19, Glock 17, Beretta 92F and SIG P226 are 0.93, 0.88, 0.72 and 0.70 respectively, a similarity list is generated and sorted in descending order, and the top four candidate types are selected. In S500 stage, the four type identifiers are used as query keys to extract corresponding detailed information from the database, such as "Glock 19, 9mm caliber, barrel length 102mm, Austrian Glock company, police law enforcement application, reference image number IMG_00123", and the result is returned, realizing the complete process from image collection to high-precision firearm type identification.
[0026] Through the above technical scheme, the present application can extract firearm features under limited firearm image samples by using ResNet-50 deep convolutional backbone network combined with ImageNet transfer learning, enabling high-robustness identification of firearm images under complex background, different angles, different lighting and partial occlusion; through global average pooling to generate a fixed-dimensional high-dimensional feature vector and L2 normalization, the subsequent cosine similarity calculation is stable and reliable, avoiding errors caused by feature amplitude differences, while supporting fast vector retrieval and candidate sorting, improving identification speed and accuracy; detailed information data including firearm model, caliber, barrel length, manufacturer, application scenario and reference image identifier are stored based on the feature vector database, so that not only type identification can be realized, but also structured and traceable detailed information can be output simultaneously, meeting the requirements of law enforcement, security and police applications for information integrity and traceability.
[0027] The application further proposes that when a plurality of feature vector sets are extracted based on a deep convolutional neural network model and according to standard firearm images, the method comprises: A plurality of standard firearm images are obtained, each corresponding to a known firearm type; forward propagation calculation is performed on each standard firearm image, and a multi-level feature response map is extracted through a convolutional layer of the deep convolutional neural network model; the multi-level feature response map is input into a global average pooling layer for compression processing of the spatial dimension, thereby generating a high-dimensional feature vector of fixed dimension; a unique reference image file identifier is generated for each standard firearm image, the reference image file identifier being composed of a firearm type identifier and an image serial number; the high-dimensional feature vector is data-bound with the corresponding firearm type identifier; and the firearm name, specification parameter, manufacturer information, applicable scene classification label, and reference image file identifier are taken as detailed information data, and are collectively formed into a structured feature vector set with the high-dimensional feature vector.
[0028] Specifically, in the process of extracting a firearm feature vector set based on a deep convolutional neural network model, standard firearm images covering different angles, lighting conditions, and resolutions are collected, each image corresponding to a known firearm type and being strictly labeled, including the firearm name, model, caliber, barrel length, weight, manufacturer, and applicable scene classification label (such as civilian, law enforcement, and military), to ensure that the training data is comprehensive and diverse; each image is input into the deep convolutional neural network model, and a multi-level feature response map is extracted step by step using the convolutional layer, including low-level edge texture, high-level structural shape, and complex component information, to capture rich features of the firearm appearance; the multi-level feature response map extracted is input into a global average pooling (GAP) layer for compression of the spatial dimension, thereby generating a high-dimensional feature vector of fixed dimension; to ensure the unique traceability of each image, a reference image file identifier is generated for it, the identifier being composed of a firearm type identifier and an image serial number, and the high-dimensional feature vector is bound with the corresponding firearm type identifier, while the aforementioned detailed information data is integrated with the feature vector to form a structured feature vector set entry; each set contains not only a high-dimensional feature vector that can be used to measure similarity, but also detailed metadata information that can be queried.
[0029] As a preferred embodiment, the scheme of the present application is implemented as follows: for example, a public security department needs to establish a handgun identification database. In the standard handgun image acquisition stage, high-definition images of multiple models of handguns including Glock 17, Glock 19, Beretta 92F, Colt 1911, etc. are collected, each image covers different shooting angles (front, side, top view), different lighting conditions and different backgrounds, each image corresponds to a known handgun type and is labeled with detailed information such as handgun name, model, caliber, barrel length, weight, manufacturer and application scenario (civilian, law enforcement, military), ensuring that the training data is rich and diverse; input each image into a convolutional neural network, and the convolutional layer extracts multi-level feature response maps of low-level texture, high-level structure and component combination step by step, then input these multi-level feature response maps into a global average pooling layer to compress the spatial dimension and generate a fixed-dimension 2048-dimensional high-dimensional feature vector; generate a unique reference image file identifier for each standard image, which is composed of a handgun type identifier and an image serial number, for example, "GLOCK19_IMG_001", and bind the high-dimensional feature vector with the corresponding handgun type identifier, and integrate the handgun name, specification parameters, manufacturer information, application scenario classification label and reference image identifier to form a complete structured feature vector set entry, and finally build a handgun feature database that can be used for subsequent fast retrieval and accurate identification.
[0030] Through the above technical scheme, the present application can accurately capture the low-level texture features, high-level structure information and component combination patterns of the handgun through multi-level feature extraction of the standard handgun image by the deep convolutional neural network model, thereby realizing robust identification of handguns of different models, different shooting angles and lighting conditions; the global average pooling is used to generate a fixed-dimension high-dimensional feature vector, which can ensure the stability and consistency of subsequent cosine similarity calculation and avoid matching errors caused by numerical amplitude differences; by generating a unique reference image file identifier for each standard image and binding it with the handgun type identifier, and integrating the detailed information data such as the handgun name, specification parameters, manufacturer information and application scenario classification label into the structured feature vector set, the unified management of feature representation and metadata information is realized.
[0031] The present application further proposes that when constructing a handgun feature database according to a plurality of feature vector sets, it includes: A feature vector data table is created in the handgun feature database, which is provided with a feature vector field, a handgun type identifier primary key field and a detailed information field; the high-dimensional feature vector in each feature vector set is stored in the feature vector field in the form of a floating-point number array; the handgun type identifier in the feature vector set is stored in the handgun type identifier primary key field as a unique identifier value; and the detailed information data in the feature vector set is stored in the detailed information field in a structured data format.
[0032] Specifically, in the process of constructing a firearm feature database according to a plurality of feature vector sets, a dedicated feature vector data table is created in the database, which includes at least three core fields: a feature vector field, a firearm type identifier primary key field, and a detailed information field. The feature vector field is used to store the high-dimensional feature vector of each firearm image extracted by the deep convolutional neural network and subjected to global average pooling, usually recording the value of each dimension in the form of a floating-point array to ensure the continuity and calculability of the features. The firearm type identifier primary key field is used to store the corresponding unique identifier, which is accurately bound to each feature vector and ensures that there are no duplicate entries in the database, facilitating subsequent fast retrieval and index construction. The detailed information field stores the metadata information corresponding to each feature vector in a structured data format, including the firearm name, model, caliber, barrel length, weight, manufacturer, applicable scene classification label, and reference image file identifier, thereby realizing the close association of feature data and actual application information. In the data writing process, all feature vector sets are written into the database one by one through batch import or incremental update, ensuring data consistency and integrity.
[0033] As a preferred embodiment, the scheme of the present application is implemented as follows: For example, a certain public security department needs to construct a handgun recognition database, a dedicated data table is created in the firearm feature database, which has a feature vector field, a firearm type identifier primary key field, and a detailed information field. For the standard image of Glock 17, the 2048-dimensional high-dimensional feature vector extracted by the convolutional neural network and subjected to global average pooling is stored in the feature vector field in the form of a floating-point array, and "GLOCK17" is stored as the firearm type identifier in the primary key field to ensure uniqueness and indexability. The detailed information field stores the complete metadata of this model of firearm, including the name "GLOCK17", the caliber "9mm", the barrel length "114mm", the weight "625g", the manufacturer "Austrian Glock Company", the applicable scene label "law enforcement / military", and the reference image identifier "GLOCK17_IMG_001", ensuring that each feature vector can be used not only for similarity calculation but also for providing complete background information. Similarly, for Beretta 92F, Colt 1911, and other models of handguns, their high-dimensional feature vectors, unique firearm type identifiers, and detailed information data are stored in the database in a structured form. In this way, a complete firearm feature database covering multiple models, multiple angles, and multiple scenes is constructed.
[0034] By the technical solution, the application creates a special feature vector data table in the database, and sets a feature vector field, a firearm type identifier primary key field and a detailed information field, realizes systematic storage and unified management of high-dimensional feature vectors, unique identifiers and detailed information, so that each data entry can be used for fast retrieval and provide complete firearm meta-information support; the high-dimensional feature vector is stored in the feature vector field in the form of a floating-point number array, which guarantees the accuracy and stability of feature calculation and supports multiple vector similarity calculation methods such as cosine similarity and Euclidean distance, thereby improving the accuracy of firearm type matching; by storing the firearm type identifier as a unique primary key field, not only is the database consistency improved, but also the index is easily established. The application further proposes that when performing a preprocessing operation on the to-be-identified firearm image, the preprocessing operation comprises: The preprocessing operation comprises input size adjustment, Gaussian filtering, contrast adjustment, data format normalization and channel order adjustment.
[0035] Specifically, when performing a preprocessing operation on the to-be-identified firearm image, the input size of the original image is adjusted, the image is uniformly scaled or cropped to a fixed size required by a deep convolutional neural network model, such as 224x224 pixels, to ensure the consistency of subsequent network input and avoid feature extraction deviation caused by size difference; Gaussian filtering is performed, random noise, light spots or background interference generated in the shooting process are suppressed by applying a Gaussian smoothing kernel to the image, thereby improving the stability of feature extraction; a contrast adjustment operation is performed, the brightness and darkness difference of the image is enhanced, so that the firearm contour, edge and component structure are clearer, thereby improving the perception ability of the convolutional network to local details; the data format of the image is normalized, the pixel value is mapped to the range of 0 to 1 or -1 to 1, so as to eliminate the influence of different light conditions or sensor acquisition differences on model input and speed up network convergence; the channel order of the image is adjusted according to the model requirements, for example, the RGB channel arrangement is adjusted to BGR or other specific order, to ensure that the input data is consistent with the channel configuration of the pre-trained network weight, and to avoid feature extraction errors caused by channel mismatch.
[0036] As a preferred embodiment, the scheme of the present application is implemented as follows: for example, a public security department needs to identify a hand gun image taken at a law enforcement scene, the original image has a resolution of 4000x3000 pixels, the lighting condition is uneven and there are background clutter. The image is scaled and cropped to 224x224 pixels to meet the input requirements of the model and ensure that the network can process data of uniform size; Gaussian filtering is applied to the image, a Gaussian kernel with a standard deviation of 1.0 is used to smooth the image, remove random noise and light spots that occur during shooting, and make the gun outline smoother and clearer; contrast enhancement is performed to increase the difference between the bright and dark parts of the image, making the details such as the barrel, grip, and safety mechanism more prominent, which helps the deep convolutional neural network to capture key features; the pixel values are normalized to the range of 0 to 1 to eliminate the effects of lighting intensity differences and speed up network inference speed and stability; according to the model training requirements, the original RGB channel is adjusted to BGR order to make the input channel consistent with the weight of the transfer learning model, avoiding feature extraction errors.
[0037] Through the above technical scheme, the present application unifies images of different sources and different resolutions into a fixed size required by the network through input size adjustment, so that the convolutional neural network can stably process each image, avoiding feature extraction deviation caused by size differences, thereby improving the input consistency and recognition accuracy of the model; Gaussian filtering can remove shooting noise, light spots and background interference, making the gun outline and key structure clearer, thereby enhancing the perception ability of the deep convolutional neural network to local details and improving the reliability of feature extraction; contrast adjustment enhances the bright-dark difference to make the edges, grips, barrels and other key components of the gun more prominent, which helps the model to capture significant features and reduces the probability of misidentification; data format normalization can eliminate the effects of differences in different lighting conditions and image acquisition devices, speed up the convergence speed of network training and inference, and improve the stability of the model in multiple scenarios; channel order adjustment ensures that the input image channel arrangement is consistent with the pre-trained model weight, avoiding feature extraction errors caused by channel mismatch.
[0038] The present application further proposes that when the deep convolutional neural network model adopts the ResNet-50 architecture and performs transfer learning on the ImageNet dataset, it includes: During transfer learning, the learnable parameters of the batch normalization layer in the deep convolutional neural network model are fine-tuned; the preprocessed gun image to be identified is input into the adjusted deep convolutional neural network model, and the final feature map is generated through the convolutional layer; a global average pooling operation is performed on the final feature map to compress the spatial dimension to 1x1 to obtain a high-dimensional feature vector; and L2 norm normalization processing is performed on the high-dimensional feature vector.
[0039] Specifically, taking ResNet-50 pre-trained on ImageNet as the backbone, keeping its entire convolutional and residual block structure to make full use of large-scale visual features (including conv1~layer4 and the corresponding BatchNorm layers), most of the weight parameters of these convolutional layers are frozen at initialization to preserve the general visual representation (usually conv1~layer3 are frozen, only a few layers after layer4 are unfrozen depending on the data volume), but the learning parameters (scale γ and shift β) of all BatchNorm layers and the update strategy of their runtime mean / variance are set to a fine-tunable state to adapt to the statistical differences of gun images; the original pre-trained model's top fully connected classifier and softmax layer are removed and replaced with an output structure suitable for retrieval / metric learning tasks, such as connecting a linear projection layer (optional Bottleneck: 2048→512 fully connected + BatchNorm + ReLU + Dropout) after GAP to reduce dimensionality and enhance discriminability, or directly using the GAP output as a feature vector for retrieval; during the transfer learning training phase, use a fine-tuning strategy (such as setting a larger learning rate of 1e-4~1e-3 for layer4 and the projection head, and a learning rate of 0 or a very small value for frozen layers), the optimizer can be AdamW or SGD+momentum, and use moderate weight decay and learning rate scheduling (Cosine / Step); the training target can be either classification cross-entropy (when there are sufficient labeled categories for supervised pre-training) or parallel use of metric learning loss (Triplet, ArcFace / CosFace) to directly improve the separability of the feature space; for forward inference, the preprocessed image to be recognized is fed into the adjusted model, and the final feature map (usually shaped as CxHxW, C=2048) is generated by consecutive convolutional layers; apply global average pooling (GAP) to the feature map to compress the spatial dimension to 1x1, thereby obtaining a fixed-dimension high-dimensional feature vector (default 2048 dimensions, or 512 dimensions after projection); to ensure the numerical stability and scale invariance of similarity calculation, perform L2 norm normalization on the vector, and consider using half-precision (FP16) or quantization during deployment to reduce latency and memory usage; at the same time, correctly manage the training state of BatchNorm when switching between training / validation (in small sample scenarios, the mean and variance can be frozen or use cumulative statistics), and record the model version, projection head configuration, and normalization process to ensure consistency between online retrieval and offline feature library construction, thereby preserving the general feature capability of ImageNet, and converting ResNet-50 output into a high-quality feature representation with strong discriminability, fixed dimension, and suitable for cosine similarity retrieval through fine-tuning BN, projection head design, and appropriate loss function.
[0040] As a preferred embodiment, the scheme of the present application is implemented as follows: for example, a hand gun image obtained from a monitoring camera needs to be recognized, first, the image is preprocessed and input into an adjusted ResNet-50 model, the model retains all the convolutional layers pre-trained on ImageNet to extract general visual features, and the weights of these convolutional layers are frozen to prevent overfitting on the limited gun dataset; the original fully connected classification layer is removed and replaced with an output layer adapted to the gun recognition task, such as a linear projection layer for generating a 2048-dimensional feature vector. During transfer learning, the learnable parameters of the batch normalization layer are fine-tuned to enable the model to adapt to the differences in illumination, background, and texture of gun images. After the image is processed by the convolutional layer, the final feature map (C×H×W, e.g. 2048×7×7) is obtained, the spatial dimension is compressed to 1×1 by global average pooling, and a fixed-dimensional high-dimensional feature vector is obtained, then the feature vector is L2 norm normalized to make the vector length 1, so that the numerical value is stable and the scale is consistent when comparing with the existing features in the gun feature database, and the cosine similarity is compared.
[0041] Through the above technical scheme, by using the ResNet-50 architecture and performing transfer learning on the ImageNet dataset, the present application retains the entire convolutional layer structure of the pre-trained model and freezes its weight parameters, and uses the general visual features such as edges, textures and shape information learned from large-scale image data, so that high recognition performance can still be maintained under limited gun image data; the original fully connected layer and classification layer are removed and replaced with an output layer adapted to gun recognition, so that the model can generate a high-dimensional feature vector suitable for a specific task, enhancing the discrimination ability of the gun type difference; during transfer learning, the learnable parameters of the batch normalization layer are fine-tuned to cope with the changes in gun image illumination, shooting angle and background complexity, improving the stability and robustness of feature extraction; the global average pooling operation compresses the spatial dimension to 1×1 to obtain a fixed-dimensional high-dimensional feature vector, making the feature representation more compact and easy to calculate the similarity with the stored features in the database; L2 norm normalization processing ensures that the feature vector has consistent scale when calculating the cosine similarity, avoiding numerical bias, improving the accuracy and efficiency of gun type recognition, and reducing the computational complexity.
[0042] The present application further proposes that when performing L2 norm normalization processing on the high-dimensional feature vector, it includes: performing a square operation on each element in the high-dimensional feature vector, accumulating and summing all square operation results, performing a square root operation on the accumulated sum result to obtain an L2 norm value, and dividing each element in the high-dimensional feature vector by the L2 norm value to obtain a normalized high-dimensional feature vector.
[0043] Specifically, in the process of performing L2 norm normalization processing on the high-dimensional feature vector, each element in the vector is squared one by one, that is, each feature value is multiplied by itself to obtain the square value of the element; the square values of all elements are accumulated and summed to obtain the sum of squares of the entire high-dimensional feature vector, which is a key step for calculating the length of the vector; the result of the accumulation and summation is square rooted to obtain the L2 norm of the vector, that is, the length of the vector in the Euclidean space; each element in the original high-dimensional feature vector is divided by the L2 norm value to realize the normalization processing of the vector, so that the length of the normalized feature vector is 1, while the relative proportions between the elements remain unchanged.
[0044] The application further proposes to calculate the cosine similarity of the high-dimensional feature vector and all feature vector sets in the firearm feature database, and generate a similarity value list, including: The feature vector field is read from the feature vector data table of the firearm feature database; the high-dimensional feature vector in each feature vector record field is extracted; the vector inner product of the high-dimensional feature vector of the to-be-identified firearm image and the high-dimensional feature vector in the firearm feature database is calculated, and the Euclidean norm of the high-dimensional feature vector of the to-be-identified firearm image and the high-dimensional feature vector in the firearm feature database is calculated; and the similarity value is calculated according to the vector inner product and the Euclidean norm; the similarity value and the corresponding firearm type identifier are mapped; and all mapping relationships are constructed into a similarity value list containing the corresponding relationship between the firearm type identifier and the similarity value according to the calculation order.
[0045] Specifically, in calculating the cosine similarity of the high-dimensional feature vector of the to-be-identified firearm image and all feature vector sets in the firearm feature database, the feature vector field in the feature vector data table of the firearm feature database is read, and each stored high-dimensional feature vector is extracted one by one; the vector inner product of the high-dimensional feature vector of the to-be-identified image and each feature vector in the database is calculated, which is used to quantify the degree of similarity of the two vectors in the feature space; the Euclidean norm of the to-be-identified vector and the database vector is calculated for standardizing the vector length; the vector inner product is divided by the product of the Euclidean norms of the two vectors to obtain the cosine similarity value, which reflects the similarity degree of the to-be-identified image and each firearm type feature vector in the database; each similarity value and the corresponding firearm type identifier are mapped to ensure that the specific firearm type corresponding to each similarity can be accurately tracked; all mapping relationships are summarized according to the calculation order to form a similarity value list containing the corresponding relationship between the firearm type identifier and the similarity value.
[0046] As a preferred embodiment, the scheme of the present application is implemented as follows: for example, the high-dimensional feature vectors of different types of guns such as AK-47, M16, G36, etc. are stored in a gun feature database, each feature vector consisting of a 512-dimensional floating-point array. When an image of an M16 gun to be identified is collected, it is first input into a deep convolutional neural network model to extract the corresponding 512-dimensional high-dimensional feature vector, and then the feature vector field of each gun type is read from the database, for example, the vector of AK-47 is V1, the vector of M16 is V2, and the vector of G36 is V3. Then, the vector inner product of the feature vector of the image to be identified and V1, V2, and V3 is calculated respectively, and the respective Euclidean norms are calculated, and then the inner product is divided by the product of the norms to obtain the similarity values by the cosine formula, such as 0.76 for V1, 0.92 for V2, and 0.68 for V3. Then, a mapping relationship is established between each similarity value and the corresponding gun type identifier, i.e. 0.76 corresponds to AK-47, 0.92 corresponds to M16, and 0.68 corresponds to G36, and finally all the mapping relationships are summarized in the order of calculation to generate a list containing the gun type identifier and the similarity value, facilitating subsequent selection of the M16 with the highest similarity as the identification result.
[0047] Through the above technical scheme, the present application reads the feature vector field of each record from the feature vector data table of the gun feature database, and extracts the corresponding high-dimensional feature vector, such as a 512-dimensional floating-point number. The high-dimensional feature vector of the image to be identified is calculated with the vector inner product of each feature vector in the database, and the respective Euclidean norms are calculated, and then the inner product is divided by the product of the norms to obtain the similarity values by the cosine similarity formula, which reflect the similarity of the image to be identified and each gun type in the feature space. Each similarity value is mapped with the corresponding gun type identifier, and all the mapping relationships are summarized in the order of calculation to form a complete list of similarity values.
[0048] The present application further proposes that when the top four gun type identifiers with the highest similarity values are selected according to the descending order of the similarity value list, it includes: The quicksort algorithm is performed on the similarity values in the similarity value list, and an ordered list is generated in the order of similarity values from high to low; when the number of gun type identifiers in the ordered list is greater than or equal to four, the top four entries corresponding to the gun type identifiers are selected; when the number of gun type identifiers in the ordered list is less than four, all valid entries corresponding to the gun type identifiers are selected.
[0049] Specifically, in the process of descending order sorting according to the similarity value list and selecting the top four gun type identifiers with the highest similarity values, each entry in the generated similarity value list is sorted quickly, which contains the similarity values corresponding to each gun feature vector in the database and the associated gun type identifiers of the gun image to be identified. The quicksort algorithm divides the list into two parts by selecting a reference value, recursively sorts each part, and arranges all similarity values from high to low in an average time complexity of O, ensuring that the most similar feature vectors are at the front of the list. After sorting, check the number of gun type identifiers in the ordered list. If the number is greater than or equal to four, directly select the top four gun type identifiers as the candidate result to ensure that the most relevant identification information is obtained; if the number is less than four, select all valid gun type identifiers to ensure that even if the database records are limited, the complete candidate set can still be output.
[0050] As a preferred embodiment, the scheme of the present application is implemented as follows: in actual application, after completing the cosine similarity calculation of the gun image and all feature vectors in the database, a list containing gun type identifiers and their corresponding similarity values is generated, such as the list containing entries "pistol_001: 0.92, submachine gun_003: 0.88, rifle_007: 0.85, pistol_002: 0.83, and shotgun_005: 0.79". The quicksort algorithm is executed on the list to arrange the similarity values from high to low, obtaining the ordered list "pistol_001, submachine gun_003, rifle_007, pistol_002, and shotgun_005". Then check the length of the list. If the number of gun type identifiers is greater than or equal to four, such as five entries in this example, select the top four gun type identifiers "pistol_001, submachine gun_003, rifle_007, and pistol_002" as the candidate identification result; if the length of the list is less than four, such as only three records, directly select all valid entries to ensure that no possible matches are missed. Through this process, the top four types most similar to the gun image to be identified can be quickly screened, providing a reliable basis for subsequent extraction of detailed information from the database and generation of the final identification result.
[0051] Through the above technical scheme, the quicksort algorithm is executed on the similarity values in the list, arranging all entries from high to low similarity to form an ordered list. After sorting, the length of the ordered list is determined. If the number of gun type identifiers is greater than or equal to four, select the top four identifiers with the highest similarity as the candidate identification result; if the length of the list is less than four, directly select all valid entries to ensure that all possible matching information is retained. This processing method not only ensures the high accuracy and reliability of the identification result, but also reduces data redundancy.
[0052] The application further proposes that when the first four gun type identifiers are extracted from the gun feature database, the detailed information data corresponding to the gun type identifiers is included: The selected gun type identifier is used as a query condition to perform a database query operation; and the corresponding detailed information data is extracted from the gun feature database.
[0053] Specifically, when the first four gun type identifiers with the highest similarity have been selected through similarity sorting, the four identifiers will be used as query conditions in turn to initiate a structured query request to the gun feature database. After receiving the query request, the database will match the record entries corresponding to each identifier in the feature vector data table, and extract the detailed information data bound thereto, including gun name, model, caliber, manufacturer, application scenario classification label, and reference image file identifier, etc. structured information. For example, when the first four selected identifiers are "rifle_003, pistol_001, submachine gun_005, and pistol_002", the database is searched in turn to read out the detailed information fields of each record and integrate them into a unified query result set.
[0054] Through the above technical solution, the application realizes direct mapping from abstract high-dimensional feature vectors to specific gun attributes through the query and extraction operation, so that the recognition result is not limited to type judgment, and the accuracy and information integrity of the recognition are improved.
[0055] Based on the above another preferred manner, referring to Figure 2 The embodiment provides a gun type recognition system based on deep learning, which is used for applying the above-mentioned gun type recognition method based on deep learning, and includes: A database construction unit configured to extract a plurality of feature vector sets based on a deep convolutional neural network model and according to standard gun images, each feature vector set including a corresponding gun type identifier and detailed information data; and construct a gun feature database according to the plurality of feature vector sets; An image acquisition unit configured to acquire a to-be-recognized gun image and perform a preprocessing operation on the to-be-recognized gun image; A first processing unit configured to input the preprocessed to-be-recognized gun image into a deep convolutional neural network model, the deep convolutional neural network model adopts a ResNet-50 architecture and performs transfer learning on an ImageNet dataset, retains all convolutional layer structures of an ImageNet pre-trained model, and freezes weight parameters of the convolutional layers; removes a fully connected layer and a classification layer of the pre-trained deep convolutional neural network model, and replaces them with an output layer adapted to a gun recognition task to output a high-dimensional feature vector with a fixed dimension; a second processing unit configured to calculate cosine similarities between the high-dimensional feature vector and all sets of feature vectors in the firearm feature database, and generate a list of similarity values; sort the list of similarity values in descending order, and select the top four firearm type identifiers with the highest similarity values; a data extraction unit configured to extract detailed information data corresponding to the top four firearm type identifiers from the firearm feature database based on the top four firearm type identifiers.
[0056] In summary, based on the deep convolutional neural network (ResNet-50 transfer learning), multi-level high-dimensional features are extracted, and global average pooling and L2 normalization are combined to generate feature vectors with discriminative power and fixed length, thereby improving the consistency and discrimination of firearm type representation, and improving the recognition accuracy and robustness. By using the transfer learning and freezing the convolutional layer, and fine-tuning the batch normalization parameter strategy, the general visual features of the large-scale ImageNet pre-trained model can be utilized, the dependence on a large amount of labeled firearm data is reduced, and efficient fine-tuning can be performed for the firearm recognition task, thereby shortening the training time and reducing the labeling cost. By packaging the high-dimensional features, type identifiers and structured detailed information of standard firearm images into a feature vector set and establishing a feature vector data table, the synchronous management of features and metadata is facilitated, and the query efficiency and scalability of the database are improved. Using the L2 normalization based on cosine similarity retrieval method, the similarity calculation can be simplified to vector inner product and norm operation, the calculation overhead is controllable, and efficient similarity retrieval (including combination with approximate nearest neighbor index) can be realized in the actual system, thereby meeting the real-time or near real-time recognition requirements. The output of the top four candidate firearm types with the highest similarity and the return of the corresponding structured detailed information (name, specification, manufacturer, application scenario and reference image identifier) not only improves the interpretability of the recognition result, but also provides rich context information for manual review, evidence chain preservation or subsequent decision-making, thereby reducing the risk of misjudgment. The image preprocessing process (size adjustment, Gaussian filtering, contrast adjustment, format normalization and channel order adjustment) enhances the adaptability of the model to input image illumination, noise and scale changes, and improves the stability and generalization ability of the system in complex field environments.
[0057] Those skilled in the art will appreciate that embodiments of the application can be provided as methods, systems or computer program products. Accordingly, the application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the application can be embodied in the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage media, etc.) having computer usable program code embodied therein.
[0058] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart or flowsheet block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks
[0059] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart or flowsheet block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks
[0060] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart or flowsheet block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks
[0061] Finally, it should be noted that the above-mentioned embodiments are merely intended for describing the technical solutions of the present application, but not for limiting it. Although the present application is described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.
Claims
1. A deep learning-based method for firearm type identification, characterized in that, include: Based on a deep convolutional neural network model, several feature vector sets are extracted from standard firearm images. Each feature vector set includes a corresponding firearm type identifier and detailed information data. A firearm feature database is constructed based on the several feature vector sets. Acquire images of the firearms to be identified and perform preprocessing operations on the images; The preprocessed images of firearms to be identified are input into a deep convolutional neural network model. The deep convolutional neural network model adopts the ResNet-50 architecture and performs transfer learning on the ImageNet dataset. All convolutional layer structures of the ImageNet pre-trained model are retained, and the weight parameters of the convolutional layers are frozen. The fully connected layers and classification layers of the pre-trained deep convolutional neural network model are removed and replaced with output layers adapted to the firearms recognition task, outputting high-dimensional feature vectors with fixed dimensions. Calculate the cosine similarity between the high-dimensional feature vector and all feature vector sets in the firearm feature database to generate a similarity value list; sort the similarity value list in descending order and select the top four firearm type identifiers with the highest similarity values; Based on the first four firearm type identifiers, the detailed information data corresponding to the first four firearm type identifiers is extracted from the firearm feature database.
2. The firearm type recognition method based on deep learning according to claim 1, characterized in that, When extracting several feature vector sets based on a deep convolutional neural network model and standard firearm images, the following are included: Acquire several standard firearm images, each corresponding to a known firearm type; perform forward propagation computation on each standard firearm image, extracting multi-level feature response maps through the convolutional layers of a deep convolutional neural network model; input the multi-level feature response maps into a global average pooling layer to compress the spatial dimension, generating a fixed-dimensional high-dimensional feature vector; generate a unique reference image file identifier for each standard firearm image, the reference image file identifier being a combination of a firearm type identifier and an image sequence number; data-bind the high-dimensional feature vector with the corresponding firearm type identifier; and combine the firearm name, specifications, manufacturer information, applicable scenario classification labels, and reference image file identifier as detailed information data with the high-dimensional feature vector to form a structured feature vector set.
3. The firearm type recognition method based on deep learning according to claim 2, characterized in that, When constructing a firearms feature database based on several sets of feature vectors, the following is included: A feature vector data table is created in the firearms feature database. The feature vector data table has a feature vector field, a firearms type identifier primary key field, and a detailed information field. The high-dimensional feature vectors in each feature vector set are stored in the feature vector field as floating-point arrays. The firearms type identifier in the feature vector set is stored as a unique identifier in the firearms type identifier primary key field. The detailed information data in the feature vector set is stored in the detailed information field in a structured data format.
4. The firearm type recognition method based on deep learning according to claim 3, characterized in that, When performing preprocessing operations on the image of the firearm to be identified, the following are included: The preprocessing includes input size adjustment, Gaussian filtering, contrast adjustment, data format normalization, and channel order adjustment.
5. The firearm type recognition method based on deep learning according to claim 4, characterized in that, When using a deep convolutional neural network model with the ResNet-50 architecture and performing transfer learning on the ImageNet dataset, the following are included: During the transfer learning process, the learnable parameters of the batch normalization layer in the deep convolutional neural network model are fine-tuned; the preprocessed image of the firearm to be identified is input into the adjusted deep convolutional neural network model, and the final feature map is generated through the convolutional layer; the final feature map is subjected to global average pooling to compress the spatial dimension to 1×1 to obtain a high-dimensional feature vector; and L2 norm normalization is performed on the high-dimensional feature vector.
6. The firearm type recognition method based on deep learning according to claim 5, characterized in that, When performing L2 norm normalization on high-dimensional eigenvectors, the following steps are included: Perform a squaring operation on each element of the high-dimensional feature vector, sum all the squaring results, perform a square root operation on the summation result to obtain the L2 norm value, divide each element of the high-dimensional feature vector by the L2 norm value to obtain the normalized high-dimensional feature vector.
7. The firearm type recognition method based on deep learning according to claim 6, characterized in that, When calculating the cosine similarity between the high-dimensional feature vector and all feature vector sets in the firearms feature database, and generating a list of similarity values, the following steps are included: Read the feature vector field from the feature vector data table of the firearm feature database; extract the high-dimensional feature vector from the feature vector field of each feature vector; calculate the dot product between the high-dimensional feature vector of the firearm image to be identified and the high-dimensional feature vector in the firearm feature database, and calculate the Euclidean norm between the high-dimensional feature vector of the firearm image to be identified and the high-dimensional feature vector in the firearm feature database; calculate the similarity value based on the dot product and the Euclidean norm; establish a mapping relationship between the similarity value and the corresponding firearm type identifier; construct a similarity value list containing the correspondence between firearm type identifier and similarity value by arranging all mapping relationships in the order of calculation.
8. The firearm type recognition method based on deep learning according to claim 7, characterized in that, When sorting the similarity value list in descending order, the top four firearm type identifiers with the highest similarity values are selected, including: A quick sorting algorithm is executed on the similarity values in the similarity value list to generate an ordered list arranged from high to low similarity values. When the number of firearm type identifiers in the ordered list is greater than or equal to four, the firearm type identifiers corresponding to the first four entries are selected. When the number of firearm type identifiers in the ordered list is less than four, the firearm type identifiers corresponding to all valid entries are selected.
9. The firearm type recognition method based on deep learning according to claim 8, characterized in that, When extracting the detailed information data corresponding to the first four firearm type identifiers from the firearm feature database, the following is included: Using the selected firearm type identifier as the query condition, a database query operation is performed; the corresponding detailed information data is extracted from the firearm feature database.
10. A deep learning-based firearm type recognition system, used to apply the deep learning-based firearm type recognition method as described in any one of claims 1-9, characterized in that, include: The database construction unit is configured to extract several feature vector sets based on the deep convolutional neural network model and standard firearm images, each feature vector set including a corresponding firearm type identifier and detailed information data; and to construct a firearm feature database based on the several feature vector sets. The image acquisition unit is configured to acquire images of firearms to be identified and to perform preprocessing operations on the images of firearms to be identified. The first processing unit is configured to input the preprocessed image of the firearm to be identified into a deep convolutional neural network model. The deep convolutional neural network model adopts the ResNet-50 architecture and performs transfer learning on the ImageNet dataset. It retains all the convolutional layer structures of the ImageNet pre-trained model and freezes the weight parameters of the convolutional layers. It removes the fully connected layers and classification layers of the pre-trained deep convolutional neural network model and replaces them with output layers adapted to the firearm recognition task, outputting a high-dimensional feature vector with fixed dimensions. The second processing unit is configured to calculate the cosine similarity between the high-dimensional feature vector and all feature vector sets in the firearm feature database, generate a similarity value list, sort the similarity value list in descending order, and select the top four firearm type identifiers with the highest similarity values. The data extraction unit is configured to extract the detailed information data corresponding to the first four firearm type identifiers from the firearm feature database based on the first four firearm type identifiers.
Citation Information
Patent Citations
Transfer learning-based multi-view commodity image retrieval and identification method
CN107908685A
A method for footprint image retrieval
CN111177446A
Target fine granularity identification transfer learning method based on image retrieval
CN117726796A
Deep learning-based qualification image classification method and system
CN117788957A