Dangerous article detection method and device, computer equipment and storage medium
By iteratively fine-tuning the initial representation vector in the open vocabulary detection model, the problems of inefficient and lack of universality of hazardous goods detection in the prior art are solved, and more efficient and flexible hazardous goods detection is achieved.
Patent Information
- Application Number
- CN202510088961.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, hazardous goods detection relies on manual inspections to be inefficient and susceptible to subjective factors, and traditional object detection models lack universality and flexibility, so they cannot detect new hazardous goods.
By determining the category of hazardous goods to be detected, the initial representation vector is obtained, and input it into the open vocabulary detection model for iterative fine-tuning until the model converges, the fine-tuned text representation vector is obtained, which is used to predict the image to be detected.
It improves the accuracy, versatility and flexibility of hazardous goods detection, avoids the drawbacks of manual testing and the limitations of traditional detection models, and does not require a large amount of manual labeling, and can be expanded to new hazardous goods categories.
Smart Images

Figure CN119992056A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and computer vision technology, and in particular to a dangerous goods detection method, device, computer equipment and storage medium. Background Art
[0002] Traditional methods of dangerous goods detection mainly rely on manual inspection and target detection models. Manual inspection is often inefficient and easily affected by subjective factors. Inspectors may miss or misdetect items due to fatigue, lack of experience, etc. For example, in large logistics warehouses, manually checking whether goods contain dangerous goods one by one is a huge workload and time-consuming.
[0003] Traditional target detection models are mostly trained for specific types of dangerous goods, requiring a large amount of data annotation, and can only detect fixed categories of targets, lacking versatility and flexibility. Once a new category of dangerous goods that does not appear in the training set is encountered, the model cannot accurately detect it. Summary of the invention
[0004] Based on this, it is necessary to provide a dangerous goods detection method, device, computer equipment and storage medium to address the above technical issues, so as to solve at least one problem existing in the above-mentioned prior art.
[0005] In a first aspect, a method for detecting dangerous goods is provided, comprising:
[0006] Determine the category of dangerous goods to be tested;
[0007] Obtaining an initial characterization vector corresponding to the category of the dangerous goods to be detected, and using the initial characterization vector as a text characterization vector;
[0008] Inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector;
[0009] The image to be detected is input into the trained open vocabulary detection model to predict the image to be detected by using the fine-tuned text representation vector.
[0010] In one embodiment, obtaining an initial characterization vector corresponding to the category of the dangerous goods to be detected as a text characterization vector includes:
[0011] Concatenate the characters of the dangerous goods names corresponding to each category of dangerous goods to be detected to obtain a concatenated string;
[0012] The concatenated character string is input into the multimodal visual basic model to obtain the initial characterization vector corresponding to each category of dangerous goods to be detected.
[0013] In one embodiment, the concatenated character string is input into a multimodal visual basic model to obtain an initial characterization vector corresponding to each category of dangerous goods to be detected, including:
[0014] Convert the concatenated string into a target text vector;
[0015] Performing attention calculation on the target text vector, and performing nonlinear transformation on the target text vector after the attention calculation;
[0016] The target text vector after nonlinear transformation is normalized and residually connected to obtain the initial representation vector.
[0017] In one embodiment, the step of inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain a fine-tuned text representation vector includes:
[0018] Obtaining image training samples, wherein the image training samples are annotated with dangerous goods category labels;
[0019] Performing image encoding on the image training sample to obtain an image representation vector;
[0020] The text representation vector and the image representation vector are input into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector.
[0021] In one embodiment, the text representation vector and the image representation vector are input into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain a fine-tuned text representation vector, including:
[0022] Step a: Based on the image representation vector and the text representation vector, a prediction result is obtained through forward propagation;
[0023] Step b: Calculating the loss value based on the dangerous goods category label and the prediction result;
[0024] Step c: fine-tuning the text representation vector based on the loss value through back propagation;
[0025] Step d: Based on the image representation vector and the fine-tuned text representation vector, a prediction result is obtained through forward propagation;
[0026] Repeat the above steps bd until the model converges;
[0027] In each iteration, the model parameters of the open vocabulary detection model are frozen.
[0028] In one embodiment, fine-tuning the text representation vector based on the loss value by back propagation includes:
[0029] Based on the loss value, calculating the gradient of the parameter associated with the text representation vector;
[0030] Based on the gradient of the parameter associated with the text representation vector, the parameter associated with the text representation vector is adjusted with the goal of reducing the loss value.
[0031] In one embodiment, obtaining a prediction result through forward propagation based on the image representation vector and the text representation vector includes:
[0032] Performing feature fusion on the image representation vector and the text representation vector to obtain a fused feature map;
[0033] Generating a plurality of candidate regions in the fused feature map by sliding a window;
[0034] Determine the predicted probability of dangerous goods category corresponding to each candidate area;
[0035] Sorting the predicted probabilities of the dangerous goods categories in ascending or descending order;
[0036] Based on the sorting result, the prediction result is obtained.
[0037] In a second aspect, a dangerous goods detection device is provided, comprising:
[0038] A unit for determining the category of dangerous goods to be detected, used for determining the category of dangerous goods to be detected;
[0039] A text representation vector acquisition unit, used to acquire an initial representation vector corresponding to the category of the dangerous goods to be detected, and use the initial representation vector as a text representation vector;
[0040] A training fine-tuning unit, used for inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector;
[0041] The actual prediction unit is used to input the image to be detected into the trained open vocabulary detection model to predict the image to be detected through the fine-tuned text representation vector.
[0042] In a third aspect, a computer device is provided, comprising a memory, a processor, and computer-readable instructions stored in the memory and running on the processor, wherein the processor implements the dangerous goods detection method as described above when executing the computer-readable instructions.
[0043] In a fourth aspect, a readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the dangerous goods detection method as described above is implemented.
[0044] The above-mentioned dangerous goods detection method, device, computer equipment and storage medium, the method implementation includes: determining the category of dangerous goods to be detected; obtaining the initial representation vector corresponding to the category of dangerous goods to be detected, and using the initial representation vector as a text representation vector; inputting the text representation vector and the image training sample into the open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain the fine-tuned text representation vector; inputting the image to be detected into the trained open vocabulary detection model to predict the image to be detected through the fine-tuned text representation vector. In the embodiment of the present application, by determining the category of dangerous goods to be detected and obtaining the corresponding initial representation vector, more targeted semantic information is provided for the model. Iteratively fine-tuning the initial representation vector in the open vocabulary detection model can enable the model to better learn the characteristics of dangerous goods and improve the detection ability of various types of dangerous goods. Finally, the fine-tuned text representation vector is used to predict the image to be detected, which greatly enhances the accuracy, versatility and flexibility of the detection model, and effectively avoids the disadvantages of manual detection and the limitations of traditional detection models. Moreover, it does not require a large amount of manual labeling and can be expanded to new categories of hazardous goods without retraining the model. It is significantly better than the zero-sample method in cross-domain and complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0046] Figure 1 is a schematic diagram of a process of a dangerous goods detection method in one embodiment of the present invention;
[0047] Figure 2 is a schematic diagram of an application environment of a dangerous goods detection method in one embodiment of the present invention;
[0048] Figure 3 is a structural schematic diagram of a dangerous goods detection device in one embodiment of the present invention;
[0049] Figure 4 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] In one embodiment, a method for detecting dangerous goods is provided, comprising the following steps:
[0052] In step S110, the category of the dangerous goods to be detected is determined;
[0053] In an embodiment of the present application, multiple images to be labeled can be collected, and the categories of dangerous goods can be labeled in the images to be labeled through target detection and labeling, such as 30 images of paint buckets and fire extinguishers, as image training samples. For example, a bounding box labeling method can be used to select the location of the dangerous goods, and then label each labeled bounding box with the category of the dangerous goods. It should be noted that all images to be labeled are uniformly labeled according to the same specifications to ensure that the labeling results are stable and reliable. The position and category of the labeled bounding box must be highly accurate, and any deviation may mislead model training.
[0054] It is understandable that the outline of dangerous goods in the image can be preliminarily identified through manual labeling methods or image labeling software, such as edge detection algorithms, and then the outline can be adjusted by labelers, or dangerous goods can be identified through labeling models, such as the Faster R-CNN model, and dangerous goods can be labeled based on the identification results.
[0055] Among them, the categories of dangerous goods may include paint buckets, fire extinguishers, corrosives, flammables, fragile items, etc. By statistically analyzing the labeled results, all the categories of dangerous goods to be tested can be determined.
[0056] In step S120, an initial characterization vector corresponding to the category of the dangerous goods to be detected is obtained, and the initial characterization vector is used as a text characterization vector;
[0057] In the embodiment of the present application, a multimodal visual basic model can be used to obtain the initial representation vector. The multimodal visual basic model is pre-trained on a very large-scale data set. Therefore, it exhibits better generalization ability than the traditional visual model and can be migrated to different fields and downstream tasks at low cost. In addition, since the visual basic model usually has two modal inputs, visual and text, it is often referred to as a multimodal model. Representative visual basic models include CLIP, GLIP, DINO, etc.
[0058] Optionally, the names of the categories of hazardous materials to be detected can be concatenated, for example, by a separator "," to form a concatenated string. Then, the concatenated string can be input into a multimodal visual basic model, such as contrastive language-image pretraining (CLIP), for processing to obtain the initial characterization vector. For example, if the categories of hazardous materials are paint buckets and fire extinguishers, they are concatenated and input into the multimodal visual basic model, and the initial feature vectors corresponding to the paint buckets and fire extinguishers can be output respectively. Therefore, the number of initial characterization vectors is consistent with the number of hazardous materials categories, that is, each hazardous materials category can correspond to an initial characterization vector.
[0059] In step S130, the text representation vector and the image training sample are input into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector;
[0060] Among them, open vocabulary detection models include but are not limited to GLIP, Grounding-DINO, and YOLO-World models.
[0061] Optionally, the image training sample can be encoded to obtain an image representation vector. The image representation vector and the text representation vector are then feature fused, and a prediction result is obtained through forward calculation based on the fused features. A loss value is then calculated based on the prediction result and the pre-labeled true label. The gradients of various parameters associated with the text representation vector are calculated through back propagation based on the loss value, and various parameters associated with the text representation vector are adjusted based on the gradients of various parameters to achieve fine-tuning of the text representation vector in each iteration. The fine-tuned text representation vector and the image representation vector are then predicted again through the open vocabulary detection model, and the fine-tuned text representation vector is adjusted again. The above steps are repeated until the model converges, and the final fine-tuned text adjustment vector can be obtained.
[0062] It should be noted that the weight parameters of the open vocabulary detection model are fixed in each iteration process, that is, the open vocabulary detection model does not change or update the model parameters during the iteration process, and only fine-tunes the text adjustment vector.
[0063] In step S140, the image to be detected is input into the trained open vocabulary detection model to predict the image to be detected by using the fine-tuned text representation vector.
[0064] Optionally, the image to be detected is encoded to obtain the features of the image to be detected, and then the fine-tuned text representation vector is fused with the features of the image to be detected. As mentioned above, the text representation vector contains semantic information related to the category of dangerous goods to be detected. Through feature fusion, image features can be associated with text semantic information to determine which areas in the image may correspond to specific categories of dangerous goods. For example, for a text representation vector representing "paint bucket", the model will look for matching feature patterns in the image features to determine whether there is a paint bucket in the image.
[0065] Based on the fused features after fusion, the location and category of dangerous goods that may exist in the image to be detected can be predicted. For example, a series of bounding boxes can be output, each of which represents a possible location of dangerous goods, and each bounding box is assigned a category label, such as "paint bucket", "fire extinguisher", etc. At the same time, a confidence score for each prediction result is given to indicate the degree of certainty of the model for the prediction. Finally, the prediction results can be filtered according to the set confidence threshold to discard prediction results with lower confidence. Finally, the filtered prediction results are displayed in a visual way on the original image to be detected, for example, the bounding box of the detected dangerous goods is drawn on the image, and the corresponding category and confidence score are marked, so that the user can intuitively understand the detection of dangerous goods in the image.
[0066] See also Figure 2 , the overall process is as follows: first determine the category of dangerous goods, and input the dangerous goods category text into the multimodal basic model for processing to obtain the original representation vector, that is, the text representation vector. Then input the training sample and the text representation vector into the open vocabulary target detection model to iteratively fine-tune the text representation vector until the model converges, and then the fine-tuned representation vector can be obtained, and then the new sample is predicted based on the fine-tuned representation vector. It should be noted that the open vocabulary target detection model needs to freeze the parameters of the open vocabulary target detection model (such as Grounding-DINO) during the fine-tuning of the text representation vector. This means that these parameters will not be updated during the fine-tuning process. By fine-tuning the representation vector to better adapt to specific target detection tasks, while avoiding large-scale changes to the existing model structure.
[0067] The embodiment of the present application provides a method for detecting dangerous goods, including: determining the category of dangerous goods to be detected; obtaining the initial representation vector corresponding to the category of dangerous goods to be detected, and using the initial representation vector as a text representation vector; inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain a fine-tuned text representation vector; inputting the image to be detected into the trained open vocabulary detection model to predict the image to be detected through the fine-tuned text representation vector. In the embodiment of the present application, by determining the category of dangerous goods to be detected and obtaining the corresponding initial representation vector, more targeted semantic information is provided for the model. Iteratively fine-tuning the initial representation vector in the open vocabulary detection model can enable the model to better learn the characteristics of dangerous goods and improve the detection ability of various types of dangerous goods. Finally, the fine-tuned text representation vector is used to predict the image to be detected, which greatly enhances the accuracy, versatility and flexibility of the detection model, and effectively avoids the disadvantages of manual detection and the limitations of traditional detection models. Moreover, it does not require a large amount of manual labeling and can be expanded to new categories of hazardous goods without retraining the model. It is significantly better than the zero-sample method in cross-domain and complex scenarios.
[0068] In an embodiment of the present application, the step of obtaining an initial characterization vector corresponding to the category of the dangerous goods to be detected as a text characterization vector includes:
[0069] Concatenate the characters of the dangerous goods names corresponding to each category of dangerous goods to be detected to obtain a concatenated string;
[0070] The concatenated character string is input into the multimodal visual basic model to obtain the initial characterization vector corresponding to each category of dangerous goods to be detected.
[0071] Optionally, the names of the categories of dangerous goods to be detected can be concatenated, for example, they can be concatenated by a separator "," to form a concatenated string. Then, the concatenated string can be input into a multimodal visual basic model, such as Contrastive Language-Image Pretraining (CLIP), for processing to obtain the initial representation vector. For example, the English representations of a paint bucket and a fire extinguisher are "paint bucket" and "extinguisher", respectively. They are concatenated into a string "paint bucket, extinguisher" separated by English commas and input into the multimodal visual basic model CLIP, and the initial representation vectors corresponding to the paint bucket and the fire extinguisher can be output respectively.
[0072] It should be noted that the names of dangerous goods corresponding to all the categories of dangerous goods to be detected can be concatenated in sequence, or the names of dangerous goods in the same category can be concatenated.
[0073] In one embodiment of the present application, the concatenated character string is input into the multimodal visual basic model to obtain the initial characterization vector corresponding to each category of dangerous goods to be detected, including:
[0074] Convert the concatenated string into a target text vector;
[0075] Performing attention calculation on the target text vector, and performing nonlinear transformation on the target text vector after the attention calculation;
[0076] The target text vector after nonlinear transformation is normalized and residually connected to obtain the initial representation vector.
[0077] Optionally, the concatenated string can be embedded to obtain the target text vector, and then the text encoder inside the CLIP model, which has a multi-head attention mechanism, can perform attention calculation on the target text vector to capture the semantic information in the text, the association between words, etc. Then, the target text vector after the multi-head attention calculation can be input into the feedforward neural network for nonlinear transformation. The feedforward neural network usually consists of two fully connected layers with a nonlinear activation function in the middle, such as ReLU (Rectified Linear Unit). After being processed by the first fully connected layer, it is calculated by the activation function and finally processed by the second fully connected layer and output. The target text vector after nonlinear transformation is normalized to make the data distribution more stable. At the same time, in order to avoid problems such as gradient disappearance, residual connection is also used. The target text vector processed by the multi-head attention mechanism and the target text vector after the feedforward neural network and normalization can be residually connected to obtain the initial representation vector.
[0078] For example, for the input string of "paint bucket, fire extinguisher", the text encoder will analyze the semantics of each hazardous material name and the combined semantic relationship between them. Through layer-by-layer calculation and feature extraction, the semantic information of the text is converted into a vector representation of a specific dimension to obtain the initial representation vector.
[0079] In one embodiment of the present application, the step of inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain a fine-tuned text representation vector includes:
[0080] Obtaining image training samples, wherein the image training samples are annotated with dangerous goods category labels;
[0081] Performing image encoding on the image training sample to obtain an image representation vector;
[0082] The text representation vector and the image representation vector are input into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector.
[0083] Alternatively, see Figure 3 , obtain image training samples, and encode the image training samples through an image encoder to obtain an image representation vector. Then, the image representation vector and the text representation vector can be input into the open vocabulary detection model for iteration to achieve iterative fine-tuning of the text representation vector until the model converges, for example, the number of iterations reaches a preset number, such as 1000 times, or the loss value is less than a preset threshold, etc., and the iteration ends. At this time, the fine-tuned text representation vector can be obtained for subsequent prediction processing of new samples.
[0084] It should be noted that the image encoder can adopt a convolutional neural network or a Transformer architecture. Taking a convolutional neural network as an example, it includes a convolutional layer, a pooling layer, and a fully connected layer. First, the image is convolved by the convolution kernel of the convolution layer, and each convolution kernel will learn different features, such as edges, textures, etc. When a convolution kernel convolves an image, it generates a feature map. Through the operation of multiple convolution kernels, multiple feature maps will be obtained. These feature maps are combined to preliminarily extract the local features of the image. Then, the local features of the extracted image can be processed by maximum pooling or average pooling through the pooling layer, which can reduce the dimension of the data, reduce the amount of calculation, and have a certain robustness to small changes such as displacement and rotation of the image. Then, in the fully connected layer, a linear combination is performed through the weight matrix, and a nonlinear transformation is performed through the activation function (such as ReLU, Sigmoid, etc.), and a fixed-dimensional image representation vector can be obtained.
[0085] Taking the Transformer architecture as an example, the image samples are first divided into multiple small blocks of fixed size, and these small blocks are linearly projected (through a fully connected layer) into a lower-dimensional vector space. This process is like encoding each small block to obtain the embedding vector (Embedding) of each small block.
[0086] Then, position encoding (sine-cosine position encoding) can be added to each small block so that the model can distinguish small blocks at different positions. The small block embedding vector after position encoding will enter multiple Transformer Encoder layers. In the Transformer Encoder layer, there are mainly two parts: multi-head attention mechanism and feed-forward neural network. First, the attention calculation is performed on each small block through the multi-head attention mechanism, and then the output of the attention mechanism is nonlinearly transformed through the feed-forward neural network, and then the image representation vector is output through average pooling or aggregation.
[0087] In one embodiment of the present application, the text representation vector and the image representation vector are input into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain a fine-tuned text representation vector, including:
[0088] Step a: Based on the image representation vector and the text representation vector, a prediction result is obtained through forward propagation;
[0089] Step b: Calculating the loss value based on the dangerous goods category label and the prediction result;
[0090] Step c: fine-tuning the text representation vector based on the loss value through back propagation;
[0091] Step d: Based on the image representation vector and the fine-tuned text representation vector, a prediction result is obtained through forward propagation;
[0092] Repeat the above steps bd until the model converges;
[0093] In each iteration, the model parameters of the open vocabulary detection model are frozen.
[0094] Optionally, an image representation vector is randomly selected from the image representation vectors and inputted into the open vocabulary detection model together with the text representation vector. The image representation vector and the text representation vector are first predicted by forward propagation. After the prediction result is obtained, the loss value corresponding to the prediction result can be calculated based on the real label of the dangerous goods category pre-annotated in the image representation vector and the loss function. Then, the gradient of the associated parameters of the text representation vector can be calculated based on the loss value, and the associated parameters of the text representation vector can be fine-tuned in the opposite direction of the gradient, so that the loss value is gradually reduced. Then, an image representation vector can be selected again and inputted into the open vocabulary detection model together with the fine-tuned text representation vector for prediction again, and the above steps are repeated until the model converges, and the fine-tuned text representation vector can be obtained.
[0095] It is understandable that the loss function may include multiple types, such as cross entropy loss function, L1 loss function or SmoothL1 loss function. Among them, the cross entropy loss function can be used to calculate the loss value for classification loss, and the L1 loss or Smooth L1 loss can be used to calculate the loss value for the location loss of the dangerous goods.
[0096] It should be noted that during the iteration of the open vocabulary detection model, the model parameters need to be frozen, which means that these parameters will not be updated during the fine-tuning process. This is because the main architecture of the model has learned some common knowledge and patterns through pre-training or other means, so the representation vector can be fine-tuned to better adapt to specific object detection tasks while avoiding large-scale changes to the existing model structure.
[0097] In one embodiment of the present application, fine-tuning the text representation vector based on the loss value by back propagation includes:
[0098] Based on the loss value, calculating the gradient of the parameter associated with the text representation vector;
[0099] Based on the gradient of the parameter associated with the text representation vector, the parameter associated with the text representation vector is adjusted with the goal of reducing the loss value.
[0100] Optionally, based on the calculated loss value, the chain rule can be used to calculate the gradient of the parameters that need to be updated in the model (i.e., the parameters related to the text representation vector). This process calculates the derivative of the loss function for each trainable parameter. These derivatives indicate the direction of change of the parameter, which reduces the value of the loss function. Based on the calculated gradient, the gradient descent algorithm can be used to update the parameters related to the representation vector. Common gradient descent algorithms include stochastic gradient descent (SGD), Adagrad, Adadelta, Adam, etc. Taking the Adam algorithm as an example, it adjusts the learning rate based on the first-order moment estimate and the second-order moment estimate of the gradient, and adaptively updates the parameters.
[0101] In an embodiment of the present application, obtaining a prediction result through forward propagation based on the image representation vector and the text representation vector includes:
[0102] Performing feature fusion on the image representation vector and the text representation vector to obtain a fused feature map;
[0103] Generating a plurality of candidate regions in the fused feature map by sliding a window;
[0104] Determine the predicted probability of dangerous goods category corresponding to each candidate area;
[0105] The predicted probabilities of the dangerous goods categories are sorted in ascending or descending order, and the predicted results are obtained based on the sorting results.
[0106] Optionally, the image representation vector and the text representation vector are cross-modally fused, such as by using an attention mechanism to perform feature fusion. For example, a multi-head attention mechanism can be used to calculate the attention weights between the image representation vector and the text representation vector. For each position in the image feature and each element in the text representation vector, the attention weights are determined by calculating the similarity scores, and then the image features and the text representation vectors are weighted and summed according to these weights to achieve feature fusion. In this way, the fused features contain both the visual information of the image and the semantic information of the text, and can be better used for target detection.
[0107] Based on the fused feature map, a small window can be slid in the feature map through the Region Proposal Network (RPN), and the region can be judged whether the region contains dangerous goods based on the features in the window. For example, the feature score in the window can be calculated and compared with the preset score threshold. If it exceeds the preset score threshold, it means that it exists. If it exists, it can be used as a candidate region, so that multiple candidate regions can be obtained. These candidate regions can be represented in the form of bounding boxes, including location (such as the coordinates of the upper left corner and the lower right corner) and size information.
[0108] For each candidate region, the classifier can be used to calculate the probability that each candidate region belongs to different dangerous goods categories, such as [0.1 (background), 0.7 (paint bucket), 0.2 (fire extinguisher)], which means that the probability of this region belonging to the paint bucket is the highest. A threshold can be set according to the category probability. For example, only candidate regions with a category probability higher than 0.5 are retained, and the retained candidate regions are sorted in ascending or descending order from high to low according to the category probability. Based on the sorting results, the final prediction results can be determined, and the detected dangerous goods, such as the location and category of paint buckets and fire extinguishers, can be directly marked on the image based on the prediction results.
[0109] In an embodiment of the present application, by determining the category of dangerous goods to be detected and obtaining the corresponding initial representation vector, more targeted semantic information is provided to the model. Iterative fine-tuning of the initial representation vector in the open vocabulary detection model enables the model to better learn the characteristics of dangerous goods and improve the detection capabilities of various types of dangerous goods. Finally, the fine-tuned text representation vector is used to predict the image to be detected, which greatly enhances the accuracy, versatility and flexibility of the detection model, and effectively avoids the drawbacks of manual detection and the limitations of traditional detection models. Moreover, it does not require a large amount of manual annotation, and can be expanded to new categories of dangerous goods without retraining the model. The effect is significantly better than the zero-sample method in cross-domain and complex scenarios.
[0110] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0111] In one embodiment, a dangerous goods detection device is provided, which corresponds one-to-one to the dangerous goods detection method in the above embodiment. Figure 3 As shown, the dangerous goods detection device includes a dangerous goods category determination unit 10, a text representation vector acquisition unit 20, a training fine-tuning unit 30 and an actual prediction unit 40. The functional modules are described in detail as follows:
[0112] The dangerous goods category determination unit 10 is used to determine the category of the dangerous goods to be detected;
[0113] A text representation vector acquisition unit 20, used to acquire an initial representation vector corresponding to the category of the dangerous goods to be detected, and use the initial representation vector as a text representation vector;
[0114] A training fine-tuning unit 30, configured to input the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector;
[0115] The actual prediction unit 40 is used to input the image to be detected into the trained open vocabulary detection model to predict the image to be detected by using the fine-tuned text representation vector.
[0116] In one embodiment of the present application, the text representation vector acquisition unit 20 is further used to:
[0117] Concatenate the characters of the dangerous goods names corresponding to each category of dangerous goods to be detected to obtain a concatenated string;
[0118] The concatenated character string is input into the multimodal visual basic model to obtain the initial characterization vector corresponding to each category of dangerous goods to be detected.
[0119] In one embodiment of the present application, the text representation vector acquisition unit 20 is further used to:
[0120] Convert the concatenated string into a target text vector;
[0121] Performing attention calculation on the target text vector, and performing nonlinear transformation on the target text vector after the attention calculation;
[0122] The target text vector after nonlinear transformation is normalized and residually connected to obtain the initial representation vector.
[0123] In one embodiment of the present application, the training fine-tuning unit 30 is further used to:
[0124] Obtaining image training samples, wherein the image training samples are annotated with dangerous goods category labels;
[0125] Performing image encoding on the image training sample to obtain an image representation vector;
[0126] The text representation vector and the image representation vector are input into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector.
[0127] In one embodiment of the present application, the training fine-tuning unit 30 is further used to:
[0128] Step a: Based on the image representation vector and the text representation vector, a prediction result is obtained through forward propagation;
[0129] Step b: Calculating the loss value based on the dangerous goods category label and the prediction result;
[0130] Step c: fine-tuning the text representation vector based on the loss value through back propagation;
[0131] Step d: Based on the image representation vector and the fine-tuned text representation vector, a prediction result is obtained through forward propagation;
[0132] Repeat the above steps bd until the model converges;
[0133] In each iteration, the model parameters of the open vocabulary detection model are frozen.
[0134] In one embodiment of the present application, the training fine-tuning unit 30 is further used to:
[0135] Based on the loss value, calculating the gradient of the parameter associated with the text representation vector;
[0136] Based on the gradient of the parameter associated with the text representation vector, the parameter associated with the text representation vector is adjusted with the goal of reducing the loss value.
[0137] In one embodiment of the present application, the training fine-tuning unit 30 is further used to:
[0138] Performing feature fusion on the image representation vector and the text representation vector to obtain a fused feature map;
[0139] Generating a plurality of candidate regions in the fused feature map by sliding a window;
[0140] Determine the predicted probability of dangerous goods category corresponding to each candidate area;
[0141] The predicted probabilities of the dangerous goods categories are sorted in ascending or descending order, and the predicted results are obtained based on the sorting results.
[0142] In an embodiment of the present application, by determining the category of dangerous goods to be detected and obtaining the corresponding initial representation vector, more targeted semantic information is provided to the model. Iterative fine-tuning of the initial representation vector in the open vocabulary detection model enables the model to better learn the characteristics of dangerous goods and improve the detection capabilities of various types of dangerous goods. Finally, the fine-tuned text representation vector is used to predict the image to be detected, which greatly enhances the accuracy, versatility and flexibility of the detection model, and effectively avoids the drawbacks of manual detection and the limitations of traditional detection models. Moreover, it does not require a large amount of manual annotation, and can be expanded to new categories of dangerous goods without retraining the model. The effect is significantly better than the zero-sample method in cross-domain and complex scenarios.
[0143] For the specific definition of the dangerous goods detection device, please refer to the definition of the dangerous goods detection method above, which will not be repeated here. Each module in the above-mentioned dangerous goods detection device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0144] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a method for detecting dangerous goods is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0145] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the above-mentioned dangerous goods detection method are implemented.
[0146] In an embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of the dangerous goods detection method as described above are implemented.
[0147] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through computer-readable instructions, and the computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they may include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0148] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0149] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for detecting dangerous goods, characterized in that: The method comprises: Determine the category of dangerous goods to be tested; Obtaining an initial characterization vector corresponding to the category of the dangerous goods to be detected, and using the initial characterization vector as a text characterization vector; Inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector; The image to be detected is input into the trained open vocabulary detection model to predict the image to be detected by using the fine-tuned text representation vector.
2. The method for detecting dangerous goods according to claim 1, characterized in that: The step of obtaining an initial characterization vector corresponding to the category of the dangerous goods to be detected as a text characterization vector includes: Concatenate the characters of the dangerous goods names corresponding to each category of dangerous goods to be detected to obtain a concatenated string; The concatenated character string is input into the multimodal visual basic model to obtain the initial characterization vector corresponding to each category of dangerous goods to be detected.
3. The method for detecting dangerous goods according to claim 2, characterized in that: The concatenated character string is input into the multimodal visual basic model to obtain the initial characterization vector corresponding to each category of dangerous goods to be detected, including: Convert the concatenated string into a target text vector; Performing attention calculation on the target text vector, and performing nonlinear transformation on the target text vector after the attention calculation; The target text vector after nonlinear transformation is normalized and residually connected to obtain the initial representation vector.
4. The method for detecting dangerous goods according to claim 1, wherein: The step of inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain a fine-tuned text representation vector includes: Obtaining image training samples, wherein the image training samples are annotated with dangerous goods category labels; Performing image encoding on the image training sample to obtain an image representation vector; The text representation vector and the image representation vector are input into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector.
5. The method for detecting dangerous goods according to claim 4, characterized in that: Inputting the text representation vector and the image representation vector into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges to obtain a fine-tuned text representation vector, including: Step a: Based on the image representation vector and the text representation vector, a prediction result is obtained through forward propagation; Step b: Calculating the loss value based on the dangerous goods category label and the prediction result; Step c: fine-tuning the text representation vector based on the loss value through back propagation; Step d: Based on the image representation vector and the fine-tuned text representation vector, a prediction result is obtained through forward propagation; Repeat the above steps bd until the model converges; In each iteration, the model parameters of the open vocabulary detection model are frozen.
6. The method for detecting dangerous goods according to claim 5, characterized in that: The fine-tuning of the text representation vector based on the loss value by back propagation includes: Based on the loss value, calculating the gradient of the parameter associated with the text representation vector; Based on the gradient of the parameter associated with the text representation vector, the parameter associated with the text representation vector is adjusted with the goal of reducing the loss value.
7. The method for detecting dangerous goods according to claim 4, characterized in that: The obtaining a prediction result through forward propagation based on the image representation vector and the text representation vector includes: Performing feature fusion on the image representation vector and the text representation vector to obtain a fused feature map; Generating a plurality of candidate regions in the fused feature map by sliding a window; Determine the predicted probability of dangerous goods category corresponding to each candidate area; Sorting the predicted probabilities of the dangerous goods categories in ascending or descending order; Based on the sorting result, the prediction result is obtained.
8. A dangerous goods detection device, characterized in that: The device comprises: A unit for determining the category of dangerous goods to be detected, used for determining the category of dangerous goods to be detected; A text representation vector acquisition unit, used to acquire an initial representation vector corresponding to the category of the dangerous goods to be detected, and use the initial representation vector as a text representation vector; A training fine-tuning unit, used for inputting the text representation vector and the image training sample into an open vocabulary detection model to iteratively fine-tune the text representation vector until the model converges, thereby obtaining a fine-tuned text representation vector; The actual prediction unit is used to input the image to be detected into the trained open vocabulary detection model to predict the image to be detected through the fine-tuned text representation vector.
9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executed on the processor, characterized in that: When the processor executes the computer-readable instructions, the dangerous goods detection method according to any one of claims 1 to 7 is implemented.
10. A readable storage medium having computer readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the dangerous goods detection method according to any one of claims 1 to 7 is implemented.