Article detection method, electronic equipment and computer readable medium
By combining the autoencoder model with the convolutional layer, combined with the sample feature list and cluster analysis, the accuracy and efficiency problems of traditional security gates are solved, efficient and accurate object detection is achieved, and it is suitable for the rapid identification of new types of objects.
Patent Information
- Application Number
- CN202410309098.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional security inspection gates have limited accuracy, limited judgment ability, low efficiency, and poor anti-interference ability. In addition, security inspection gates based on CNN models take a long time to train, are highly dependent on high-performance hardware, and have poor on-site flexibility.
An autoencoder model is used for unsupervised learning. By adding a convolutional layer before the encoder and combining the sample feature list with cluster analysis, the item type can be determined, reducing manual labeling and computational complexity, and improving detection accuracy and efficiency.
Significantly save model training time and costs, improve the accuracy and efficiency of object detection, adapt to the detection of new types of objects, reduce dependence on hardware, and enhance anti-interference capabilities.
Smart Images

Figure CN120670872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of security inspection technology, and in particular to an object detection method, an electronic device and a computer-readable medium. Background Art
[0002] Security gates are devices used to detect contraband when passing through them. Traditional security gates detect contraband by measuring the amplitude and phase of signals. These gates suffer from limitations such as limited accuracy, limited capabilities, low efficiency, and poor anti-interference capabilities. In response, a new generation of security gates based on convolutional neural network (CNN) models has emerged. However, CNN models take a long time to train and rely heavily on high-performance hardware. Summary of the Invention
[0003] The present invention aims to solve one of the technical problems in the related art to a certain extent. To this end, the present invention provides an object detection method, an electronic device and a computer-readable medium.
[0004] The present invention first provides an object detection method, wherein the object detection method includes:
[0005] Process the original detection data collected by the security door to obtain detection demodulation data;
[0006] Inputting the detected demodulated data into a pre-trained autoencoder model to obtain the detected low-dimensional features output by the encoder of the autoencoder model; wherein the autoencoder model includes a convolutional layer, the encoder, and a decoder cascaded in sequence, and the hyperparameters of the encoder and the hyperparameters of the decoder are associated with the hyperparameters of the convolutional layer;
[0007] When the locally stored sample feature list is not empty, determining the similarity between the features in the sample feature list and the detection low-dimensional features; wherein the sample feature list includes a mapping relationship between type information of multiple items and multiple sample features;
[0008] When a target feature exists in the sample feature list, the type of the item of the target feature is used as the type of the item corresponding to the detection original data, wherein the target feature is a feature whose similarity with the detection low-dimensional feature meets a preset condition;
[0009] When the locally stored sample feature list is empty, and / or when the target feature does not exist in the sample feature list, cluster analysis is performed on the detected low-dimensional features to obtain corresponding label information, and the label information is used as the type of the object corresponding to the original detection data.
[0010] Optionally, performing cluster analysis on the detected low-dimensional features to obtain corresponding label information includes:
[0011] Clustering the detected low-dimensional features to obtain feature categories of the detected low-dimensional features;
[0012] The label information for detecting the low-dimensional feature is determined based on a mapping relationship between the feature category and the label information used to characterize the type of the item.
[0013] Optionally, in the step of performing cluster analysis on the detected low-dimensional features to obtain corresponding label information, the detected low-dimensional features are input into a pre-trained Gaussian mixture model, so that the pre-trained Gaussian mixture model performs the step of clustering the detected low-dimensional features and the step of determining the label information of the detected low-dimensional features based on the mapping relationship between the feature category and the label information used to characterize the item type; wherein,
[0014] The Gaussian mixture model includes multiple single Gaussian models, and the multiple single Gaussian models correspond one-to-one to multiple label information used to characterize the type of item.
[0015] Optionally, the object detection method further includes:
[0016] Upon receiving a test result correction instruction, obtaining type information of the item included in the test result correction instruction;
[0017] Adding the type information of the item carried in the detection result correction instruction to the detection low-dimensional feature;
[0018] The detected low-dimensional features after adding the type information of the object are used as sample features and stored in the sample feature list.
[0019] Optionally, the object detection method further includes:
[0020] When the total number of features in the sample feature list is greater than a first preset number threshold, the features in the sample feature list are sent to the server so that the server can store the features in the sample feature list in a local clustering database according to the type information of the items carried.
[0021] Optionally, determining the similarity between the features in the sample feature list and the low-dimensional features includes:
[0022] Calculating the mean square error between the features in the sample feature list and the detection low-dimensional features, wherein the mean square error is used to represent the similarity;
[0023] The target feature is a feature whose mean square error with the detected low-dimensional feature is not greater than a preset similarity threshold.
[0024] Optionally, the security gate is provided with at least a first number of transmitting coils and a second number of receiving coils which are positioned independently of each other; the first number of transmitting coils respectively use different frequencies; the original detection data collected by the security gate includes a first number of transmitting signals and a second number of receiving signals.
[0025] Optionally, the detected demodulated data includes multiple channels of vector data, each channel of the vector data corresponds to one channel of the transmitted signal and one channel of the received signal, and the transmitted signals and received signals corresponding to any two channels of the vector data are not all the same;
[0026] In the step of inputting the detection demodulation data into a pre-trained autoencoder model to obtain the detection low-dimensional features output by the encoder of the autoencoder model, the convolutional layer performs the following steps:
[0027] Sorting the multi-channel vector data according to the corresponding received signals;
[0028] Among the sorted multi-channel vector data, each group of multi-channel vector data corresponding to the same received signal is sorted again according to the corresponding transmitted signal;
[0029] Performing convolution processing on the re-sorted multi-channel vector data to obtain detection convolution features;
[0030] The detection convolution feature is passed to the encoder, so that the encoder outputs the detection low-dimensional feature according to the detection convolution feature.
[0031] Furthermore, the present invention also provides an electronic device, wherein the electronic device includes:
[0032] one or more processors;
[0033] A memory having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement the object detection method provided by the present invention.
[0034] Furthermore, the present invention also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the object detection method provided by the present invention.
[0035] In the object detection method provided by the present invention, the detection raw data collected by the security gate is processed to obtain detection demodulation data; the detection demodulation data is input into a pre-trained autoencoder model to obtain the detection low-dimensional features output by the encoder of the autoencoder model; the autoencoder model is used for the detection of the type of object, and there is no need to manually label the data when training the model, which directly avoids the manual workload; by adding a convolution layer in front of the encoder, the convolution layer of the trained autoencoder model can extract the detection convolution features from the detection demodulation data, thereby enhancing the feature expression ability, reducing the amount of parameter data and the computational complexity; the encoder of the autoencoder model can map the detection convolution features to a low-dimensional space to obtain the detection low-dimensional features, thereby realizing the compression and dimensionality reduction of the detection convolution features and reducing the storage and computational costs; the autoencoder model can learn the important features of the detection demodulation data and remove redundant features. The residual information improves the accuracy of the detection results of the object type; by determining the similarity between the features in the sample feature list and the detection low-dimensional features when the locally stored sample feature list including the mapping relationship between the type information of multiple objects and multiple sample features is not empty, and when there is a target feature in the sample feature list whose similarity with the detection low-dimensional features meets the preset conditions, the type of the object of the target feature is used as the type of the object corresponding to the detection original data; and when the locally stored sample feature list is empty, and / or when there is no target feature in the sample feature list, cluster analysis is performed on the detection low-dimensional features to obtain the corresponding label information, and the label information is used as the type of the object corresponding to the detection original data, which can significantly save model training time and training costs, speed up the speed of object detection, thereby improving the efficiency of object detection and improving the accuracy of object detection. At the same time, this method can also be used for a large number of security inspection doors in the same application scenarios, widely improving the efficiency and accuracy of object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention will be further described below in conjunction with the accompanying drawings:
[0037] Figure 1 This is a schematic diagram of the principle of the autoencoder model in related technologies;
[0038] Figure 2 This is a flow chart of an implementation of the object detection method provided by an embodiment of the present invention;
[0039] Figure 3 This is a flow chart of an implementation method of performing cluster analysis on detected low-dimensional features to obtain corresponding label information, as provided in an embodiment of the present invention;
[0040] Figure 4 1 is a flow chart of another embodiment of the object detection method provided in an embodiment of the present invention;
[0041] Figure 5 This is a flowchart of an implementation method for determining the similarity between sample features and detecting low-dimensional features provided by an embodiment of the present invention;
[0042] Figure 6 It is a structural diagram of a security door in the related art;
[0043] Figure 7 Schematic diagram of the distribution of a receiving coil implementation provided by an embodiment of the present invention;
[0044] Figure 8 1 is a flow chart of an implementation of the steps performed by the convolutional layer provided in an embodiment of the present invention;
[0045] Figure 9 This is a flowchart of an implementation of a method for training a Gaussian mixture model provided in an embodiment of the present invention;
[0046] Figure 10 This is a flowchart of an implementation of a method for training an autoencoder model provided in an embodiment of the present invention;
[0047] Figure 11 Schematic diagram of the structure of an autoencoder model provided by an embodiment of the present invention;
[0048] Figure 12 is a structural diagram of a detection system provided by an embodiment of the present invention;
[0049] Figure 13 A schematic diagram of a workflow provided by an embodiment of the present invention;
[0050] Figure 14 This is a module diagram of an embodiment of the electronic device provided by the present invention;
[0051] Figure 15 It is a module diagram of the computer-readable medium provided by the present invention.
[0052] Description of Reference Numerals
[0053] 101: Processor 102: Memory
[0054] 103: I / O interface 104: bus DETAILED DESCRIPTION
[0055] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described in the embodiments are intended to explain the present invention and are not to be construed as limiting the present invention.
[0056] References in this specification to "one embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment itself can be included in at least one embodiment disclosed herein. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.
[0057] A security gate is a device used to detect whether a person passing through the gate is carrying prohibited items. Traditional security gate object detection methods detect prohibited items based on the amplitude and phase of the signal. This method has some problems:
[0058] 1. Limited accuracy: Traditional object detection methods for security gates rely on preset amplitude and phase thresholds to determine object attributes. This is relatively simple and often cannot effectively cope with complex scenarios and a wide variety of prohibited items. For example, when prohibited items are placed together with everyday items or when the shape, size, and material of prohibited items change, it may not be possible to make correct judgments. This is not ideal for scenarios that require more precise identification.
[0059] 2. Limited judgment capabilities: Traditional methods for detecting items at security gates cannot provide specific information about the types of prohibited items. The detection results only indicate whether the items carried by the person passing through the gate are prohibited or non-prohibited. This makes it difficult for security personnel to accurately understand the nature and risk level of the items carried by the person passing through the gate when handling alarms. Therefore, they may need to conduct additional manual inspections or use other equipment for further identification, which increases operational complexity and time costs.
[0060] 3. Low efficiency: Traditional methods for detecting objects at security gates require a lot of manual intervention before deployment. For example, constant testing and adjustment of preset thresholds are required to achieve a high detection accuracy, resulting in low efficiency.
[0061] 4. Poor anti-interference ability: Traditional security door object detection methods are more sensitive to electromagnetic interference from the external environment. For example, metal materials, electronic equipment or other electromagnetic sources may affect the amplitude of the coil, resulting in false alarms or missed alarms. This poor anti-interference ability limits its reliability and stability in complex environments.
[0062] At present, the related technology generally uses the convolutional neural network (CNN) model for security door object detection. Although the CNN model has many advantages over the above-mentioned traditional security door object detection methods, it also has huge disadvantages: the training time of the CNN model is long and it is highly dependent on high-performance hardware.
[0063] The training method of the CNN model is essentially a supervised learning method, which requires the use of a large amount of labeled data and a long training time to learn and adjust the model parameters. In addition, in order to train and run the CNN model, strong computing resources and high-performance hardware such as graphics processing units (GPUs) are required.
[0064] In light of this, the inventors of this invention propose using unsupervised learning methods for object detection at security gates. Unsupervised learning methods, such as autoencoder models, do not require labeled data for model training. Instead, they attempt to learn low-dimensional feature representations of the input data, allowing for easy generation or reconstruction of the original data based on these low-dimensional feature representations.
[0065] In machine learning and computer vision, feature extraction involves extracting important information or attributes from raw data. These extracted features can be used to describe certain aspects of the raw data. Feature extraction typically involves mapping the raw data into a new space, often with a higher or lower dimensionality than the original data.
[0066] The dimension of a feature represents the number of elements contained in the feature vector, that is, the length of the feature vector. A feature vector can be thought of as the coordinates of a point in a high-dimensional space, where each dimension represents a different feature.
[0067] In the fields of machine learning and computer vision, low-level features and high-level features are used to describe feature representations of image data at different levels of abstraction. Low-level features are primitive or basic features extracted directly from image data, including color, texture, edges, corners, etc., at a lower level of abstraction, and can directly reflect the basic attributes and local structure of the image. High-level features are feature representations derived through higher-level calculations and analysis based on low-level features. High-level features typically include object shape, posture, contextual information, semantic labels, etc., have a higher level of abstraction, and can represent more complex semantics and concepts in the image.
[0068] like Figure 1As shown, the autoencoder model in the related art usually consists of two parts, one is the encoder and the other is the decoder. As the name suggests, the encoder performs an encoding operation on the input data, that is, the input data is mapped into a low-dimensional space to obtain low-dimensional high-level features; the decoder performs a decoding operation on the low-dimensional high-level features passed by the encoder, that is, these low-dimensional high-level features are remapped to the original high-dimensional space to obtain output data. During the training process, the autoencoder model adjusts the parameters of the encoder and the decoder so that the output data obtained by its own reconstruction (or restoration) is close to the input data, that is, it tries to minimize the difference between the input data and the output data obtained by its own reconstruction in order to learn an effective feature representation of the input data.
[0069] The primary goal of an autoencoder model is to learn effective feature representations of the data. This doesn't necessarily require that the reconstructed output data be exactly the same as the input data. Therefore, during training, the model calculates the error between the predicted and true values using a loss function (for example, the most basic type of loss function is the mean squared error (MSE) function) and gradually reduces this error to improve the model's accuracy and generalization. After training, the feature representations learned by the encoder can be used for various tasks, such as image classification. Autoencoder models offer the following advantages: 1. Unsupervised Learning: Autoencoders are an unsupervised learning method that does not require data labeling. Instead, they can automatically learn effective feature representations from unlabeled data, making them applicable to analysis and modeling of large datasets. 2. Data Compression and Dimensionality Reduction: Autoencoders can map high-dimensional input data into a low-dimensional space, achieving data compression and dimensionality reduction, reducing storage and computational costs. 3. Feature Learning and Representation Learning: Autoencoders can capture important features of the data and remove redundant information by learning a latent space representation of the data.
[0070] However, the inventors of the present invention also found that the data collected by the security gate has the disadvantages of weak feature expression ability, large data volume, and complex calculation. In other words, if an autoencoder model that only includes an encoder and a decoder is directly used, it is very likely that the feature representation of the data cannot be learned completely and accurately.
[0071] In view of this, the inventors of the present invention further proposed to add a convolutional layer before the encoder of the autoencoder model, and integrate the frequency information of each data channel through the convolutional layer, so as to enhance the feature expression ability, reduce the amount of parameter data and computational complexity.
[0072] When using the CNN model, the detection results of the item type can be directly output. However, when using a pre-trained autoencoder model, the low-dimensional feature representation of the input data is learned through the convolutional layer and the encoder, and the detection results of the item type are not directly output. In view of this, the inventors of the present invention further proposed that the sample feature list can be stored locally. After obtaining the low-dimensional features based on the autoencoder model, the low-dimensional features are matched with the features in the sample feature list. When the similarity between the features in the sample feature list and the low-dimensional features is high, the detection results of the item type are directly determined. When the sample feature list is empty and the similarity between the features in the sample feature list and the low-dimensional features is low, the detection results of the item type are determined by performing cluster analysis on the low-dimensional features, thereby effectively ensuring the accuracy of the detection.
[0073] Accordingly, as a first aspect of an embodiment of the present invention, a method for detecting an object is provided, wherein: Figure 2 As shown, the method may include the following steps:
[0074] In step S110, the detection raw data collected by the security gate is processed to obtain detection demodulated data;
[0075] In step S120, the detection demodulated data is input into a pre-trained autoencoder model to obtain the detection low-dimensional features output by the encoder of the autoencoder model; wherein the autoencoder model includes a convolutional layer, the encoder, and a decoder cascaded in sequence, and the hyperparameters of the encoder and the hyperparameters of the decoder are associated with the hyperparameters of the convolutional layer;
[0076] In step S130, if the locally stored sample feature list is not empty, the similarity between the features in the sample feature list and the detection low-dimensional features is determined; wherein the sample feature list includes mapping relationships between type information of multiple items and multiple sample features;
[0077] In step S140, if a target feature exists in the sample feature list, the type of the item of the target feature is used as the type of the item corresponding to the detection original data, wherein the target feature is a feature whose similarity with the detection low-dimensional feature meets a preset condition;
[0078] In step S150, when the locally stored sample feature list is empty, and / or when the target feature does not exist in the sample feature list, cluster analysis is performed on the detected low-dimensional features to obtain corresponding label information, and the label information is used as the type of the object corresponding to the detected original data.
[0079] Among them, after the detection demodulation data is input into the pre-trained autoencoder model, the convolution layer of the autoencoder model performs convolution processing on the detection demodulation data to obtain detection convolution features and pass them to the encoder of the autoencoder model. The encoder of the autoencoder model encodes the detection convolution features and outputs detection low-dimensional features.
[0080] The autoencoder model is pre-trained based on the initial model. The initial model includes an initial convolutional layer, an initial encoder, and an initial decoder that are cascaded in sequence. During the pre-training process, the initial convolutional layer performs convolution processing on the training demodulation data to obtain training convolution features; the initial encoder performs encoding processing on the training convolution features to obtain training low-dimensional features; the initial decoder performs decoding processing on the training low-dimensional features to obtain output features; according to the loss function between the input features (i.e., the training demodulation data input to the initial model) and the output features, the parameters of the initial convolutional layer, the initial encoder, and the initial decoder are updated. The hyperparameters of the initial convolutional layer can be directly determined according to the dimension of the training demodulation data, and the hyperparameters of the initial encoder and the hyperparameters of the initial decoder are all associated with the hyperparameters of the initial convolutional layer. After the training is completed, the initial model, the initial convolutional layer, the initial encoder, and the initial decoder correspond to the autoencoder model, the convolutional layer, the encoder, and the decoder, respectively.
[0081] In the embodiment of the present invention, the raw data collected by the security gate when executing the article detection method is referred to as detection raw data, the raw data collected by the security gate when pre-training the autoencoder model is referred to as training raw data, and the detection demodulated data and the training demodulated data, the detection convolutional features and the training convolutional features, and the detection low-dimensional features and the training low-dimensional features are similar. It can be understood that in the present invention, although the detection raw data is compared with the training raw data, the detection demodulated data is compared with the training demodulated data, the detection convolutional features are compared with the training convolutional features, and the detection low-dimensional features are compared with the training low-dimensional features, although the content may be different, there is no difference in data dimension and data length.
[0082] The features in the locally stored sample feature list are the sample features used to match the detection low-dimensional features. If the sample feature list is not empty, the similarity between each sample feature and the detection low-dimensional feature is first determined. Furthermore, if there is a feature in the sample feature list whose similarity with the detection low-dimensional feature meets the preset conditions, this feature can be regarded as a target feature that is relatively similar to the detection low-dimensional feature. In this case, the type of the item of the target feature can be used as the type of the item corresponding to the detection raw data.
[0083] When the sample feature list is empty, the corresponding label information is obtained by clustering the low-dimensional features of the detection, and the label information is used as the type of object corresponding to the original detection data.
[0084] When the similarity between the features in the sample feature list and the detection low-dimensional features does not meet the preset conditions, it means that there is no target feature that is relatively similar to the detection low-dimensional feature. At this time, the corresponding label information is obtained by clustering analysis of the detection low-dimensional features, and the label information is used as the type of object corresponding to the detection original data to obtain a detection result of the object type with higher accuracy.
[0085] It should be noted that, compared with the traditional CNN model, the autoencoder model and cluster analysis algorithm have the advantages of fast training speed and low requirements for computing resources and hardware. Even when there is no target feature that is relatively similar to the low-dimensional feature to be detected, cluster analysis is performed on the low-dimensional feature to be detected. This can significantly save model training time and training costs, speed up the speed of object detection, thereby improving the efficiency of object detection and the accuracy of object detection. In an embodiment of the present invention, the autoencoder model includes a convolutional layer, an encoder and a decoder that are cascaded in sequence. When processing the original detection data, only the convolutional layer and the encoder are used, and the data output by the encoder does not need to be input into the decoder for decoding. When training the autoencoder model, the decoder is required for verification.
[0086] The tag information obtained by clustering the detected low-dimensional features is used to characterize the item type. However, in the embodiments of the present invention, there are no specific restrictions on the item types. The specific item types may depend on the specific application scenario of the security door. For example, as a widely applicable implementation, the item types may include: mobile phones, knives, daily necessities, etc. There is no specific restriction on the format of the tag information. For example, the tag information can be a string of characters, a number, or a combination thereof.
[0087] In the object detection method provided by the embodiment of the present invention, the detection raw data collected by the security gate is processed to obtain detection demodulation data; the detection demodulation data is input into a pre-trained autoencoder model to obtain the detection low-dimensional features output by the encoder of the autoencoder model; the autoencoder model is used for the detection of the type of object, so there is no need to manually label the data when training the model, thus avoiding manual workload. By adding a convolutional layer before the encoder, the convolutional layer of the trained autoencoder model can extract detection convolution features from the detection demodulation data, thereby enhancing the feature expression capability and reducing the amount of parameter data and computational complexity. The encoder of the autoencoder model can map the detection convolution features to a low-dimensional space to obtain detection low-dimensional features, thereby achieving compression and dimensionality reduction of the detection convolution features and reducing storage and computational costs. The autoencoder model can learn to detect important features of the demodulated data and remove redundant information, thereby improving the accuracy of the detection results of the object type; by determining the similarity between the features in the sample feature list and the detection low-dimensional features when the locally stored sample feature list, which includes the mapping relationship between the type information of multiple objects and multiple sample features, is not empty, and when there is a target feature in the sample feature list whose similarity with the detection low-dimensional features meets the preset conditions, the type of the object of the target feature is used as the type of the object corresponding to the detection original data; and when the locally stored sample feature list is empty, and / or when the target feature does not exist in the sample feature list, cluster analysis is performed on the detection low-dimensional features to obtain corresponding label information, and the label information is used as the type of the object corresponding to the detection original data, which can significantly save model training time and training costs, speed up the speed of object detection, thereby improving the efficiency of object detection and improving the accuracy of object detection. At the same time, this method can also be used for a large number of security inspection doors in the same application scenarios, widely improving the efficiency and accuracy of object detection.
[0088] Detecting low-dimensional features can effectively reflect the attributes of items. If multiple items are of the same type, there will be certain similarities between the detected low-dimensional features. If the item types are different, there will be certain differences between the detected low-dimensional features. Mapping different feature categories of the detected low-dimensional features to different item types can quickly identify the corresponding item types. Accordingly, in some embodiments, such as Figure 3 As shown, the cluster analysis of the detected low-dimensional features to obtain corresponding label information can be performed as follows:
[0089] In step S210, clustering is performed on the detected low-dimensional features to obtain feature categories of the detected low-dimensional features;
[0090] In step S220, the label information for detecting the low-dimensional feature is determined based on the mapping relationship between the feature category and the label information for characterizing the type of the item.
[0091] In an embodiment of the present invention, as one of the ways to improve clustering accuracy and reduce clustering cost, a Gaussian mixture model can be used to perform cluster analysis on the detected low-dimensional features to obtain corresponding label information.
[0092] Accordingly, in some embodiments, in the step of performing cluster analysis on the detected low-dimensional features to obtain corresponding label information, the detected low-dimensional features are input into a pre-trained Gaussian mixture model, so that the pre-trained Gaussian mixture model performs the step of clustering the detected low-dimensional features and the step of determining the label information of the detected low-dimensional features based on the mapping relationship between the feature category and the label information used to characterize the item type; wherein,
[0093] The Gaussian mixture model includes multiple single Gaussian models, and the multiple single Gaussian models correspond one-to-one to multiple label information used to characterize the type of item.
[0094] Each single Gaussian model represents a feature category and has a label used to characterize the item type. The Gaussian mixture model calculates the posterior probability of the low-dimensional feature corresponding to each single Gaussian model. The low-dimensional feature is then associated with the single Gaussian model with the highest posterior probability, thereby determining the corresponding label used to characterize the item type.
[0095] In an embodiment of the present invention, the Gaussian mixture model can be trained based on the low-dimensional features of the sample, and the low-dimensional features of the sample can be extracted from training demodulated data with multiple item types using a pre-trained autoencoder model.
[0096] The Gaussian Mixture Model (GMM) can effectively model and fit complex data distributions, including multiple clusters, clusters of different shapes and sizes, and so on. Unlike hard clustering algorithms (such as K-means), GMM provides soft clustering capabilities, that is, the clustering result is the probability of assigning each data point to each category. This is more applicable to data points with fuzzy boundaries or overlapping areas, providing more detailed and flexible clustering results. Because GMM models the probability density function of each category, it can better handle the impact of outliers. Outliers usually have a large impact on the mean and variance. However, in GMM, each Gaussian distribution (i.e., each single Gaussian model) is represented by only a portion of the samples, so the impact of outliers is relatively small. In addition, the GMM model is a lightweight model that can be deployed on embedded devices or mobile terminals, and the detection results it provides are more real-time.
[0097] Furthermore, the inventors of the present invention have discovered that, in addition to the drawbacks of long training times and a high reliance on high-performance hardware, the CNN model suffers from another significant drawback: poor on-site flexibility. In some application scenarios, security gate operators will perform a secondary verification of the item type detection results. If the verification result is inaccurate, it indicates that a new type of item has likely appeared on-site. This requires the recollection and labeling of a large amount of data to retrain the CNN model, which takes a long time. During this period, the security gate can only rely on manual detection of the new item, and lacks the flexibility to respond on-site and the ability to quickly adapt and adjust.
[0098] In light of this, the inventors of this invention have further proposed that whenever a detection result correction instruction is received, the low-dimensional detection features and the correct object type information indicated in the detection result correction instruction can be stored in the sample feature list as sample features. In this way, when an item of the same type passes through the security gate again, it can be quickly detected.
[0099] Accordingly, in some embodiments, Figure 4 As shown, the method may further comprise the following steps:
[0100] In step S160, upon receiving the detection result correction instruction, obtaining the type information of the item carried in the detection result correction instruction;
[0101] In step S170, the type information of the object carried in the detection result correction instruction is added to the detection low-dimensional feature;
[0102] In step S180, the detected low-dimensional features after adding the type information of the object are used as sample features and stored in the sample feature list.
[0103] The embodiments of the present invention do not impose any special restrictions on how to receive the detection result correction instruction. For example, the security gate operator can perform a secondary check and send a detection result correction instruction to this end after discovering a detection error. It can also be carried out by connecting to other detection equipment to perform a secondary check and receive the detection result correction instruction sent by other detection equipment after identifying a detection error.
[0104] The object detection method provided by the embodiment of the present invention may occasionally cause detection errors during actual application. For example, the object carried by the object passing through the door is a mobile phone, but because the features in the locally stored sample feature list are not rich enough (that is, the sample features are insufficient) and / or the clustering analysis algorithm is not accurate enough, the detection result of the object type is determined to be a knife. At this time, it is only necessary to update the correct object type information and the detection low-dimensional features to the locally stored sample feature list. The next time the same mobile phone passes through the security gate, it can be quickly and correctly detected as a mobile phone.
[0105] In the object detection method provided in an embodiment of the present invention, upon receiving a detection result correction instruction, the method obtains the object type information contained in the detection result correction instruction; adds the object type information contained in the detection result correction instruction to the detection low-dimensional features; and stores the detection low-dimensional features, after adding the object type information, as sample features in the sample feature list. This method enables rapid adaptive adjustments without retraining the autoencoder model, and has strong generalization capabilities to cope with unexpected scenarios involving new types of objects.
[0106] When the total number of features in the sample feature list gradually increases, consideration may be given to sending the features in the sample feature list to the server to prevent the sample feature list from becoming too large and affecting the efficiency of object detection. Accordingly, in some embodiments, the method may further include the following step: when the total number of features in the sample feature list exceeds a first preset threshold, sending the features in the sample feature list to the server, so that the server can store the features in the sample feature list in a local clustering database according to the type information of the carried objects.
[0107] In which, the features in the clustering database are used to train the initial mixture model to obtain a trained Gaussian mixture model; when the total number of new features added by the server in the clustering database reaches a second preset number threshold, the initial mixture model can be re-trained according to the clustering database to obtain the latest trained Gaussian mixture model.
[0108] Since the newly added features in the clustering database are the features in the sample feature list received by the server, when the total number of newly added features in the clustering database reaches the second preset threshold, it means that the number of detection errors exceeds a certain number, that is, the number of new items exceeds a certain number. At this time, the server re-trains the initial mixture model based on the clustering database and can obtain a Gaussian mixture model with better clustering ability.
[0109] In the training process of the autoencoder model, the mean square error function is usually used as the loss function. In view of this, the inventors of the present invention further propose that the mean square error between the detected low-dimensional features and each feature in the locally stored sample feature list can be calculated to represent the similarity between the detected low-dimensional features and the sample features. Accordingly, in some embodiments, such as Figure 5 As shown, the determining of the similarity between the features in the sample feature list and the detected low-dimensional features may be performed as follows:
[0110] In step S310, the mean square error between the features in the sample feature list and the detection low-dimensional features is calculated, and the mean square error is used to represent the similarity;
[0111] The target feature is a feature whose mean square error with the detected low-dimensional feature is not greater than a preset similarity threshold.
[0112] It is understood that when calculating the mean square error (MSE) to represent similarity, a smaller MSE indicates a higher similarity, and conversely, a larger MSE indicates a lower similarity. Therefore, when the MSE is no greater than a preset similarity threshold, it indicates that the similarity is greater than a certain similarity threshold and can be considered sufficiently similar.
[0113] In addition, the inventors of the present invention have also discovered that the accuracy of the detection results of the item type is closely related to the original detection data collected by the security gate used for item detection.
[0114] like Figure 6 The figure shows the structure of the security door in the related art. The security door is usually composed of two door panels on the left and right and a top box. The two door panels on the left and right are respectively a transmitting door panel and a receiving door panel. A transmitting coil is set in the transmitting door panel, and a receiving coil is set in the receiving door panel. Key electronic components such as board power supply are usually set in the top box.
[0115] The inventors discovered that in related art, the transmitting coils of security gates typically use a fixed frequency. This means the signal emitted by the transmitting coil is fixed in frequency, and the receiving coil only needs to return a single received signal. However, different objects respond differently to signals of different frequencies. Some objects may not be sensitive to the fixed frequency used by the transmitting coil, making them difficult to detect.
[0116] In light of this, the inventors of the present invention further propose that the different responses of objects to signals of different frequencies are caused by their electromagnetic characteristics and material properties. Different objects have different frequency dependencies for the transmission, reflection, and absorption of electromagnetic waves. This frequency dependency allows the use of multiple frequency points in security gates to improve the accuracy of object type detection, as signals of different frequencies can reveal different characteristics and properties of an object. By simultaneously using multiple transmission frequencies and multiple, independently located receiving coils to feedback received signals for transmitted signals of different frequencies, the raw detection data collected by the security gate can cover a wider range of object characteristics and properties. Detection based on this raw detection data can improve the accuracy of the security gate's detection of various objects.
[0117] Accordingly, in some embodiments, the security gate is provided with at least a first number of transmitting coils and a second number of receiving coils which are positioned independently of each other; the first number of transmitting coils respectively use different frequencies; the original detection data collected by the security gate includes a first number of transmitting signals and a second number of receiving signals.
[0118] It should be noted that the second number of receiving coils are positioned independently of each other, so each receiving coil will feed back a receiving signal independently of each other, and therefore the number of received signals is also the second number.
[0119] In the present invention, there is no particular limitation on the first number and the second number. For example, as a specific implementation method that takes into account both cost savings and improved accuracy, the first number and the second number can be 3 and 8 respectively.
[0120] For example, the security door has 3 transmitting coils, each transmitting 3 signals of different frequencies, and 8 receiving coils, which feed back 8 receiving signals in total. Figure 7 As shown in FIG, a distribution diagram of a receiving coil of a receiving door panel provided by an embodiment of the present invention is provided, wherein the entire receiving door panel is evenly divided into 8 defense zones according to height ( Figure 8 There are 8 defense zones from top to bottom in the figure, and two parallel grids from left to right in each row represent one defense zone. Each defense zone uses an independent receiving coil to collect signals, that is, each defense zone has an independent inverted eight-shaped coil responsible for detection, thereby obtaining receiving signals from eight coils in the receiving door panel.
[0121] Setting up multiple independent defense zones and then setting up independent receiving coils within the defense zones has the following benefits: 1. Improving signal strength: By dividing the receiving coils into multiple defense zones, the received signal strength can be increased. The coil in each defense zone only needs to receive the signal at a specific location and reduces mutual interference with other coils. This separation can improve signal strength and stability and reduce signal attenuation caused by interference. 2. Reducing noise interference: When each defense zone uses an independent coil for detection, signal interference caused by environmental noise and interference can be reduced. Because each coil only focuses on the signal at a specific location and does not have to deal with interference sources at other locations, this can improve the signal-to-noise ratio between the signal and noise, making it easier to distinguish and analyze the target signal. 3. Improving anti-interference ability: Using independent coil detection in different zones can improve the system's anti-interference ability. When one coil is interfered with, the other coils can still operate normally. This redundant design can reduce the failure of the entire system due to the failure or interference of a single coil, thereby improving safety and reliability. 4. Optimizing signal demodulation: Independent coil detection makes it easier to demodulate and analyze the signal in each defense zone, which can improve the accuracy and efficiency of the signal processing algorithm, thereby improving the performance and response speed of the system.
[0122] In the object detection method provided by an embodiment of the present invention, at least a first number of transmitting coils and a second number of receiving coils, which are positioned independently of each other, are provided on the security inspection door; the first number of transmitting coils respectively use different frequencies; the raw data collected by the security inspection door includes a first number of transmission signals and a second number of receiving signals, thereby improving the feature and attribute representation capabilities of the raw detection data collected by the security inspection door, thereby further improving the accuracy of the detection results of the object type.
[0123] Here is a brief description of the process of collecting and detecting raw data at security inspection gates: When a person passes through a key-operated gate, the charged particles (such as electrons) in their body and the metal objects they carry will cut the magnetic field around the receiving coil, causing a change in the magnetic field. This change in the magnetic field will cause an induced electromotive force (electromotive force) to be generated in the receiving coil. The magnitude of the electromagnetic force (EMF) is related to the rate of change of the magnetic field and the material of the receiving coil. Since the movement speed of the object passing through the door is relatively slow and the number of charged particles in the body is small, the electromotive force induced in the receiving coil for the object passing through the door is relatively weak. Most of the induced electromotive force is the induced electromotive force generated by the metal objects carried by the object passing through the door cutting the magnetic field. In order to convert the induced electromotive force into a processable electrical signal, a series of amplification, filtering and conversion operations are required: first, the induced electromotive force is converted into a differential voltage signal by a differential amplifier. Then, the differential voltage signal is input into a preamplifier for further amplification and filtering. It can also be further amplified and filtered by a multi-stage amplifier to improve the signal-to-noise ratio and sensitivity. After the necessary signal processing, the signal is finally sent to the analog-to-digital converter (ADC) for digitization. The ADC samples the signal and converts it into a digital signal.
[0124] The resulting raw detection data is the continuous digital signals from the second number of receiving coils (referred to as received signals) and the digital signals from the first number of transmitting coils (referred to as transmitted signals). Each piece of raw detection data is stored in the (M, L) dimension, where M is the sum of the first and second numbers, indicating that each piece of raw detection data contains M channels of data, including the first number of transmitted signals and the second number of received signals. The first number of transmitted signals correspond to the first number of different frequencies. L represents the data length, which is related to the time it takes for the object to pass through the gate. When the ADC sampling frequency is fixed, faster transitions result in shorter L, and vice versa.
[0125] Understandably, the raw data collected by security gates often suffers from noise interference and poor quality. After the security gates collect the raw data, they need to process it to obtain demodulated data before it can be used as input for the pre-trained autoencoder model.
[0126] Accordingly, in some embodiments, the processing of the detection raw data collected by the security gate to obtain detection demodulation data can be performed as follows: the detection raw data collected by the security gate is sequentially subjected to bandpass filtering, downsampling, Hilbert transform, mixing, and low-pass filtering to obtain multi-channel vector data; wherein the vector data includes an in-phase component and an orthogonal component, and each channel of the vector data corresponds to one channel of the transmitted signal and one channel of the received signal; for each channel of the vector data, its amplitude and phase are determined based on the in-phase component and the orthogonal component included therein; and the detection demodulation data is determined based on the multi-channel vector data and its amplitude and phase.
[0127] It is understood that the transmitted signals and received signals corresponding to any two paths of vector data are not always identical (if they were identical, the vector data would be identical, resulting in data redundancy). The total number of the multiple paths of vector data is the product of the first number (of the transmitting coil) and the second number (of the receiving coil).
[0128] The processing of the raw data collected by security inspection gates described above primarily involves orthogonal demodulation. The goal is to extract the required information from the complex signal and make it easier to analyze and process. Through orthogonal demodulation, a signal with multiple frequency components can be converted into a set of mutually orthogonal basic waveforms. From these basic waveforms, the signal's amplitude and phase characteristics can be extracted, and redundant information can be removed, making it easier for the pre-trained autoencoder model to process the signal. The following briefly describes the execution of this processing.
[0129] The detected raw data, i.e., the raw signal, can be divided into two parts. One part includes the continuous digital signals of the second number of receiving coils, referred to as the receiving signal, and the other part includes the digital signals of the first number of transmitting coils, referred to as the transmitting signal or reference signal.
[0130] First, bandpass filtering is used to reduce noise interference in the received signal and reference signal, filtering out noise components that do not match the target signal's frequency, thereby improving the quality of the received and reference signals. Furthermore, downsampling is used to reduce the data volume, improve data processing efficiency, and conserve storage resources. Furthermore, a Hilbert transform is performed on the reference signal, which is equivalent to creating a 90-degree phase shift in the frequency domain. The resulting signal is called the envelope signal. The reference signal and the envelope signal correspond to the two orthogonal local oscillators (commonly known as the I and Q paths) in quadrature demodulation, where the I path is the reference signal and the Q path is the envelope signal. Furthermore, the I and Q paths are mixed (or multiplied) with the received signal. Furthermore, these two mixed signals are low-pass filtered to remove high-frequency components, retaining only the baseband signal (eliminating the influence of the transmitting coil and the effect of the gate object). This yields the in-phase component (I path) and the quadrature component (Q path), respectively. Finally, the in-phase component (I path) and the quadrature component (Q path) are respectively subtracted from the I path and Q path of the empty field to finally obtain IQ data, also called vector data.
[0131] Since there are a first number of reference signals and a second number of received signals, for each piece of detected original data, a total of vector data of the product of the first number and the second number is obtained. For example, when there are 3 reference signals and 8 received signals, a total of 24 vector data are obtained.
[0132] Then, the amplitude and phase of each vector data are calculated by the formula. The amplitude is The phase is actan(Q / I).
[0133] Finally, based on each channel of vector data and its amplitude and phase, detection and demodulation data with a dimension of (4, N, X) is obtained, where 4 represents the I channel, Q channel, amplitude, and phase, N represents a total of N channels of vector data, and X represents the data length of the vector data.
[0134] However, it should be noted that the data length of the detection raw data is related to the gate time (speed), so the data lengths of different detection raw data may be different. That is to say, although the data lengths of the multi-channel vector data obtained by processing one detection raw data are the same, the data lengths of different multi-channel vector data obtained by processing different detection raw data may also be different. However, the data input to the autoencoder model needs to maintain a fixed data length. In view of this, the inventors of the present invention propose that after processing the multi-channel vector data, it is also necessary to ensure that the length of the multi-channel vector data is the standard data length.
[0135] Accordingly, in some embodiments, before determining the amplitude and phase of each channel of vector data based on the in-phase component and quadrature component included therein, the method may further include the following steps: when the data length of the multi-channel vector data does not match the standard data length, truncating or padding the multi-channel vector data so that the data length of the multi-channel vector data is the standard data length; accordingly, determining the amplitude and phase of each channel of vector data based on the in-phase component and quadrature component included therein may include the following steps: when the data length of the multi-channel vector data is the standard data length, determining the amplitude and phase of each channel of vector data based on the in-phase component and quadrature component included therein.
[0136] In the present invention, there is no special limitation on how to determine the standard data length. For example, when pre-training the initial model to obtain the autoencoder model, the data lengths of the multi-channel vector data (that is, multiple groups of multi-channel vector data) corresponding to all the original training data are counted and sorted, and the data length of a certain ranking position is determined as the standard data length. For example, when sorting the groups of multi-channel vector data in order from short to long according to the data length, the ranking position of each group of multi-channel vector data increases from 0% to 100%, and the data length of the multi-channel vector data with a ranking position of 90% can be determined as the standard data length.
[0137] It is understandable that whether it is the detection demodulation data input into the pre-trained autoencoder model during object detection, or the training demodulation data input into the initial model during model training, the data length should remain consistent, for example, both should be maintained at the standard data length.
[0138] As mentioned above, by adding a convolutional layer before the encoder of the autoencoder model, it is possible to enhance the feature expression capability, reduce the amount of parameter data and computational complexity. Accordingly, in some embodiments, the detected demodulated data includes multiple channels of vector data, each channel of the vector data corresponds to one channel of the transmitted signal and one channel of the received signal, and the corresponding transmitted signals and received signals between any two channels of the vector data are not respectively the same;
[0139] In the step of inputting the detection demodulation data into a pre-trained autoencoder model to obtain the detection low-dimensional features output by the encoder of the autoencoder model, as Figure 8 As shown, the convolutional layer performs the following steps:
[0140] In step S410, the multi-channel vector data is sorted according to the corresponding received signals;
[0141] In step S420, in the sorted multi-channel vector data, each group of multi-channel vector data corresponding to the same received signal is sorted again according to the corresponding transmitted signal;
[0142] In step S430, convolution processing is performed on the re-sorted multi-channel vector data to obtain detection convolution features;
[0143] In step S440, the detection convolution feature is passed to the encoder, so that the encoder outputs the detection low-dimensional feature according to the detection convolution feature.
[0144] For example, when the dimension of the detection demodulation data is (4, 24, X) and the hyperparameters of the convolution layer are set to (in_channels = 4, out_channels = 32, kernel_size = 3, stride = 3, bias = False), the convolution layer arranges the 24-channel vector data of the detection demodulation data in the following order: three-channel vector data between the first received signal and the first, second, and third transmitted signals (i.e., a group of multi-channel vector data corresponding to the same received signal), three-channel vector data between the second received signal and the first, second, and third transmitted signals... three-channel vector data between the eighth received signal and the first, second, and third transmitted signals; then, convolution processing is performed through the convolution layer to integrate the frequency information of each data channel (i.e., each group of multi-channel vector data corresponding to the same received signal) to obtain a detection convolution feature with a dimension of (32, 8, X / 3), thereby enhancing the feature expression capability and reducing the amount of parameter data and computational complexity. Afterwards, the detection convolutional features with a dimension of (32, 8, X / 3) are passed to the encoder so that the encoder can detect low-dimensional features based on the detection convolutional feature output.
[0145] As a second aspect of an embodiment of the present invention, a training method for a Gaussian mixture model is also provided, wherein, Figure 9 As shown, the method includes multiple iterative processes, and each iterative process may include the following steps:
[0146] In step S510, the sample low-dimensional features are used as training data to train the initial mixture model of the current iteration process to obtain the fitting mixture model of the current iteration process;
[0147] In step S520, a clustering result of the low-dimensional features of the sample is determined according to the fitted mixture model of the current iteration process;
[0148] In step S530, the clustering accuracy of the current iteration process is determined based on the clustering results and the label information carried by the sample low-dimensional features;
[0149] In response to the end of each iteration process, the method may further include the following steps:
[0150] In step S540, when the clustering accuracy of the current iteration process is less than or equal to the clustering accuracy of the previous iteration process, the fitted mixture model of the previous iteration process is used as the trained Gaussian mixture model, and the iteration is terminated;
[0151] The Gaussian mixture model includes multiple single Gaussian models, each of which corresponds one-to-one to multiple label information used to characterize the type of item. The Gaussian mixture model can output one of the multiple label information as the type of item corresponding to the original detection data collected by the security gate by calculating the posterior probability of the input detection low-dimensional feature corresponding to each single Gaussian model. The low-dimensional detection feature is output by the encoder of a pre-trained autoencoder model based on the detection demodulation data input to the autoencoder model, and the detection demodulation data is obtained by processing the original detection data.
[0152] In step S550, when the clustering accuracy of the current iterative process is greater than the clustering accuracy of the previous iterative process, the number of initial single models included in the initial mixed model of the current iterative process is increased to obtain the initial mixed model of the next iterative process, and the next iterative process is executed.
[0153] In the Gaussian mixture model training method provided in the second aspect of the embodiments of the present invention, for the sake of distinction, the low-dimensional features used to train the Gaussian mixture model are referred to as sample low-dimensional features, and the low-dimensional features to which the trained Gaussian mixture model is applied are referred to as detection low-dimensional features, both of which are distinct from the training low-dimensional features used in the autoencoder model training process. The sample low-dimensional features can be extracted from training demodulated data of multiple item types using a pre-trained autoencoder model.
[0154] In an embodiment of the present invention, multiple training raw databases can be pre-established. A large amount of historical raw data of different item types collected by security inspection gates is stored as training raw data in the training raw databases. All training raw data stored in each training raw database carries the same item type information. All training raw data in each training raw database is processed to obtain corresponding training demodulated data, thereby establishing multiple training demodulated databases. All training demodulated data stored in each training demodulated database carries the same item type information. For each training demodulated database, all training demodulated data stored therein is input into a pre-trained autoencoder model to obtain sample low-dimensional features output by the autoencoder model, thereby establishing multiple clustering databases. All sample low-dimensional features stored in each clustering database carry the same item type information. In this way, the features in these multiple clustering databases can be used to train an initial mixture model to obtain a Gaussian mixture model.
[0155] In an embodiment of the present invention, when the initial mixed model of the iterative process includes multiple initial single models, the fitted mixed model of the iterative process obtained through the above step S510 then includes multiple fitted single models. However, after obtaining the fitted mixed model of the iterative process, it is only clear that each fitted single model represents a cluster, and it is not clear what specific category the cluster represented by each fitted single model is. At this time, multiple fitted single models can be matched one-to-one with multiple label information used to characterize item types.
[0156] Specifically, the fitted mixed model obtained in the current iterative process can be used to cluster the low-dimensional features of the samples to obtain the clustering results of the low-dimensional features of the samples; since the type information of the items carried by all the low-dimensional features of the samples is known, all possible situations are traversed, and the accuracy corresponding to each possible situation is calculated. The highest accuracy is used as the clustering accuracy of the current iterative process, which is used to make a one-to-one correspondence between multiple fitted single models and multiple label information used to characterize the type of items.
[0157] For example, each sample low-dimensional feature can be uniquely mapped to one of the fitted single models. When the fitted mixed model of this iterative process includes three fitted single models, the fitted single model 1, the fitted single model 2 and the fitted single model 3 respectively cluster the sample low-dimensional features of the first category, the sample low-dimensional features of the second category and the sample low-dimensional features of the third category, and the label information includes mobile phones, knives and daily necessities. Then a total of six possible situations are traversed: ① The first category, the second category and the third category are mobile phones, knives and daily necessities respectively; ② The first category, the second category and the third category are mobile phones, daily necessities and knives respectively; ③ The first category, the second category and the third category are knives, daily necessities and mobile phones respectively; ④ The first category, the second category and the third category are knives, mobile phones and daily necessities respectively; ⑤ The first category, the second category and the third category are daily necessities, mobile phones and knives respectively; ⑥ The first category, the second category and the third category are daily necessities, knives and mobile phones respectively. The accuracy of each possible situation is calculated. For example, the accuracy of the first possible situation is calculated by the following formula: (the total number of sample low-dimensional features of the items carried in the first category whose type information is mobile phones + the total number of sample low-dimensional features of the items carried in the second category whose type information is knives + the total number of sample low-dimensional features of the items carried in the third category whose type information is daily necessities) / (the total number of sample low-dimensional features in the first category + the total number of sample low-dimensional features in the second category + the total number of sample low-dimensional features in the third category).
[0158] For another example, the accuracy of the fifth possible case mentioned above is the highest. This highest accuracy is used as the clustering accuracy of the current iterative process. Then, the label information of daily necessities, mobile phones and knives can be added to the above-mentioned fitting single model 1, fitting single model 2 and fitting single model 3 in sequence, and multiple fitting single models can be successfully matched one by one with multiple label information used to characterize the types of items.
[0159] It is understandable that when executing the above-mentioned step S540, when the clustering accuracy of the current iterative process is less than or equal to the clustering accuracy of the previous iterative process, the fitting mixed model of the previous iterative process is used as the trained Gaussian mixture model, and the multiple fitting single models included in the fitting mixed model of the previous iterative process are respectively used as the multiple single Gaussian models included in the trained Gaussian mixture model. In order to save computing resources, corresponding label information can be added to the multiple single Gaussian models according to the clustering accuracy of the previous iterative process at the end of the iteration. Of course, corresponding label information can also be added to the multiple fitting single models included in the fitting mixed model obtained in the current iterative process according to the clustering accuracy of the current iterative process each time step S530 is executed.
[0160] It can be understood that in response to the end of the first iteration process, since there is no "clustering accuracy of the previous iteration process", the "clustering accuracy of the previous iteration process" can be initialized to 0, and then the clustering accuracy of the first iteration process is bound to be greater than the clustering accuracy of the previous iteration process, and the above step S550 will be executed.
[0161] The inventors of the present invention propose that although a fitted mixture model is obtained through the above-mentioned step S510, the clustering ability of the fitted mixture model does not necessarily reach the optimal limit. For example, although the fitted mixture model can cluster the low-dimensional features of all samples into three categories: mobile phones, knives, and daily necessities, knives are also divided into different types. For example, swords and knives are both knives, but swords are double-edged and knives are single-edged. If the fitted mixture model can further distinguish between these two different types of knives, it will obviously have better object type detection capabilities and can obtain more accurate object type detection results. Therefore, by executing the above-mentioned step S550, the number of initial single models included in the initial mixture model of the current iteration process is increased to obtain the initial mixture model of the next iteration process, and the next iteration process is executed to attempt to obtain a fitted mixture model with optimal clustering ability as the trained Gaussian mixture model.
[0162] Accordingly, in the above step S540, when the clustering accuracy of the current iteration process is less than or equal to the clustering accuracy of the previous iteration process, it indicates that the clustering ability of the fitting mixture model obtained in the previous iteration process has reached the optimal limit.
[0163] In the training method of the Gaussian mixture model provided in an embodiment of the present invention, by executing step S540 or step S550 at the end of each iterative process, i.e., steps S510-S530, a fitted mixture model with optimal clustering ability can be obtained as a trained Gaussian mixture model, thereby significantly improving the accuracy of the detection results of the object type.
[0164] In some embodiments, the step of using the sample low-dimensional features as training data to train the initial mixture model of the current iteration process to obtain the fitted mixture model of the current iteration process (ie, step S510) can be performed as follows:
[0165] Initializing the parameters of a plurality of initial single models included in the initial mixed model of the current iteration process; wherein the parameters of the initial single model include a mean, a covariance matrix, and a mixing coefficient, and the mixing coefficient represents the weight of the corresponding initial single model;
[0166] According to the expectation maximization algorithm (EM) and the low-dimensional features of the sample, a fitting calculation is performed on the initialized initial mixture model of the current iterative process to obtain the fitted mixture model of the current iterative process; the fitting calculation may specifically include the following steps:
[0167] Utilizing the probability density function and Bayesian theorem, calculating the posterior probability of the sample low-dimensional features corresponding to each initialized initial single model;
[0168] updating the parameters of the current multiple initial single models according to the calculated posterior probabilities;
[0169] When the updated initial mixed model reaches the convergence condition, the fitting is stopped and the updated initial mixed model is used as the fitted mixed model of the current iteration process;
[0170] If the updated initial hybrid model does not meet the convergence condition, jump to the step of using the probability density function and Bayes' theorem to calculate the posterior probability of the sample low-dimensional features corresponding to each initialized initial single model.
[0171] The convergence condition may include: a parameter change of the updated initial hybrid model is less than a preset threshold, a preset number of iterations is reached, and the like.
[0172] As described above, in the object detection method provided in the first aspect of the embodiment of the present invention and the Gaussian mixture model training method provided in the second aspect of the embodiment of the present invention, a pre-trained autoencoder model is used. Accordingly, as a third aspect of the embodiment of the present invention, a training method for an autoencoder model is also provided, wherein, Figure 10 As shown, the method may include the following steps:
[0173] In step S610, the original training data collected by the security gate is processed to obtain demodulated training data;
[0174] In step S620, the initial model is trained using the acquired training demodulated data as training data to obtain a trained autoencoder model, wherein the initial model includes an initial convolutional layer, an initial encoder, and an initial decoder that are cascaded in sequence, and the hyperparameters of the initial encoder and the hyperparameters of the initial decoder are associated with the hyperparameters of the initial convolutional layer. The encoder of the trained autoencoder model can output and detect low-dimensional features.
[0175] Specifically, how to process the original training data collected by the security inspection door to obtain the training demodulated data is consistent with the principle of processing the original detection data collected by the security inspection door to obtain the detection demodulated data mentioned above, and will not be repeated here.
[0176] Among them, after inputting the detection demodulation data into the trained autoencoder model, the encoder of the trained autoencoder model can output detection low-dimensional features, based on which the detection results of the item type of the original detection data collected by the security gate can be further obtained.
[0177] The hyperparameters of the initial convolutional layer can be directly determined according to the dimension of the training demodulated data, while the hyperparameters of the initial encoder and the hyperparameters of the initial decoder are both related to the hyperparameters of the initial convolutional layer. For example, when the dimension of the training demodulated data is (4, 24, X), the hyperparameters of the initial convolutional layer can be set to (in_channels = 4, out_channels = 32, kernel_size = 3, stride = 3, bias = False). In the initial convolutional layer, the 24 vector data of the training demodulated data are first arranged in the following order: three-way vector data between the first received signal and the first, second, and third transmitted signals (i.e., a group of multi-way vector data corresponding to the same received signal), three-way vector data between the second received signal and the first, second, and third transmitted signals, and... three-way vector data between the eighth received signal and the first, second, and third transmitted signals. Then, feature extraction is performed through the initial convolutional layer, integrating the frequency information of each data channel (i.e., each group of multi-way vector data corresponding to the same received signal) to obtain initial convolution features of dimension (32, 8, X / 3). This enhances feature expression capability and reduces the amount of parameter data and computational complexity.
[0178] In an embodiment of the present invention, an autoencoder model is obtained based on the initial model training, and the initial model includes an initial convolutional layer, an initial encoder and an initial decoder that are cascaded in sequence. During the training process, the initial convolutional layer performs convolution processing on the training demodulated data to obtain training convolution features; the initial encoder performs encoding processing on the training convolution features to obtain training low-dimensional features; the initial decoder performs decoding processing on the training low-dimensional features to obtain output features; according to the loss function between the input features (i.e., the training demodulated data input to the initial model) and the output features, the parameters of the initial convolutional layer, the initial encoder and the initial decoder are updated. The hyperparameters of the initial convolutional layer can be directly determined according to the dimension of the training data, and the hyperparameters of the initial encoder and the hyperparameters of the initial decoder are all associated with the hyperparameters of the initial convolutional layer. After the training is completed, the initial model, the initial convolutional layer, the initial encoder and the initial decoder correspond to the autoencoder model, the convolutional layer, the encoder and the decoder respectively.
[0179] It should be noted that the convolution operation can not only process image problems, but also data problems (such as financial data). It is just that it currently performs better in image processing. In addition, the demodulated data obtained by the embodiment of the present invention based on the original data processing (whether it is training demodulation data or detection demodulation data) is consistent with the image data in dimension (the dimension of image data is (C, H, W) and the dimension of demodulation data is (4, 24, X)). Therefore, the convolution layer can be reasonably used to perform convolution processing on the demodulated data to obtain convolution features.
[0180] In the training method of the autoencoder model provided in the embodiment of the present invention, the training raw data collected by the security gate is processed to obtain training demodulated data; the obtained training demodulated data is used as training data to train the initial model to obtain a trained autoencoder model, wherein the initial model includes an initial convolution layer, an initial encoder and an initial decoder cascaded in sequence, and the encoder of the trained autoencoder model can output detection low-dimensional features. It can be seen that there is no need to manually label the data when training the model, which directly avoids manual workload; by adding an initial convolution layer before the initial encoder of the initial model, the convolution layer of the trained autoencoder model can extract detection convolution features from the detection demodulated data, thereby enhancing feature expression capabilities, reducing the amount of parameter data and computational complexity; the encoder of the trained autoencoder model can map the detection convolution features to a low-dimensional space to obtain detection low-dimensional features, thereby achieving compression and dimensionality reduction of the detection convolution features and reducing storage and computational costs; the autoencoder model can learn important features of the detection demodulated data and remove redundant information. Subsequently, the autoencoder model is applied to the object detection method to improve the accuracy of the object type detection results.
[0181] At the same time, the training method of the autoencoder model provided by the third aspect of the present invention, the training method of the Gaussian mixture model improved by the second aspect of the present invention, and the object detection method provided by the first aspect of the present invention can also be used for different security gate systems and customized and optimized, thereby widely improving the detection efficiency and the accuracy of the detection results.
[0182] In some embodiments, the training demodulated data carries item type information, and the training method of the autoencoder model may further include the following steps:
[0183] Processing the training demodulated data according to the convolutional layer and encoder of the trained autoencoder model to obtain low-dimensional features of the samples;
[0184] According to the low-dimensional features of each sample and the type information of the items it carries, the low-dimensional features of each sample are stored in a clustering database of the corresponding item type information; wherein the low-dimensional features of the samples in each clustering database are used to train the initial mixture model to obtain a trained Gaussian mixture model.
[0185] In the present invention, there is no special limitation on the types of items. The specific types of items may depend on the specific application scenarios of the security door. For example, as a widely applicable implementation method, the types of items may include knives, mobile phones and daily necessities.
[0186] As a specific implementation, multiple training raw databases can be pre-established. A large amount of historical raw data of different item types collected by security inspection gates is stored as training raw data in these databases. All training raw data stored in each training raw database carries the same item type information. All training raw data in each training raw database is processed to obtain corresponding training demodulated data, thereby establishing multiple training demodulated databases. All training demodulated data stored in each training demodulated database carries the same item type information. The initial model is trained based on the training demodulated data in each training demodulated database to obtain a trained autoencoder model.
[0187] Then, for each training demodulation database, all the training demodulation data stored therein is input into a pre-trained autoencoder model to obtain the sample low-dimensional features output by the autoencoder model's encoder, thereby establishing multiple cluster databases; all the sample low-dimensional features stored in each cluster database carry the same item type information. In this way, the features in these multiple cluster databases can be used to train the initial mixture model to obtain a Gaussian mixture model.
[0188] Here, sample low-dimensional features are extracted directly from the training demodulated data used for training the autoencoder model to train the Gaussian mixture model. However, it should be noted that the present invention is not limited to this. Sample low-dimensional features can also be extracted from other sample data other than the training demodulated data to train the Gaussian mixture model, as long as the other sample data is consistent with the training demodulated data in terms of other irrelevant variables such as data dimension and data length.
[0189] By simultaneously using multiple transmitting frequencies and multiple independently positioned receiving coils to feed back received signals for transmitted signals of different frequencies, the raw training data collected by the security gate can cover a wider range of object features and attributes. Training the model based on this raw training data can enable the autoencoder model to learn more comprehensive feature representations.
[0190] Accordingly, in some embodiments, the security gate is provided with at least a first number of transmitting coils and a second number of receiving coils that are independently positioned from each other; the first number of transmitting coils respectively use different frequencies; and the raw training data collected by the security gate includes the first number of transmitting signals and the second number of receiving signals.
[0191] The training demodulation data includes multiple channels of vector data and amplitudes and phases thereof. The multiple channels of vector data have the same data length, and each channel of the vector data corresponds to one channel of the transmitted signal and one channel of the received signal.
[0192] It should be noted that the second number of receiving coils are positioned independently of each other, so each receiving coil will feed back a receiving signal independently of each other, and therefore the number of received signals is also the second number.
[0193] In the present invention, there is no particular limitation on the first number and the second number. For example, as a specific implementation method that takes into account both cost savings and improved accuracy, the first number and the second number can be 3 and 8 respectively.
[0194] Signals of different frequencies can reveal different characteristics and properties of an object. By simultaneously using multiple frequencies and independently positioned receiving coils to provide feedback for transmitted signals at different frequencies, the raw training data collected by the security gate can cover a wider range of object characteristics and properties. Training the model based on this raw training data allows the autoencoder model to learn more effective feature representations.
[0195] In addition, the raw training data covers more application scenarios and can also enable the autoencoder model to learn more effective feature representations. Accordingly, in some embodiments, the raw training data corresponds to multiple item types, multiple item location distribution states, multiple collection speeds, and multiple item quantity states.
[0196] Multiple item location distribution states mean that the distribution of items carried by a person passing through the security gate can vary; multiple collection speeds mean that the speed at which items carried by a person passing through the security gate can vary; and multiple item quantity states mean that the number of items carried by a person passing through the security gate can vary. In short, the amount of raw training data must be sufficient, the types of items covered in the raw training data must meet the application scenarios of the security gate, and the raw training data must cover most variables such as collection speed, item location distribution, and item quantity.
[0197] The following reference Figure 11 The training method of the autoencoder model provided in the embodiment of the present invention is described in detail.
[0198] like Figure 11As shown, it is a structural diagram of an initial model and an autoencoder model provided by an embodiment of the present invention. The training demodulated data is input to the initial convolution layer (Conv) through the input layer (Input). On the initial convolution layer, the training demodulated data is convolved to obtain training convolution features. The training convolution features are passed to the initial encoder (Encoder (ViT)). The initial encoder compresses the training convolution features into a lower-dimensional feature representation. On the initial encoder, the training convolution features are encoded to obtain training low-dimensional features (which can also be directly referred to as encoding). The output of the initial encoder, i.e., the training low-dimensional features, is passed to the initial decoder (Decoder (ViT)) as input. The initial decoder corresponds to the initial encoder, and the number of neurons in the last layer of the initial decoder matches the dimension of the input layer. The initial decoder gradually increases the dimension of the training low-dimensional features, and tries to restore the training low-dimensional features to the input (i.e., the training demodulated data). On the initial decoder, the training low-dimensional features are decoded to obtain output features (which can also be directly referred to as decoding).
[0199] During the training process, the goal of the autoencoder model is to minimize the difference between the input (i.e., training demodulated data) and the output features. The mean square error (MSE) function can be used as the loss function to measure this difference. The parameters of the initial convolutional layer, initial encoder, and initial decoder of the initial model are updated through backpropagation and gradient descent to minimize the difference. The training process is iterated multiple times until the preset termination training conditions are reached, such as reaching the maximum number of iterations, the loss function converges, and so on.
[0200] Through training, eventually, the encoder learns the low-dimensional representation of the convolutional features, and the decoder learns how to reconstruct the input data.
[0201] As can be seen, the backbone of the autoencoder model provided by the embodiment of the present invention, namely the encoder and decoder, uses the ViT (Vision Transformer) structure. ViT is a graph recognition model based on the Transformer architecture. Although the raw data collected by the security gate in the present invention is essentially a digital signal and not image data, the demodulated data (whether training demodulated data or detection demodulated data) obtained by processing the raw data in the embodiment of the present invention is consistent in dimension with the image data, and the autoencoder model can be used to process the demodulated data.
[0202] The dimensions of image data are (C, H, W), where C represents the three channels of the image, including R, G, and B. RGB refers to the color model of the image, which stands for red, green, and blue, respectively. This is a method used to describe and represent color images. In the RGB color model, the color of each pixel is represented by the intensity values of the three components red, green, and blue; H represents the height of the image; and W represents the width of the image.
[0203] As described above, the dimensions of the original data are (M, L), where M is the sum of the first and second quantities, indicating that each piece of original data contains M channels of data, including the first number of received signals and the second number of transmitted signals. L represents the data length of the received and transmitted signals. After processing the original data, the dimensions of the demodulated data are (4, N, X), where 4 represents the I channel, Q channel, amplitude, and phase. Each piece of demodulated data contains multiple channels of vector data. N represents the total number of vector data channels, and X represents the data length of the vector data.
[0204] It can be seen that in terms of dimension, 4 (I path, Q path, amplitude, phase) of the demodulated data can correspond to C (R, G, B) of the image data, N of the demodulated data can correspond to H of the image data, and X of the demodulated data can correspond to W of the image data, that is, the demodulated data is consistent with the image data in dimension.
[0205] The following reference Figure 12 The object detection method provided by the first aspect of the present invention is described in detail.
[0206] like Figure 12 As shown, it is a structural diagram of a detection system provided in an embodiment of the present invention. The detection system may include a signal acquisition module, a signal processing module, an encoding module, a clustering module and a sample feature list, which is used to execute the object detection method provided in the first aspect of the embodiment of the present invention.
[0207] In the signal acquisition module, signals of the objects passing through the security door and the items they carry are collected to obtain receiving signals and transmitting signals from the receiving coil and transmitting coil. In other words, the signal acquisition module can be the receiving coil and transmitting coil of the security door, or it can be a door panel or security door including the receiving coil and transmitting coil.
[0208] In the signal processing module, the raw detection data collected by the security gate is processed to obtain detection demodulated data. Specifically, the raw signal collected by the signal acquisition module undergoes orthogonal demodulation and other processing. The purpose is to extract the required information from the complex signal and make it easier to analyze and process. Through orthogonal demodulation, a signal with multiple frequency components can be converted into a set of mutually orthogonal basic waveforms. From this basic waveform, the signal's amplitude and phase characteristics can be extracted, and redundant information can be removed, making it easier for the pre-trained autoencoder model to process the signal.
[0209] In the encoding module, the detection and demodulation data is input into a pre-trained autoencoder model to obtain the detection low-dimensional features output by the autoencoder model's encoder. The autoencoder model includes a cascade of convolutional layers, an encoder, and a decoder. In other words, the encoding module includes at least the convolutional layers and encoder of the pre-trained autoencoder model.
[0210] In the clustering module, when the locally stored sample feature list is empty and / or when the target feature does not exist in the sample feature list, a clustering analysis algorithm (such as a pre-trained Gaussian mixture model) is used to distinguish the detection low-dimensional features corresponding to different item types.
[0211] In the sample feature list, if there is a target feature whose similarity with the detection low-dimensional feature meets the preset conditions, the type of the object of the target feature is used as the type of the object corresponding to the detection original data; and most importantly, whenever a detection result correction instruction is received, the type information of the object carried in the detection result correction instruction is obtained, the type information of the object is added to the detection low-dimensional feature, and the detection low-dimensional feature after the type information of the object is added is used as a sample feature and stored in the sample feature list.
[0212] The following reference Figure 13 This article will provide an overall description of the object detection method, the training method of the Gaussian mixture model, and the training method of the autoencoder model provided in the embodiments of the present invention.
[0213] like Figure 13 The figure shows a collaborative process diagram provided by an embodiment of the present invention. The security gate performs object detection, and the server performs Gaussian mixture model training and autoencoder model training.
[0214] The server side obtains a large amount of training original data collected by the security gate, stores it into the training original database according to the type information of the items carried, processes the data in the training original database, obtains training demodulation data and stores it into the training demodulation database according to the type information of the items carried; trains the initial model based on the data in the training demodulation database to obtain a trained autoencoder model; uses the trained autoencoder model to encode the data in the training demodulation database to obtain low-dimensional features of the samples; stores the type information of the items carried by the sample low-dimensional features into the clustering database; trains the initial hybrid model based on the data in the clustering database to obtain a trained GMM model.
[0215] When someone passes through the security gate, the security gate collects the original detection data, processes it to obtain detection demodulation data, and uses the pre-trained autoencoder model to encode the detection demodulation data to obtain detection low-dimensional features.
[0216] If a target feature exists in the sample feature list and its similarity with the low-dimensional feature being tested meets a preset condition, the type of the item corresponding to the target feature is used as the type of the item corresponding to the original data being tested. The target feature is a feature whose mean square error with the low-dimensional feature being tested is no greater than a preset similarity threshold.
[0217] When the locally stored sample feature list is empty, and / or when the target feature does not exist in the sample feature list, the detected low-dimensional feature is input into a pre-trained GMM model to obtain the label information output by the GMM model, and the label information is used as the detection result of the object type corresponding to the original data.
[0218] Upon receiving a detection result correction instruction, the security gate stores the detected low-dimensional features and the type of item included in the detection result correction instruction in the sample feature list. Furthermore, when the amount of data in the sample feature list reaches a threshold, the security gate sends the features in the sample feature list to the server, which then stores the received features in the clustering database. When the number of newly added features in the clustering database reaches a threshold, the server retrains the initial hybrid model based on the clustering database to obtain the latest trained GMM model, which is then used to update the GMM model at the security gate.
[0219] As a second aspect of an embodiment of the present invention, an electronic device is provided, wherein, Figure 14 As shown, the electronic device includes:
[0220] One or more processors 101;
[0221] The memory 102 stores one or more computer programs. When the one or more computer programs are executed by the one or more processors 101, the one or more processors 101 implement the object detection method provided in the first aspect of the embodiment of the present invention.
[0222] The electronic device may further include one or more I / O interfaces 103 connected between the processor 101 and the memory 102 and configured to implement information exchange between the processor 101 and the memory 102 .
[0223] Among them, the processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) is connected between the processor and the memory, and can realize information exchange between the processor and the memory, including but not limited to a data bus (Bus), etc.
[0224] In some embodiments, the processor 101 , the memory 102 , and the I / O interface 103 are connected to each other via a bus 104 , and further connected to other components of the computing device.
[0225] like Figure 15 As shown, as a third aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the object detection method provided by the first aspect of the embodiment of the present invention is implemented.
[0226] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. Accordingly, the computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can implement the method of any of the above-mentioned embodiments. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the embodiments of the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0227] The above are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Those skilled in the art should understand that the present invention includes but is not limited to the contents described in the drawings and the above specific embodiments. Any modifications that do not deviate from the functional and structural principles of the present invention are intended to be included within the scope of the claims.
Claims
1. A method for detecting an object, characterized in that: The object detection method comprises: Process the original detection data collected by the security door to obtain detection demodulation data; Inputting the detected demodulated data into a pre-trained autoencoder model to obtain the detected low-dimensional features output by the encoder of the autoencoder model; wherein the autoencoder model includes a convolutional layer, the encoder, and a decoder cascaded in sequence, and the hyperparameters of the encoder and the hyperparameters of the decoder are associated with the hyperparameters of the convolutional layer; When the locally stored sample feature list is not empty, determining the similarity between the features in the sample feature list and the detection low-dimensional features; wherein the sample feature list includes a mapping relationship between type information of multiple items and multiple sample features; When a target feature exists in the sample feature list, the type of the item of the target feature is used as the type of the item corresponding to the detection original data, wherein the target feature is a feature whose similarity with the detection low-dimensional feature meets a preset condition; When the locally stored sample feature list is empty, and / or when the target feature does not exist in the sample feature list, cluster analysis is performed on the detected low-dimensional features to obtain corresponding label information, and the label information is used as the type of the object corresponding to the original detection data.
2. The object detection method according to claim 1, characterized in that: The cluster analysis of the detected low-dimensional features to obtain corresponding label information includes: Clustering the detected low-dimensional features to obtain feature categories of the detected low-dimensional features; The label information for detecting the low-dimensional feature is determined based on a mapping relationship between the feature category and the label information used to characterize the type of the item.
3. The object detection method according to claim 2, characterized in that: In the step of performing cluster analysis on the detected low-dimensional features to obtain corresponding label information, the detected low-dimensional features are input into a pre-trained Gaussian mixture model so that the pre-trained Gaussian mixture model performs the step of clustering the detected low-dimensional features and the step of determining the label information of the detected low-dimensional features based on the mapping relationship between the feature category and the label information used to characterize the item type; wherein, The Gaussian mixture model includes multiple single Gaussian models, and the multiple single Gaussian models correspond one-to-one to multiple label information used to characterize the type of item.
4. The object detection method according to any one of claims 1 to 3, characterized in that: The object detection method further includes: Upon receiving a test result correction instruction, obtaining type information of the item included in the test result correction instruction; Adding the type information of the item carried in the detection result correction instruction to the detection low-dimensional feature; The detected low-dimensional features after adding the type information of the object are used as sample features and stored in the sample feature list.
5. The object detection method according to claim 4, characterized in that: The object detection method further includes: When the total number of features in the sample feature list is greater than a first preset number threshold, the features in the sample feature list are sent to the server so that the server can store the features in the sample feature list in a local clustering database according to the type information of the items carried.
6. The object detection method according to any one of claims 1 to 3, characterized in that: Determining the similarity between the features in the sample feature list and the low-dimensional features includes: Calculating the mean square error between the features in the sample feature list and the detection low-dimensional features, wherein the mean square error is used to represent the similarity; The target feature is a feature whose mean square error with the detected low-dimensional feature is not greater than a preset similarity threshold.
7. The object detection method according to claim 1, characterized in that: The security door is provided with at least a first number of transmitting coils and a second number of receiving coils which are positioned independently of each other; the first number of transmitting coils respectively use different frequencies; the original detection data collected by the security door includes a first number of transmitting signals and a second number of receiving signals.
8. The object detection method according to claim 7, characterized in that: The detected demodulated data includes multiple channels of vector data, each channel of the vector data corresponds to one channel of the transmitted signal and one channel of the received signal, and the transmitted signal and the received signal corresponding to any two channels of the vector data are not respectively the same; In the step of inputting the detection demodulation data into a pre-trained autoencoder model to obtain the detection low-dimensional features output by the encoder of the autoencoder model, the convolutional layer performs the following steps: Sorting the multi-channel vector data according to the corresponding received signals; Among the sorted multi-channel vector data, each group of multi-channel vector data corresponding to the same received signal is sorted again according to the corresponding transmitted signal; Performing convolution processing on the re-sorted multi-channel vector data to obtain detection convolution features; The detection convolution feature is passed to the encoder, so that the encoder outputs the detection low-dimensional feature according to the detection convolution feature.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A memory having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement the object detection method according to any one of claims 1 to 8.
10. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the object detection method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
A pattern recognition method based on a variable sample stack type self-coding network
CN109558873A
Article identification method, device and equipment
CN113344012A
Deep semi-supervised learning network intrusion detection method based on self-supervised variational LSTM
CN113569243A
Data processing method and device, electronic equipment and storage medium
CN114707174A
Community discovery method and system based on improved depth sparse automatic encoder
CN115049008A