Inspection method, inspection equipment, and computer-readable medium

By employing deep learning techniques, including CNNs for image analysis and BRNNs for natural language processing, the security inspection process is streamlined, addressing inefficiencies and inaccuracies in current methods and enhancing the detection of prohibited items.

JP7678661B2Active Publication Date: 2025-05-16NUCTECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2019565493
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-09-18
Filing Date
2018-09-17
Publication Date
2025-05-16
Estimated Expiration
2038-09-17

AI Technical Summary

Technical Problem

Current security inspection methods in the distribution industry, such as those used in the 'National National Online Shopping' boom and the 'Belt and Road' policy, face challenges in efficiently and accurately verifying the contents of packages against declared information, often leading to resource wastage and inspector burden.

Method used

The proposed solution involves using deep learning-based methods, specifically convolutional neural networks (CNNs) for image processing and bidirectional recurrent neural networks (BRNNs) for natural language processing, to automatically inspect packages by comparing X-ray images with declared information, ensuring accurate and efficient security inspections.

Benefits of technology

This approach significantly improves the accuracy and speed of security inspections, reducing the reliance on manual methods and minimizing the risk of prohibited items being undetected, thereby enhancing the efficiency and reliability of security checks in the distribution industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007678661000001
    Figure 0007678661000001
  • Figure 0007678661000002
    Figure 0007678661000002
  • Figure 0007678661000003
    Figure 0007678661000003
Patent Text Reader

Abstract

An inspection method, inspection equipment, and computer-readable medium are provided. The method includes: scanning an object to be inspected using X-rays to obtain an image of the object; processing the image using a first neural network to obtain a semantic expression of the object; reading text information from a manifest of the object to be inspected; processing the text information from the manifest of the object to obtain a semantic feature of the object to be inspected using a second neural network; and determining whether to allow the object to pass based on the semantic expression and the semantic feature. This method can ensure the accuracy of the inspection while significantly increasing the speed of the inspection and significantly improving the efficiency of security inspection.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present application relates to security inspection, and more particularly to a radiation imaging based inspection method, inspection equipment, and computer readable medium. [Background technology]

[0002] The distribution industry is becoming more and more important at present. The distribution industry plays an extremely important role in both the "online shopping for the whole nation" boom and the "Belt and Road" initiative being promoted by the government. However, there are frequent cases of illegal actors taking advantage of the convenience of the distribution chain to falsely declare mail and actually mail prohibited items such as drugs, explosives, and firearms. This poses a huge threat to social security.

[0003] Article 85 of China's Anti-Terrorism Law stipulates that "distribution operators of railway, road, waterway, and air transport, postal, and delivery services that fail to implement security inspection systems, fail to verify the identity of clients, or fail to implement security inspections and open inspections of transported and mailed goods in accordance with regulations, and fail to implement a registration system for the identity of clients requesting transport and mail and the information of goods, shall be punished." Therefore, delivery companies must carry out 100% inspection before packaging, and must fully implement a 100% real-name system and 100% X-ray security inspection system.

[0004] In this case, if manual inspection methods are used to check whether the declared information matches the actual goods, this will inevitably lead to problems such as a large amount of resource waste and reduced efficiency, etc. In addition, manual inspections also place a huge burden on inspectors. Summary of the Invention [Means for solving the problem]

[0005] In view of one or more of the problems of the above-mentioned prior art, embodiments of the present disclosure propose an inspection method, an inspection equipment, and a computer-readable medium, which can automatically inspect an object to be inspected, ensure the inspection accuracy rate, and greatly improve the inspection speed.

[0006] In one aspect of the present disclosure, an inspection method is proposed, including the steps of scanning an object to be inspected using X-rays to obtain an image of the object to be inspected, processing the image using a first neural network to obtain a semantic expression of the object to be inspected, reading text information of a manifest of the object to be inspected, processing the text information of the manifest of the object to be inspected using a second neural network to obtain a semantic feature of the object to be inspected, and determining whether to allow the object to pass based on the semantic expression and the semantic feature.

[0007] According to an embodiment of the present disclosure, the first neural network is a convolutional neural network, or a convolutional neural network based on a candidate region, or a convolutional neural network based on a fast candidate region, and the second neural network is a recurrent neural network or a bidirectional recurrent neural network.

[0008] According to an embodiment of the present disclosure, a set of pre-constructed image-word pairs is used to train the first neural network.

[0009] According to an embodiment of the present disclosure, before processing the image using the first neural network, the method further includes the steps of binarizing the image of the inspected object, calculating an average value for the binarized image, and subtracting the average value from each pixel value of the binarized image.

[0010] According to an embodiment of the present disclosure, the step of determining whether to allow the passage of the inspected object based on the semantic expression and the semantic feature includes calculating a distance between a first vector representing the semantic expression and a second vector representing the semantic feature, and allowing the passage of the inspected object if the calculated distance is smaller than a threshold value.

[0011] According to an embodiment of the present disclosure, during the training process of the first neural network, a correspondence is established between a plurality of region features contained in a sample image and a plurality of words contained in the manifest information of the sample image.

[0012] According to an embodiment of the present disclosure, the scalar product between the feature vector representing the region feature and the semantic vector representing the word is taken as the similarity between the region feature and the word, and the weighted sum of similarities between multiple region features of the sample image and multiple words included in its manifest information is taken as the similarity between the sample image and its manifest information.

[0013] In another aspect of the present disclosure, we propose an inspection equipment including a scanning device that uses X-rays to scan an object to be inspected and obtain a scanned image, an input device that inputs manifest information of the object to be inspected, and a processor, wherein the processor is configured to process the image using a first neural network to obtain a semantic expression of the object to be inspected, process text information of the manifest of the object to be inspected using a second neural network to obtain semantic features of the object to be inspected, and determine whether to allow the object to pass based on the semantic expression and the semantic features.

[0014] In yet another aspect of the present disclosure, a method for implementing a method of the present invention, when executed by a processor, comprising: utilizing a first neural network to process the X-ray image of the inspected object to obtain a semantic representation of the inspected object; utilizing a second neural network to process the text information of the manifest of the inspected object to obtain semantic features of the inspected object; The present invention proposes a computer-readable medium having a computer program stored thereon, the computer program realizing a step of determining whether to allow the passage of the inspected object based on the semantic expression and the semantic features.

[0015] By utilizing the technical solutions in the above embodiments, the inspection accuracy can be ensured, the inspection speed can be greatly increased, and the efficiency of security inspection can be greatly improved.

[0016] In order to make the present disclosure easier to understand, the present disclosure will be described in detail with reference to the following drawings. [Brief description of the drawings]

[0017] [Figure 1] FIG. 1 is a schematic diagram of an inspection fixture according to an embodiment of the present disclosure. [Diagram 2] FIG. 2 is a schematic diagram showing the internal configuration of a computer for image processing in the embodiment shown in FIG. 1. [Diagram 3] FIG. 2 is a schematic diagram illustrating an artificial neural network used in the inspection apparatus and method of the present disclosure; [Figure 4] FIG. 2 is a schematic diagram illustrating another artificial neural network used in the inspection apparatus and method of the present disclosure; [Diagram 5] FIG. 1 is a schematic diagram illustrating the process of aligning images and meanings according to an embodiment of the present disclosure. [Figure 6] 4 is a flowchart showing the construction of an image-word semantic model in the inspection equipment and inspection method according to an embodiment of the present invention. [Figure 7] 1 is a flowchart illustrating a process of performing a security inspection on an object to be inspected using an inspection method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] Hereinafter, specific examples of the present invention will be described in detail. It should be noted that the examples described here are illustrative and are not intended to limit the present invention. In the following description, numerous specific details are described in order to make the present invention easier to understand. However, it is obvious to those skilled in the art that the present invention does not necessarily need to be practiced by adopting these specific details. In other embodiments, detailed descriptions of well-known structures, materials, or methods are omitted so as not to unnecessarily obscure the present invention.

[0019] In view of the problem of low inspection efficiency in the prior art, the embodiment of the present disclosure proposes an inspection technology based on deep learning, which can intelligently complete the comparison of the image list of the goods security inspection machine. By using the technical solution of the embodiment of the present disclosure, it is possible to find an area in the luggage image of the goods security inspection machine that does not match the data of the customs declaration form on which the goods are declared. This area may be falsely declared or concealed and not declared. Falsely declared items are generally prohibited items among circulating goods or dangerous items that pose a threat to the safety of distribution and transportation, and in order to avoid inspection, the luggage is disguised as a safe luggage for transportation on the customs declaration form. Concealed and undeclared items are generally small in number and volume, also known as "concealed carrying", which is a common method used to smuggle prohibited items.

[0020] In addition, in the scanned image of the goods security inspection machine, when a box contains many goods among the goods in circulation, this is a problem that cannot be solved by conventional image processing technology. To be precise, this is a problem of complex division of ambiguity in data supervision of customs declarations, which is affected by the inconsistency of equipment. For example, in different equipment, the effect of algorithms is necessarily different, and the data format of the customs declaration gives multiple supervision values ​​(e.g., how many types of luggage there are, the categories and unit weights of each type, etc.), and each pixel in the image may belong to multiple luggage, etc. The method disclosed herein adopts a method based on deep learning to solve this problem, and no matter how many types of goods are stored in a box, there is no need to manually obtain features, and uses a convolutional neural network and large-scale training data to obtain trained image models and semantic models, and further generates corresponding image-semantic expressions for the test images, thereby accurately completing image list comparison when a box contains many goods.

[0021] Fig. 1 is a schematic diagram showing the configuration of an inspection equipment according to an embodiment of the present disclosure. The inspection equipment 10 shown in Fig. 1 includes an X-ray source 11, a detector module 15, a collection circuit 16, a controller 17, and a data processing computer 18. The radiation source 11 includes one or more X-ray generators, and may perform single-energy transmission scanning or dual-energy transmission scanning.

[0022] As shown in FIG. 1, an object 14 to be inspected, such as an item of luggage, is placed on a conveyor 13 and passes through a scanning area between a radiation source 11 and a detector module 15. In some embodiments, the detector module 15 and the acquisition circuitry 16 are, for example, an integrated modular detector and data collector, such as a multi-row detector, which detects radiation transmitted through the object 14 to obtain analog signals, and converts the analog signals into digital signals to output an X-ray transmission image of the object 14 to be inspected. In the case of a dual energy mode, for example, one row of detectors may be provided for high energy radiation and another row of detectors may be provided for low energy radiation, or the high energy radiation and the low energy radiation may use the same row of detectors in a time-shared manner. The controller 17 controls each part of the entire system to operate synchronously. The data processing computer 18 processes the data collected by the data acquisition circuitry 16, processes the image data, and outputs the results. For example, the data processing computer 18 executes an image processing program, analyzes and learns the scanned image to obtain a semantic expression of the image, and then compares the obtained semantic expression with the semantic features contained in the manifest information of the baggage item to determine whether the declared information matches the object of the baggage item. If there is a match, the baggage item is allowed to pass, and if not, an alarm is issued to notify security personnel that the baggage item is suspicious.

[0023] According to this embodiment, the detector module 15 and the acquisition circuit 16 are used to obtain the transmission data of the inspected object 14. The acquisition circuit 16 includes a data amplifying and forming circuit that operates in a (current) integration mode or a pulse (count) mode. The data output cable of the acquisition circuit 16 is connected to a controller 17 and a data processing computer 18, and the acquired data is stored in the data processing computer 18 according to a trigger command.

[0024] In some embodiments, the detector module 15 includes multiple detection units and receives the X-rays transmitted through the inspected object 14. The acquisition circuit 16 is coupled to the detector module 15 and converts the signal generated by the detector module 16 into detection data. The controller 17 is connected to the radiation source 11 via a control line CTRL1 and to the detector module 15 via a control line CTRL2, and is connected to the acquisition circuit 16 to control one or more X-ray generators in the radiation source 11 to perform a single-energy scan on the inspected object 14 or a dual-energy scan on the inspected object 14, so that the X-rays are emitted and transmitted through the inspected object 14 as the inspected object 14 moves. The controller 17 also controls the detector module 15 and the acquisition circuit 16 to obtain corresponding transmission data, such as single-energy transmission data or dual-energy transmission data. The data processing computer 18 obtains an image of the inspected object 14 based on the transmission data, processes the image, and judges whether the two are consistent based on the manifest information of the inspected object 14.

[0025] Fig. 2 is a block diagram of the data processing computer shown in Fig. 1. As shown in Fig. 2, the data processing computer 20 includes a storage means 21, a ROM (Read Only Memory) 22, a RAM (Random Access Memory) 23, an input device 24, a processor 25, a display means 26, an interface unit 27, and a bus 28.

[0026] The data collected by the collection circuit 16 is stored in the storage means 21 via the interface unit 27 and the bus 28. The ROM (Read Only Memory) 22 stores the layout information and programs of the computer data processor. The RAM (Random Access Memory) 23 is for temporarily storing various data during the operation of the processor 25. The storage means 21 further stores computer programs for data processing. The storage means 21, ROM 22, RAM 23, input device 24, processor 25, display means 28 and interface unit 27 are connected to the internal bus 28.

[0027] After a user inputs an operation command through an input device 24 such as a keyboard or a mouse, the instruction code of the computer program instructs the processor 25 to execute a data processing algorithm, and after obtaining the data processing result, the data processing result is displayed on a display device 27 such as an LCD display, or the processing result is directly output in the form of a hard copy such as a printout.

[0028] For example, the radiation source 11 may be a radioisotope (eg, Co-60), a low energy X-ray machine or a high energy X-ray accelerator, or the like.

[0029] For example, the detector module 15 may be a gas detector, a scintillator detector, or a solid-state detector when classified by material, and may be a single-row, double-row, or multi-row, and a single-layer detector or a double-layer high-low energy detector when classified by array arrangement.

[0030] Although the above has been described as the case where the object 14 to be inspected, such as an item of luggage, moves along the conveyor 13 and passes through the inspection area, it will be appreciated by those skilled in the art that the object 14 to be inspected may be stationary and the radiation source and detector array may be moved to accomplish the scanning process.

[0031] In order to recognize features in a transparent image, the embodiment of the present disclosure proposes to use a convolutional neural network CNN to recognize features in an image. Hereinafter, the convolutional neural network 30 according to the embodiment of the present disclosure will be described in detail with reference to FIG. 3. FIG. 3 is a schematic diagram showing a convolutional neural network 30 according to the embodiment of the present disclosure. As shown in FIG. 3, the convolutional neural network 30 may generally include multiple convolutional layers 32 and 34, and these convolutional layers 32 and 34 are generally sets of partially overlapping small neurons (in a mathematical sense, they are also called convolutional kernels, and hereinafter, unless otherwise specified, these two terms can be used interchangeably). Additionally, throughout this disclosure, unless otherwise stated, for any two layers in the convolutional neural network 30, one layer closer to the input data (or input layer, e.g., input layer 31 in FIG. 3 ) is referred to as the “front” or “lower” layer, and the other layer closer to the output data (or output layer, e.g., output layer 37 in FIG. 3 ) is referred to as the “back” or “upper” layer. Additionally, during training, validation, and / or use, the direction from the input layer (e.g., input layer 31 in FIG. 3 ) to the output layer (e.g., output layer 37 in FIG. 3 ) is referred to as the forward direction, and the direction from the output layer (e.g., output layer 37 in FIG. 3 ) to the input layer (e.g., input layer 31 in FIG. 3 ) is referred to as the backward direction.

[0032] Taking the first convolutional layer 32 shown in FIG. 3 as an example, these small neurons can process each local part of the input image. Then, the outputs of these small neurons are merged and arranged as one output (called feature mapping, for example, a rectangle in the first convolutional layer 32) to obtain an output image that can better represent a certain feature in the original image. At the same time, the partially overlapping arrangement between adjacent neurons also allows the convolutional neural network 30 to have a certain shift tolerance for the feature in the original image. In other words, even if the feature in the original image changes its position by shifting it with a certain tolerance, the convolutional neural network 30 can accurately recognize this feature. The details of the convolutional layer will be described later and will not be described in detail here.

[0033] The next layer is an optional pooling layer, i.e., the first pooling layer 33, which mainly performs downsampling on the output data of the previous convolutional layer 32 while preserving features, reducing the amount of computation and preventing overfitting.

[0034] The next layer is also a convolutional layer, namely the second convolutional layer 34, which can perform further feature sampling on the output data generated by the first convolutional layer 32 and downsampled by the pooling layer 33. Intuitively, the features it learns are larger in overallity than the features learned by the first convolutional layer. Similarly, any subsequent convolutional layers generalize over the features of the previous convolutional layer.

[0035] A convolutional layer (e.g., the first and second convolutional layers 32 and 34) is the core structural unit of a CNN (e.g., a convolutional neural network 30). The parameters of this layer consist of a set of learnable convolutional kernels (or simply called convolutional kernels), each of which has a very small receptive field but extends over the entire depth of the input data. In the forward process, each convolutional kernel is convolved along the width and height of the input data, and the scalar product between the elements of the convolutional kernel and the input data is calculated to generate a two-dimensional active mapping of the convolutional kernel. As a result, the network can learn a convolutional kernel that is activated only when it sees a specific type of feature at a certain input spatial location.

[0036] The active mappings of all convolution kernels are stacked along the depth direction to form the full output data of the convolution layer. Each element of the output data can then be interpreted as the output of a convolution kernel that has seen a small region in the input and shares parameters with other convolution kernels in the same active mapping.

[0037] The depth of the output data controls the number of convolution kernels that connect to the same region of the input data in a layer. For example, as shown in FIG. 3, the depth of the first convolution layer 32 is 4, and the depth of the second convolution layer 34 is 6. All these convolution kernels learn to be activated for different features in the input. For example, when the first convolution layer 32 takes an original image as input, different convolution kernels along the depth dimension (i.e., different corners in FIG. 3) can be activated when various directional edges or grayscale blocks appear in the input data.

[0038] The training process is a very important part in deep learning. In order to ensure that the network can be trained effectively, random gradient descent can be adopted. For example, the Nesterov optimization algorithm can be adopted to solve the problem. In some embodiments, the initial learning rate can be set to start from 0.01 and gradually decrease until an optimal value is found. In some embodiments, for the initial value of the weight, a Gaussian random process with small variance can be used to initialize the weight value of each convolution kernel. In some embodiments, the image training set can be images of items with the feature locations in the images marked.

[0039] According to an embodiment of the present disclosure, a dense image representation is generated in an image by using a CNN (Convolution Neural Network) as shown in FIG. 3 to represent the information of an item in an image of an item security inspection machine. In particular, when using a convolutional neural network model to extract image features, an extraction method of a convolutional neural network based on a candidate region or a convolutional neural network based on a fast candidate region (Faster-RCNN) can be adopted. For example, a dense image representation is generated in an image by using a CNN (Convolution Neural Network) to represent the information of an item in an image of a small item security inspection machine. Since the phrase representation of an image generally relates to objects and their attributes in an image, an RCNN (Region Convolutional Neural Network) is used to detect objects in each image. A series of items and their confidence levels are obtained by detecting in the entire image, and the 10 most confident detection positions are used, and an image representation is calculated from all pixels in each boundary block. Those skilled in the art should understand that other artificial neural networks can be used to recognize and learn features in a transparent image. The embodiments of the present disclosure are not limited thereto.

[0040] To learn the manifest information or to identify the information contained in the manifest, the embodiment of the present disclosure employs recurrent neural networks (RNNs) that are applicable to natural language learning. Figure 4 shows a schematic diagram of an artificial neural network used in the inspection equipment and inspection method of the embodiment of the present disclosure.

[0041] As shown in Figure 4, an RNN typically contains input units, and the input set is {x0, x1, ..., x t , x t+1 , ...}, and the output set of the output units is denoted as {y0, y1, ..., y t , y t+1 The RNN further includes hidden units, whose output set is denoted as {s0, s1, ..., s t , s t+1 , ...}, these hidden units play the most important role. In Figure 4, there is one unidirectional information stream from the input units to the hidden units, and another unidirectional information stream from the hidden units to the output units. In some cases, the RNN can break the latter restriction and direct information back from the output units to the hidden units, which are called "Back Projections". And the input of the hidden layer also contains the state of the previous hidden layer. That is, the nodes in the hidden layer can be self-connected or connected to each other.

[0042] The recurrent neural network is expanded into a full neural network, as shown in Figure 4. For example, for a phrase that contains five words, the expanded network becomes a five-layer neural network, with each layer representing one word.

[0043] In this way, natural language is converted into machine-recognizable codes so that it can be easily quantified in the machine learning process. Since words are the basis for understanding and processing natural language, it is necessary to quantify words, and the embodiment of the present disclosure proposes to use word representation. A word vector refers to representing one word using a real vector v having a predetermined length. A one-hot vector may be used to represent a word. That is, a vector of |V|*1 is generated according to the number of words |V|, and when one bit is "1", the other bit is "0", and this vector represents one word.

[0044] In an RNN, each layer shares the parameters U, V, and W after completing one step of input, which means that each step in an RNN is doing the same thing, just the input is different, greatly reducing the parameters to be learned in the network.

[0045] Bidirectional RNN (Bidirectional Recurrent Neural Network) has an improvement over RNN in that the current output (the output of the tth step) does not depend only on the previous series, but also on the following series. For example, to predict the missing word in a phrase, it is necessary to predict it from the context. Bidirectional RNN is a relatively simple RNN, which is made up of two RNNs stacked one on top of the other. The output is determined by the state of the hidden layers of these two RNNs.

[0046] In the embodiment of the present disclosure, as shown in FIG. 4, a Recurrent Neural Networks (RNN) model is adopted for the customs declaration information or delivery list item information corresponding to an image. For example, in the embodiment of the present disclosure, a bidirectional RNN (BRNN) method is adopted. A BRNN inputs a series of n words, and each word is coded one-hot. DThe input data is then input to the network, converting each word into a constant h-dimensional vector. The representation of the word is enriched by using the context of the word, which has varying girths. Those skilled in the art will appreciate that other artificial neural networks can be used to recognize and learn features in the manifest information. The embodiments of the present disclosure are not limited thereto.

[0047] According to the embodiment of the present disclosure, the inspection process involves three parts: 1) building a database of semantic information of scanned images and corresponding customs declarations, 2) building a semantic model corresponding to the image model of the luggage item, and 3) building an image list matching model to complete intelligent image list matching. The database of semantic information of scanned images and corresponding customs declarations includes three parts: image collection, image pre-processing, and sample pre-processing.

[0048] The construction of a database of images and corresponding semantic information of customs declaration forms is mainly divided into three parts: image collection, image pre-processing, and sample pre-processing. The construction process is as follows: (1) Image collection: A certain number of scanned images of goods are collected by goods security inspection machines, so that the image data database contains various kinds of goods images. It should be noted that at this time, the images include normal goods and prohibited goods. (2) Image pre-processing: The images obtained by scanning and collection have noise information, so the scanned images can be pre-processed. Since the binary image has a uniform physical resolution and can be conveniently combined and used with multiple types of algorithms, the binary image is adopted in this application. It should be noted that in order to ensure the generalization performance of the model, the average value of all images can be subtracted from each image.

[0049] For example, a number of product images scanned by security inspection machines are collected so that the image data base contains various product images. It should be noted that the images include both normal and prohibited products. After that, the images obtained by scanning and collecting contain noise information, so the scanned images need to be pre-processed. Since the binarized image has a uniform physical resolution and can be conveniently combined with multiple types of algorithms, the binarized image is adopted in this application. After the binarized image is obtained, the average grayscale value of the image is calculated, and the average value of all images is subtracted from each image to ensure the generalization performance of the model, and the obtained result is used as the image input of the model.

[0050] According to an embodiment of the present disclosure, the semantic information of the customs declaration is adopted to mark the sample. The image and the semantic information form a complete information pair, which facilitates the training of the network. It should be noted that the semantic information here also includes the product information listed in the delivery transportation list, etc. For example, the semantic information of the customs declaration is adopted to mark the sample. The image and the semantic information form a complete information pair, which facilitates the training of the network. In order to make the model highly sensitive to keywords, it is necessary to process non-keyword words in the description of the product, for example, words that are not keywords for the product, such as "one type" and "one variety", are deleted.

[0051] The process of constructing an image list contrast model involves learning the inherent modal correspondence between language and visual data in a data set consisting of images and their corresponding phrase representations, as shown in Figure 5. As shown in Figure 5, the disclosed method develops a multi-modal code based on a novel model combination scheme and a structuring goal. D The input of the two modes is aligned using a matching model.

[0052] Image features and semantic features are combined into a multi-modal code. DThe image list comparison model is finally obtained by random gradient descent training, i.e., RCNN52 and BRNN53 as shown in FIG. 5. For example, if the calculated result shows that the image corresponds to the semantic information of the declared item, the item will automatically pass through the item security inspection machine. If the calculated result shows that the image does not correspond to the semantic information of the declared item, the item security inspection machine will issue an alarm to notify the operator that there is an abnormality, so that the relevant processing can be performed. In this disclosure, the model is based on a hypothesis, for example, D It needs to learn only from training vocabulary, without relying on input models, rules or categories.

[0053] In this way, these large (image-phrase) data sets are used to consider the linguistic representations of the images as weak markings. Sequentially segmented words in these phrases correspond to some specific but unknown locations in the image. Neural networks 52 and 53 are used to infer these "alignments" and apply them to a learning representation generation model. Specifically, as described above, the present disclosure employs a deep neural network model to infer potential alignment relationships between segment phrases and their corresponding representation image regions. Such a model can be modeled on a common, multi-modal code. D The two modes are related by a matching space and a structuring goal. A multi-mode recurrent neural network architecture is adopted to input an image, generate a corresponding text representation, perform keyword matching between the generated image representation and the image marking information, and judge the similarity between the generated phrase representation and the marking information. Experiments show that the generated text representation phrase has obvious advantages over search-based methods. The model is trained on the inferred correspondences and its effectiveness is tested on a novel locally marked dataset.

[0054] As shown in Figure 5, the multimode cord DWe use a matching model to align image features with semantic features, and finally obtain an image list comparison model through random gradient descent training.

[0055] For example, the constructed RCNN52 and BRNN53 can convert each scanned image of a small item security inspection machine and the corresponding phrase representation into a common set of h-dimensional vectors. Although the supervised vocabulary has granularity across the whole image and phrase, (image-phrase) can be regarded as a function of (domain-word) scores. Intuitively, for an (image-phrase) combination, if some of its words are supported by enough objects or attributes in the image, they should get a high matching score. The vector v of the i-th image domain i and the t-th word vector s t The scalar product between can be interpreted as a scale of similarity and can be used to define a score between an image and a phrase, with a higher score indicating a closer correspondence between the image and the customs declaration.

[0056] In this way, a complete (image-semantic) model can be trained. Note that during actual inspection, the difference between the generated semantic expression and the semantic expression of the actual item can also be part of the loss function. During inspection, after inputting an image to be processed according to the training model, a semantic expression corresponding to the image is obtained, and then it is matched with the actual customs declaration to obtain a confidence score, which is used to determine whether the image list is consistent.

[0057] According to an embodiment of the present disclosure, during training of a neural network, a correspondence is established between a plurality of region features included in a sample image and a plurality of words included in the manifest information of the sample image. For example, a scalar product between a feature vector representing a region feature and a semantic vector representing a word is taken as a similarity between the region feature and the word, and a weighted sum of similarities between the plurality of region features of the sample image and a plurality of words included in the manifest information is taken as a similarity between the sample image and its manifest information.

[0058] FIG. 6 is a flow chart showing the construction of an image-word model in the inspection equipment and inspection method according to the embodiment of the present invention. As shown in FIG. 6, in step S61, a pre-prepared labeled image is input to the convolutional neural network 30 for training. In step S62, a trained image-word model, i.e., a pre-trained convolutional neural network, is obtained. In order to improve the accuracy of the model, in step S63, a test image is input to test the network, and in step S64, a predicted score and a difference with the label are calculated. For example, a test image is input to the neural network to obtain a predicted word expression and a difference between the word expression and the label. For example, two vectors respectively represent a word expression and a label, and the difference between the two vectors represents the difference between the two. In step S65, it is determined whether the difference is smaller than a threshold value. If it is larger, in step S66, the network parameters are updated to adjust the network. If the difference is smaller than the threshold value, in step S67, a network model is constructed. That is, the training for the network is completed.

[0059] FIG. 7 is a flowchart illustrating a process of performing a security inspection on an object to be inspected using an inspection method according to an embodiment of the present disclosure.

[0060] In the actual security inspection process, in step S71, the inspection equipment shown in FIG. 1 is used to scan the inspected object 14, and a transmission image of the inspected object is obtained. In step S74, the manifest information of the inspected object is input to the data processing computer 18 by manual recording, a barcode scanner, or other methods. In the data processing computer 18, in step S72, a first neural network, such as a convolutional neural network or RCNN, is used to process the transmission image, and in step S73, a semantic expression of the inspected object is obtained. In step S75, a bidirectional recursive neural network is used to process the character information of the manifest, and a semantic feature of the inspected object is obtained. Then, in step S76, it is determined whether the semantic feature matches the semantic expression obtained from the image, and if not, an alarm is issued in step S77. If there is a match, in step S78, the inspected object is allowed to pass. According to some embodiments, after calculating a distance between a first vector representing a semantic expression and a second vector representing a semantic feature, if the calculated distance is smaller than a threshold, the object to be inspected may be allowed to pass. Here, the distance between the two vectors may be represented by a sum of absolute values ​​of differences between elements of the two vectors, or may be represented by the Euclidean distance of the two vectors. The embodiments of the present disclosure are not limited thereto.

[0061] This disclosure uses image processing and deep learning methods to intelligently inspect whether declared items match the actual items, greatly improving work efficiency and realizing "speedy passage" while reducing the side effects of various subjective factors and realizing "reliable management". Therefore, it is an important measure for the intelligence of current security inspections and has huge market potential.

[0062] In the above detailed description, various embodiments of the testing method and testing apparatus have been described using schematic diagrams, flow charts, and / or examples. Where such schematic diagrams, flow charts, and / or examples include one or more functions and / or operations, those skilled in the art will appreciate that each function and / or operation in such schematic diagrams, flow charts, or examples may be implemented individually and / or jointly by various configurations, hardware, software, firmware, or substantially any combination thereof. In one embodiment, some portions of the subject matter described in the embodiments of the present invention may be implemented in an Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), Digital Signal Processing (DSP), or other integrated format. However, those skilled in the art will appreciate that some aspects of the embodiments disclosed herein may be equivalently implemented in whole or in part in an integrated circuit. For example, the mechanisms may be implemented as one or more computer programs running on one or more computers (e.g., one or more programs running on one or more computer systems), one or more programs running on one or more processors (e.g., one or more programs running on one or more microprocessors), firmware, or substantially any combination of the above. In addition, one skilled in the art will be capable of designing circuitry and / or writing software and / or firmware code based on this disclosure. In addition, one skilled in the art will understand that the mechanisms described in this disclosure may be distributed as a program product in multiple forms, and that the exemplary embodiments of the present disclosure remain applicable regardless of the specific type of signal-bearing medium that effectuates the distribution.Examples of signal carrier media include, but are not limited to, recordable recording media such as floppy disks (FDs), HDDs, compact discs (CDs), digital versatile discs (DVDs), digital tape, computer memory, and the like, and transport media such as digital and / or analog communications media (e.g., fiber optic cables, wave guides, wired communications links, wireless communications links, and the like).

[0063] Although the present invention has been described above with reference to some typical embodiments of the present invention, it should be understood that the terms used are for the purpose of explanation and illustration, and are not intended to limit the present invention. Furthermore, since the present invention can be specifically implemented in various forms without departing from the spirit and scope of the invention, the above-described embodiments should not be limited to the above details, but should be broadly interpreted within the spirit and scope limited by the claims. All modifications and improvements within the scope of the claims or equivalents are included in the scope of the claims. [Explanation of symbols]

[0064] 10 Inspection Equipment 11 X-ray source 13 Conveyor 14 Inspected object 15 Detector Module 16 Acquisition circuit 17 Controller 18 Data Processing Computer

Claims

1. scanning an inspected object using X-rays to obtain an image of the inspected object; processing the image using a first neural network to obtain first semantic features of the inspected object; reading character information of a manifest of the object to be inspected; utilizing a second neural network to process text information of the manifest of the inspected object to obtain a second semantic feature of the inspected object; aligning the first semantic feature with the second semantic feature using a multimodal coding model, and determining whether to allow the inspection object to pass based on an image list matching model; the multi-modal coding model is configured to associate the first semantic features obtained by processing the image and the second semantic features obtained by processing the character information together in a common multi-modal coding space; The image list comparison model is constructed by a process in which a correspondence relationship is established between a plurality of area features included in a sample image and a plurality of words included in manifest information of the sample image, a scalar product between a feature vector representing the plurality of area features and a semantic vector representing the plurality of words is set as a similarity between the plurality of area features and the plurality of words, and a weighted sum of similarities between the plurality of area features of the sample image and the plurality of words included in the manifest information of the sample image is set as a similarity between the sample image and the manifest information of the sample image, The first neural network is a convolutional neural network, or a candidate region based convolutional neural network, or a candidate region based fast convolutional neural network, and the second neural network is a recurrent neural network or a bidirectional recurrent neural network. Testing method.

2. prior to processing an image using said first neural network; binarizing the image of the inspected object; calculating an average value for the binarized image; and subtracting the average value from each pixel value of the binarized image. The inspection method according to claim 1 .

3. The step of determining whether to permit passage of the inspection object based on the first semantic feature and the second semantic feature includes: Calculating a distance between a first vector representing the first semantic feature and a second vector representing the second semantic feature; and allowing the inspected object to pass if the calculated distance is less than a threshold value. The inspection method according to claim 1 .

4. a scanning device that scans an object to be inspected using X-rays to obtain a scanned image; an input device for inputting manifest information of the object to be inspected; A processor, The processor, processing the image utilizing a first neural network to obtain a first semantic feature of the inspected object; Utilizing a second neural network to process text information of the manifest of the inspected object to obtain a second semantic feature of the inspected object; aligning the first semantic feature with the second semantic feature using a multimodal coding model, and determining whether to allow the inspection object to pass based on an image list matching model; the multi-modal coding model is configured to associate the first semantic features obtained by processing the image and the second semantic features obtained by processing the character information together in a common multi-modal coding space; The image list comparison model is constructed by a process in which a correspondence relationship is established between a plurality of area features included in a sample image and a plurality of words included in manifest information of the sample image, a scalar product between a feature vector representing the plurality of area features and a semantic vector representing the plurality of words is set as a similarity between the plurality of area features and the plurality of words, and a weighted sum of similarities between the plurality of area features of the sample image and the plurality of words included in the manifest information of the sample image is set as a similarity between the sample image and the manifest information of the sample image, the first neural network is a convolutional neural network, or a candidate region based convolutional neural network, or a candidate region based fast convolutional neural network; the second neural network is a recurrent neural network or a bidirectional recurrent neural network; Inspection equipment.

5. The processor further comprises: prior to processing an image using said first neural network; binarizing the image of the inspected object; Calculate the average value of the binarized image; arranged to subtract said average value from each pixel value of the binarized image; 5. The inspection facility according to claim 4.

6. The processor further comprises: Calculating a distance between a first vector representing the first semantic feature and a second vector representing the second semantic feature; if the calculated distance is less than a threshold, then permitting passage of the inspected object.

5. The inspection facility according to claim 4.

7. When executed by a processor, processing an X-ray image of an inspected object utilizing a first neural network to obtain a first semantic feature of the inspected object; utilizing a second neural network to process text information of the manifest of the inspected object to obtain a second semantic feature of the inspected object; a step of aligning the first semantic features and the second semantic features using a multimodal coding model and determining whether to allow passage of the inspected object based on an image list comparison model, the multimodal coding model being configured to associate together the first semantic features obtained by processing the X-ray image and the second semantic features obtained by processing the character information through a common multimodal coding space; The image list comparison model is constructed by a process in which a correspondence relationship is established between a plurality of area features included in a sample image and a plurality of words included in manifest information of the sample image, a scalar product between a feature vector representing the plurality of area features and a semantic vector representing the plurality of words is set as a similarity between the plurality of area features and the plurality of words, and a weighted sum of similarities between the plurality of area features of the sample image and the plurality of words included in the manifest information of the sample image is set as a similarity between the sample image and the manifest information of the sample image, the first neural network is a convolutional neural network, or a candidate region based convolutional neural network, or a candidate region based fast convolutional neural network; the second neural network is a recurrent neural network or a bidirectional recurrent neural network; Computer-readable medium.

Citation Information

Patent Citations

  • Object identifying device

    JP1995160665A

  • Management system for substrate mounted component

    JP2009075744A

  • X-ray inspection system that integrates manifest data into imaging / detection processing

    JP2014525594A

  • Methods, apparatus, and products for semantic processing of text

    JP2015515674A

  • Inspection method for cargo and its system

    JP2017097853A