Cigarette parcel identification system and method based on multi-modal feature fusion
By combining a multimodal feature fusion method with dual-energy computed tomography and millimeter wave detection, the accuracy problem of cigarette package identification was solved, and cigarette package identification and quantity judgment were achieved under non-destructive testing.
Patent Information
- Application Number
- CN202510826298.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
AI Technical Summary
Existing package inspection technologies have difficulty accurately identifying cigarette packages and their quantity without opening the packages, especially when they are disguised.
A method combining dual-energy computed tomography and millimeter wave detection is used to obtain the density distribution matrix, equivalent material order matrix and three-dimensional dielectric constant tensor of the package. Cigarette packages are then identified through multimodal feature fusion and machine learning algorithms.
The accuracy and reliability of cigarette package identification are improved, and it is possible to determine whether a package is a cigarette package and the quantity under non-destructive testing.
Smart Images

Figure CN120708009A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a cigarette package recognition system and method based on multimodal feature fusion. Background Art
[0002] During the logistics transportation and mailing process, tobacco packages need to be accurately identified and supervised to ensure that the quantity of tobacco products such as cigarettes transported and mailed complies with regulations.
[0003] However, in actual operations, tobacco products such as cigarettes may be cleverly disguised in other items. Therefore, it is usually difficult to identify tobacco packages without opening the packages during traditional package inspections, and it is even more difficult to determine whether the quantity complies with regulations.
[0004] Based on this, it is necessary to study a system and method for identifying tobacco packages, so as to achieve accurate identification and quantity judgment of tobacco packages under the premise of non-destructive testing, and provide effective technical support for the supervision of tobacco products in logistics transportation and mailing. Summary of the Invention
[0005] To solve the above problems, one aspect of the embodiments of this specification provides a cigarette package recognition method based on multimodal feature fusion, the method comprising:
[0006] Acquiring dual-energy computed tomography data and millimeter wave detection data collected for the target package, wherein the dual-energy computed tomography data is used to reflect the structural characteristics of the interior of the target package, and the millimeter wave detection data is used to reflect the dielectric properties of the material inside the target package;
[0007] Obtaining a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data, and obtaining a three-dimensional dielectric constant tensor corresponding to the target package based on the millimeter wave detection data;
[0008] Performing feature extraction on the density distribution matrix and the equivalent material ordinal matrix to obtain a first eigenvector;
[0009] Performing feature extraction on the three-dimensional dielectric constant tensor to obtain a second eigenvector;
[0010] Performing feature fusion on the first feature vector and the second feature vector to obtain a multimodal feature fusion vector;
[0011] The multimodal feature fusion vector is processed by a pre-trained package recognition model to determine whether the target package is a cigarette package.
[0012] In some embodiments, obtaining a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data includes:
[0013] Performing image reconstruction based on the dual-energy computed tomography data to obtain a target three-dimensional reconstructed image;
[0014] Density analysis and equivalent material ordinal analysis are performed on each voxel unit in the area corresponding to the target package in the target three-dimensional image according to the dual-energy computed tomography data to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package.
[0015] In some embodiments, performing density analysis and equivalent material ordinal analysis on each voxel unit in the region corresponding to the target package in the target three-dimensional image according to the dual-energy computed tomography data to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package includes:
[0016] For each target voxel cell, do the following:
[0017] Step 1, determining the first attenuation coefficient and the second attenuation coefficient of the target voxel unit under two different X-ray energies;
[0018] Step 2: Substituting the first attenuation coefficient and the second attenuation coefficient into a preset material decomposition equation according to a dual-energy attenuation model to calculate a physical density value of the target voxel unit;
[0019] Step 3, calculating the equivalent material ordinal value corresponding to the target voxel unit according to the physical density value of the target voxel unit and the preset mapping relationship between the equivalent material ordinal number and the physical density;
[0020] The physical density values and equivalent material ordinal values corresponding to all target voxel units are respectively combined according to the spatial relationship in the target package to obtain the density distribution matrix and equivalent material ordinal matrix corresponding to the target package.
[0021] In some embodiments, obtaining a three-dimensional dielectric constant tensor corresponding to the target package based on the millimeter wave detection data includes:
[0022] Performing phase unwrapping and scattered field inversion on the millimeter wave detection data to generate a three-dimensional complex permittivity distribution corresponding to the target package;
[0023] extracting a real permittivity component from the complex permittivity distribution;
[0024] A three-dimensional coordinate system is constructed according to the spatial relationship between multiple voxel units contained in the target package, and the extracted real dielectric constant component is mapped to the corresponding position in the three-dimensional coordinate system to obtain a three-dimensional dielectric constant tensor corresponding to the target package.
[0025] In some embodiments, the output of the package recognition model is a confidence value that the target package is a cigarette package, and the processing of the multimodal feature fusion vector by the pre-trained package recognition model to determine whether the target package is a cigarette package includes:
[0026] When the confidence value is greater than or equal to a first preset threshold, determining that the target package is a cigarette package and triggering a cigarette quantity identification process;
[0027] When the confidence value is greater than or equal to the second preset threshold and less than the first preset threshold, prompting relevant staff to perform manual verification;
[0028] When the confidence value is less than the second preset threshold, the target package is determined to be a non-cigarette package.
[0029] In some embodiments, the cigarette quantity identification process includes:
[0030] performing three-dimensional contour fitting based on the target three-dimensional reconstructed image corresponding to the dual-energy computed tomography data, and calculating similarity between the fitted three-dimensional contour and a preset reference three-dimensional contour to determine at least one cigarette state present in the target package;
[0031] According to the number of target objects corresponding to each cigarette state and the number of reference cigarettes corresponding to each cigarette state, an estimation result of the number of cigarettes corresponding to the target package is obtained.
[0032] In some embodiments, the cigarette quantity identification process includes:
[0033] performing three-dimensional contour fitting based on the target three-dimensional reconstructed image corresponding to the dual-energy computed tomography data, and calculating similarity between the fitted three-dimensional contour and a preset reference three-dimensional contour to determine at least one cigarette state present in the target package;
[0034] Mapping the at least one cigarette state in the target package into a third feature vector, and performing feature fusion on the third feature vector, the first feature vector, and the second feature vector to obtain a fourth feature vector;
[0035] The fourth eigenvector is processed by a pre-trained cigarette quantity recognition model to obtain an estimation result of the cigarette quantity corresponding to the target package.
[0036] In some embodiments, the cigarette quantity recognition model includes a first convolution unit, a second convolution unit, a first pooling unit, and a first fully connected neural network. The cigarette quantity recognition model is trained based on the following method:
[0037] Obtaining dual-energy computed tomography (DECT) sample data and millimeter wave detection sample data collected for a sample package, as well as a cigarette quantity label corresponding to the sample package, wherein the DECT sample data is used to reflect the structural characteristics of the interior of the sample package, the millimeter wave detection sample data is used to reflect the dielectric properties of the material inside the sample package, and the cigarette quantity label is used to reflect the number of cigarettes contained in the sample package;
[0038] Obtaining a sample density distribution matrix and a sample equivalent material ordinal matrix corresponding to the sample package based on the dual-energy computed tomography sample data, and obtaining a sample three-dimensional dielectric constant tensor corresponding to the sample package based on the millimeter wave detection sample data;
[0039] Inputting the sample density distribution matrix and the sample equivalent material ordinal matrix into the first convolution unit for feature extraction to obtain a first sample feature vector;
[0040] Inputting the sample three-dimensional dielectric constant tensor into the second convolution unit for feature extraction to obtain a second sample feature vector;
[0041] performing three-dimensional contour fitting based on a target three-dimensional reconstructed image corresponding to the dual-energy computed tomography sample data, and calculating a similarity between the fitted three-dimensional contour and a preset reference three-dimensional contour to determine at least one cigarette state present in the sample package, and then mapping the at least one cigarette state present in the sample package using a preset mapping rule to obtain a third sample feature vector;
[0042] Performing feature fusion on the third sample feature vector, the first sample feature vector, and the second sample feature vector to obtain a fourth sample feature vector;
[0043] Processing the fourth sample feature vector using the first pooling unit and the first fully connected neural network to obtain a predicted number of cigarettes;
[0044] A first sum of the differences between the predicted number of cigarettes and the cigarette number labels corresponding to all sample packages is calculated, and the parameters of the first convolution unit, the second convolution unit, the first pooling unit, and the first fully connected neural network are iteratively adjusted with the goal of reducing the first sum until the trained cigarette number recognition model is obtained when the preset training conditions are met.
[0045] In some embodiments, the package recognition model includes a third convolutional unit, a fourth convolutional unit, a second pooling unit, and a second fully connected neural network. The package recognition model is trained based on the following method:
[0046] Obtaining dual-energy computed tomography sample data and millimeter wave detection sample data collected for a sample package, as well as a package label corresponding to the sample package, wherein the dual-energy computed tomography sample data is used to reflect the structural characteristics of the interior of the sample package, the millimeter wave detection sample data is used to reflect the dielectric properties of the material inside the sample package, and the package label is used to indicate whether the sample package is a cigarette package;
[0047] Obtaining a sample density distribution matrix and a sample equivalent material ordinal matrix corresponding to the sample package based on the dual-energy computed tomography sample data, and obtaining a sample three-dimensional dielectric constant tensor corresponding to the sample package based on the millimeter wave detection sample data;
[0048] Inputting the sample density distribution matrix and the sample equivalent material ordinal matrix into the third convolution unit for feature extraction to obtain a first sample feature vector;
[0049] Inputting the sample three-dimensional dielectric constant tensor into the fourth convolution unit for feature extraction to obtain a second sample feature vector;
[0050] Performing feature fusion on the first sample feature vector and the second sample feature vector to obtain a sample multimodal feature fusion vector;
[0051] Processing the sample multimodal feature fusion vector through the second pooling unit and the second fully connected neural network to obtain a predicted package recognition result;
[0052] A second sum of the differences between the predicted package identification results and the package labels corresponding to all sample packages is calculated, and the parameters of the third convolution unit, the fourth convolution unit, the second pooling unit, and the second fully connected neural network are iteratively adjusted with the goal of reducing the second sum until the trained package identification model is obtained when the preset training conditions are met.
[0053] Another aspect of the embodiments of this specification further provides a cigarette package recognition system based on multimodal feature fusion, the system comprising:
[0054] an acquisition module, configured to acquire dual-energy computed tomography data and millimeter-wave detection data collected for a target package, wherein the dual-energy computed tomography data is used to reflect the structural characteristics of the interior of the target package, and the millimeter-wave detection data is used to reflect the dielectric properties of the material inside the target package;
[0055] a first processing module, configured to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data, and obtain a three-dimensional dielectric constant tensor corresponding to the target package based on the millimeter wave detection data;
[0056] A first feature extraction module is used to extract features from the density distribution matrix and the equivalent material ordinal matrix to obtain a first feature vector;
[0057] A second feature extraction module is used to extract features from the three-dimensional dielectric constant tensor to obtain a second feature vector;
[0058] a multimodal feature fusion module, configured to perform feature fusion on the first feature vector and the second feature vector to obtain a multimodal feature fusion vector;
[0059] The second processing module is used to process the multimodal feature fusion vector using a pre-trained package recognition model to determine whether the target package is a cigarette package.
[0060] The beneficial effects that may be brought about by the cigarette package identification system and method based on multimodal feature fusion provided in the embodiments of this specification include at least: by combining millimeter wave detection data with dual-energy computed tomography data, more comprehensive and accurate feature information and judgment basis can be provided for cigarette package identification, thereby improving the accuracy and reliability of cigarette package identification results to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbers represent the same structures, wherein:
[0062] Figure 1 is an exemplary flow chart of a cigarette package recognition method based on multimodal feature fusion according to some embodiments of this specification;
[0063] Figure 2 is an exemplary structural diagram of a cigarette quantity recognition model according to some embodiments of this specification;
[0064] Figure 3 is a schematic diagram of an exemplary structure of a package identification model according to some embodiments of this specification;
[0065] Figure 4 This is an exemplary module diagram of a cigarette package recognition system based on multimodal feature fusion according to some embodiments of this specification. DETAILED DESCRIPTION
[0066] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0067] It should be understood that the terms "system," "device," "unit," and / or "module" used in this specification are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0068] As used in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but also include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0069] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0070] Tobacco products include cigarettes, cigars, and shredded tobacco, among which cigarettes are the most common. Because cigarettes are the most common and their characteristics are more distinct than those of other tobacco products, the present invention provides a cigarette package recognition system and method based on multimodal feature fusion, specifically targeting cigarettes. This system uses dual-energy computed tomography data and millimeter-wave detection data to identify cigarette packages, thereby improving the accuracy and reliability of the recognition results and enabling identification of the number of cigarettes in a package without destructive testing.
[0071] The cigarette package recognition system and method based on multimodal feature fusion provided in the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0072] Figure 1This is an exemplary flow chart of a method for identifying cigarette packages based on multimodal feature fusion according to some embodiments of this specification. In some embodiments, the method for identifying cigarette packages based on multimodal feature fusion can be executed by processing logic, which can include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (instructions running on a processing device to perform hardware simulation), etc., or any combination thereof. In some embodiments, Figure 1 One or more operations in the flowchart of the multimodal feature fusion-based cigarette package identification method can be implemented by a processing device and / or a terminal device. For example, the method can be stored in the form of a computer program and / or instructions in a storage device and invoked and / or executed by the processing device and / or terminal device.
[0073] Reference Figure 1 The cigarette package recognition method based on multimodal feature fusion provided in the embodiment of the present application may include the following steps S110 to S160:
[0074] Step S110: Acquire dual-energy computed tomography data and millimeter wave detection data collected for the target package. In the embodiment of the present application, step S110 may be executed by the acquisition module 210 mentioned below.
[0075] In the embodiment of the present application, the target package may refer to a package that needs to be identified as a cigarette package, which may be a transport package during logistics transportation, or a package that is inspected by law enforcement agencies during an inspection process.
[0076] It is understandable that the currently used package detection technologies (such as X-ray detection) have certain limitations when dealing with cigarette packages. For example, traditional single-modal detection methods may not be able to accurately distinguish cigarettes from other items with similar structures, and are even less able to identify disguised cigarette packages and the number of cigarettes involved.
[0077] In response to the above problems, in an embodiment of the present application, in order to more accurately identify cigarette packages, dual-energy computed tomography data and millimeter wave detection data collected for the target package can be obtained, and then these data can be analyzed and processed through multimodal feature fusion and machine learning algorithms to extract complementary information from data of different modalities, thereby improving the accuracy and reliability of subsequent recognition results.
[0078] Specifically, in an embodiment of the present application, the dual-energy computed tomography data can be obtained by scanning the target package using a dual-energy computed tomography device. The dual-energy computed tomography device can use two different energies of X-rays to image the target package, thereby providing information reflecting the structural characteristics of the target package (such as the density of the material inside the target package and its corresponding equivalent material number). The millimeter wave detection data can be obtained by transmitting millimeter waves and receiving their reflected waves using a millimeter wave detection device. In an embodiment of the present application, the millimeter wave detection data can be used to reflect the dielectric properties of the material inside the target package.
[0079] It should be noted that different substances will exhibit different absorption characteristics under different energy X-rays. By analyzing the differences in these absorption characteristics, we can obtain approximate information about the internal structure and material composition of the target package, thereby helping to identify the approximate outline and shape of the items in the target package. Millimeter waves have strong penetrability and can penetrate some non-conductive materials, and have better detection capabilities for items hidden inside the package. In an embodiment of the present application, the millimeter wave detection data is time series data, which can reflect the dielectric properties and differences of the materials corresponding to different depths inside the target package through the time series characteristics of the received data (i.e., the reflected wave). Therefore, in an embodiment of the present application, by combining the millimeter wave detection data with the above-mentioned dual-energy computed tomography data, it is possible to further supplement the information that the dual-energy computed tomography data cannot cover, thereby providing a more comprehensive and accurate judgment basis for cigarette package identification, thereby improving the accuracy and reliability of the cigarette package identification results obtained in the subsequent process to a certain extent.
[0080] Step S120: Determine the density distribution matrix and equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data, and determine the three-dimensional permittivity tensor corresponding to the target package based on the millimeter-wave detection data. In this embodiment of the present application, step S120 may be performed by the first processing module 220 described below.
[0081] In an embodiment of the present application, after obtaining the dual-energy computed tomography data and millimeter-wave detection data collected for the target package through the above steps, image reconstruction can be performed based on the dual-energy computed tomography data to obtain a target three-dimensional reconstructed image (this reconstruction process can be regarded as a prior art). Then, density analysis and equivalent material ordinal analysis are performed on each voxel unit in the area corresponding to the target package in the target three-dimensional image based on the dual-energy computed tomography data to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package.
[0082] In this embodiment, each voxel in the region corresponding to the target package can be understood as a tiny cubic region in the three-dimensional space corresponding to the target package. Together, these voxels constitute a complete representation of the target package in three-dimensional space. By performing density analysis and equivalent material ordinal analysis on each voxel, the internal structural characteristics of the target package and the distribution characteristics of the internal materials can be accurately reflected.
[0083] Specifically, in the embodiment of the present application, for each target voxel unit, the following operations may be performed to determine its corresponding physical density value and equivalent material ordinal value:
[0084] Step 1: Determine the first attenuation coefficient and the second attenuation coefficient of the target voxel unit under two different X-ray energies.
[0085] In an embodiment of the present application, the first attenuation coefficient and the second attenuation coefficient can be determined as follows: First, the projection data of the target voxel unit under two different energy X-ray irradiations are obtained from the dual-energy computed tomography data. Then, based on the existing mathematical mapping model of attenuation coefficient and projection data, the obtained projection data are combined to calculate the first attenuation coefficient and the second attenuation coefficient of the target voxel unit under the two different X-ray energies. In an embodiment of the present application, the mathematical mapping model can be constructed based on the attenuation law of X-rays in matter. More calculation details about the first attenuation coefficient and the second attenuation coefficient can be regarded as prior art and will not be discussed in detail in this specification.
[0086] Step 2: According to the dual-energy attenuation model, the first attenuation coefficient and the second attenuation coefficient are substituted into a preset material decomposition equation for calculation to obtain the physical density value of the target voxel unit.
[0087] In the embodiment of the present application, the dual-energy attenuation model is a mathematical model for calculating the physical density of a substance based on the attenuation of two X-rays of different energies in the substance. In the embodiment of the present application, the dual-energy attenuation model can be trained based on the attenuation characteristics of X-rays when interacting with the substance.
[0088] In an embodiment of the present application, a material decomposition equation can be obtained based on the dual-energy attenuation model. By substituting the first attenuation coefficient and the second attenuation coefficient into the equation, after a series of mathematical operations, the physical density value corresponding to the target voxel unit can be obtained. It should be noted that, in an embodiment of the present application, the material decomposition equation can be regarded as a specific mathematical expression obtained by training the dual-energy attenuation model, and its coefficients and parameters can be determined based on the training data of the dual-energy attenuation model. In some embodiments, the material decomposition equation can be a processing function (such as an activation function) corresponding to the hidden layer of the dual-energy attenuation model.
[0089] It should be pointed out that in the embodiment of the present application, by combining the attenuation coefficients at two different X-ray energies, the attenuation differences of different materials at different X-ray energies can be fully utilized, thereby providing a richer and more accurate data basis for the physical density calculation of each target voxel unit.
[0090] Step 3: Calculate the equivalent material ordinal value corresponding to the target voxel unit according to the physical density value of the target voxel unit and the preset mapping relationship between the equivalent material ordinal number and the physical density.
[0091] In the embodiments of the present application, the mapping relationship between the equivalent substance ordinal number and the physical density can be constructed using a large amount of experimental data. For example, in some embodiments, X-ray scanning experiments of different energies can be performed on multiple known substances (each known substance corresponds to an equivalent substance ordinal number), and then the corresponding physical densities can be calculated using the above method. The physical densities can then be combined with the equivalent substance ordinal numbers corresponding to the known substances to construct a mapping table, thereby obtaining the mapping relationship between the equivalent substance ordinal number and the physical density.
[0092] In an embodiment of the present application, after obtaining the physical density value of the target voxel unit through the above steps, the physical density value of the target voxel unit can be substituted into the above mapping table to obtain its corresponding equivalent material ordinal value, that is, the equivalent material ordinal value corresponding to the target voxel unit.
[0093] It should be noted that in some embodiments of the present application, after obtaining the physical density value of the target voxel unit through the above steps, a target physical density value closest to the physical density value of the target voxel unit can be searched in the mapping table, and then the equivalent material ordinal number corresponding to the target physical density value is used as the equivalent material ordinal value corresponding to the target voxel unit.
[0094] In an embodiment of the present application, the physical density value and equivalent material ordinal value corresponding to each target voxel unit can be calculated in the above manner, and then the physical density values and equivalent material ordinal values corresponding to all target voxel units are combined according to their spatial relationship in the target package (i.e., their position relationship in the spatial coordinate system) to obtain the density distribution matrix and equivalent material ordinal matrix corresponding to the target package.
[0095] Furthermore, in an embodiment of the present application, a three-dimensional dielectric constant tensor corresponding to the target package can be obtained based on the millimeter wave detection data. This process may include the following steps:
[0096] First, phase unwrapping and scattered field inversion are performed on the millimeter wave detection data to generate a three-dimensional complex dielectric constant distribution corresponding to the target wrapping.
[0097] In the embodiments of this application, phase unwrapping refers to the process of restoring the true phase from the phase that has been truncated within the main value interval in the millimeter-wave detection data. Since the phase data collected during the millimeter-wave detection process is usually truncated within a limited interval, resulting in discontinuous phase information, phase unwrapping technology can restore these truncated phases to the continuous true phase, thereby providing accurate phase information for subsequent scattered field inversion.
[0098] Specifically, in the embodiment of the present application, a three-dimensional least squares phase unwrapping algorithm can be performed on the millimeter wave detection data in a Cartesian coordinate system, and the phase difference matrix between adjacent detection points can be solved. (where S m , S n is the signal amplitude of the adjacent detection point), and then the continuous phase field is reconstructed based on the Poisson equation, so as to perform phase unwrapping processing on the millimeter wave detection data.
[0099] Scattering field inversion refers to the use of unwrapped phase and amplitude data, through specific algorithms and models, to infer the complex dielectric constant distribution inside the target package from the detected scattered field information, thereby obtaining the three-dimensional complex dielectric constant distribution corresponding to the target package.
[0100] In some embodiments, an electromagnetic scattering integral equation may be established, and then a conjugate gradient method may be used to solve the dielectric constant perturbation distribution, and the complex dielectric constant corresponding to each voxel unit may be calculated using the perturbation relationship.
[0101] Just as an example, in some embodiments of the present application, the electromagnetic scattering integral equation can be expressed as follows:
[0102] E sc (r)=∫∫∫ V G(r,r')·χ(r')Einc (r')dr'
[0103] Among them, E sc (r) represents the scattered electric field at the observation point r (in complex form); r represents the position coordinate of the observation point (three-dimensional vector); r' represents the position coordinate of the source point (the source point can refer to any point in the spatial region corresponding to the target package) (three-dimensional vector); V represents the spatial region occupied by the target package; G(r,r') represents the Green's function, which is used to describe the propagation of electromagnetic waves from the source point r' to the observation point r; χ(r') represents the dielectric constant perturbation at position r', which is defined as (where ε0 is the dielectric constant of vacuum); E inc (r') represents the incident electric field at position r'.
[0104] In some embodiments, the iterative formula for solving the dielectric constant perturbation distribution by the conjugate gradient method can be expressed as:
[0105]
[0106] Among them, X (n) The dielectric constant perturbation distribution (X (n+1) Similarly); α n Indicates the step size of the nth iteration; represents the complex conjugate of the incident electric field; G * represents the complex conjugate of Green's function; E meas (r) represents the scattered electric field actually measured (obtained through the millimeter wave detection data); Indicates that based on the current χ (n) The scattered electric field calculated by the above electromagnetic scattering integral equation; Re[*] represents the real part operator.
[0107] Furthermore, the above disturbance relation can be expressed as:
[0108] ε r (r)=ε bkg +χ(r)
[0109] Among them, ε r (r) represents the relative dielectric constant at position r (complex form, also known as the complex dielectric constant mentioned later); ε bkg Represents the ambient dielectric constant.
[0110] Through the above processing, the complex dielectric constant corresponding to each voxel unit can be obtained.
[0111] Furthermore, after obtaining the complex permittivity corresponding to each voxel unit through the above steps, the real permittivity component can be extracted from the complex permittivity distribution. Then, a three-dimensional coordinate system is constructed according to the spatial relationship between the multiple voxel units contained in the target package, and the extracted real permittivity components are mapped to corresponding positions in the three-dimensional coordinate system, thereby obtaining a three-dimensional permittivity tensor corresponding to the target package.
[0112] It should be pointed out that in the embodiment of the present application, the three-dimensional dielectric constant tensor corresponding to the target package can reflect the dielectric properties of the material at different positions inside the target package, thereby providing richer and more reliable feature information for cigarette package identification.
[0113] Step S130, feature extraction is performed on the density distribution matrix and the equivalent material ordinal matrix to obtain a first feature vector. In the embodiment of the present application, step S130 can be performed by the first feature extraction module 230 mentioned later. Specifically, in some embodiments of the present application, step S130 can be performed by Figure 3 The third convolutional unit in the package recognition model shown performs
[0114] In an embodiment of the present application, the density distribution matrix and the equivalent material ordinal matrix may be subjected to feature extraction by a third convolution unit in a pre-trained package recognition model to obtain a first feature vector.
[0115] It can be understood that in the embodiment of the present application, the third convolution unit can learn the key characteristic patterns in the density distribution matrix and the equivalent material ordinal matrix through training, and extract representative features from the density distribution matrix and the equivalent material ordinal matrix, thereby obtaining the first eigenvector, providing feature information with certain utilization value for subsequent cigarette package identification.
[0116] Step S140, extracting features from the three-dimensional dielectric constant tensor to obtain a second feature vector. In the embodiment of the present application, step S140 can be performed by the second feature extraction module 240 mentioned later. Specifically, in some embodiments of the present application, step S140 can be performed by Figure 3 The fourth convolutional unit in the package recognition model shown performs
[0117] Similarly, in this embodiment of the present application, a fourth convolution unit in a pre-trained package recognition model can be used to extract features from the three-dimensional permittivity tensor to obtain a second eigenvector. This fourth convolution unit can be trained to learn key characteristic patterns in the three-dimensional permittivity tensor and extract representative features from the three-dimensional permittivity tensor to obtain the second eigenvector, providing valuable feature information for subsequent cigarette package recognition.
[0118] More details about the package recognition model and the training process of the third and fourth convolutional units can be found later and will not be discussed in detail here.
[0119] Step S150: Perform feature fusion on the first feature vector and the second feature vector to obtain a multimodal feature fusion vector. In the embodiment of the present application, step S150 may be performed by the multimodal feature fusion module 250 mentioned below.
[0120] Continue to refer to Figure 3 In an embodiment of the present application, after obtaining the first feature vector and the second feature vector through the above steps, feature fusion can be performed on the first feature vector and the second feature vector to obtain a multimodal feature fusion vector.
[0121] In this embodiment of the present application, the method for fusing the first feature vector and the second feature vector may include, but is not limited to, concatenation. That is, in this embodiment of the present application, the first feature vector and the second feature vector may be concatenated end to end in a certain order to form a longer vector, thereby completely preserving all information of the two vectors.
[0122] Step S160: Process the multimodal feature fusion vector using a pre-trained package recognition model to determine whether the target package is a cigarette package. In this embodiment of the present application, step S160 may be performed by the second processing module 260 mentioned below.
[0123] In an embodiment of the present application, the multimodal feature fusion vector obtained in the above steps can be processed by a pre-trained package recognition model to determine whether the target package is a cigarette package.
[0124] For details, please refer to Figure 3In some embodiments of the present application, the package recognition model may include, in addition to the third and fourth convolutional units described above, a second pooling unit and a second fully connected neural network. After obtaining the multimodal feature fusion vector, the multimodal feature fusion vector may be processed by the second pooling unit and the second fully connected neural network in the pre-trained package recognition model to obtain a package recognition result for the target package. In this embodiment of the present application, the package recognition result may reflect whether the target package is a cigarette package.
[0125] In the embodiment of the present application, the second pooling unit can be used to perform a downsampling operation on the multimodal feature fusion vector obtained by processing the third and fourth convolution units, reducing the feature dimension while retaining the main feature information, thereby reducing the complexity of subsequent calculations and the risk of overfitting of the model. In other words, in the embodiment of the present application, through the processing of the second pooling unit, more representative feature information can be extracted, so that the subsequent second fully connected neural network can more accurately determine whether the target package is a cigarette package based on these features.
[0126] In this embodiment of the present application, the second fully connected neural network may include multiple hidden layers and an output layer. The hidden layer can be used to perform nonlinear transformation and feature extraction on the feature information processed by the second pooling unit. The activation functions of different neurons can be used to explore the complex relationships between features, further improving the expressive power of the features. Based on the feature information transmitted by the hidden layer, the output layer can output a prediction result indicating whether the target package is a cigarette package.
[0127] For example, in some embodiments of the present application, the output layer may use a sigmoid activation function to map the output value to a range from 0 to 1 to measure the confidence that the target package is a cigarette package. A larger output value indicates a higher likelihood that the target package is a cigarette package, and vice versa, a lower likelihood indicates a lower likelihood that the target package is a cigarette package.
[0128] In some embodiments of the present application, when the confidence value output by the output layer is greater than or equal to a first preset threshold (e.g., 0.9), the target package can be determined to be a cigarette package, and the cigarette quantity identification process can be triggered. When the confidence value output by the output layer is greater than or equal to a second preset threshold (e.g., 0.6) and less than the first preset threshold, relevant staff can be prompted to perform manual verification; when the confidence value output by the output layer is less than the second preset threshold, the target package can be determined to be a non-cigarette package.
[0129] Specifically, in some embodiments of the present application, the above-mentioned cigarette quantity identification process may include the following steps:
[0130] First, a 3D contour fitting is performed based on the target 3D reconstructed image corresponding to the dual-energy computed tomography data. Similarity is calculated between the fitted 3D contour and a preset reference 3D contour to determine at least one cigarette state present in the target package. Then, based on the number of target objects corresponding to each cigarette state and the number of reference cigarettes corresponding to each cigarette state, an estimated number of cigarettes in the target package is obtained.
[0131] In this embodiment of the present application, the preset reference 3D profile refers to a representative standard 3D profile obtained by analyzing and statistically analyzing a large number of 3D reconstructed images of cigarette package samples. This profile covers the profile characteristics of various common cigarette states, such as packed in cartons, boxes, and bulk, as well as states consisting of cigarettes of varying numbers and / or states stacked together. In this embodiment of the present application, by calculating the similarity between the fitted 3D profile and the preset reference 3D profile (e.g., using methods such as Euclidean distance and cosine similarity), the possible cigarette states can be accurately identified from the 3D reconstructed image of the target package.
[0132] It is understood that each state can correspond to a certain number of cigarettes (i.e., a reference number of cigarettes). For example, in a carton state, each carton of cigarettes generally contains 200 cigarettes; in a box state, each box of cigarettes generally contains 20 cigarettes; and in a bulk state, the number of cigarettes can be estimated based on the number of identified individual cigarettes or the accumulated volume of multiple cigarettes. Based on this, in some embodiments of the present application, the number of target objects corresponding to each cigarette state and the reference number of cigarettes corresponding to each cigarette state can be used to calculate the estimated number of cigarettes corresponding to the target package.
[0133] In some embodiments of the present application, the aforementioned cigarette count recognition process may misidentify objects with similar structures as cigarettes, resulting in an overestimation of the final cigarette count estimate. Based on this, in other embodiments of the present application, at least one cigarette state identified from the target 3D reconstructed image may be combined with the aforementioned first and second eigenvectors. This reduces the probability of misidentification based on the structural characteristics and material dielectric properties of the target package as reflected by the first and second eigenvectors, thereby ensuring the accuracy of the cigarette count estimate.
[0134] Specifically, in an embodiment of the present application, three-dimensional contour fitting can be performed based on the target three-dimensional reconstructed image corresponding to the dual-energy computed tomography data, and the similarity between the fitted three-dimensional contour and the preset reference three-dimensional contour is calculated to determine at least one cigarette state present in the target package; then, the at least one cigarette state present in the target package is mapped to a third eigenvector, and the third eigenvector is feature fused with the first eigenvector and the second eigenvector to obtain a fourth eigenvector; finally, the fourth eigenvector is processed by a pre-trained cigarette quantity recognition model to obtain a cigarette quantity estimation result corresponding to the target package.
[0135] Reference Figure 2 In an embodiment of the present application, the cigarette quantity recognition model may include a first convolution unit, a second convolution unit, a first pooling unit, and a first fully connected neural network, wherein the first convolution unit functions similarly to the third convolution unit, and is used to extract features from the density distribution matrix and the equivalent material ordinal matrix to obtain a first eigenvector (the first eigenvector may be the same as or different from the first eigenvector obtained by the third convolution unit). The second convolution unit functions similarly to the fourth convolution unit, and may be used to extract the three-dimensional dielectric constant tensor to obtain a second eigenvector (the second eigenvector may be the same as or different from the second eigenvector obtained by the fourth convolution unit).
[0136] Furthermore, in this embodiment of the present application, the at least one cigarette condition present in the target package can be mapped to a third feature vector. For example, the textual information describing the at least one cigarette condition present in the target package can be converted into a third feature vector in numerical form using a word vector model such as Word2Vec.
[0137] Furthermore, in this embodiment of the present application, the third feature vector can be fused with the first feature vector processed by the first convolution unit and the second feature vector processed by the second convolution unit to obtain a fourth feature vector. This fourth feature vector is then processed by the first pooling unit and the first fully connected network to ultimately obtain a cigarette quantity recognition result.
[0138] In an embodiment of the present application, the specific function of the first pooling unit can refer to the above-mentioned second pooling unit, and the composition structure and function of the first fully connected neural network can also refer to the above-mentioned second fully connected neural network. The difference between the two is that the output layer of the first fully connected neural network can output a numerical value for representing the number of cigarettes contained in the target package (for example, using a linear activation function), while the output layer of the second fully connected neural network outputs a confidence value between 0 and 1 for reflecting that the target package is a cigarette package.
[0139] It is understood that in the embodiment of the present application, the cigarette quantity recognition model can be trained to acquire the above-mentioned ability to recognize the number of cigarettes contained in the target package. Figure 2 The following briefly introduces the training process of the cigarette quantity recognition model:
[0140] In an embodiment of the present application, dual-energy computed tomography sample data and millimeter wave detection sample data collected for a sample package, as well as a cigarette quantity label corresponding to the sample package, can be obtained, wherein the dual-energy computed tomography sample data is used to reflect the structural characteristics inside the sample package, the millimeter wave detection sample data is used to reflect the dielectric properties of the material inside the sample package, and the cigarette quantity label is used to reflect the number of cigarettes contained in the sample package.
[0141] Furthermore, in an embodiment of the present application, the sample density distribution matrix and the sample equivalent material ordinal matrix corresponding to the sample package can be obtained based on the dual-energy computed tomography sample data, and the sample three-dimensional dielectric constant tensor corresponding to the sample package can be obtained based on the millimeter wave detection sample data (the specific processing process can refer to the relevant contents of the above-mentioned density distribution matrix, equivalent material ordinal matrix and three-dimensional dielectric constant tensor).
[0142] Furthermore, in an embodiment of the present application, the sample density distribution matrix and the sample equivalent material number matrix can be input into the first convolution unit for feature extraction to obtain a first sample feature vector. The sample three-dimensional dielectric constant tensor can be input into the second convolution unit for feature extraction to obtain a second sample feature vector.
[0143] Furthermore, three-dimensional contour fitting is performed based on the target three-dimensional reconstructed image corresponding to the dual-energy computed tomography sample data, and the similarity between the fitted three-dimensional contour and the preset reference three-dimensional contour is calculated to determine at least one cigarette state present in the sample package, and then the at least one cigarette state present in the sample package is mapped through a preset mapping rule (such as the above-mentioned Word2Vec and other word vector models) to obtain a third sample feature vector.
[0144] Furthermore, in the embodiment of the present application, the third sample feature vector may be subjected to feature fusion with the first sample feature vector and the second sample feature vector to obtain a fourth sample feature vector.
[0145] Furthermore, in the embodiment of the present application, the fourth sample feature vector may be processed by the first pooling unit and the first fully connected neural network to obtain the predicted number of cigarettes.
[0146] Furthermore, in an embodiment of the present application, a first sum of the differences between the predicted number of cigarettes and the cigarette number labels corresponding to all sample packages can be calculated, and the parameters of the first convolution unit, the second convolution unit, the first pooling unit and the first fully connected neural network can be iteratively adjusted with the goal of reducing the first sum until the preset training conditions are met (for example, the first sum is reduced to a preset value) to obtain the trained cigarette number recognition model.
[0147] More details about the training process of the cigarette quantity recognition model can be regarded as prior art and will not be discussed in detail in this specification.
[0148] The following is the package identification model involved in the above content (i.e. Figure 3 The training process of the package recognition model shown in Figure 1 is also briefly introduced:
[0149] In an embodiment of the present application, dual-energy computed tomography sample data and millimeter wave detection sample data collected for a sample package, as well as a package label corresponding to the sample package, can be obtained, wherein the dual-energy computed tomography sample data is used to reflect the structural characteristics inside the sample package, the millimeter wave detection sample data is used to reflect the dielectric properties of the material inside the sample package, and the package label is used to reflect whether the sample package is a cigarette package.
[0150] Furthermore, in an embodiment of the present application, the sample density distribution matrix and the sample equivalent material ordinal matrix corresponding to the sample package can be obtained based on the dual-energy computed tomography sample data, and the sample three-dimensional dielectric constant tensor corresponding to the sample package can be obtained based on the millimeter wave detection sample data (the specific processing process can refer to the relevant contents of the above-mentioned density distribution matrix, equivalent material ordinal matrix and three-dimensional dielectric constant tensor).
[0151] Furthermore, in an embodiment of the present application, the sample density distribution matrix and the sample equivalent material number matrix can be input into the third convolution unit for feature extraction to obtain a first sample feature vector. The sample three-dimensional dielectric constant tensor can be input into the fourth convolution unit for feature extraction to obtain a second sample feature vector.
[0152] Furthermore, in an embodiment of the present application, feature fusion can be performed on the first sample feature vector and the second sample feature vector to obtain a sample multimodal feature fusion vector, and the sample multimodal feature fusion vector can be processed by the second pooling unit and the second fully connected neural network to obtain a predicted package recognition result.
[0153] Furthermore, in an embodiment of the present application, a second sum of the differences between the predicted package identification results and the package labels corresponding to all sample packages can be calculated, and the parameters of the third convolution unit, the fourth convolution unit, the second pooling unit, and the second fully connected neural network can be iteratively adjusted with the goal of reducing the second sum until the preset training condition is met (for example, the second sum is reduced to a preset value) to obtain the trained package identification model.
[0154] More details about the training process of the package recognition model can be regarded as prior art and will not be discussed in detail in this specification.
[0155] Figure 4 This is a module diagram of a cigarette package recognition system based on multimodal feature fusion according to some embodiments of this specification. In some embodiments, Figure 4 The cigarette package identification system 200 based on multimodal feature fusion shown can be implemented in software and / or hardware. For example, it can be configured in the form of software and / or hardware to a processing device and / or terminal device for processing dual-energy computed tomography data and millimeter wave detection data collected for a target package, and determining whether the target package is a cigarette package.
[0156] Reference Figure 4 In some embodiments, the cigarette package identification system 200 based on multimodal feature fusion may include an acquisition module 210, a first processing module 220, a first feature extraction module 230, a second feature extraction module 240, a multimodal feature fusion module 250, and a second processing module 260. Among them:
[0157] The acquisition module 210 may be configured to acquire dual-energy computed tomography data and millimeter wave detection data collected for a target package.
[0158] The first processing module 220 can be used to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data, and obtain a three-dimensional dielectric constant tensor corresponding to the target package based on the millimeter wave detection data.
[0159] The first feature extraction module 230 may be configured to perform feature extraction on the density distribution matrix and the equivalent material ordinal matrix to obtain a first feature vector.
[0160] The second feature extraction module 240 may be configured to perform feature extraction on the three-dimensional dielectric constant tensor to obtain a second feature vector.
[0161] The multimodal feature fusion module 250 may be configured to perform feature fusion on the first feature vector and the second feature vector to obtain a multimodal feature fusion vector.
[0162] The second processing module 260 may be configured to process the multimodal feature fusion vector using a pre-trained package recognition model to determine whether the target package is a cigarette package.
[0163] For more details about the above modules, please refer to other places in this manual (for example Figures 1 to 3 part and its related description), which will not be repeated here.
[0164] It should be understood that Figure 4 The illustrated cigarette package identification system 200 based on multimodal feature fusion and its modules can be implemented in various ways. For example, in some embodiments, the system and its modules can be implemented using hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or specially designed hardware. Those skilled in the art will appreciate that the above-described systems and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The systems and their modules described herein can be implemented not only using hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips or transistors, or programmable hardware devices such as field programmable gate arrays or programmable logic devices, but can also be implemented using software executed by various types of processors, or a combination of the above-described hardware circuits and software (e.g., firmware).
[0165] It should be noted that the above description of the cigarette package identification system 200 based on multimodal feature fusion is provided for illustrative purposes only and is not intended to limit the scope of this specification. It is understood that those skilled in the art can, based on the description of this specification, arbitrarily combine the modules or form subsystems connected with other modules without departing from the principles of this specification. For example, Figure 4The acquisition module 210, first processing module 220, first feature extraction module 230, second feature extraction module 240, multimodal feature fusion module 250, and second processing module 260 described above may be different modules in a system, or a single module may implement the functions of two or more of the aforementioned modules. Such variations are within the scope of protection of this application.
[0166] In summary, the beneficial effects that may be brought about by the embodiments of this specification include but are not limited to: (1) In the cigarette package identification system and method based on multimodal feature fusion provided in some embodiments of this specification, by combining millimeter wave detection data with dual-energy computed tomography data, more comprehensive and accurate feature information and judgment basis can be provided for cigarette package identification, thereby improving the accuracy and reliability of cigarette package identification results to a certain extent; (2) In the cigarette package identification system and method based on multimodal feature fusion provided in some embodiments of this specification, by obtaining the density distribution matrix and equivalent material ordinal matrix corresponding to the target package based on dual-energy computed tomography data, the target package is obtained through millimeter wave detection data. The corresponding three-dimensional dielectric constant tensor can more accurately and comprehensively reflect the structural characteristics and material dielectric properties inside the target package, thereby providing a more reliable judgment basis for the subsequent identification process; (3) In the cigarette package identification system and method based on multimodal feature fusion provided in some embodiments of this specification, by combining the third eigenvector obtained based on at least one cigarette state existing in the target package with the first eigenvector obtained based on the density distribution matrix and equivalent material ordinal matrix corresponding to the target package, and the second eigenvector obtained based on the three-dimensional dielectric constant tensor corresponding to the target package to realize cigarette quantity recognition, the probability of misidentification can be reduced to a certain extent, thereby improving the accuracy of the cigarette quantity recognition result.
[0167] It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects that may be produced may be any one or a combination of the above, or any other possible beneficial effects.
[0168] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit this specification. Although not explicitly stated herein, various modifications, improvements, and revisions to this specification may be made by those skilled in the art. Such modifications, improvements, and revisions are suggested in this specification and remain within the spirit and scope of the exemplary embodiments of this specification.
[0169] This specification also uses specific terms to describe the embodiments of this specification. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "one embodiment," "an embodiment," or "an alternative embodiment" two or more times in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics of one or more embodiments of this specification may be appropriately combined.
[0170] In addition, it will be understood by those skilled in the art that various aspects of this specification may be illustrated and described by a number of patentable categories or situations, including any new and useful process, machine, product or combination of substances, or any new and useful improvements thereto. Accordingly, various aspects of this specification may be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software may be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". In addition, various aspects of this specification may be represented as a computer product located in one or more computer-readable media, which includes computer-readable program code.
[0171] A computer storage medium may include a propagated data signal embodying the computer program code, for example, in baseband or as part of a carrier wave. The propagated signal may be in a variety of forms, including electromagnetic, optical, or any suitable combination thereof. A computer storage medium may be any computer-readable medium other than a computer-readable storage medium that can be connected to an instruction execution system, apparatus, or device to communicate, propagate, or transfer the program for use. The program code on the computer storage medium may be transmitted via any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of these.
[0172] The computer program codes required for the operation of the various parts of this specification can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C, Visual Basic, Fortran2003, Perl, COBOL2002, PHP, ABAP, dynamic programming languages such as Python, Ruby and Groovy, or other programming languages. The program code can be run entirely on the user's computer, or as a separate software package on the user's computer, or partly on the user's computer and partly on a remote computer, or entirely on a remote computer or processing device. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).
[0173] In addition, unless expressly stated in the claims, the order of the processing elements and sequences, the use of alphanumeric characters, or the use of other names described in this specification are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some embodiments of the invention that are currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the spirit and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing processing device or mobile device.
[0174] Similarly, it should be noted that, in order to simplify the presentation of this specification and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.
[0175] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.
Claims
1. A cigarette package recognition method based on multimodal feature fusion, characterized in that: include: Acquiring dual-energy computed tomography data and millimeter wave detection data collected for the target package, wherein the dual-energy computed tomography data is used to reflect the structural characteristics of the interior of the target package, and the millimeter wave detection data is used to reflect the dielectric properties of the material inside the target package; Obtaining a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data, and obtaining a three-dimensional dielectric constant tensor corresponding to the target package based on the millimeter wave detection data; Performing feature extraction on the density distribution matrix and the equivalent material ordinal matrix to obtain a first eigenvector; Performing feature extraction on the three-dimensional dielectric constant tensor to obtain a second eigenvector; Performing feature fusion on the first feature vector and the second feature vector to obtain a multimodal feature fusion vector; The multimodal feature fusion vector is processed by a pre-trained package recognition model to determine whether the target package is a cigarette package.
2. The cigarette package identification method based on multimodal feature fusion according to claim 1, characterized in that: The obtaining of a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data includes: Performing image reconstruction based on the dual-energy computed tomography data to obtain a target three-dimensional reconstructed image; Density analysis and equivalent material ordinal analysis are performed on each voxel unit in the area corresponding to the target package in the target three-dimensional image according to the dual-energy computed tomography data to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package.
3. The cigarette package identification method based on multimodal feature fusion according to claim 2, characterized in that: The method of performing density analysis and equivalent material ordinal analysis on each voxel unit in the region corresponding to the target package in the target three-dimensional image according to the dual-energy computed tomography data to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package includes: For each target voxel cell, do the following: Step 1, determining the first attenuation coefficient and the second attenuation coefficient of the target voxel unit under two different X-ray energies; Step 2: Substituting the first attenuation coefficient and the second attenuation coefficient into a preset material decomposition equation according to a dual-energy attenuation model to calculate a physical density value of the target voxel unit; Step 3, calculating the equivalent material ordinal value corresponding to the target voxel unit according to the physical density value of the target voxel unit and the preset mapping relationship between the equivalent material ordinal number and the physical density; The physical density values and equivalent material ordinal values corresponding to all target voxel units are respectively combined according to the spatial relationship in the target package to obtain the density distribution matrix and equivalent material ordinal matrix corresponding to the target package.
4. The cigarette package identification method based on multimodal feature fusion according to claim 1, characterized in that: The obtaining of a three-dimensional dielectric constant tensor corresponding to the target package based on the millimeter wave detection data includes: Performing phase unwrapping and scattered field inversion on the millimeter wave detection data to generate a three-dimensional complex permittivity distribution corresponding to the target package; extracting a real permittivity component from the complex permittivity distribution; A three-dimensional coordinate system is constructed according to the spatial relationship between multiple voxel units contained in the target package, and the extracted real dielectric constant component is mapped to the corresponding position in the three-dimensional coordinate system to obtain a three-dimensional dielectric constant tensor corresponding to the target package.
5. The cigarette package identification method based on multimodal feature fusion according to claim 1, characterized in that: The output of the package recognition model is a confidence value that the target package is a cigarette package. The multimodal feature fusion vector is processed by the pre-trained package recognition model to determine whether the target package is a cigarette package, including: When the confidence value is greater than or equal to a first preset threshold, determining that the target package is a cigarette package and triggering a cigarette quantity identification process; When the confidence value is greater than or equal to the second preset threshold and less than the first preset threshold, prompting relevant staff to perform manual verification; When the confidence value is less than the second preset threshold, the target package is determined to be a non-cigarette package.
6. The cigarette package identification method based on multimodal feature fusion according to claim 5, characterized in that: The cigarette quantity identification process includes: performing three-dimensional contour fitting based on the target three-dimensional reconstructed image corresponding to the dual-energy computed tomography data, and calculating similarity between the fitted three-dimensional contour and a preset reference three-dimensional contour to determine at least one cigarette state present in the target package; According to the number of target objects corresponding to each cigarette state and the number of reference cigarettes corresponding to each cigarette state, an estimation result of the number of cigarettes corresponding to the target package is obtained.
7. The cigarette package identification method based on multimodal feature fusion according to claim 5, characterized in that: The cigarette quantity identification process includes: performing three-dimensional contour fitting based on the target three-dimensional reconstructed image corresponding to the dual-energy computed tomography data, and calculating similarity between the fitted three-dimensional contour and a preset reference three-dimensional contour to determine at least one cigarette state present in the target package; Mapping the at least one cigarette state in the target package into a third feature vector, and performing feature fusion on the third feature vector, the first feature vector, and the second feature vector to obtain a fourth feature vector; The fourth eigenvector is processed by a pre-trained cigarette quantity recognition model to obtain an estimation result of the cigarette quantity corresponding to the target package.
8. The cigarette package identification method based on multimodal feature fusion according to claim 7, characterized in that: The cigarette quantity recognition model includes a first convolution unit, a second convolution unit, a first pooling unit, and a first fully connected neural network. The cigarette quantity recognition model is trained based on the following method: Obtaining dual-energy computed tomography (DECT) sample data and millimeter wave detection sample data collected for a sample package, as well as a cigarette quantity label corresponding to the sample package, wherein the DECT sample data is used to reflect the structural characteristics of the interior of the sample package, the millimeter wave detection sample data is used to reflect the dielectric properties of the material inside the sample package, and the cigarette quantity label is used to reflect the number of cigarettes contained in the sample package; Obtaining a sample density distribution matrix and a sample equivalent material ordinal matrix corresponding to the sample package based on the dual-energy computed tomography sample data, and obtaining a sample three-dimensional dielectric constant tensor corresponding to the sample package based on the millimeter wave detection sample data; Inputting the sample density distribution matrix and the sample equivalent material ordinal matrix into the first convolution unit for feature extraction to obtain a first sample feature vector; Inputting the sample three-dimensional dielectric constant tensor into the second convolution unit for feature extraction to obtain a second sample feature vector; performing three-dimensional contour fitting based on a target three-dimensional reconstructed image corresponding to the dual-energy computed tomography sample data, and calculating a similarity between the fitted three-dimensional contour and a preset reference three-dimensional contour to determine at least one cigarette state present in the sample package, and then mapping the at least one cigarette state present in the sample package using a preset mapping rule to obtain a third sample feature vector; Performing feature fusion on the third sample feature vector, the first sample feature vector, and the second sample feature vector to obtain a fourth sample feature vector; Processing the fourth sample feature vector using the first pooling unit and the first fully connected neural network to obtain a predicted number of cigarettes; A first sum of the differences between the predicted number of cigarettes and the cigarette number labels corresponding to all sample packages is calculated, and the parameters of the first convolution unit, the second convolution unit, the first pooling unit, and the first fully connected neural network are iteratively adjusted with the goal of reducing the first sum until the trained cigarette number recognition model is obtained when the preset training conditions are met.
9. The cigarette package recognition method based on multimodal feature fusion according to any one of claims 1 to 8, characterized in that: The package recognition model includes a third convolutional unit, a fourth convolutional unit, a second pooling unit, and a second fully connected neural network. The package recognition model is trained based on the following method: Obtaining dual-energy computed tomography sample data and millimeter wave detection sample data collected for a sample package, as well as a package label corresponding to the sample package, wherein the dual-energy computed tomography sample data is used to reflect the structural characteristics of the interior of the sample package, the millimeter wave detection sample data is used to reflect the dielectric properties of the material inside the sample package, and the package label is used to indicate whether the sample package is a cigarette package; Obtaining a sample density distribution matrix and a sample equivalent material ordinal matrix corresponding to the sample package based on the dual-energy computed tomography sample data, and obtaining a sample three-dimensional dielectric constant tensor corresponding to the sample package based on the millimeter wave detection sample data; Inputting the sample density distribution matrix and the sample equivalent material ordinal matrix into the third convolution unit for feature extraction to obtain a first sample feature vector; Inputting the sample three-dimensional dielectric constant tensor into the fourth convolution unit for feature extraction to obtain a second sample feature vector; Performing feature fusion on the first sample feature vector and the second sample feature vector to obtain a sample multimodal feature fusion vector; Processing the sample multimodal feature fusion vector through the second pooling unit and the second fully connected neural network to obtain a predicted package recognition result; A second sum of the differences between the predicted package identification results and the package labels corresponding to all sample packages is calculated, and the parameters of the third convolution unit, the fourth convolution unit, the second pooling unit, and the second fully connected neural network are iteratively adjusted with the goal of reducing the second sum until the trained package identification model is obtained when the preset training conditions are met.
10. A cigarette package recognition system based on multimodal feature fusion, characterized in that: include: an acquisition module, configured to acquire dual-energy computed tomography data and millimeter-wave detection data collected for a target package, wherein the dual-energy computed tomography data is used to reflect the structural characteristics of the interior of the target package, and the millimeter-wave detection data is used to reflect the dielectric properties of the material inside the target package; a first processing module, configured to obtain a density distribution matrix and an equivalent material ordinal matrix corresponding to the target package based on the dual-energy computed tomography data, and obtain a three-dimensional dielectric constant tensor corresponding to the target package based on the millimeter wave detection data; A first feature extraction module is used to extract features from the density distribution matrix and the equivalent material ordinal matrix to obtain a first feature vector; A second feature extraction module is used to extract features from the three-dimensional dielectric constant tensor to obtain a second feature vector; a multimodal feature fusion module, configured to perform feature fusion on the first feature vector and the second feature vector to obtain a multimodal feature fusion vector; The second processing module is used to process the multimodal feature fusion vector using a pre-trained package recognition model to determine whether the target package is a cigarette package.