Intelligent multi-mode classification detection system and method
Through the intelligent multimodal classification detection system, the multi-sensor network and multimodal data fusion technology are used to solve the problem of poor adaptability of the classification detection system in the existing technology, and the high accuracy identification and positioning of prohibited items are achieved, and security inspection efficiency is improved.
Patent Information
- Application Number
- CN202510415280.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing classification detection system is detected by electromagnetic induction and ion migration spectrum technology, it is easily disturbed by environmental metal objects, resulting in poor adaptability, false alarms and missed alarms.
The intelligent multimodal classification detection system is adopted, including a detection unit module, a data fusion module, an item classification module, a prohibited positioning module and a detection report generation module. It identifies and locates prohibited items through multi-sensor network, multi-modal data feature extraction and fusion, and preset classification recognition algorithms.
It improves the adaptability and accuracy of the classification detection system, reduces false alarms and missed reports, quickly identify the location of prohibited items, and reduces manual inspection time and labor.
Smart Images

Figure CN120296472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data detection, and in particular, to an intelligent multi-modal classification detection system and method. Background Art
[0002] A classification detection system refers to a technical system that can identify, classify, and detect target objects. Its purpose is to distinguish different types of objects from complex signals or data and mark or alarm them. In the field of security inspection, the classification detection system can effectively identify potential dangerous items such as explosives, drugs, weapons, etc., thereby improving the public security level.
[0003] Currently, the classification detection system detects through the principle of electromagnetic induction and ion mobility spectrometry technology. This method can only detect target contraband items singly and is easily interfered by environmental metal items, resulting in false alarms, which leads to poor adaptability of the classification detection system and is prone to false alarms and missed detections. Summary of the Invention
[0004] The present invention provides an intelligent multi-modal classification detection system and method, whose main purpose is to improve the adaptability and accuracy of the classification detection system.
[0005] To achieve the above object, an intelligent multi-modal classification detection system provided by the present invention includes: a detection unit module, a data fusion module, an item classification module, a contraband positioning module, and a detection report generation module;
[0006] The material determination module is used to clarify the detection requirements of the classification detector, configure the multi-sensors of the classification detector according to the detection requirements, construct a sensor network of the multi-sensors, and construct a detection unit of the classification detector based on the sensors and the sensor network;
[0007] The data fusion module is used to collect the item multi-modal data and face images of the corresponding detection objects of the classification detector according to the detection unit, extract the multi-modal data features of the item multi-modal data, perform data fusion on the item multi-modal data based on the multi-modal data features to obtain item fusion data, extract the image features of the face images, and analyze the personal information of the detection objects based on the image features;
[0008] The item classification module is used to analyze the item features of the items carried by the detection object according to the item fusion data, identify the item types of the carried items by using a preset classification and recognition algorithm based on the item features, analyze the risk coefficient of the carried items according to the item types, and determine the prohibited level of the carried items based on the risk coefficient, where the prohibited levels include: regular items, mildly prohibited items, moderately prohibited items, and severely prohibited items;
[0009] The prohibited item positioning module is used to mark the carried items according to the prohibited level to obtain prohibited marked items, construct a human detection model of the classification detector, and map the prohibited marked items into the human detection model to obtain the positions of the prohibited items;
[0010] The detection report generation module is used to integrate the multi-modal detection system of the classification detector and generate a detection report of the multi-modal detection system based on the positions of the prohibited items, the prohibited levels, the item types, and the person information.
[0011] Optionally, constructing the sensor network of the multi-sensors includes:
[0012] Analyze the node density, network distribution range, and transmission rate requirements of the multi-sensors;
[0013] Determine the structural layout of the multi-sensors according to the node density and the network distribution range;
[0014] Determine the network communication protocol and network topology of the multi-sensors based on the transmission rate requirements;
[0015] Configure the network parameters of the multi-sensors, where the network parameters include: node ID, communication frequency, and data transmission rate;
[0016] Integrate the sensor network of the multi-sensors according to the structural layout, the network communication protocol, the network topology, and the network parameters.
[0017] Optionally, constructing the detection unit of the classification detector based on the sensors and the sensor network includes:
[0018] Construct the detection unit framework of the classification detector and determine the sensor interface of the detection unit framework based on the sensors;
[0019] Construct the communication interface of the detection unit framework according to the sensor network;
[0020] Plan the signal processing circuit of the communication interface and configure the microcontroller of the detection unit framework;
[0021] Integrate the microcontroller, the signal processing circuit, the sensor interface, and the communication interface into the detection unit framework to obtain an initial detection unit;
[0022] Test the detection performance of the initial detection unit. When the detection performance meets the preset detection performance standard, use the initial detection unit as the detection unit of the classification detector.
[0023] Optionally, the data fusion of the item multimodal data based on the multimodal data features to obtain item fusion data includes:
[0024] Based on the multimodal data features, determine the characteristic probability distribution of the item multimodal data;
[0025] Analyze the modal data weights of the item multimodal data;
[0026] According to the characteristic probability distribution and the modal data weights, calculate the fusion probability distribution of the item multimodal data;
[0027] According to the fusion probability distribution, perform data fusion on the item multimodal data to obtain item fusion data.
[0028] Optionally, the extraction of the image features of the face image includes:
[0029] Locate the face area of the face image and detect the landmark points of the face area;
[0030] Based on the landmark points, align the face image with the preset standard face position to obtain an aligned image;
[0031] Use the preset Laplacian algorithm to detect the face edge features of the aligned image;
[0032] Calculate the LBP value of the corresponding pixels of the aligned image, and based on the LBP value, extract the texture features of the aligned image;
[0033] According to the texture features and the face edge features, determine the image features of the aligned image.
[0034] Optionally, the analysis of the item features of the items carried by the detection object according to the item fusion data includes:
[0035] Construct an initialization global analysis model of the classification detector corresponding to the detection object, and extract the initialization model parameters of the initialization global analysis model;
[0036] According to the initialized model parameters, train the local node analysis model of the classification detector using a preset local training set;
[0037] Calculate the loss function of the local node analysis model, and based on the loss function, determine the local model parameters of the local node analysis model;
[0038] According to the local model parameters, calculate the global model parameters of the initialized global analysis model using a preset aggregation algorithm;
[0039] Based on the global model parameters, update the initialized global analysis model to obtain an updated global model;
[0040] Analyze the model performance of the updated global model. When the model performance meets the preset model performance standard, use the updated global model as the global analysis model of the classification detector;
[0041] Based on the item fusion data, analyze the item characteristics of the carried item using the global analysis model.
[0042] Optionally, the analyzing the risk coefficient of the carried item according to the item type includes:
[0043] Based on the item type, determine the potential dangerous characteristics of the carried item, and calculate the exposure frequency of the carried item;
[0044] According to the potential dangerous characteristics, analyze the probability of danger occurrence, the consequences of danger, and the effectiveness of control measures of the carried item;
[0045] Determine the risk consequence weight of the risk consequence;
[0046] According to the exposure frequency, the probability of danger occurrence, the consequences of danger, the risk consequence weight, and the effectiveness of control measures, calculate the risk coefficient of the carried item.
[0047] Optionally, the mapping the prohibited marked item to the human body detection model to obtain the prohibited item position includes:
[0048] Construct a three-dimensional coordinate system of the human body detection model;
[0049] Determine the two-dimensional position of the prohibited item of the prohibited marked item;
[0050] Obtain the imaging geometric parameters of the imaging device corresponding to the prohibited marked item;
[0051] According to the imaging geometric parameters, convert the two-dimensional position of the prohibited item into the three-dimensional space coordinates of the prohibited item;
[0052] Map the three-dimensional spatial coordinates into the human detection model according to the three-dimensional coordinate system to obtain the contraband position.
[0053] Optionally, generating a detection report of the multi-modal detection system based on the contraband position, the contraband level, the item type, and the person information, including:
[0054] Determine the alarm level of the classification detector according to the contraband level and the item type;
[0055] Determine the sound and light alarm mode of the classification detector according to the alarm level;
[0056] Determine the target person of the classification detector corresponding to the detection object according to the person information;
[0057] Based on the contraband position, determine the specific hiding position of the contraband corresponding to the target person;
[0058] Determine the detailed description and image evidence of the contraband;
[0059] Based on the sound and light alarm mode, the target person, the specific hiding position, the detailed description of the contraband, and the image evidence, construct the detection report of the classification detector.
[0060] An intelligent multi-modal classification detection method, characterized in that the method includes:
[0061] Clarify the detection requirements of the classification detector, configure the multi-sensors of the classification detector according to the detection requirements, construct the sensor network of the multi-sensors, and construct the detection unit of the classification detector based on the sensors and the sensor network;
[0062] According to the detection unit, collect the item multi-modal data and face images of the classification detector corresponding to the detection object, extract the multi-modal data features of the item multi-modal data, based on the multi-modal data features, perform data fusion on the item multi-modal data to obtain item fusion data, extract the image features of the face image, and analyze the person information of the detection object based on the image features;
[0063] According to the item fusion data, analyze the item features of the items carried by the detection object, based on the item features, use a preset classification and recognition algorithm to identify the item type of the carried items, according to the item type, analyze the danger coefficient of the carried items, and based on the danger coefficient, determine the contraband level of the carried items, where the contraband level includes: conventional items, slightly contraband items, moderately contraband items, and severely contraband items;
[0064] Mark the carried items according to the prohibited level to obtain prohibited marked items, construct a human detection model of the classification detector, and map the prohibited marked items into the human detection model to obtain the positions of prohibited items;
[0065] Integrate the multi-modal detection system of the classification detector, and generate a detection report of the multi-modal detection system based on the positions of prohibited items, the prohibited levels, the types of items, and the personal information.
[0066] In the embodiment of the present invention, by constructing the sensor network of the multi-sensors, a wider area can be covered, the monitoring of a large-range environment can be realized, and at the same time, fast data collection and processing can be achieved, improving the response speed of the system; optionally, in the embodiment of the present invention, by constructing the sensor network of the multi-sensors, a wider area can be covered, the monitoring of a large-range environment can be realized, and at the same time, fast data collection and processing can be achieved, improving the response speed of the system; in the embodiment of the present invention, by using a preset classification and recognition algorithm to identify the type of the carried item based on the item features, rich item features (such as appearance, material, function, behavior features) can be extracted, and the classification and recognition algorithm can more accurately identify the type of item; in the embodiment of the present invention, by using a preset classification and recognition algorithm to identify the type of the carried item based on the item features, rich item features (such as appearance, material, function, behavior features) can be extracted, and the classification and recognition algorithm can more accurately identify the type of item; in the embodiment of the present invention, by mapping the prohibited marked items into the human detection model, the positions of prohibited items can be accurately determined, the security inspection process is accelerated, the time and labor required for manual inspection are reduced, and finally, in the embodiment of the present invention, by generating a detection report of the multi-modal detection system based on the positions of prohibited items, the prohibited levels, the types of items, and the personal information, detailed prohibited item position information can be provided, enabling security inspection personnel to quickly and accurately locate and confirm prohibited items. Therefore, the adaptability and accuracy of the classification detection system are improved. Description of the Drawings
[0067] Figure 1 It is a functional module diagram of an intelligent multi-modal classification detection system provided by an embodiment of the present invention;
[0068] Figure 2 It is a flowchart of an intelligent multi-modal classification detection method provided by an embodiment of the present invention;
[0069] The realization, functional features and advantages of the purpose of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Detailed Embodiments
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0071] In addition, the sequence of steps in the following method embodiments is only an example and is not strictly limited.
[0072] In fact, the server device deployed by the intelligent multi-modal classification detection system may be composed of one or more devices. The above intelligent multi-modal classification detection system can be implemented as: a service instance, a virtual machine, or a hardware device. For example, the intelligent multi-modal classification detection system can be implemented as a service instance deployed on one or more devices in a cloud node. Simply put, the intelligent multi-modal classification detection system can be understood as a software deployed on a cloud node for providing intelligent multi-modal classification detection services to each client. Alternatively, the intelligent multi-modal classification detection system can also be implemented as a virtual machine deployed on one or more devices in a cloud node. An application software for managing each client is installed in the virtual machine. Or, the intelligent multi-modal classification detection system can also be implemented as a server composed of many identical or different types of hardware devices, and one or more hardware devices are set to provide intelligent multi-modal classification detection services to each client.
[0073] In terms of implementation form, the intelligent multi-modal classification detection system and the client adapt to each other. That is, if the intelligent multi-modal classification detection system is an application installed on a cloud service platform, then the client is a client that establishes a communication connection with the application; or if the intelligent multi-modal classification detection system is implemented as a website, then the client is implemented as a web page; or if the intelligent multi-modal classification detection system is implemented as a cloud service platform, then the client is implemented as a small program in an instant messaging application.
[0074] Refer to Figure 1 As shown, it is a functional module diagram of the intelligent multi-modal classification detection system provided by an embodiment of the present invention.
[0075] The intelligent multi-modal classification and detection system 100 according to the present invention can be set in a cloud server. In terms of implementation form, it can be used as one or more service devices, or can be installed as an application on the cloud (such as a server for intelligent multi-modal classification and detection, a server cluster, etc.), or can also be developed into a website. According to the functions achieved, the intelligent multi-modal classification and detection system 100 includes a detection unit module 101, a data fusion module 102, an item classification module 103, a prohibited item positioning module 104, and a detection report generation module 105.
[0076] In the embodiments of the present invention, in the tracking based on intelligent multi-modal classification and detection, each of the above modules can be independently implemented and called with other modules. Here, the call can be understood as that a certain module can be connected to multiple modules of another type and provide corresponding services for the multiple modules it is connected to. In the intelligent multi-modal classification and detection system provided by the embodiments of the present invention, without modifying the program code, the applicable scope of the intelligent multi-modal classification and detection architecture can be adjusted by adding modules and directly calling, so as to achieve cluster-level horizontal expansion, so as to achieve the purpose of quickly and flexibly expanding the intelligent multi-modal classification and detection system. In practical applications, the above modules can be set in the same device or different devices, or can also be set in virtual devices, such as service instances in a cloud server.
[0077] Next, specific embodiments will be used to illustrate each component and the specific working process of the intelligent multi-modal classification and detection system respectively.
[0078] The detection unit module 101 is used to clarify the detection requirements of the classification detector, configure the multi-sensors of the classification detector according to the detection requirements, construct a sensor network of the multi-sensors, and construct a detection unit of the classification detector based on the sensors and the sensor network.
[0079] In the embodiments of the present invention, by clarifying the detection requirements of the classification detector, corresponding sensors can be configured to achieve high recognition rate and low false alarm rate, and improve the accuracy of detection results. Among them, the detection requirements refer to a series of technical indicators and functional requirements that need to be specifically clarified when designing, developing, and deploying the classification detector.
[0080] In the embodiments of the present invention, by configuring the multi-sensors of the classification detector according to the detection requirements, different types of information of the target (such as vision, thermal infrared, radar, etc.) can be collected simultaneously, so as to obtain more comprehensive data, improve the recognition ability of the target, and reduce the limitations that may exist in a single sensor. Among them, the multi-sensors refer to a variety of different types or functions of sensors integrated to achieve specific detection and classification tasks, such as: optical sensors, chemical sensors, biological sensors, etc.
[0081] In the embodiments of the present invention, by constructing the sensor network of the multi-sensors, a wider area can be covered, the monitoring of a large-scale environment can be realized, and at the same time, rapid data collection and processing can be achieved, improving the response speed of the system. Among them, the sensor network refers to a system composed of multiple sensor nodes, and these nodes are interconnected by wired or wireless means and cooperate with each other to collect, process and transmit information.
[0082] As an embodiment of the present invention, the construction of the sensor network of the multi-sensors includes:
[0083] Analyze the node density, network distribution range and transmission rate requirements of the multi-sensors;
[0084] Determine the structural layout of the multi-sensors according to the node density and the network distribution range;
[0085] Based on the transmission rate requirements, determine the network communication protocol and network topology of the multi-sensors;
[0086] Configure the network parameters of the multi-sensors, where the network parameters include: node ID, communication frequency and data transmission rate;
[0087] Integrate the sensor network of the multi-sensors according to the structural layout, the network communication protocol, the network topology, and the network parameters.
[0088] Among them, the node density refers to the number of sensor nodes per unit area or volume. The network distribution range refers to the spatial range covered by the sensor network. The transmission rate requirement refers to the speed requirement for data transmission in the sensor network, which determines the speed at which data is transmitted from the sensor node to the data processing center. The structural layout refers to the spatial arrangement and configuration method of each sensor node in the sensor network. The network communication protocol refers to the format of data exchange, transmission method, error detection and correction mechanism, and other agreements for communication between network devices in the sensor network, such as ZigBee, 6LoWPAN, Wi-Fi, etc. The network topology refers to the physical layout and connection method of each node (sensor, processor, communication device, etc.) in the sensor network, such as star topology, ring topology, tree topology, etc. The network parameters refer to the key parameters that need to be configured and optimized during network communication. The node ID refers to the unique identifier of each sensor node in the network. The communication frequency refers to the radio frequency used by the sensor node for wireless communication. The data transmission rate refers to the speed of data transmission in the network.
[0089] Optionally, determining the structural layout of the multi-sensor according to the node density and the network distribution range can be determined by a genetic algorithm. For example, through the genetic algorithm, natural selection and genetic mechanisms are simulated to optimize the sensor layout.
[0090] In the embodiment of the present invention, by constructing the detection unit of the classification detector based on the sensor and the sensor network, multiple physical parameters can be monitored simultaneously, and the data collected by the sensor can be quickly responded to and processed, providing a rich information basis for classification. Among them, the detection unit refers to an independent module responsible for sensing, collecting, processing, and transmitting information in the classification detector system.
[0091] As an embodiment of the present invention, constructing the detection unit of the classification detector based on the sensor and the sensor network includes:
[0092] Constructing the detection unit framework of the classification detector, and determining the sensor interface of the detection unit framework based on the sensor;
[0093] Constructing the communication interface of the detection unit framework according to the sensor network;
[0094] Planning the signal processing circuit of the communication interface and configuring the microcontroller of the detection unit framework;
[0095] Integrating the microcontroller, the signal processing circuit, the sensor interface, and the communication interface into the detection unit framework to obtain an initial detection unit;
[0096] Testing the detection performance of the initial detection unit, and when the detection performance meets the preset detection performance standard, using the initial detection unit as the detection unit of the classification detector.
[0097] Among them, the detection unit framework refers to the physical structure and electronic architecture of the detection unit. The sensor interface refers to the hardware part in the detection unit for connecting the sensor and exchanging data with the sensor. The communication interface refers to the hardware and software components for data transmission and interaction between the device and the external network, such as a wireless communication interface, a fiber optic communication interface, etc. The signal processing circuit refers to an electronic circuit for operations such as receiving, amplifying, filtering, converting, encoding, decoding, modulating, and demodulating, such as a signal amplifier, a filter, a signal converter, etc. The microcontroller refers to a microcomputer system integrating a central processing unit (CPU), memory (including random access memory RAM and read-only memory ROM), input / output (I / O) interfaces, and other auxiliary functions.
[0098] Optionally, testing the detection performance of the initial detection unit includes:
[0099] Construct a variety of detection modules for the initial detection unit;
[0100] Use the oscilloscope corresponding to the variety of detection modules to detect the signal waveform of the initial detection unit;
[0101] Determine the circuit operating state of the initial detection unit according to the signal waveform;
[0102] Use the signal generator corresponding to the variety of detection modules to test the performance of the electronic components of the initial detection unit;
[0103] Use the spectrum analyzer corresponding to the variety of detection modules to analyze the electromagnetic signals generated during the operation of the initial detection unit;
[0104] Detect the interference signals of the electromagnetic signals and determine the electromagnetic compatibility of the initial detection unit under the influence of the interference signals;
[0105] Determine the detection performance of the initial detection unit according to the circuit operating state, the performance of the electronic components, and the electromagnetic compatibility.
[0106] Among them, the variety of detection modules refer to a set of functional modules used for multi-dimensional performance detection and evaluation of the initial detection unit, such as oscilloscopes, signal generators, spectrum analyzers, etc. The oscilloscope refers to an electronic test instrument used for observing, measuring, and analyzing the waveforms of electrical signals. The signal waveform refers to the graphical representation of the variation of electrical signals (such as voltage, current) over time. The circuit operating state refers to the operating conditions and performance of the circuit under specific conditions, reflecting whether the circuit is operating normally, whether there are faults or abnormalities, and whether the performance of the circuit meets the design requirements, such as normal operating state, abnormal operating state, and fault state, etc. The signal generator refers to an electronic test instrument used for generating electrical signals with specific frequencies, amplitudes, and waveforms. The performance of the electronic components refers to the electrical characteristics and functional performance of electronic components (such as sensors, amplifiers, filters, etc.) under specific operating conditions, reflecting whether the components can operate normally, whether they meet the design requirements, and whether their functions in the circuit are effective. The spectrum analyzer refers to an electronic measuring instrument used for analyzing the spectrum structure of electrical signals, which can measure parameters such as signal distortion, modulation degree, spectral purity, frequency stability, and intermodulation distortion. The electromagnetic signal refers to the information carried by electromagnetic waves, which can propagate in space in the form of waves or be transmitted through conductors (such as cables). The interference signal refers to those signals that are unexpectedly mixed or superimposed on the useful signal during signal transmission or reception, such as electromagnetic interference, electrostatic interference, etc. The electromagnetic compatibility refers to the ability of the initial detection unit to operate normally in its electromagnetic environment and not generate unacceptable electromagnetic interference.
[0107] The data fusion module 102 is configured to collect multi-modal data of an item and a face image of a detection object corresponding to the classification detector according to the detection unit, extract multi-modal data features of the multi-modal data of the item, fuse the multi-modal data of the item based on the multi-modal data features to obtain fused item data, extract image features of the face image, and analyze the personal information of the detection object based on the image features.
[0108] In an embodiment of the present invention, by collecting the multi-modal data of an item and a face image of a detection object corresponding to the classification detector according to the detection unit, items can be more accurately identified and classified. Multiple data sources provide more comprehensive information and reduce the uncertainty that may be brought by a single data source. Among them, the multi-modal data of the item refers to various types of data about the item collected from different sources and different sensors, such as visual data, infrared data, electromagnetic data, etc. The face image refers to a visual representation of an individual's facial features obtained through an image capture device (such as a camera).
[0109] In an embodiment of the present invention, by extracting the multi-modal data features of the multi-modal data of the item, the use of computing resources and storage resources can be optimized, unnecessary data processing requirements can be reduced, and misidentifications caused by the limitations of single-modal data can be reduced. Among them, the multi-modal data features refer to representative attributes or metrics extracted from item information collected from different sensors or data sources, such as visual features, text features, electromagnetic features, etc.
[0110] Optionally, the extraction of the multi-modal data features of the multi-modal data of the item can be performed through end-to-end learning techniques, such as a convolutional neural network (CNN) combined with a recurrent neural network (RNN) or a Transformer, to automatically extract multi-modal data features.
[0111] In an embodiment of the present invention, by fusing the multi-modal data of the item based on the multi-modal data features to obtain fused item data, the dimension of information can be increased, enabling the classifier to capture more subtle feature differences, thereby improving the discrimination ability. Among them, the fused item data refers to a data set obtained by integrating and merging data about the same item from different sensors or different modalities through data fusion techniques.
[0112] As an embodiment of the present invention, the fusion of the multi-modal data of the item based on the multi-modal data features to obtain fused item data includes:
[0113] Determining the feature probability distribution of the multi-modal data of the item based on the multi-modal data features;
[0114] Analyze the modality data weights of the multi-modal data of the item;
[0115] According to the feature probability distribution and the modality data weights, use the following formula to calculate the fusion probability distribution of the multi-modal data of the item:
[0116]
[0117] where, G r (h) represents the fusion probability distribution, D represents the normalization factor, exp represents the exponential function with base e, n represents the number of modalities of the multi-modal data of the item, q i represents the modality data weight corresponding to the i-th modality of the multi-modal data of the item, ln represents the logarithmic function with base e, H represents the multi-modal data feature, h represents a data feature h in the multi-modal data feature, G i (h) represents the feature probability distribution of the feature h in the i-th modality data of the multi-modal data of the item;
[0118] According to the fusion probability distribution, perform data fusion on the multi-modal data of the item to obtain the fused data of the item.
[0119] Among them, the feature probability distribution refers to the probability distribution of data features in each modality data. The modality data weight refers to the weight value assigned to each modality data, which is used to represent the importance or reliability of the modality in the fusion process. The fusion probability distribution refers to the unified probability distribution obtained by integrating the probability distributions from different modality data. The normalization factor is a constant used to ensure that the fused probability distribution satisfies the basic properties of probability, that is, the sum of the probabilities of all possible values is 1.
[0120] Optionally, the analysis of the modality data weights of the multi-modal data of the item can learn the weights of each modality through a neural network (such as a fully connected network).
[0121] In the embodiment of the present invention, by extracting the image features of the face image, the accuracy of the face recognition system can be significantly improved, enabling it to accurately identify individuals under different lighting, pose, expression, and occlusion conditions. Among them, the image feature refers to the image data containing the face.
[0122] As an embodiment of the present invention, the extraction of the image features of the face image includes:
[0123] Locate the face area of the face image and detect the landmark points of the face area;
[0124] Based on the landmark points, align the face image with the preset standard face position to obtain the aligned image;
[0125] Detect the facial edge features of the aligned image using a preset Laplacian algorithm;
[0126] Calculate the LBP values of the corresponding pixels of the aligned image, and based on the LBP values, extract the texture features of the aligned image;
[0127] Determine the image features of the aligned image according to the texture features and the facial edge features.
[0128] Among them, the facial region refers to the part of the image that contains a face. The landmark points refer to the key points with clear semantic meanings in the facial image, such as eyes, nose, mouth, etc. The aligned image refers to the image after adjusting the detected facial region to a standard position and posture through image transformation (such as affine transformation or perspective transformation). The preset Laplacian algorithm refers to an edge detection algorithm for image processing, which detects edges and details in the image by calculating the second derivative of the image. The facial edge features refer to the edge information extracted from the facial image, which is used to describe the contour and structure of the face. The LBP value refers to a simple and effective feature representation method for describing the local texture features of an image. The texture features refer to the information extracted from the image that can describe the local or global texture patterns of the image.
[0129] Optionally, the facial region of the facial image can be located by a face detection algorithm, such as Haar feature, MTCNN, YOLO, etc.
[0130] In the embodiment of the present invention, by analyzing the personal information of the detected object based on the image features, it can be used in an identity authentication system to confirm personal identity. Among them, the personal information refers to various data and information related to a person that can be obtained through image analysis, such as basic identity information, biometric features, appearance features, etc.
[0131] The item classification module 103 is used to analyze the item features of the items carried by the detected object according to the item fusion data, and based on the item features, use a preset classification and recognition algorithm to identify the item type of the carried items. According to the item type, analyze the risk coefficient of the carried items, and based on the risk coefficient, determine the prohibited level of the carried items. Among them, the prohibited levels include: regular items, slightly prohibited items, moderately prohibited items, and severely prohibited items.
[0132] In an embodiment of the present invention, by analyzing the item features of the items carried by the detection object according to the item fusion data, the features of the items can be more comprehensively described by fusing multi-modal data, thereby improving the accuracy of item recognition. Among them, the item features refer to the information extracted from the items that can describe the attributes, functions or states of the items, such as appearance features, material features, function features, etc.
[0133] As an embodiment of the present invention, the analyzing the item features of the items carried by the detection object according to the item fusion data includes:
[0134] Construct an initial global analysis model for the classification detector corresponding to the detection object, and extract the initial model parameters of the initial global analysis model;
[0135] According to the initial model parameters, train the local node analysis model of the classification detector using a preset local training set;
[0136] Calculate the loss function of the local node analysis model, and based on the loss function, determine the local model parameters of the local node analysis model;
[0137] According to the local model parameters, calculate the global model parameters of the initial global analysis model using a preset aggregation algorithm;
[0138] Based on the global model parameters, update the initial global analysis model to obtain an updated global model;
[0139] Analyze the model performance of the updated global model. When the model performance meets the preset model performance standard, use the updated global model as the global analysis model of the classification detector;
[0140] Based on the item fusion data, analyze the item features of the carried items using the global analysis model.
[0141] Among them, the initialized global analysis model refers to a global feature analysis model initialized by the central server in the federated learning framework, which is used for distributed training and model aggregation among multi-node classification detectors. The initialized model parameters refer to the initial parameters set by the central server for the global analysis model, and these parameters will be distributed to all local nodes as the starting point for local training. The preset local training set refers to the private data sets owned by each local node classification detector in the federated learning framework. These data sets usually contain multi-modal fusion data and are used to train the local node analysis model. The local node analysis model refers to the machine learning model running on each local node classification detector in the federated learning framework. The loss function refers to the mathematical function used to analyze the difference between the predicted value and the actual value of the model during the machine learning process of the local node analysis model. The local model parameters refer to the specific parameters of the machine learning model trained on each local node classification detector in the federated learning framework. The preset aggregation algorithm refers to the algorithm used in federated learning to merge the model parameters of multiple local nodes to construct or update the global model, such as federated averaging algorithm, weighted federated learning, etc. The global model parameters refer to a set of parameters obtained by aggregating the model parameters of multiple local nodes in federated learning, and these parameters define the structure and behavior of the global model. The updated global model refers to the new version of the global model obtained after multiple rounds of local model training and parameter aggregation in the federated learning framework. The model performance refers to the performance level of the updated global model in the item feature analysis task, which is usually measured by a series of evaluation indicators, such as accuracy, precision, recall, etc. The preset model performance standard refers to a series of criteria or indicators used to measure and compare the performance of machine learning models. The global analysis model refers to a unified model constructed by aggregating the parameters and knowledge of multiple local models or node models in the federated learning framework.
[0142] Optionally, the loss function of the local node analysis model can be calculated by a distributed optimization algorithm, such as asynchronous stochastic gradient descent, distributed SGD, etc.
[0143] In the embodiment of the present invention, the item type of the carried item can be identified by using a preset classification and recognition algorithm based on the item features. By extracting rich item features (such as appearance, material, function, behavior features), the classification and recognition algorithm can more accurately identify the item type. Among them, the preset classification and recognition algorithm refers to the machine learning or deep learning algorithm pre-selected and configured for classifying and recognizing item types in the item detection and recognition task. The item type refers to the category or type of items recognized by the algorithm or system during the classification and recognition process, such as knives, lighters, explosives, etc.
[0144] In an embodiment of the present invention, by analyzing the risk coefficient of the carried item according to the item type, the dependence on manual inspection can be reduced, the inspection efficiency and accuracy can be improved, and human errors can be reduced. Among them, the risk coefficient is a quantitative index used to evaluate the degree of danger that an item may pose to human health, the environment, or property.
[0145] As an embodiment of the present invention, the analyzing the risk coefficient of the carried item according to the item type includes:
[0146] Based on the item type, determining the potential dangerous characteristics of the carried item and calculating the exposure frequency of the carried item;
[0147] According to the potential dangerous characteristics, analyzing the probability of danger occurrence, the consequences of danger, and the effectiveness of control measures of the carried item;
[0148] Determining the risk consequence weight of the risk consequence;
[0149] According to the exposure frequency, the probability of danger occurrence, the consequences of danger, the risk consequence weight, and the effectiveness of control measures, using the following formula to calculate the risk coefficient of the carried item:
[0150]
[0151] Among them, θ represents the risk coefficient, F represents the probability of danger occurrence, m represents the number of consequences of the risk consequence, K a represents the a-th risk consequence, V a represents the risk consequence weight corresponding to the a-th risk consequence, B represents the exposure frequency, and P represents the effectiveness of control measures.
[0152] Among them, the potential dangerous characteristics refer to the characteristics of the item itself that cause injury, health problems, environmental damage, or property loss, such as flammability, toxicity, explosiveness, etc. The exposure frequency refers to the number of times that people, property, or the environment are exposed to a specific danger or risk source within a specific time period. For example, if 10 events of controlled knives are detected within a month, the exposure frequency is 10. The probability of danger occurrence refers to the possibility of a dangerous event (such as an accident, explosion, poisoning, etc.) occurring under specific conditions. The consequences of danger refer to the negative results or impacts caused when a dangerous event occurs. The effectiveness of control measures refers to the efficacy of the control measures implemented to reduce or eliminate a specific risk in achieving its predetermined goal, and in this application, it can be set to 0.8. The risk consequence weight is an index used to quantify the severity of the consequences caused by potential dangerous events.
[0153] Exemplarily, in the present application, the probability of the occurrence of the hazard is 0.01. After the occurrence of the hazardous event, there are 3 types of hazardous consequences, namely personal injury, property damage, and social order disorder. According to the severity of the hazardous consequences, the hazardous consequences are quantified as 8 for personal injury, 5 for property damage, and 3 for social order disorder. The corresponding weights of the 3 types of hazardous consequences are determined to be 0.5, 0.4, and 0.1 respectively. The exposure rate is 10, and the effectiveness of the control measure is 0.8. Then the hazard coefficient of the carried item is 0.01*(8*0.5 + 5*0.4 + 3*0.1)*10 / 0.8 = 0.9.
[0154] In an embodiment of the present invention, by determining the prohibited level of the carried item based on the hazard coefficient, the hazard of the item can be quantitatively evaluated, and the safety risk can be managed more effectively to ensure the safety of personnel and facilities. Among them, the prohibited level refers to a classification system that classifies carried items into different levels according to the hazard coefficient of the item or other safety evaluation criteria. The conventional items refer to those items that are commonly used in daily life, are not restricted by law, and can be freely carried. These items usually do not pose a threat to public safety, such as personal electronic products, clothing, books, food, etc. The mildly prohibited items refer to those items that may be restricted in a specific environment but usually do not cause serious harm. These items may require special packaging or quantity restrictions, or can only be carried under specific conditions. For example, certain liquids, gels, small tools, etc. The moderately prohibited items refer to those items that may pose a certain threat to public safety or may be dangerous in specific situations. Carrying these items may be subject to stricter restrictions or require special permits. For example, certain types of knives, ammunition, matches, lighters, etc. The severely prohibited items refer to those items that are extremely likely to cause serious injury, damage, or safety risks. These items are usually completely prohibited from being carried, and violating the regulations may result in serious legal consequences. For example, explosives, guns, drugs, weapons of mass destruction, etc.
[0155] Optionally, determining the prohibited level of the carried item based on the hazard coefficient can be determined by threshold technology. For example, the hazard coefficient threshold of conventional items can be set to 0, the hazard coefficient threshold of mildly prohibited items can be set to 0 to 30, the hazard coefficient threshold of moderately prohibited items can be set to 30 to 70, and the hazard coefficient threshold of severely prohibited items can be set to above 70.
[0156] The prohibited item positioning module 104 is configured to mark the carried item according to the prohibited level to obtain a prohibited marked item, construct a human detection model of the classification detector, and map the prohibited marked item into the human detection model to obtain the position of the prohibited item.
[0157] In the embodiments of the present invention, by marking the carried items according to the prohibited level, the prohibited marked items can be obtained, which can quickly identify potential dangerous items, ensure that dangerous items are not misjudged or omitted, reduce the inspection time, and improve the efficiency of the overall security inspection process. Among them, the prohibited marked items refer to prohibited items that are specially marked for easy identification, classification, tracking or management.
[0158] Optionally, the marking of the carried items according to the prohibited level to obtain the prohibited marked items can be performed by a machine learning model. For example, a machine learning model is trained for items with different prohibited levels, and the machine learning model marks items with different prohibited levels with different colors.
[0159] In the embodiments of the present invention, by constructing a human body detection model of the classification detector, prohibited items can be accurately identified and located, achieving high-precision identification of prohibited items and reducing false alarms and missed reports. Among them, the human body detection model refers to a digital representation that simulates the three-dimensional shape and structure of a real human body.
[0160] Optionally, the construction of the human body detection model of the classification detector can be performed by three-dimensional scanning technology.
[0161] In the embodiments of the present invention, by mapping the prohibited marked items into the human body detection model, the prohibited item position can be obtained, and the specific position of the prohibited item on the human body can be accurately determined, accelerating the security inspection process and reducing the time and labor required for manual inspection. Among them, the prohibited item position refers to the specific position of the target prohibited item on the body of the person being inspected.
[0162] As an embodiment of the present invention, the mapping of the prohibited marked items into the human body detection model to obtain the prohibited item position includes:
[0163] Construct a three-dimensional coordinate system of the human body detection model;
[0164] Determine the two-dimensional position of the prohibited item of the prohibited marked item;
[0165] Obtain the imaging geometric parameters of the imaging device corresponding to the prohibited marked item;
[0166] According to the imaging geometric parameters, convert the two-dimensional position of the prohibited item into three-dimensional space coordinates of the prohibited item;
[0167] According to the three-dimensional coordinate system, map the three-dimensional space coordinates into the human body detection model to obtain the prohibited item position.
[0168] Among them, the three-dimensional coordinate system refers to a mathematical system used to describe and locate points in three-dimensional space, such as the Cartesian coordinate system. The two-dimensional position of the prohibited item refers to the planar coordinate position corresponding to the prohibited item in the image. The imaging geometric parameters refer to the parameters used to describe the geometric relationship between the imaging device (such as X-ray scanner, millimeter-wave scanner, CT scanner, etc.) and the object to be scanned. The three-dimensional space coordinates refer to a set of numerical values used to determine the position of an object in three-dimensional space.
[0169] Optionally, the conversion of the two-dimensional position of the prohibited item into the three-dimensional space coordinates of the prohibited item according to the imaging geometric parameters can be performed through multi-view imaging technology. For example, multi-view imaging technology is used to capture image data of the prohibited item from different angles, and according to the image data, the depth information of the prohibited item is determined through parallax calculation, thereby determining the three-dimensional space coordinates of the prohibited item.
[0170] The detection report generation module 105 is configured to integrate the multi-modal detection system of the classification detector and generate a detection report of the multi-modal detection system based on the prohibited item position, the prohibited level, the item type, and the person information.
[0171] In the embodiment of the present invention, by integrating the multi-modal detection system of the classification detector, multiple detection means (such as X-ray, metal detection, explosive detection, chemical detection, etc.) can be combined, and the system can more accurately identify and classify different prohibited items, reducing false alarms and missed detections. Among them, the multi-modal detection system refers to a comprehensive system integrating multiple sensors and detection technologies, which can simultaneously or sequentially use different detection methods to detect and analyze the target object.
[0172] Optionally, the integration of the multi-modal detection system of the classification detector includes:
[0173] Construct an acoustic-optic alarm unit of the classification detector, and integrate an initial detection system of the classification detector according to the detection unit corresponding to the classification detector and the acoustic-optic alarm unit;
[0174] Simulate the actual test scenario of the initial detection system, and based on the actual test scenario, use a preset metal test sample to test the initial detection system to obtain detection test data;
[0175] According to the detection test data, analyze the false alarm rate and anti-interference performance of the initial detection system;
[0176] According to the false alarm and the anti-interference performance, analyze the system performance of the initial detection system;
[0177] When the system performance meets the preset system performance standard, use the initial detection system as the multi-modal detection system of the classification detector.
[0178] Among them, the acoustic-optic alarm unit refers to a component in a security detection system used to issue visual and auditory alarms. When contraband or other items that require special attention are detected, this unit will be activated. The initial detection system refers to the system to be tested by initially integrating a detection unit and an acoustic-optic alarm unit. The actual test scenario refers to the specific environmental conditions and application scenarios simulated or set up to evaluate and verify the performance of the initial detection system, such as high-traffic detection environments, electromagnetic interference environments, etc. The preset metal test samples refer to a collection of metal items specially selected or made to test and evaluate the performance of classification detectors, such as coins, metal ornaments, mock metal weapons, etc. The detection test data refers to a series of performance indicators and result information collected during the test of the initial detection system, such as alarm records, test sample recognition rates, test sample positioning, etc. The false alarm rate refers to the probability that the initial detection system incorrectly identifies non-contraband items as contraband during the detection process. The anti-interference performance refers to the ability of the initial detection system to maintain stable operation and accurately detect target objects in the presence of various interference factors. The system performance refers to the performance of the initial detection system during the system test process, such as accuracy, detection efficiency, stability, etc. The preset system performance standards refer to a series of performance indicators and thresholds set in advance to ensure that the system meets specific operation requirements and user expectations.
[0179] In an embodiment of the present invention, by generating a detection report of the multimodal detection system based on the contraband position, the contraband level, the item type, and the person information, detailed contraband position information can be provided, enabling security inspection personnel to quickly and accurately locate and confirm contraband items. Among them, the detection report refers to a document that details and summarizes the information found during the security inspection process.
[0180] As an embodiment of the present invention, generating the detection report of the multimodal detection system based on the contraband position, the contraband level, the item type, and the person information includes:
[0181] Determine the alarm level of the multimodal detection system according to the contraband level and the item type;
[0182] Determine the acoustic-optic alarm mode of the multimodal detection system according to the alarm level;
[0183] Determine the target person corresponding to the detection object of the multimodal detection system according to the person information;
[0184] Based on the contraband position, determine the specific hiding position of the contraband corresponding to the target person;
[0185] Determine the detailed description and image evidence of the contraband;
[0186] Based on the acoustic and optical alarm mode, the target person, the specific hiding location, the detailed description of the contraband, and the image evidence, construct a detection report for the classification detector.
[0187] Among them, the alarm level refers to a series of levels preset according to factors such as the type, danger level, and legal regulations of the detected contraband. The acoustic and optical alarm mode refers to a security alarm system that combines sound and visual signals and is used to issue an alarm when a specific threat or abnormal situation is detected. For example, different alarm levels may use different tones or volumes to distinguish the severity of the threat, and red, yellow, or green lights are used, with different colors representing different alarm levels. The target person refers to an individual who is specifically concerned or tracked in activities such as security, monitoring, and investigation. The specific hiding location refers to the exact hiding location of the contraband on the target person or in the items they carry. The detailed description of the contraband refers to a detailed description of the item prohibited from being carried or transported, such as name, geometric features, item characteristics, etc. The image evidence refers to the visual record used to prove the existence and characteristics of the contraband.
[0188] In the embodiments of the present invention, by constructing the sensor network of the multi-sensors, a wider area can be covered, realizing the monitoring of a large-scale environment, and at the same time realizing fast data acquisition and processing, improving the response speed of the system; optionally, in the embodiments of the present invention, by constructing the sensor network of the multi-sensors, a wider area can be covered, realizing the monitoring of a large-scale environment, and at the same time realizing fast data acquisition and processing, improving the response speed of the system; in the embodiments of the present invention, by using the preset classification and recognition algorithm to identify the type of the carried item based on the item characteristics, rich item characteristics (such as appearance, material, function, behavior characteristics) can be extracted, and the classification and recognition algorithm can more accurately identify the item type; in the embodiments of the present invention, by using the preset classification and recognition algorithm to identify the type of the carried item based on the item characteristics, rich item characteristics (such as appearance, material, function, behavior characteristics) can be extracted, and the classification and recognition algorithm can more accurately identify the item type; in the embodiments of the present invention, by mapping the contraband marked item to the human body detection model, the position of the contraband can be accurately determined, accelerating the security inspection process, reducing the time and labor required for manual inspection. Finally, in the embodiments of the present invention, by generating a detection report for the multi-modal detection system based on the position of the contraband, the contraband level, the item type, and the person information, detailed contraband position information can be provided, enabling security inspection personnel to quickly and accurately locate and confirm the contraband. Therefore, the present invention can improve the adaptability and accuracy of the classification detection system.
[0189] As shown Figure 2 in the figure, it is a schematic flowchart of an intelligent multi-modal classification detection method provided by an embodiment of the present invention. In this embodiment, the intelligent multi-modal classification detection method includes:
[0190] Clarify the detection requirements of the classification detector. According to the detection requirements, configure the multi-sensors of the classification detector, construct a sensor network of the multi-sensors, and based on the sensors and the sensor network, construct a detection unit of the classification detector;
[0191] According to the detection unit, collect the item multi-modal data and face images of the detection object corresponding to the classification detector, extract the multi-modal data features of the item multi-modal data, based on the multi-modal data features, perform data fusion on the item multi-modal data to obtain item fusion data, extract the image features of the face image, and based on the image features, analyze the personal information of the detection object;
[0192] According to the item fusion data, analyze the item features of the items carried by the detection object. Based on the item features, use a preset classification and recognition algorithm to identify the item types of the carried items. According to the item types, analyze the risk coefficients of the carried items. Based on the risk coefficients, determine the prohibited levels of the carried items, where the prohibited levels include: regular items, mildly prohibited items, moderately prohibited items, and severely prohibited items;
[0193] According to the prohibited levels, mark the carried items to obtain prohibited marked items, construct a human body detection model of the classification detector, and map the prohibited marked items into the human body detection model to obtain the prohibited item positions;
[0194] Integrate the multi-modal detection system of the classification detector, and based on the prohibited item positions, the prohibited levels, the item types, and the personal information, generate a detection report of the multi-modal detection system.
[0195] In several embodiments provided by the present invention, it should be understood that the provided systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0196] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An intelligent multi-modal classification and detection system, characterized in that The intelligent multi-modal classification detection of the system includes: The material determination module is used to clarify the detection requirements of the classification detector, configure the multi-sensors of the classification detector according to the detection requirements, construct the sensor network of the multi-sensors, and construct the detection unit of the classification detector based on the sensors and the sensor network; The data fusion module is used to collect the multi-modal data of the items and the face images of the detection objects corresponding to the classification detector according to the detection unit, extract the multi-modal data features of the multi-modal data of the items, fuse the multi-modal data of the items based on the multi-modal data features to obtain the item fusion data, extract the image features of the face images, and analyze the personal information of the detection objects based on the image features; The item classification module is used to analyze the item features of the items carried by the detection objects according to the item fusion data, identify the item types of the carried items by using a preset classification and recognition algorithm based on the item features, analyze the risk coefficients of the carried items according to the item types, and determine the prohibited levels of the carried items based on the risk coefficients, where the prohibited levels include: conventional items, mildly prohibited items, moderately prohibited items, and severely prohibited items; The prohibited item positioning module is used to mark the carried items according to the prohibited levels to obtain the prohibited marked items, construct the human body detection model of the classification detector, and map the prohibited marked items into the human body detection model to obtain the positions of the prohibited items; The detection report generation module is used to integrate the multi-modal detection system of the classification detector and generate the detection report of the multi-modal detection system based on the positions of the prohibited items, the prohibited levels, the item types, and the personal information; 2. The intelligent multi-modal classification detection system according to claim 1, wherein The construction of the sensor network of the multi-sensors includes: Analyze the node density, network distribution range, and transmission rate requirements of the multi-sensors; Determine the structural layout of the multi-sensors according to the node density and the network distribution range; Determine the network communication protocol and network topology of the multi-sensors based on the transmission rate requirements; Configure the network parameters of the multi-sensors, where the network parameters include: node ID, communication frequency, and data transmission rate; Integrate the sensor network of the multi-sensors according to the structural layout, the network communication protocol, the network topology, and the network parameters; 3. The intelligent multimodal classification detection system according to claim 1, wherein, The construction of the detection unit of the classification detector based on the sensors and the sensor network includes: Construct the detection unit framework of the classification detector, and determine the sensor interface of the detection unit framework based on the sensors; Construct the communication interface of the detection unit framework according to the sensor network; Plan the signal processing circuit of the communication interface and configure the microcontroller of the detection unit framework; Integrate the microcontroller, the signal processing circuit, the sensor interface, and the communication interface into the detection unit framework to obtain the initial detection unit; Test the detection performance of the initial detection unit. When the detection performance meets the preset detection performance standard, use the initial detection unit as the detection unit of the classification detector.
4. The intelligent multi-modal classification and detection system according to claim 1, characterized in that, Based on the multi-modal data features, perform data fusion on the item multi-modal data to obtain item fusion data, including: Based on the multi-modal data features, determine the characteristic probability distribution of the item multi-modal data; Analyze the modal data weights of the item multi-modal data; According to the characteristic probability distribution and the modal data weights, calculate the fusion probability distribution of the item multi-modal data; According to the fusion probability distribution, perform data fusion on the item multi-modal data to obtain item fusion data.
5. The intelligent multi-modal classification and detection system according to claim 1, wherein The extraction of the image features of the face image includes: Locate the face area of the face image and detect the landmark points in the face area; Based on the landmark points, align the face image with the preset standard face position to obtain the aligned image; Use the preset Laplacian algorithm to detect the face edge features of the aligned image; Calculate the LBP value of the corresponding pixels of the aligned image, and based on the LBP value, extract the texture features of the aligned image; According to the texture features and the face edge features, determine the image features of the aligned image.
6. The intelligent multimodal classification detection system according to claim 1, characterized in that, According to the item fusion data, analyze the item features of the items carried by the detection object, including: Construct an initialization global analysis model for the classification detector corresponding to the detection object, and extract the initialization model parameters of the initialization global analysis model; According to the initialization model parameters, use the preset local training set to train the local node analysis model of the classification detector; Calculate the loss function of the local node analysis model, and based on the loss function, determine the local model parameters of the local node analysis model; According to the local model parameters, use the preset aggregation algorithm to calculate the global model parameters of the initialization global analysis model; Based on the global model parameters, update the initialization global analysis model to obtain an updated global model; Analyze the model performance of the updated global model. When the model character meets the preset model performance standard, use the updated global model as the global analysis model of the classification detector; Based on the item fusion data, use the global analysis model to analyze the item features of the carried items.
7. The intelligent multimodal classification detection system according to claim 1, characterized in that According to the item type, analyze the risk coefficient of the carried item, including: Based on the item type, determine the potential dangerous characteristics of the carried item and calculate the exposure frequency of the carried item; According to the potential dangerous characteristics, analyze the risk occurrence probability, risk consequences and effectiveness of control measures of the carried item; Determine the risk consequence weight of the risk consequence; According to the exposure frequency, the risk occurrence probability, the risk consequences, the risk consequence weight and the effectiveness of control measures, calculate the risk coefficient of the carried item.
8. The intelligent multimodal classification detection system according to claim 7, wherein Map the prohibited marked item to the human body detection model to obtain the prohibited item position, including: Construct a three-dimensional coordinate system for the human body detection model; Determine the two-dimensional position of the prohibited marked item; Obtain the imaging geometric parameters of the imaging device corresponding to the prohibited marked item; According to the imaging geometric parameters, convert the two-dimensional position of the prohibited item into the three-dimensional space coordinates of the prohibited item; According to the three-dimensional coordinate system, map the three-dimensional space coordinates into the human body detection model to obtain the position of the prohibited item.
9. The intelligent multimodal classification and detection system according to claim 1, wherein Based on the position of the prohibited item, the prohibited level, the item type, and the person information, generate a detection report for the multi-modal detection system, including: Determine the alarm level of the multi-modal detection system according to the prohibited level and the item type; Determine the audible and visual alarm mode of the multi-modal detection system according to the alarm level; Determine the target person of the multi-modal detection system corresponding to the detection object according to the person information; Based on the position of the prohibited item, determine the specific hiding position of the prohibited item corresponding to the target person; Determine the detailed description and image evidence of the prohibited item; Based on the audible and visual alarm mode, the target person, the specific hiding position, the detailed description of the prohibited item, and the image evidence, construct a detection report for the classification detector.
10. An intelligent multi-modal classification detection method, characterized in that, The method includes: Clarify the detection requirements of the classification detector, configure the multi-sensors of the classification detector according to the detection requirements, construct a sensor network of the multi-sensors, and construct a detection unit of the classification detector based on the sensors and the sensor network; According to the detection unit, collect the multi-modal data of the items and the face images of the detection object corresponding to the classification detector, extract the multi-modal data features of the multi-modal data of the items, fuse the multi-modal data of the items based on the multi-modal data features to obtain item fusion data, extract the image features of the face images, and analyze the person information of the detection object based on the image features; According to the item fusion data, analyze the item features of the items carried by the detection object, identify the item type of the carried items using a preset classification and recognition algorithm based on the item features, analyze the risk coefficient of the carried items according to the item type, and determine the prohibited level of the carried items based on the risk coefficient, where the prohibited level includes: conventional items, mildly prohibited items, moderately prohibited items, and severely prohibited items; Mark the carried items according to the prohibited level to obtain prohibited marked items, construct a human body detection model of the classification detector, and map the prohibited marked items into the human body detection model to obtain the position of the prohibited item; Integrate the multi-modal detection system of the classification detector, and generate a detection report for the multi-modal detection system based on the position of the prohibited item, the prohibited level, the item type, and the person information.
Citation Information
Patent Citations
Multilayer multidimensional three-dimensional smart security check system
CN108345050A
Security check image recognition method and device, electronic equipment and storage medium
CN112712093A
Rapid human body security check method based on vortex electromagnetic wave modal scanning
CN113567975A
Face attribute recognition method, device and equipment and readable storage medium
CN118379776A
Human body security check method and system based on terahertz imaging
CN118430014A