Fast Detection Method for Multiple Target Liquid Types Based on Millimeter-Wave and Vision Fusion

Through the multimodal detection method of the integration of millimeter wave radar and camera, the accuracy and adaptability of liquid recognition under complex container conditions are solved, and efficient and accurate liquid recognition under a variety of container materials and structures are achieved.

CN120198789BActive Publication Date: 2025-07-25NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510686776.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-25
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing liquid detection technology is difficult to achieve accurate identification under complex container conditions, and is disturbed by factors such as container material, shape and color, and a single modal detection solution cannot meet the accuracy, adaptability and reliability requirements in actual applications.

Method used

A multimodal detection method that integrates millimeter wave radar and camera is used to obtain millimeter wave reflected signals and image data, target detection and feature extraction are performed, feature decoupling is used with a dual-branch encoder, and liquid classification is performed through the Transformer module to achieve rapid identification of liquid types.

Benefits of technology

It improves the accuracy and robustness of liquid recognition, is highly adaptable, and can accurately identify liquids under complex backgrounds and diverse container conditions. It is suitable for a variety of container materials and structures, achieving contactless rapid detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198789B_ABST
    Figure CN120198789B_ABST
Patent Text Reader

Abstract

This application belongs to the field of liquid type perception and multi-modal sensing technology, and discloses a multi-target liquid type rapid detection method based on millimeter wave-vision fusion, including: first, obtaining the millimeter wave reflection signals and image data of multi-targets, performing target detection and single-target feature extraction on the image data, and separating the single-target millimeter wave original signals at the same time. Then, the original signals are subjected to segmented preprocessing and fast Fourier transform and then input into two encoders to extract the features of the liquid and the container respectively. During training, the container features extracted from the image are used as auxiliary labels, and a decoupling constraint term is introduced between the two millimeter wave features extracted to achieve decoupling. Finally, the single-target classification results are output according to the liquid features extracted from the millimeter wave signals, and all single-target classification results are uniformly integrated and output in a visual form. This application can identify the liquid types in multiple containers at the same time, improving the recognition accuracy, robustness and detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of liquid type sensing and multimodal sensing, and specifically relates to a multi-target liquid type rapid detection method based on millimeter-wave-vision fusion. Background Technique

[0002] Liquid detection technology has wide application value and important significance in the fields of food safety, public safety, medical diagnosis, etc. In the field of food safety, the accurate identification of liquid components is directly related to food quality monitoring and packaging anti-counterfeiting; in the field of public safety, the rapid identification of dangerous liquids can effectively improve the security inspection efficiency and prevention and control capabilities; in the medical field, body fluid component analysis provides an important basis for disease diagnosis. With the increasing demand for non-contact detection, the development of efficient and accurate liquid identification technology has become the focus of industry attention.

[0003] Existing liquid detection methods are mainly divided into three categories: special instruments, vision detection, and wireless sensing technology. Although special instruments such as spectral analyzers have high detection accuracy, they have limitations such as expensive equipment, the need for professional operation, and possible sample destruction; vision detection methods based on cameras can obtain appearance features such as the color and transparency of liquids, but are limited by the opacity of containers and the similarity of liquid appearances such as water and alcohol, making it difficult to achieve accurate identification; wireless sensing technology uses wireless signals with the advantage of being able to penetrate containers, but the collected signals are easily interfered by factors such as container shape and material, resulting in a low signal-to-noise ratio and difficult feature extraction. These methods each have technical bottlenecks that are difficult to overcome. Generally speaking, single-modal detection schemes often cannot meet the comprehensive requirements of accuracy, adaptability, and reliability in practical applications. Summary of the Invention

[0004] To solve the above technical problems, this application provides a multi-target liquid type rapid detection method based on millimeter-wave-vision fusion. This method makes full use of the multimodal information of millimeter-wave radar and cameras, overcomes the interference of physical characteristics such as container material, color, and shape in traditional detection technologies, is applicable to various materials such as glass and plastic and containers with complex structures, effectively solves the interference problem brought by container packaging to liquid identification, improves the accuracy and robustness of liquid identification, and has wide adaptability and application prospects.

[0005] To achieve the above object, this application is implemented through the following technical solutions:

[0006] This application is a multi-target liquid type rapid detection method based on millimeter-wave-vision fusion. The multi-target liquid type rapid detection method includes the following steps:

[0007] Step 1, Obtain multimodal data: In the detection scenario, use a millimeter-wave radar to obtain millimeter-wave reflection signals of multiple targets, i.e., millimeter-wave radar data, and synchronously use a camera to obtain image data of multiple target images;

[0008] Step 2, Image target detection and generation of auxiliary labels: Perform target detection on the image data of the multiple target images obtained in Step 1 to generate the center coordinates of the single-target position bounding box, and simultaneously extract the container features of each target image , and the container features are the auxiliary labels for the weakly supervised training in Step 5;

[0009] Step 3, Separation of multiple targets from millimeter-wave radar signals: Use the center coordinates of the single-target position bounding box obtained in Step 2 to assist in millimeter-wave radar signal beamforming, achieve the separation of multiple targets from millimeter-wave radar signals, and obtain the separated single-target millimeter-wave raw signals;

[0010] Step 4, Processing of single-target millimeter-wave radar data: Perform segmentation processing on each chirp of the separated single-target millimeter-wave raw signals. For each sub-chirp signal, use the fast Fourier transform to obtain the spectral characteristics of the single target in the distance dimension, thereby perform grid processing on the detected distance range of the radar, extract the amplitude spectra and phase spectra of a total of n grids before and after the grid corresponding to the position where the target is located, n is an integer from 5 to 10, and construct a 2×5×4×20 tensor for a single target, where 2 refers to the two dimensions of amplitude and phase, and both of these two dimensions are three-dimensional tensors of 5×4×20, to obtain the processed single-target millimeter-wave radar data;

[0011] Step 5, Extraction of double-branch features of single-target millimeter-wave radar: Parallelly input the processed single target, i.e., the millimeter-wave radar data of the th target, which is a 2×5×4×20 tensor, into two encoders with the same structure but independent parameters, i.e., the first encoder and the second encoder, perform feature extraction and then output liquid classification features and container features . When training the second encoder, the container features of each target image generated in Step 2 will be used as auxiliary labels to assist in weakly supervised training, so that the features can contain more packaging-related information;

[0012] Step 6, Feature decoupling based on mutual information: Introduce a decoupling constraint term between the liquid classification features and the container features to achieve the decoupling of the liquid classification features and the container features ;

[0013] Step 7, Based on the liquid classification features Output single-object classification result: Based on the liquid classification features output in step 5 As the input, it is sent into a lightweight Transformer module with a multi-head self-attention mechanism. After being encoded by the Transformer module, the output features are non-linearly mapped through a fully connected classifier to generate prediction scores for the corresponding liquid types, and finally the single-object classification result is output in text form;

[0014] Step 8, Integration and output of multiple-object classification: Integrate all single-object classification results and output them in a visual form.

[0015] A further improvement of this application is that: Step 1 specifically includes the following steps:

[0016] Step 1.1, Obtain the millimeter-wave reflection signals of multiple objects, that is, millimeter-wave radar data, through a millimeter-wave radar. The reflection signals of the multiple objects include the amplitude and phase information of the millimeter-wave radar data. The millimeter-wave radar is a frequency-modulated continuous-wave radar with a transmission frequency range of 77 GHz - 81 GHz;

[0017] Step 1.2, At the same time, the camera obtains the image data of multiple-object images synchronized with the millimeter-wave radar data at a frame rate of 30 fps. Among them, the camera and the millimeter-wave radar need to cover the same target area in the same spatial scene.

[0018] A further improvement of this application is that: Step 2 specifically includes the following steps:

[0019] Step 2.1, Adjust the size of the images of multiple-object images; obtain images with a size specification of 640×640;

[0020] Step 2.2, Use the YOLOv11 model to perform object detection on the size-adjusted multiple-object images and output the center coordinates of the bounding box centers of all single objects;

[0021] Step 2.3, Use the feature extractor of the YOLOv11 model to extract the container features of each object in the multiple-object images to obtain the container features of the th object image

[0022] A further improvement of this application is that: Step 3 specifically includes the following steps:

[0023] Step 3.1, Spatial coordinate alignment: Convert the center coordinates of the single-object position bounding box obtained in step 2 to the radar polar coordinate system through a binocular vision-radar joint calibration matrix, establish the mapping relationship between the center coordinates of the single-object position bounding box and the radar beam direction, and obtain the single-object azimuth angle information;

[0024] Step 3.2, Adaptive beamforming: Based on the single-target azimuth information obtained in Step 3.1, digital beamforming technology is used to generate a spatially selective receiving beam, with the main lobe of each receiving beam aligned with a certain target liquid in the environment, while suppressing the interference signals from objects other than the target liquid in the environment;

[0025] Step 3.3, Multi-channel signal separation: Realize multi-beam parallel processing on the FPGA hardware platform, perform spatial domain filtering on the millimeter-wave reflection signals of multiple targets obtained by the millimeter-wave radar in Step 1, generate independent signal channels for each single target, and obtain the separated single-target millimeter-wave original signals.

[0026] A further improvement of this application is that in Step 4, each chirp of the separated single-target millimeter-wave original signal is segmented, and fast Fourier transform is performed on each sub-chirp, specifically including the following steps:

[0027] Step 4.1, The millimeter-wave radar uses frequency-modulated continuous-wave technology to sense multiple surrounding targets, continuously emits millimeter-wave radar signals within the fast time period of a chirp, and the frequency of the emitted signal is:

[0028]

[0029] where, is the starting frequency, is the frequency modulation slope, is the time;

[0030] The frequency of the multi-target reflection signal is:

[0031]

[0032] where, is the distance from the radar to the target, is the speed of light;

[0033] Step 4.2, After the multi-target reflection signal is received by the millimeter-wave radar, the transmitted signal and the received signal, i.e., the reflection signal, are mixed to obtain an intermediate-frequency signal, and the frequency of the intermediate-frequency signal is:

[0034]

[0035] where, is the frequency modulation slope;

[0036] Step 4.3, When the frequency modulation slope is fixed, the frequency of the intermediate-frequency signal only depends on the distance from the radar to the target, and by the frequency of the intermediate-frequency signal Perform a fast Fourier transform to obtain the spectral characteristics of a single target in the range dimension, thereby performing grid division on the range detected by the radar, and extracting the amplitude spectra and phase spectra of a total of n grids before and after the grid corresponding to the position of the target. n is an integer from 5 to 10, and a 2×5×4×20 tensor is constructed for a single target, where 2 refers to the two dimensions of amplitude and phase, and both of these two dimensions are three-dimensional tensors of 5×4×20, obtaining the processed single-target millimeter-wave radar data.

[0037] When the target is a liquid, the received signal strength (RSS) of the liquid is:

[0038]

[0039] where, is the transmitted signal strength, is the transmitting antenna gain, is the receiving antenna gain, is the millimeter-wave wavelength, is the reflection coefficient of the liquid;

[0040] Step 4.4, According to the Fresnel reflection formula, the relationship between the reflection coefficient of the liquid, the refractive index of the liquid, and the refractive index of the container material is expressed as:

[0041]

[0042] The relationship between the refractive index of the container material and the dielectric constant of the liquid is expressed as:

[0043]

[0044] where, is the real part of the dielectric constant, is the imaginary part of the dielectric constant, under high-frequency conditions;

[0045] The dielectric constant of the liquid and the electromagnetic frequency are related by the double Debye model and expressed as:

[0046]

[0047] where, is the dielectric constant at radio high frequencies, is the dielectric constant at static i.e., zero frequency, is the dielectric constant at the intermediate frequency, and are the two relaxation time constants respectively.

[0048] A further improvement of this application lies in: the dual-branch feature extraction in step 5 is specifically as follows: the first encoder extracts features from the internal liquid of the target, and performs supervised training with the liquid type as the label, and outputs the liquid classification features of the th target , the second encoder extracts features from the multi-target image of the target, and performs weakly supervised training under the supervision of the auxiliary label provided in step 2, and outputs the container features of the th target .

[0049] A further improvement of this application lies in: step 6 specifically includes the following steps:

[0050] Step 6.1, the output th liquid classification feature and the container feature are respectively and , where is the number of samples, is the feature dimension, then the orthogonal loss is expressed as :

[0051]

[0052] Among them, is the Frobenius norm, which measures the overall non-orthogonal degree between the liquid classification feature and the container feature ;

[0053] Step 6.2, minimize the orthogonal loss , so that the liquid classification feature and the container feature are kept as orthogonal as possible in the feature space, that is, independent of each other, and the decoupling of the liquid classification feature and the container feature is realized. From the perspective of information theory, this orthogonality constraint can also be regarded as an indirect constraint on mutual information:

[0054]

[0055] Among them, represents the mutual information, which is used to measure the dependence or correlation between the liquid classification feature and the container feature .

[0056] A further improvement of this application lies in: the Transformer module in step 7 includes a single-layer encoder, a multi-head attention sub-layer and a feed-forward neural network sub-layer, where the multi-head sub-attention sub-layer:

[0057]

[0058] in, is each attention head:

[0059]

[0060] in, is the query matrix, which means we need to find similarities from other features. is the key matrix, representing the importance of the feature. is a value matrix, representing the actual content of the feature, It is a linear transformation matrix used to reduce the dimension or linearly combine the concatenated features;

[0061] Feedforward Neural Network Sublayers:

[0062]

[0063] in, is the input feature, i.e. the liquid classification feature vector, is the first layer weight matrix, is the second layer weight matrix, is the first layer bias vector, is the second layer bias vector;

[0064] Introducing residual connections and layer normalization mechanisms between the multi-head attention sublayer and the feedforward neural network sublayer to improve the ability to express non-local features and obtain feature representations to be fed into the classifier , expressed as:

[0065]

[0066]

[0067] in, is the feature after multi-head self-attention (MHSA) and residual connection, It is the feature after the feedforward neural network (FFN) and residual connection, that is, the final feature representation.

[0068] A further improvement of the present application is that step 8 specifically includes the following steps:

[0069] The single target classification results of step 8.1 and step 7 are aligned and integrated with the individual position bounding boxes of multiple targets in step 2 to form target recognition records containing spatial position information and classification attributes. Each record is structured as follows:

[0070]

[0071] Among them, is the center coordinate of the single-object position bounding box of the target in the image, is the class confidence of the single-object bounding box selection, is the single-object liquid classification result described in step 7, is the confidence of the single-object liquid classification result; For the confidence of the single-object liquid classification result;

[0072] Step 8.2, Result Visualization: Generate a visualization interface of all single-object classification results on the terminal device, and use the structured representation in step 8.1 as the label of the visualization output to ensure the correspondence of the output content, and intuitively display the images, liquid classifications, and recognition confidence information of each target.

[0073] The beneficial effects of this application are:

[0074] This application innovatively proposes a fusion detection scheme for millimeter-wave radar and vision collaborative perception. By deeply integrating the penetration detection ability of millimeter waves and the advantage of visual appearance feature extraction, a multi-modal information complementary mechanism is constructed: visual data captures the features of the liquid container and is used to assist in decoupling the millimeter-wave signal features, decouple and separate the two parts of the features related to the liquid and the container in the millimeter-wave signal, and then focus on the liquid-related features and input them into the deep learning model for inference, thereby realizing the classification of liquids.

[0075] The entire model architecture with feature decoupling as the core proposed in this application gives full play to the complementary advantages of millimeter waves and vision in information perception, not only enhancing the robustness of the system under complex backgrounds and diverse container conditions, but also improving the perception ability of the essential features of the internal liquid. Through information collaborative learning between different modalities, the overall recognition accuracy and system stability are effectively improved, and it is applicable to dynamic and diverse actual scenarios.

[0076] This application has strong adaptability, is not affected by factors such as container material, shape, and color, and is applicable to liquid detection of various complex-structured containers.

[0077] The non-contact detection of this application avoids contaminating or interfering with the liquid, ensuring the safe detection of volatile, toxic, and corrosive liquids.

[0078] This application can detect multiple targets simultaneously, and the detection process does not require sample processing. It can obtain the liquid type information of multiple targets in real time, with a fast detection speed and high efficiency.

[0079] The overall detection frame rate of this application can reach 15 frames per second, and the processing delay is less than 200 milliseconds. It can achieve efficient and safe identification of liquids without damaging the samples, and is widely applicable to fields such as security inspection, industrial quality inspection, medical monitoring, smart home, and vehicle-mounted detection, with significant social benefits and application prospects.

[0080] This application can be widely applied to scenarios such as airport and station security inspections in the field of public safety, automated quality inspection production lines in industrial production, infusion monitoring and sample analysis in the medical and health industry, liquid recognition household appliances in smart homes, and vehicle-mounted dangerous goods detection in the field of intelligent transportation. Its highly compatible system architecture is convenient for integration into various intelligent terminals, which can not only promote the technological upgrading of related industries, but also ensure public safety by improving the detection efficiency of dangerous liquids. At the same time, it has significant environmental protection value in the classification and treatment of waste liquids and resource recovery, showing broad social and economic benefits and sustainable development potential. Brief Description of the Drawings

[0081] Figure 1 is the flowchart of this application.

[0082] Figure 2 is the flowchart of this application.

[0083] Figure 3 is the application scenario diagram of this application.

[0084] Figure 4 is the schematic diagram of the chirp sub - segment interception of the millimeter - wave radar signal of this application.

[0085] Figure 5 is the architecture diagram of the Transformer module of this application.

[0086] Figure 6 is the sub - chirp amplitude diagram of the signal at the millimeter - wave radar receiving antenna 1 of this application.

[0087] Figure 7 is the sub - chirp phase diagram of the signal at the millimeter - wave radar receiving antenna 1 of this application. Detailed Embodiments

[0088] The following will disclose the embodiments of the present invention with diagrams. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.

[0089] As Figures 1-3 shown, this application is a method for rapid detection of multiple - target liquid types based on millimeter - wave - vision fusion. The rapid detection method includes the following steps:

[0090] Step 1. Obtain multimodal data: In the detection scenario, use a millimeter-wave radar to obtain the millimeter-wave reflection signals of multiple targets, i.e., millimeter-wave radar data, and simultaneously use a camera to obtain the image data of multiple target images. The specific steps are as follows:

[0091] Step 1.1. Use the IWR1642 millimeter-wave radar from Texas Instruments to obtain the millimeter-wave reflection signals of multiple targets, i.e., millimeter-wave radar data. The reflection signals of the multiple targets include the amplitude and phase information of the millimeter-wave radar data. The millimeter-wave radar is a frequency-modulated continuous-wave radar with a transmit frequency range of 77 GHz - 81 GHz;

[0092] Step 1.2. Meanwhile, the camera obtains the image data of multiple target images synchronized with the millimeter-wave radar data at a frame rate of 30 fps, where the camera and the millimeter-wave radar need to cover the same target area in the same spatial scene.

[0093] Step 2. Image target detection and generation of auxiliary labels: Perform target detection on the image data of the multiple target images obtained in Step 1 to generate the center coordinates of the single-target position bounding box, and simultaneously extract the container features of each target image , and the container features are the auxiliary labels for the weakly supervised training in Step 5. The specific steps are as follows:

[0094] Step 2.1. Resize the image of the multiple target images to obtain an image with a size of 640×640;

[0095] Step 2.2. Use the YOLOv11 model to perform target detection on the resized multiple target images and output the center coordinates of all single-target position bounding boxes;

[0096] Step 2.3. Use the feature extractor of the YOLOv11 model to extract the container features of each target in the multiple target images to obtain the th multiple target container feature .

[0097] Step 3. Separation of multiple targets from millimeter-wave radar signals: Use the center coordinates of the single-target position bounding boxes obtained in Step 2 to assist in millimeter-wave radar signal beamforming, achieve the separation of multiple targets from millimeter-wave radar signals, and obtain the separated single-target millimeter-wave raw signals. The specific steps are as follows:

[0098] Step 3.1. Spatial coordinate alignment: Convert the center coordinates of the single-target position bounding boxes obtained in Step 2 to the radar polar coordinate system through a binocular vision-radar joint calibration matrix, establish the mapping relationship between the center coordinates of the single-target position bounding boxes and the radar beam pointing, and obtain the single-target azimuth angle information;

[0099] Step 3.2, Adaptive beamforming: Based on the single-target azimuth information obtained in Step 3.1, digital beamforming technology is used to generate a spatially selective receiving beam. The main lobe of each receiving beam is aligned with a certain target liquid in the environment, while suppressing the interference signals from objects other than the target liquid in the environment.

[0100] Step 3.3, Multi-channel signal separation: Implement multi-beam parallel processing on the FPGA hardware platform, perform spatial domain filtering on the millimeter-wave reflection signals of multiple targets obtained by the millimeter-wave radar in Step 1, generate independent signal channels for each single target, and obtain the separated single-target millimeter-wave raw signals.

[0101] Step 4, Millimeter-wave radar data processing of single target: Each chirp of the separated single-target millimeter-wave raw signal is segmented, as Figure 4 shown. For each sub-chirp signal, fast Fourier transform is used to obtain the spectral characteristics of the single target in the distance dimension, so as to perform grid processing on the detected distance range of the radar, extract the amplitude spectra and phase spectra of a total of n grids before and after the grid corresponding to the position of the target, where n is an integer from 5 to 10. In this embodiment, n = 5 is taken, and a 2×5×4×20 tensor is constructed for a single target, where 2 refers to the two dimensions of amplitude and phase, and both of these two dimensions are three-dimensional tensors of 5×4×20, and the processed single-target millimeter-wave radar data is obtained.

[0102] In this step, each chirp of the separated single-target millimeter-wave raw signal is segmented, and fast Fourier transform is performed on each sub-chirp, which specifically includes the following steps:

[0103] Step 4.1, The millimeter-wave radar uses frequency-modulated continuous-wave technology to sense multiple surrounding targets, continuously emits millimeter-wave radar signals within the fast time period of a chirp, and the frequency of the emitted signal is:

[0104]

[0105] where, is the starting frequency, is the frequency modulation slope, is the time;

[0106] The frequency of the multi-target reflection signal is:

[0107]

[0108] where, is the distance from the radar to the target, is the speed of light;

[0109] Step 4.2: After the multi-target reflected signal is received by the millimeter-wave radar, the transmitted signal and the received signal, i.e., the reflected signal, are mixed to obtain an intermediate-frequency signal, and the frequency of the intermediate-frequency signal is:

[0110]

[0111] wherein, is the frequency modulation slope;

[0112] Step 4.3: When the frequency modulation slope is fixed, the frequency of the intermediate-frequency signal is only related to the distance from the radar to the target. By performing a fast Fourier transform on the frequency of the intermediate-frequency signal, the processed single-target millimeter-wave radar data is obtained.

[0113] When the target is a liquid, the reflection signal strength (RSS) of the liquid is:

[0114]

[0115] wherein, is the transmitted signal strength, is the transmitting antenna gain, is the receiving antenna gain, is the millimeter-wave wavelength, is the reflection coefficient of the liquid;

[0116] Step 4.4: According to the Fresnel reflection formula, the relationship between the reflection coefficient of the liquid, the refractive index of the liquid, and the refractive index of the container material is expressed as:

[0117]

[0118] The relationship between the refractive index of the container material and the dielectric constant of the liquid is expressed as:

[0119]

[0120] wherein, is the real part of the dielectric constant, is the imaginary part of the dielectric constant under high-frequency conditions;

[0121] The relationship between the dielectric constant of the liquid and the electromagnetic frequency is represented by the double Debye model as:

[0122]

[0123] Among them, is the dielectric constant at radio frequency, is the dielectric constant at static, i.e., zero frequency, is the dielectric constant at intermediate frequency, and are two relaxation time constants respectively.

[0124] The double Debye model is used to describe the frequency response behavior of polarized materials under the action of electromagnetic waves. The present invention is used to model the frequency response of liquids (high-frequency dielectrics) in the millimeter-wave band.

[0125] As Figure 1 shown, the signal frequencies of different chirp sub-fragments are different, and the peak signal intensities after performing fast Fourier transform on the corresponding sub-fragments are also different. Steps 4.3 - 4.4 also reflect that different liquids have different frequency responses to millimeter-wave signals, which is the fundamental theory for distinguishing different liquids in this application.

[0126] In the embodiment, each chirp of the separated single-target millimeter-wave raw signal is segmented, and fast Fourier transform is performed on each sub-chirp to extract the amplitude spectrum and phase spectrum of a total of five range boxes before and after the target box. The data dimension is 10×4×20, where 10 represents 2 data modalities, i.e., amplitude and phase × 5 range boxes, 4 represents the number of channels, and 20 represents the number of sub-chirps intercepted in a single chirp.

[0127] Step 5: Dual-branch feature extraction of a single-target millimeter-wave radar: The processed single-target, i.e., the th target millimeter-wave radar data, i.e., a 2×5×4×20 tensor, is parallelly input into two encoders with the same structure but independent parameters, i.e., the first encoder and the second encoder. The dual-branch feature extraction is specifically as follows: The first encoder extracts features of the internal liquid of the target and performs supervised training with the liquid type as the label, and outputs the th target's liquid classification features , and the second encoder extracts features of the multi-target image of the target and performs weakly supervised training under the supervision of the auxiliary label provided in step 2, and outputs the th target's container features . When training the second encoder, each target image container feature generated in step 2 will be used as an auxiliary label to assist weakly supervised training, so that the feature can contain more packaging-related information.

[0128] Step 6: Feature decoupling based on mutual information: In the liquid classification features and the container features A decoupling constraint term is introduced between them. This constraint term is represented in the form of an orthogonal loss function, which is used to minimize the projection coincidence degree of the two feature representations, suppress their overlapping distribution in the feature space, and then enhance the feature diversity and complementarity, indirectly achieving the approximate minimization goal of mutual information. Thus, the model is encouraged to learn independent features with complementarity to achieve the decoupling of liquid classification features and container features , which specifically includes the following steps:

[0129] Step 6.1, the th liquid classification feature and container feature are respectively and , where is the number of samples, is the feature dimension, and the orthogonal loss is expressed as :

[0130]

[0131] where is the Frobenius norm, which measures the overall non-orthogonal degree between the liquid classification feature and the container feature ;

[0132] Step 6.2, minimize the orthogonal loss , so that the liquid classification feature and the container feature are kept as orthogonal as possible (i.e., independent) in the feature space, realizing the decoupling of the liquid classification feature and the container feature . From the perspective of information theory, this orthogonality constraint can also be regarded as an indirect constraint on mutual information:

[0133]

[0134] where represents the mutual information, which is used to measure the dependence or correlation between the liquid classification feature and the container feature .

[0135] Step 7, output a single-object classification result based on the liquid classification feature : Based on the liquid classification feature output in step 5 As input, it is fed into a lightweight Transformer module with a multi-head self-attention mechanism. After being encoded by the Transformer module, the output features are non-linearly mapped through a fully-connected classifier to generate prediction scores corresponding to the types of liquids. Finally, the single-object classification results are output in text form. This Transformer module has the characteristics of a small number of parameters and high computational efficiency, adapts to the deployment requirements of edge computing devices, and can achieve fast, efficient, and non-contact identification of liquid types while ensuring the recognition accuracy.

[0136] As Figure 5 shown, the Transformer module includes a single-layer encoder, a multi-head attention sub-layer, and a feed-forward neural network sub-layer. Among them, the multi-head sub-attention sub-layer:

[0137]

[0138] Among them, is each attention head:

[0139]

[0140] Among them, is the query matrix, representing the need to find similarities from other features, is the key matrix, representing the importance identification of features, is the value matrix, representing the actual content of features, is a linear transformation matrix used to reduce the dimension or linearly combine the concatenated features;

[0141] The feed-forward neural network sub-layer:

[0142]

[0143] Among them, is the input feature, i.e., the liquid classification feature vector, is the first-layer weight matrix, is the second-layer weight matrix, is the first-layer bias vector, is the second-layer bias vector;

[0144] A residual connection and a layer normalization mechanism are introduced between the multi-head attention sub-layer and the feed-forward neural network sub-layer to enhance the non-local feature expression ability and obtain the feature representation to be fed into the classifier , expressed as:

[0145]

[0146]

[0147] in, is the feature after multi-head self-attention (MHSA) and residual connection, It is the feature after the feedforward neural network (FFN) and residual connection, that is, the final feature representation.

[0148] The final feature representation It has stronger non-local modeling and context capture capabilities. After Transformer encoding, the output features are nonlinearly mapped through a fully connected classifier to generate a prediction score for the corresponding liquid type as the basis for determining the target category. This classification module has the characteristics of small parameters and high computational efficiency. It is suitable for the deployment requirements of edge computing devices and can achieve fast, efficient, and contactless recognition of liquid types while ensuring recognition accuracy.

[0149] Step 8, integrated output of multiple target classifications: All single target classification results are unified and integrated and output in a visual form. Specifically, the liquid type prediction results obtained by the Transformer and the fully connected classifier for each target in step seven are summarized into a target category list according to the target number. Combined with the spatial position information and confidence index provided in the image detection stage, each target is assigned a "position-type-confidence" triple label to achieve accurate recognition and expression of multiple targets in the scene. This application packages these triplet data into a structured output, which can be used for subsequent upper-level application calls, such as liquid hazardous goods identification, intelligent sorting or warehouse management tasks. In addition, the output results can also be visualized and superimposed with the original image to assist users in intuitively understanding the recognition results and improve the interpretability and practicality of the system. This step not only completes the classification integration of multiple targets, but also marks the closure of the complete processing flow from raw data acquisition to the output of structured recognition results.

[0150] Step 8 specifically includes the following steps:

[0151] The single target classification results of step 8.1 and step 7 are aligned and integrated with the individual position bounding boxes of multiple targets in step 2 to form target recognition records containing spatial position information and classification attributes. Each record is structured as follows:

[0152]

[0153] in, is the target in the image The center coordinates of the single target position bounding box, is the category confidence of the single target box selection, is the single target liquid classification result described in step 7, It is the confidence of single target liquid classification result;

[0154] Step 8.2, Result Visualization: Generate a visualization interface for all single-object classification results on the terminal device, and use the structured representation in Step 8.1 as the label for the visualization output to ensure the correspondence of the output content, and intuitively display the information of the images, liquid classification, and recognition confidence of each object.

[0155] Figure 6 It is the sub-chirp amplitude diagram of the signal at the millimeter-wave radar receiving antenna 1 of the present application. According to Figure 6 It can be seen that under the condition of the same container and the same position, the amplitudes of the sub-chirps extracted from the millimeter-wave reflection signals of different liquids have different frequency responses to the signals, specifically manifested as different trends and directions of the amplitudes with the change of the signal frequency, indicating that the sub-chirp amplitude information of the millimeter-wave signal can reflect the unique characteristics of the liquid.

[0156] Figure 7 It is the sub-chirp phase diagram of the signal at the millimeter-wave radar receiving antenna 1 of the present application. According to Figure 6 It can be seen that under the condition of the same container and the same position, the phases of the sub-chirps extracted from the millimeter-wave reflection signals of different liquids are different, indicating that the sub-chirp phase information of the millimeter-wave signal can also reflect the characteristic differences between liquids.

[0157] The present application decouples the reflection signal of the millimeter-wave radar into internal liquid characteristics and external packaging characteristics, uses the visual data provided by the camera as the auxiliary supervision for the extraction of external packaging characteristics, and then ensures the decoupling effect of the internal liquid information and the external packaging information through the orthogonal loss function, and then focuses on the internal liquid characteristics of the container extracted by the first encoder, and realizes the recognition and classification of the liquid based on this. The present application can realize the synchronous recognition, accurate classification, and efficient output of multiple liquid targets in complex scenarios, has good practicability and system scalability, and is particularly suitable for multi-object recognition scenarios such as logistics, security inspection, and industrial quality inspection.

[0158] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A rapid multi-target liquid type detection method based on millimeter-wave and vision fusion, characterized in that: The multi-target liquid type rapid detection method includes the following steps: Step 1, Obtain multi-modal data: In the detection scenario, use a millimeter-wave radar to obtain the millimeter-wave reflection signals of multi-targets, and synchronously obtain the image data of the multi-target images; Step 2, Image Target Detection and Generation of Auxiliary Labels: Perform target detection on the image data of the multi-target image obtained in Step 1 to generate the center coordinates of the single-target position bounding box, and at the same time extract the feature of each target image container ; Step 3, Separation of multi-targets in millimeter-wave radar signals: Use the center coordinates of the single-target position bounding box obtained in Step 2 to assist in millimeter-wave radar signal beamforming, realize the separation of multi-targets in millimeter-wave radar signals, and obtain the separated single-target millimeter-wave original signals; Step 4, Processing of millimeter-wave radar data of single-targets: Perform segmentation processing on each chirp of the separated single-target millimeter-wave original signals. For each sub-chirp signal, use the fast Fourier transform to obtain the spectral characteristics of the single-target in the range dimension. Perform grid division on the range range detected by the millimeter-wave radar, extract the amplitude spectra and phase spectra of a total of n grids before and after the grid corresponding to the position of the single-target, and construct a tensor for the single-target to obtain the processed single-target millimeter-wave radar data; Step 5: Dual-branch feature extraction of single-target millimeter-wave radar: The th processed single-target millimeter-wave radar data, i.e., tensors, are parallelly input into two encoders with the same structure but independent parameters, namely the first encoder and the second encoder, and after feature extraction, liquid classification features and container features are output. When training the second encoder, each target image container feature generated in Step 2 is used as an auxiliary label for weakly supervised training; Step 6: Feature decoupling based on mutual information: Introduce a decoupling constraint term between the liquid classification features and the container features to achieve the decoupling of the liquid classification features and the container features ; Step 7: Based on the liquid classification features Output the single-target liquid classification result: Based on the liquid classification features output in Step 5 As the input, it is sent into the Transformer module. After being encoded by the Transformer module, the output features are non-linearly mapped through a fully connected classifier to generate the prediction scores for the corresponding liquid types, and finally the single-target liquid classification result is output in text form; Step 8, Integrated output of multiple target classifications: Uniformly integrate all single-target liquid classification results and output them in a visual form.

2. The multi-object liquid type rapid detection method based on millimeter wave-vision fusion according to claim 1, characterized in that: The specific steps of Step 1 include the following steps: Step 1.1, Obtain the millimeter-wave reflection signals of multi-targets through a millimeter-wave radar. The reflection signals of the multi-targets include the amplitude and phase information of the millimeter-wave radar data. The millimeter-wave radar is a frequency-modulated continuous-wave radar with a transmission frequency range of 77 GHz - 81 GHz; Step 1.2, At the same time, the camera obtains the image data of the multi-target images synchronized with the millimeter-wave radar data at a frame rate of 30 fps, where the camera and the millimeter-wave radar need to cover the same target area in the same spatial scene.

3. The multi-target liquid type rapid detection method based on millimeter-wave-vision fusion according to claim 1, characterized in that: The specific steps of Step 2 include the following steps: Step 2.1, Adjust the size of the image data of the multi-target images; Step 2.2, Use the YOLOv11 model to perform target detection on the resized multi-target images and output the center coordinates of all single-target position bounding boxes; Step 2.3: Use the feature extractor of the YOLOv11 model to extract the container features of each target in the multi-target image, and obtain the container feature of the th target image 4. The multi-object liquid type rapid detection method based on millimeter wave-vision fusion according to claim 1, characterized in that: The specific steps of Step 3 include the following steps: Step 3.1, Spatial coordinate alignment: Convert the center coordinates of the single-target position bounding box obtained in Step 2 to the radar polar coordinate system through a binocular vision-radar joint calibration matrix, establish the mapping relationship between the center coordinates of the single-target position bounding box and the radar beam pointing, and obtain the single-target azimuth information; Step 3.2, Adaptive beamforming: Based on the single-target azimuth information obtained in Step 3.1, generate a receiving beam with spatial selectivity. Each main lobe of the receiving beam is aligned with a target liquid, and at the same time, suppress the interference signals from objects other than the target liquid in the environment; Step 3.3, Multi-channel signal separation: Perform spatial domain filtering on the millimeter-wave reflection signals of multi-targets obtained by the millimeter-wave radar in Step 1, generate independent signal channels for each single-target, and obtain the separated single-target millimeter-wave original signals.

5. The multi-target liquid type rapid detection method based on millimeter wave-vision fusion according to claim 1, characterized in that: In Step 4, perform segmentation processing on each chirp of the separated single-target millimeter-wave original signals and perform fast Fourier transform on each sub-chirp, which specifically includes the following steps: Step 4.1: The millimeter-wave radar uses frequency-modulated continuous-wave technology to sense multiple surrounding targets, continuously emits millimeter-wave radar signals within a fast time period of a chirp, and the frequency of the emitted signal is as follows: ; Among them, is the starting frequency, is the frequency modulation slope, is the time; The frequency of the multi-target reflection signal is: ; Among them, is the distance from the radar to the target, is the speed of light; Step 4.

2. After the multi-target reflection signal is received by the millimeter-wave radar, the transmitted signal is mixed with the received signal, i.e., the reflection signal, to obtain an intermediate-frequency signal, and the frequency of the intermediate-frequency signal is: ; Among them, is the frequency modulation slope; Step 4.

3. When the frequency modulation slope is fixed, the frequency of the intermediate frequency signal is only related to the distance from the millimeter wave radar to the target . By performing a fast Fourier transform on the frequency of the intermediate frequency signal , the processed single-target millimeter wave radar data is obtained. When the target is a liquid, the reflection signal intensity of the liquid is: ; Among them, is the transmission signal strength, is the transmission antenna gain, is the receiving antenna gain, is the millimeter wave wavelength, is the reflection coefficient of the liquid; Step 4.4, Reflectivity of the liquid , Refractive index of the liquid and refractive index of the container material are related as follows: ; Refractive Index of Container Material and Dielectric Constant of Liquid It is expressed as: ; wherein, is the real part of the dielectric constant, is the imaginary part of the dielectric constant, under high-frequency conditions; Dielectric constant of a liquid and electromagnetic frequency is expressed as: ; Among them, is the dielectric constant at high frequencies, is the dielectric constant at static, i.e., zero frequency, is the dielectric constant at intermediate frequencies, and are two relaxation time constants respectively.

6. The multi-target liquid type rapid detection method based on millimeter wave-vision fusion according to claim 1, wherein: The double-branch feature extraction in step 5 is specifically as follows: The first encoder extracts features from the internal liquid of the target, performs supervised training with the liquid type as the label, and outputs the th liquid classification feature . The second encoder extracts features from the multi-target image of the target, performs weakly supervised training under the supervision of the auxiliary label provided in step 2, and outputs the th container feature .

7. The multi-target liquid type rapid detection method based on millimeter wave-vision fusion according to claim 6, characterized in that: Step 6 specifically includes the following steps: Step 6.1, the th liquid classification feature and the container feature are respectively and , then the orthogonal loss is expressed as : ; Among them, is the number of samples, is the feature dimension, is the Frobenius norm, which is used to measure the liquid classification features and the container features represents the overall non-orthogonal degree between them; Step 6.2, Minimize the orthogonal loss , so that the liquid classification features and the container features remain orthogonal (i.e., independent of each other) in the feature space, achieving the decoupling of the liquid classification features and the container features : ; Among them, represents mutual information, which is used to measure the characteristics of liquid classification and the characteristics of the container to measure the dependence or correlation between them.

8. The multi-object liquid type rapid detection method based on millimeter wave-vision fusion according to claim 1, wherein: The Transformer module in Step 7 includes a single-layer encoder, a multi-head attention sub-layer, and a feed-forward neural network sub-layer, where the multi-head attention sub-layer: ; Among them, is each attention head: ; Among them, is the query matrix, is the key matrix, is the value matrix, is a linear transformation matrix; The feed-forward neural network sub-layer: ; Among them, is the input feature, i.e., the liquid classification feature vector, is the first-layer weight matrix, is the second-layer weight matrix, is the first-layer bias vector, is the second-layer bias vector; Introduce a residual connection and a layer normalization mechanism between the multi-head attention sublayer and the feed-forward neural network sublayer to obtain the feature representation to be fed into the classifier : ; ; Among them, is the feature that has passed through multi-head self-attention and residual connection.

9. The multi-target liquid type rapid detection method based on millimeter wave-vision fusion according to claim 1, characterized in that: Step 8 specifically includes the following steps: Step 8.1: Align and integrate the single-object classification result of Step 7 with the individual position bounding boxes of multiple objects in Step 2 to form an object recognition record containing spatial position information and classification attributes. Each record is structured as follows: ; Among them, is the center coordinate of the single-object position bounding box of the target in the image , is the confidence of the single-object bounding box selection is the single-object liquid classification result described in step 7 is the confidence of the single-object liquid classification result; Step 8.2: Result visualization: Generate a visualization interface for all single-object liquid classification results on the terminal device, and use the structured representation in Step 8.1 as the label for the visualization output to ensure the correspondence of the output content and intuitively display the information of the image, liquid classification, and recognition confidence of each object.

Citation Information

Patent Citations

  • Probability target detection method based on image and millimeter wave radar fusion

    CN117746204A

  • Millimeter wave contraband detection method based on multi-head attention and BIFPN improved YOLO v7

    CN118261900A