A method, device and medium for fusion of camera and radar signals of hidden targets
By combining millimeter radar and depth cameras, using neural network models to extract and fusion the feature of cameras and radar signals, the problem that cameras cannot directly detect metals under the shadows in the prior art are solved, and accurate detection and positioning of metals are achieved, and the accuracy and speed of detection are improved.
Patent Information
- Application Number
- CN202510113519.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-01-24
AI Technical Summary
In the prior art, cameras cannot directly detect metal under the shield. Millimeter wave sensors can detect the existence of metal, but are relatively abstract, lack intuitive visual spatial information, and the data characteristics of cameras and radars are different, and there is no targeted algorithm for data fusion.
By combining millimeter radar and depth cameras, neural network models are used to extract and fusion the feature of camera and radar signals, a joint feature vector is generated, and a binary segmentation mask of feature representation is generated through a dual-channel convolution decoder to determine the position of the hidden target.
Accurate detection and positioning of metals under the shield is achieved, more intuitive visual spatial information is provided, the accuracy and speed of detection is improved, and the situation of misjudgment and misjudgment are reduced.
Smart Images

Figure CN119575367B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of sensor fusion technology, and in particular to a hidden target detection method, device and medium by fusing camera and radar signals. Background Art
[0002] Traditional metal detection technology often has limitations. For example, a single metal detector may have poor metal positioning accuracy, especially in complex crowded scenes or when there are multiple interference factors. It also requires close-up detection of the subjects one by one, which is slow. In addition, it is difficult to accurately determine the specific location of the metal and related spatial information in the case of close-up detection.
[0003] With the development of sensor technology, cameras can provide rich visual information. However, cameras cannot directly detect metal hidden under clothing or other coverings. Millimeter wave sensors can penetrate non-metallic materials to a certain extent and generate reflection signals to metal objects, thereby detecting the presence of metal. However, millimeter wave data is relatively abstract and lacks intuitive visual spatial background information, making it difficult to accurately determine the specific position of the metal in the overall scene and the object it belongs to. In addition, in the existing technology, the data features of cameras and radars are different. To extract effective features from these data and match them, complex algorithms need to be designed.
[0004] Through the above analysis, the problems and defects of the prior art are as follows:
[0005] Cameras in existing technologies cannot directly detect metal under obstructions. Millimeter wave sensors can detect the presence of metal, but this is relatively abstract and requires close inspection, lacking intuitive visual spatial information. In addition, the data characteristics of cameras and radars are different, and there is no targeted algorithm for data fusion. Summary of the invention
[0006] The embodiments of the present application provide a method, device and medium for fusing camera and radar signals of a hidden target, which can solve the problem in the prior art that cameras cannot directly detect metal under obstructions. Millimeter wave sensors can detect the presence of metal, but are relatively abstract and lack intuitive visual spatial information. In addition, the data features of cameras and radars are different, and there is no targeted algorithm for data fusion.
[0007] In a first aspect, an embodiment of the present application provides a method for fusing camera and radar signals of a hidden target, the method comprising: detecting a hidden target within a preset distance through a millimeter radar and a depth camera, and obtaining a first radar signal and a first TOF image, respectively; preprocessing the first radar signal and the first TOF image, and performing feature extraction and fusion through a neural network model to obtain a joint feature vector, wherein the neural network model includes a dual-channel convolutional encoder and a dual-channel convolutional decoder; performing deep feature amplification and deep feature extraction on the joint feature vector to obtain a feature representation; generating a binary segmentation mask of the feature representation through a dual-channel convolutional decoder to determine the position of the hidden target.
[0008] In one implementation of the present application, the first radar signal and the first TOF image are preprocessed, specifically including: determining the resolution of the first TOF image according to the input layer of the neural network model; traversing the depth values of the pixels of the first TOF image, and replacing the pixel values with a depth value of zero with the pixel values with the smallest non-zero depth value in the TOF image; mapping the depth value and the radar signal within a preset range to obtain a second radar signal and a second TOF image.
[0009] In one implementation of the present application, a dual-channel convolution encoder includes a radar signal encoder and a TOF image encoder; and feature extraction and fusion are performed through a neural network model, specifically including: based on the radar signal encoder including a convolutional long short-term memory layer and a radar convolutional layer, the time characteristics of the second radar signal are extracted through the convolutional long short-term memory layer and converted into a first feature map; the first feature map is convolved through the radar convolution layer using a first preset step size, and features are extracted and combined through the receptive field to obtain a second feature map.
[0010] In one implementation of the present application, feature extraction and fusion are performed through a neural network model, specifically including: based on the TOF encoder including an image convolution layer, convolution is performed through the image convolution layer using a second preset step size to obtain a TOF feature map; the second feature map is converted into the same form as the TOF feature map to obtain a third feature map; the TOF feature map and the third feature map are feature-fused along the depth axis to obtain a first joint feature vector.
[0011] In one implementation of the present application, deep feature amplification and feature extraction are performed on the joint feature vector to obtain a feature representation, specifically including: based on the receptive field including the first receptive field and the second receptive field, the local features of the first joint feature vector are extracted through a deep separable convolution layer and the first receptive field, and linearly combined in the depth direction to obtain a second joint feature vector; the second joint feature vector is incremented through the second receptive field to extract feature dependencies to obtain a third joint feature vector; the third joint feature vector is two-dimensionally upsampled, and the local features of the second joint feature vector are extracted through the first receptive field, and the features are aggregated by the second receptive field to obtain a feature representation.
[0012] In one implementation of the present application, a binary segmentation mask of the feature representation is generated through a dual-channel convolution decoder to determine the position of the hidden target, specifically including: converting the feature representation into a pixel probability map of the hidden target through a dual-channel convolution decoder and a sigmoid activation function; performing threshold calculation on a preset neighborhood of each pixel according to the pixel probability map to obtain a local threshold; converting the pixel probability map into a binary segmentation mask according to the local threshold to obtain the area where the metal object of the hidden target is located and the background area.
[0013] In one implementation of the present application, the method also includes: using a real data set marked with hidden target positions to jointly train a dual-channel convolutional encoder and a dual-channel convolutional decoder; using a loss function to calculate the difference between the binary segmentation mask and the true annotation, and updating the parameters through a back-propagation algorithm.
[0014] In one implementation of the present application, the method further includes: during the training process, presetting a discard rate for the neural network model to prevent overfitting; and using batch normalization and ReLU activation function for all convolutional layers.
[0015] In a second aspect, an embodiment of the present application also provides a device for fusing camera and radar signals of a hidden target, the device comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: detect hidden targets within a preset distance through a millimeter radar and a depth camera, and obtain a first radar signal and a first TOF image, respectively; preprocess the first radar signal and the first TOF image, and perform feature extraction and fusion through a neural network model to obtain a joint feature vector, wherein the neural network model comprises a dual-channel convolutional encoder and a dual-channel convolutional decoder; perform deep feature amplification and deep feature extraction on the joint feature vector to obtain a feature representation; generate a binary segmentation mask of the feature representation through the dual-channel convolutional decoder to determine the position of the hidden target.
[0016] In a third aspect, an embodiment of the present application also provides a non-volatile computer storage medium for the fusion of camera and radar signals of a hidden target, storing computer executable instructions, wherein the computer executable instructions are configured to: detect hidden targets within a preset distance through a millimeter radar and a depth camera, and obtain a first radar signal and a first TOF image, respectively; preprocess the first radar signal and the first TOF image, and perform feature extraction and fusion through a neural network model to obtain a joint feature vector, wherein the neural network model includes a dual-channel convolutional encoder and a dual-channel convolutional decoder; perform deep feature amplification and deep feature extraction on the joint feature vector to obtain a feature representation; generate a binary segmentation mask of the feature representation through a dual-channel convolutional decoder to determine the position of the hidden target.
[0017] The embodiments of the present application provide a method, device and medium for the fusion of camera and radar signals for hidden targets, which fuse the camera and millimeter wave sensor. The millimeter wave radar can use the radio frequency signal to detect the reflectivity of the object, and the depth camera can provide scene structure information, which effectively utilizes the advantages of both. At the same time, dual sensor fusion technology and deep learning neural network are used to achieve hidden object detection while protecting the privacy of users through the architecture hhCameraRadar. The binary segmentation mask can provide more detailed target boundary information, and the use of local thresholds helps to reduce misjudgments and missed judgments. It can detect multiple people at the same time to improve the detection speed and quality. Low-cost, portable sensor fusion technology provides a paradigm for hidden metal detection, which is of great significance to fields such as public safety monitoring. The research trend of multimodal data fusion has also promoted the development of this technology, aiming to meet the increasingly complex security detection and monitoring needs by integrating data from different types of sensors. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 A flow chart of a method for fusing camera and radar signals of a hidden target provided in an embodiment of the present application;
[0020] Figure 2 An overall logical architecture diagram of a method for fusing camera and radar signals of a hidden target provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of a binary segmentation mask for a method of fusing camera and radar signals of a hidden target provided in an embodiment of the present application;
[0022] Figure 4A schematic diagram of the internal structure of a device for fusing camera and radar signals for a hidden target provided in an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0024] The embodiments of the present application provide a method, device and medium for the fusion of camera and radar signals of a hidden target, which solves the problem in the prior art that cameras cannot directly detect metal under obstructions. Millimeter wave sensors can detect the presence of metal, but are relatively abstract and lack intuitive visual spatial information. In addition, the data features of cameras and radars are different, and there is no targeted algorithm for data fusion.
[0025] The technical solution proposed in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0026] Figure 1 A flow chart of a method for fusing camera and radar signals of a hidden target provided in an embodiment of the present application. Figure 1 As shown, a method for fusing camera and radar signals of a hidden target provided in an embodiment of the present application specifically includes the following steps:
[0027] Step 10: Detect hidden targets within a preset distance using the millimeter radar and the depth camera to obtain a first radar signal and a first TOF (Time of Flight) image, respectively.
[0028] In this step, for example, a 24GHz millimeter-wave radar and a depth camera, the millimeter-wave radar has two receiving antennas and one transmitting antenna, and the communication distance is about 15 meters. The depth camera uses a pair of ultra-wide sensors 50 mm apart to calculate the depth of the stereo image. When the target appears within the acquisition range of the depth camera and the high-density wave radar, a detection area of about 5×5 meters can be demarcated in the venue and marked with markers or fences. The subjects are required to stand in a row in this area and remove all metal objects on their bodies. Then the metal objects are hidden under the clothes at the chest position and a green mark is attached to the surface of the clothes. The mark is only used to assist positioning in the RGB image and is invisible in the depth image. The radar and depth camera are started to start collecting data at the same time. During the collection process, the subject moves randomly in the detection area at a normal walking speed, changes position and posture, and the collection system records the radar's I / Q signals, depth images and RGB images in real time. The sensors collect data and input them into their respective neural networks to start the data calculation and prediction process.
[0029] Step 20: Preprocess the first radar signal and the first TOF image, and perform feature extraction and fusion through a neural network model to obtain a joint feature vector, wherein the neural network model includes a dual-channel convolution encoder and a dual-channel convolution decoder.
[0030] In an embodiment of the present application, a neural network architecture for data processing is proposed, named hhCameraRadar architecture, which is a dual-channel convolutional encoder-decoder structure designed to learn the combined potential features focused on radar and TOF modalities to extract the position of the target object in the TOF image. The specific workflow is as follows.
[0031] As an optional embodiment, preprocessing the first radar signal and the first TOF image may specifically include: Step 201: determining the resolution of the first TOF image according to the input layer of the neural network model.
[0032] In this step, the resolution is adjusted so that the data can fit the input layer of the neural network, ensuring that the network can correctly extract features when processing images. For example, if the input layer of the hhCameraRadar architecture requires the resolution of the TOF image to be 256x256 pixels, and the resolution of the original TOF image is 512x512 pixels, the original TOF image needs to be downsampled and its resolution adjusted to 256x256 pixels to meet the requirements of the model input layer.
[0033] Step 202: traverse the depth values of the pixels of the first TOF image, and replace the pixel value with a depth value of zero with the pixel value with the minimum non-zero depth value in the TOF image.
[0034] In this step, zero-depth pixel processing is performed to replace the zero-depth pixels with the minimum pixel value in the scene to create a smooth image, avoiding the appearance of discontinuities, so that the neural network can learn the feature patterns in the image more stably during the training process, thereby improving the effectiveness and accuracy of the training.
[0035] Step 203: Map the depth value and the radar signal into a preset range to obtain a second radar signal and a second TOF image.
[0036] In this step, data standardization is performed. The radar data is collected in the original I / Q (I represents the in-phase component and Q represents the orthogonal component) form and normalized to the [0,1] interval. The purpose is to make the data between different features or different samples comparable and avoid the adverse effects of some eigenvalues being too large or too small on subsequent calculations. The TOF image data is converted into a form with specific statistical characteristics (within the [0,1] range) to make the data distribution between different images more consistent and speed up the training of the neural network.
[0037] As an optional embodiment, based on the dual-channel convolution encoder including a radar signal encoder and a TOF image encoder; and performing feature extraction and fusion through a neural network model, it can specifically include: step 204: based on the radar signal encoder including a convolutional long short-term memory layer and a radar convolutional layer, extracting the time feature of the second radar signal through the convolutional long short-term memory layer, and converting it into a first feature map;
[0038] In this step, the radar signal encoder consists of two 1D convolutional long short-term memory layers followed by three 1D convolutional (Conv1D) layers, as shown in Figure 2 As shown in the figure, data with a dimension of 1*256*4 is input into the first one-dimensional LSTM network layer, and a convolution operation is performed on the time dimension of the data. At the same time, the gating mechanism of LSTM (Long Short-Term Memory) is combined to extract the time-related features in the input data sequence. These time features can reflect the motion state of the target object.
[0039] Step 205: convolve the first feature map using a radar convolution layer using a first preset step size, and extract and combine features through a receptive field to obtain a second feature map.
[0040] In this step, the features output by the first LSTM network layer are passed to the second network layer, which further processes the data and converts them into convolution-compatible features, so that the data can better adapt to subsequent convolution layer operations, ensuring the smooth transmission and processing of data between different layers, while further integrating and optimizing time-related features, laying the foundation for the subsequent extraction of higher-level features.
[0041] The next two Conv1D (one-dimensional convolution) layers can convolve the data with a stride of 2. At the same time, by setting a large receptive field (1,7) of the convolution kernel, while reducing the size of the feature map, it is able to incorporate more time-related information and larger spatial context to calculate the features required for the next layer, extract more representative and abstract features from the radar data, reduce data redundancy, and highlight the key information related to the detection of hidden metal objects.
[0042] The third Conv1D layer learns the final embedding of the radar data, and its output filter mapping is 16→32→64→64→128. All layers use the ReLU activation function to introduce nonlinear factors to enhance the network's expressiveness and enable the network to learn complex relationships in the data.
[0043] As an optional embodiment, feature extraction and fusion are performed through a neural network model, which may specifically include: Step 206: Based on the TOF encoder including an image convolution layer, convolution is performed through the image convolution layer using a second preset step size to obtain a TOF feature map.
[0044] In this step, the TOF encoder consists of 4 Conv2D layers with zero padding and ReLU activation functions. Each convolution layer performs convolution operations on the image using different receptive field strides. By setting a large receptive field (7,7), the relationship between large objects in the camera field of view can be captured by taking advantage of the fact that depth images are smoother and have less intensity transitions than RGB images. The raw TOF image data is converted into latent features with specific feature representations. The final output is a feature map with a shape of 4×4×128, which can be fused with radar features and used together for subsequent hidden metal object detection tasks.
[0045] Step 207: Convert the second characteristic map into the same form as the TOF characteristic map to obtain a third characteristic map.
[0046] In this step, the output of the Conv1D layer is reshaped into a form similar to the shape of the TOF feature output so that it can be fused with the features of the TOF image in the subsequent steps.
[0047] Step 208: Perform feature fusion on the TOF feature map and the third feature map along the depth axis to obtain a first joint feature vector.
[0048] In this step, the features of the radar signal and the depth image signal are fused along their respective depth axes and combined into a new data object.
[0049] Step 30: Perform deep feature amplification and deep feature extraction on the joint feature vector to obtain feature representation.
[0050] As an optional embodiment, deep feature amplification and feature extraction are performed on the joint feature vector to obtain feature representation, which may specifically include: Step 301: based on the receptive field including the first receptive field and the second receptive field, local features of the first joint feature vector are extracted through a depth-separable convolution layer and the first receptive field, and linearly combined in the depth direction to obtain a second joint feature vector;
[0051] In the embodiment of the present application, multiple receptive fields are used. The first receptive field, i.e., a small receptive field (e.g., 1×1), is used to extract local features, and the second receptive field, i.e., a large receptive field (e.g., 4×4), is used to learn more contextual relationships. In actual operation, both can be determined according to actual conditions.
[0052] In this step, radar signals are better at detecting the presence of metal objects, while TOF images can provide spatial location and structural information of objects. The fused features can reflect both advantages, which helps to improve the accuracy of detection and positioning. Let the radar feature be expressed as (where d is the dimension of the radar feature), the TOF feature is expressed as , the joint embedding feature vector after feature fusion operation is:
[0053] Step 302: Increment the second joint feature vector through the second receptive field to extract feature dependencies and obtain a third joint feature vector;
[0054] In this step, the fused vector is sent to the Deep Feature Modulation (DFM) block for processing. The fused feature vector J passes through a deep separable convolution layer, in which individual features are learned for all 256 concatenated filter maps in the feature representation using a 3×3 spatial receptive field, followed by a 1×1 convolution in the feature depth direction. The purpose is to aggregate the information from the two modalities into a learned joint representation, thereby preliminarily integrating the data features of the two modalities. A 4×4 convolution with an expanded receptive field can be used to learn larger spatial relationships between data, connect dynamic structures in the feature space in a larger area, provide greater global context information for the features, and further expand the depth and breadth of mining feature relationships.
[0055] At the same time, a dropout layer with a dropout rate of 0.3 can be used in the entire DFM block to prevent the model from overfitting and ensure the generalization and stability of the model. In addition, all convolutional layers use batch normalization operations, followed by ReLU activation functions. Such standardization and activation operations are used to optimize the model training process and feature representation effect.
[0056] Step 303: perform two-dimensional upsampling on the third joint feature vector, extract local features of the second joint feature vector through the first receptive field, and aggregate the features through the second receptive field to obtain feature representation.
[0057] In this step, the output of the DFM block is firstly upsampled in two dimensions to make the feature map larger in the spatial dimension and restore a certain resolution. The upsampled features are then passed to the Feature Extraction and Embedding (FEE) The FEE block consists of a series of convolution operations, including 1×1 and 4×4 convolutions. Similar to the DFM block, a small 1×1 receptive field is used first to build local features on all filter maps, while keeping the model's parameter scale small, avoiding problems such as overfitting due to the model being too complex, and mining as many feature details as possible under limited parameters. Then, a large 4×4 receptive field size is used to aggregate features in a larger spatial context. Since the location of hidden objects is usually only a few pixels in size, this method of first extracting small receptive field features and then aggregating them in a larger spatial context can help locate hidden objects and improve positioning accuracy by observing features from multiple receptive field angles. Similarly, a dropout layer with a dropout rate of 0.3 is used throughout the FEE block to prevent overfitting, and all convolutional layers use a batch normalization operation followed by a ReLU activation function to ensure model training results and effective feature representation.
[0058] Step 40: Generate a binary segmentation mask of the feature representation through a dual-channel convolution decoder to determine the location of the hidden target.
[0059] As an optional embodiment, a binary segmentation mask of the feature representation is generated through a dual-channel convolution decoder to determine the position of the hidden target, which may specifically include: step 401: converting the feature representation into a pixel probability map of the hidden target through a dual-channel convolution decoder and a sigmoid activation function;
[0060] In this step, the output of the FEE block is fed into a convolutional layer, which uses a sigmoid activation function to generate a pixel-by-pixel probability map of the location of hidden metal objects in the TOF scene based on the fused feature information learned previously. , indicating the location The sigmoid function then converts the output of the convolution layer into a probability value for the existence of a hidden metal object at each pixel position, so that the system can express the prediction of the position of the hidden metal object in the form of probability, thereby more intuitively evaluating the possibility of each pixel belonging to a hidden metal object.
[0061] Step 402: According to the pixel probability map, a threshold value is calculated for a preset neighborhood of each pixel to obtain a local threshold value.
[0062] In this step, for each pixel in the image, the grayscale value distribution in its neighborhood, such as 3x3, 5x5, is calculated. According to the distance between the pixels in the neighborhood and the central pixel, the grayscale value is weighted averaged using a Gaussian function, and then a constant is subtracted as the threshold.
[0063] Step 403: According to the local threshold, the pixel probability map is converted into a binary segmentation mask to obtain the area where the metal object of the hidden target is located and the background area.
[0064] In this step, the probability map can be converted into a binary segmentation mask to clearly distinguish the area where the hidden metal object is located from the background area, thereby realizing the detection and positioning tasks of the hidden metal object. Figure 3 shown.
[0065] More specifically, the output result is: where T represents the set threshold, Indicates location Is the location judged as a hidden metal object? When When there is a target object, Time display position No target object is found.
[0066]
[0067] As an optional embodiment, the method may also include: using a real data set marked with hidden target positions to jointly train a dual-channel convolutional encoder and a dual-channel convolutional decoder; using a loss function to calculate the difference between the binary segmentation mask and the true annotation, and updating the parameters through a back-propagation algorithm.
[0068] As an optional embodiment, the method also includes: during the training process, presetting a discard rate for the neural network model to prevent overfitting; and using batch normalization and ReLU activation function for all convolutional layers.
[0069] In summary, the embodiment of the present application is implemented by combining millimeter-wave radar and depth camera technology with a neural network architecture. Through the neural network (hhCameraRadar architecture), the encoder in the model extracts features from the radar signal and the TOF image respectively. The neural network architecture uses a convolutional long short-term memory (LSTM) module to process the radar signal, and uses a convolution operation to process the depth image signal, embeds it into a high-dimensional latent space (feature fusion), and then processes it through a deep feature magnification (DFM) block and a feature extraction and embedding (FEE) block, and uses deep feature magnification technology to amplify the combined latent features. Finally, a convolutional decoder and a sigmoid activation function are used to generate a binary segmentation mask for the entire image, that is, each pixel position is predicted as a probability value for the existence of a hidden metal object, which is converted into a binary form to obtain the final binary segmentation mask, which can determine the position of the hidden object.
[0070] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a device for fusing camera and radar signals of a hidden target, and its structure is as follows: Figure 4 shown.
[0071] Figure 4 A schematic diagram of the internal structure of a device for fusing camera and radar signals for a hidden target provided in an embodiment of the present application. Figure 4 As shown, the device includes:
[0072] at least one processor 401;
[0073] and, a memory 402 communicatively coupled to the at least one processor;
[0074] Among them, the memory 402 stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor 401 so that the at least one processor 401 can: detect hidden targets within a preset distance through millimeter radar and depth camera, and obtain a first radar signal and a first TOF image respectively; preprocess the first radar signal and the first TOF image, and extract and fuse features through a neural network model to obtain a joint feature vector, wherein the neural network model includes a dual-channel convolutional encoder and a dual-channel convolutional decoder; perform deep feature amplification and deep feature extraction on the joint feature vector to obtain a feature representation; generate a binary segmentation mask of the feature representation through a dual-channel convolutional decoder to determine the position of the hidden target.
[0075] Some embodiments of the present application provide corresponding Figure 1 A non-volatile computer storage medium for fusing camera and radar signals of a hidden target, storing computer executable instructions, wherein the computer executable instructions are configured to: detect a hidden target within a preset distance by using a millimeter radar and a depth camera, and obtain a first radar signal and a first TOF image respectively; preprocess the first radar signal and the first TOF image, and extract and fuse features by using a neural network model to obtain a joint feature vector, wherein the neural network model includes a dual-channel convolution encoder and a dual-channel convolution decoder; perform deep feature amplification and deep feature extraction on the joint feature vector to obtain a feature representation; generate a binary segmentation mask of the feature representation by using the dual-channel convolution decoder to determine the position of the hidden target.
[0076] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the IoT device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0077] The system and medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the system and medium also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the system and medium will not be repeated here.
[0078] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0079] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0080] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0082] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0083] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0084] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0085] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0086] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for fusing camera and radar signals of a hidden target, characterized in that: The method comprises: Detect hidden targets in a preset area through a millimeter radar and a depth camera to obtain a first radar signal and a first TOF image respectively; Preprocessing the first radar signal and the first TOF image, and extracting and fusing features through a neural network model to obtain a joint feature vector and a second feature map, wherein the neural network model includes a dual-channel convolution encoder and a dual-channel convolution decoder; Feature extraction and fusion are performed through the neural network model, including: Based on the TOF image encoder including an image convolution layer, convolution is performed by the image convolution layer using a second preset step size to obtain a TOF feature map; Converting the second characteristic map into the same form as the TOF characteristic map to obtain a third characteristic map; Perform feature fusion on the TOF feature map and the third feature map along the depth axis to obtain a first joint feature vector; Performing deep feature amplification and deep feature extraction on the joint feature vector to obtain feature representation; The joint feature vector is subjected to deep feature amplification and feature extraction to obtain feature representation, specifically including: Based on the receptive field including the first receptive field and the second receptive field, extracting local features of the first joint feature vector through the depth-separable convolution layer and the first receptive field, and performing linear combination in the depth direction to obtain a second joint feature vector; Incrementing the second joint feature vector through the second receptive field to extract feature dependencies, thereby obtaining a third joint feature vector; Performing two-dimensional upsampling on the third joint feature vector, extracting local features of the second joint feature vector through the first receptive field, and aggregating the features through the second receptive field to obtain a feature representation; The binary segmentation mask of the feature representation is generated by the dual-channel convolution decoder to determine the position of the hidden target.
2. The method for fusing camera and radar signals of a hidden target according to claim 1, characterized in that: The preprocessing of the first radar signal and the first TOF image specifically includes: Determining the resolution of the first TOF image according to the input layer of the neural network model; Traversing the depth values of the pixels of the first TOF image, and replacing the pixel value with a depth value of zero with the pixel value with the minimum non-zero depth value in the TOF image; The depth value and the radar signal are mapped into a preset range to obtain a second radar signal and a second TOF image.
3. The method for fusing camera and radar signals of a hidden target according to claim 2, characterized in that: Based on the dual-channel convolution encoder including a radar signal encoder and the TOF image encoder; The feature extraction and fusion are performed through the neural network model, specifically including: Based on the radar signal encoder including a convolutional long short-term memory layer and a radar convolutional layer, extracting the time feature of the second radar signal through the convolutional long short-term memory layer and converting it into a first feature map; The first feature map is convolved by the radar convolution layer using a first preset step size, and features are extracted and combined through a receptive field to obtain a second feature map.
4. The method for fusing camera and radar signals of a hidden target according to claim 1, characterized in that: Generating a binary segmentation mask of the feature representation by the dual-channel convolution decoder to determine the position of the hidden target specifically includes: The feature representation is converted into a pixel probability map of the hidden target through the dual-channel convolution decoder and the sigmoid activation function; According to the pixel probability map, a threshold value is calculated for a preset neighborhood of each pixel to obtain a local threshold value; According to the local threshold, the pixel probability map is converted into a binary segmentation mask to obtain the area where the metal object of the hidden target is located and the background area.
5. The method for fusing camera and radar signals of a hidden target according to claim 4, characterized in that: The method further comprises: Using a real data set with hidden target locations annotated, jointly training the dual-channel convolutional encoder and the dual-channel convolutional decoder; A loss function is used to calculate the difference between the binary segmentation mask and the true annotation, and the parameters are updated through the back-propagation algorithm.
6. The method for fusing camera and radar signals of a hidden target according to claim 5, characterized in that: The method further comprises: During the training process, a discard rate is preset for the neural network model to prevent overfitting; Batch normalization and ReLU activation function are used for all convolutional layers.
7. A device for fusing camera and radar signals for a hidden target, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Execute the steps of a method for fusing camera and radar signals of a hidden target as described in any one of claims 1-6.
8. A non-volatile computer storage medium for fusion of camera and radar signals for a hidden target, storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Execute the steps of a method for fusing camera and radar signals of a hidden target as described in any one of claims 1-6.
Citation Information
Patent Citations
Camera and millimeter wave radar front fusion road surface target detection method
CN116259024A
Human body modeling system and method based on fusion of ToF camera and millimeter wave radar
CN117036600A