Multi-target detection method
Through the multi-object detection method, the image and radar data are calibrated and fuzzy inference are used to use fuzzy neural networks to solve the problems of high hardware costs, low detection accuracy and insufficient real-time performance, low-cost and efficient target detection are achieved, and cyclists' safety perception in complex traffic scenarios is improved.
Patent Information
- Application Number
- CN202510883941.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-30
AI Technical Summary
The existing target detection scheme that integrates radar and camera data has problems such as high hardware deployment cost, low detection accuracy and low efficiency, especially in complex environments, it is difficult to meet the real-time nature of riding scenarios and the specific target classification needs of non-motor vehicles.
The multi-objective detection method is adopted to obtain image data and radar data for data calibration, and the pre-trained fuzzy neural network model is used to perform fuzzy inference, output the matching degree of image data and each target in the radar data, and filter out the targets whose matching degree exceeds the preset threshold, including spatial alignment and time alignment processing, reducing hardware costs and improving detection efficiency.
While reducing hardware costs, it improves detection accuracy and efficiency. It is suitable for edge computing platforms with limited computing power. It has real-time and interpretability, solves the problems of low detection accuracy and insufficient real-time performance of traditional methods in complex environments, and improves the safety perception ability of cyclists in complex traffic scenarios.
Smart Images

Figure CN120387144A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and in particular to a multi-object detection method. Background Art
[0002] In recent years, non-motorized vehicle travel modes represented by bicycles and electric bicycles have developed rapidly globally, and the scale of the cycling population has continued to expand. However, the resulting traffic safety hazards cannot be ignored. The problem of cycling safety is becoming increasingly prominent due to perception delays or blind spots, especially in complex traffic scenarios such as at night or in bad weather, and the accident casualty rate is relatively high, urgently requiring an improvement in perception ability.
[0003] In the field of object perception and detection in cycling traffic scenarios, multi-sensor data fusion strategies have effectively improved detection accuracy by integrating complementary information. The current mainstream solutions generally adopt a heterogeneous sensor combination of radar and camera. Among them, radar has the ability to work all-weather and the characteristics of accurate distance and speed measurement, but it also faces challenges such as sparse point clouds and specular reflections; the camera can identify the categories of targets such as bicycles and electric vehicles through image texture features, but it is prone to missed detections and false detections at night with low illuminance or in bad weather such as rain and snow, and it cannot provide accurate three-dimensional spatial positioning. The collaborative fusion of the two can build an all-weather perception system that complements the radar's motion parameter detection characteristics and the visual target recognition ability, thereby reducing the risk of traffic accidents.
[0004] However, the current object detection methods using radar and camera data fusion still have some defects. For example: Zhou et al. proposed a method for realizing feature-level fusion of radar and camera in a bird's-eye view (Reference: Zhou, X., et al., "Radar and Camera Fusion in Bird's Eye View for Autonomous Driving," IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 5, 2022, pp. 1234-1245). This method uses radar to obtain the distance and speed information of the target, and at the same time extracts image features through the camera; then projects the two types of data onto the bird's-eye view plane for fusion, and finally completes object detection through a deep learning model. However, this method relies on high-precision sensors and powerful computing platforms, resulting in high hardware costs and complex deployment, which is not suitable for resource-constrained non-motorized vehicle scenarios; and its algorithm complexity is high and the interpretability is poor; in addition, there is also a problem of insufficient real-time performance, resulting in the inability to meet the requirements of the cycling scenario for rapid response.
[0005] For another example, Xu et al. proposed an Adaptive Fusion Network (AFnet) based on the attention mechanism to achieve the fusion of radar and camera data through a deep learning model. This method was trained on the NuScenes and CARLA datasets with the aim of improving object detection accuracy. (Reference: Xu C, Zhao H, Xie H, et al. Multi-sensor Decision-level Fusion Network Based on Attention Mechanism for Object Detection[J]. IEEE Sensors Journal, 2024.). However, this method relies on large-scale datasets for model training, and the network structure is complex, requiring a high-performance GPU to support real-time inference, which significantly increases the hardware cost and makes it difficult to be deployed in resource-constrained non-motor vehicle scenarios; moreover, due to the high computational complexity of the model, the inference speed is limited and cannot meet the real-time requirements of the non-motor vehicle multi-object detection system; in addition, the system's reliance on high-precision sensors for radar and cameras further exacerbates the cost and deployment difficulty, making it difficult to meet the low-cost constraint.
[0006] For another example, Alai H. and Rajamani R. proposed a system for detecting and tracking target vehicles by fusing data from low-cost 2-D radars and monocular cameras. The system locates the corners of the target vehicle by analyzing the spatial data of the radar and the image information of the camera, and infers the vehicle's pose based on this. Then, the system uses a high-gain observer to process the noise in the data, thereby stably calculating the position, speed, and direction of the target vehicle. However, on the one hand, the data points generated by the 2-D radar are relatively sparse and may not be able to completely capture the target in a complex environment, resulting in incomplete detection results; on the other hand, although the algorithm of the high-gain observer improves the accuracy, the computational amount is large, which may affect the real-time response ability of the system; and in rainy, foggy, or low-light weather conditions, the performance of the monocular camera will decline, thereby affecting the detection accuracy of the entire system.
[0007] In summary, the existing object detection solutions for fusing radar and camera data mainly suffer from the problems of high hardware deployment cost, low detection accuracy, and low efficiency. Summary of the Invention
[0008] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a multi-object detection method that can improve the detection accuracy and detection efficiency while reducing the hardware cost.
[0009] The purpose of the present invention can be achieved through the following technical solutions: A multi-object detection method includes the following steps: S1. Obtain image data and radar data and perform data calibration; S2. Obtain a visual feature vector and a radar feature vector based on the image data and radar data after data calibration; S3. Use a pre-trained fuzzy neural network model to perform fuzzy inference on the visual feature vector and the radar feature vector, and output a fuzzy inference result, where the fuzzy inference result includes the matching degree between each target in the image data and the radar data; S4. Screen out the targets with a matching degree exceeding a preset threshold from the fuzzy inference result.
[0010] Further, in step S1, the image data is obtained by a camera, and the radar data is obtained by a radar. In step S1, data calibration includes spatial alignment and time alignment of the data. Specifically, spatial alignment is achieved by correcting the yaw angle and pitch angle of the camera and the radar, so as to unify the image data and the radar data into the same coordinate system, realizing spatial alignment of the data; Time alignment is specifically achieved by synchronously starting the camera and the radar through an external trigger signal, sorting the data by time stamps, taking the time stamp of the camera as the standard, setting a time window, and selecting the radar data with the smallest time difference and the image data for time alignment; Or vice versa, taking the time stamp of the radar as the standard, setting a time window, and selecting the image data with the smallest time difference and the radar data for time alignment.
[0011] Further, the specific process of correcting the yaw angle and pitch angle of the camera and the radar is as follows: According to the radar yaw angle output by the radar and the radar pitch angle , determine the camera yaw angle and the camera pitch angle of the camera through the following formula:
[0012]
[0013]
[0014] where, ( , ) is the central pixel coordinate of the target box in the camera, 、 is the focal length of the camera, 、 is the optical center of the camera, ( , ) is the coordinate of the target in the normalized plane; Afterwards, the difference between the radar yaw angle and the camera yaw angle is used as a compensation value for correcting the radar yaw angle, and the difference between the radar pitch angle and the camera pitch angle is used as a compensation value for correcting the radar pitch angle, so as to unify the image data and the radar data into the camera coordinate system; Or vice versa, unify the image data and the radar data into the radar coordinate system.
[0015] Furthermore, the step S2 specifically processes the image data through the YOLO algorithm to obtain a visual feature vector suitable for input to the fuzzy neural network ; and performs preprocessing on the point cloud data of the radar data to obtain a radar feature vector .
[0016] Furthermore, the visual feature vector x i is specifically: , wherein, represents the th target in the image data, ( , ) are the coordinates of the th target in the image data, and are the width and height of the target bounding box respectively, is the target category, is the confidence of the target; The radar feature vector y j is specifically: , wherein, represents the th target in the radar data, ( , ) are the coordinates of the th target in the radar data, is the target distance detected by the radar, is the radar yaw angle, is the target speed detected by the radar, is the target movement trend detected by the radar.
[0017] Furthermore, the pre-trained fuzzy neural network model in the step S3 includes an input layer, a fuzzification layer, a rule layer, a normalization layer and an output layer; The input layer is used to input the visual feature vector and the radar feature vector into the fuzzification layer; The fuzzification layer is used to semantically describe the input quantity through the membership function to map the input quantity to the fuzzy set; The rule layer is used to reason and calculate the fuzzified input quantity according to the fuzzy rules; The normalization layer is used to convert the result of fuzzy inference and calculation into an accurate fuzzy score; The output layer is used to output the matching degree between each target in the image data and the radar data.
[0018] Furthermore, the specific working process of the fuzzification layer is as follows: Describe the distance between the th target in the image data and the th target in the radar data: , Obtain 3 fuzzy sets: matching, partially matching, not matching, and describe them using a piecewise linear function: ; Among them, is the fuzzy score corresponding to the distance fuzzy matching; Describe the target speed detected by the radar, and obtain 5 fuzzy sets: low, slightly low, medium, slightly fast, fast, and describe them using a piecewise linear function: , , , , ; Among them, , , , , are the piecewise linear functions corresponding to the 5 fuzzy sets of "low, slightly low, medium, slightly fast, fast", respectively; The reason for dividing the target speed detected by the radar into five fuzzy sets is to effectively reflect the motion characteristics of different targets. According to traffic regulations, the maximum speed of an electric bicycle cannot exceed 25 km / h , and the speed of a motor vehicle on an urban road cannot exceed 50 km / h, Therefore, taking 0~15 km / h , 10~25 km / h , 20~35 km / h, 30 - 45 km / h , 40 - 45+ km / h Divide the interval. For this interval, when v is less than 10 km / h , it is considered that the speed is low, so the fuzzy membership degree is 1. When 10 < v ≤ 15, use the formula to linearly describe the membership degree of the speed to this fuzzy set. When v > 15, the membership degree is 0, which conforms to the triangular membership function in the fuzzy neural network; Similarly, when v < 40, the membership degree is 0. When 40 < v ≤ 45, use the formula to linearly describe the membership degree of the speed to this fuzzy set. When v > 45, the membership degree is 1; For , , these three intervals, use the trapezoidal membership function in the fuzzy neural network to achieve smooth transition through linear increase / decrease for easy description; By describing the change rate of the target bounding box in the image data, 2 fuzzy sets are obtained: approaching, moving away, and are described using the following membership functions: , , where is the membership function corresponding to the "approaching" fuzzy set, is the membership function corresponding to the "moving away" fuzzy set; Since the change rate of the target bounding box may be affected by factors such as illumination and occlusion, the Gaussian membership function in the fuzzy neural network is used to smooth these uncertainties to effectively handle the slight differences in the change rate of the bounding box.
[0019] Furthermore, the fuzzy rules include: Rule 1: If "speed = fast" and "radar movement trend = approaching" and "target type = car", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; where is the fuzzy score of Rule 1, corresponds to "speed = fast", corresponds to "radar movement trend = approaching", corresponds to "target type = car"; Rule 2: If "speed = slightly fast" and "radar movement trend = approaching" and "target type = electric bicycle", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 2, corresponding to "speed = slightly fast", corresponding to "radar movement trend = approaching", corresponding to "target type = electric bicycle"; Rule 3: If "speed = low" and "radar movement trend = moving away" and "target type = electric bicycle", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 3, corresponding to "speed = low", corresponding to "radar movement trend = moving away", corresponding to "target type = electric bicycle"; Rule 4: If "speed = slightly low" and "radar movement trend = approaching" and "target type = bicycle", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 4, corresponding to "speed = slightly low", corresponding to "radar movement trend = approaching", corresponding to "target type = bicycle"; Rule 5: If "speed = medium" and "radar movement trend = moving away" and "target type = bicycle", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 5, corresponding to "speed = medium", corresponding to "radar movement trend = moving away", corresponding to "target type = bicycle"; Rule 6: If "speed = fast" and "radar movement trend = moving away" and "target type = pedestrian", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 6, Corresponds to "Speed = Fast", Corresponds to "Radar movement trend = Away", Corresponds to "Target type = Pedestrian"; Rule 7: If "Radar movement trend = Approaching" and "Visual detection movement trend = Approaching", then "Matching possibility = High", and the corresponding reasoning and calculation formula are: ; Wherein, is the fuzzy score of Rule 7, Corresponds to "Visual detection movement trend = Approaching", Corresponds to "Radar movement trend = Approaching"; Rule 8: If "Radar movement trend = Away" and "Visual detection movement trend = Away", then "Matching possibility = High", and the corresponding reasoning and calculation formula are: ;
[0020] Wherein, is the fuzzy score of Rule 8, Corresponds to "Visual detection movement trend = Away", Corresponds to "Radar movement trend = Away".
[0021] The above Rules 1 to 6 are designed based on the fuzzy speed of the target, the radar movement trend , and the target type judged by the vision module . According to experience, it is considered that the target that can approach at a high speed is a car accelerating from behind, so Rule 1 is designed; the target that can approach at a relatively high speed is an electric vehicle accelerating from behind, so Rule 2 is designed; the target that can move away at a low speed is an electric vehicle with a speed slightly lower than its own behind, so Rule 3 is designed; the target that can approach at a relatively low speed is a bicycle accelerating from behind, so Rule 4 is designed; the target that can move away at a medium speed is a bicycle with a speed lower than its own behind, so Rule 5 is designed; the target behind moving away at a high speed is a pedestrian with a speed much lower than its own (the relative speed between the two is large), so Rule 6 is designed.
[0022] The above Rules 7 and 8 are designed based on the movement trend fed back by the radar and the fuzzy movement trend obtained according to the target box change rate . When both are approaching, they are regarded as matching, and when both are away, they are regarded as matching.
[0023] Furthermore, the fuzzy score output by the normalization layer is specifically: , , , wherein, is the fuzzy score obtained by matching between the i th visual target and the j th radar target, , , are the activation weights of the three types of fuzzy matching results of distance fuzzy matching , speed fuzzy matching , and motion trend fuzzy matching respectively, is the maximum value selected from the fuzzy scores obtained from Rule 1 to Rule 6, is the maximum value selected from the fuzzy scores obtained from Rule 7 and Rule 8.
[0024] Furthermore, the output layer specifically outputs a matching matrix between the image data and each target in the radar data based on the fuzzy score and the multi-target matching matrix.
[0025] Compared with the prior art, the present invention has the following advantages: After acquiring the image data and the radar data, the present invention first performs data calibration, and then based on the image data and the radar data after data calibration, obtains a visual feature vector and a radar feature vector. Then, a pre-trained fuzzy neural network model is used to perform fuzzy inference on the visual feature vector and the radar feature vector, and outputs a fuzzy inference result. Among them, the fuzzy inference result includes the matching degree between the image data and each target in the radar data. Finally, the targets with a matching degree exceeding a preset threshold are screened out from the fuzzy inference result. Thus, the fuzzy neural network is used to fuse the visual feature vector and the radar feature vector data, fuzzify the quantified input features, and perform matching through fuzzy rules. This method has high interpretability and can effectively reduce the demand for input data in model training. At the same time, the number of structural layers of the fuzzy neural network is lower than that of deep learning frameworks such as convolutional neural networks. Therefore, its network structure is relatively simple, and the required computing power is much lower than that of other algorithms, which is suitable for edge computing platforms with limited computing power and effectively reduces the deployment cost. The present invention can effectively solve the problems of low detection accuracy, high hardware cost, insufficient real-time performance, and lack of classification of specific non-motor vehicle targets in traditional fusion methods.
[0026] The present invention performs data calibration on image data and radar data, including spatial alignment and temporal alignment. Spatial alignment corrects the yaw and pitch angles of the camera and radar, thereby unifying the image and radar data into the same coordinate system and achieving spatial alignment. Temporal alignment sorts the data by timestamp, sets a time window based on the camera's timestamp, and selects the radar data with the smallest time difference for time alignment with the image data. Conversely, sets a time window based on the radar's timestamp and selects the image data with the smallest time difference for time alignment with the radar data. This data calibration ensures that subsequent processing accurately obtains visual feature vectors and radar feature vectors.
[0027] On the one hand, the present invention processes image data through the YOLO algorithm to obtain visual feature vectors suitable for fuzzy neural network input, which can greatly improve detection efficiency while ensuring accuracy as much as possible, reducing the computing power requirements of the deployment platform; on the other hand, point cloud data preprocessing is performed on radar data to obtain radar feature vectors, which can extract key features, identify and eliminate outliers or noise data, and improve the detection accuracy of low-cost radar.
[0028] In the present invention, the pre-trained fuzzy neural network model includes an input layer, a fuzzification layer, a rule layer, a normalization layer, and an output layer. The fuzzification layer semantically describes the input quantity through a preset membership function and maps the input quantity to a fuzzy set. The rule layer infers and calculates the fuzzified input quantity according to fuzzy rules. The normalization layer converts the results of the fuzzy reasoning and calculation into an accurate fuzzy score. Finally, the output layer outputs the degree of match between the image data and each target in the radar data. The present invention utilizes fuzzy neural networks and weight distribution to fuse visual and radar data, introduces expert knowledge to design membership functions, compiles fuzzy rules, fuzzifies quantized input features, and matches them using fuzzy rules. This method can achieve highly robust and low-cost target detection and recognition, and has the advantages of real-time and interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Schematic diagram of the method flow of the present invention; Figure 2 Schematic diagram of the application process of the embodiment; Figure 3 Schematic diagram of the field of view of the camera and radar in the embodiment; Figure 4 Schematic diagram of the working process of the fuzzy neural network model in the embodiment; Figure 5 Schematic diagram comparing the data fusion time of this solution and other algorithms in the embodiment. DETAILED DESCRIPTION
[0030] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] Embodiment
[0032] As Figure 1 shown, a multi-object detection method includes the following steps: S1. Obtain image data and radar data, and perform data calibration; S2. Based on the image data and radar data after data calibration, obtain visual feature vectors and radar feature vectors; S3. Use a pre-trained fuzzy neural network model to perform fuzzy inference on the visual feature vectors and radar feature vectors, and output a fuzzy inference result, where the fuzzy inference result includes the matching degree between each target in the image data and the radar data; S4. Screen out the targets whose matching degree exceeds a preset threshold from the fuzzy inference result.
[0033] This embodiment applies the above solution. As Figure 2 shown, the main contents are as follows: First, obtain image data through a camera and obtain radar data through a radar. In this embodiment, a millimeter-wave radar is selected. The radar is installed under the seat of a non-motor vehicle or a bicycle, perpendicular to the ground and ensuring no occlusion. The radar can directly output the yaw angle, pitch angle, speed, distance, and radar movement trend (approaching / leaving) of the detected targets. The camera is also installed under the seat of the non-motor vehicle or the bicycle, installed side by side with the radar. The field of view of the camera covers 120° to the rear. As Figure 3 shown, the camera outputs the types (vehicles, bicycles, electric bicycles, pedestrians, etc.), confidence levels (0-1), and target box area change rates of the detected targets through the YOLO (such as YOLOv8n) algorithm.
[0034] In this embodiment, the execution entity of the multi-object detection method is a software or a hardware device, including but not limited to at least one of the following: user equipment, network equipment, etc. Among them, the user equipment may include but not be limited to computers, smartphones, personal digital assistants (Personal Digital Assistant, abbreviated as: PDA) and other electronic devices, etc. The network equipment may include but not be limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, consisting of a super virtual computer composed of a group of loosely coupled computers.
[0035] In this embodiment, an edge computing platform is specifically used to execute the above multi-object detection method, such as the RK3576 embedded platform. This platform runs a Fuzzy Neural Networks (FNN) model and the YOLOv8n algorithm. It is used to process the data obtained by the camera and the radar and output the matching results of the targets in the two types of data. Those with a matching degree exceeding the preset threshold can be determined as real targets, otherwise they are ignored. In this embodiment, the original network backbone of YOLOv8n is replaced by MobileNetV3, and a Convolutional Block Attention Module (CBAM) is added for optimization. The C3Ghost module is used in the detection head part to further reduce the number of parameters, which can greatly improve the detection efficiency while ensuring the accuracy as much as possible, reducing the computing power requirements for the deployment platform. The original computing amount of YOLOv8n is 8.2 GFLOPS, and the computing amount of the improved new network is 4.6 GFLOPS.
[0036] After acquiring the image data and radar data, data calibration is performed, which specifically includes spatial alignment and time alignment of the data.
[0037] In terms of spatial alignment, by correcting the yaw angle and pitch angle of the camera and the radar, the image data and radar data are unified into the same coordinate system to achieve spatial alignment of the data. In terms of time alignment, the camera and the radar are synchronously started through an external trigger signal, and the data is sorted by timestamp. Taking the timestamp of the camera as the standard, a time window is set, and the radar data with the smallest time difference is selected for time alignment with the image data; or vice versa, taking the timestamp of the radar as the standard, a time window is set, and the image data with the smallest time difference is selected for time alignment with the radar data.
[0038] Among them, the steps of correcting the yaw angle and pitch angle of the camera and the radar specifically include: The radar directly outputs the radar yaw angle and the radar pitch angle , and the camera yaw angle of the camera is determined through the following formula and the camera pitch angle ;
[0039]
[0040]
[0041] Among them, ( , ) is the central pixel coordinate of the target box in the camera. 、 is the focal length of the camera, 、 is the optical center of the camera, ( , ) are the coordinates of the target in the normalized plane; Take the difference between the radar yaw angle and the camera yaw angle as the compensation value for correcting the radar yaw angle, and take the difference between the radar pitch angle and the camera pitch angle as the compensation value for correcting the radar pitch angle. The specific operation is to directly add the corresponding compensation value to the radar yaw angle and the radar pitch angle, so as to unify the radar data into the camera coordinate system. Of course, in other embodiments, the image data can also be unified into the radar coordinate system.
[0042] It is easy to understand that there may be multiple different targets in the image data. Therefore, the camera pitch angle and the camera yaw angle of each target can be obtained respectively, and then take the average value of, so as the compensation value for correcting the radar yaw angle, take the average value of as the compensation value for correcting the radar pitch angle.
[0043] The pre-trained fuzzy neural network model includes an input layer, a fuzzification layer, a rule layer, a normalization layer and an output layer, as Figure 4 shown. Among them, the input layer is used to input the visual feature vector and the radar feature vector into the fuzzification layer; the fuzzification layer is used to semantically describe the input quantity through the membership function and map the input quantity to the fuzzy set; the rule layer is used to reason and calculate the fuzzified input quantity according to the fuzzy rules; the normalization layer is used to convert the result of the fuzzy reasoning and calculation into an accurate fuzzy score; the output layer is used to output the matching degree between the image data and each target in the radar data. The number of structural layers of the fuzzy neural network is lower than that of deep learning frameworks such as convolutional neural networks. Therefore, its network structure is relatively simple, and the required computing power is only 3.2 MFLOPS, which is much lower than other algorithms. It can be well applied to the edge computing platform with limited computing power, effectively reducing the deployment cost and improving the operation efficiency. To verify the effectiveness of this solution, in this embodiment, in the same experimental environment, for the visual module data and radar data collected within the same time period, the method proposed in this solution is compared and tested with other existing algorithms (including four multi-modal target detection fusion algorithms, namely AVOD, MV3D, F-PointNet and ContFuse). Each round is repeated 100 times and the average value of the fusion time is taken, and a total of 56 rounds of experiments are carried out. The specific experimental comparison results are as Figure 5 shown, Figure 5In the method proposed in this solution, FNN Figure 5 The abscissa corresponds to the number of running rounds, and the ordinate corresponds to the average fusion time. Experimental data shows that the time required for data fusion in this solution is 0.18 ms, which is 82% faster than other algorithms.
[0044] 1. Input layer The image data obtained by the camera cannot be directly input into the fuzzy neural network. Therefore, the YOLO algorithm is used to process the image data to obtain a visual feature vector suitable for input into the fuzzy neural network. x i 。
[0045] Among them, , represents the th target in the image data. ([[]] , ) is the coordinate of the th target in the image data. and are the width and height of the target bounding box respectively. is the target category. is the confidence of the target; after preprocessing the point cloud data of the radar data, a radar feature vector is formed : Among them, represents the th target in the radar data. ([[]] , ) is the coordinate of the th target in the radar data. is the distance of the target detected by the radar. is the yaw angle of the radar. is the speed of the target detected by the radar. is the movement trend of the target detected by the radar, including approaching and moving away.
[0046] 2. Fuzzification layer 1) By describing the distance between the th target in the image data and the th target in the radar data, 3 fuzzy sets are obtained, namely matching, partial matching, and non-matching. Among them , a piecewise linear function is used for description: 。
[0047] 2) By the target speed detected by the radar Describe to obtain five fuzzy sets: low, slightly low, medium, slightly fast, and fast, which are described using piecewise linear functions: , , , , ; 3) Describe the change rate of the target bounding box in the image data to obtain two fuzzy sets: approaching and moving away, which are described using the following membership functions:
[0048] , .
[0049] 3. Rule layer The fuzzy rules include: Rule 1: If "speed = large" and "radar movement trend = approaching" and "target type = car", then "matching possibility = high", and the reasoning and calculation formula are: .
[0050] Rule 2: If "speed = relatively large" and "radar movement trend = approaching" and "target type = electric bicycle", then "matching possibility = high", and the reasoning and calculation formula are: .
[0051] Rule 3: If "speed = small" and "radar movement trend = moving away" and "target type = electric bicycle", then "matching possibility = high", and the reasoning and calculation formula are: .
[0052] Rule 4: If "speed = relatively small" and "radar movement trend = approaching" and "target type = bicycle", then "matching possibility = high", and the reasoning and calculation formula are: .
[0053] Rule 5: If "speed = medium" and "radar movement trend = moving away" and "target type = bicycle", then "matching possibility = high", and the reasoning and calculation formula are: .
[0054] Rule 6: If "speed = large" and "radar movement trend = moving away" and "target type = pedestrian", then "matching possibility = high", and the reasoning and calculation formula are: 。
[0055] Rule 7: If "radar movement trend = approaching" and "visual detection movement trend = approaching", then "matching possibility = high". The reasoning and calculation formula are as follows: 。
[0056] Rule 8: If "radar movement trend = moving away" and "visual detection movement trend = moving away", then "matching possibility = high". The reasoning and calculation formula are as follows: 。
[0057] 4. Normalization layer The normalization layer is implemented through the following formula:
[0058] , , , where 、 、 are the activation weights of the three types of fuzzy matching results of distance fuzzy matching 、speed fuzzy matching 、movement trend fuzzy matching respectively. is the total activation weight. During model training, its initial value is set according to expert experience. The specific training process of the model is common knowledge and will not be elaborated in this embodiment.
[0059] 5. Output layer The output layer outputs the matching matrix between each target in the image data and the radar data through fuzzy scoring and a multi-target matching matrix. For N radar targets and M visual targets, an N×M matching matrix is generated, and the Hungarian algorithm is used to extract the optimal matching. In this embodiment, there are 2 radar targets and 3 visual targets, and the FNN outputs a 2×3 matching matrix:
[0060] For example, if the preset threshold is set to 0.8, when the matching degree between radar target 1 and visual target 1 and the matching degree between radar target 2 and visual target 3 both exceed 0.8, it indicates that radar target 1 matches visual target 1 and radar target 2 matches visual target 3. This shows that radar target 1 and visual target 1 are the same and are real detected targets, radar target 2 and visual target 3 are the same and are real detected targets, and visual target 2 is a detection error and not a real target.
[0061] In summary, this solution uses a fuzzy neural network and weight allocation to fuse visual and radar data, achieving high-robustness and low-cost object detection and recognition. It is applicable to cycling traffic, can solve the problems of low detection accuracy, high hardware cost, insufficient real-time performance, and lack of specific object classification for non-motor vehicles in traditional fusion methods, effectively improving the safety perception ability of cyclists in complex traffic scenarios, and having the advantages of real-time performance and interpretability.
Claims
1. A multi-object detection method, characterized in that, It includes the following steps: Step S1: Obtain image data and radar data, and perform data calibration; Step S2: Based on the image data and radar data after data calibration, obtain visual feature vectors and radar feature vectors; Step S3: Use a pre-trained fuzzy neural network model to perform fuzzy inference on the visual feature vectors and radar feature vectors, and output fuzzy inference results. Among them, the fuzzy inference results include the matching degree between each target in the image data and the radar data; Step S4: Screen out the targets with a matching degree exceeding a preset threshold from the fuzzy inference results.
2. The multi-object detection method according to claim 1, characterized in that In step S1, the image data is obtained by a camera, and the radar data is obtained by a radar. In step S1, the data calibration includes spatial alignment and time alignment of the data. Among them, the spatial alignment is specifically achieved by correcting the yaw angle and pitch angle of the camera and the radar, so as to unify the image data and the radar data into the same coordinate system and realize the spatial alignment of the data; The time alignment is specifically achieved by synchronously starting the camera and the radar through an external trigger signal, sorting the data according to the time stamp, taking the time stamp of the camera as the standard, setting a time window, and selecting the radar data with the smallest time difference and the image data for time alignment; Or vice versa, taking the time stamp of the radar as the standard, setting a time window, and selecting the image data with the smallest time difference and the radar data for time alignment.
3. A multi-object detection method according to claim 2, characterized in that, The specific process of correcting the yaw angle and pitch angle of the camera and the radar is as follows: According to the radar yaw angle output by the radar and the radar pitch angle , determine the camera yaw angle of the camera through the following formula and the camera pitch angle : Among them, ( , ) are the central pixel coordinates of the target box in the camera, 、 is the focal length of the camera, 、 is the optical center of the camera, ( , ) are the coordinates of the target in the normalized plane; Afterwards, the difference between the radar yaw angle and the camera yaw angle is used as the compensation value for correcting the radar yaw angle, and the difference between the radar pitch angle and the camera pitch angle is used as the compensation value for correcting the radar pitch angle, so as to unify the image data and the radar data into the camera coordinate system; Or vice versa, unify the image data and the radar data into the radar coordinate system.
4. A multi-object detection method according to claim 1, characterized in that Specifically, in step S2, the YOLO algorithm is used to process the image data to obtain a visual feature vector suitable for input to the fuzzy neural network ; And perform point cloud data preprocessing on the radar data to obtain radar feature vectors .
5. A multi-object detection method according to claim 4, characterized in that, The visual feature vector x i Specifically: , Among them, represents the th target in the image data, ( , ) are the coordinates of the th target in the image data, and are the width and height of the target bounding box respectively, is the target category, is the confidence of the target; The radar feature vector y j Specifically: , Among them, represents the th target in the radar data, ( , ) is the coordinate of the th target in the radar data, is the target distance detected by the radar, is the radar yaw angle, is the target speed detected by the radar, is the target motion trend detected by the radar.
6. A multi-object detection method according to claim 5, characterized in that, The pre-trained fuzzy neural network model in step S3 includes an input layer, a fuzzification layer, a rule layer, a normalization layer, and an output layer; The input layer is used to input the visual feature vectors and radar feature vectors into the fuzzification layer; The fuzzification layer is used to semantically describe the input quantity through a membership function to map the input quantity to a fuzzy set; The rule layer is used to perform inference and calculation on the fuzzified input quantity according to fuzzy rules; The normalization layer is used to convert the results of fuzzy inference and calculation into accurate fuzzy scores; The output layer is used to output the matching degree between each target in the image data and the radar data.
7. A multi-object detection method according to claim 6, characterized in that The specific working process of the fuzzification layer is as follows: Describe the distance between the th target in the image data and the th target in the radar data as follows: , Obtain 3 fuzzy sets: matching, partial matching, and non-matching, which are described using piecewise linear functions: ; Describe the target speed detected by the radar to obtain five fuzzy sets: low, slightly low, medium, slightly fast, and fast, and use piecewise linear functions to describe them: , , , , ; Among them, , , , , are piecewise linear functions corresponding to the 5 fuzzy sets of "low, slightly low, medium, slightly fast, fast", respectively; By describing the change rate of the target bounding box in the image data Two fuzzy sets, approaching and moving away, are obtained and described using the following membership functions: , , wherein, is the membership function corresponding to the "close" fuzzy set, is the membership function corresponding to the "far" fuzzy set.
8. A multi-object detection method according to claim 7, characterized in that, The fuzzy rules include: Rule 1: If "speed = fast" and "radar movement trend = approaching" and "target type = car", then "matching possibility = high", and the corresponding inference and calculation formula is: ; Among them, is the fuzzy score of Rule 1, corresponding to "speed = fast", corresponding to "radar movement trend = approaching", corresponding to "target type = car"; Rule 2: If "speed = slightly fast" and "radar movement trend = approaching" and "target type = electric bicycle", then "matching possibility = high", and the corresponding inference and calculation formula is: ; Among them, is the fuzzy score of Rule 2, corresponding to "speed = slightly faster", corresponding to "radar movement trend = approaching", corresponding to "target type = electric bicycle"; Rule 3: If "speed = low" and "radar movement trend = moving away" and "target type = electric bicycle", then "matching possibility = high", and the corresponding inference and calculation formula is: ; Among them, is the fuzzy score of Rule 3, corresponding to "speed = low", corresponding to "radar movement trend = away", corresponding to "target type = electric bicycle"; Rule 4: If "speed = slightly low" and "radar movement trend = approaching" and "target type = bicycle", then "matching possibility = high", and the corresponding inference and calculation formula is: ; Among them, is the fuzzy score of Rule 4, corresponding to "Speed = Slightly low", corresponding to "Radar movement trend = Approaching", corresponding to "Target type = Bicycle"; Rule 5: If "speed = medium" and "radar movement trend = away" and "target type = bicycle", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 5, corresponding to "Speed = Medium", corresponding to "Radar movement trend = Away", corresponding to "Target type = Bicycle"; Rule 6: If "speed = fast" and "radar movement trend = away" and "target type = pedestrian", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 6, corresponding to "Speed = Fast", corresponding to "Radar movement trend = Away", corresponding to "Target type = Pedestrian"; Rule 7: If "radar movement trend = approaching" and "visual detection movement trend = approaching", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 7, corresponding to "Visual detection movement trend = approaching", corresponding to "Radar movement trend = approaching"; Rule 8: If "radar movement trend = away" and "visual detection movement trend = away", then "matching possibility = high", and the corresponding reasoning and calculation formula are: ; Among them, is the fuzzy score of Rule 8, corresponding to "Visual detection movement trend = Away", corresponding to "Radar movement trend = Away".
9. A multi-object detection method according to claim 8, characterized in that, The specific fuzzy score output by the normalization layer is: , , , Among them, is the fuzzy score obtained by matching between the i th visual target and the j th radar target. , , are the activation weights of the three types of fuzzy matching results of distance fuzzy matching , speed fuzzy matching , and motion trend fuzzy matching respectively. is the maximum value selected from the fuzzy scores obtained from Rule 1 to Rule 6, is the maximum value selected from the fuzzy scores obtained from Rule 7 and Rule 8.
10. A multi-object detection method according to claim 6, characterized in that, The output layer specifically outputs the matching matrix between each target in the image data and the radar data based on the fuzzy score and the multi-target matching matrix.
Citation Information
Patent Citations
Calibration device and radar and camera combined calibration method and system
CN110310339A
Target classification method based on fuzzy logic reasoning
CN111474538A
Target detection method and device based on Leiyu fusion, storage medium and product
CN119205750A
Hilly and mountainous area tracked vehicle intelligent chassis system and algorithm based on environment perception and road surface recognition
CN119414830A
Radar and camera detection target matching method and system in highway scene
CN119888646A