Radar and Video Fusion Based Target Detection Method, Device, Equipment and Medium

Through the improved radar and video fusion module in the YOLO model, the radar and video data are deeply integrated, which solves the problem of low detection accuracy in harsh environments in traditional methods, and achieves higher target detection accuracy and robustness.

CN119229348BActive Publication Date: 2025-05-27HEBEI DEGUROON ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411541729.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-27
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

Traditional target detection methods that rely on single video data perform poorly in bad weather or low lighting conditions, and the radar data lacks intuitive images, resulting in low accuracy, poor accuracy and low robustness of target detection.

Method used

Through the radar and video fusion module based on the improved YOLO model, the radar point cloud data and video data are preprocessed, features are extracted, feature splicing, fusion, residual processing, secondary splicing and convolution processing are performed to achieve deep fusion of radar and video data.

Benefits of technology

Improve the accuracy and robustness of target detection, especially in conditions such as severe weather, low lighting or target occlusion, which reduces dependence on environmental conditions, reduces the consumption of computing resources, and maintains rapid response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229348B_ABST
    Figure CN119229348B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of electronic digital data processing technology, and provides a target detection method, device, equipment and medium based on radar and video fusion, the method comprising: obtaining radar point cloud data and video data of a target to be detected; processing the radar point cloud data and video data based on the radar and video fusion module in the improved YOLO model to obtain a target detection result; wherein the radar and video fusion module is used to pre-process the radar point cloud data and video data to obtain a radar image and a video image, and perform feature extraction, feature splicing, feature fusion, residual processing, secondary splicing and convolution processing on the radar image and the video image. The present application can effectively fuse radar and video data to improve the accuracy and robustness of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of electronic digital data processing, and particularly relates to a target detection method, device, equipment and medium based on radar and video fusion. Background Technique

[0002] Currently, the environmental perception devices used in intelligent transportation mainly include cameras and millimeter-wave radars. Millimeter-wave radars have the detection capabilities of target distance, speed, azimuth angle, etc., and at the same time, quantitatively test the target reflection ability at a specific incident angle, and give an estimate of the target size from the echo scattering characteristics of the millimeter-wave radar.

[0003] Cameras often use multi-sensors for information fusion, and use multi-sensor information fusion technology to achieve complementary advantages of each sensor, and accurately collect dimension information such as the color, size, dimensions, and projection contour of the target to the greatest extent, improving the comprehensiveness and accuracy of the environmental perception system.

[0004] In target detection technology, it is crucial to accurately detect targets. However, traditional methods relying on single video data perform poorly in bad weather or low lighting conditions. While millimeter-wave radars, as supplementary sensors, can provide stable target information, they lack intuitive images. Although there are already methods that combine radar and video for target detection, the accuracy of target detection is not high, the accuracy is poor, and the robustness is also low. Summary of the Invention

[0005] The embodiments of this application provide a target detection method, device, equipment and medium based on radar and video fusion to effectively fuse radar and video data and improve the accuracy and robustness of target detection.

[0006] This application is implemented through the following technical solutions:

[0007] In a first aspect, the embodiments of this application provide a target detection method based on radar and video fusion, including:

[0008] Obtain the radar point cloud data and video data of the target to be detected.

[0009] Based on the processing results of the radar point cloud data and video data by the radar and video fusion module in the improved YOLO model, obtain the target detection result; wherein, the radar and video fusion module is used to preprocess the radar point cloud data and video data to obtain a radar image and a video image, and perform feature extraction, feature splicing, feature fusion, residual processing, secondary splicing and convolution processing on the radar image and the video image.

[0010] In combination with the first aspect, in some possible implementation manners, performing feature extraction, feature splicing, feature fusion, residual processing, secondary splicing and convolution processing on the radar image and the video image includes:

[0011] Perform a 3x3 convolution operation on the radar image to extract the first radar image feature; perform a 3x3 convolution operation on the video image to extract the first video image feature.

[0012] Concatenate the first radar image feature and the first video image feature through a concat operation to obtain the first concatenated feature.

[0013] Extract and fuse the first concatenated feature through a transformer Encoder layer to obtain the first pre-output feature.

[0014] Perform residual processing and a 3x3 convolution operation on the radar image and the video image to obtain the second pre-output feature.

[0015] Perform a concat channel concatenation on the first pre-output feature and the second pre-output feature to obtain the third pre-output feature.

[0016] Perform a 1x1 convolution operation on the third pre-output feature to obtain the output result of the radar and video fusion module.

[0017] Combined with the first aspect, in some possible implementation manners, the improved YOLO model further includes a feature fusion module;

[0018] Based on the processing results of the radar and video fusion module in the improved YOLO model for the radar point cloud data and the video data, obtain the target detection result, including:

[0019] Process the processing results of the radar and video fusion module in the improved YOLO model for the radar point cloud data and the video data through at least one feature fusion module to obtain the target detection result; wherein, the feature fusion module is used to perform residual processing, convolution, depthwise separable convolution processing, concat fusion and add fusion operation processing on the input image.

[0020] Combined with the first aspect, in some possible implementation manners, performing residual processing, convolution, depthwise separable convolution processing, concat fusion and add fusion operation processing on the input image includes:

[0021] Perform residual processing on the input image to obtain the first feature.

[0022] Perform a 1xk convolution operation on the input image to obtain the second feature.

[0023] Perform a kx1 convolution operation on the input image to obtain the third feature.

[0024] Perform Ghost convolution and 3x3 depthwise separable convolution on the input image to obtain the fourth feature.

[0025] Concatenate and fuse the second feature, the third feature, and the fourth feature to obtain the first fused feature.

[0026] Perform an add operation on the first feature and the first fused feature to obtain the fourth pre-output feature.

[0027] Successively perform Ghost convolution and 1x1 convolution on the fourth pre-output feature to obtain the output result of the feature fusion module.

[0028] Combined with the first aspect, in some possible implementation manners, the input image is subjected to residual processing, convolution, depthwise separable convolution processing, concatenate fusion, and add fusion operation processing, including:

[0029] Perform residual processing on the input image to obtain the first feature.

[0030] Perform a 1xk convolution operation on the input image to obtain the second feature

[0031] Perform a kx1 convolution operation on the input image to obtain the third feature;

[0032] Perform Ghost convolution and 3x3 depthwise separable convolution on the input image to obtain the fourth feature.

[0033] Concatenate and fuse the second feature, the third feature, and the fourth feature to obtain the first fused feature.

[0034] Repeat the above steps multiple times to obtain multiple first fused features.

[0035] Perform an add operation on the first feature and multiple first fused features to obtain the fourth pre-output feature.

[0036] Successively perform Ghost convolution and 1x1 convolution on the fourth pre-output feature to obtain the output result of the feature fusion module.

[0037] Combined with the first aspect, in some possible implementation manners, the improved YOLO model includes: a Backbone layer, a Neck layer, and a Head layer; the feature fusion module includes: a first feature fusion module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, a fifth feature fusion module, a sixth feature fusion module, a seventh feature fusion module, and an eighth feature fusion module.

[0038] The Backbone layer includes: a radar and video fusion module, a first CBS module, a first feature fusion module, a second CBS module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, and an SPPF module connected in sequence.

[0039] The Neck layer includes: a PSA sampling attention module, an Upsample module, a first Concat module, a fifth feature fusion module, a second Concat module, a sixth feature fusion module, a third CBS module, a third Concat module, a seventh feature fusion module, a fourth CBS module, a fourth Concat module, and an eighth feature fusion module connected in sequence.

[0040] The Head layer includes: a first Head module, a second Head module, and a third Head module.

[0041] The second CBS module is also connected to the second Concat module; the third feature fusion module is also connected to the first Concat module; the SPPF module is also connected to the PSA sampling attention module; the PSA sampling attention module is also connected to the fourth Concat module; the fifth feature fusion module is also connected to the third Concat module; the sixth feature fusion module is also connected to the first Head module; the seventh feature fusion module is also connected to the second Head module; the eighth feature fusion module is also connected to the third Head module.

[0042] In combination with the first aspect, in some possible implementation manners, the radar and video fusion module is used to preprocess radar point cloud data and video data to obtain a radar image and a video image, including:

[0043] The radar and video fusion module converts the radar point cloud data into an image to obtain a radar image.

[0044] The video data is intercepted frame by frame through the OpenCV method to obtain a video image.

[0045] In a second aspect, an embodiment of the present application provides an object detection device based on radar and video fusion, including:

[0046] A data acquisition module for acquiring radar point cloud data and video data of a target to be detected.

[0047] A result output module for obtaining an object detection result based on the processing results of the radar point cloud data and the video data by the radar and video fusion module in the improved YOLO model; wherein, the radar and video fusion module is used to preprocess the radar point cloud data and the video data to obtain a radar image and a video image, and perform feature extraction, feature splicing, feature fusion, residual processing, secondary splicing, and convolution processing on the radar image and the video image.

[0048] In a third aspect, an embodiment of the present application provides a terminal device, including: a processor and a memory, where the memory is used to store a computer program, and the processor implements the object detection method based on radar and video fusion according to any one of the first aspect when executing the computer program.

[0049] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the radar and video fusion-based target detection method according to any one of the first aspect.

[0050] It can be understood that the beneficial effects of the above second aspect to fourth aspect can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0051] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0052] In the present application, the radar and video fusion module processes radar point cloud data and video data, realizing the deep fusion of the two types of data and achieving efficient and accurate target detection. Compared with the prior art, in the present invention, by fusing the rich texture information of video data and the accurate ranging ability of radar data, significant performance advantages are shown under conditions such as bad weather, low illumination, or target occlusion, reducing the dependence on environmental conditions and improving the accuracy of target detection; also, through efficient data processing and advanced feature extraction techniques, the consumption of computing resources is reduced, and a fast response ability is maintained, enhancing the real-time processing ability.

[0053] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0055] Figure 1 is a schematic flowchart of a radar and video fusion-based target detection method provided by an embodiment of the present application;

[0056] Figure 2 is a schematic structural diagram of an improved YOLO model provided by an embodiment of the present application;

[0057] Figure 3 is a schematic structural diagram of a radar and video fusion module provided by an embodiment of the present application;

[0058] Figure 4 is a schematic structural diagram of a Transformer Encoder module provided by an embodiment of the present application;

[0059] Figure 5 It is a schematic structural diagram of the MLP module provided by an embodiment of the present application;

[0060] Figure 6 It is a schematic structural diagram of the feature fusion module provided by an embodiment of the present application;

[0061] Figure 7 It is a schematic structural diagram of the target detection device based on radar and video fusion provided by an embodiment of the present application;

[0062] Figure 8 It is a schematic structural diagram of the terminal device provided by an embodiment of the present application. Detailed implementation manners

[0063] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented in order to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0064] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0065] It should also be understood that the term " / and / " as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0066] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.

[0067] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0068] Reference to "one embodiment" or "some embodiments" described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0069] An embodiment of this application proposes a target detection method based on radar and video fusion. Figure 1 It is a schematic flowchart of the target detection method based on radar and video fusion provided by an embodiment of this application. Referring to Figure 1 , the details of this method are as follows:

[0070] Step 101, obtain the radar point cloud data and video data of the target to be detected.

[0071] Step 102, based on the processing results of the radar point cloud data and video data by the radar and video fusion module in the improved YOLO model, obtain the target detection result; wherein, the radar and video fusion module is used to preprocess the radar point cloud data and video data to obtain a radar image and a video image, and perform feature extraction, feature splicing, feature fusion, residual processing, secondary splicing, and convolution processing on the radar image and the video image.

[0072] Exemplarily, as Figure 2 shown, the improved YOLO model includes: a Backbone layer, a Neck layer, and a Head layer; the feature fusion module includes: a first feature fusion module GI1, a second feature fusion module GI2, a third feature fusion module GI3, a fourth feature fusion module GI4, a fifth feature fusion module GI5, a sixth feature fusion module GI6, a seventh feature fusion module GI7, and an eighth feature fusion module GI8.

[0073] The Backbone layer includes: a radar and video fusion module TF, a first CBS module CBS1, a first feature fusion module GI1, a second CBS module CBS2, a second feature fusion module GI2, a third feature fusion module GI3, a fourth feature fusion module GI4, and an SPPF module connected in sequence.

[0074] The Neck layer includes: a PSA sampling attention module, an Upsample module, a first Concat module Concat1, a fifth feature fusion module GI5, a second Concat module Concat2, a sixth feature fusion module GI6, a third CBS module CBS3, a third Concat module Concat3, a seventh feature fusion module GI7, a fourth CBS module CBS4, a fourth Concat module Concat4, and an eighth feature fusion module GI8, which are connected in sequence.

[0075] The Head layer includes: a first Head module Head1, a second Head module Head2, and a third Head module Head3.

[0076] The second CBS module CBS2 is also connected to the second Concat module Concat2; the third feature fusion module GI3 is also connected to the first Concat module Concat1; the SPPF module is also connected to the PSA sampling attention module; the PSA sampling attention module is also connected to the fourth Concat module Concat4; the fifth feature fusion module GI5 is also connected to the third Concat module Concat3; the sixth feature fusion module GI6 is also connected to the first Head module Head1; the seventh feature fusion module GI7 is also connected to the second Head module Head2; the eighth feature fusion module GI8 is also connected to the third Head module Head3.

[0077] Specifically, in the Backbone layer: The radar and video fusion module TF: This is the initial processing stage of the framework, responsible for fusing the input millimeter-wave radar data and video data to provide a unified data format for subsequent feature extraction. The CBS module (CBS1, CBS2, CBS3, and CBS4): This module consists of multiple convolutional layers, batch normalization layers, and activation function layers, and is used to extract and enhance image features. The feature fusion module (GI1, GI2, GI3, GI4, GI5, GI6, GI7, and GI8): Responsible for further processing and fusing features to enhance the network's adaptability and robustness to different types of image features. The SPPF module: A spatial pyramid pooling fusion module used to integrate features of different scales and enhance the model's ability to recognize targets of different sizes.

[0078] In the Neck layer: Feature fusion modules (GI1, GI2, GI3, GI4, GI5, GI6, GI7, and GI8): Appearing again in the Neck part, indicating their importance in the feature extraction process, used to further refine and fuse features from the Backbone. Concat module: Responsible for merging feature maps from different sources or different processing stages, increasing the dimension and information content of features, and providing support for more accurate object detection. Upsample module: Used to increase the resolution of the feature map, which is particularly important in object detection as it helps the model to more clearly identify small objects or detailed parts. PSA sampling attention module PSA: A sampling attention module that may be used to improve the model's attention to the target area, thereby enhancing the detection accuracy.

[0079] In the Head layer: The Head modules (Head1, Head2, and Head3) are responsible for the final detection tasks, such as bounding box regression and class classification. It may contain several fully connected layers or convolutional layers, as well as activation functions and loss functions for outputting the final detection results.

[0080] Through the collaborative work of these modules, the entire improved YOLO framework realizes the complete process from data fusion to feature extraction and then to final object detection. Each module is optimized for specific tasks, enabling the entire system to effectively perform object detection in complex environments.

[0081] Exemplarily, for radar images and video images, feature extraction, feature stitching, feature fusion, residual processing, secondary stitching, and convolutional processing are performed, including:

[0082] Perform a 3x3 convolution operation on the radar image to extract the first radar image feature; perform a 3x3 convolution operation on the video image to extract the first video image feature.

[0083] Stitch the first radar image feature and the first video image feature through a concat operation to obtain the first stitched feature.

[0084] Perform feature extraction and fusion on the first stitched feature through the transformer Encoder layer to obtain the first pre-output feature.

[0085] Perform residual processing and a 3x3 convolution operation on the radar image and the video image to obtain the second pre-output feature.

[0086] Perform a concat channel stitch on the first pre-output feature and the second pre-output feature to obtain the third pre-output feature.

[0087] Perform a 1x1 convolution operation on the third pre-output feature to obtain the output result of the radar and video fusion module.

[0088] Exemplarily, the radar and video fusion module is used to preprocess radar point cloud data and video data to obtain radar images and video images, including:

[0089] The radar and video fusion module converts the radar point cloud data into an image to obtain a radar image.

[0090] The video data is intercepted frame by frame through the OpenCV method to obtain video images.

[0091] Specifically, as Figure 3 shown, the TF module first converts the point cloud data of the millimeter-wave radar into a two-dimensional image, and at the same time uses OpenCV technology to capture video images frame by frame. Then, these two types of images are both processed through a 3x3 convolution kernel to extract image features. Subsequently, these features are merged in the channel dimension to generate a fused feature map.

[0092] The fused feature map is further processed through the Transformer Encoder layer. As Figure 4 shown, this layer consists of multiple key components: for example, the Layer Norm layer is responsible for normalizing the features to maintain the stability of network training; the Multi-Head Attention layer enables the model to simultaneously focus on multiple regions of the image; the Dropout layer randomly discards some features to avoid overfitting; as Figure 5 shown, the MLP Block layer contains a Linear layer, a GELU activation function, and a Dropout layer, which together enhance the model's ability to handle nonlinear problems. In addition, the Transformer Encoder layer adopts a residual structure, and by adding the input to the output of the MLP Block, it ensures the effective transmission of information in the deep network.

[0093] After being processed by the Encoder layer, the original radar and video images will go through a 3x3 convolution operation and residual processing again. This step further extracts and fuses the image information. Finally, the output of the Encoder layer is concatenated with the features after residual processing, and feature extraction and channel number adjustment are performed through a 1x1 convolution kernel to ensure that the output features can be smoothly connected to the subsequent layers of the network.

[0094] The following is a further elaboration of the detailed working process of the relevant part of the TF module:

[0095] 1. Image preprocessing:

[0096] The original point cloud data of the millimeter-wave radar is first converted into an image format so that the radar data can be processed and analyzed in the form of an image.

[0097] Video data is intercepted frame by frame through OpenCV, converting the continuous video stream into a series of individual image frames.

[0098] 2. Preliminary Convolution Operation:

[0099] Apply a 3x3 convolutional kernel to the radar image and the video image respectively for convolution operation. This step can be expressed as:

[0100]

[0101]

[0102] Here represents the 3x3 convolution operation for extracting image features.

[0103] 3. Feature Concatenation:

[0104] Concatenate the feature maps of the radar image and the video image after convolution operation and along the channel dimension to obtain the fused feature map :

[0105]

[0106] 4. Transformer Encoder Layer:

[0107] The fused feature map is fed into the Transformer Encoder layer for further feature extraction and fusion. This layer consists of multiple components:

[0108] Layer Norm: Normalize the feature map. The formula is:

[0109]

[0110] where is the input feature, are the learnable parameters.

[0111] Multi-Head Attention: Through the multi-head attention mechanism, the model can focus on different parts of the image. The formula is

[0112]

[0113] where represent the query, key, and value matrices respectively.

[0114] Dropout: Randomly discard a part of the features to prevent overfitting, usually with a certain probability p during the training process.

[0115] MLP Block: It contains two linear transformation layers, with a GELU activation function and a Dropout layer in the middle. The formula is:

[0116]

[0117] Residual connection: Add the output of the Encoder layer to the input. The formula is:

[0118]

[0119] 5. Residual processing and further convolutional operations:

[0120] Perform residual processing on the original radar images and video images, and apply the 3x3 convolutional operation again to further extract and fuse information:

[0121]

[0122]

[0123] 6. Final feature fusion:

[0124] The features processed by the Transformer Encoder layer are concatenated with the features of the radar and video images after residual processing and :

[0125]

[0126] Finally, use the 1x1 convolutional operation to perform feature extraction and channel number adjustment on the concatenated features to meet the requirements of the subsequent network layers:

[0127]

[0128] Here represents the 1x1 convolutional kernel, which is used to extract the final features and adjust the channel number.

[0129] With this design, the TF module not only realizes the effective fusion of radar and video images, but also significantly improves the accuracy and efficiency of object detection through in-depth feature extraction and fusion of the Transformer Encoder layer.

[0130] Exemplarily, the improved YOLO model also includes a feature fusion module;

[0131] Based on the processing results of radar point cloud data and video data by the radar and video fusion module in the improved YOLO model, the target detection results are obtained, including:

[0132] The processing results of the radar point cloud data and video data by the radar and video fusion module in the improved YOLO model are processed by at least one feature fusion module to obtain the target detection results; wherein, the feature fusion module is used to perform residual processing, convolution, depthwise separable convolution processing, concat fusion, and add fusion operation processing on the input image.

[0133] Exemplarily, performing residual processing, convolution, depthwise separable convolution processing, concat fusion, and add fusion operation processing on the input image includes:

[0134] Performing residual processing on the input image to obtain the first feature.

[0135] Performing a 1xk convolution operation on the input image to obtain the second feature.

[0136] Performing a kx1 convolution operation on the input image to obtain the third feature.

[0137] Performing Ghost convolution and 3x3 depthwise separable convolution on the input image to obtain the fourth feature.

[0138] Performing concat fusion on the second feature, third feature, and fourth feature to obtain the first fusion feature.

[0139] Performing an add operation on the first feature and the first fusion feature to obtain the fourth pre-output feature.

[0140] Performing Ghost convolution and 1x1 convolution on the fourth pre-output feature in sequence to obtain the output result of the feature fusion module.

[0141] Exemplarily, performing residual processing, convolution, depthwise separable convolution processing, concat fusion, and add fusion operation processing on the input image includes:

[0142] Performing residual processing on the input image to obtain the first feature.

[0143] Performing a 1xk convolution operation on the input image to obtain the second feature

[0144] Performing a kx1 convolution operation on the input image to obtain the third feature;.

[0145] Performing Ghost convolution and 3x3 depthwise separable convolution on the input image to obtain the fourth feature.

[0146] Performing concat fusion on the second feature, third feature, and fourth feature to obtain the first fusion feature.

[0147] Repeat the above steps multiple times to obtain multiple first fusion features.

[0148] Perform an add operation on the first feature and the multiple first fusion features to obtain a fourth pre-output feature.

[0149] Successively perform Ghost convolution and 1x1 convolution on the fourth pre-output feature to obtain the output result of the feature fusion module.

[0150] The design of the GI module aims to significantly improve the accuracy and efficiency of object detection through highly customized convolution operations and feature fusion strategies. Specifically, as Figure 6 shown, the GI module first dispatches the input image to four independent processing branches. In the first branch, the image undergoes residual processing, which enhances the flow of information in the deep network by adding a residual connection between the input and the output of the convolutional layer, helping to alleviate the vanishing gradient problem and enabling the network to learn more complex features. The second and third branches respectively apply 1xk and kx1 convolutional operations, which focus on extracting features of the image in specific directions while effectively controlling the computational cost by using a smaller convolutional kernel size. The fourth branch employs Ghost convolution, an innovative convolutional method that can increase the network's parameters and capacity without significantly increasing the computational burden, and then further refines feature extraction through 3x3 depthwise separable convolution (DW convolution).

[0151] In the feature fusion stage, the outputs of branches two and three are first concatenated (concat) in the channel dimension to integrate feature information from different directions. Then, this fused feature is concatenated again with the feature processed by Ghost convolution and DW convolution in branch four to achieve deep feature fusion. To further enhance the expressiveness of the features, the entire process of feature extraction and fusion can be repeated n times as needed, where n is a configurable parameter, enabling the network to adjust its complexity and performance according to the specific task. After multiple iterations, the residual feature of branch one is added (add) to the fused feature to achieve deeper feature fusion. Finally, further feature extraction and parameter reduction operations are performed through Ghost convolution, which uses 1x1 convolution to adjust the number of channels to ensure the compatibility of the output features with subsequent network layers.

[0152] The detailed working mechanism of the GI module is as follows:

[0153] 1. Input image processing:

[0154] Input image is fed into the GI module for feature extraction preparation.

[0155] 2. Branch 1 (Residual Processing):

[0156] Image is processed through a residual block RR, which contains convolutional operations and possibly activation functions:

[0157]

[0158] Residual connection helps the flow of information in the network and reduces the vanishing gradient problem.

[0159] 3. Branches 2 and 3 (Directional Convolutions):

[0160] Branch 2 applies a 1xk convolution to the image :

[0161]

[0162] Branch 3 applies a kx1 convolution to the image :

[0163]

[0164] where and represent the 1xk and kx1 convolutional kernels, respectively, which extract the features of the image in the horizontal and vertical directions.

[0165] 4. Branch 4 (Ghost Convolution and DW Convolution):

[0166] The Ghost convolution G first processes the image :

[0167] Then, a 3x3 depthwise separable convolution DW is applied:

[0168] The Ghost convolution enhances the feature extraction ability by increasing the network's parameters and capacity without significantly increasing the computational burden.

[0169] 5. Feature Fusion:

[0170] The outputs of Branches 2 and 3 are first concatenated in the channel dimension:

[0171]

[0172] This fused feature is then concatenated again with the feature of Branch 4:

[0173]

[0174] 6. Repeated Operations:

[0175] The entire feature extraction and fusion process can be repeated n times, where n is a configurable parameter to further enhance the expressiveness of the features.

[0176] 7. Residual Feature Fusion:

[0177] Perform an addition operation on the residual features of Branch 1 and the fused features:

[0178]

[0179] 8. Further Feature Extraction and Parameter Reduction:

[0180] Further process using Ghost convolution :

[0181]

[0182] Finally, adjust the number of channels through 1x1 convolution:

[0183]

[0184] Among them, represents a 1x1 convolution kernel, which is used to extract the final features and adjust the number of channels to match the requirements of subsequent network layers.

[0185] The design of the GI module improves the performance of the network in object detection tasks through highly customized convolution operations and feature fusion strategies. Through flexible configuration and efficient feature extraction mechanisms, the GI module enhances the adaptability and robustness of the network to different types of image features, providing strong technical support for object detection in complex environments. The design of this module enables the network to effectively perform object detection under various conditions, such as different lighting and weather conditions, thereby improving the accuracy and reliability of detection.

[0186] The above object detection method based on radar and video fusion processes radar point cloud data and video data through a radar and video fusion module, achieving deep fusion of the two types of data, and realizing efficient and accurate object detection through an improved TF - YOLO architecture. Compared with the existing technology, because the present invention fuses the rich texture information of video data and the accurate ranging ability of radar data, it shows significant performance advantages under conditions such as bad weather, low illumination, or target occlusion, reduces the dependence on environmental conditions, and improves the accuracy of object detection; it also reduces the consumption of computing resources through efficient data processing and advanced feature extraction techniques, and maintains a fast response ability, enhancing the real - time processing ability.

[0187] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0188] Corresponding to the target detection method based on radar and video fusion described in the above embodiments, Figure 5 The block diagram of the target detection device based on radar and video fusion provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0189] See Figure 7 , the target detection device based on radar and video fusion in the embodiments of the present application may include:

[0190] The data acquisition module 201 is configured to acquire the radar point cloud data and video data of the target to be detected.

[0191] The result output module 202 is configured to obtain the target detection result based on the processing results of the radar point cloud data and video data by the radar and video fusion module in the improved YOLO model; wherein, the radar and video fusion module is configured to preprocess the radar point cloud data and video data to obtain a radar image and a video image, and perform feature extraction, feature splicing, feature fusion, residual processing, secondary splicing, and convolution processing on the radar image and the video image.

[0192] Exemplarily, performing feature extraction, feature splicing, feature fusion, residual processing, secondary splicing, and convolution processing on the radar image and the video image includes:

[0193] Performing a 3x3 convolution operation on the radar image to extract the first radar image feature; performing a 3x3 convolution operation on the video image to extract the first video image feature.

[0194] Performing feature splicing on the first radar image feature and the first video image feature through the concat operation to obtain the first spliced feature.

[0195] Performing feature extraction and fusion on the first spliced feature through the transformer Encoder layer to obtain the first pre-output feature.

[0196] Performing residual processing and a 3x3 convolution operation on the radar image and the video image to obtain the second pre-output feature.

[0197] Performing concat channel splicing on the first pre-output feature and the second pre-output feature to obtain the third pre-output feature.

[0198] Performing a 1x1 convolution operation on the third pre-output feature to obtain the output result of the radar and video fusion module.

[0199] Exemplarily, the improved YOLO model further includes a feature fusion module;

[0200] The result output module 202 can also be used for:

[0201] Processing the processing results of the radar point cloud data and video data by the radar and video fusion module in the improved YOLO model through at least one feature fusion module to obtain a target detection result; wherein, the feature fusion module is used to perform residual processing, convolution, depthwise separable convolution processing, concat fusion, and add fusion operation processing on the input image.

[0202] Exemplarily, performing residual processing, convolution, depthwise separable convolution processing, concat fusion, and add fusion operation processing on the input image includes:

[0203] Performing residual processing on the input image to obtain a first feature.

[0204] Performing a 1xk convolution operation on the input image to obtain a second feature.

[0205] Performing a kx1 convolution operation on the input image to obtain a third feature.

[0206] Performing Ghost convolution and 3x3 depthwise separable convolution on the input image to obtain a fourth feature.

[0207] Performing concat fusion on the second feature, the third feature, and the fourth feature to obtain a first fusion feature.

[0208] Performing an add operation on the first feature and the first fusion feature to obtain a fourth pre-output feature.

[0209] Performing Ghost convolution and 1x1 convolution on the fourth pre-output feature in sequence to obtain the output result of the feature fusion module.

[0210] Exemplarily, performing residual processing, convolution, depthwise separable convolution processing, concat fusion, and add fusion operation processing on the input image includes:

[0211] Performing residual processing on the input image to obtain a first feature.

[0212] Performing a 1xk convolution operation on the input image to obtain a second feature

[0213] Performing a kx1 convolution operation on the input image to obtain a third feature;

[0214] Performing Ghost convolution and 3x3 depthwise separable convolution on the input image to obtain a fourth feature.

[0215] Concatenate and fuse the second, third, and fourth features to obtain the first fused feature.

[0216] Repeat the above steps multiple times to obtain multiple first fused features.

[0217] Perform an add operation on the first feature and multiple first fused features to obtain the fourth pre-output feature.

[0218] Successively perform Ghost convolution and 1x1 convolution on the fourth pre-output feature to obtain the output result of the feature fusion module.

[0219] Exemplarily, the improved YOLO model includes: a Backbone layer, a Neck layer, and a Head layer; the feature fusion module includes: a first feature fusion module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, a fifth feature fusion module, a sixth feature fusion module, a seventh feature fusion module, and an eighth feature fusion module.

[0220] The Backbone layer includes: a radar and video fusion module, a first CBS module, a first feature fusion module, a second CBS module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, and an SPPF module connected in sequence.

[0221] The Neck layer includes: a PSA sampling attention module, an Upsample module, a first Concat module, a fifth feature fusion module, a second Concat module, a sixth feature fusion module, a third CBS module, a third Concat module, a seventh feature fusion module, a fourth CBS module, a fourth Concat module, and an eighth feature fusion module connected in sequence.

[0222] The Head layer includes: a first Head module, a second Head module, and a third Head module.

[0223] The second CBS module is also connected to the second Concat module; the third feature fusion module is also connected to the first Concat module; the SPPF module is also connected to the PSA sampling attention module; the PSA sampling attention module is also connected to the fourth Concat module; the fifth feature fusion module is also connected to the third Concat module; the sixth feature fusion module is also connected to the first Head module; the seventh feature fusion module is also connected to the second Head module; the eighth feature fusion module is also connected to the third Head module.

[0224] Exemplarily, the radar and video fusion module is used to preprocess radar point cloud data and video data to obtain radar images and video images, including:

[0225] The radar and video fusion module converts radar point cloud data into an image to obtain a radar image.

[0226] The video data is intercepted frame by frame through the OpenCV method to obtain video images.

[0227] It should be noted that for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of this application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.

[0228] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0229] The embodiment of the present application also provides a terminal device. Refer to Figure 8 , the terminal device 300 may include: at least one processor 310, a memory 320. The memory 320 is used to store a computer program 321, and the processor 310 is used to call and run the computer program 321 stored in the memory 320 to implement the steps in any of the above method embodiments, such as Figure 1 the steps 101 to 102 in the embodiment shown. Or, when the processor 310 executes the computer program, it implements the functions of each module / unit in the above device embodiments, such as Figure 7 the functions of each module shown.

[0230] Exemplarily, the computer program 321 can be divided into one or more modules / units. One or more modules / units are stored in the memory 320 and executed by the processor 310 to complete this application. The one or more modules / units can be a series of computer program segments capable of completing specific functions, and this program segment is used to describe the execution process of the computer program in the terminal device 300.

[0231] Those skilled in the art can understand that Figure 8These are merely examples of terminal devices, which do not constitute a limitation on terminal devices. They may include more or fewer components than shown in the figures, or combine certain components, or have different components, such as input / output devices, network access devices, buses, etc.

[0232] The processor 310 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0233] The memory 320 may be an internal storage unit of the terminal device or an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 320 is used to store the computer program and other programs and data required by the terminal device. The memory 320 may also be used to temporarily store data that has been output or is to be output.

[0234] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0235] The target detection method based on radar and video fusion provided by the embodiments of this application can be applied to terminal devices such as computers, wearable devices, vehicle-mounted devices, tablet computers, laptop computers, and netbooks. The embodiments of this application do not impose any restrictions on the specific types of terminal devices.

[0236] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in each of the above-described embodiments of the object detection method based on radar and video fusion can be implemented.

[0237] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal is caused to execute the steps in each of the above-described embodiments of the object detection method based on radar and video fusion.

[0238] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps in each of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.

[0239] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0240] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0241] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0242] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0243] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A target detection method based on radar and video fusion, characterized in that: include: Obtain radar point cloud data and video data of the target to be detected; Based on the processing result of the radar point cloud data and the video data by the radar and video fusion module in the improved YOLO model, a target detection result is obtained; wherein the radar and video fusion module is used to pre-process the radar point cloud data and the video data to obtain a radar image and a video image, and perform feature extraction, feature splicing, feature fusion, residual processing, secondary splicing and convolution processing on the radar image and the video image; The performing feature extraction, feature splicing, feature fusion, residual processing, secondary splicing and convolution processing on the radar image and the video image includes: Performing a 3x3 convolution operation on the radar image to extract a first radar image feature; performing a 3x3 convolution operation on the video image to extract a first video image feature; Perform feature splicing on the first radar image feature and the first video image feature through a concat operation to obtain a first splicing feature; Extracting and fusing the first concatenated features through a transformer encoder layer to obtain first pre-output features; Performing residual processing and a 3x3 convolution operation on the radar image and the video image to obtain a second pre-output feature; Perform concat channel concatenation on the first pre-output feature and the second pre-output feature to obtain a third pre-output feature; Performing a 1x1 convolution operation on the third pre-output feature to obtain an output result of the radar and video fusion module; The improved YOLO model also includes a feature fusion module; The target detection result is obtained by processing the radar point cloud data and the video data by the radar and video fusion module in the improved YOLO model, including: The processing result of the radar point cloud data and the video data by the radar and video fusion module in the improved YOLO model is processed by at least one of the feature fusion modules to obtain a target detection result; wherein the feature fusion module is used to perform residual processing, convolution, depth-separable convolution processing, concat fusion and add fusion operation processing on the input image; The improved YOLO model includes: a Backbone layer, a Neck layer and a Head layer; the feature fusion module includes: a first feature fusion module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, a fifth feature fusion module, a sixth feature fusion module, a seventh feature fusion module and an eighth feature fusion module; The Backbone layer includes: a radar and video fusion module, a first CBS module, the first feature fusion module, a second CBS module, the second feature fusion module, the third feature fusion module, the fourth feature fusion module and an SPPF module connected in sequence; The Neck layer includes: a PSA sampling attention module, an Upsample module, a first Concat module, the fifth feature fusion module, a second Concat module, the sixth feature fusion module, a third CBS module, a third Concat module, the seventh feature fusion module, a fourth CBS module, a fourth Concat module and the eighth feature fusion module connected in sequence; The Head layer includes: a first Head module, a second Head module and a third Head module; The second CBS module is also connected to the second Concat module; the third feature fusion module is also connected to the first Concat module; the SPPF module is also connected to the PSA sampling attention module; the PSA sampling attention module is also connected to the fourth Concat module; the fifth feature fusion module is also connected to the third Concat module; the sixth feature fusion module is also connected to the first Head module; the seventh feature fusion module is also connected to the second Head module; the eighth feature fusion module is also connected to the third Head module.

2. The target detection method based on radar and video fusion as claimed in claim 1, characterized in that: The performing residual processing, convolution, depth-separable convolution processing, concat fusion and add fusion operation processing on the input image includes: Perform residual processing on the input image to obtain the first feature; Perform 1xk convolution operation on the input image to obtain the second feature; Perform a kx1 convolution operation on the input image to obtain the third feature; Perform Ghost convolution and 3x3 depth-separable convolution on the input image to obtain the fourth feature; Concat the second feature, the third feature and the fourth feature to obtain a first fused feature; Performing an add operation on the first feature and the first fusion feature to obtain a fourth pre-output feature; Ghost convolution and 1x1 convolution are performed on the fourth pre-output feature in sequence to obtain the output result of the feature fusion module.

3. The target detection method based on radar and video fusion as claimed in claim 1, characterized in that: The performing residual processing, convolution, depth-separable convolution processing, concat fusion and add fusion operation processing on the input image includes: Perform residual processing on the input image to obtain the first feature; Perform 1xk convolution operation on the input image to obtain the second feature; Perform a kx1 convolution operation on the input image to obtain the third feature; Perform Ghost convolution and 3x3 depth-separable convolution on the input image to obtain the fourth feature; Concat the second feature, the third feature and the fourth feature to obtain a first fused feature; Repeat the above steps multiple times to obtain multiple first fusion features; Performing an add operation on the first feature and the multiple first fusion features to obtain a fourth pre-output feature; Ghost convolution and 1x1 convolution are performed on the fourth pre-output feature in sequence to obtain the output result of the feature fusion module.

4. The target detection method based on radar and video fusion as claimed in claim 1, characterized in that: The radar and video fusion module is used to pre-process the radar point cloud data and the video data to obtain a radar image and a video image, including: The radar and video fusion module converts the radar point cloud data into an image to obtain the radar image; The video data is intercepted frame by frame by using the OpenCV method to obtain the video image.

5. A target detection device based on radar and video fusion, characterized in that: include: A data acquisition module is used to acquire radar point cloud data and video data of the target to be detected; A result output module is used to obtain a target detection result based on the processing result of the radar point cloud data and the video data by the radar and video fusion module in the improved YOLO model; wherein the radar and video fusion module is used to pre-process the radar point cloud data and the video data to obtain a radar image and a video image, and perform feature extraction, feature splicing, feature fusion, residual processing, secondary splicing and convolution processing on the radar image and the video image; The performing feature extraction, feature splicing, feature fusion, residual processing, secondary splicing and convolution processing on the radar image and the video image includes: Performing a 3x3 convolution operation on the radar image to extract a first radar image feature; performing a 3x3 convolution operation on the video image to extract a first video image feature; Perform feature splicing on the first radar image feature and the first video image feature through a concat operation to obtain a first splicing feature; Extracting and fusing the first concatenated features through a transformer encoder layer to obtain first pre-output features; Performing residual processing and a 3x3 convolution operation on the radar image and the video image to obtain a second pre-output feature; Perform concat channel concatenation on the first pre-output feature and the second pre-output feature to obtain a third pre-output feature; Performing a 1x1 convolution operation on the third pre-output feature to obtain an output result of the radar and video fusion module; The improved YOLO model also includes a feature fusion module; The result output module is also used for: The processing result of the radar point cloud data and the video data by the radar and video fusion module in the improved YOLO model is processed by at least one of the feature fusion modules to obtain a target detection result; wherein the feature fusion module is used to perform residual processing, convolution, depth-separable convolution processing, concat fusion and add fusion operation processing on the input image; The improved YOLO model includes: a Backbone layer, a Neck layer and a Head layer; the feature fusion module includes: a first feature fusion module, a second feature fusion module, a third feature fusion module, a fourth feature fusion module, a fifth feature fusion module, a sixth feature fusion module, a seventh feature fusion module and an eighth feature fusion module; The Backbone layer includes: a radar and video fusion module, a first CBS module, the first feature fusion module, a second CBS module, the second feature fusion module, the third feature fusion module, the fourth feature fusion module and an SPPF module connected in sequence; The Neck layer includes: a PSA sampling attention module, an Upsample module, a first Concat module, the fifth feature fusion module, a second Concat module, the sixth feature fusion module, a third CBS module, a third Concat module, the seventh feature fusion module, a fourth CBS module, a fourth Concat module and the eighth feature fusion module connected in sequence; The Head layer includes: a first Head module, a second Head module and a third Head module; The second CBS module is also connected to the second Concat module; the third feature fusion module is also connected to the first Concat module; the SPPF module is also connected to the PSA sampling attention module; the PSA sampling attention module is also connected to the fourth Concat module; the fifth feature fusion module is also connected to the third Concat module; the sixth feature fusion module is also connected to the first Head module; the seventh feature fusion module is also connected to the second Head module; the eighth feature fusion module is also connected to the third Head module.

6. A terminal device, comprising: A processor and a memory, wherein the memory stores a computer program that can be run on the processor, wherein when the processor executes the computer program, the target detection method based on radar and video fusion as described in any one of claims 1 to 4 is implemented.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the target detection method based on radar and video fusion as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Foreign matter intrusion detection method and device based on convolutional neural network, and storage medium

    CN117237613A

  • Dual-modal target detection method and system based on cross-modal attention mechanism fusion

    CN117422971A