Multi-sensor target detection method and device, electronic equipment and storage medium

By fusing the confidence scores and spatial locations of multiple sensors, the problem of fusion between sensor types has been solved, improving the target detection accuracy in autonomous driving, especially the fusion effect of millimeter-wave radar, lidar, and image sensors.

CN114677655BActive Publication Date: 2025-11-28SIWAVE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210136782.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-11-28
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

Different types of sensors cannot be effectively fused when detecting targets in autonomous driving, resulting in reduced detection accuracy and efficiency.

Method used

By determining the target detection results collected by at least two types of sensors, including confidence level and spatial location, the confidence level and spatial location are fused to generate a fused confidence level and spatial location, which is then used to perform autonomous driving operations.

Benefits of technology

It improves the accuracy of target detection, overcomes the limitations of single-type sensors, realizes the complementary fusion of different types of sensors, and enhances detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677655B_ABST
    Figure CN114677655B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a multi-sensor target detection method, device, electronic equipment and storage medium. The method comprises the following steps: determining first target detection results collected by at least two types of sensors respectively; the target detection result comprises a confidence degree and a spatial position, and the spatial position is represented by a target center point coordinate and a size; confidence degrees are fused and spatial positions are fused according to the first target detection results collected by the sensors respectively to obtain fused confidence degrees and fused spatial positions; and the second target detection result is determined according to the fused confidence degrees and the fused spatial positions, so as to execute an automatic driving operation. The influence of the limitation of a single type of sensor on target detection accuracy is solved, different types of sensors are complementarily fused, and the target detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of target detection of automatic driving technology, and particularly relates to a multi-sensor target detection method and device, electronic equipment and a storage medium. BACKGROUND

[0002] Target detection algorithm is one of the key research directions of computer vision, and the target detection algorithm can promote the exploration of environmental perception technology and promote the development of automatic driving technology.

[0003] In recent years, with the continuous development of target detection, many researches on 2D and 3D object detection, semantic segmentation and object tracking have been inspired. Since each type of sensor has certain limitations, therefore, the automatic driving car will be equipped with different types of sensors to improve the accuracy of target detection by using their complementary characteristics. However, with the increase of the number of sensors, the fusion of different types of sensors such as radar and image increases the difficulty and accuracy of sensor data fusion, and reduces the accuracy and efficiency of target detection. SUMMARY

[0004] The present application provides a multi-sensor target detection method, device, electronic equipment and storage medium to solve the problem that different types of sensor target detection cannot be well fused.

[0005] According to an aspect of the present application, a multi-sensor target detection method is provided, comprising:

[0006] determining first target detection results respectively collected by at least two types of sensors; the target detection result comprises a confidence and a spatial position, and the spatial position is represented by a target center point coordinate and a size;

[0007] performing confidence fusion and spatial position fusion according to the first target detection results collected by each sensor to obtain a fused confidence and a fused spatial position;

[0008] determining a second target detection result according to the fused confidence and the fused spatial position, and using the second target detection result to perform an automatic driving operation.

[0009] According to another aspect of the present application, a multi-sensor target detection device is provided, comprising:

[0010] a collection module configured to determine first target detection results respectively collected by at least two types of sensors; the target detection result comprises a confidence and a target position feature, and the target position feature is represented by a target center point coordinate and a size;

[0011] A fusion module is configured to perform confidence fusion and spatial position fusion on the first target detection results collected by the sensors respectively, to obtain fused confidence and fused spatial position.

[0012] A detection module is configured to determine a second target detection result according to the fused confidence and the fused spatial position, to perform an automatic driving operation.

[0013] According to another aspect of the present application, an electronic device is provided, which comprises:

[0014] at least one processor; and

[0015] a memory connected to the at least one processor in communication; wherein,

[0016] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the multi-sensor target detection method according to any one of the embodiments of the present application.

[0017] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to perform the multi-sensor target detection method according to any one of the embodiments of the present application when executed.

[0018] The technical solution of the embodiments of the present application determines the first target detection results collected by at least two types of sensors respectively, the target detection result comprises confidence and spatial position, the spatial position is represented by target center point coordinates and size, confidence fusion and spatial position fusion are performed on the first target detection results collected by the sensors respectively, to obtain fused confidence and fused spatial position, the second target detection result is determined according to the fused confidence and the fused spatial position, to perform an automatic driving operation, the influence of the limitation of a single type of sensor on target detection accuracy is solved, complementary fusion of different types of sensors is realized, and target detection accuracy is improved.

[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to make the technical solutions in the embodiments of the present application clearer, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative effort based on these drawings.

[0021] Figure 1 is a flow chart of a multi-sensor target detection method according to an embodiment of the present application;

[0022] Figure 2 is a multi-sensor target detection architecture diagram according to an embodiment of the present application;

[0023] Figure 3 is a confidence fusion diagram in multi-sensor target detection according to an embodiment of the present application;

[0024] Figure 4 is a spatial position fusion diagram in multi-sensor target detection according to an embodiment of the present application;

[0025] Figure 5 is a whole architecture diagram of spatial position fusion in multi-sensor target detection according to an embodiment of the present application;

[0026] Figure 6 is a spatial position fusion diagram in multi-sensor target detection according to an embodiment of the present application;

[0027] Figure 7 is another spatial position fusion diagram in multi-sensor target detection according to an embodiment of the present application;

[0028] Figure 8 is a structural schematic diagram of a multi-sensor target detection device according to an embodiment of the present application;

[0029] Figure 9 is a structural schematic diagram of an electronic device implementing a multi-sensor target detection method according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the technical solutions in the embodiments of the present application clearer, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some of the embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative effort based on these drawings.

[0031] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present application and the above-described accompanying drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] The full-link interface test method, device, electronic equipment and storage medium provided in the application will be described in detail below through the following various embodiments and their optional schemes.

[0033] Figure 1 For the flowchart of the multi-sensor target detection method provided in the embodiments of the application, the embodiments can be applicable to the case of detecting targets in front of an autonomous vehicle. The method can be executed by a multi-sensor target detection device, which can be realized in the form of hardware and / or software, and can be configured in any electronic device with network communication function. As shown in the figure, the method can include the following steps: Figure 1

[0034] S110, determining a first target detection result collected by at least two types of sensors respectively; wherein the target detection result includes confidence and spatial position, and the spatial position is represented by target center point coordinates and size.

[0035] The multi-sensor target detection method of the present application scheme can be applied to electronic equipment, which can include but is not limited to terminal equipment and servers for detecting targets in front of an autonomous vehicle. The terminal equipment can include but is not limited to mobile phones, tablets, vehicle-mounted computers and other terminals.

[0036] In the process of autonomous driving, pre-configured sensors can be used to detect targets in the environment in front of the vehicle. The targets can include but are not limited to obstacles, other vehicles and pedestrians in front of the vehicle. The at least two types of sensors can include millimeter wave radar sensors, laser radar sensors and image sensors.

[0037] ​The target detection result of each type of sensor can include the spatial position of the detected target in front of the vehicle and the corresponding confidence. The confidence can be a value in the interval [0, 1], and the confidence can indicate the possibility of the existence of a target in front of the vehicle detected by the sensor. The greater the value of the confidence, the higher the confidence, and the greater the possibility of the existence of a target in front of the vehicle. The spatial position can be represented by the coordinates of the center point of the target and the size of the target. The center point of the target is represented by the coordinates of three points, i.e., the horizontal coordinate x, the vertical coordinate y, and the vertical coordinate z. The size of the target is represented by the length, width, and height w, l, and h.

[0038] S120, respectively according to the first target detection results collected by each sensor, confidence fusion and spatial position fusion are performed to obtain fused confidence and fused spatial position.

[0039] It is considered that each type of sensor has certain limitations. For example, the millimeter wave radar can provide accurate 3D measurement, but the point cloud generated thereby becomes sparse at a long distance, thereby reducing the ability to accurately detect a target at a long distance. The image provides rich appearance features, but is not a good source of depth prediction information. However, the laser radar has a relatively high requirement for weather, and the ability to accurately detect a target at a long distance will be reduced in some extreme weather.

[0040] Therefore, the autonomous vehicle can be equipped with different types of sensors to utilize the complementary characteristics of different types of sensors to complementally fuse the confidence and the spatial position in the target detection results collected by different types of sensors, and to realize target detection by using the complementally fused confidence and the fused spatial position, thereby reducing the great influence of the accuracy of target detection by using a single type of sensor.

[0041] In an optional solution of the embodiment, according to the first target detection results collected by each sensor, confidence fusion and spatial position fusion are performed to obtain fused confidence and fused spatial position, which can include the following steps A1-A3:

[0042] Step A1, the confidence and the spatial position in the first target detection results collected by different types of sensors are spliced to obtain a spliced matrix feature.

[0043] Referring to Figure 2For example, taking millimeter wave radar, laser radar and image sensor as examples, the first target detection results collected by different types of sensors are described as follows: the first target detection result of the millimeter wave radar is denoted as x1, y1, z1, w1, l1, h1, c1, wherein c is the confidence size; the first target detection result of the laser radar is denoted as x2, y2, z2, w2, l2, h2, c2; and the first target detection result of the image sensor is denoted as x3, y3, 0, w3, l3, 0, c3. Since the detection result obtained by the image sensor is two-dimensional, z and h are 0.

[0044] Referring to Figure 2 The first target detection results collected by each type of sensor can include 7 element values, and the feature splicing of the first target detection results collected by the three types of sensors can obtain a 3x7-dimensional matrix feature.

[0045] Optionally, the confidence and the spatial position in the first target detection results collected by different types of sensors are spliced to obtain a feature result, including: normalizing the confidence and the spatial position in the first target detection results collected by different types of sensors, and splicing the normalized confidence and the spatial position corresponding to each first target detection result to obtain a feature result.

[0046] Step A2, adjusting the confidence in each first target detection result according to the difference between each spatial position in the spliced matrix feature to obtain a fused confidence.

[0047] Referring to Figure 2 After obtaining the spliced matrix feature, for the confidence of any row, the confidence of the row can be complementarily fine-tuned by analyzing the difference between the spatial position of the row and the spatial positions of other rows to implement a feature extraction operation, so that the fused confidence C1, C2, C3 can be obtained after the feature extraction operation, and the dimension is 3x1. The feature extraction can be implemented by a series of convolution layers or a certain neural network model. For example, if the feature extraction is set to 1x7 convolution kernel operation, the 3x7-dimensional matrix feature can obtain the 3x1-dimensional fused feature after convolution operation.

[0048] In an optional example of the embodiment, adjusting the confidence in each first target detection result according to the difference between each spatial position in the spliced matrix feature to obtain a fused confidence can include the following steps:

[0049] input the spliced matrix feature into a preset confidence fusion extraction model to obtain fused confidence; the confidence fusion extraction model is used for analyzing and judging whether a same target can be detected by different types of sensors within a preset range at a same position, and adjusting the confidence in each first target detection result according to the analysis and judgment result; when a same target is detected by different types of sensors within a preset range at a same position, the confidence is increased after the confidence fusion is triggered.

[0050] Referring to Figure 3 Taking millimeter wave radar, laser radar and image sensor as an example, the 3x7 matrix feature is obtained after the preliminary detection results of the three sensors are spliced. The confidence fusion extraction model can be implemented by a series of convolution layers or a certain neural network model. For example, if the confidence fusion extraction model is set to 1x7 convolution kernel operation, the fused feature with a dimension of 3x1 can be obtained after the 3x7 dimensional feature is subjected to convolution operation. The 3x1 dimensional result is the confidence of the detection result of the three sensors respectively.

[0051] Referring to Figure 3 The confidence fusion extraction model is used for judging whether a same target is detected by different sensors within a preset range at a same position based on each spatial position, and adjusting the confidence in each first target detection result according to the judgment result. The confidence fusion can be that when a same target is detected by different types of sensors within a preset range at a same position, the confidence of the sensors should be higher, that is, it is more likely that the target exists. On the contrary, if only one sensor detects the target, the confidence will be reduced.

[0052] Step A3, adjusting the spatial position in each first target detection result according to the size of each confidence in the spliced matrix feature to obtain the fused spatial position.

[0053] Referring to Figure 2 and Figure 3 Taking millimeter wave radar, laser radar and image sensor as an example, the spatial transformation matrix T with a dimension of 3x2x3x4 can be obtained after the spliced feature is subjected to feature extraction operation. The spatial transformation matrix can be used to change the position of the detected target, so as to obtain new fused position coordinates. The dimension of the fused position information is still 3x6. For example, the millimeter wave radar position fusion result is recorded as X1, Y1, Z1, W1, L1, H1; the laser radar position fusion result is recorded as X2, Y2, Z2, W2, L2, H2; and the image sensor position fusion result is recorded as X3, Y3, Z3, W3, L3, H3. Here, the fused position information of the image sensor is no longer a two-dimensional result, and the spatial position after spatial position transformation can have three dimensions.

[0054] In an optional example of the embodiment, the spatial positions in each first target detection result are adjusted according to the confidence levels in the spliced matrix feature, to obtain fused spatial positions, which can include the following steps B1-B2:

[0055] Step B1, input the spliced matrix feature into a preset spatial position fusion extraction model to obtain a spatial transformation matrix; the spatial position fusion extraction model is used to analyze the spatial transformation matrix for changing and adjusting the spatial positions in the target detection result based on the confidence levels in the spliced matrix feature.

[0056] Step B2, convert the spatial positions in the first target detection result by the affine transformation matrix and the translation transformation matrix included in the spatial transformation matrix, to obtain the fused spatial positions.

[0057] Referring to Figure 4 In the spatial position fusion, the spatial position fusion extraction model can be implemented by a series of convolution operations or neural network models, and the spatial transformation matrix including the affine transformation matrix and the translation transformation matrix can be obtained through feature extraction (for example, the 3x3 affine transformation matrix A and the 3x1 translation transformation matrix S shown in FIG. 3). Figure 4 Since the spatial position information of each sensor includes not only 3 center point coordinates but also 3 size information, 2 spatial transformation matrices are obtained for each sensor, i.e., the dimension is 2x3x4. The dimension of the spatial transformation matrix T corresponding to the three sensors is 3x2x3x4. After the spatial transformation matrix, the position information in the original detection result is updated to the fused position information, and the dimension is 3x6.

[0058] Referring to Figure 5 FIG. 3 shows the overall architecture of the spatial position fusion, which includes the transformation processes of the spatial position information of the millimeter wave radar, the laser radar, and the image sensor. In the spatial position fusion process, the new spatial position coordinates X, Y, and Z can be obtained by matrix transformation of the original spatial position coordinates x, y, and z. As shown in the following formula, the spatial transformation matrix includes the affine transformation matrix A and the translation transformation matrix S. The affine transformation matrix A is a 3x3 matrix composed of a-i in the formula, and the translation transformation matrix S is a 3x1 matrix composed of j-l in the formula.

[0059]

[0060] Referring to Figure 6, the figure shows the process of spatial position fusion of millimeter wave radar or laser radar. Among them, the basic dimension of the spatial transformation matrix T is 3x4, the left three columns of the spatial transformation matrix T correspond to the 3x3 part of the affine transformation matrix S, and the right column corresponds to the 3x1 part of the translation transformation matrix S. In the process of spatial position fusion implementation, the original spatial position information can be first subjected to affine transformation matrix, and then the translation transformation is added to obtain new spatial position information. Of course, the original spatial position information can also be first subjected to translation transformation, and then the affine transformation is added to obtain new spatial position information.

[0061] Referring to Figure 7 , the figure shows the process of fusion of the spatial position of the image sensor. The fusion process is similar to the spatial position fusion process of the radar. The difference lies in that in the fusion implementation process, the original spatial position information must be first subjected to translation transformation, and then the affine transformation is added to obtain new spatial position information. If the affine transformation is performed first, then due to the lack of one dimension in the original position information of the image sensor, the parameters of the corresponding position in the affine matrix will be invalid.

[0062] S130, according to the fused confidence and the fused spatial position, determine a second target detection result, which is used to perform an automatic driving operation.

[0063] According to the fused confidence and the fused spatial position, the second target detection result can be determined, which can include the following steps: performing non-maximum suppression on the fused confidence and the fused spatial position to obtain the second target detection result.

[0064] After confidence fusion and spatial position fusion, the new confidence and spatial position are subjected to non-maximum suppression (NMS) to obtain the final three-dimensional target detection result, so that the final second target detection result can be used to indicate the automatic driving operation.

[0065] According to the technical scheme of the embodiment of the present application, the first target detection result collected by at least two types of sensors is determined, the target detection result includes a confidence degree and a spatial position, the spatial position is expressed by a target center point coordinate and a size, the confidence degree fusion and the spatial position fusion are respectively performed according to the first target detection result collected by each sensor, the fused confidence degree and the fused spatial position are obtained, and the second target detection result is determined according to the fused confidence degree and the fused spatial position, so as to perform the automatic driving operation. The influence of the limitation of a single type of sensor on the target detection precision is solved, the complementary fusion of different types of sensors is realized, the target detection precision is improved, especially the millimeter wave radar, the laser radar and the image detection result can be fused, and the target detection precision is improved. Meanwhile, the fusion scheme has scalability, is not limited to the fusion of millimeter wave radar, laser radar and image three types of sensor data, and is also applicable to the fusion process of two types of sensors or more than three types of sensors.

[0066] Figure 8 A structural block diagram of a multi-sensor target detection device provided in the embodiment of the present application is shown in the figure. The embodiment can be applied to the case of detecting the target in front of the automatic driving vehicle. The multi-sensor target detection device can be realized in the form of hardware and / or software, and can be configured in any electronic device with network communication function. As shown in the figure, the device can include an acquisition module 810, a fusion module 820 and a detection module 830. Among them, Figure 8

[0067] The acquisition module 810 is configured to determine the first target detection result collected by at least two types of sensors. The target detection result includes a confidence degree and a target position feature, and the target position feature is expressed by a target center point coordinate and a size.

[0068] The fusion module 820 is configured to perform confidence degree fusion and spatial position fusion according to the first target detection result collected by each sensor, and obtain the fused confidence degree and the fused spatial position.

[0069] The detection module 830 is configured to determine the second target detection result according to the fused confidence degree and the fused spatial position, and perform the automatic driving operation.

[0070] On the basis of the above-mentioned embodiment, the fusion module 820 can optionally include:

[0071] The feature splicing unit is configured to splice the confidence degree and the spatial position in the first target detection result collected by different types of sensors to obtain the spliced matrix feature.

[0072] ​The confidence fusion unit is configured to adjust the confidence in each first target detection result according to the difference between each spatial position in the spliced matrix feature, to obtain fused confidence.

[0073] The spatial position fusion unit is configured to adjust the spatial position in each first target detection result according to the confidence in each spatial position in the spliced matrix feature, to obtain fused spatial position.

[0074] Optionally, the confidence fusion unit comprises:

[0075] The spliced matrix feature is input into a preset confidence fusion extraction model to obtain fused confidence; the confidence fusion extraction model is configured to analyze and determine whether a same target can be detected by different types of sensors within a preset range at a same position, and adjust the confidence in each first target detection result according to the analysis and determination result; when the same target is detected by different types of sensors within the preset range at the same position, the confidence is increased after the confidence fusion is triggered.

[0076] Optionally, the spatial position fusion unit comprises:

[0077] The spliced matrix feature is input into a preset spatial position fusion extraction model to obtain a spatial transformation matrix; the spatial position fusion extraction model is configured to analyze and determine the spatial transformation matrix for changing and adjusting the spatial position in the target detection result based on the confidence in each spatial position in the spliced matrix feature.

[0078] The spatial position in the first target detection result is converted by an affine transformation matrix and a translation transformation matrix included in the spatial transformation matrix, to obtain fused spatial position.

[0079] Optionally, the detection module 830 comprises:

[0080] The fused confidence and the fused spatial position are subjected to non-maximum suppression to obtain the second target detection result.

[0081] Optionally, the confidence and the spatial position in the first target detection result collected by different types of sensors are spliced to obtain a feature result, comprising:

[0082] The confidence and the spatial position in the first target detection result collected by different types of sensors are normalized, and the normalized confidence and the spatial position corresponding to each first target detection result are spliced to obtain a feature result.

[0083] On the basis of the above-mentioned embodiments, optionally, the at least two types of sensors include a millimeter wave radar sensor, a laser radar sensor, and an image sensor.

[0084] The multi-sensor target detection device provided in the embodiments of the present application can execute the multi-sensor target detection method provided in any of the embodiments of the present application, has the corresponding functions and advantages of executing the multi-sensor target detection method, and the detailed process is referred to the related operations of the multi-sensor target detection method in the foregoing embodiments.

[0085] Figure 9 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0086] As shown in Figure 9 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0087] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, a loudspeaker, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0088] The processor 11 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the multi-sensor target detection method.

[0089] In some embodiments, the multi-sensor target detection method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the multi-sensor target detection method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the multi-sensor target detection method by any other suitable means, such as by means of firmware.

[0090] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0091] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a machine or a remote machine or a server.

[0092] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0093] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0094] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), blockchain network, and the Internet.

[0095] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0096] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0097] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A multi-sensor target detection method, characterized in that, include: Determine the detection results of the first target collected by at least two types of sensors; The target detection results include confidence level and spatial location, where the spatial location is represented by the coordinates and size of the target's center point; the confidence level indicates the probability that a target is detected ahead of the vehicle by the sensor. Based on the first target detection results collected by each sensor, confidence fusion and spatial location fusion are performed respectively to obtain the fused confidence and fused spatial location; The second target detection result is determined based on the fused confidence level and the fused spatial location, and then used to perform autonomous driving operations; The process involves fusing confidence scores and spatial locations based on the first target detection results collected by each sensor, resulting in fused confidence scores and fused spatial locations, including: The confidence level and spatial location feature results from the first target detection results collected by different types of sensors are concatenated to obtain the concatenated matrix feature; Based on the differences between various spatial locations in the spliced ​​matrix features, the confidence scores of each first target detection result are adjusted to obtain the fused confidence scores. Based on the confidence level of each feature in the spliced ​​matrix, the spatial position of each first target detection result is adjusted to obtain the fused spatial position.

2. The method according to claim 1, characterized in that, Based on the differences between spatial locations in the concatenated matrix features, the confidence scores of each first target detection result are adjusted to obtain the fused confidence scores, including: The concatenated matrix features are input into a preset confidence fusion extraction model to obtain the fused confidence score. The confidence fusion extraction model is used to analyze and determine whether the same target can be detected by different types of sensors within a preset range at the same location, and adjusts the confidence score in each first target detection result according to the analysis and judgment results. When the same target is detected by different types of sensors within a preset range at the same location, the confidence score will increase after the confidence fusion is triggered.

3. The method according to claim 1, characterized in that, Based on the confidence levels of each feature in the concatenated matrix, the spatial positions of each first target detection result are adjusted to obtain the fused spatial positions, including: The concatenated matrix features are input into a preset spatial location fusion extraction model to obtain a spatial transformation matrix. The spatial location fusion extraction model is used to parse the spatial transformation matrix that adjusts the spatial location of the target detection result based on the confidence level of each feature in the concatenated matrix. The spatial position in the first target detection result is transformed by the affine transformation matrix and the translation transformation matrix included in the spatial transformation matrix to obtain the fused spatial position.

4. The method according to claim 1, characterized in that, The second target detection result is determined based on the fused confidence level and the fused spatial location, including: The second target detection result is obtained by applying nonmaximum suppression to the fused confidence score and the fused spatial location.

5. The method according to claim 1, characterized in that, The confidence level and spatial location feature results from the first target detection results collected by different types of sensors are concatenated, including: The confidence and spatial location of the first target detection results collected by different types of sensors are normalized, and the normalized confidence and spatial location of each first target detection result are spliced ​​together as feature results.

6. The method according to any one of claims 1-5, characterized in that, The at least two types of sensors include millimeter-wave radar sensors, lidar sensors, and image sensors.

7. A multi-sensor target detection device, characterized in that, include: The acquisition module is used to determine the first target detection results acquired by at least two types of sensors; The target detection result includes confidence level and target position features, where the target position features are represented by the coordinates of the target's center point and its size; the confidence level is used to indicate the probability that a target is detected ahead of the vehicle by the sensor. The fusion module is used to perform confidence fusion and spatial location fusion based on the first target detection results collected by each sensor, so as to obtain the fused confidence and fused spatial location. The fusion module includes: The feature stitching unit is used to stitch together the confidence and spatial location feature results from the first target detection results collected by different types of sensors to obtain the stitched matrix features. The confidence fusion unit is used to adjust the confidence of each first target detection result according to the differences between the spatial locations in the features of the spliced ​​matrix, so as to obtain the fused confidence. The spatial location fusion unit is used to adjust the spatial location of each first target detection result according to the confidence level of each feature in the spliced ​​matrix to obtain the fused spatial location; The detection module is used to determine the detection result of the second target based on the fused confidence level and the fused spatial location, in order to perform autonomous driving operations.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multi-sensor target detection method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multi-sensor target detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Three-dimensional target detection method and device based on multi-sensor information fusion

    CN110929692A