Methods and systems for target detection
The method and system for target detection using radar and visible light cameras improve detection capabilities by adaptively fusing image features, addressing the limitations of single sensors in complex environments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-04-16
AI Technical Summary
Single image acquisition devices, such as visible light cameras, struggle to handle complex and changing situations in target detection, necessitating a more comprehensive approach using multiple image acquisition devices like radar and visible light cameras.
A method and system for target detection that involves acquiring and processing images from both radar and visible light cameras, employing a region-adaptive feature-level fusion to generate and combine image features, determining target regions, and generating a detection result based on these features.
Enhances detection effectiveness by leveraging the vertical and horizontal field of view advantages of radar and visible light cameras, ensuring consistent performance across all regions.
Smart Images

Figure CN2025085618_16042026_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR TARGET DETECTIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to Chinese Patent Application No. 202411402384. X, filed on October 9, 2024, the entire contents of which are incorporated herein by reference.TECHNICAL FIELD
[0002] The present disclosure generally relates to the field of target detection, and in particular, to a method and a system for target detection and a storage medium.BACKGROUND
[0003] In the field of target detection, a single image acquisition device is usually used as a sensor, and with the development of technology and the increase in demand, the single image acquisition device is no longer able to solve complex and changing situations in target detection. For example, in the field of intelligent transportation, a single visible light camera is usually used as a sensor, but the single visible light camera may no longer solve complex and changing situations in intelligent transportation. Radar cameras may be an effective complement to visible light cameras. How to comprehensively utilize multiple image acquisition devices, e.g., radar cameras and visible light cameras, has become a new trend and challenge in the field of target detection.
[0004] Accordingly, there is a need to provide a method and a system for target detection.SUMMARY
[0005] One of the embodiments of the present disclosure provides a method for target detection. The method may include acquiring a first image and a second image, the first image and the second image representing a same scene at a same time point; generating a first image feature by performing feature extraction on the first image, the first image feature including multiple first image sub-features of multiple first regions; generating a second image feature by performing feature extraction on the second image, the second image feature including multiple second image sub-features of multiple second regions, each of the multiple first regions corresponds to one of the multiple second regions; determining a target image feature based on the first image feature and the second image feature, the target image feature including multiple target regions, a target image sub-feature for each of the multiple target regions being determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region; and generating a target detection result of a specific targets based on the target image feature.
[0006] One of the embodiments of the present disclosure provides a system for target detection. The system may include an acquisition module configured to acquire a first image and a second image, the first image and the second image representing a same scene at a same time point; a first generation module configured to generate a first image feature by performing feature extraction on the first image, the first image feature including multiple first image sub-features of multiple first regions; a second generation module configured to generate a second image feature by performing feature extraction on the second image, the second image feature including multiple second image sub-features of multiple second regions, each of the multiple first regions corresponds to one of the multiple second regions; a determination module configured to determine a target image feature based on the first image feature and the second image feature, the target image feature including multiple target regions, a target image sub-feature of each of the multiple target regions being determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region; and a third generation module configured to generate a target detection result of a specific targets based on the target image feature.
[0007] One of the embodiments of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, and when a computer reads the computer instructions in the storage medium, the computer performs the following method: acquiring a first image and a second image, the first image and the second image representing a same scene at a same time point; generating a first image feature by performing feature extraction on the first image, the first image feature including multiple first image sub-features of multiple first regions; generating a second image feature by performing feature extraction on the second image, the second image feature including multiple second image sub-features of multiple second regions, each of the multiple first regions corresponds to one of the multiple second regions; determining a target image feature based on the first image feature and the second image feature, the target image feature including multiple target regions, a target image sub-feature for each of the multiple target regions being determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region; and generating a target detection result of a specific targets based on the target image feature.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The present disclosure will be further illustrated by way of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbering denotes the same structure, wherein:
[0009] FIG. 1 is a schematic diagram illustrating an application scenario of a system for target detection according to some embodiments of the present disclosure;
[0010] FIG. 2 is a schematic diagram illustrating exemplary hardware and / or software components of an exemplary computing device according to some embodiments of the present disclosure;
[0011] FIG. 3 is an exemplary modular diagram illustrating a system for target detection according to some embodiments of the present disclosure;
[0012] FIG. 4 is an exemplary flowchart illustrating a method for target detection according to some embodiments of the present disclosure;
[0013] FIG. 5 is a flowchart illustrating an exemplary process for generating a target image sub-feature of a target region according to some embodiments of the present disclosure;
[0014] FIG. 6 is an exemplary schematic diagram illustrating a process for determining a target image sub-feature according to some embodiments of the present disclosure;
[0015] FIG. 7 is a schematic diagram illustrating a target detection model according to some embodiments of the present disclosure;
[0016] FIG. 8 is an exemplary flowchart illustrating a process for target detection implemented based on a target detection model according to some embodiments of the present disclosure;
[0017] FIG. 9 is an exemplary flowchart illustrating a process for target detection according to some other embodiments of the present disclosure;
[0018] FIG. 10 is an exemplary flowchart illustrating a process for target detection according to some other embodiment of the present disclosure;
[0019] FIG. 11 is a schematic diagram illustrating an exemplary device for target detection according to some embodiments of the present disclosure; and
[0020] FIG. 12 is a schematic diagram illustrating an exemplary non-transitory computer-readable storage medium according to some embodiments of the present disclosure.DETAILED DESCRIPTION
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be used in the description of the embodiments are briefly described below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present disclosure, and it is possible for a person of ordinary skill in the art to apply the present disclosure to other similar scenarios in accordance with these drawings without creative labor. Unless obviously obtained from the context or the context illustrates otherwise, the same numeral in the drawings refers to the same structure or operation.
[0022] It should be understood that as used herein, the terms "system" , "device" , "unit" and / or "module" as used herein is a way to distinguish between different components, elements, parts, sections or assemblies at different levels. However, the words may be replaced by other expressions if other words accomplish the same purpose.
[0023] As shown in the present disclosure and in the claims, unless the context clearly suggests an exception, the words "one, " "a" , "an" , and / or "the" do not refer specifically to the singular, but may also include the plural. In general, the terms "including" and "comprising" only suggest the inclusion of explicitly identified steps and elements that do not constitute an exclusive list, and the method or devices may also include other steps or elements.
[0024] Flowcharts are used in the present disclosure to illustrate operations performed by a system according to embodiments of the present disclosure. It should be appreciated that the preceding or following operations are not necessarily performed in an exact sequence. Instead, steps may be processed in reverse order or simultaneously. Also, it is possible to add other operations to these processes or remove a step or steps from them.
[0025] Taking multiple image acquisition devices as visible light cameras and radar cameras as an example, the radar-vision fusion methods may include three types: 1) data-level fusion methods; 2) decision-level fusion methods; and 3) feature-level fusion methods. The data-level fusion methods, such as ROI generation through radar points, may be performed based on a count of effective radar points. The decision-level fusion methods, such as combining the results of different target detections, may be performed based on the effectiveness of a single sensor (e.g., an image acquisition device) and the fusion method. The advantage of a radar camera lies in a vertical field of view, and the advantage of a visible camera lies in a horizontal field of view, and there are a crossover and differences in the advantages of the radar camera and the visible camera. A feature-level fusion method may be performed by fusing features of image data from different data sources (e.g., the radar camera and the visible camera) and the fused features at different regions are both determined from different data sources, resulting in less effective applications than single sensors in some regions.
[0026] Thus, some embodiments of the present disclosure use a region-adaptive feature-level fusion method, which may effectively ameliorate the above problems, so that the effect of ray-vision fusion is no less effective than the detection effect of a single sensor in all regions.
[0027] FIG. 1 is a schematic diagram illustrating an application scenario of a system for target detection according to some embodiments of the present disclosure.
[0028] In some embodiments, the application scenario 100 of the system for target detection may include a processing device 110, a network 120, a first image acquisition device 130, and a second image acquisition device 140, as shown in FIG. 1.
[0029] The processing device 110 may process data and / or information obtained from other devices or system components. The processing device 100 may execute program instructions based on the data, information, and / or processing results to perform one or more functions described in this application. For example, the processing device 110 may obtain a first image and a second image and determine a target detection result based on the first image and the second image. More details regarding the first image, the second image, the target detection result, or the like may be found in other contents of the present disclosure (e.g., FIG. 4 and the descriptions thereof) .
[0030] In some embodiments, the processing device 110 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processing device 110 may be local or remote. In some embodiments, the processing device 110 may be implemented on a cloud platform. Merely by way of example, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an inter-cloud, a multi-cloud, or the like, or any combination thereof. In some embodiments, the processing device 110 may be implemented by a computing device 200 having one or more components illustrated in FIG. 2.
[0031] The network 120 may include any suitable network that may facilitate the exchange of information and / or data. For example, the processing device 110, the first image acquisition device 130, and / or the second image acquisition device 140 may pass information and / or data via the network 120.
[0032] In some embodiments, the network may be or include a public network (e.g., the Internet) , a private network (e.g., a local area network (LAN) ) , a wired network, a wireless network (e.g., an 802.11 network, a Wi-Fi network) , a frame relay network, a virtual private network (VPN) , a satellite network, a telephone network, routers, hubs, switches, server computers, and / or any combination thereof. For example, the network may include a cable network, a wireline network, a fiber-optic network, a telecommunications network, an intranet, a wireless local area network (WLAN) , a metropolitan area network (MAN) , a public telephone switched network (PSTN) , a BluetoothTM network, a ZigBeeTM network, a near field communication (NFC) network, or the like, or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include wired and / or wireless network access points such as base stations and / or internet exchange points through which one or more components of the application scenario 100 of the system for target detection may be connected to the network to exchange data and / or information.
[0033] The first image acquisition device 130 and the second image acquisition device 140 refer to devices configured to acquire the first image and the second image, respectively. For example, the first image acquisition device 130 may be a visible light camera and the second image acquisition device 140 may be a radar camera. As another example, the first image acquisition device 130 may be a visible light camera and the second image acquisition device 140 may be another light camera.
[0034] In some embodiments, the application scenario 100 of the system for target detection may also include one or more other devices, for example, a terminal device (not shown in the figures) , a storage device (not shown in the figures) , or the like.
[0035] In some embodiments, a user may interact with the system for target detection via the terminal device. Exemplary terminal devices may include mobile devices, tablets, laptops, etc., or any combination thereof. In some embodiments, the processing device 110 may be part of the terminal device.
[0036] The storage device may store data, instructions, and / or any other information. In some embodiments, the storage device may store data and / or instructions related to target detection. For example, the storage device may store the first image and the second image. As another example, the storage device may store instructions for processing the first image and the second image for target detection. In some embodiments, the storage device may include a mass storage, removable storage, a volatile read-and-write memory, a read-only memory (ROM) , or the like, or any combination thereof. In some embodiments, the storage device may be implemented on a cloud platform. In some embodiments, the storage device may be integrated into the processing device 110 and / or the terminal device.
[0037] In some embodiments, the storage device may be connected to the network 120 to communicate with one or more other components (e.g., the processing device 110, etc. ) of the application scenario 100 of the system for target detection. One or more of the components of the application scenario 100 of the system for target detection may access data or instructions stored in the storage device over the network 120. In some embodiments, the storage device may be part of the processing device 110.
[0038] It should be noted that the above description regarding the application scenario 100 of the system for target detection is merely provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For persons having ordinary skills in the art, multiple variations and modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure. In some embodiments, the application scenario 100 of the system for target detection may include one or more additional components and / or one or more components of the application scenario 100 of the system for target detection described in the present disclosure. Additionally or alternatively, two or more components of the application scenario 100 of the system for target detection may be integrated into a single component. A component of the application scenario 100 of the system for target detection may be implemented on two or more sub-components.
[0039] FIG. 2 is a schematic diagram illustrating exemplary hardware and / or software components of an exemplary computing device according to some embodiments of the present disclosure. In some embodiments, the processing device 110 and / or the terminal device may be implemented on the computing device 200. As illustrated in FIG. 2, the computing device 200 may include a processor 210, a storage 220, an input / output (I / O) 230, and a communication port 240.
[0040] The processor 210 may execute computer instructions (e.g., program code) and perform functions of the processing device 110 in accordance with techniques as described herein. The computer instructions may include, for example, routines, programs, objects, components, data structures, procedures, modules, and functions, which perform particular functions as described herein.
[0041] In some embodiments, the processor 210 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC) , an application specific integrated circuits (ASICs) , an application-specific instruction-set processor (ASIP) , a central processing unit (CPU) , a graphics processing unit (GPU) , a physics processing unit (PPU) , a microcontroller unit, a digital signal processor (DSP) , a field programmable gate array (FPGA) , an advanced RISC machine (ARM) , a programmable logic device (PLD) , any circuit or processor capable of executing one or more functions, or the like, or any combinations thereof.
[0042] Merely for illustration, only one processor is described in the computing device 200. However, it should be noted that the computing device 200 in the present disclosure may also include multiple processors, thus operations and / or method operations that are performed by one processor as described in the present disclosure may also be jointly or separately performed by the multiple processors. For example, if in the present disclosure the processor of the computing device 200 executes both operation A and operation B, it should be understood that operation A and operation B may also be performed by two or more different processors jointly or separately in the computing device 200 (e.g., a first processor executes operation A and a second processor executes operation B, or the first and second processors jointly execute operations A and B) .
[0043] The storage 220 may store data obtained from one or more components of a system for target detection 300. In some embodiments, the storage 220 may include a mass storage device, a removable storage device, a volatile read-and-write memory, a read-only memory (ROM) , or the like, or any combination thereof. In some embodiments, the storage 220 may store one or more programs and / or instructions to perform exemplary methods described in the present disclosure. For example, the storage 220 may store a program for the processing device 110 to execute to compress the machine learning model.
[0044] The I / O 230 may input and / or output signals, data, information, etc. In some embodiments, the I / O 230 may enable a user interaction with the processing device 110. In some embodiments, the I / O 230 may include an input device and an output device. The input device may include a keyboard, a touch screen, a speech input, an eye tracking input, a brain monitoring system, or any other comparable input mechanism. The input information received through the input device may be transmitted to another component (e.g., the processing device 110) via, for example, a bus, for further processing. Other types of the input devices may include a cursor control device, such as a mouse, a trackball, or cursor direction keys, etc. The output device may include a display (e.g., a liquid crystal display (LCD) , a light-emitting diode (LED) -based display, a flat panel display, a curved screen, a television device, a cathode ray tube (CRT) , a touch screen) , a speaker, a printer, or the like, or a combination thereof.
[0045] The communication port 240 may be connected to a network (e.g., the network 160) to facilitate data communications. The communication port 240 may establish connections between the processing device 110 and the terminal device. The connection may be a wired connection, a wireless connection, any other communication connection that may enable data transmission and / or reception, and / or any combination of these connections. The wired connection may include, for example, an electrical cable, an optical cable, a telephone wire, or the like, or any combination thereof. The wireless connection may include, for example, a BluetoothTM link, a Wi-FiTM link, a WiMaxTM link, a WLAN link, a ZigBeeTM link, a mobile network link (e.g., 3G, 4G, 5G) , or the like, or a combination thereof. In some embodiments, the communication port 240 may be and / or include a standardized communication port, such as RS232, RS485, etc. In some embodiments, the communication port 240 may be a specially designed communication port.
[0046] FIG. 3 is an exemplary modular diagram illustrating a system for target detection according to some embodiments of the present disclosure.
[0047] In some embodiments, as shown in FIG. 3, the system for target detection 300 may include an acquisition module 310, a first generation module 320, a second generation module 330, a determination module 340, and a third generation module 350.
[0048] The acquisition module 310 may be configured to acquire a first image and a second image. The first image and the second image may represent a same scene at a same time point.
[0049] In some embodiments, when at least one of the first image and the second image is acquired based on radar data, the acquisition module 310 may be further configured to convert the radar data into a grayscale image to acquire at least one of the first image and the second image.
[0050] The first generation module 320 may be configured to generate a first image feature by performing feature extraction on the first image. The first image feature includes multiple first image sub-features of multiple first regions.
[0051] In some embodiments, the first generation module 320 may be further configured to generate, based on the first image, the first image feature by using a first sub-model in a target detection model.
[0052] The second generation module 330 may be configured to generate a second image feature by performing feature extraction on the second image. The second image feature includes multiple second image sub-features of multiple second regions, and each of the multiple first regions correspondes to one of the multiple second regions.
[0053] In some embodiments, the second generation module 330 may be further configured to generate, based on the second image, the second image feature by a second sub-model in the target detection model.
[0054] The determination module 340 may be configured to determine a target image feature based on the first image feature and the second image feature. The target image feature includes multiple target regions. A target image sub-feature of each of the multiple target regions may be determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region.
[0055] In some embodiments, the determination module 340 may be further configured to generate a third image feature based on the first image feature and the second image feature. The third image feature includes multiple third image sub-features of multiple third regions, and each of the multiple third regions corresponds to one of the multiple first regions, one of the multiple second regions, and one of the multiple target regions. And for each of the multiple target regions, the determination module 340 may be configured to generate the target image sub-feature corresponding to the target region, based on at least one of the first image sub-feature of the first region corresponding to the target region, the second image sub-feature of the second region corresponding to the target region, and the third image sub-feature of the third region corresponding to the target region.
[0056] In some embodiments, the determination module 340 may be further configured to designate one of the first image sub-feature, the second image sub-feature, and the third image sub-feature, as the target image sub-feature of the target region.
[0057] In some embodiments, the determination module 340 may be further configured to in response to determining that a first condition is satisfied, designate the first image sub-feature as the target image sub-feature of the target region. The first condition includes: a first target confidence level corresponding to the first region being greater than or equal to a first threshold, and a second target confidence level corresponding to the second region being less than a second threshold. The determination module 340 may be further configured to in response to determining that a second condition is satisfied, designate the second image sub-feature as the target image sub-feature of the target region. The second condition includes: the first target confidence level corresponding to the first region being less than the first threshold, and the second target confidence level corresponding to the second region being greater than or equal to the second threshold. And The determination module 340 may be further configured to in response to determining that the first condition and the second condition are not satisfied, designate the third image sub-feature as the target image sub-feature of the target region.
[0058] In some embodiments, the determination module 340 may be further configured to obtain a first detection result of the first image by performing a target detection based on the first image feature. The first detection result includes first position information of one or more specific targets in the first image and a first confidence level corresponding to each of the one or more specific targets. And the determination module 340 may be further configured to determine the first target confidence level based on the first confidence level and the first position information.
[0059] In some embodiments, the determination module 340 may be further configured to obtain a second detection result of the second image by performing a target detection based on the second image feature. The second detection result includes second position information of one or more specific targets in the second image and a second confidence level corresponding to each of the one or more specific targets. And the determination module 340 may be further configured to determine the second target confidence level based on the second confidence level and the second position information.
[0060] In some embodiments, the determination module 340 may be further configured to obtain at least one historical first detection result of at least one first historical image, with each of the at least one historical first detection result including a historical first confidence level; and determine the first target confidence level based on the at least one historical first confidence level. The first image and the at least one first historical image may be acquired by the first image acquisition device during a continuous time period.
[0061] In some embodiments, the determination module 340 may be further configured to, in response to determining that a posture of the first image acquisition device acquiring the first image of a scene that is the same as a posture of the first image acquisition device acquiring the at least one first historical image of the scene, determine, based on the historical first confidence level, the first target confidence level.
[0062] In some embodiments, the determination module 340 may be further configured to obtain at least one historical second detection result of the at least one second historical image, with each of the at least one historical second detection result including a historical second confidence level; and determine the second target confidence level based on the at least one historical second confidence level. The second image and the at least one second historical image may be acquired by the second image acquisition device during a continuous time period.
[0063] In some embodiments, the determination module 340 may be further configured to, in response to determining that a posture of the second image acquisition device acquiring the second image of a scene that is the same as a posture of the second image acquisition device acquiring the at least one second historical image of the scene, determine, based on the historical second confidence level, the second target confidence level.
[0064] In some embodiments, the determination module 340 may be further configured to generate, based on the first image feature and the second image feature, the third image feature by using a third sub-model of the target detection model; and generate the target image feature based on the first image feature, the second image feature and the third image feature.
[0065] The third generation module 350 may be configured to generate a target detection result based on the target image feature.
[0066] In some embodiments, the third generation module 350 may be further configured to generate the target detection result, based on the target image feature, by using the third sub-model.
[0067] In some embodiments, the system for target detection 300 may further include a training module (not shown in the figures) , and the training module may be configured to acquire a sample set including multiple groups of first sample images and second sample images. Each group of the multiple groups of first sample images and second sample images represent a same scene. The training module may be further configured to train an initial first sub-model based on the multiple groups of first sample images to obtain the first sub-model; train an initial second sub-model based on the multiple groups of second sample images to obtain the second sub-model; and train the third sub-model based on the sample set, during the training process of the third sub-model, parameters of the first sub-model and parameters of the second sub-model are unchanged.
[0068] More details regarding the acquisition module 310, the first generation module 320, the second generation module 330, the determination module 340, and the third generation module 350 may be found in other contents of the present disclosure (e.g., descriptions in connection with FIG. 4 -FIG. 10) .
[0069] It should be noted that the above descriptions of the system for target detection 300 are provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For persons having ordinary skills in the art, various modifications and changes in the forms and details of the application of the above method and system may occur without departing from the principles of the present disclosure. In some embodiments, the system for target detection 300 may include one or more other modules and / or one or more modules described above may be omitted. Additionally or alternatively, two or more modules may be integrated into a single module and / or a module may be divided into two or more units. However, those variations and modifications also fall within the scope of the present disclosure.
[0070] FIG. 4 is an exemplary flowchart illustrating a method for target detection according to some embodiments of the present disclosure.
[0071] In some embodiments, process 400 may be executed by the system for target detection 300. For example, process 400 may be implemented as a set of instructions stored in a storage device. In some embodiments, the processing device 110 (e.g., the processor 210 of the computing device 200 and / or one or more modules illustrated in FIG. 3) may execute the set of instructions and may accordingly be directed to perform the process 400. The operations of the illustrated process presented below are intended to be illustrative. In some embodiments, the process 400 may be accomplished with one or more additional operations not described and / or without one or more of the operations discussed. Additionally, the order of the operations of process 400 illustrated in FIG. 4 and described below is not intended to be limiting.
[0072] In 410, a first image and a second image may be acquired. Operation 410 may be performed by the acquisition module 310.
[0073] The first image and the second image refer to data obtained by imaging a same scene by two different image acquisition devices. The first image and the second image may represent the same scene at the same time point. The image acquisition devices may include a visible light camera, a radar camera, an infrared camera, an ultraviolet camera, an X-ray camera, an ultrasonic camera, or the like. For example, the first image is an image captured by the visible light camera, and the second image is an image transformed from point cloud data captured by the radar camera. As another example, the first image is an image captured by a visible light camera, and the second image is an image captured by another visible light camera. The scene refers to a specific physical area or environment being located within a monitoring range of each of the image acquisition devices. For example, the scene in intelligent transportation includes elements such as a road, a traffic sign, a pedestrian, a vehicle, etc.
[0074] In some embodiments, the different image acquisition devices may include the first image acquisition device 130 and the second image acquisition device 140.
[0075] In some embodiments, the processing device may generate a first image based on first image data captured by the first image acquisition device for the scene, and generate a second image based on second image data captured by the second image acquisition device for the scene.
[0076] In some embodiments, the processing device may obtain the first image from the first image acquisition device, and obtain the second image from the second image acquisition device.
[0077] One of the first image and the second image may be a two-dimensional image, a three-dimensional image, or the like. When the dimensions of the first image and the second image are different, the processing device may convert a high-dimensional image among the first image and the second image to a low-dimensional image through a process such as orthographic projection, perspective projection, image rendering, etc.; or the processing device may convert the low-dimensional image among the first image and the second image to the high-dimensional image through approaches such as light field reconstruction, machine learning modeling, manual modeling, or the like, to make the converted first image and the converted second image being in the same dimensions. Understandably, converting the low-dimensional image to the high-dimensional image uses depth estimation, which is less accurate than converting the high-dimensional image to the low-dimensional image, and thus converting the high-dimensional image to the low-dimensional image is mostly used.
[0078] The size of the first image and the second image may be the same or different.
[0079] When acquisition moments of the first image acquisition device and the second image acquisition device are one-to-one, in other words, the first acquisition frequency of the first image acquisition device and the second acquisition frequency of the second image acquisition device are the same, the processing device may take an image acquired by the first image acquisition device as the first image; an image acquired by the second image acquisition device at the same acquisition moment of the first image as the second image.
[0080] When the first acquisition frequency of the first image acquisition device is different from the second acquisition frequency of the second image acquisition device, and a set including multiple acquisition moments corresponding to a lower acquisition frequency among the first acquisition frequency and the second acquisition frequency is a subset of a set including multiple acquisition moments corresponding to a higher acquisition frequency among the first acquisition frequency and the second acquisition frequency, the processing device may take an image acquired by the first image acquisition device and an image acquired by the second image acquisition device at any one of the multiple acquisition moments corresponding to the lower acquisition frequency as the first image and the second image, respectively. For example, taking the first image acquisition device as a radar camera and the second image acquisition device as a visible light camera as an example, since the acquisition frequency of the radar camera is lower than the acquisition frequency of the visible light camera, if a set including multiple acquisition moments corresponding to the acquisition frequency of the radar camera (e.g., T1, T3, T5, ..., Tn) is satisfied to be a subset of a set including multiple acquisition moments corresponding to the acquisition frequency of the visible light camera (e.g., T1, T2, T3, T4, T5, ..., Tn) , the processing device may designate an image acquired by the radar camera and an image acquired by the visible light camera at any one of the multiple acquisition moments corresponding to the acquisition frequency of the radar camera (e.g., T1, T3, T5, ..., Tn) as the first image and the second image, respectively, to ensure that the first image and the second image are acquired at the same moment (or time point) .
[0081] When the first acquisition frequency of the first image acquisition device is different from the second acquisition frequency of the second image acquisition device, the processing device may perform a time synchronization operation of the first image data and / or the second image data to generate the first image and the second image.
[0082] The time synchronization operation refers to a process that the first image data and the second image data are processed to acquire the first image and the second image acquired at the same moment. As the first acquisition frequency of the first image acquisition device may be different from the second acquisition frequency of the second image acquisition device, to ensure that the acquisition moment of the first image and the acquisition moment of the second image are the same, it is necessary to carry out the time synchronization operation.
[0083] The processing device may perform the time synchronization operation by using a timestamp marking technique, a frame alignment technique, an interpolation algorithm, or the like. The interpolation algorithm may include a linear interpolation algorithm, a spline interpolation algorithm, a polynomial interpolation algorithm, etc., or any combination thereof. For example, the processing device may mark each image frame captured by the first image acquisition device and the second image acquisition device with a timestamp, and then, based on the timestamps corresponding to the image frames captured by the image acquisition devices, map the mage frames from different image acquisition devices to the same unified time coordinate system. When the first acquisition frequency of the first image acquisition device is different from the second acquisition frequency of the second image acquisition device, and frame rates of the first image acquisition device and the second image acquisition device are not accurately aligned, the processing device may use the interpolation algorithm to interpolate image frames in the time domain to achieve time synchronization of the first image acquisition device and the second image acquisition device. Exemplarily, when the acquisition moments of the first image acquisition device include t1, t3, t5, ....., tm, the acquisition moments of the second image acquisition device include t1, t2, t3, t4, t5, ....., tn, then the processing device needs to determine, by using the interpolation algorithm, image frames at t2, t4, ....., tn based on image frames acquired at t1, t3, t5, ....., tm.
[0084] In some embodiments, the processing device may preprocess the first image data and / or the second image data to generate the first image and the second image by other ways. Exemplarily, the ways for realizing the preprocessing include, but are not limited to, a filtering operation, a denoising operation, a coordinate system transformation operation, a scale adjustment operation, or the like, or any combination thereof.
[0085] The filtering operation refers to a process of processing a signal, data, or image in a way that removes unwanted components or enhances components of interest. The filtering operation may be useful for removing noise, smoothing data, and improving the quality of a signal. In some embodiments, the filtering operation may include using a low-pass filtering algorithm, a high-pass filtering algorithm, a band-pass filtering algorithm, a mean filtering algorithm, a median filtering algorithm, or the like.
[0086] The denoising operation refers to a process of removing or reducing noise from data such as signals, images, audio or video. The denoising operation serves a purpose of retaining the valid information of a signal while removing as many interfering components as possible to improve the quality of the signal. In some embodiments, the denoising operation may include a filtering algorithm, an averaging algorithm, a wavelet transform, or the like.
[0087] The coordinate system transformation operation refers to a process of converting or mapping between different coordinate systems. Through the coordinate system transformation operation, the first image data and the second image data may be converted to the same coordinate system, which facilitates the subsequent feature fusion.
[0088] The processing device may implement the coordinate system transformation operation through a coordinate system conversion matrix. For example, taking the first image acquisition device is a radar camera and the second image acquisition device is a visible light camera as an example, the first image data acquired by the first image acquisition device is radar data and the second image data acquired by the second image acquisition device is image data, the processing device may convert the radar data from a polar coordinate system of the radar data to a Cartesian coordinate system of the image data; or convert the image data from the Cartesian coordinate system to the polar coordinate system; or convert both the image data and the radar data to an additional coordinate system (e.g., a geographic coordinate system) .
[0089] In some embodiments, when the size of the first image and the size of the second image are different, the processing device may crop the image of the larger size to make the first image and the second image have the same size.
[0090] In some embodiments, when at least one of the first image and the second image is acquired based on the radar data, the processing device may further convert the radar data to a grayscale image to acquire at least one of the first image and the second image.
[0091] The radar data refers to information obtained through radar technology related to position, distance, speed, and other characteristics of an object. The coordinate system of the radar data includes the polar coordinate system. In some embodiments, the radar data may be denoted as point cloud data.
[0092] The grayscale image refers to an image that has only grayscale information and no color information. In some embodiments, the grayscale value of a grayscale image converted from the radar data represents the intensity of the reflection.
[0093] In some embodiments, the processing device may obtain the radar data acquired by the radar camera on the scene, convert the radar data from the polar coordinate system to the Cartesian coordinate system to obtain a third image, and convert the third image into a grayscale image, and the grayscale image may be designated as the first image. For example, the processing device may first convert distances and angles (i.e., each data point in the radar data) in polar coordinates to horizontal and vertical coordinates in the Cartesian coordinate system by using one or more formulas; scale and translate coordinates in the Cartesian coordinate system to map all data points onto a rectangular image region to generate the third image, for each data point in the third image or the radar data, normalize reflection intensity values of the radar camera within a range of 0-255. The normalized intensity values are placed at the third image, resulting in the grayscale image (i.e., the first image) . The manner of implementing the normalization may include linear normalization, etc.
[0094] In some embodiments, since the radar camera is not able to directly acquire image data, the radar data acquired by the radar camera is first converted into a coordinate system, i.e., converted from the polar coordinate system into the Cartesian coordinate system, to obtain the third image, and then perform grayscale processing on the third image to obtain the first image, i.e., the first image is a grayscale image. The image grayscale values in the first image represent the reflection intensity. The third image refers to an image after the coordinate system transformation of the radar data.
[0095] In some embodiments of the present disclosure, the radar data itself being a set of data points cannot be directly divided into regions. By converting the radar data to the grayscale image, it becomes easier for subsequent region division.
[0096] In 420, a first image feature may be generated by performing feature extraction on the first image. Operation 420 may be performed by the first generation module 320.
[0097] The feature extraction refers to a process of selecting, separating, and / or identifying image features from raw image data that are useful for analysis purposes.
[0098] As used herein, an image feature (e.g., the first image feature or the second image feature) refers to a set of various properties of an image (e.g., the first image, the second image) , including color, shape, texture, etc. For example, the image feature may include a feature or characteristic of an object in the image that distinguishes the object from other classes of objects, or a collection of such features and characteristics. In some embodiments, the various properties of an image may include brightness, contrast, color, texture, shape, image moments, histogram, principal components, key point features, edge features, or the like, or any combination thereof. The image feature (e.g., the first image feature or the second image feature) extracted from an image (e.g., the first image or the second image) may be denoted as a matrix or an image including multiple elements and each of the multiple elements may correspond to a region or a pixel of the image. Each of the multiple elements may represent various properties of the corresponding region or pixel of the image. For example, each element of the first image feature may represent various properties of the corresponding region or pixel of the first image. As used herein, an element of the image feature extracted from an image corresponding to the region or the pixel of the image refers to that the element and the region or the pixel having the same position in the image feature and the image.
[0099] In some embodiments, the processing device may perform the feature extraction in multiple ways, e.g., using a feature extraction algorithm, etc. Exemplary feature extraction algorithm includes a Histogram of Oriented Gradient (HOG) feature extraction algorithm, a Local Binary Pattern (LBP) feature extraction algorithm, a Haar feature extraction algorithm, or the like. As another example, the processing device may extract the first image feature by using a first sub-model.
[0100] The first sub-model refers to a model used to perform the detection of the first image.
[0101] In some embodiments, the first sub-model may include a first feature extraction network. The first feature extraction network refers to a model for performing feature extraction on the first image. For example, the first sub-model may include a multi-layer network, e.g., the first sub-model may be a Convolutional Neural Network (CNN) , a Region Proposal Network (RPN) , etc. The multi-layer network may include a convolutional layer, a pooling layer, a normalization layer, an activation function layer, or the like.
[0102] More details regarding the first sub-model may be found in other contents of the present disclosure (e.g., description in connection with FIG. 7) .
[0103] In some embodiments, the first image feature includes multiple first image sub-features of multiple first regions. Each of the multiple first image sub-features may correspond to one of the multiple first regions. For example, the first image feature may be denoted by a first matrix and the first matrix may include the multiple first regions, and each of the multiple first regions may include a first image sub-feature.
[0104] In some embodiments, the processing device may divide the first image feature into the multiple first regions by a division pattern. The division pattern may be defined by division parameters including, e.g., a direction of each of the division lines and a count of the division lines, an interval between two adjacent division lines. For example, the first image feature may be denoted as the first matrix or an image. The division pattern of the first image feature may be defined by two vertical lines parallel to the column direction of the first image feature, two horizontal lines parallel to the row direction of the first image feature, and the interval between any two adjacent division lines. As a further example, the processing device divides the first image feature into N x M first regions equally. N and M are integers greater than 0, which are preset according to the actual needs when used. Exemplarily, as shown in FIG. 6, the processing device divides the first image feature 610 into 3×3 first regions equally.
[0105] In some embodiments, the division pattern of the first image feature may be set by the user based on experience or set by default by the system. The accuracy requirement for target detection may be proportional to a count of regions divided. During a division process, the size of the regions must be kept moderate, neither too large nor too small. If the size of the regions are too large, target image sub-features at each position is not possible to accurately determine, and the regions may only be roughly differentiated, resulting in poor accuracy; conversely, if the regions are too small, it may result in insufficient details covered to reveal the key features of a specific targets, potentially missing important feature associations, or the size of the specific targets may cover one or more regions, which is not conducive to subsequent detection. It should be noted that the division of the image needs to take into account the position where the specific targets is located in the first image, for example, the division of the first image feature needs to be made such that the magnitude of the divided regions includes at least one specific targets; when the first image acquisition device is a visible light camera, an angle of the visible light camera needs to be considered when dividing the image; when the first image acquisition device is a radar camera, a distance of the radar camera needs to be considered when dividing the image.
[0106] The first image sub-feature refers to a partial feature of the first image feature corresponding to the first region. For example, as shown in FIG. 6, the first image feature 610 is divided into nine first regions, and the nine first regions in the first image feature 610 correspond to nine first image sub-features respectively. Exemplarily, the first region in the first row and first column of the first image feature 610 corresponds to the first image sub-feature A1.
[0107] In 430, a second image feature may be generated by performing feature extraction on the second image. Operation 430 may be performed by the second generation module 330.
[0108] In some embodiments, the processing device may perform feature extraction on the second image to obtain the second image feature, for example, using a feature extraction algorithm, etc. As another example, the processing device may, by using a second sub-model, extract the second image feature.
[0109] The second sub-model refers to a model used to perform detection on the second image.
[0110] In some embodiments, the second sub-model may include a second feature extraction network. The second feature extraction network refers to a model for performing feature extraction on the second image. For example, the second sub-model may include a multi-layer network, e.g., the second sub-model may be a CNN, a RPN, etc. The multi-layer network may include a convolutional layer, a pooling layer, a normalization layer, an activation function layer, or the like.
[0111] In some embodiments, the first sub-model and the second sub-model may be the same model or may be different models of the same structure.
[0112] More details regarding the feature extraction algorithm may be found in the corresponding description above. More details regarding the second sub-model may be found in other contents of the present disclosure (e.g., description in connection with FIG. 7) .
[0113] The processing device may use a same feature extraction technique to perform feature extraction on the first image and the second image, or may use different feature extraction techniques to perform feature extraction on the first image and the second image.
[0114] In some embodiments, the second image feature includes multiple second image sub-features of multiple second regions. Each of the multiple second image sub-features may correspond to one of the multiple second regions. For example, the second image feature may be denoted by a second matrix and the second matrix may include the multiple second regions, and each of the multiple second regions may include a second image sub-feature.
[0115] Each of the multiple first regions may correspond to one of the multiple second regions. As used herein, two corresponding regions in the first image feature and the second image feature refer to a first region in the first image feature and a second region in the second image feature having the same position in the first image feature and in the second image feature. For example, the first image feature and the second image feature may be denoted as the first matrix and the second matrix, respectively, the corresponding first region and second region are located in the same position of the first matrix and the second matrix.
[0116] In some embodiments, to ensure that each first region corresponds to one second region, the division patterns of the first image feature and the second image feature are the same. The division patterns of the first image feature and the second image feature being the same refers to that each division parameter of the first image feature and each division parameter of the second image feature are the same. For example, the direction of each of the division lines, the count of the division lines, and the intervals between two adjacent division lines of the first image feature and the second image feature are the same. The division patterns being the same such that the processing device divides the first image feature and the second image feature into an equal count of first regions and second regions, respectively, and that the multiple first regions corresponding to the first image feature and the multiple second regions corresponding to the second image feature are one-to-one.
[0117] More details regarding the division pattern may be found in the corresponding description above.
[0118] The second image sub-feature refers to portions of the second image feature corresponding to the second region. For example, as shown in FIG. 6, the second image feature 620 is divided into nine second regions, and the nine second regions in the second image feature 620 correspond to nine second image sub-features respectively. Exemplarily, the second region in the first column and the first row of the second image feature 620 corresponds to the second image sub-feature A2.
[0119] In some embodiments, each of the first image feature and the second image feature may be represented in a matrix having the same size as the first image and the second image. Elements in the matrix representing each of the first image feature and the second image feature correspond one-to-one with pixels in each of the first image and the second image. The elements in the matrix and the corresponding pixels in the images represent the same position in the scene.
[0120] In 440, a target image feature may be determined based on the first image feature and the second image feature. Operation 440 may be performed by the determination module 340.
[0121] The target image feature includes multiple target image sub-features of multiple target regions. Each of the multiple target image sub-features may correspond to one of the multiple target regions. For example, the target image feature may be denoted by a target matrix and the target matrix may include the multiple target regions, and each of the multiple target regions may include a target image sub-feature.
[0122] Each of the multiple target regions may correspond to one of the multiple first regions and one of the multiple second regions. As used herein, corresponding regions in the target image feature and each of the first image feature and the second image feature refers to a target region in the target image feature and each of a first region in the first image feature and a second region in the second image feature having the same position in the target image feature, the first image feature, and the second image feature. For example, the first image feature and the second image feature may be denoted as a first matrix and a second matrix, respectively, the target image feature may be denoted as a target matrix, the first matrix, the second matrix, and the target matrix have the same size, the corresponding target region, first region, and second region are located in the same position of the first matrix, the second matrix, and the target matrix.
[0123] A target image sub-feature for each of the multiple target regions is determined based on at least one of a first image sub-feature of a first region corresponding to the target region and a second image sub-feature of a second region corresponding to the target region.
[0124] Target regions refer to portions of the target image feature.
[0125] The multiple target image sub-features of the multiple target regions may be stitched together as the target image feature. For example, as shown in FIG. 6, the target image feature 650 may include nine target image sub-features.
[0126] The processing device may generate the target image sub-feature of the target region in multiple ways, based on at least one of the first image sub-feature of the first region and the second image sub-feature of the second region. For example, the processing device may determine a first target confidence level of a first image-sub-feature of the first region corresponding to the target sub-region and a second target confidence level of a second image sub-feature of the second region corresponding to the target region; if the first target confidence level is greater than the second target confidence level, the first image sub-feature corresponding to the first target confidence level is determined as the target image sub-feature; if the first target confidence level is less than the second target confidence level, the second image sub-feature corresponding to the second target confidence level is determined as the target image sub-feature; if the first target confidence level is equal to the second target confidence level, at least one of the first image sub-feature or the second image sub-feature is determined as the target image sub-feature (e.g., a weighted sum of the first image sub-feature and the second image sub-feature) . More details regarding determining the first target confidence level and the second target confidence level may be found in other contents of the present disclosure (e.g., description in connection with FIG. 5) .
[0127] As another example, the processing device may generate a third image feature including multiple third image sub-features based on the first image feature and the second image feature, and, based on the first image sub-feature of the first region corresponding to the target region, the second image sub-feature of the second region corresponding to the target region, and a third image sub-feature of a third region corresponding to the target region, generate the target image sub-feature of the target region.
[0128] The third image feature may be obtained after fusing the first image feature and the second image feature. For example, the third image feature 630 is a horizontal line-filled portion as shown in FIG. 6. More details regarding this section may be found in other contents of the present disclosure (e.g., description in connection with FIG. 5) .
[0129] In some embodiments, the target image feature may be determined by using a third sub-model. For example, the first image feature and the second image feature may be inputted into the third sub-model and the third sub-model may generate the target image feature according to one of the above processes. More details regarding this section may be found in other contents of the present disclosure (e.g., description in connection with FIG. 7) .
[0130] In 450, a target detection result may be generated based on the target image feature. Operation 450 may be performed by the third generation module 350.
[0131] Target detection refers to detecting one or more subjects of interest, determining a position of each of the one or more subjects, determining a classification of the one or more subjects, and / or determining whether the one or more subjects belong to a specific target or classification. The target detection result includes the target position information and / or target classification information of each of one or more subjects represented in the scene monitored by the first image acquisition device and the second image acquisition device.
[0132] The target position information of a subject may include a target position of the subject represented in the scene. In some embodiments, the processing device may mark the target position of the subject in at least one of the first image and the second image. For example, the processing device may use a bounding box (also referred to as a detection box) enclosing the subject to mark the subject in at least one of the first image and the second image. The bounding box refers to a rectangular region enclosing and localizing a subject in the image during a target detection task. The bounding box may include a rectangular frame, a circular frame, a polygonal frame, or the like. For example, assuming that the subject is a vehicle, the target detection result may include: the target position information of vehicle 1 is the upper left corner coordinates of a rectangular bounding box enclosing vehicle 1 as (x1, y1) , the lower right corner coordinates of the rectangular bounding box as (x2, y2) , and the target confidence level of vehicle 1 is 0.5; the target position information of vehicle 2 is the upper left corner coordinates of a bounding box enclosing vehicle 2 as (x3, y3) , the lower right corner coordinates of the bounding box as (x4. y4) , and the target confidence level of vehicle 2 is 0.8.
[0133] In some embodiments, the target position information may also be represented in other forms, e.g., the target position information may include position information for a segmentation mask.
[0134] The target classification information of a subject may include a target classification of the subject. The target classification may indicate a classification or a specific target that the subject belongs to. For example, the target classification may include a vehicle, a car, a pedestrian, an animal, etc. The target confidence level of a detected subject represents the probability that the subject belongs to one of a vehicle, a car, a pedestrian, or an animal. As another example, the target classification may include an anomaly result or a normal result. The target detection result may include a target confidence level corresponding to each of the one or more detected subject. The target confidence level for a detected subject refers to a probability that the subject belongs to a target classification or a specific target. For example, the target confidence level for a detected subject refers to a probability that the predicted bounding box for enclosing the detected subject contains a specific target (e.g., a vehicle, a car, a pedestrian, or an animal) or the detected subject belongs to the specific target (e.g., a vehicle, a car, a pedestrian, or an animal) . The target confidence level may be expressed as a real value ranging from 0 to 1.
[0135] In some embodiments, the target detection result may be determined by using a target detection model including the first sub-model, the second sub-model, and the third sub-model. More details regarding the target detection model may be found in other contents of the present disclosure (e.g., description in connection with FIG. 7) .
[0136] In some embodiments, the processing device may stitch all the target image sub-features corresponding to the multiple target regions, thereby obtaining the target image feature, and then based on the target image feature, determine the target detection result. For example, as shown in FIG. 6, the processing device stitches the target image sub-features corresponding to the nine target regions to obtain the target image feature 650, and then determine, based on the target image feature 650, the target detection result of the specific targets.
[0137] In some embodiments of the present disclosure, by dividing the first image feature and the second image feature into the multiple first image sub-features and the multiple second image sub-features, respectively, and determine the target image sub-features of different target regions based on the first image sub-features and second image sub-features of first regions and second regions corresponding to the target regions, thereby the target image sub-features of different target regions may be determined adaptively according to confidence levels of first image sub-features and second image sub-features of the first regions and the second regions. In other words, target image sub-features of the different target regions would not be determined according to the same confidence level, but determined according to different confidence levels of different regions, thereby improving the accuracy of the fused image feature, further improving the accuracy of target detection. Subsequently, based on the target image sub-feature including the target image features, an accurate target detection result of the specific targets can be generated so that the post-target detection effect is not inferior to the detection effect of a single image acquisition device in all regions.
[0138] It should be noted that the above description of the process 400 is merely provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For persons having ordinary skills in the art, multiple variations or modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure.
[0139] FIG. 5 is a flowchart illustrating an exemplary process for generating a target image sub-feature of a target region according to some embodiments of the present disclosure.
[0140] In some embodiments, process 500 may be executed by the system for target detection 300. For example, the process 500 may be implemented as a set of instructions stored in a storage device. In some embodiments, the processing device 110 (e.g., the processor 210 of the computing device 200 and / or one or more modules illustrated in FIG. 3) may execute the set of instructions and may accordingly be directed to perform the process 500. The operations of the illustrated process presented below are intended to be illustrative. In some embodiments, the process 500 may be accomplished with one or more additional operations not described and / or without one or more of the operations discussed. Additionally, the order of the operations of process 500 illustrated in FIG. 5 and described below is not intended to be limiting.
[0141] In 510, a third image feature may be generated based on a first image feature and a second image feature. The third image feature includes multiple third image sub-features of multiple third regions.
[0142] The first image feature is extracted from a first image and the second image feature is extracted from a second image. The first image and the second image represent the same scene at the same time point. The first image feature includes multiple first image sub-features of multiple first regions. Each of the multiple first image sub-features may correspond to one of the multiple first regions. For example, the first image feature may be denoted by a first matrix and the first matrix may include the multiple first regions, and each of the multiple first regions may include a first image sub-feature. The second image feature includes multiple second image sub-features of multiple second regions. Each of the multiple second image sub-features may correspond to one of the multiple second regions. For example, the second image feature may be denoted by a second matrix and the second matrix may include the multiple second regions, and each of the multiple second regions may include a second image sub-feature.
[0143] In some embodiments, the third image feature includes multiple third image sub-features of multiple third regions. Each of the multiple third image sub-features may correspond to one of the multiple third regions. For example, the third image feature may be denoted by a third matrix and the third matrix may include the multiple third regions, and each of the multiple third regions may include a third image sub-feature.
[0144] Each of the multiple third regions corresponds to one of multiple first regions, one of multiple second regions. As used herein, corresponding regions in the third image feature, the first image feature, and the second image feature refer to a third region in the third image feature, a first region in the first image feature, and a second region in the second image feature having the same position in the third image feature, the first image feature, and the second image feature. For example, the first image feature, the second image feature, and the third image feature may be denoted as a first matrix, a second matrix, and a third matrix, respectively. The first matrix, the second matrix, and the third matrix have the same size, and the corresponding first region, second region, and third region are located in the same position of the first matrix, the second matrix, and the target matrix.
[0145] In some embodiments, the processing device divides the third image feature into the multiple third regions in a same division pattern as the first image feature or the second image feature. More descriptions for the division pattern may be found in FIG, 4 and the descriptions thereof. For example, a count and a direction of division lines of the third image feature are the same as the count and the direction of division lines of the first image feature and the second image feature.
[0146] The third image sub-feature refers to a partial feature of the third image feature corresponding to the third region. For example, as shown in FIG. 6, the third image feature 630 is divided into nine third regions, respectively, and each of the nine third regions in the third image feature 630 corresponds to one of nine third image sub-features. Exemplarily, the third region in the first row and first column of the third image feature 630 corresponds to the third image sub-feature A3.
[0147] In some embodiments, the first image feature, the second image feature, and the third image feature represent a same scene, and the processing device may use a same division pattern with the first image feature or the second image feature to divide the third image feature into the multiple third regions, which ensures that each of the multiple third regions correspond to one of the multiple first regions, one of the multiple second regions, and one of the multiple target regions. Dividing the third image feature into the multiple third regions is similar to dividing the first image feature into the multiple first regions, and may be referred to in FIG. 4 and its corresponding description.
[0148] More details regarding the first image feature, the second image feature, the division pattern, or the like may be found in other contents of the present disclosure (e.g., description in connection with FIG. 4) .
[0149] The processing device may generate the third image feature based on the first image feature and the second image feature in multiple ways. For example, the processing device may determine the third image feature by fusing the first image feature and the second image feature based on the weight of the first image feature and the weight of the second image feature. The weight of the first image feature and the weight of the second image feature may be determined based on a first target confidence level of the first image feature and a second target confidence level of the second image feature. The ratio of the weight of the first image feature and the weight of the second image feature may be same as a ratio of the first target confidence level and the second target confidence level. As another example, the processing device may fuse a first image and a second image to obtain a third image and the third image feature is extracted from the third image.
[0150] In some embodiments, the processing device may input the first image feature and the second image feature into a fusion unit of a third sub-model to obtain the third image feature. The fusion unit may process the first image feature and the second image feature in multiple ways to obtain the third image feature. For example, the fusion unit may fuse the first image feature and the second image feature by stitching the first image feature and the second image feature in a channel dimension to obtain a stitched image feature, and then feed the stitched image feature into an attention mechanism module of the fusion unit such as a squeeze-and-excitation attention (SE) module, a convolutional block attention module (CBAM) , etc., to obtain the third image feature.
[0151] The fusion unit refers to a network layer of the third sub-model used to generate the third image feature. The fusion unit is a portion of the third sub-model. More details regarding the fusion unit and the third sub-model may be found in other contents of the present disclosure (e.g., description in connection with FIG. 7) .
[0152] In 520, a target image feature may be determined based on the first image feature, the second image feature, and the third image feature. In some embodiments, operation 520 may be performed by the determination module 340.
[0153] The target image feature includes multiple target image sub-features of multiple target regions. Each of the multiple target image sub-features may correspond to one of the multiple target regions. For example, the target image feature may be denoted by a target matrix and the target matrix may include the multiple target regions, and each of the multiple target regions may include a target image sub-feature.
[0154] Each of the multiple target regions may correspond to one of the multiple first regions, one of the multiple second regions, and one of the multiple third regions. As used herein, corresponding regions in the target image feature and each of the first image feature, the second image feature, and the third image feature refers to a target region in the target image feature and each of a first region in the first image feature, a second region in the second image feature, and a third region in the third image feature having the same position in the target image feature, the first image feature, the second image feature, and the third image feature. For example, the first image feature, the second image feature, and the third image feature may be denoted as a first matrix, a second matrix, and a third matrix respectively, the target image feature may be denoted as a target matrix, the first matrix, the second matrix, the third matrix, and the target matrix have the same size, the corresponding target region, first region, the second region. and third region are located in the same position of the first matrix, the second matrix, the third matrix, and the target matrix.
[0155] For each of the multiple target regions, a target image sub-feature corresponding to the target region may be generated based on at least one of the first image sub-feature of a first region corresponding to the target region corresponding to the target region, the second image sub-feature of the second region corresponding to a target region, and the third image sub-feature of a third region corresponding to the target region.
[0156] In some embodiments, the processing device may designate one of the first image sub-feature, the second image sub-feature, and the third image sub-feature as the target image sub-feature of the target region.
[0157] For example, for each of the multiple target regions, the processing device may select one of the corresponding first image sub-feature, the second image sub-feature, and the third image sub-feature as the target image sub-feature according to a preset rule. The preset rule may be a random selection, i.e., for any one of the target regions, the processing device randomly selects one of the first image sub-feature, the second image sub-feature, and the third image sub-feature as the target image sub-feature of the target region. As another example, the processing device may set up a selection matrix for selecting one of the first image sub-feature, the second image sub-feature, or the third image sub-feature as the target image sub-feature of each target region . The selection matrix may include multiple elements each of which corresponds to one of the multiple target regions of the target image feature. An element in the selection matrix represents selecting which one of the corresponding first image sub-feature, the second image sub-feature, and the third image sub-feature as the target image sub-feature of a target region. The selection matrix may be a default setting of the system according to a first target confidence level and a second target confidence level (e.g., set according to a first condition, a second condition) , or according to a setting of the first image acquisition device and the second image acquisition device, or set by an operator according to experience. As a further example, if an element of the selection matrix corresponding to a target region is value 1, a first image sub-feature of a first region corresponding to the target region may be selected as the target image sub-feature of the target region. As still a further example, the processing device divides the first image feature, the second image feature, and the third image feature into 3×3 grids of first regions, second regions, and third regions, respectively; for the target region in the first column of the first row, the first image sub-feature in the first region is set as the target image sub-feature corresponding to the target region; for the target region in the second column of the first row, the third image sub-feature in the third region is set as the target image sub-feature corresponding to the target region; and for the target region in the third column of the first row, the second image sub-feature in the second region is set as the target image sub-feature corresponding to the target region. By analogy, the processing device also sets corresponding target image sub-features for the target regions at the other six positions, and when target detection is performed in the subsequent period, selects the corresponding target image sub-feature for the target region at each position according to the pre-set rule.
[0158] In some embodiments, the processing device, in response to determining that a first condition is satisfied, designates the first image sub-feature as the target image sub-feature of the target region; in response to determining that a second condition is satisfied, designates the second image sub-feature as the target image sub-feature of the target region; and in response to determining that the first condition and the second condition are not satisfied, designates the third image sub-feature as the target image sub-feature of the target region.
[0159] The first condition is configured to determine whether the first image sub-feature of each first region is capable of being used as the target image sub-feature of a target region corresponding to the first region.
[0160] For a first region and a second region corresponding to the first region, the first condition includes that the first target confidence level of the first region is greater than or equal to a first threshold and a second target confidence level of the second region corresponding to the first region is less than a second threshold.
[0161] The second condition is configured to determine whether the second image sub-feature of each second region is capable of being used as the target image sub-feature of a target region corresponding to the second region.
[0162] For a first region and a second region corresponding to the first region, the second condition includes that the first target confidence level of the first region is less than the first threshold, and the second target confidence level of the second region is greater than or equal to the second threshold.
[0163] The first threshold and the second threshold may be or not be equal, and the magnitude of the first threshold and the second threshold may be set according to the actual situation.
[0164] In some embodiments, the processing device also divides the first image and the second image according to the division pattern of the first image feature, the second image feature, and the third image feature. The processing device divides the first image into multiple first regions, and divides the second image into multiple second regions. Each of the first regions in the first image may correspond to a first region in the first image feature. Each of the second regions in the second image may correspond to a second region in the second image feature.
[0165] For example, the processing device divides the first image feature, the second image feature, and the third image feature all into N×M regions, and correspondingly, divides the first image into N×M first regions, and divides the second image into N×M second regions, wherein N is a count of rows and M is a count of columns.
[0166] The first target confidence level of a first region in the first image feature refers to a confidence level characterizing the presence of one or more specific targets in the first region of the first image. The higher the first target confidence level of the first region in the first image feature, the higher the probability of the first region (or a first detection box in the first region) of the first image including one or more specific targets, i.e., the first image sub-feature in the first region of the first image feature is more conducive to target detection. When there is no specific target in the first region of the first image, the first target confidence level is 0.
[0167] In some embodiments, the processing device obtains a first detection result of the first image based on the first image feature. The first detection result includes first position information of one or more specific targets in the first image and a first confidence level corresponding to each of the one or more specific targets. For example, each of the one or more specific targets in the first image may be marked with a first bounding box (also referred to as a first detection box) and the first position information may include a position of each of the first detection boxes in the first image. The first confidence level corresponding to each of the one or more specific targets refers to a probability that a predicted first detection box contains a specific target. The processing device determines the first target confidence level based on the first confidence level corresponding to each of the one or more specific targets and the first position information of the one or more specific targets.
[0168] The processing device performs target detection based on the first image feature to obtain the first detection boxes of the one or more specific targets and first confidence level of each of the first detection boxes, thereby obtaining the first detection result. In some embodiments, the first detection result may include only one first detection box and the first confidence level corresponding to the only one first detection box, or may include multiple first detection boxes and multiple first confidence levels corresponding to the multiple first detection boxes, depending on a count of the one or more specific targets present in the first region of the first image.
[0169] For determining a first target confidence level of a first region in the first image feature, the processing device may determine one or more first detection boxes in the first region of the first image based on the first position information of the first detection result and determine the first target confidence level based on the first confidence levels of the one or more first detection boxes in the first region of the first image. For example, the processing device may determine an average of the first confidence levels of the one or more first detection boxes in the first region of the first image as the first target confidence level of the first region. As another example, the processing device may determine the first target confidence level by other ways, e.g., the processing device may use a median, a plurality, a maximum, or a minimum of the first confidence levels of the first detection boxes in the first region of the first image, as the first target confidence level of the first region in the first image feature.
[0170] In some embodiments of the present disclosure, by performing the target detection based on the first image feature, the first target confidence level can be determined more accurately; at the same time, performing the target detection improves the speed of the image processing, which is conducive to guaranteeing the safety of the intelligent driving.
[0171] In some embodiments, in response to determining that a posture of a first image acquisition device acquiring the first image of a scene that is the same as a posture of the first image acquisition device acquiring at least one first historical image of the scene, the processing device may obtain at least one historical first detection result including at least one historical first confidence level of the at least one first historical image; and determine the first target confidence level based on at least one historical first confidence level. The first image and the at least one first historical image may be acquired by the first image acquisition device during a continuous time period. Each of the at least one historical first detection result includes one or more historical first confidence levels.
[0172] The posture refers to a position and an attitude of an image acquisition device in space. The position refers to a position of the image acquisition device in three-dimensional space, usually expressed in coordinates (x, y, z) , and the attitude refers to an orientation or rotational state of the image acquisition device, usually expressed in an angle.
[0173] The position may be determined in multiple ways, e.g., the position may be calibrated by a known calibration object (e.g., a checkerboard grid) , determined by a sensor, acquired by direct measurement, etc.
[0174] The first image and the at least one first historical image are acquired by the first image acquisition device during the continuous time period, and the posture of the first image acquisition device acquiring the first image that is the same as the posture of the first image acquisition device acquiring the at least one first historical image, means that the position and the attitude of the first image acquisition device remains unchanged during the continuous time period.
[0175] The first historical image refers to an image obtained based on historical data acquired by the first image acquisition device at a historical moment. The confidence level of the first region of the first historical image including one or more specific targets also reflects the confidence level of the first region of the first image including the one or more specific targets. The first target confidence level of the first region may thus be determined based on the first historical confidence level of the first region of the first historical image including the one or more specific targets.
[0176] The historical first detection result refers to a detection result for performing target detection on the first region in the first historical image.
[0177] For determining a first target confidence level of a first region in the first image feature, the processing device may determine one or more first detection boxes in the first region of each of the at least one first historical image based on first position information of each of the at least one historical first detection result and determine the first target confidence level based on the historical first confidence levels of the one or more first detection boxes in the first region of each of the at least one first historical image. For example, the first target confidence level of a first region in the first image feature may be an average of the historical first confidence levels of the one or more first detection boxes in the first region of each of the at least one first historical image. In some embodiments, the first target confidence level of a first region in the first image feature may be a median, a maximum or a minimum of the historical first confidence levels of the one or more first detection boxes in the first region of each of the at least one first historical image.
[0178] In some embodiments of the present disclosure, when the postures at the time of acquiring the images are consistent, the angle that the first image acquisition device acquires the images remains unchanged, and the first target confidence level can be determined directly by using the historical first confidence level to reduce the computation and to improve the efficiency of image processing.
[0179] The second target confidence level of a second region in the second image feature refers to a confidence level characterizing the presence of the one or more specific targets in the second region of the second image. The higher the second target confidence level of the second region in the second image feature, the higher the probability of the second region (or a second detection box in the second region) of the second image including the one or more specific targets, i.e., the second image sub-feature in the second region of the second image feature is more conductive to target detection. When there is no specific target in the second region of the second image, the second target confidence level is 0.
[0180] In some embodiments, the processing device obtains a second detection result of the second image based on the second image feature. The second detection result includes second position information of the one or more specific targets in the second image and a second confidence level corresponding to each of the one or more specific targets. For example, each of the one or more specific targets in the second image may be marked with a second bounding box (also referred to as a second detection box) and the second position information may include a position of each of the second detection boxes in the second image. The second confidence level corresponding to each of the one or more specific targets refers to a probability that a predicted second detection box contains a specific target. The processing device determines the second target confidence level based on the second confidence level corresponding to each of the one or more specific targets and the second position information of the one or more specific targets.
[0181] The second detection result is similar to the first detection result, the second position information is similar to the first position information, and the second confidence level is similar to the first confidence level, and the second detection result is obtained in the same manner as the first detection result, as may be found above and in the corresponding description.
[0182] In some embodiments, the second detection result may include only one second detection box and the second confidence level corresponding to the only one second detection box, or may include multiple second detection boxes and multiple second confidence levels corresponding to the multiple second detection boxes, depending on a count of the one or more specific targets present in the second region of the second image.
[0183] The way of determining the second target confidence level based on the second confidence level is similar to the way of determining the first target confidence level based on the first confidence level, and may be found above and in its corresponding description.
[0184] In some embodiments of the present disclosure, by performing the target detection on the second image feature, the second target confidence level can be determined more accurately; at the same time, performing the target detection improves the speed of the image processing, which is conducive to safeguarding the safety of the intelligent driving.
[0185] In some embodiments, in response to determining that a posture of a second image acquisition device acquiring the second image of a scene that is the same as a posture of the second image acquisition device acquiring at least one second historical image of the scene, the processing device may obtain at least one historical second detection result including at least one historical second confidence level of the at least one second historical image; and determine the second target confidence level based on the at least one historical second confidence level. The second image and the at least one second historical image may be acquired by the second image acquisition device during a continuous time period. Each of the at least one historical second detection result includes one or more historical second confidence levels.
[0186] The second image and the at least one second historical image are acquired by the second image acquisition device during the continuous time period, the posture of the second image acquisition device acquiring the second image that is the same as the posture of the second image acquisition device acquiring the at least one second historical image, means that the position, and the attitude of the second image acquisition device remains unchanged during the continuous time period.
[0187] The second historical image refers to an image obtained based on historical data acquired by the second image acquisition device at a historical moment. The confidence level of the second region of the second historical image including one or more specific targets also reflects the confidence level of the second region of the second image including the one or more specific targets. The second target confidence level may thus be determined based on the second historical confidence level of the second region of the second historical image including the one or more specific targets.
[0188] The historical second detection result refers to a detection result for performing target detection on the second region of the second historical image.
[0189] For determining a second target confidence level of a second region in the second image feature, the processing device may determine one or more second detection boxes in the second region of each of the at least one second historical image based on second position information of each of the at least one historical second detection result and determine the second target confidence level based on the historical second confidence levels of the one or more second detection boxes in the second region of each of the at least one second historical image. For example, the second target confidence level of a second region in the second image feature may be an average of the historical second confidence levels of the one or more second detection boxes in the second region of each of the at least one second historical image. In some embodiments, the second target confidence level of a second region in the second image feature may be a median, a maximum or a minimum of the historical second confidence levels of the one or more second detection boxes in the second region of each of the at least one second historical image.
[0190] In some embodiments, the processing device may determine the first detection result and a second detection result based on a first sub-model and a second sub-model in a target detection model, respectfully.
[0191] In some embodiments, the first sub-model further includes a first detection network, and the second sub-model further includes a second detection network. The first detection network refers to a model for performing detection on the first image sub-feature. The second detection network refers to a model for performing detection on the second image sub-feature.
[0192] More details regarding the first sub-model and the second sub-model may be found in the other contents of the present disclosure (e.g., description in connection with FIG. 7) .
[0193] In some embodiments of the present disclosure, when the postures at the time of acquiring the images are consistent, the angle that the second image acquisition device acquires the image remains unchanged, and the second target confidence level can be determined directly by using the historical second confidence level to reduce the computation and improve the efficiency of the image processing.
[0194] In some embodiments, when the first target confidence level of the first region of the first image feature is greater than or equal to the first threshold and the second target confidence level corresponding to the second region is less than the second threshold (i.e., when the first condition is satisfied) , it indicates that there is a higher probability of the presence of the one or more specific targets in the first region of the first image, and there is a lower probability of the presence of the one or more specific targets in the second region of the second image, so that the first image sub-feature in the first region includes more features of the one or more specific targets, and the second image sub-feature in the second region includes fewer features of the one or more specific targets, and thus the processing device may designate the first image sub-feature as the target image sub-feature of the target region.
[0195] In some embodiments, when the first target confidence level corresponding to the first region is less than the first threshold and the second target confidence level corresponding to the second region is greater than or equal to the second threshold (i.e., when the second condition is satisfied) , it indicates that there is a lower probability of the presence of the one or more specific targets in the first region of the first image, and there is a higher probability of the presence of the one or more specific targets in the second region of the second image, so that the first image sub-feature in the first region includes fewer features of the one or more specific targets and the second image sub-feature in the second region includes more features of the one or more specific targets, and thus the processing device may designate the second image sub-feature as the target image sub-feature of the target region.
[0196] In some embodiments, when neither the first condition nor the second condition is satisfied, it indicates that there is a lower probability of the presence of the one or more specific targets in the first region of the first image, and there is a lower probability of the presence of the one or more specific targets in the second region of the second image; or, there is a higher probability of the presence of the one or more specific targets in the first region of the first image, and there is a higher probability of the presence of the one or more specific targets in the second region of the second image; or, the probability of the presence of the one or more specific targets in the first region of the first image is the same as the probability of the presence of the one or more specific targets in the second region of the second image, and at this time, it is more reasonable to perform target detection based on the third image sub-feature determined by fusing the first image sub-feature and the second image sub-feature, so the processing device may designate the third image sub-feature as the target image sub-feature corresponding to the target region.
[0197] For example, the processing device may divide the first image feature, the second image feature, and the third image feature into N x M grids of first regions, second regions, and third regions, respectively, with N being a count of rows and M being a count of columns. For each first region of the multiple first regions, the first target confidence level corresponding to the first region may be represented by Avg_Radar_conij; for each second region of the multiple second regions, the second target confidence level corresponding to the second region may be represented by Avg_Vision_conij, and the processing device may, based on Equation (1) , determine a trustworthy feature class Regionij of the target region: wherein, i denotes the ith row, j denotes the jth column, i is an integer between 1 and N, and j is an integer between 1 and M; Regionij denotes the trustworthy feature class of a target region (ij) ; Avg_Radar_conij denotes the first target confidence level; Avg_Vision_conij denotes the second target confidence level; t1 denotes the first threshold; t2 denotes the second threshold; Regionij being equal to 0 denotes the first image sub-feature is credible, Regionij being equal to 1 denotes the second image sub-feature is credible, and Regionij being equal to 2 denotes the third image sub-feature is credible.
[0198] Merely by way of example, as shown in FIG. 6, the processing device may divide the first image feature, the second image feature and the third image feature into 3×3 grids of first regions, second regions and third regions, respectively; if Regionij of a target region (ij) is 0, then the target image sub-feature of the target region is the first image sub-feature, i.e., the target region is filled with dots; if Regionij of a target region (ij) is 1, then the target image sub-feature of the target region is the second image sub-feature, i.e., the target region is filled with diagonal lines; if Regionij of a target region (ij) is 2, the target image sub-feature of the target region is the third image sub-feature, i.e., the target region is filled with horizontal lines. Taking the target region in the first row and first column as an example, if the first row and first column of the trustworthy feature class 640 of the target region in the first row and first column is 2, the third image sub-feature is credible, and thus the third image sub-feature A3 is designated as the target image sub-feature A4 , i.e., the target region is filled with horizontal lines.
[0199] In some embodiments, the processing device may determine the target image sub-features in other ways. For example, the processing device may determine a third target confidence level based on the third image sub-feature, and generate the target image sub-feature based on the first image sub-feature, the second image sub-feature, the third image sub-feature, the first target confidence level, the second target confidence level and the third target confidence level. The approach of obtaining the third target confidence level is similar to the approaches of obtaining the first target confidence level and the second target confidence level, and may be found above and its corresponding description.
[0200] The processing device may compare the first target confidence level, the second target confidence level, and the third target confidence level with the first threshold, the second threshold, and a third threshold, respectively. When only one of the first target confidence level, the second target confidence level, and the third target confidence level is greater than its corresponding threshold (e.g., the first threshold, the second threshold, and a third threshold) , the processing device may designate the image sub-feature that is greater than its corresponding threshold as the target image sub-feature of the target region. Exemplarily, when the first target confidence level is greater than the first threshold, and the second target confidence level and the third target confidence level are both less than the corresponding thresholds, the processing device designates the first image sub-feature as the target image sub-feature of the target region. The magnitude of the third threshold and the first threshold and the second threshold may be equal or unequal, and the magnitude of the third threshold may be set according to the actual situation.
[0201] When two of the first target confidence level, the second target confidence level, and the third target confidence level are greater than the corresponding thresholds, the processing device may perform a weighted summation of the two image sub-features that are greater than the corresponding thresholds, and designate the weighted sum as the target image sub-feature of the target region. Exemplarily, when the first target confidence level is greater than the first threshold and the second target confidence level is greater than the second threshold, and the third target confidence level is less than the third threshold, the processing device designates a weighted sum of the first image sub-feature and the second image sub-feature as the target image sub-feature of the target region.
[0202] When the first target confidence level, the second target confidence level, and the third target confidence level are greater than the corresponding thresholds or less than the corresponding thresholds, the processing device may perform a weighted summation of the first image sub-feature, the second image sub-feature, and the third image sub-feature, and designate a weighted sum as the target image sub-feature of the target region.
[0203] In some embodiments, the processing device may set a weight of the third image sub-feature larger to make the weighted sum more accurate. The weights of the weighted summation are related to the quality of the data source. The quality of the data source depends on the image acquisition device used. Taking radar cameras and visible light cameras as an example, the visible light cameras have a higher imaging quality than the radar cameras when there is sufficient light.
[0204] In some embodiments of the present disclosure, the third target confidence level is also considered when determining the target image sub-features, which can make the determined target image sub-features more comprehensive and accurate, and further can improve the accuracy of the target detection model.
[0205] In some embodiments of the present disclosure, the process of determining the target image sub-features can be made more accurate based on the comparison of the first target confidence level corresponding to the first region and the second target confidence level corresponding to the second region with the first threshold and the second threshold; by the predetermined first condition and second condition, the accuracy of determining the target image sub-features can be further improved.
[0206] In some embodiments of the present disclosure, in the process of generating the target image features, for each of the multiple target regions, one of the first image feature, the second image feature, and the third image feature is selected to be the target image sub-feature corresponding to the target region, i.e., the target image feature combines the first image feature, the second image feature, and the third image feature at the same time, so that the target image feature can integrate the advantages of the first image feature, the second image feature, and the third image feature, and target detection based on the target image feature can ultimately ensure the accuracy of the detection.
[0207] It should be noted that the above description of the process 500 is merely provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For persons having ordinary skills in the art, multiple variations or modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure.
[0208] FIG. 7 is a schematic diagram illustrating a target detection model according to some embodiments of the present disclosure; and FIG. 8 is an exemplary flowchart illustrating a method for target detection implemented based on a target detection model according to some embodiments of the present disclosure.
[0209] To increase the speed of image processing, the processing device may perform target detection by the target detection model. In some embodiments, an exemplary process 800 for implementing the method for target detection based on the target detection model includes the following operations, as shown in FIG. 8. The process 800 may be performed by the processing device.
[0210] In 810, a first image and a second image may be acquired. Operation 810 may be found in the corresponding description of operation 410 in FIG. 4.
[0211] In 820, based on the first image, a first image feature may be generated. For example, the first image feature may be generated by using a first sub-model of a target detection model as described elsewhere in the present disclosure.
[0212] In 830, based on the second image, a second image feature may be generated. For example, the second image feature may be generated by using a second sub-model of the target detection model.
[0213] In 840, the first image feature may be divided into multiple first regions and the second image feature may be divided into multiple second regions. The division patterns of the first image feature and the second image feature are the same. Operation 840 may be found in the corresponding description of operation 430 in FIG. 4. For example, the second image feature and the first image feature may be divided by using a third sub-model of the target detection model.
[0214] In 850, for each of multiple target regions, based on a first image sub-feature and a second image sub-feature, a target image sub-feature corresponding to the target region. The third image feature may be generated by using the third sub-model of the target detection model. For example, the third sub-model may generate a third image feature by fusing the first image feature and the second image feature, determine a first target confidence level of each of the first regions and a second target confidence level of each of the second regions, and determine the target image sub-feature corresponding to each target region based on the third image sub-feature, the first image sub-feature, and the second image sub-feature.
[0215] In 860, a target detection result may be generated based on the target image sub-features of the multiple target regions. For example, the third sub-model may be used to generate the target detection result.
[0216] More details regarding the target detection model, the first sub-model, the second sub-model, and the third sub-model may be found below.
[0217] In some embodiments, the target detection model may include a first sub-model, a second sub-model, and a third sub-model. The first sub-model may be referred to as a first detection sub-model and the second sub-model may be referred to as a second detection sub-model.
[0218] In some embodiments, the first sub-model 730 may include a first feature extraction network 731 and a first detection network 732. More details regarding the first sub-model may be found in other contents of the present disclosure (e.g., description in connection with FIG. 4) .
[0219] In some embodiments, the processing device may input the first image 710 into the first sub-model 730 to output the first image feature, as shown in FIG. 7. After inputting the first image 710 into the first sub-model 730, the first feature extraction network 731 performs feature extraction on the first image 710 to obtain the first image feature; and after receiving the first image feature output by the first feature extraction network 731, the first detection network 732 performs target detection on the first image feature to obtain a first detection result of the first image. The first detection result of the first image may include first position information of one or more specific targets (or subjects) in the first image and a first confidence level corresponding to each of the one or more specific targets (or subjects) .
[0220] In some embodiments, the second sub-model 740 may include a second feature extraction network 741 and a second detection network 742. More details regarding the second sub-model may be found in other contents of the present disclosure (e.g., description in connection with FIG. 4) .
[0221] In some embodiments, the processing device may input the second image 720 into the second sub-model 740 to output the second image feature, as shown in FIG. 7. After inputting the second image 720 into the second sub-model 740, the second feature extraction network 741 performs feature extraction on the second image 720 to obtain the second image feature; and after receiving the second image feature output by the second feature extraction network 741, the second detection network 742 performs target detection on the second image feature to obtain a second detection result of the second image. The second detection result of the second image may include second position information of one or more specific targets (or subjects) in the second image and a second confidence level corresponding to each of the one or more specific targets (or subjects) .
[0222] In some embodiments, the first sub-model and the second sub-model may be integrated into one single sub-model for processing the first image and the second image. In other words, one of the first sub-model or the second sub-model is included in the target detection model to process both the first image and the second image , and a training process of the only one of the first sub-model or the second sub-model may utilize first sample images and second sample images to train the one of the first sub-model or the second sub-model.
[0223] In some embodiments, as shown in FIG. 7, the third sub-model may include a fusion unit 751, a region-adaptive unit 752, and a detection unit 753. The fusion unit 751 may be an attention mechanism module such as an SE module, a CBAM module, etc. The region-adaptive unit 752 may include a CNN, an RPN, etc. The detection unit 753 may include a CNN, an RPN, etc.
[0224] The first image feature outputted by the first feature extraction network 731 and the second image feature output by the second feature extraction network 741 may also be inputted into the fusion unit 751 for example, by the first feature extraction network 731 (or the first detection network 732) and the second feature extraction network 741 (or the second detection network 742) .
[0225] The fusion unit 751 fuse the first image feature and the second image feature to generate and output a third image feature. For example, the fusion unit 751 may fuse the first image feature and the second image feature by stitching the first image feature and the second image feature in a channel dimension to obtain a stitched image feature, and then feed the stitched image feature into an attention mechanism module of the fusion unit such as a squeeze-and-excitation attention (SE) module, a convolutional block attention module (CBAM) , etc., to obtain the third image feature.
[0226] The first image feature, the second image feature, and the third image feature may be inputted into the region-adaptive unit 752 to generate and output a target image feature. For example, the region-adaptive unit 752 may determine the target image feature based on the first image feature, the second image feature, the third image feature, the first detection result, and the second detection result. The region-adaptive unit 752 may determine a first target confidence level of each first region of the first image feature based on the first detection result and determine a second target confidence level of each second region of the second image feature based on the second detection result. The region-adaptive unit 752 may determine the target image feature based on the first image feature, the second image feature, the third image feature, the first target confidence level, and the second target confidence level. The region-adaptive unit 752 may determine whether at least one of a first condition or a second condition is satisfied and determine a target image sub-feature of each target region from one of a first image sub-feature, a second image sub-feature, and a third image-feature corresponding to the target region. More descriptions for determining the target image feature may be found elsewhere in the present disclosure.
[0227] The target image feature output by the region-adaptive unit 752 may be input into the detection unit 753 to generate and output a target detection result. It is to be understood that in the present embodiment, by fusing the first image feature and the second image feature, not directly fusing the first image sub-feature and the second sub-feature, image sub-features of the surrounding regions of each first region and each second region are considered to be used to determine the fused feature (i.e., the third image feature) of the corresponding third region, the richness and accuracy of the features are higher, and the accuracy of the region-adaptive unit is higher.
[0228] In some embodiments, the third sub-model may include the region-adaptive unit and the detection unit. The region-adaptive unit and the detection unit may be a CNN, an RPN, or the like.
[0229] The target detection model may be obtained by training in one of multiple ways. For example, the first sub-model, the second sub-model, and the third sub-model are obtained by joint training.
[0230] Merely by way of example, a sample set may be obtained. The sample set may include multiple groups of first sample images and second sample images. Each group may include a first sample image and a second sample image represent the same scene. Each of at least a portion of the first sample images includes a first label and a target label. The first label includes first position information and a first confidence level of each specific target in the first sample image. The first label includes target position information and a target confidence level of each specific target in the first sample image. Each of at least a portion of the second sample images includes a second label and a target label. The second label includes target position information and a target confidence level of each specific target in the second sample image. The training of the target detection model may be performed by performing multiple iterations. In each iteration, each first sample image may be inputted into an initial first sub-model to obtain an output of the initial first sub-model. The output of the initial first sub-model may include a first detection result and / or first image data corresponding to each first sample image. Each second sample image may be inputted into an initial second sub-model to obtain an output of the initial second sub-model. The output of the initial second sub-model may include a second detection result and / or second image data corresponding to each second sample image. The output of the initial first sub-model and the output of the initial second sub-model may be inputted into an initial third sub-model to obtain an output of the initial third sub-model. The output of the initial third sub-model may include an estimated detection result (including positions of one or more subjects represented in the first sample image and / or the second sample image and / or an estimated confidence level of each of the one or more subjects) . A first loss function may be constructed based on the output of the initial first sub-model and the corresponding first label, a second loss function may be constructed based on the output of the initial second sub-model and the corresponding second label, and a third loss function may be constructed based on the output of the initial third sub-model and the corresponding third label. Parameters of the initial first sub-model, the initial second sub-model, and the initial third sub-model may be updated based on the first loss function, the second loss function, and the third loss function in response to a termination condition is not satisfied. The termination condition includes the first loss function, the second loss function, and the third loss function converging, a count of iterations reaching a threshold, or the like.
[0231] It is to be understood that joint training requires the first detection network, the second detection network, and the detection unit to be interconnected with each other, and therefore the model is trained with better accuracy.
[0232] As another example, the target detection model may be obtained by training each sub-model independently.
[0233] For example, the processing device trains the initial first sub-model based on the first sample images to obtain the first sub-model; train the initial second sub-model based on the second sample images to obtain the second sub-model; and train the third sub-model based on the outputs of the first sub-model and the second sub-model generating in the training of the first sub-model and the second sub-model. During the training of the third sub-model, parameters of the first sub-model and parameters of the second sub-model are unchanged.
[0234] It should be noted that in the process of training, the processing device also divides the sample set into a training set and a validation set. The training set refers to a sample set used to train the target detection model; and the validation set refers to a sample set used to evaluate the performance of the target detection model.
[0235] In some embodiments, as shown in FIG. 7, the processing device trains the first sub-model 730 and the second sub-model 740 using the sample set, and freezes the parameters of the trained first sub-model 730 and the trained second sub-model 740, i.e., the first sub-model 730 and the second sub-model 740 may not be trained again in the subsequent training process of the third sub-model. Freezing means keeping the parameters of the first sub-model and the parameters of the second sub-model unchanged during the training of the third sub-model. For example, the processing device saves the weight of the first sub-model and the weight of the second sub-model with the highest accuracy determined based on the validation set.
[0236] The first sub-model and the second sub-model may be trained synchronously or not. The training of the first sub-model and the second sub-model may include multiple rounds. Each round may include multiple iterations using at least a portion of the multiple groups of first sample images and second samples images. In each iteration, each first sample image may be inputted into an initial first sub-model to obtain an output of the initial first sub-model. The output of the initial first sub-model may include a first detection result and / or first image data corresponding to each first sample image. Each second sample image may be inputted into an initial second sub-model to obtain an output of the initial second sub-model. The output of the initial second sub-model may include a second detection result and / or second image data corresponding to each second sample image. A first loss function may be constructed based on the output of the initial first sub-model and the corresponding first label, a second loss function may be constructed based on the output of the initial second sub-model and the corresponding second label, and a third loss function may be constructed based on the output of the initial third sub-model and the corresponding third label. Parameters of the initial first sub-model may be updated based on the first loss function in response to determining that a first termination condition is not satisfied, and parameters of the initial second sub-model may be updated based on the second loss function in response to a second termination condition is not satisfied. The first termination condition includes the first loss function converging, a count of iterations reaching a threshold, or the like. The second termination condition includes the second loss function converging, a count of iterations reaching a threshold, or the like. For each iteration, an output (i.e., the first detection result) of the initial first sub-model and an output (i.e., the second detection result) of the initial second sub-model generated based on the first sample image and the second sample image in the same group may be stored into a same set (also referred to as an output set) . The first image feature and the second image feature extracted from the first sample image and the second sample image in the same group may be stored into the same set (also referred to as a feature set) . The output set and the feature set may be used as a sample for training of the third sub-model and the target label may be used as a label of the sample.
[0237] In some embodiments, the validation set may be used to determine the performance of the trained initial first sub-model and the trained initial second sub-model generated in each round. A trained initial first sub-model and a trained initial second sub-model generated in a round with the best performance may be determined as the final first sub-model and the final second sub-model. The trained initial first sub-model and the trained initial second sub-model generated in each round may be independently evaluated. In other words, the final first sub-model and the final second sub-model may be derived from different rounds.
[0238] In each round of the multiple rounds of training, the processing device saves the parameters of the corresponding trained sub-model; validates the trained sub-model in each round by the validation set to determine the model accuracy; and then selects the sub-model with the highest model accuracy as the final trained sub-model. Taking the first sub-model as an example, assuming that the training set includes one thousand first sample images, the processing device may divide the training set into ten subsets, each of which includes one hundred first sample images, conduct ten rounds of training on the first sub-model, and save the parameters of the corresponding trained first sub-model in each round of training. Then the processing device may validate each round of the trained first sub-model by using the validation set to obtain the respective model accuracy; select the first sub-model with the highest model accuracy as the final trained first sub-model, and retain the weight of the final trained first sub-model. The training process of the second sub-model by using the training set and the validation set is similar to the training process of the second sub-model by using the training set and the validation set, with the difference being that the training set includes one thousand second sample images. The first sample images correspond one-to-one with the second sample images, as may be found in the previous description. Merely by way of example, if ten rounds of training are both performed, and the first sub-model obtained from round 6 of the ten rounds of training has the highest accuracy, and the second sub-model obtained from round 8 of the ten rounds of training has the highest accuracy, then the weight of the first sub-model in round 6 and the weight of the second sub-model in round 8 are retained.
[0239] In some embodiments, the processing device may keep the parameters of the first sub-model and the parameters of the second sub-model unchanged, input the first sample images in the sample set into the first sub-model to determine the first sample image features and the first sample detection result, and input the second sample images in the sample set into the second sub-model to determine the second sample image features and the second sample detection result; and then, train the third sub-model based on the first sample image features, the first sample detection result, the second sample image features, the second sample detection result, and the target labels corresponding to the sample set.
[0240] For example, in each iteration of the training of the third sub-model, the processing device may input the first sample image feature, the first sample detection result, the second sample image feature, and the second sample detection result generated based on the first sample image and the second sample image in the same group into the initial third sub-model to obtain the output of the initial third sub-model; construct a third loss function based on the output of the initial third sub-model and the target label; update the parameters of the initial third sub-model based on the third loss function; and until a third termination condition is satisfied, complete the training and obtain the trained third sub-model. The third termination condition includes the third loss function converging, a count of iterations reaching a threshold, or the like.
[0241] It is to be understood that the processing device may construct the third loss function, based on a similarity between the target position information of the target detection result and the sample position information of the sample detection result, and a difference between the target confidence level of the target detection result and the sample confidence level of the sample detection result.
[0242] More details regarding the first image, the second image, the first image feature, the second image feature, the target detection result, the first image sub-features, the second image sub-features, the first regions, the second regions, the third regions, the target regions, the target image sub-features and the target detection result may be found in other contents of the present disclosure (e.g., descriptions in connection with FIG. 4 -FIG. 6) .
[0243] In some embodiments of the present disclosure, it is relatively easy to optimize the target detection model by training the model independently, as the loss proportions of the three sub-models are unlikely to be identical, and the focus needs to be on considering the loss of the third sub-model.
[0244] In some embodiments, the inputs of the third sub-model may include multiple types. For example, the inputs of the third sub-model may include the first image sub-features, the first target confidence level, the second image sub-features, and the second target confidence level; as another example, the inputs of the third sub-model may include the first image sub-features, the second image sub-features, and a trustworthy feature class determined based on the first target confidence level and the second target confidence level. When the inputs of the third sub-model are transformed, corresponding transformations are also required when the third sub-model is trained alone.
[0245] In some embodiments of the present disclosure, by using a multi-layered target detection model, the data can be processed in greater detail, enabling the capture of finer visual details; the layered structure of the model helps subsequent debugging and testing to locate the model in a timely manner; at the same time, the layered target detection model is more adaptable and can be applied to a wide range of scenarios, and also effectively reduces the time wasted in the target detection process.
[0246] It should be noted that the above description is merely provided for the purposes of illustration, and not intended to limit the scope of the present disclosure. For persons having ordinary skills in the art, multiple variations or modifications may be made under the teachings of the present disclosure. However, those variations and modifications do not depart from the scope of the present disclosure.
[0247] FIG. 9 is an exemplary flowchart illustrating a method for target detection according to some other embodiments of the present disclosure. In some embodiments, process 900 includes the following operations, as shown in FIG. 9. The process 900 may be performed by the processing device.
[0248] In 910, a first image and a second image may be acquired. Operation 910 may be found in the corresponding description of operation 410 in FIG. 4.
[0249] In 920, the first image may be divided into multiple first regions and the second image may be divided into multiple second regions. The division patterns of the first image and the second image are the same. More details regarding dividing the first image and the second image may be found in other contents of the present disclosure (e.g., description in connection with FIG. 5) .
[0250] In 930, a first image feature is extracted from the first image and a second image is extracted from the second image. The first image feature may include multiple first image sub-features each of which corresponds to one of the first regions of the first image and the second image feature may include multiple second image sub-features each of which corresponds to one of the second regions of the second image.
[0251] In some embodiments, the first image feature may be obtained by performing feature extraction on the multiple first regions and the second image feature may be obtained by performing feature extraction on the multiple second regions. In other words, feature extraction is performed on each first region in the multiple first sub-images and each second region in the multiple second sub-images respectively to obtain the first image sub-feature for each first region and the second image sub-feature for each second region. The approach of performing the feature extraction on each first region to obtain the corresponding first image sub-feature and performing feature extraction on each second region to obtain the corresponding second image sub-feature is similar to the approach of performing feature extraction on the first image to obtain the first image features, which may be found in the corresponding description above.
[0252] In 940, for each of multiple target regions, based on at least one of the first image sub-feature corresponding to the first region and the second image sub-feature corresponding to the second region, a target image sub-feature corresponding to the target region may be generated.
[0253] In 950, based on the target image sub-features of the multiple target regions, a target detection result may be generated.
[0254] Operation 940 and operation 950 may be found in the corresponding descriptions of operation 440 and operation 450 in FIG. 4.
[0255] In some embodiments of the present disclosure, the applicability of the method for target detection can be effectively expanded by adopting the approach of dividing the region first and then performing feature extraction.
[0256] FIG. 10 is an exemplary flowchart illustrating a method for target detection according to some other embodiments of the present disclosure. In some embodiments, process 1000 includes the following operations, as shown in FIG. 10. The process 1000 may be performed by the processing device.
[0257] In 1010, a first image and a second image may be acquired. Operation 1010 may be found in the corresponding description of operation 410 in FIG. 4.
[0258] In 1020, a first image feature may be generated by performing feature extraction on the first image. Operation 1020 may be found in the corresponding description of operation 820 in FIG. 8.
[0259] In 1030, a second image feature may be generated by performing feature extraction on the second image. Operation 1030 may be found in the corresponding description of operation 830 in FIG. 8.
[0260] In 1040, a third image feature may be generated based on the first image feature and the second image feature. More details regarding generating the third image feature may be found in other contents of the present disclosure (e.g., description in connection with FIG. 5) . The process of generating the third image feature based on the first image feature and the second image feature by the third sub-model of the target detection model is similar to the process of generating the third image sub-feature based on a first sub-image feature and a second sub-image feature by the fusion unit of the third sub-model of the target detection model, and may be found in other contents of the present disclosure (e.g., description in connection with FIG. 8) .
[0261] In 1050, a target image feature may be generated based on the first image feature, the second image feature, and the third image feature. More details regarding the target image feature may be found in other contents of the present disclosure (e.g., description in connection with FIG. 5) .
[0262] In 1060, a target detection result may be generated based on the target image feature. Operation 1060 may be found in the corresponding description of operation 860 in FIG. 8.
[0263] In some embodiments, as shown in FIG. 11, some embodiments of the present disclosure further provide a device for target detection 1110 including a memory 1120 and a processor 1130 which are interconnected with each other, the memory 1120 being used to storing computer programs, when executed by the processor 1130, the computer programs being used to implement a method in any of the above embodiments. The memory 1120 may be a storage device in FIG. 1, and the processor 1130 may be the processing device 110 in FIG. 1.
[0264] In some embodiments, the device for target detection 1110 is capable of receiving data captured by a first image acquisition device and data captured by a second image acquisition device, and processing the data captured by the first image acquisition device to obtain a first image, and process the data acquired by the second image acquisition device to obtain a second image. In some embodiments, the device for target detection 1110 is independent of both the first image acquisition device and the second image acquisition device. In other embodiments, the device for target detection 1110 and at least one of the first image acquisition device, and the second image acquisition device may be simultaneously integrated on the same device.
[0265] In some embodiments, as shown in FIG. 12, some embodiments of the present disclosure further provide a non-transitory computer-readable storage medium 1210, the non-transitory computer-readable storage medium 1210 stores computer programs 1220, and the computer programs 1220 are capable of being executed by a processor to implement operations as described in any of the method embodiments described above.
[0266] In some embodiments, after the computer reads the computer instructions in the non-transitory computer-readable storage medium 1210, the computer performs a method as following: acquiring a first image and a second image, the first image and the second image representing a same scene at a same time point; generating a first image feature by performing feature extraction on the first image, the first image feature including multiple first image sub-features of multiple first regions; generating a second image feature by performing feature extraction on the second image, the second image feature including multiple second image sub-features of multiple second regions, each of the multiple first regions corresponding to one of the multiple second region; determining a target image feature based on the first image feature and the second image feature, the target image feature including multiple target regions, a target image sub-feature for each of the multiple target regions being determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region; and generating a target detection result based on the target image feature.
[0267] The non-transitory computer-readable storage medium 1210 may specifically be a device that may store the computer programs 1220 such as a USB flash drive, a removable hard disk, a read-only memory (ROM) , a random access memory (RAM) , a magnetic disk, or an optical disk, or may also be a server that has the stored computer programs 1220, the server may send the stored computer programs 1220 to other devices to run, or may also self-run the stored computer programs 1220.
[0268] Beneficial effects that may be brought about by the embodiments of the present disclosure include, but are not limited to: (1) by dividing the first image feature and the second image feature into multiple regions respectively, and based on the target image sub-features of the multiple regions, accurate target detection results can be generated, so as to make the effect of the post-target detection not lower than that of the detection effect of a single image acquisition device in all the regions; (2) based on the comparison between the first target confidence level and the second target confidence level corresponding to regions, and the first threshold and the second threshold, it can make the process of determining target image sub-features more precise; and by the first condition and the second condition that are preset, it is possible to further improve the accuracy of identifying the target image sub-features; (3) through the use of a multilayered target detection model, the data can be processed in more detail so that finer visual details can be captured; the hierarchical structure of the model helps subsequent debugging and testing to locate problems with the model in a timely manner; at the same time, the hierarchical target detection model is more adaptable and can be applied to multiple scenarios, and it can also effectively reduce the waste of time in the target detection process.
[0269] The basic concepts have been described above, and it is apparent to those skilled in the art that the foregoing detailed disclosure serves only as an example and does not constitute a limitation of the present disclosure. While not expressly stated herein, a person skilled in the art may make various modifications, improvements, and amendments to the present disclosure. Those types of modifications, improvements, and amendments are suggested in the present disclosure, so those types of modifications, improvements, and amendments are still within the spirit and scope of the exemplary embodiments of the present disclosure.
[0270] Also, the specification uses specific words to describe embodiments of the specification. such as "an embodiment" , "one embodiment" , and / or "some embodiment" means a feature, structure, or characteristic associated with at least one embodiment of the present disclosure. Accordingly, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" referred to two or more times in different positions in the present disclosure do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics of one or more embodiments of the present disclosure may be suitably combined.
[0271] Furthermore, unless expressly stated in the claims, the order of the processing elements and sequences, the use of numerical letters, or the use of other names as described in the present disclosure are not intended to qualify the order of the processes and methods of the present disclosure. While some embodiments of the invention that are currently considered useful are discussed in the foregoing disclosure by way of various examples, it should be appreciated that such details serve only illustrative purposes, and that additional claims are not limited to the disclosed embodiments, rather, the claims are intended to cover all amendments and equivalent combinations that are consistent with the substance and scope of the embodiments of the present disclosure. For example, although the implementation of various components described above may be embodied in a hardware device, it may also be implemented as a software only solution, e.g., an installation on an existing server or mobile device.
[0272] Similarly, it should be noted that in order to simplify the presentation of the disclosure of the specification, and thereby aid in the understanding of one or more embodiments of the invention, the foregoing descriptions of embodiments of the specification sometimes combine a variety of features into a single embodiment, accompanying drawings, or descriptions thereof. However, this method of disclosure does not imply that the objects of the present disclosure require more features than those mentioned in the claims. Rather, claimed subject matter may lie in less than all features of a single foregoing disclosed embodiment.
[0273] Some embodiments use numbers to describe the number of components and attributes, and it should be understood that such numbers used in the description of the embodiments are modified in some examples by the modifiers "approximately" , or "substantially" , or "proximately" is used in some examples. Unless otherwise noted, the terms "approximately, " "substantially" , or "proximately" indicates that a ±20%variation in the stated number is allowed. Correspondingly, in some embodiments, the numerical parameters used in the specification and claims are approximations, which can change depending on the desired characteristics of individual embodiments. In some embodiments, the numerical parameters should take into account the specified number of valid digits and employ general place-keeping. While the numerical domains and parameters used to confirm the breadth of their ranges in some embodiments of the present disclosure are approximations, in specific embodiments such values are set to be as precise as possible within a feasible range.
[0274] For each of the patents, patent applications, patent application disclosures, and other materials cited in the present disclosure, such as articles, books, specification sheets, publications, documents, and the like, are hereby incorporated by reference in their entirety into the present disclosure. Application history documents that are inconsistent with or conflict with the contents of the present disclosure are excluded, as are documents (currently or hereafter appended to the present disclosure) that limit the broadest scope of the claims of the present disclosure. It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or use of terms in the materials appended to the present disclosure and those set forth herein, the descriptions, definitions and / or use of terms in the present disclosure shall control.
[0275] Finally, it should be understood that the embodiments described in the present disclosure are only used to illustrate the principles of the embodiments of the present disclosure. Other deformations may also fall within the scope of the present disclosure. As such, alternative configurations of embodiments of the present disclosure may be viewed as consistent with the teachings of the present disclosure as an example, not as a limitation. Correspondingly, the embodiments of the present disclosure are not limited to the embodiments expressly presented and described herein
Claims
1.A method for target detection, comprising:acquiring a first image and a second image, the first image and the second image representing a same scene at a same time point;generating a first image feature by performing feature extraction on the first image, the first image feature including multiple first image sub-features of multiple first regions;generating a second image feature by performing feature extraction on the second image, the second image feature including multiple second image sub-features of multiple second regions, each of the multiple first regions corresponding to one of the multiple second regions;determining a target image feature based on the first image feature and the second image feature, the target image feature including multiple target regions, a target image sub-feature for each of the multiple target regions being determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region; andgenerating a target detection result based on the target image feature.2.The method of claim 1, wherein the determining a target image feature based on the first image feature and the second image feature includes:generating a third image feature based on the first image feature and the second image feature, the third image feature including multiple third image sub-features of multiple third regions, and each of the multiple third regions corresponding to one of the multiple first regions, one of the multiple second regions, and one of the multiple target regions; andfor each of the multiple target regions, generating the target image sub-feature corresponding to the target region based on at least one of the first image sub-feature of the first region corresponding to the target region, the second image sub-feature of the second region corresponding to the target region, and the third image sub-feature of the third region corresponding to the target region.3.The method of claim 2, wherein the generating the target image sub-feature corresponding to the target region based on at least one of the first image sub-feature of the first region corresponding to the target region, the second image sub-feature of the second region corresponding to the target region, and the third image sub-feature of the third region corresponding to the target region includes:designating one of the first image sub-feature, the second image sub-feature, and the third image sub-feature as the target image sub-feature of the target region.4.The method of claim 3, wherein the designating one of the first image sub-feature, the second image sub-feature, and the third image sub-feature as the target image sub-feature of the target region includes:in response to determining that a first condition is satisfied, designating the first image sub-feature as the target image sub-feature of the target region, the first condition including: a first target confidence level corresponding to the first region being greater than or equal to a first threshold, and a second target confidence level corresponding to the second region being less than a second threshold;in response to determining that a second condition is satisfied, designating the second image sub-feature as the target image sub-feature of the target region, the second condition including: the first target confidence level corresponding to the first region being less than the first threshold, and the second target confidence level corresponding to the second region being greater than or equal to the second threshold; andin response to determining that the first condition and the second condition are not satisfied, designating the third image sub-feature as the target image sub-feature of the target region.5.The method of claim 4, wherein the first target confidence level is determined according to operations including:obtaining a first detection result of the first image by performing a target detection based on the first image feature, wherein the first detection result includes first position information of one or more specific targets in the first image and a first confidence level corresponding to each of the one or more specific targets; anddetermining the first target confidence level based on the first confidence level and the first position information.6.The method of claim 4 or 5, wherein the second target confidence level is determined according to operations including:obtaining a second detection result of the second image by performing a target detection based on the second image feature, wherein the second detection result includes second position information of one or more specific targets in the second image and a second confidence level corresponding to each of the one or more specific targets; anddetermining the second target confidence level based on the second confidence level and the second position information.7.The method of claim 4 or 6, wherein the first target confidence level is determined according to operations including:obtaining at least one historical first detection result of at least one first historical image, each of the at least one historical first detection result including at least one historical first confidence level; anddetermining the first target confidence level based on the at least one historical first confidence level, the first image and the at least one first historical image being acquired by a first image acquisition device during a continuous time period.8.The method of claim 7, further comprising:in response to determining that a posture of the first image acquisition device acquiring the first image of a scene that is the same as a posture of the first image acquisition device acquiring the at least one first historical image of the scene, determining, based on the historical first confidence level, the first target confidence level.9.The method of claim 4 or 5, wherein the second target confidence level is determined according to operations including:obtaining at least one historical second detection result of at least one second historical image, each of the at least one historical second detection result including at least one historical second confidence level; anddetermining the second target confidence level based on at least one historical second confidence level, the second image and the at least one second historical image being acquired by a second image acquisition device during a continuous time period.10.The method of claim 9, further comprising:in response to determining that a posture of the second image acquisition device acquiring the second image of a scene that is the same as a posture of the second image acquisition device acquiring the at least one second historical image of the scene, determining, based on the historical second confidence level, the second target confidence level.11.The method of claims 1-10, wherein the generating a first image feature by performing feature extraction on the first image, and the generating a second image feature by performing feature extraction on the second image include:generating, based on the first image, the first image feature by using a first sub-model in a target detection model;generating, based on the second image, the second image feature by using a second sub-model in the target detection model; andthe determining a target image feature based on the first image feature and the second image feature includes:generating, based on the first image feature and the second image feature, the third image feature by using a third sub-model of the target detection model;generating the target image feature based on the first image feature, the second image feature, and the third image feature; andthe generating a target detection result of a target to be detected based on the target image feature includes:generating the target detection result of the target to be detected, based on the target image feature, by the third sub-model.12.The method of claim 11, a training process of the target detection model including:obtaining a sample set including multiple groups of first sample images and second sample images, each group of the multiple groups of first sample images and second sample images representing a same scene;training an initial first sub-model based on the multiple groups of first sample images to obtain the first sub-model;training an initial second sub-model based on the multiple groups of second sample images to obtain the second sub-model; andtraining the third sub-model based on the sample set, during the training process of the third sub-model, parameters of the first sub-model and parameters of the second sub-model are unchanged.13.The method of claim 1, wherein at least one of the first image and the second image is acquired based on radar data, and the acquiring a first image and a second image includes:converting the radar data into a grayscale image to acquire at least one of the first image and the second image.14.A system for target detection, comprising:an acquisition module configured to acquire a first image and a second image, the first image and the second image representing a same scene at a same time point;a first generation module configured to generate a first image feature by performing feature extraction on the first image, the first image feature including multiple first image sub-features of multiple first regions;a second generation module configured to generate a second image feature by performing feature extraction on the second image, the second image feature including multiple second image sub-features of multiple second regions, each of the multiple first regions corresponds to one of the multiple second regions;a determination module configured to determine a target image feature based on the first image feature and the second image feature, the target image feature including multiple target regions, a target image sub-feature of each of the multiple target regions being determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region; anda third generation module configured to generate a target detection result of a target to be detected based on the target image feature.15.The system of claim 14, wherein the determination module is further configured to:generate a third image feature based on the first image feature and the second image feature, the third image feature including multiple third image sub-features of multiple third regions, each of the multiple third regions corresponding to one of the multiple first regions, one of the multiple second regions, and one of the multiple target regions; andfor each of the multiple target regions, generate the target image sub-feature corresponding to the target region based on at least one of the first image sub-feature of the first region corresponding to the target region, the second image sub-feature of the second region corresponding to the target region, and the third image sub-feature of the third region corresponding to the target region.16.The system of claim 15, wherein the determination module is further configured to:designate one of the first image sub-feature, the second image sub-feature, and the third image sub-feature as the target image sub-feature of the target region.17.The system of claim 16, wherein the determination module is further configured to:in response to determining that a first condition is satisfied, designate the first image sub-feature as the target image sub-feature of the target region, the first condition including: a first target confidence level corresponding to the first region being greater than or equal to a first threshold, and a second target confidence level corresponding to the second region being less than a second threshold;in response to determining that a second condition is satisfied, designate the second image sub-feature as the target image sub-feature of the target region, the second condition including: the first target confidence level corresponding to the first region being less than the first threshold, and the second target confidence level corresponding to the second region being greater than or equal to the second threshold; andin response to determining that the first condition and the second condition are not satisfied, designate the third image sub-feature as the target image sub-feature of the target region.18.The system of claim 17, wherein the determination module is further configured to:obtain a first detection result of the first image by performing a target detection based on the first image feature, wherein the first detection result includes first position information of one or more specific targets in the first image and a first confidence level corresponding to each of the one or more specific targets; anddetermine the first target confidence level based on the first confidence level and the first position information.19.The system of claim 18, wherein the determination module is further configured to:obtain a second detection result of the second image by performing a target detection based on the second image feature, wherein the second detection result includes second position information of one or more specific targets in the second image and a second confidence level corresponding to each of the one or more specific targets; anddetermine the second target confidence level based on the second confidence level and the second position information.20.A non-transitory computer-readable storage medium storing computer instructions, and when a computer reads the computer instructions in the storage medium, the computer performs the following method:acquiring a first image and a second image, the first image and the second image representing a same scene at a same time point;generating a first image feature by performing feature extraction on the first image, the first image feature including multiple first image sub-features of multiple first regions;generating a second image feature by performing feature extraction on the second image, the second image feature including multiple second image sub-features of multiple second regions, each of the multiple first regions corresponds to one of the multiple second regions;determining a target image feature based on the first image feature and the second image feature, the target image feature including multiple target regions, a target image sub-feature for each of the multiple target regions being determined based on at least one of the first image sub-feature of the first region corresponding to the target region and the second image sub-feature of the second region corresponding to the target region; andgenerating a target detection result based on the target image feature.
Citation Information
Patent Citations
Image processing method and device
CN110378952A
Target detection method and device, model training method and device, equipment and storage medium
CN115909253A
Tunneling working face target detection method based on multi-sensor data fusion
CN118425955A
Target detection method, target detection device and computer readable storage medium
CN118918319A
Systems and methods for medical diagnosis
US20210304896A1