Target detection method, model training method, device, equipment, vehicle and medium
The event camera captures the event flow of light changes and combines spatial information to extract space-time features, which solves the problem of the autonomous driving perception system identifying obstacles in harsh environments, and realizes accurate detection and recognition of occluded dynamic targets.
Patent Information
- Application Number
- CN202510094268.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult for autonomous driving to accurately identify pedestrians, vehicles or other obstacles in severe weather, night driving or complex urban environments, increasing the risk of traffic accidents.
A target detection method is adopted to capture the event stream of light changes in the driving scene through the event camera, perform spatial and temporal feature extraction, and conduct detection in combination with spatial information to achieve accurate detection of occluded dynamic targets.
In low-light scenarios, the light changes triggered by the movement of the blocked dynamic target can be quickly detected, which improves the accuracy of identifying potential dynamic targets and reduces the risk of traffic accidents.
Smart Images

Figure CN120014600A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent driving technology, and specifically to a target detection method, model training method, device, equipment, vehicle and medium. Background Art
[0002] The perception system of autonomous driving technology primarily relies on a variety of sensors, such as cameras, lidar, infrared thermal imaging, and ultrasonic sensors, to help the vehicle understand its surroundings. These systems typically include vision-based object detection algorithms, lidar obstacle detection, and infrared sensors that sense temperature differences. This technology provides good environmental perception under normal weather conditions. However, in inclement weather, during nighttime driving, or in complex urban environments, these technologies struggle to accurately identify pedestrians, vehicles, or other obstacles, increasing the risk of traffic accidents. Summary of the Invention
[0003] Embodiments of the present application provide a target detection method, model training method, device, equipment, vehicle and medium.
[0004] The technical solutions adopted in the embodiments of this application are as follows:
[0005] A target detection method includes: obtaining a light change event stream in a vehicle driving scene; the light change event stream is captured by an event camera; performing spatiotemporal feature extraction on the light change event stream to obtain a first target feature representing the light change in the driving scene; based on the first target feature, detecting potential occluded dynamic targets in the driving scene to obtain a first target detection result; and the movement of the occluded dynamic targets triggers light changes in the driving scene.
[0006] According to the above technical approach, first, an event stream of light changes in a driving scene captured by an event camera is acquired. Second, spatiotemporal features are extracted from this light change event stream to obtain a first target feature representing the light changes in the driving scene. Finally, based on the first target feature, occluded dynamic targets in the driving scene that could potentially trigger light changes in the driving scene are detected to obtain a first target detection result. In this way, by leveraging the event camera's sensitivity to light changes and capturing subtle light changes in the driving scene, light changes in the driving scene triggered by the motion of the occluded dynamic target can be quickly detected, achieving accurate detection of potential occluded dynamic targets in low-light scenarios.
[0007] Furthermore, the above method also includes: obtaining spatial information of the driving scene; performing spatiotemporal feature extraction on the light change event stream to obtain a first target feature characterizing the light changes in the driving scene, including: performing spatiotemporal feature extraction on the light change event stream and spatial information to obtain a first target feature characterizing the light changes in the vehicle driving scene.
[0008] Using these technical approaches, we acquire spatial information about the driving scene and extract spatiotemporal features from the light change event stream and spatial information to obtain a primary target feature that characterizes the light changes within the vehicle driving scene. This combined analysis of spatial information and the light change event stream improves the accuracy of detecting obscured dynamic objects.
[0009] Furthermore, the first target feature includes an event density feature and a target motion feature; spatiotemporal feature extraction is performed on the light change event stream and spatial information to obtain a first target feature characterizing the light change in the vehicle driving scene, including: spatiotemporal feature extraction on the light change event stream to obtain an event density feature; spatiotemporal feature extraction is performed on the light change event stream and spatial information to obtain a target motion feature; the target motion feature characterizes the motion state of a potential obscured dynamic target.
[0010] Using these techniques, we extract spatiotemporal features from the light change event stream to generate event density features. We then extract spatiotemporal features from the light change event stream and spatial information to generate target motion features. These target motion features characterize the motion state of potentially obscured dynamic targets. This allows for more accurate perception of occluded dynamic targets, based on the light change frequency represented by the event density features and the motion state of the occluded dynamic targets represented by the target motion features.
[0011] Furthermore, when the spatial information includes at least one of the following: image information collected by an image acquisition device; point cloud information collected by a laser radar; infrared thermal imaging information collected by an infrared camera.
[0012] According to the above technical means, spatial information includes at least one of the following: image information collected by image acquisition equipment; point cloud information collected by lidar; and infrared thermal imaging information collected by infrared cameras. In this way, data collected by different sensors can be applied to different environments and lighting conditions, providing more comprehensive and accurate spatial information.
[0013] Furthermore, the above method also includes: obtaining spatial information of the driving scene; performing feature extraction on the spatial information to obtain a second target feature; based on the first target feature, detecting potential occluded dynamic targets in the driving scene to obtain a first target detection result, including: based on the first target feature and the second target feature, detecting potential occluded dynamic targets in the driving scene to obtain a first target detection result.
[0014] Using the above technical means, spatial information of the driving scene is acquired; features are extracted from this spatial information to obtain a second target feature; and based on the first and second target features, potential occluded dynamic targets in the driving scene are detected to obtain a first target detection result. This provides richer spatial information based on the first and second target features, improving the accuracy of identifying occluded dynamic targets in the driving scene.
[0015] Furthermore, based on the first target feature and the second target feature, potential occluded dynamic targets in the driving scene are detected to obtain a first target detection result, including: converting the first target feature and the second target feature into the target feature space respectively to obtain a third target feature and a fourth target feature; fusing the third target feature and the fourth target feature to obtain a target fusion feature; based on the target fusion feature, potential occluded dynamic targets in the driving scene are detected to obtain a first target detection result.
[0016] According to the above technical means, the first target feature and the second target feature are respectively converted to the target feature space to obtain the third target feature and the fourth target feature; the third target feature and the fourth target feature are fused to obtain the target fusion feature; based on the target fusion feature, potential occluded dynamic targets in the driving scene are detected to obtain the first target detection result. In this way, converting the first target feature and the second target feature to the target feature space can reduce the difference between the first target feature and the second target feature in different spaces, improve the consistency of the first target feature and the second target feature, and fuse the third target feature and the fourth target feature to analyze the data from different angles, thereby obtaining more comprehensive and accurate detection results.
[0017] Furthermore, when the spatial information includes image information captured by an image acquisition device, the second target feature includes image features obtained by extracting features from the image information; when the spatial information includes point cloud information captured by a lidar, the second target feature includes point cloud features obtained by extracting features from the point cloud information; when the spatial information includes infrared thermal imaging information captured by an infrared camera, the second target feature includes heat source distribution features obtained by extracting features from the infrared thermal imaging information.
[0018] According to the above technical means, when the spatial information includes image information captured by an image acquisition device, the second target features include image features obtained by extracting features from the image information; when the spatial information includes point cloud information captured by a lidar, the second target features include point cloud features obtained by extracting features from the point cloud information; when the spatial information includes infrared thermal imaging information captured by an infrared camera, the second target features include heat source distribution features obtained by extracting features from the infrared thermal imaging information. In this way, data collected by multimodal sensors can provide more comprehensive and accurate spatial information.
[0019] Furthermore, the first target detection result includes the speed, position and movement direction of the potential occluded dynamic target; the method also includes: based on the speed, position and movement direction of the occluded dynamic target, predicting the movement trajectory of the occluded dynamic target to obtain the target movement trajectory.
[0020] According to the above technical means, the motion trajectory of the occluded dynamic target is predicted based on the speed, position and motion direction of the occluded dynamic target to obtain the target motion trajectory. In this way, the motion trajectory of the occluded dynamic target can be predicted more accurately.
[0021] Performing spatiotemporal feature extraction on a light change event stream to obtain a first target feature that characterizes light changes in a driving scene, including: utilizing a feature extraction network in a target model to perform spatiotemporal feature extraction on the light change event stream to obtain a first target feature; and detecting potential occluded dynamic targets in the driving scene based on the first target feature to obtain a first target detection result, including: utilizing a detection network in the target model to detect potential occluded dynamic targets in the driving scene based on the first target feature to obtain a first target detection result.
[0022] Using these technical approaches, the feature extraction network in the target model extracts spatiotemporal features from the light change event stream to obtain the first target feature. The detection network in the target model then detects potentially occluded dynamic targets in the driving scene to obtain the first target detection result. In this way, the feature extraction and detection networks in the target model enable more accurate perception of occluded objects.
[0023] A model training method comprises: obtaining a training sample set; the training sample set comprises a plurality of training samples with target detection result labels, the training samples comprising a historical light change event stream in a vehicle driving scene; the historical light change event stream is captured by an event camera; utilizing a feature extraction network of a to-be-trained model to extract spatiotemporal features from the historical light change event stream in the training samples to obtain a fifth target feature characterizing light changes in the driving scene; utilizing a detection network of the to-be-trained model to detect potential occluded dynamic targets in the driving scene based on the fifth target feature to obtain a second target detection result corresponding to the training sample; the movement of the occluded dynamic target triggers light changes in the driving scene; and performing at least one parameter update on the feature extraction network and the detection network in the to-be-trained model based on the second target detection result and the target detection result label corresponding to each training sample to obtain a target model.
[0024] According to the above technical means, a training sample set is obtained; the training sample set includes multiple training samples with target detection result labels, and the training samples include a historical light change event stream captured by an event camera in a vehicle driving scene; the feature extraction network of the model to be trained is used to extract spatiotemporal features from the historical light change event stream in the training samples to obtain a fifth target feature that characterizes the light changes in the driving scene; the detection network of the model to be trained is used to detect potential occluded dynamic targets in the driving scene based on the fifth target feature to obtain a second target detection result corresponding to the training sample; the movement of the occluded dynamic target triggers light changes in the driving scene; based on the second target detection result and target detection result label corresponding to each training sample, the feature extraction network and detection network in the model to be trained are updated with at least one parameter to obtain a target model. In this way, the sensitivity of the event camera to light changes is utilized to capture the weak light change event stream in the driving scene, and the model to be trained is trained to obtain a target model that can more accurately detect potential dynamic targets in low-light scenes.
[0025] A target detection device, comprising:
[0026] A first acquisition module is configured to acquire a light change event stream in a vehicle driving scene; the light change event stream is captured by an event camera;
[0027] A spatiotemporal feature extraction module is used to extract spatiotemporal features from the light change event stream to obtain a first target feature representing the light change in the driving scene;
[0028] The detection module is used to detect potential obscured dynamic targets in the driving scene based on the first target feature to obtain a first target detection result; the movement of the obscured dynamic target triggers light changes in the driving scene.
[0029] An embodiment of the present application provides a computer device including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0030] An embodiment of the present application provides a vehicle, comprising the above-mentioned computer device.
[0031] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements some or all of the steps in the above method when executed by a processor.
[0032] Beneficial effects of this application:
[0033] (1) By utilizing the sensitivity of the event camera to light changes and capturing subtle light changes in driving scenes, the light changes in driving scenes triggered by the motion of occluded dynamic targets can be quickly detected, thus achieving accurate detection of potential occluded dynamic targets in low-light scenes.
[0034] (2) By acquiring the spatial information of the driving scene and extracting the spatiotemporal features of the light change event stream and spatial information, the spatial information and the light change event stream are combined for analysis, thereby improving the accuracy of the perception of obscured dynamic targets.
[0035] (3) By extracting the spatiotemporal features of the light change event stream separately to obtain the event density features, and extracting the spatiotemporal features of the light change event stream and spatial information to obtain the target motion features, we can achieve more accurate perception of the occluded dynamic target based on the light change frequency represented by the event density features and the motion state of the occluded dynamic target represented by the target motion features.
[0036] (4) The data collected by different sensors such as image acquisition equipment, lidar, and infrared cameras can be applied to different environments and lighting conditions, providing more comprehensive and accurate spatial information.
[0037] (5) The spatial information features of the driving scene are extracted to obtain the second target features. Based on the first target features and the second target features, potential occluded dynamic targets in the driving scene are detected, which can provide richer spatial information and improve the accuracy of identifying occluded dynamic targets in the driving scene.
[0038] (6) Converting the first target feature and the second target feature into the target feature space can reduce the difference between the first target feature and the second target feature in different spaces, improve the consistency between the first target feature and the second target feature, and fuse the third target feature and the fourth target feature. The data can be analyzed from different angles, thereby obtaining more comprehensive and accurate detection results.
[0039] (7) Feature extraction is performed on the data collected by the multimodal sensors to provide more comprehensive and accurate spatial information for detecting occluded dynamic targets in driving scenes.
[0040] (8) According to the speed, position and movement direction of the occluded dynamic target, the motion trajectory of the occluded dynamic target is predicted to achieve a more accurate prediction of the motion trajectory of the occluded dynamic target.
[0041] (9) By utilizing the feature extraction network and detection network in the target model, more accurate perception of occluded objects can be achieved.
[0042] (10) Utilizing the sensitivity of the event camera to light changes, the weak light change event stream in the driving scene is captured and used to train the training model to obtain a target model that can more accurately detect potential dynamic targets in low-light scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of the implementation process of a target detection method proposed in this application;
[0044] Figure 2 A schematic diagram of the implementation process of a model training method proposed in this application;
[0045] Figure 3 A schematic diagram of the structure of an end-to-end large model proposed in this application;
[0046] Figure 4 A schematic diagram of a 360-degree omnidirectional perception system proposed in this application, which includes infrared thermal imaging sensors and visual cameras installed around a vehicle;
[0047] Figure 5 A schematic diagram of the environment perception of an autonomous vehicle at an intersection equipped with a multimodal perception system proposed in this application;
[0048] Figure 6 This is a schematic diagram of the principle of an infrared thermal imaging sensor and a visual camera working together to capture surrounding environment information proposed in this application;
[0049] Figure 7 This is a schematic diagram of the detection results based on infrared thermal imaging technology proposed in this application;
[0050] Figure 8 This is a schematic diagram of the detection results after the fusion of visual perception and infrared thermal imaging proposed in this application;
[0051] Figure 9 A schematic diagram of the structure of a target detection device proposed in this application;
[0052] Figure 10 This is a hardware entity diagram of a computer device proposed in this application. DETAILED DESCRIPTION
[0053] The following will describe the embodiments of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for the purpose of illustrating the present application and are not intended to limit the scope of protection of the present application.
[0054] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.
[0055] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0056] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0058] The embodiment of the present application proposes a target detection method, which can be executed by a processor of a computer device. When implemented, the computer device may include but is not limited to at least one of a car computer, a robot, a laptop computer, a tablet computer, a desktop computer, a large-screen device, a server, a mobile device, etc. Figure 1 As shown, the target detection method includes the following steps S101 to S103, wherein:
[0059] Step S101: Acquire a light change event stream in a vehicle driving scene; the light change event stream is captured by an event camera;
[0060] Here, the light change event stream is a light change event stream in a driving scene of a vehicle captured by an event camera within a target time range.
[0061] In some embodiments, event cameras may be mounted on the front, rear, left, and right sides of the vehicle.
[0062] In some implementations, an event camera is used to capture a light change event stream in scenes with insufficient lighting or complex weather, such as inclement weather or nighttime driving.
[0063] In some embodiments, the event camera can capture subtle light changes triggered by the motion of a dynamic target obscured at a turning intersection, as well as subtle light changes triggered by the motion of a dynamic target obscured in bushes.
[0064] Step S102: extracting spatiotemporal features from the light change event stream to obtain a first target feature representing light changes in the driving scene;
[0065] Here, spatiotemporal feature extraction includes the process of extracting key information of time and space dimensions from the light change event stream.
[0066] The first target feature includes but is not limited to a time density feature.
[0067] In some embodiments, spatiotemporal features are extracted from the light change event stream using a feature extraction network in the target model, where the feature extraction network in the target model includes but is not limited to at least one of a backbone network and a spatiotemporal network.
[0068] In some embodiments, a spatiotemporal network extracts spatiotemporal features from the light change event stream in both temporal and spatial dimensions. For example, the spatiotemporal network can be at least one of a long short-term memory network (LSTM) and a convolutional neural network (CNN). By extracting spatiotemporal features from the light change event stream using the spatiotemporal network, a first target feature is obtained. Based on the first target feature, occluded dynamic targets in the driving scene can be detected.
[0069] Step S103: Based on the first target feature, a potential obscured dynamic target in the driving scene is detected to obtain a first target detection result; the movement of the obscured dynamic target triggers a change in light in the driving scene.
[0070] In some embodiments, potential obscured dynamic targets may include, but are not limited to, vehicles obscured at a turning intersection, and / or pedestrians holding luminous objects behind obstructions. The luminous objects may include, but are not limited to, at least one of a flashlight, a mobile phone, a candle, and the like.
[0071] The obstruction may include but is not limited to at least one of a roadside building, a stationary vehicle, bushes, etc.
[0072] In some implementations, a detection network in the target model is used to detect potential occluded dynamic targets in the driving scene based on the first target feature.
[0073] In some embodiments, the above method may further include: determining a driving strategy of the vehicle based on the target detection result; and controlling the vehicle based on the driving strategy.
[0074] In some implementations, event cameras can be used to supplement dynamic perception. These cameras can capture image changes with microsecond temporal resolution, making them ideal for handling high-speed, dynamic scenes like corners. At corners, conventional perception systems struggle to capture rapidly appearing targets (e.g., pedestrians rapidly crossing the intersection or vehicles rapidly turning a corner). By detecting changes in the brightness of each pixel, event cameras can quickly capture the instantaneous motion of targets. This is particularly true for objects that quickly emerge from behind obstructions, allowing them to accurately capture their motion information and promptly detect potential dangers.
[0075] In some embodiments, the event camera can capture the moment of appearance of the target through a timestamp mark, and combined with the heat source data of infrared thermal imaging, it can determine the movement direction and speed of the dynamic target that is blocked in the shortest time, and quickly adjust the vehicle's driving path or speed to avoid collision.
[0076] In some implementations, the target model integrates data from event cameras, infrared cameras, and other sensors (such as lidar and image acquisition devices). This multimodal data fusion enables accurate identification and prediction of the behavior of occluded dynamic targets, generating optimal avoidance strategies and controlling the vehicle accordingly. For example, if an occluded dynamic target is predicted to intersect the vehicle's path, the target model can preemptively plan lane changes, speed reduction, and other maneuvers to ensure safe passage through a curve.
[0077] In some implementations, a vehicle detects subtle light changes at a turning intersection. An event camera captures this light change information in real time, generating a stream of light change events within a target time range. This stream of light change events is then processed to detect dynamic targets that are obscured at the turning intersection but trigger light changes in the driving scene.
[0078] In some embodiments, in an environment with dense bushes or trees, there is usually subtle movement of objects (such as pedestrians, animals, etc.). The event camera can capture subtle light changes through the gaps in the bushes or trees. Therefore, the light change event stream captured by the event camera within a certain time window can be used to detect the obscured dynamic target.
[0079] In some implementations, infrared thermal imaging can effectively detect faint heat sources behind dense bushes or trees, particularly living organisms like pedestrians or animals. Even if a dynamic target is partially obscured or its thermal radiation is weak, infrared thermal imaging can still identify it through temperature comparison. Therefore, spatiotemporal feature extraction can be performed on the light change event stream captured by the event camera and the heat source information collected by the infrared camera to detect obscured dynamic targets.
[0080] In the embodiment of the present application, first, an event stream of light changes in a driving scene captured by an event camera is acquired. Second, spatiotemporal features are extracted from the light change event stream to obtain a first target feature representing the light changes in the driving scene. Finally, based on the first target feature, an obscured dynamic target in the driving scene that potentially triggers the light changes in the driving scene is detected to obtain a first target detection result. In this way, by utilizing the event camera's sensitivity to light changes and capturing subtle light changes in the driving scene, light changes in the driving scene triggered by the motion of the obscured dynamic target can be quickly detected, achieving accurate detection of potential obscured dynamic targets in low-light scenarios.
[0081] In some embodiments, the target detection method further includes the following step S104, wherein:
[0082] Step S104: Acquire spatial information of the driving scene.
[0083] In some embodiments, the vehicle may also be equipped with at least one of an image acquisition device (ie, a visual camera), a laser radar, an infrared camera, etc.
[0084] Vision cameras are used to capture image information and identify conventional targets (such as vehicles and pedestrians). They provide clear visible light images, helping to detect and identify pedestrians, vehicles, traffic signs, and road markings on the road. Spatial information such as the boundaries, shape, and position of occluded dynamic targets can be extracted from the image information.
[0085] LiDAR provides precise depth information, helping to detect obstacles and their spatial distribution in the environment, improving the depth perception shortcomings of pure vision algorithms. Depth information is used to understand the positional relationship between obscured dynamic objects and the vehicle in the driving scene, providing more complete spatial information within the scene.
[0086] Infrared cameras are used to capture temperature differences in driving scenes under low light, nighttime, or obstructed conditions, and to detect dynamic targets (such as pedestrians and animals) that are difficult for visual cameras to identify. Using temperature differences in driving scenes, the distribution of heat sources can be determined, and the distribution of heat sources in driving scenes can be used to reflect spatial information.
[0087] The above step S102 includes the following step S1021, wherein:
[0088] Step S1021: performing spatiotemporal feature extraction on the light change event stream and the spatial information to obtain a first target feature characterizing the light change in the vehicle driving scene.
[0089] Here, the spatial information provides the driving road information and the spatial information of the obscured dynamic target.
[0090] In some embodiments, a spatiotemporal network is used to extract spatiotemporal features of light change event streams and spatial information to obtain first target features. The first target features may include at least one of event density features, target motion features, optical features, etc.
[0091] In some embodiments, cross-modal spatiotemporal feature extraction is performed on the light change event stream and spatial information to obtain a first target feature containing spatial information and light change information.
[0092] In the above embodiment, spatial information of the driving scene is acquired, and spatiotemporal features are extracted from the light change event stream and spatial information to obtain a first target feature that characterizes the light changes in the vehicle driving scene. This combined analysis of spatial information and the light change event stream improves the accuracy of the perception of obscured dynamic objects.
[0093] In some embodiments, the first target feature includes an event density feature and a target motion feature.
[0094] The above step S1021 includes the following steps S10211 and S10212, wherein:
[0095] Step S10211: extracting spatiotemporal features from the light change event stream to obtain the event density features;
[0096] Here, the event density feature may reflect the frequency and / or density of light change events.
[0097] In some embodiments, the event density feature may include but is not limited to a first density feature characterizing the frequency of occurrence of light change events in the time dimension, and / or a second density feature characterizing the density of occurrence of light change events in the spatial dimension, etc.
[0098] Step S10212: performing spatiotemporal feature extraction on the light change event stream and the spatial information to obtain the target motion feature; the target motion feature represents the motion state of the potential obscured dynamic target.
[0099] Here, the target motion feature may include light change features that can reflect the lane, speed and / or motion direction of the blocked dynamic target.
[0100] In some embodiments, the lane where the obscured dynamic target is located can be determined based on the position corresponding to the light change represented by the target movement; the speed of the obscured dynamic target can be determined based on the rate of light change represented by the target movement; and the direction of movement of the obscured dynamic target can be determined based on the position distribution change corresponding to the light change represented by the target movement.
[0101] For example, when the position distribution corresponding to the light change changes from darker to brighter, the movement direction of the obscured dynamic target is determined to be approaching the vehicle; when the position distribution corresponding to the light change changes from brighter to darker, the movement direction of the obscured dynamic target is determined to be away from the vehicle.
[0102] In some embodiments, spatiotemporal feature extraction can also be performed on the light change event stream and spatial information to obtain target optical features. The target optical features can reflect at least one of the intensity distribution of light, the projection of occluded dynamic targets in the light, etc.
[0103] In the above embodiment, spatiotemporal feature extraction is performed on the light change event stream to obtain event density features, and spatiotemporal feature extraction is performed on the light change event stream and spatial information to obtain target motion features. The target motion features characterize the motion state of potential occluded dynamic targets. In this way, based on the light change frequency characterized by the event density features and the motion state of the occluded dynamic targets characterized by the target motion features, more accurate perception of occluded dynamic targets is achieved.
[0104] In some embodiments, the spatial information includes at least one of the following:
[0105] Image information collected by an image acquisition device;
[0106] Point cloud information collected by LiDAR;
[0107] Infrared thermal imaging information collected by infrared cameras.
[0108] In some embodiments, the image information provides rich color and texture information for identifying objects, people, traffic signs, etc. in the environment, and is suitable for daytime and well-lit environments.
[0109] Point cloud information provides information such as the location, shape, and size of dynamic targets, and is applicable to various lighting conditions. It can detect and identify obstacles, including vehicles, pedestrians, and roadblocks, providing precise environmental perception.
[0110] Infrared thermal imaging information can detect hidden heat sources, such as obscured pedestrians or vehicles, improving the stealthiness of environmental perception, and is generally not limited even at night or in bad weather.
[0111] In the above embodiments, spatial information includes at least one of the following: image information captured by an image acquisition device; point cloud information collected by a lidar; and infrared thermal imaging information collected by an infrared camera. In this way, data collected by different sensors can be adapted to different environments and lighting conditions, providing more comprehensive and accurate spatial information.
[0112] In some embodiments, the target detection method further includes the following steps S105 and S106, wherein:
[0113] Step S105: Acquire spatial information of the driving scene;
[0114] Here, the spatial information includes but is not limited to at least one of image information, point cloud information, infrared thermal imaging information, and the like.
[0115] Step S106: performing feature extraction on the spatial information to obtain a second target feature.
[0116] The backbone network in the feature extraction network is used to extract features from spatial information to obtain the second target feature.
[0117] In some implementations, the backbone network can be used to perform feature extraction on image information, point cloud information, and infrared thermal imaging information to obtain image features, point cloud features, and thermal imaging features.
[0118] In some embodiments, the second target feature includes but is not limited to at least one of an image feature, a point cloud feature, a heat source distribution feature, and the like.
[0119] In some embodiments, the second target feature may include, but is not limited to, features of unobstructed dynamic targets in the driving scene, features of static targets, and / or features of environmental information.
[0120] The above step S103 includes the following step S1031, wherein:
[0121] Step S1031: Based on the first target feature and the second target feature, detect potential blocked dynamic targets in the driving scene to obtain a first target detection result.
[0122] In some embodiments, the first target feature and the second target feature may be fused to obtain a target fusion feature; based on the target fusion feature, potential occluded dynamic targets in the driving scene are detected to obtain a first target detection result.
[0123] In some embodiments, when the first target feature includes an event density feature and the second target feature includes an image feature, the event density feature and the image feature are fused to obtain a target fusion feature; when the first target feature includes an event density feature and the second target feature includes a point cloud feature, the event density feature and the point cloud feature are fused to obtain a target fusion feature; when the first target feature includes an event density feature and the second target feature includes a heat source distribution feature, the event density feature and the heat source distribution feature are fused to obtain a target fusion feature.
[0124] In some embodiments, when the first target feature includes an event density feature and the second target feature includes an image feature and a point cloud feature, the event density feature, the image feature, and the point cloud feature are fused to obtain a target fusion feature; when the first target feature includes an event density feature and the second target feature includes an image feature and a heat source distribution feature, the event density feature, the image feature, and the heat source distribution feature are fused to obtain a target fusion feature; when the first target feature includes an event density feature and the second target feature includes a point cloud feature and a heat source distribution feature, the event density feature, the point cloud feature, and the heat source distribution feature are fused to obtain a target fusion feature.
[0125] In some embodiments, when the first target feature includes event density features and target motion features, and the second target feature includes image features and point cloud features, the event density features, target motion features, image features, and point cloud features are fused to obtain a target fusion feature; when the first target feature includes event density features and target motion features, and the second target feature includes image features and heat source distribution features, the event density features, target motion features, image features, and heat source distribution features are fused to obtain a target fusion feature; when the first target feature includes event density features and target motion features, and the second target feature includes point cloud features and heat source distribution features, the event density features, target motion features, point cloud features, and heat source distribution features are fused to obtain a target fusion feature.
[0126] In some embodiments, when the first target feature includes event density features and target motion features, and the second target feature includes image features, point cloud features, and heat source distribution features, the event density features, target motion features, image features, point cloud features, and heat source distribution features are fused to obtain target fusion features.
[0127] In the above embodiment, spatial information of the driving scene is acquired; features are extracted from the spatial information to obtain a second target feature; and based on the first and second target features, potential occluded dynamic targets in the driving scene are detected to obtain a first target detection result. In this way, the first and second target features provide richer spatial information, improving the accuracy of identifying occluded dynamic targets in the driving scene.
[0128] In some embodiments, the above step S1031 includes the following steps S10311 to S10313, wherein:
[0129] Step S10311: converting the first target feature and the second target feature into a target feature space respectively to obtain a third target feature and a fourth target feature;
[0130] In some embodiments, the first target feature and the second target feature are converted into a Bird's-Eye View (BEV) space to achieve unification of the first target feature and the second target feature in terms of spatial perspective.
[0131] Step S10312: fusing the third target feature and the fourth target feature to obtain a target fusion feature;
[0132] In some embodiments, a feature fusion module is used to fuse the third target feature and the fourth target feature, and the features extracted by different networks are integrated to obtain a target fusion feature. The occluded dynamic target is detected based on the target fusion feature to obtain a comprehensive detection result, thereby improving the accuracy of detecting the occluded dynamic target.
[0133] The feature fusion module extracts different features from the input data of each sensor, performs joint analysis through a multi-layer deep learning model, and generates a unified target fusion feature.
[0134] In some embodiments, the third target feature and the fourth target feature may be fused using a variety of fusion methods, including but not limited to at least one of weighted fusion, feature-level fusion, decision-level fusion, and the like.
[0135] Step S10313: Based on the target fusion feature, detect potential occluded dynamic targets in the driving scene to obtain a first target detection result.
[0136] In the above embodiment, the first target feature and the second target feature are respectively converted to the target feature space to obtain the third target feature and the fourth target feature; the third target feature and the fourth target feature are fused to obtain the target fusion feature; based on the target fusion feature, potential occluded dynamic targets in the driving scene are detected to obtain the first target detection result. In this way, converting the first target feature and the second target feature to the target feature space can reduce the differences between the first target feature and the second target feature in different spaces, improve the consistency of the first target feature and the second target feature, and fuse the third target feature and the fourth target feature to analyze the data from different angles, thereby obtaining more comprehensive and accurate detection results.
[0137] In some embodiments, when the spatial information includes image information acquired by an image acquisition device, the second target feature includes an image feature obtained by extracting features from the image information;
[0138] In a case where the spatial information includes point cloud information collected by a laser radar, the second target feature includes a point cloud feature obtained by extracting features from the point cloud information;
[0139] In a case where the spatial information includes infrared thermal imaging information collected by an infrared camera, the second target feature includes a heat source distribution feature obtained by extracting features from the infrared thermal imaging information.
[0140] In some implementations, feature extraction is performed on data collected by multimodal sensors to adapt to environmental detection in different scenarios and improve the generalization ability of the target model.
[0141] In the above embodiment, when the spatial information includes image information captured by an image acquisition device, the second target features include image features extracted from the image information; when the spatial information includes point cloud information captured by a lidar, the second target features include point cloud features extracted from the point cloud information; and when the spatial information includes infrared thermal imaging information captured by an infrared camera, the second target features include heat source distribution features extracted from the infrared thermal imaging information. In this way, data collected by multimodal sensors can provide more comprehensive and accurate spatial information.
[0142] In some embodiments, the first target detection result includes the speed, position and movement direction of the potential occluded dynamic target;
[0143] The target detection method further includes the following step S107:
[0144] Step S107: Based on the speed, position and movement direction of the blocked dynamic target, the movement trajectory of the blocked dynamic target is predicted to obtain the target movement trajectory.
[0145] In some embodiments, the speed of the obscured dynamic target can be predicted based on the rate of light change represented by the first target feature; the lane of the obscured dynamic target can be predicted based on the position corresponding to the light change represented by the first target feature; and the direction of movement of the obscured dynamic target can be predicted based on the change in position distribution corresponding to the light change represented by the first target feature.
[0146] For example, when the position distribution corresponding to the light change represented by the first target feature changes from darker to brighter, it is determined that the obscured dynamic target is close to the vehicle; when the position distribution corresponding to the light change represented by the first target feature changes from brighter to darker, it is determined that the obscured dynamic target is far away from the vehicle.
[0147] In some embodiments, based on the speed, position, and motion direction of the occluded dynamic target, a prediction network is used to predict the motion trajectory of the occluded dynamic target to obtain the target motion trajectory.
[0148] In the above embodiment, the motion trajectory of the occluded dynamic target is predicted based on the speed, position and motion direction of the occluded dynamic target to obtain the target motion trajectory. In this way, a more accurate prediction of the motion trajectory of the occluded dynamic target can be achieved.
[0149] In some embodiments, the above step S102 includes the following step S1022, wherein:
[0150] Step S1022: Utilizing the feature extraction network in the target model, performing spatiotemporal feature extraction on the light change event stream to obtain the first target feature.
[0151] In some embodiments, the target model is a large end-to-end model that has been pre-trained and learned, has strong reasoning and self-learning capabilities, and can make comprehensive judgments on weak signals and environmental changes in complex scenes.
[0152] The target model includes a feature extraction network that can be used to extract spatiotemporal features from the light change event stream.
[0153] In some implementations, the target model can combine data from multimodal sensors to perform deep temporal reasoning. For example, when the motion of an obscured dynamic object at a corner triggers a subtle change in light, the target model can infer, based on historical data, vehicle trajectories, and traffic regulations, that this signal may represent a potentially dangerous target. The target model not only relies on pixel information from static images but can also reason about dynamic environments, thereby anticipating potential dangers and taking proactive evasive measures.
[0154] In some embodiments, the target model can autonomously learn from complex environments and continuously adjust its parameters through online learning and feedback mechanisms to adapt to different driving scenarios. When encountering new complex scenarios (such as obscured targets or subtle changes in light sources), the target model can continuously optimize its recognition capabilities through self-learning. For example, when encountering a human figure behind bushes for the first time, it may not be immediately recognized. However, as data accumulates and the target model is optimized, the perception system can more accurately predict the target through infrared thermal imaging and dynamic motion signals.
[0155] The above step S103 includes the following step S1032, wherein:
[0156] Step S1032: Utilizing the detection network in the target model, based on the first target feature, detect potential occluded dynamic targets in the driving scene to obtain the first target detection result.
[0157] In some embodiments, the target model has a powerful cross-modal data fusion capability and can perform in-depth target detection by utilizing data collected by sensors such as integrated image information, point cloud information, and infrared thermal imaging information from the detection network.
[0158] In the above embodiment, the feature extraction network in the target model extracts spatiotemporal features from the light change event stream to obtain first target features. The detection network in the target model then detects potentially occluded dynamic targets in the driving scene to obtain first target detection results. In this way, the feature extraction network and detection network in the target model enable more accurate perception of occluded objects.
[0159] In some embodiments, the feature extraction network in the target model includes a spatiotemporal network, and step S1022 may include:
[0160] Step S10221: Utilize the spatiotemporal network to extract spatiotemporal features of the light change event stream to obtain a first target feature.
[0161] In some embodiments, the feature extraction network in the target model further includes a backbone network, and the target detection method further includes:
[0162] Step S108, obtaining spatial information of the driving scene;
[0163] Step S109: Using the backbone network, extract features from the spatial information to obtain second target features.
[0164] The above step S1032 may include:
[0165] Step S10321: Utilize the detection network in the target model to detect potential occluded dynamic targets in the driving scene based on the first target feature and the second target feature to obtain a first target detection result.
[0166] In some embodiments, the detection network includes a feature alignment module, a feature fusion module, and a detection module.
[0167] The above step S10321 may include:
[0168] Step S103211: using a feature alignment module, converting the first target feature and the second target feature into a target feature space to obtain a third target feature and a fourth target feature;
[0169] Step S103212: using a feature fusion module, fusing the third target feature and the fourth target feature to obtain a target fusion feature;
[0170] Step S10323: Using a detection module, based on the target fusion feature, detect potential occluded dynamic targets in the driving scene to obtain a first target detection result.
[0171] In some embodiments, the target model further includes a prediction network, and the first target detection result includes the speed, position, and motion direction of the potential occluded dynamic target;
[0172] The target detection method further includes the following step S110:
[0173] Step S110: using the prediction network, based on the speed, position and movement direction of the occluded dynamic target, predicting the motion trajectory of the occluded dynamic target to obtain the target motion trajectory.
[0174] The embodiment of the present application proposes a model training method, which can be executed by a processor of a computer device. When implemented, the computer device may include but is not limited to at least one of a car computer, a robot, a laptop computer, a tablet computer, a desktop computer, a large-screen device, a server, a mobile device, etc. Figure 2 As shown, the model training method includes the following steps S201 to S204, wherein:
[0175] Step S201: Acquire a training sample set; the training sample set includes a plurality of training samples with target detection result labels, and the training samples include a historical light change event stream in a vehicle driving scene; the historical light change event stream is captured by an event camera;
[0176] Step S202: Using the feature extraction network of the model to be trained, extracting spatiotemporal features from the historical light change event stream in the training sample to obtain a fifth target feature representing light changes in the driving scene;
[0177] Here, the feature extraction network includes but is not limited to at least one of a backbone network and a spatiotemporal network. The backbone network is used to extract features from spatial information; the spatiotemporal network is used to extract spatiotemporal features from the light change event stream, or to extract spatiotemporal features from the light change event stream and spatial information.
[0178] In some embodiments, the model to be trained is learned through a large amount of historical data, including lighting change patterns and environmental changes at different intersections, and will associate subtle light changes with potential obscured dynamic targets (such as vehicles and pedestrians approaching from a distance), thereby enabling accurate detection in actual scenarios.
[0179] In some embodiments, when encountering a new scenario, the target model can quickly adjust model parameters through self-learning functions so that it can better adapt to the new environment.
[0180] Step S203: Detecting potential occluded dynamic targets in the driving scene based on the fifth target feature using the detection network of the to-be-trained model to obtain a second target detection result corresponding to the training sample; the movement of the occluded dynamic target triggers a change in light in the driving scene;
[0181] Step S204: Based on the second target detection result and the target detection result label corresponding to each of the training samples, update the parameters of the feature extraction network and the detection network in the to-be-trained model at least once to obtain a target model.
[0182] In the above embodiment, a training sample set is obtained; the training sample set includes multiple training samples with target detection result labels, and the training samples include a historical light change event stream captured by an event camera in a vehicle driving scene; the feature extraction network of the model to be trained is used to extract spatiotemporal features from the historical light change event stream in the training samples to obtain a fifth target feature that characterizes the light change in the driving scene; the detection network of the model to be trained is used to detect potential occluded dynamic targets in the driving scene based on the fifth target feature to obtain a second target detection result corresponding to the training sample; the movement of the occluded dynamic target triggers light changes in the driving scene; based on the second target detection result and target detection result label corresponding to each training sample, the feature extraction network and detection network in the model to be trained are updated with at least one parameter to obtain a target model. In this way, the sensitivity of the event camera to light changes is utilized to capture the weak light change event stream in the driving scene, and the model to be trained is trained to obtain a target model that can more accurately detect potential dynamic targets in low-light scenes.
[0183] In some embodiments, the feature extraction network of the model to be trained includes the spatiotemporal network to be trained; the training samples include historical spatial information in the vehicle's driving scene; and the fifth target feature includes historical event density features and historical target motion features.
[0184] The above model training method further includes the following steps S205 and S206:
[0185] Step S205: using the spatiotemporal network to be trained, extracting spatiotemporal features from the historical light change event stream to obtain historical event density features;
[0186] Step S206: Using the spatiotemporal network to be trained, perform spatiotemporal feature extraction on the historical light change event stream and historical spatial information to obtain historical target motion features; the historical target motion features represent the historical motion state of potential occluded dynamic targets.
[0187] In some embodiments, the training samples include historical spatial information in the vehicle's driving scenes, and the feature extraction network of the model to be trained includes a backbone network to be trained.
[0188] The above model training method also includes:
[0189] Step S207: Using the backbone network to be trained, extract features from the historical spatial information to obtain a sixth target feature;
[0190] The above step S203 includes the following step S2031:
[0191] Step S2031: Using the detection network of the model to be trained, based on the fifth target feature and the sixth target feature, detect potential occluded dynamic targets in the driving scene to obtain a second target detection result corresponding to the training sample.
[0192] In some embodiments, the detection network of the model to be trained includes a feature alignment module, a feature fusion module and a detection module.
[0193] The above step S2031 includes the following steps S20311 to S20313:
[0194] Step S20311: using the feature alignment module, respectively converting the fifth target feature and the sixth target feature into the target feature space to obtain the seventh target feature and the eighth target feature;
[0195] Step S20312: using the feature fusion module, fusing the seventh target feature and the eighth target feature to obtain a first fused feature corresponding to the training sample;
[0196] Step S20313: Using the detection module of the model to be trained, based on the first fusion features corresponding to the training samples, detect potential occluded dynamic targets in the driving scene to obtain second target detection results corresponding to the training samples.
[0197] In some embodiments, the second object detection result includes the speed, position, and motion direction of the potential occluded dynamic object.
[0198] The model training method further includes step S208:
[0199] Step S208: using the prediction network of the model to be trained, based on the speed, position and movement direction of the occluded dynamic target, predict the motion trajectory of the occluded dynamic target to obtain a first motion trajectory corresponding to the training sample.
[0200] In some embodiments, the second target detection result includes the speed, position and movement direction of the potential occluded dynamic target; the target detection result label includes a speed label corresponding to the speed of the occluded dynamic target, a position label corresponding to the position of the occluded dynamic target, and a movement direction label corresponding to the movement direction of the occluded dynamic target.
[0201] The above step S204 includes the following step S2041:
[0202] Step S2041: Based on the speed, position and movement direction of the occluded dynamic target, as well as the speed label, position label and movement direction label, the parameters of the feature extraction network and the detection network in the training model are updated at least once to obtain the target model.
[0203] The following describes the application of the embodiments of the present application in actual scenarios.
[0204] In the field of autonomous driving technology, the perception system acts as the "eyes" of the entire vehicle, and its accuracy and stability directly determine its safety. The primary task of the perception system is to perceive the surrounding environment in real time and provide accurate obstacle detection, target tracking, and path planning information. Currently, autonomous driving perception systems primarily rely on visual cameras and lidar, which provide good environmental perception in normal weather conditions. However, the performance of these technologies is often limited in inclement weather, night driving, or complex urban environments, affecting the perception system's ability to accurately identify pedestrians, vehicles, and other obstacles, thereby increasing the risk of traffic accidents.
[0205] In recent years, the autonomous driving field has gradually shifted to the use of end-to-end large models. These models not only process multimodal input data (such as images, point clouds, and infrared thermal imaging), but also perform deep reasoning and dynamic target prediction, significantly enhancing the perception capabilities of autonomous driving systems in complex environments.
[0206] When the light level is low, the quality of the images captured by the visual camera will be significantly reduced, which will affect the recognition effect of the system. For example, when driving at night, even with the assistance of car lights and street lights, the camera may still cause the detection accuracy to decrease due to insufficient light. In addition, strong light at night (such as the headlights of oncoming vehicles) or reflective objects may interfere with the normal operation of the camera, causing glare or overexposure, thereby affecting the autonomous driving system's judgment of road conditions. In addition, when faced with an occluded scene, the visual camera has difficulty identifying the obscured target, which in turn causes the autonomous driving system to be unable to accurately assess the potential dangers in the environment.
[0207] LiDAR performs poorly in extreme weather conditions such as heavy rain and snow. Rain and snow can scatter and absorb the laser beam, causing signal attenuation and increasing detection errors, ultimately affecting the radar's detection accuracy. While LiDAR can still operate in rain and snow, its effective detection range and resolution are significantly reduced, making it difficult to ensure system reliability. Furthermore, LiDAR's laser signal is weakly reflected by certain materials (such as glass and metal), which can prevent accurate detection of these objects.
[0208] Infrared thermal imaging technology typically has a lower resolution and cannot provide the same image detail as high-definition cameras, which affects the perception system's ability to accurately identify and classify target objects. Infrared thermal imaging technology relies on temperature differences between target objects, so detection effectiveness may be limited for objects with a similar temperature to the background environment. Furthermore, the high cost of infrared thermal imaging cameras may increase the overall cost of autonomous driving systems.
[0209] Based on the above description, the perception system provided in the embodiment of the present application adopts an end-to-end large model (corresponding to the target model in the aforementioned embodiment) for multimodal data fusion, and combines reasoning and self-learning capabilities, thereby greatly improving the system's target recognition, dynamic prediction and safety decision-making capabilities in complex environments.
[0210] Existing visual perception methods struggle to effectively detect subtle lighting changes at intersections and behind obstacles, potentially dangerous targets. The challenge in these complex scenarios lies primarily in the fact that visual perception systems typically rely on pixel-level information processing, making it difficult to effectively capture dynamic or subtle environmental changes. However, large, end-to-end models can combine data from multiple sensors (such as lidar, infrared thermal imaging, and event cameras) for comprehensive analysis, overcoming the limitations of visual algorithms.
[0211] The embodiments of the present application illustrate the target detection method using an end-to-end large model by detecting dynamic targets in scenes with faint light changes at turning intersections and in scenes with occlusion (corresponding to the potential occluded dynamic targets in the aforementioned embodiments).
[0212] In scenarios with subtle lighting changes at intersections, event cameras capture these subtle changes and combine them with the spatial information provided by lidar to predict targets. Through reasoning, the large end-to-end model can combine these subtle lighting changes with contextual information (such as lane markings and traffic lights) to infer possible dynamic targets (such as oncoming vehicles or pedestrians) and predict their trajectories. This approach overcomes the limitations of vision systems, which are often unable to make judgments due to subtle or momentary lighting changes.
[0213] In obstructed scenarios, infrared thermal imaging sensors capture changes in heat sources, which, combined with visual imagery and lidar spatial data, can more accurately identify obscured targets. Even for dynamic, low-contrast targets, the perception system can effectively infer their presence through temporal reasoning and target behavior prediction.
[0214] In the embodiment of the present application, an end-to-end large model is used to make comprehensive judgments on weak signals and environmental changes in complex scenes. Figure 3As shown, the end-to-end large model can include a spatiotemporal network 301, a backbone network 302, a BEV feature conversion module 303, a feature fusion module 304, and a decoder 305. The spatiotemporal network 301 is used to extract spatiotemporal features from the image information 306 collected by the visual camera, the point cloud information 307 collected by the lidar, the infrared thermal imaging information 308 collected by the infrared camera, and the light change event stream 309 collected by the event camera to obtain first target features, thereby enabling the detection of targets in scenes with faint light changes and occlusions at turning intersections. The backbone network 302 is used to extract features from the image information 306 collected by the visual camera, the point cloud information 307 collected by the lidar, the infrared thermal imaging information 308 collected by the infrared camera, and the light change event stream 309 collected by the event camera to obtain second target features to supplement the spatial information. The BEV feature conversion module 303 converts the first target feature output by the spatiotemporal network 301 into the BEV space to obtain the third target feature, and converts the second target feature output by the backbone network 302 into the BEV space to obtain the fourth target feature, thereby achieving feature alignment. The feature fusion module 304 is used to fuse the third and fourth target features to obtain a fused target feature. The decoder 305 is used to obtain the target detection result based on the fused target feature.
[0215] The end-to-end large model can combine data from multimodal sensors for deep temporal reasoning. For example, when the perception system detects a subtle change in light at a corner, the end-to-end model can infer that this signal represents a potentially dangerous target based on historical data, vehicle trajectories, and traffic regulations. The end-to-end model not only relies on pixel information from static images but also reasoning about dynamic environments, enabling it to anticipate potential dangers and take proactive evasive action.
[0216] The end-to-end large model is capable of autonomous learning from complex environments. Through online learning and feedback mechanisms, the system can continuously adjust model parameters to adapt to different driving scenarios. When the system encounters new complex scenarios (such as obscured targets or subtle changes in light sources), the model can continuously optimize its recognition capabilities through self-learning. For example, when the system first encounters a human figure appearing behind bushes, it may not be able to recognize it immediately. However, as data accumulates and the model is optimized, the system can more accurately predict the target through infrared thermal imaging and dynamic motion signals.
[0217] The end-to-end large-scale model boasts powerful cross-modal data fusion capabilities, integrating image information, point cloud information, infrared thermal imaging information, and light change event streams for in-depth target detection and prediction. In complex scenarios (including those with subtle lighting changes at corners and occlusions), target recognition relies not only on visual data but also leverages the strengths of LiDAR and infrared sensors to further enhance perception accuracy.
[0218] In this application, the reasoning and self-learning capabilities of a large end-to-end model are combined with multimodal data fusion technology to achieve accurate target recognition and prediction in complex environments that are difficult for perception systems in related technologies to handle through real-time processing and analysis of various sensor data. The specific implementation method includes the following key steps:
[0219] (1) Perception system architecture and multimodal data fusion;
[0220] The perception system improves perception accuracy and robustness by integrating data from multiple sensors in real time. Specifically, it integrates the following sensors: a visual camera (corresponding to the image acquisition device in the previous embodiment), a laser radar (LiDAR), an infrared thermal imaging sensor (corresponding to the infrared camera in the previous embodiment), and an event camera.
[0221] In the embodiment of the present application, an event camera is used to capture light change information at an intersection in real time (corresponding to the light change event stream in the aforementioned embodiment). Based on the light change information, dynamic targets in subtle light change scenes and occlusion scenes at turning intersections are detected. At the same time, visual cameras, lidar, and / or infrared thermal imaging sensors are used to supplement spatial information, wherein:
[0222] Vision cameras can provide high-resolution image information and identify the shape, color and texture characteristics of objects.
[0223] LiDAR measures the position, shape, and distance of objects by emitting laser beams and receiving reflected signals. It is therefore able to generate accurate three-dimensional environmental models and provide highly accurate environmental perception data for autonomous vehicles.
[0224] Infrared thermal imaging technology detects temperature differences by detecting infrared radiation emitted by objects. This allows it to effectively detect pedestrians and other heat sources even in environments where visible light is limited, such as at night, in dense fog, or during heavy rain. Infrared thermal imaging technology does not rely on external light sources and can operate in completely dark or extremely low-light environments.
[0225] Figure 4 This is a schematic diagram of the 360-degree all-round perception system proposed in this application, which includes infrared thermal imaging sensors and visual cameras installed around the vehicle. Figure 4As shown, multiple sensor groups 502 are deployed around vehicle 501. These sensors include event cameras, visual cameras, and infrared thermal imaging sensors, forming a comprehensive perception system covering the entire perimeter of the vehicle. Sensors 502's perception range is defined as a sector-shaped area 401. The visual cameras primarily capture image information under visible light, while the infrared thermal imaging sensors are used to detect heat-generating objects within the vehicle's field of view, such as pedestrians and vehicles. In extreme weather conditions, such as heavy fog and rainstorms, the infrared thermal imaging sensors can complement the visual cameras' limitations and effectively detect obscured pedestrians or other obstacles.
[0226] Data fusion method: In the subtle lighting changes and occlusion scenarios at intersections, vision-based object detection algorithms, lidar obstacle detection, and infrared sensor temperature sensing have difficulty capturing potential dynamic targets. Therefore, the data collected by each sensor is input into an end-to-end large model (such as a deep learning model based on a convolutional neural network (CNN)). The end-to-end large model extracts different features from the input data of each sensor and jointly analyzes these features through a multi-layer deep learning model to generate a unified target recognition and prediction output.
[0227] (2) Prediction of weak light changes and dynamic target recognition;
[0228] In complex environments like corners, visual camera systems struggle to capture dynamic targets caused by local lighting changes. In particular, the faint flickering of lights from distant vehicles or pedestrians can make it difficult to accurately judge the visual image, easily causing the autonomous driving system to miss potential dangers.
[0229] In this embodiment of the present application, when a vehicle approaches a curve, the event camera can capture real-time information about changes in light at the intersection, such as the flashing of headlights from distant vehicles and changes in intersection traffic lights, and output the data with extremely high temporal resolution (microseconds). By processing this event data, even subtle changes in intersection traffic lights, such as the flashing of lights from distant vehicles, can be quickly detected.
[0230] Figure 5 A schematic diagram of the environment perception of an autonomous vehicle at an intersection equipped with a multimodal perception system proposed in this application is shown in FIG. Figure 5As shown, the front of vehicle 501 is equipped with multiple sensors 502, including a visual camera, an infrared thermal imaging sensor, and an event camera. The multimodal perception system integrates data from various sensors to perceive the surrounding environment in real time. There is an obstruction on the left side of the intersection, and target vehicle 503 is at a turning. The visual camera installed on vehicle 501 cannot directly perceive target vehicle 504. As target vehicle 503 moves, the light 504 emitted by target vehicle 503 changes. Therefore, the event camera can use this change in light 504 to determine the target vehicle's motion state. Combined with the spatial information from the lidar, it can infer the possibility that target vehicle 503 is about to enter the path of vehicle 501. The reasoning capabilities of the end-to-end large model can predict potential collision risks in advance. When the planning module detects that target vehicle 503 may enter the path of vehicle 501, it adjusts its trajectory or reduces speed to avoid a collision.
[0231] The light changes obtained from the event camera are combined with the spatial data of the lidar (corresponding to the point cloud information in the aforementioned embodiment) and analyzed through the temporal reasoning module of the end-to-end large model. The end-to-end model infers the dynamic targets that may be represented by the changes in weak light sources through historical trajectories and background knowledge (such as traffic rules, historical data, etc.). For example, if the system detects a change in light at a corner, after combining the spatial information of the lidar, the model can infer the dynamic behavior of vehicles, pedestrians or other traffic participants that may be approaching.
[0232] During the training phase, the perception system draws on a vast amount of historical data, including lighting patterns at different intersections and environmental changes. The system learns to associate subtle light changes with potential dynamic targets (such as approaching vehicles or pedestrians), enabling it to make accurate judgments in real-world scenarios.
[0233] Once a potential dynamic target is predicted, the deep learning model can be used to further infer the target's trajectory and direction of movement. For example, by detecting subtle changes in light, it can infer whether the target is approaching quickly or moving slowly, and then plan a safe driving route based on the target's trajectory.
[0234] (3) Detection and prediction of occluded targets;
[0235] In complex environments, targets may be obscured by obstacles such as buildings and bushes, making it impossible for pure vision systems to directly identify the target. This is especially true at night or in low-light environments, where even if the target is present, the recognition performance of existing perception systems will be greatly reduced.
[0236] In the embodiments of this application, the infrared thermal imaging sensor can identify heat sources caused by temperature differences in low-light environments, helping the perception system capture targets that are obscured by the visual camera. For example, when a pedestrian appears behind bushes, even though there is no direct object outline in the visual image, the infrared thermal imaging sensor can detect the pedestrian's body temperature signal and output it as a potential target.
[0237] In the occlusion scene at the turning intersection, Figure 6 This is a schematic diagram of the principle of the infrared thermal imaging sensor and visual camera proposed in this application working together to capture surrounding environment information. Figure 6 As shown, area 601 represents blind spots not covered by the visual camera. However, objects within these blind spots can be detected by the infrared thermal imaging sensor. For example, a cyclist behind vehicle 501 is located within the camera's blind spot. The infrared thermal imaging sensor captures and identifies the cyclist's thermal image, ultimately matching its position with the visual image. This multi-sensor fusion technology effectively enhances the autonomous driving system's perception capabilities in complex environments.
[0238] The following is a comparison of the test results using only infrared thermal imaging technology and the test results using a combination of visual cameras and infrared thermal imaging technology.
[0239] Figure 7 This is a schematic diagram of the detection results based on infrared thermal imaging technology proposed in this application, such as Figure 7 As shown, dynamic targets in the scene are perceived through the infrared radiation they emit, showing the positions of different heat sources (such as electric vehicle riding target 701, vehicle target 702 and pedestrian target 703, etc.) in the road scene. Since infrared thermal imaging does not rely on visible light sources, it can still clearly identify target objects in the scene even at night or in low visibility environments. It can be seen that in complex urban environments, pedestrians and vehicles are accurately identified by their temperature differences. This technology is particularly suitable for severe weather conditions such as heavy fog and heavy rain, and can provide stable and reliable perception results, making up for the shortcomings of perception systems in related technologies in these scenarios.
[0240] Figure 8 This is a schematic diagram of the detection results after the fusion of visual perception and infrared thermal imaging proposed in this application. Figure 8 As shown, a multimodal fusion algorithm is used, combining the stability of infrared thermal imaging with the detail perception capabilities of visual cameras. By fusing information from two different modalities, it can accurately identify and mark pedestrian targets 703 and vehicle targets 702 in nighttime or low-light environments. Figure 8The detection boxes and confidence values in the image demonstrate that the perception system, by combining infrared thermal imaging with visual data, can simultaneously process multiple targets, significantly improving detection accuracy and scene understanding. The auxiliary role of infrared thermal imaging ensures that the system can still obtain sufficient environmental perception information even when vision is impaired or blurred.
[0241] Furthermore, event cameras can be used to supplement dynamic perception. They can capture image changes with microsecond temporal resolution, making them ideal for handling high-speed, dynamic scenes such as turning intersections. At turning intersections, perception systems in related technologies struggle to capture rapidly appearing targets (for example, pedestrians rapidly crossing the intersection or vehicles rapidly turning a corner). By detecting changes in the brightness of each pixel, event cameras can quickly capture the instantaneous motion of targets. This is especially true for objects that quickly emerge from behind obstructions. Event cameras can accurately capture their motion information, helping the system to promptly perceive potential dangers.
[0242] The event camera can capture the moment of the target's appearance through a timestamp. Combined with the heat source data of infrared thermal imaging, it can predict the target's movement direction and speed in the shortest time and quickly adjust the driving path or speed to avoid collision.
[0243] In an obscured environment with dense bushes or trees, there are often faint objects (such as pedestrians, animals, etc.) that may be partially or completely obscured. Multimodal technologies, especially the combination of infrared thermal imaging and event cameras, can effectively identify and predict these obscured targets.
[0244] Infrared thermal imaging can track the changes in these weak heat sources and, combined with the target's motion pattern, predict its likely trajectory. By utilizing the real-time detection of dynamic changes in event cameras, the instantaneous movement of the target can be captured, providing more accurate dynamic information.
[0245] Once the infrared thermal imaging system detects a heat source, the end-to-end large model analyzes the target's motion pattern using a temporal reasoning module. Based on historical behavioral data, the end-to-end large model predicts the target's trajectory and calculates its relative speed and position to the vehicle. This prediction enables the autonomous driving system to proactively take appropriate obstacle avoidance or deceleration measures to ensure safe driving.
[0246] The end-to-end large model continuously learns dynamic target behaviors in complex scenarios. For example, if the perception system encounters a new type of occlusion (such as a different type of obstacle or a new light pattern), the end-to-end large model can quickly adjust model parameters through self-learning to better adapt to the new environment.
[0247] (4) Target identification and path planning;
[0248] In complex environments, targets are not only static but also dynamically change with the changing environment. Therefore, autonomous driving systems need not only to identify targets but also to perform accurate path planning and decision-making.
[0249] (5) Model optimization and continuous learning.
[0250] When faced with ever-changing and complex environments, autonomous driving systems need to continuously optimize and adapt to new situations, especially in unknown or unconventional scenarios. Through continuous online learning mechanisms, they can continuously gain experience from new environmental data and update the end-to-end large model. Whenever encountering new complex scenarios (such as unusual lighting changes or occlusion), the new data is automatically recorded and the model parameters are updated. Using reinforcement learning algorithms, decision-making strategies are continuously adjusted through interaction with the environment and feedback. This self-learning mechanism enables gradual optimization in diverse and complex environments, thereby improving the accuracy of target recognition and prediction.
[0251] In the embodiments of the present application, by combining the reasoning ability, self-learning ability and multimodal data fusion of the end-to-end large model, the limitations of the perception system in related technologies in scenes with weak light changes and occlusions are solved. Through the data fusion of multimodal sensors such as event cameras, infrared thermal imaging, and lidar, potential targets can be predicted in real time and accurate trajectory planning can be performed, thereby ensuring the efficient operation and safety of the autonomous driving system in complex environments.
[0252] Based on the above embodiments, this embodiment provides a target detection device, such as Figure 9 As shown, the target detection device 900 includes: a first acquisition module 901, which is used to obtain a light change event stream in a vehicle's driving scene; the light change event stream is captured by an event camera; a spatiotemporal feature extraction module 902, which is used to perform spatiotemporal feature extraction on the light change event stream to obtain a first target feature that characterizes the light change in the driving scene; a detection module 903, which is used to detect potential occluded dynamic targets in the driving scene based on the first target feature to obtain a first target detection result; the movement of the occluded dynamic target triggers the light change in the driving scene.
[0253] In some embodiments, the object detection device further includes: a second acquisition module, configured to acquire spatial information of the driving scene;
[0254] The spatiotemporal feature extraction module includes: a first spatiotemporal feature extraction unit, configured to perform spatiotemporal feature extraction on the light change event stream and the spatial information to obtain a first target feature representing light changes in the vehicle driving scene.
[0255] In some embodiments, the first spatiotemporal feature extraction unit includes: a first spatiotemporal feature extraction subunit, used to perform spatiotemporal feature extraction on the light change event stream to obtain the event density feature; a second spatiotemporal feature extraction subunit, used to perform spatiotemporal feature extraction on the light change event stream and the spatial information to obtain the target motion feature; the target motion feature characterizes the motion state of a potential occluded dynamic target.
[0256] In some embodiments, the spatial information includes at least one of the following: image information collected by an image acquisition device; point cloud information collected by a laser radar; and infrared thermal imaging information collected by an infrared camera.
[0257] In some embodiments, the target detection device also includes: a third acquisition module for acquiring spatial information of the driving scene; a feature extraction module for performing feature extraction on the spatial information to obtain a second target feature; the detection module includes: a first detection unit for detecting potential occluded dynamic targets in the driving scene based on the first target feature and the second target feature to obtain a first target detection result.
[0258] In some embodiments, the first detection unit includes: a conversion subunit, used to convert the first target feature and the second target feature into a target feature space respectively to obtain a third target feature and a fourth target feature; a fusion subunit, used to fuse the third target feature and the fourth target feature to obtain a target fusion feature; and a detection subunit, used to detect potential occluded dynamic targets in the driving scene based on the target fusion feature to obtain a first target detection result.
[0259] In some embodiments, when the spatial information includes image information captured by an image acquisition device, the second target feature includes image features obtained by extracting features from the image information; when the spatial information includes point cloud information captured by a lidar, the second target feature includes point cloud features obtained by extracting features from the point cloud information; when the spatial information includes infrared thermal imaging information captured by an infrared camera, the second target feature includes heat source distribution features obtained by extracting features from the infrared thermal imaging information.
[0260] In some embodiments, the first target detection result includes the speed, position and movement direction of the potential obscured dynamic target; the target detection device also includes a prediction module for predicting the motion trajectory of the obscured dynamic target based on the speed, position and movement direction of the obscured dynamic target to obtain the target motion trajectory.
[0261] In some embodiments, the spatiotemporal feature extraction module includes a second spatiotemporal feature extraction unit, which is used to use the feature extraction network in the target model to perform spatiotemporal feature extraction on the light change event stream to obtain the first target feature; the detection module includes a second detection unit, which is used to use the detection network in the target model to detect potential occluded dynamic targets in the driving scene to obtain the first target detection result.
[0262] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0263] This embodiment further proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements some or all of the steps in the above method when executing the program.
[0264] This embodiment further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.
[0265] This embodiment further provides a computer program, including computer-readable codes. When the computer-readable codes are run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0266] This embodiment also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, implements some or all of the steps in the above method. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0267] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.
[0268] It should be noted that the embodiment of the present application provides a hardware entity of a computer device, such as Figure 10 As shown, the hardware entities of the computer device 1000 include: a processor 1001 that generally controls the overall operation of the computer device 1000. A communication interface 1002 enables the computer device to communicate with other terminals or servers via a network. A memory 1003 is configured to store instructions and applications executable by the processor 1001, and can also cache data to be processed or processed by the processor 1001 and various modules in the computer device 1000 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented by flash memory (FLASH) or random access memory (RAM). Data can be transmitted between the processor 1001, the communication interface 1002, and the memory 1003 via a bus 1004.
[0269] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0270] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A target detection method, characterized in that: include: Acquire a light change event stream in a driving scene of a vehicle; the light change event stream is captured by an event camera; Extracting spatiotemporal features of the light change event stream to obtain a first target feature representing the light change in the driving scene; Based on the first target feature, a potential obscured dynamic target in the driving scene is detected to obtain a first target detection result; the movement of the obscured dynamic target triggers a change in light in the driving scene.
2. The target detection method according to claim 1, characterized in that: The method further comprises: Acquiring spatial information of the driving scene; The step of extracting spatiotemporal features from the light change event stream to obtain a first target feature characterizing the light change in the driving scene includes: Spatiotemporal features are extracted from the light change event stream and the spatial information to obtain a first target feature that characterizes the light change in the vehicle driving scene.
3. The target detection method according to claim 2, characterized in that: The first target feature includes an event density feature and a target motion feature; the step of extracting spatiotemporal features from the light change event stream and the spatial information to obtain the first target feature characterizing the light change in the vehicle driving scene includes: Extracting spatiotemporal features of the light change event stream to obtain the event density features; The spatiotemporal features are extracted from the light change event stream and the spatial information to obtain the target motion features; the target motion features represent the motion state of potential blocked dynamic targets.
4. The target detection method according to claim 2, characterized in that: The spatial information includes at least one of the following: Image information collected by an image acquisition device; Point cloud information collected by LiDAR; Infrared thermal imaging information collected by infrared cameras.
5. The method according to claim 1, characterized in that The target detection method further includes: Acquiring spatial information of the driving scene; Performing feature extraction on the spatial information to obtain a second target feature; The detecting, based on the first target feature, a potential blocked dynamic target in the driving scene to obtain a first target detection result includes: Based on the first target feature and the second target feature, a potential blocked dynamic target in the driving scene is detected to obtain a first target detection result.
6. The target detection method according to claim 5, characterized in that: The detecting, based on the first target feature and the second target feature, a potential blocked dynamic target in the driving scene to obtain a first target detection result includes: Respectively converting the first target feature and the second target feature into a target feature space to obtain a third target feature and a fourth target feature; fusing the third target feature and the fourth target feature to obtain a target fusion feature; Based on the target fusion feature, potential blocked dynamic targets in the driving scene are detected to obtain a first target detection result.
7. The target detection method according to claim 5, characterized in that: In the case where the spatial information includes image information acquired by an image acquisition device, the second target feature includes an image feature obtained by extracting features from the image information; In a case where the spatial information includes point cloud information collected by a laser radar, the second target feature includes a point cloud feature obtained by extracting features from the point cloud information; In a case where the spatial information includes infrared thermal imaging information collected by an infrared camera, the second target feature includes a heat source distribution feature obtained by extracting features from the infrared thermal imaging information.
8. The target detection method according to any one of claims 1 to 7, characterized in that: The first target detection result includes the speed, position and movement direction of the potential blocked dynamic target; The target detection method further includes: Based on the speed, position and movement direction of the blocked dynamic target, the movement trajectory of the blocked dynamic target is predicted to obtain the target movement trajectory.
9. The target detection method according to any one of claims 1 to 7, characterized in that: The step of extracting spatiotemporal features from the light change event stream to obtain a first target feature characterizing the light change in the driving scene includes: Using a feature extraction network in a target model, extracting spatiotemporal features of the light change event stream to obtain the first target feature; The detecting, based on the first target feature, a potential blocked dynamic target in the driving scene to obtain a first target detection result includes: Using the detection network in the target model, based on the first target feature, potential occluded dynamic targets in the driving scene are detected to obtain the first target detection result.
10. A model training method, characterized in that: include: Acquire a training sample set; the training sample set includes a plurality of training samples with target detection result labels, the training samples include a historical light change event stream in a vehicle driving scene; the historical light change event stream is captured by an event camera; Using a feature extraction network of the model to be trained, extracting spatiotemporal features from the historical light change event stream in the training sample to obtain a fifth target feature that characterizes light changes in the driving scene; Using the detection network of the model to be trained, based on the fifth target feature, a potential occluded dynamic target in the driving scene is detected to obtain a second target detection result corresponding to the training sample; the movement of the occluded dynamic target triggers a light change in the driving scene; Based on the second target detection results and target detection result labels corresponding to each of the training samples, the parameters of the feature extraction network and the detection network in the model to be trained are updated at least once to obtain a target model.
11. A target detection device, characterized in that: include: A first acquisition module is used to acquire a light change event stream in a driving scene of a vehicle; the light change event stream is captured by an event camera; A spatiotemporal feature extraction module, used to extract spatiotemporal features from the light change event stream to obtain a first target feature representing the light change in the driving scene; A detection module, configured to detect a potential occluded dynamic target in the driving scene based on the first target feature, and obtain a first target detection result; The movement of the blocked dynamic object triggers a change in light in the driving scene.
12. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 10 are implemented.
13. A vehicle, characterized in that: Comprising the computer device of claim 12.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 10 are implemented.
15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method according to any one of claims 1 to 10 are implemented.