Vehicle environment sensing method, device and system
By introducing scene recognition models to the vehicle dynamically allocating sensor weights, the perceived accuracy problem caused by fixed sensor weights is solved, and efficient perception and accurate decision-making in various driving scenarios are achieved.
Patent Information
- Application Number
- CN202510778464.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, the fixed sensor weight makes it difficult for vehicles to effectively perceive complex or rare driving scenarios, affecting the perception accuracy and decision-making accuracy of assisted driving.
A scene recognition model is introduced to identify the current driving environment type, and the weight value of sensor data is dynamically allocated according to the scene type, and the fusion perception results are generated through multi-sensor fusion technology.
It improves the vehicle's perception accuracy and the accuracy of assisted driving decisions in a variety of driving scenarios, is more adaptable, and can effectively deal with complex and rare driving scenarios.
Smart Images

Figure CN120440036A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle assisted driving, and in particular to a vehicle environment perception method, device and system. Background Art
[0002] Currently, vehicles are typically equipped with multiple sensors, and use multi-sensor fusion technology to achieve active perception of the external environment. For example, vehicles that support assisted driving functions are typically equipped with sensors such as cameras and lidar. The assisted driving module uses the various sensor data collected by these sensors to complete perception, and then performs prediction and regulation (i.e., planning and control) based on this data, thereby assisting the driver in driving the vehicle.
[0003] In related technologies, to ensure the stability of each perception result, the weights of the various sensors involved in multi-sensor fusion technology (i.e., the proportion of various sensor data in the fusion perception calculation process) are usually fixed and maintained unchanged during the vehicle design and development phase. However, fixed sensor weights make it difficult for vehicles to effectively perceive complex environments or less common driving scenarios. In other words, fixed and locked sensor weights are difficult to effectively cope with the ever-changing driving scenarios. As a result, perception accuracy in some scenarios is low, which affects the accuracy of the subsequent decision-making process of assisted driving and urgently needs to be improved. Summary of the Invention
[0004] In view of this, the present application provides a vehicle environment perception method, device and system, which solves the problems caused by fixed sensor weights in related technologies by introducing a new scene recognition model to identify the current scene type and assign weights to sensor data according to the type.
[0005] Specifically, this application is implemented through the following technical solutions:
[0006] According to a first aspect of the present application, a vehicle environment perception method is provided, which is applied to a vehicle equipped with multiple sensors, and the method includes:
[0007] Acquiring corresponding types of sensor data collected by the multiple sensors for the current driving environment;
[0008] calling a scene recognition model to identify a current scene type of the current driving environment based at least on the sensor data, and obtaining weight values assigned by the scene recognition model to various sensor data that match the current scene type;
[0009] According to various sensor data and their weight values, a fusion perception result of driving sensitive factors in the current driving environment is generated.
[0010] According to a second aspect of the present application, a vehicle environment perception device is provided, which is applied to a vehicle equipped with multiple sensors, and includes:
[0011] a data acquisition unit, configured to acquire corresponding types of sensor data collected by the multiple sensors for the current driving environment;
[0012] a weight assignment unit, configured to call a scene recognition model to identify a current scene type of the current driving environment based at least on the sensor data, and obtain weight values assigned by the scene recognition model to various sensor data that match the current scene type;
[0013] The result generating unit is used to generate a fusion perception result of the driving sensitive factors in the current driving environment according to various sensor data and their weight values.
[0014] According to a third aspect of the present application, a vehicle environment perception system is provided, the system comprising a scene recognition model, a fusion module and a plurality of candidate perception models, wherein:
[0015] The scene recognition model is configured to: identify a current scene type of a current driving environment based at least on input sensor data; and select at least one target perception model that matches the current scene type from the plurality of candidate perception models, and assign a weight value that matches the current scene type to target sensor data corresponding to each target perception model;
[0016] Each target perception model is used to: calculate the corresponding preliminary perception results based on its corresponding target sensor data;
[0017] The fusion module is used to perform weighted operations on the preliminary perception results according to corresponding weight values, and generate fused perception results for the driving sensitive factors in the current driving environment according to the operation results.
[0018] According to a fourth aspect of the present application, there is provided a vehicle, comprising:
[0019] a processor, a memory for storing instructions executable by the processor, and a variety of sensors;
[0020] The processor implements the method as described in the first aspect above by running the executable instructions.
[0021] According to a fifth aspect of the present application, a computer program and / or instructions are provided, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0022] The technical solution provided by this application may at least have the following beneficial effects:
[0023] Through the above embodiments, after obtaining the corresponding types of sensor data collected by multiple sensors installed on the vehicle, the scene recognition model will be called to identify the current scene type based on at least the sensor data, and obtain the weight values assigned by the model to various sensor data that match the current scene type; finally, according to the various sensor data and their weight values, a fusion perception result of the driving sensitive factors in the current driving environment is generated, thereby achieving effective perception of the current driving environment.
[0024] Given that different sensor types often have different operating parameters (such as different detection distances, detection ranges, and signal attenuation under different weather conditions), this solution assigns corresponding weights to the different types of sensor data collected by different sensors for fusion operations. Specifically, based on traditional multi-sensor fusion technology, this solution introduces a pre-trained scene recognition model. This model leverages its powerful reasoning capabilities to accurately identify the current scene type of the driving environment (e.g., day / night, clear / rainy / snowy / foggy / dusty weather, bumpy / curving road conditions, etc.). Based on the recognition results, each sensor data type is assigned a weight that matches the current scene type. This allows different sensor types to have different weights in the subsequent multi-sensor fusion process, effectively improving the accuracy of the perception of driving-sensitive factors in the current scene type and contributing to the accuracy of subsequent assisted driving decision-making processes. Clearly, this solution, based on the scene recognition model, can adjust the sensor data weights in real time based on the recognition results of the current scene type. Therefore, it can effectively achieve accurate perception of driving-sensitive factors in a variety of driving scenarios, achieving greater adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a schematic diagram of the hardware architecture of an environmental perception system shown in an embodiment of the present application.
[0026] Figure 2 The figure is a flow chart showing a method for environmental perception of a vehicle according to an exemplary embodiment.
[0027] Figure 3 is a schematic diagram showing an ROI adjustment effect according to an exemplary embodiment.
[0028] Figure 4 This is a software architecture diagram of an environmental perception system shown in an embodiment of the present application.
[0029] Figure 5 It is a schematic structural diagram of a vehicle shown in an exemplary embodiment.
[0030] Figure 6 The figure is a block diagram of an environment perception device for a vehicle showing an exemplary embodiment. DETAILED DESCRIPTION
[0031] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0032] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this application. In some other embodiments, the method may include more or fewer steps than those described in this application. In addition, a single step described in this application may be broken down into multiple steps for description in other embodiments; and multiple steps described in this application may be combined into a single step for description in other embodiments.
[0033] In related technologies, in order to ensure the stability of each perception result, the weights of the various sensors involved in multi-sensor fusion technology (that is, the proportion of various sensor data in the fusion perception calculation process) are usually solidified and remain unchanged during the vehicle design and development stage.
[0034] However, the number of driving scenarios in the real world is numerous and ever-changing, and the long-tail problem of driving scenarios is more prominent, that is, there are driving scenarios that are rare but overall quite numerous and relatively complex, and there are many long-tail cases (such as edge cases, corner cases, extreme cases, etc.). The preset sensor weights are usually only set according to some common scenarios. The fixed sensor weights make it difficult for the vehicle to effectively perceive complex environments or relatively rare driving scenarios, so the fusion effect in these scenarios is not good. In other words, the fixed and locked sensor weights are difficult to effectively cope with the ever-changing driving scenarios, and thus the perception accuracy in some scenarios is low, which affects the decision-making accuracy of subsequent assisted driving and needs to be improved urgently.
[0035] In response to the technical problems existing in the related technologies, this application proposes a new environmental perception method based on a scene recognition model, as well as an environmental perception system corresponding to this method, which is described in detail below in conjunction with the accompanying drawings and related embodiments.
[0036] The scene recognition model described in this application can be built based on any form of neural network framework and trained in a supervised or unsupervised manner. Taking supervised training as an example, sample data corresponding to the driving case can be obtained in advance, and any sample data contains sample environmental data collected during the vehicle's driving process (specifically including different types of sensor data collected by multiple sample sensors) and corresponding sample weight values (i.e., weight values set for the sensor data collected by the above-mentioned multiple sample sensors). The sample weight values can be manually marked by technicians or generated by other standard software and used as labels for the sample environmental data. Through the above training, the obtained scene recognition model can accurately and efficiently use environmental data to identify the scene type of the vehicle's driving environment, and output weight values assigned to various sensor data.
[0037] In one embodiment, the scene recognition model can be trained using any type of large multimodal model (Large Multimodal Model, LMM, which can recognize and process multimodal data such as vision, text, and audio) to ensure that the scene recognition model can output accurate inference results based on multimodal data.
[0038] For example, a Vision-Language Model (VLM) can be used. This model is a multimodal artificial intelligence model that combines the capabilities of computer vision (CV) and natural language processing (NLP). It can simultaneously understand and process images (or videos) and text information, and establish connections between the two to achieve more complex tasks. The scene recognition model implemented using the VLM can relatively more accurately identify the current scene type from multiple dimensions based on multimodal sensor data and output a weighted value that matches the type (i.e., consistent with the current driving environment).
[0039] This application combines traditional multi-sensor fusion technology with a scene recognition model (based on neural networks such as VLM), thereby assigning weight values that are consistent with the current scene type to the sensor data collected by various sensors through the scene recognition model, and realizing efficient and accurate perception of the vehicle's surrounding environment based on multiple types of sensor data and the above-mentioned weight values, thereby improving the vehicle's environmental perception accuracy.
[0040] In addition, the assisted driving function described in this application is used to achieve assisted driving of the vehicle, and this application does not limit the level of assisted driving (such as L2, L3, L4, etc.). However, it should be noted that the assisted driving achieved by the system should comply with the relevant laws, regulations and standards of the relevant countries and regions (such as the place where the vehicle is sold and / or used), and provide relevant personnel (such as drivers, etc.) with corresponding operation entrances for them to choose to authorize or refuse use.
[0041] Figure 1 This is a schematic diagram of the hardware architecture of an environment perception system shown in an embodiment of the present application. Figure 1 As shown, from a hardware perspective, the system may include only the vehicle 11, or may include both the vehicle 11 and the server 13. From a software perspective, the environmental perception system may include a scene recognition model, a perception model, and a fusion module. If the environmental perception system only includes the vehicle 11, then the scene recognition model, the perception model, the fusion module, etc. are all deployed locally in the vehicle 11, such as in the domain controller of the vehicle 11. Exemplarily, the scene recognition model, the perception model, and the fusion module may be deployed in the domain controller of the assisted driving domain of the vehicle 11. It is understandable that the environmental perception system at this time is an on-board system, and the system may even run offline. If the environmental perception system includes both the vehicle 11 and the server 13, then at least one of the scene recognition model, the perception model, and the fusion module may be deployed in the server 13, and the remaining models / modules may be deployed locally in the vehicle, which will not be described in detail.
[0042] In addition to domain control, the vehicle 11 can also be equipped with various types of sensors. Among them, the sensors described in this application are sensors for collecting data on the external environment of the vehicle, and the number of any type of sensors can be one or more. The embodiment of this application does not limit the number of the above sensors and their installation positions. Exemplarily, the sensors may include cameras (such as a front-view camera 111, a side-view camera 112, a rear-view camera 113, etc.), a lidar 114, a millimeter-wave radar, an ultrasonic radar, etc., and may also include a radiometer (for detecting light intensity), a rain gauge (for detecting current rainfall), etc., which will not be repeated. Of course, the vehicle can also be equipped with at least one in-vehicle sensor (such as an in-vehicle camera, an in-vehicle microphone, an in-vehicle biosensor, an in-vehicle odor sensor, etc.) at a suitable location to realize the corresponding vehicle-mounted functions, and the embodiment of this application does not limit this.
[0043] In addition, when the environmental perception system includes a server 13 or there is a need to interact with the server 13, the vehicle can establish a network connection with the remote server 13 through a wireless communication module to interact with the server 13. For example, if the scene recognition model is deployed in the server 13, the vehicle 11 can send the environmental data collected by the aforementioned sensors to the server 13, and receive the weight value allocation results of the various sensor data returned by it. For another example, if the scene recognition model, perception model and fusion module are deployed in the server 13, the vehicle 11 can send the environmental data collected by the aforementioned sensors to the server 13, and receive the fusion perception results generated by it according to the allocated weight values.
[0044] The server 13 can be a physical server containing an independent host, or a virtual server hosted by a host cluster, a cloud server, etc. Furthermore, the embodiments of the present application do not restrict the number, type, or specific interaction method of the server 13 with the vehicle. Regarding the network 10 for interaction between the vehicle 11 and the server 13, the communication can be achieved by specifically selecting a corresponding type of wireless network based on the communication methods supported by the corresponding devices, and this application does not impose any restrictions on this.
[0045] In addition, the vehicle described in this application (such as vehicle 11) can be a pickup truck, a sedan, an SUV (Sport Utility Vehicle), a motorhome, a truck, etc. in terms of its functional form; in terms of its power form, it can be a fuel vehicle or a new energy vehicle (such as a hybrid vehicle, a pure electric vehicle, a hydrogen energy vehicle, a methanol energy vehicle, etc.). The present invention does not limit the specific form of the vehicle. In addition, the passenger 12 in the cabin can be a driver sitting in the driver's seat, or it can be at least one passenger sitting in other positions. This application also does not limit the number of passengers and their seating positions. Of course, the assisted driving function described in this application is used to assist the driver in driving the vehicle, and will not be repeated here.
[0046] Figure 2 This is a flow chart of a vehicle environment perception method shown in an embodiment of the present application. The method is applied to a vehicle equipped with multiple sensors, and can be specifically applied to the aforementioned environment perception system (composed of a scene recognition model, a perception model, and a fusion module). Figure 2 As shown, the method includes the following steps 202 to 206.
[0047] Step 302: Obtain corresponding types of sensor data collected by the multiple sensors for the current driving environment.
[0048] After the vehicle is powered on and / or the engine is started, the various sensors installed in the vehicle begin to collect corresponding sensor data in real time based on the current driving environment. Subsequently, after the assisted driving function is activated (e.g., the driver triggers the activation of the function), the environmental perception system can obtain the above sensor data in real time and use this data for subsequent processing to assist the driver in controlling the vehicle. Of course, at least some of the above-mentioned various sensors may also collect corresponding sensor data in real time after the assisted driving function is activated, and not collect data when the system is not activated, thereby avoiding ineffective sensor data collection and conserving vehicle power, storage, and other resources.
[0049] In the embodiments of the present application, the current driving environment may include some driving sensitive factors (i.e., objects or information that may interfere with the vehicle's driving). This solution aims to perceive these driving sensitive factors based on the sensor data in order to provide more accurate perception results for the downstream modules of the assisted driving system (such as the prediction module, the regulation and control module, etc.). For example, when the vehicle (hereinafter referred to as the self-vehicle) is driving normally in the current lane, the preceding vehicle, the following vehicle in the same lane, the preceding / following vehicles in the adjacent lane, the traffic lights / pedestrians at the intersection ahead, the lane lines / obstacles on the road, the lateral / longitudinal distance between the self-vehicle and the adjacent vehicle, the deviation distance / angle of the self-vehicle from the lane lines, etc., are all driving sensitive factors in the current driving environment.
[0050] Among them, the driving sensitive factors may include high-risk driving sensitive factors (i.e., driving sensitive factors that may have a certain impact on the safe driving of the vehicle), such as the sudden braking of the vehicle in front of the lane may cause the vehicle to rear-end the vehicle in front, the vehicle behind may accelerate to rear-end the vehicle, and other vehicles in the adjacent lane may change lanes and cause scratches with the vehicle, etc. These vehicles are all high-risk driving sensitive factors for the vehicle. For another example, when the vehicle is driving in the current lane without a guardrail, two-wheeled vehicles (bicycles, electric vehicles, etc.) driving on the non-motorized vehicle lane next to it, and children chasing and playing on the sidewalk may suddenly invade the current lane and collide with the vehicle, so they are also high-risk driving sensitive factors for the vehicle. The environmental perception system of the present application can accurately identify the above-mentioned high-risk driving sensitive factors by assigning appropriate weight values.
[0051] In one embodiment, the various sensors installed on a vehicle may include at least two of the following: cameras (e.g., divided into front-view cameras, side-view cameras, rear-view cameras, etc. according to the installation location, and divided into short-focus cameras and long-focus cameras according to the focal length range), LiDAR (Light Detection and Ranging), millimeter-wave radar, ultrasonic radar, etc. Among them, the sensor data collected by the camera includes an environmental image (each image may include pixel information and depth information, etc.), the sensor data collected by the LiDAR includes laser data (e.g., point cloud data received after the LiDAR emits laser light and is reflected back to the radar by other objects), the sensor data collected by the millimeter-wave radar includes millimeter-wave data (e.g., millimeter-wave data received after the millimeter-wave radar emits millimeter-wave signals and is reflected back to the radar by other objects), and the sensor data collected by the ultrasonic radar includes ultrasonic data (e.g., ultrasonic data received after the ultrasonic radar emits ultrasonic signals and is reflected back to the radar by other objects). The following will not be repeated. The specific parameters of the above-mentioned various sensors, such as brand, size, and installation location, can be flexibly set according to the actual scenario, and the embodiments of the present application are not limited to this.
[0052] It's understandable that different types of sensors collect different types of sensor data, and multiple sensors of the same type collect the same type of sensor data. If a vehicle is equipped with n types of sensors, these sensors will collect n different types of sensor data based on the current driving environment. This data is then processed using multi-sensor fusion technology to perceive various driving-sensitive factors in the current driving environment.
[0053] Step 304 : Calling a scene recognition model to identify a current scene type of the current driving environment based at least on the sensor data, and obtaining weight values assigned by the scene recognition model to various sensor data that match the current scene type.
[0054] Step 306 : Generate a fusion perception result of the driving sensitive factors in the current driving environment according to various sensor data and their weight values.
[0055] After acquiring the corresponding types of sensor data collected by various sensors, they can be input into a scene recognition model. As previously mentioned, the scene recognition model can be a VLM. Given that such models have the ability to process multimodal data (i.e., they can recognize and process multiple types of data), in order to improve the model's recognition accuracy for these multiple types of sensor data, a multimodal prompt word can be constructed based on the aforementioned sensor data and input into the VLM. This will not be further described.
[0056] In one embodiment, in addition to inputting the aforementioned sensor data into the scene recognition model, non-sensor data related to the vehicle may also be obtained and input. In other words, the sensor data and non-sensor data may be input into the scene recognition model together, so that the scene recognition model can identify the current scene type based on the sensor data and the non-sensor data. The non-sensor data may include at least one of the following: weather data, lighting data, map data, positioning data of the vehicle, driving data of the vehicle, and traffic rules corresponding to the current location. Of course, other data may also be included, which will not be described one by one. It is understandable that the above-mentioned non-sensor data is data of a different dimension from the aforementioned sensor data, and thus can describe the current driving environment from more dimensions, richer and more detailed, thereby facilitating the model to make a more comprehensive and accurate comprehensive judgment on the current driving environment, thereby effectively improving the recognition accuracy of the current scene type and the accuracy of the weight value allocation.
[0057] In related technical solutions, the fusion module usually uses all the sensor data equipped on the vehicle to perform fusion operations; for example, all the preset perception models will be called to perform preliminary perception operations, and the fusion module will perform fusion operations based on the preliminary perception results of each perception module to obtain the final fusion perception result. This type of operation method uses all perception modules indiscriminately, so regardless of the complexity of the actual driving scene, it will occupy a large amount of vehicle computing and storage resources, and the overall perception efficiency is low. In this regard, based on the aforementioned current scene type accurately identified by the scene recognition model, this solution further proposes a solution for flexibly selecting appropriate perception modules to participate in fusion according to the current scene type, so as to save vehicle resources as much as possible (for example, a small number of perception models can be selected in simple scenes, and multiple perception models can be selected in complex scenes).
[0058] In one embodiment, based on the identified current scene type, the scene recognition model can also select a target perception model from a plurality of preset candidate perception models to generate a fusion perception result. For example, the scene recognition model can select at least one target perception model that matches the current scene type from a plurality of candidate perception models, and assign corresponding weight values to at least the target sensor data corresponding to each target perception model. Based on this, when obtaining the weight values assigned by the scene recognition model to various sensor data, the model selection result and weight assignment result output by the scene recognition model can be received, wherein the model selection result is used to characterize the at least one target perception model that matches the current scene type selected by the scene recognition model from a plurality of candidate perception models, and each target perception model is used to determine the driving sensitive factors in the current driving environment based on its corresponding target sensor data and generate a corresponding fusion perception result; and the weight assignment result is at least used to characterize the weight values of various target sensor data.
[0059] It can be understood that each of the above-mentioned multiple candidate perception models corresponds to at least one type of sensor data. For example, the multiple candidate perception models may include at least two of the following: a visual perception model, a stereo binocular perception model, a lidar perception model, a millimeter-wave radar perception model, an ultrasonic radar perception model, an OCC occupancy network perception model, and a BEVFormer perception model. Among them, the visual perception model corresponds to an environmental image captured by at least one camera, the stereo binocular perception model corresponds to environmental images captured by multiple cameras, the lidar perception model corresponds to laser data captured by the lidar, the millimeter-wave radar perception model corresponds to millimeter-wave data captured by the millimeter-wave radar, the ultrasonic radar perception model corresponds to ultrasonic data captured by the ultrasonic radar, the OCC (Occupied Channel Carrier) occupancy network perception model corresponds to at least two of the above-mentioned environmental images, laser data, and millimeter-wave data, and the BEVFormer (Bird's Eye View Transformer) perception model also corresponds to at least two of the above-mentioned environmental images, laser data, and millimeter-wave data. Each of the above-mentioned candidate perception models can adopt a traditional perception model. In other words, the specific perception algorithm of each of the above-mentioned candidate perception models (that is, the algorithm that uses the corresponding type of sensor data to perform calculations to obtain the preliminary perception results of a single dimension corresponding to this type of sensor) can be found in the records in the relevant technology, and the embodiments of the present application are not limited to this.
[0060] Among them, after any candidate perception model is selected as the target perception model, at least one type of sensor data corresponding to the model becomes the corresponding target sensor data, and accordingly, the preliminary perception result of the model will be used by the fusion module to generate a fused perception result. For example, each target perception model can be called to calculate the corresponding preliminary perception result based on the corresponding target sensor data; then the fusion module can be called to perform weighted operations on each preliminary perception result according to the corresponding weight value, and generate a fused perception result of the corresponding target driving sensitive factor based on the operation result. Exemplarily, after the lidar perception model is selected as the target perception model, the preliminary perception result calculated by the model based on the laser data is used by the fusion module to generate the corresponding fused perception result; and after the OCC occupancy network perception model is selected as the target perception model, the preliminary perception result calculated by the model based on its own corresponding target sensor data (i.e., at least two types of environmental images, laser data, and millimeter wave data) is used by the fusion module to generate the corresponding fused perception result, which will not be repeated.
[0061] In addition, the environmental perception model can send the weight values assigned by itself to the sensor data corresponding to each target perception module to the corresponding target perception module, so that each module perception module can send its own preliminary perception results and corresponding weight values to the fusion module. Alternatively, on the one hand, each perception module can independently perform perception operations and send its own preliminary operation results to the fusion module; on the other hand, the environmental perception module can send the weight values assigned by itself to the sensor data corresponding to each target perception module to the fusion module. After obtaining the preliminary perception results and corresponding weight values of each target perception module, the fusion module can perform a fusion operation (such as weighted summation calculation) on each preliminary perception result based on the weight value to obtain the final fusion perception result. It can be seen that the aforementioned step 306 can be executed by calling the fusion module in the environmental perception system.
[0062] It can be understood that when the target perception model only includes any one of the visual perception model, stereo binocular perception model, lidar perception model, millimeter wave radar perception model and ultrasonic radar perception model, the preliminary perception result generated by the model based only on the target sensor data corresponding to itself will be used by the fusion module as the final fusion perception result; and when the target perception model only includes the OCC occupancy network perception model or the BEVFormer perception model, the preliminary perception result generated by the model based on the multiple target sensor data corresponding to itself will be used by the fusion module as the fusion perception result (at this time, it can be regarded as fusing these multiple target sensor data to achieve perception); and when the target perception model includes multiple perception models mentioned above, the fusion module can calculate and generate the final fusion perception result based on the preliminary perception results generated by each target perception model (at this time, it can be regarded as fusing these multiple target sensor data to achieve perception).
[0063] In one embodiment, the sensor data may include environmental images collected by a camera, laser data collected by a lidar sensor, and millimeter wave data collected by a millimeter wave radar. Based on the above sensor data, different weight values can be set for different sensors according to their characteristics under the following typical scene types. For example, when the current scene type is an irregular obstacle scene, the weight value of the laser data may be greater than the weight values of other sensor data. If a collision occurs on the road ahead, objects or debris may be scattered. If the road ahead is closed, cones, water barriers, etc. may be set up. The above objects often have irregular shapes, that is, they are irregular obstacles. The recognition accuracy of lidar for such obstacles is often higher than that of other sensors. Therefore, assigning a larger weight value to laser data helps to improve the detection accuracy of irregular obstacles, thereby effectively avoiding secondary accidents or intrusions into closed roads.
[0064] For another example, if the current scene type is an extreme weather scene, the weight value of the laser data can be lower than the weight value of the millimeter wave data. The above extreme weather scenes may be rain, snow, fog, haze, hail, dust and other weather scenes with obstructed vision. In such scenes, particulate matter in the air scatters the laser signal to a large extent, causing the laser signal to attenuate rapidly, thus affecting the detection range and accuracy of the laser signal to varying degrees. Millimeter wave signals have better detection effects than lidar in such weather scenes. Therefore, assigning a larger weight value to millimeter wave data can help improve detection accuracy in such extreme weather scenes, thereby effectively coping with extreme weather.
[0065] For another example, when the current scene type is a backlit scene, the weight value of the environmental image is lower than the weight value of the laser data and / or lower than the weight value of the millimeter-wave data. Driving towards the sun, using high beams in the opposite lane, the background lights on both sides of the road are too bright, or the gantry lights above the road are too bright, etc., are all typical backlit scenes. At this time, objects other than the light source in the environmental image captured by the camera are often difficult to distinguish, while laser (invisible) and millimeter-wave signals are usually not affected by the above-mentioned high-brightness light sources. Therefore, assigning a lower weight to the environmental image and increasing the weights of the laser data and millimeter-wave data can ensure that objects ahead can be accurately identified even in backlit scenes, avoiding accidents caused by backlighting.
[0066] For example, if the current scene type is a bumpy or turning scene, the weight value of the environmental image is greater than the weight value of other sensor data. When the road is bumpy or turns at a large angle, bumps such as potholes and bumps on the road ahead often cause the vehicle to bump up and down, affecting the riding experience. In this case, other types of sensors often have difficulty effectively detecting these potholes or bumps. However, image recognition technology based on environmental images can quickly and accurately detect these bumps, allowing the vehicle to effectively respond (such as adjusting the height and hardness of the air suspension in advance) and reduce bumps.
[0067] In one embodiment, the aforementioned step 306 can be executed by calling a fusion module in the environmental perception system. For example, after the scene recognition model assigns corresponding weight values to various sensor data, it can send each weight value to the corresponding target perception model; then, each target perception model can send its own preliminary perception results (wherein the preliminary perception results of any target perception model are calculated by at least one sensor data corresponding to the model) and the corresponding weight value to the fusion module; finally, the fusion module performs a fusion operation based on the received preliminary perception results and the corresponding weight values to generate the final fused perception result. Alternatively, the scene recognition model can also send the selection results of the target perception model and the corresponding weight assignment results to the fusion module, and then the fusion module can perform a fusion operation based on the preliminary fusion results provided by each target perception module and the corresponding weight values provided by the scene recognition module to generate the fused perception result.
[0068] In one embodiment, in order to further improve the accuracy of the fusion perception results generated by the fusion module, the aforementioned part of the non-sensor data can also be provided to the fusion module, so that the module can make a comprehensive judgment based on the aforementioned preliminary perception results and weight values received and the non-sensor data, and finally generate a more accurate fusion perception result. Figure 4 As shown in the figure, on the one hand, the preliminary perception results and weight values of the perception model can be sent to the fusion module, and on the other hand, the map, positioning and vehicle information (such as speed, acceleration, turn signal on status, etc.) can be sent to the fusion module, so that the module can perform multi-dimensional fusion operations based on this information to obtain more accurate calculation results.
[0069] For example, in the aforementioned extreme weather scenario, the weight values of the laser data and the millimeter wave data can be set to 0.3 and 0.7 respectively (it is assumed that the weight values of other sensor data are all 0, the same below); at this time, if the laser data indicates that the distance between the vehicle and the vehicle in front of the same lane is 8 meters, and the millimeter wave data indicates that the distance between the vehicle and the vehicle in front of the same lane is 9 meters, then the final distance between the vehicle and the vehicle in front of the same lane is 9*0.3+8*0.7=8.3 meters. For another example, in the aforementioned backlight scenario, the weight values of the environmental image, the laser data, and the millimeter wave data can be set to 0.2, 0.4, and 0.4 respectively. If the environmental image indicates that the distance between the vehicle and the pedestrian in front is 15 meters, the laser data indicates that the distance between the vehicle and the pedestrian in front is 20 meters, and the millimeter wave data indicates that the distance between the vehicle and the pedestrian in front is 18 meters, then the final distance between the vehicle and the pedestrian in front is 15*0.2+20*0.4*18*0.4=18.2 meters.
[0070] Given that different sensor types often have different operating parameters (such as different detection distances, detection ranges, and signal attenuation under different weather conditions), this solution assigns corresponding weights to the different types of sensor data collected by different sensors for fusion operations. Specifically, based on traditional multi-sensor fusion technology, this solution introduces a pre-trained scene recognition model. This model leverages its powerful reasoning capabilities to accurately identify the current scene type of the driving environment (e.g., day / night, clear / rainy / snowy / foggy / dusty weather, bumpy / curving road conditions, etc.). Based on the recognition results, each sensor data type is assigned a weight that matches the current scene type. This allows different sensor types to have different weights in the subsequent multi-sensor fusion process, effectively improving the accuracy of the perception of driving-sensitive factors in the current scene type and contributing to the accuracy of subsequent assisted driving decision-making processes. Clearly, this solution, based on the scene recognition model, can adjust the sensor data weights in real time based on the recognition results of the current scene type. Therefore, it can effectively achieve accurate perception of driving-sensitive factors in a variety of driving scenarios, achieving greater adaptability.
[0071] When sensor data includes environmental images captured by cameras, related technologies typically perform recognition on the entire environmental image or within a fixed ROI. However, due to the complexity and variability of driving scenarios, the locations of driving-sensitive factors in the environmental image often change. Therefore, traditional factor recognition methods can suffer from low accuracy. To address this issue, this solution proposes a dynamic ROI recognition solution.
[0072] In one embodiment, when the sensor data includes an environmental image captured by a camera, the scene recognition model can identify driving sensitive factors in the current driving environment based on the sensor data and determine the location of high-risk driving sensitive factors therein, and then output corresponding ROI adjustment instructions. Thus, the environmental perception system can adjust the ROI position in the environmental image in response to the ROI adjustment instructions outputted based on the driving sensitive factors, so that the adjusted ROI covers the high-risk driving sensitive factors identified by the scene recognition model. Figure 3 As shown in the figure, at a certain moment when the vehicle is driving in the current lane, the right front vehicle closest to the vehicle is flashing its left turn signal (indicating that it wants to change lanes to the left or overtake), and the ROI position in the image now covers the vehicle, as shown in the rectangular ROI1 in the figure. After a period of time, the left turn signal of the right front vehicle goes out (indicating that it gives up changing lanes or overtaking), and at the same time, there is a left front vehicle in the left lane of the vehicle that is overtaking the vehicle, so the ROI position can be adjusted, and the adjusted ROI area covers the left front vehicle, as shown in the rectangular ROI1' in the figure. Of course, given the complexity of the current driving environment and the continuity of the camera capturing the environmental image, the above ROI adjustment may occur in Figure 3 Any image between the two images shown (such as the first frame image of the aforementioned left vehicle), and the environment image at any moment may contain multiple ROIs at the same time, which will not be described in detail.
[0073] Furthermore, for the environmental image, the resolution of the inner part of the ROI can be set to be higher than the resolution of the outer part; and / or the reduction factor of the inner part of the ROI can be set to be lower than the reduction factor of its outer part. Of course, the sampling rate of the camera corresponding to the ROI area can also be increased, that is, the camera can be controlled to capture environmental images related to the high-risk driving sensitivity factors at a higher frame rate (such as increasing the sampling frame rate from 20fps to 30fps). In this way, high-risk driving sensitivity factors can be processed with higher resolution, smaller scaling loss, and higher frame rate, reducing prediction inaccuracies that may be caused by processing delays.
[0074] Furthermore, to ensure that environmental images captured at higher resolutions, minimal scaling loss, and higher frame rates can be processed smoothly and efficiently, resulting in more accurate results, the aforementioned scene recognition model can be allocated more storage and computing resources. Furthermore, the model can be allocated resources in real time based on the number, type, and data volume of high-risk driving sensitivities, thereby improving the efficient use of local vehicle resources. Furthermore, when high-risk driving sensitivities are predicted, drivers can be alerted through methods such as highlighting or voice announcements to minimize potential risks.
[0075] In one embodiment, when there are multiple sensors of any type, the accuracy information of each sensor of that type can also be provided to the scene recognition model; accordingly, the scene recognition model can determine the weight values of each sensor that is positively correlated with its accuracy based on the accuracy information of each sensor of that type, and determine the weight value of the aforementioned sensor data of that type based on the weight value of each sensor of that type. It can be seen that when there are multiple sensors of any type, this solution allows the scene recognition model to perform a weighted sum operation based on the accuracy information of each sensor, and determine the weight value of the sensor data collected by the sensor of that type based on the operation result. Obviously, the weight value is strongly correlated with the overall accuracy of each sensor of that type. For example, the higher the sensor accuracy, the greater the weight value of the corresponding sensor data. This method helps to improve the weight of high-precision sensors in multi-sensor fusion operations, and helps to improve the credibility of the final perception results.
[0076] In one embodiment, when the vehicle has turned on the assisted driving function, it can also generate assisted driving control instructions for the current driving environment based on the fusion perception results of the driving sensitive factors, and control the vehicle to travel according to the assisted driving control instructions. Among them, the assisted driving control instructions can be generated by the driving control module associated with the fusion module and sent to the corresponding mechanical components (such as engine / motor, throttle / electric switch, brake, steering wheel, etc.) for execution to control the vehicle to achieve corresponding driving behavior at the macro level. The assisted driving control instructions can specifically be acceleration instructions, deceleration instructions, lane change instructions, turning instructions, U-turn instructions, parking instructions or flashing indicator light instructions, etc., and the embodiments of the present application are not limited to this. It can also be understood that the process of generating the above-mentioned assisted driving control instructions and controlling the vehicle's travel is the prediction and regulation stage after perception in the assisted driving function. By controlling the vehicle's travel according to the assisted driving control instructions, assisted driving of the vehicle can be achieved, that is, assisting its driver in driving the vehicle.
[0077] So far, the vehicle environment perception method has been described. Based on this method and the embodiment of the aforementioned perception model, this application also proposes a vehicle environment perception system, which includes a scene recognition model, a fusion module and multiple candidate perception models, wherein:
[0078] The scene recognition model is configured to: identify a current scene type of a current driving environment based at least on input sensor data; and select at least one target perception model that matches the current scene type from the plurality of candidate perception models, and assign a weight value that matches the current scene type to target sensor data corresponding to each target perception model;
[0079] Each target perception model is used to: calculate the corresponding preliminary perception results based on its corresponding target sensor data;
[0080] The fusion module is used to perform weighted operations on the preliminary perception results according to corresponding weight values, and generate fused perception results for the driving sensitive factors in the current driving environment according to the operation results.
[0081] As mentioned above, the scene recognition model can be VLM. Figure 4 The operation process of the system is described.
[0082] like Figure 4As shown in the figure, non-sensor data related to the vehicle (such as map / positioning information, vehicle information, weather information, and traffic regulations), as well as sensor data collected by various sensors installed on the vehicle (such as n1 cameras, n2 lidars, and n3 millimeter-wave radars), are all input into the pre-trained VLM. The VLM is fully trained and has functions such as "environmental understanding," "dynamic adjustment of ROI," "target perception model selection," and "dynamic allocation of sensor weights (KPIs)." Note that this does not mean that the VLM consists of the above four functional modules, but rather that the VLM model itself possesses multiple of these functions simultaneously.
[0083] For example, the VLM can understand the current driving environment based on the aforementioned sensor and non-sensor data inputs and adjust the ROI position based on the understanding results. Furthermore, the VLM can determine the current scene type of the current driving environment based on the understanding results, select at least one target perception model that matches the current scene type from multiple candidate perception models, and then assign corresponding weights to the target sensor data corresponding to each target perception model. Furthermore, the scene recognition model can output corresponding model selection and weight assignment results, specifying the selected target perception model and the assigned weight value.
[0084] In addition, at least part of the non-sensor data can be provided to the fusion module for fusion operation. Figure 4 As shown in the figure, in addition to sending the preliminary perception results and weight values of the fusion module to the fusion module, non-sensor data such as map, positioning and vehicle information (such as speed, acceleration, turn signal on status, etc.) can also be sent to the fusion module, so that the module can perform multi-dimensional fusion operations based on this information, which helps to obtain more accurate calculation results.
[0085] The fusion module can output a variety of data, including lane markings, obstacle types, longitudinal distances, lateral distances, and orientation information.
[0086] The following is an illustrative example of the operation process of the environment perception system using three specific scenarios.
[0087] Example 1:
[0088] Serious traffic accidents are common on highways, many of which are caused by secondary collisions. Therefore, effectively detecting the accident zone on highways and controlling subsequent vehicles to avoid it is one of the current technical challenges in assisted driving. Specifically, the vehicle's perception and recognition of the accident zone is slow or difficult to identify.
[0089] To this end, this application proposes the following solutions:
[0090] First, the VLM analyzes the environment in front of the vehicle based on sensor data collected by multiple sensors. For example, it can identify loudspeaker broadcasts beside the road, accident warnings from the gantry, behavioral analysis of the forward traffic flow, and scattered parts or debris on the ground.
[0091] Then, if a traffic accident is likely to occur ahead of the analyzed value, the ROU is adjusted accordingly for further accurate recognition. For example, if a traffic accident is predicted ahead of time on the left front, the ROI in the environment image is dynamically adjusted according to the output instructions of the VLM to expand the ROI area on the left side (at least the left front area) so that more details on the left side can be captured.
[0092] Then, when the vehicle is in an accident scenario, the VLM can call upon a target perception model that matches the scenario. For example, because the types of obstacles ahead are irregular (such as cones of various types and sizes, fences, and scattered debris), a single perception model is difficult to accurately identify. Therefore, the OCC occupancy network perception model can be selected as the primary perception model to accurately identify irregular obstacles in the area.
[0093] Finally, the VLM assigns appropriate weights (i.e., consistent with the current scenario type) to the corresponding sensors based on the aforementioned target perception model. As mentioned earlier, the OCC occupancy network perception model is selected for this scenario. In this case, the laser data collected by the LiDAR can be given a higher weight, followed by the environmental image captured by the binocular camera, while the weights of other sensor data can be set to zero (for example, millimeter-wave radars generally have difficulty identifying irregular obstacles such as those mentioned above).
[0094] Through this process, we can effectively understand the current scene type, select a more appropriate perception model, and set more reasonable weights for the corresponding sensor data. This approach enables early detection of irregular obstacles, which helps the subsequent planning module make correct decisions.
[0095] Example 2:
[0096] In rainy, foggy, snowy, or dusty conditions, visibility far ahead is poor. In such cases, the related art typically assigns the same fixed weight to each sensor. However, rain, snow, fog, and dusty conditions vary in severity (e.g., rain can be categorized as light, moderate, heavy, or torrential). Fixed weights can lead to inaccurate multi-sensor fusion perception and poor fusion performance.
[0097] To this end, this application proposes the following solutions:
[0098] A large amount of test data under different operating conditions can be collected in advance as training samples to train the aforementioned VLM. This allows the model to learn the characteristics and weighting patterns of various sensors, thereby facilitating the precise allocation of sensor weights for various extreme weather scenarios. For example, the test data shown in Table 1 below can be used to train the VLM.
[0099] Table 1
[0100]
[0101] In addition, the sensor weight parameters in the above table can also be input into the BEVFomer perception model so that the model can perform weighted summation on the sensor data from each channel during the inference phase and output more accurate perception results.
[0102] Example 3:
[0103] For continuously undulating slopes or roads with large turns, since the perception model in related technologies solidifies the ROI during the design phase, some low obstacles (pits or bumps) may not be detected, ultimately resulting in poor ride comfort.
[0104] To this end, this application proposes the following solutions:
[0105] First, when the VLM perceives a slope change in the road ahead by understanding the current driving environment, it can adjust the position of the ROI to include locations with larger curvature changes (i.e., the locations of the above-mentioned obstacles) to avoid missing key obstacles.
[0106] Then, the VLM selects a target perception model, such as calling a stereo binocular perception model to identify obstacles ahead in advance, especially when there are bumpy roads. It can also control and adjust the suspension height to improve comfort.
[0107] Finally, VLM can adjust the weight value of the data collected by the binocular camera (environmental images and depth information, etc.) to 100% (the weight value of other sensor data is zero) to respond to the scene more accurately.
[0108] In addition, given that the stereo binocular algorithm takes up a lot of system resources, the scene recognition model can control the exit of the binocular algorithm model after the vehicle passes the road section where the obstacle is located to save system resources.
[0109] It should also be noted that the above embodiment is divided into multiple steps mainly to explain the operating principle of the environmental perception system in a more organized manner. When the solution is implemented, there may be no obvious time interval or the time interval between the above steps may be very short. From a macro perspective, it can be considered that the VLM and perception model in the environmental perception system work simultaneously.
[0110] See Figure 5 , Figure 5 This is a hardware structure diagram of a vehicle in which an environmental perception device of a vehicle is located, which is shown as an exemplary embodiment. At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510, and of course may also include hardware required for other services. One or more embodiments of the present application can be implemented based on software, such as the processor 502 reading the corresponding computer program from the non-volatile memory 510 into the memory 508 and then running it. Of course, in addition to software implementation, one or more embodiments of the present application do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0111] See Figure 6 , Figure 6 This is a block diagram of an environment perception device for a vehicle, shown in an exemplary embodiment. The environment perception device for a vehicle can be applied to Figure 5 In the vehicle shown, the technical solution of the present application is implemented. The device may include:
[0112] The data acquisition unit 601 is configured to acquire corresponding types of sensor data collected by the multiple sensors for the current driving environment;
[0113] a weight assignment unit 602 configured to call a scene recognition model to identify a current scene type of the current driving environment based at least on the sensor data, and obtain weight values assigned by the scene recognition model to various sensor data that match the current scene type;
[0114] The result generating unit 603 is configured to generate a fusion perception result of the driving sensitive factors in the current driving environment according to various sensor data and their weight values.
[0115] Optionally, the weight distribution unit 602 is specifically configured to include:
[0116] inputting the sensor data and the non-sensor data into a scene recognition model, so that the scene recognition model identifies the current scene type based on the sensor data and the non-sensor data;
[0117] The non-sensor data includes at least one of the following: weather data, lighting data, map data, positioning data of the vehicle, driving data of the vehicle, and traffic rules corresponding to the current location.
[0118] Optionally, the weight allocation unit 602 is specifically configured to:
[0119] Receive the model selection result and weight assignment result output by the scene recognition model, wherein the model selection result is used to represent at least one target perception model selected by the scene recognition model from multiple candidate perception models that matches the current scene type, and each target perception model is used to determine the driving sensitive factors in the current driving environment based on its corresponding target sensor data and generate a corresponding fusion perception result; the weight assignment result is at least used to represent the weight values of various target sensor data.
[0120] Optionally, the result generating unit 603 is specifically configured to:
[0121] Call each target perception model to calculate the corresponding preliminary perception results based on the corresponding target sensor data;
[0122] The fusion module is called to perform weighted operations on each preliminary perception result according to the corresponding weight value, and generates the corresponding fusion perception result of the target driving sensitive factor based on the operation result.
[0123] Optionally, the multiple candidate perception models include at least two of the following:
[0124] Visual perception model, stereo binocular perception model, lidar perception model, millimeter wave radar perception model, ultrasonic radar perception model, OCC occupancy network perception model, and BEVFormer perception model.
[0125] Optionally, the sensor data includes an environment image captured by a camera, and the apparatus further includes a ROI adjustment unit 604 configured to:
[0126] In response to a region of interest (ROI) adjustment instruction output by the scene recognition model, a position of the ROI in the environment image is adjusted so that the adjusted ROI covers the high-risk driving sensitive factors identified by the scene recognition model.
[0127] Optional,
[0128] The resolution of the inner image portion of the ROI region is higher than the resolution of the outer image portion;
[0129] and / or,
[0130] The reduction factor of the inner image portion of the ROI region is lower than that of the outer image portion.
[0131] Optionally, the sensor data includes an environmental image collected by a camera, laser data collected by a lidar sensor, and millimeter wave data collected by a millimeter wave radar, wherein:
[0132] When the current scene type is an irregular obstacle scene, the weight value of the laser data is greater than the weight values of other sensor data;
[0133] When the current scene type is an extreme weather scene, the weight value of the laser data is lower than the weight value of the millimeter wave data;
[0134] When the current scene type is a backlit scene, the weight value of the environment image is lower than the weight value of the laser data and / or lower than the weight value of the millimeter wave data;
[0135] In a case where the current scene type is a bumpy scene or a turning scene, the weight value of the environment image is greater than the weight values of other sensor data.
[0136] Optionally, the system further includes a sensor weight determination unit 605, configured to:
[0137] When there are multiple sensors of any type, the accuracy information of each sensor of that type is obtained, and the scene recognition model is called to determine the weight values of each sensor that is positively correlated with its accuracy based on the accuracy information of each sensor of that type, and the weight value of the sensor data of that type is determined based on the weight values of each sensor of that type.
[0138] Optionally, a driving control unit 606 is further included, configured to:
[0139] When the assisted driving function is turned on for the vehicle, an assisted driving control instruction for the current driving environment is generated according to the fusion perception result of the driving sensitive factors, and the vehicle is controlled according to the assisted driving control instruction.
[0140] Optionally, the scene recognition model is a visual language model VLM.
[0141] The implementation process of the functions and effects of each unit in the device is specifically described in the implementation process of the corresponding steps in the method, which will not be repeated here.
[0142] Accordingly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the vehicle environment perception method as described in any of the above embodiments.
[0143] Accordingly, this specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the vehicle environment perception method as described in any of the above embodiments.
[0144] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0145] The systems, devices, modules, or units described in the embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0146] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0147] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0148] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0149] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0150] The foregoing description is for specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0151] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0152] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0153] The above description is merely a preferred embodiment of one or more embodiments of the present application and is not intended to limit one or more embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application shall be included in the scope of protection of one or more embodiments of the present application.
[0154] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
Claims
1. A vehicle environment perception method, characterized in that: Applied to a vehicle equipped with multiple sensors, the method includes: Obtaining corresponding types of sensor data collected by the multiple sensors for the current driving environment; calling a scene recognition model to identify a current scene type of the current driving environment based at least on the sensor data, and obtaining weight values assigned by the scene recognition model to various sensor data that match the current scene type; According to various sensor data and their weight values, a fusion perception result of driving sensitive factors in the current driving environment is generated.
2. The method according to claim 1, characterized in that The calling of the scene recognition model to identify the current scene type of the current driving environment based at least on the sensor data includes: inputting the sensor data and the non-sensor data into a scene recognition model, so that the scene recognition model identifies the current scene type based on the sensor data and the non-sensor data; The non-sensor data includes at least one of the following: weather data, lighting data, map data, positioning data of the vehicle, driving data of the vehicle, and traffic rules corresponding to the current location.
3. The method according to claim 1, characterized in that The obtaining of weight values assigned by the scene recognition model to various sensor data that match the current scene type includes: Receive the model selection result and weight assignment result output by the scene recognition model, wherein the model selection result is used to represent at least one target perception model selected by the scene recognition model from multiple candidate perception models that matches the current scene type, and each target perception model is used to determine the driving sensitive factors in the current driving environment based on its corresponding target sensor data and generate a corresponding fusion perception result; the weight assignment result is at least used to represent the weight values of various target sensor data.
4. The method according to claim 3, characterized in that Generating a fusion perception result of driving sensitive factors in the current driving environment according to various sensor data and their weight values includes: Call each target perception model to calculate the corresponding preliminary perception results based on the corresponding target sensor data; The fusion module is called to perform weighted operations on each preliminary perception result according to the corresponding weight value, and generates the corresponding fusion perception result of the target driving sensitive factor based on the operation result.
5. The method according to claim 3 or 4, characterized in that The plurality of candidate perception models include at least two of the following: Visual perception model, stereo binocular perception model, lidar perception model, millimeter wave radar perception model, ultrasonic radar perception model, OCC occupancy network perception model, and BEVFormer perception model.
6. The method according to claim 1, characterized in that The sensor data includes an environment image captured by a camera, and the method further includes: In response to a region of interest (ROI) adjustment instruction output by the scene recognition model, a position of the ROI in the environment image is adjusted so that the adjusted ROI covers the high-risk driving sensitive factors identified by the scene recognition model.
7. The method according to claim 6, characterized in that In the environment image, The resolution of the inner portion of the ROI is higher than the resolution of the outer portion; and / or, The reduction factor of the inner portion of the ROI is lower than that of the outer portion thereof.
8. The method according to claim 1, characterized in that The sensor data includes the environment image collected by the camera, the laser data collected by the lidar sensor, and the millimeter wave data collected by the millimeter wave radar, wherein: When the current scene type is an irregular obstacle scene, the weight value of the laser data is greater than the weight values of other sensor data; When the current scene type is an extreme weather scene, the weight value of the laser data is lower than the weight value of the millimeter wave data; When the current scene type is a backlit scene, the weight value of the environment image is lower than the weight value of the laser data and / or lower than the weight value of the millimeter wave data; In a case where the current scene type is a bumpy scene or a turning scene, the weight value of the environment image is greater than the weight values of other sensor data.
9. The method according to claim 1, characterized in that Also includes: When there are multiple sensors of any type, the accuracy information of each sensor of that type is obtained, and the scene recognition model is called to determine the weight values of each sensor that is positively correlated with its accuracy based on the accuracy information of each sensor of that type, and the weight value of the sensor data of that type is determined based on the weight values of each sensor of that type.
10. The method according to claim 1, characterized in that Also includes: When the assisted driving function is turned on for the vehicle, an assisted driving control instruction for the current driving environment is generated according to the fusion perception result of the driving sensitive factors, and the vehicle is controlled according to the assisted driving control instruction.
11. The method according to claim 1, wherein The scene recognition model is a visual language model VLM.
12. A vehicle environment perception device, characterized in that: Applicable to vehicles equipped with multiple sensors, the device includes: a data acquisition unit, configured to acquire corresponding types of sensor data collected by the multiple sensors for the current driving environment; a weight assignment unit, configured to call a scene recognition model to identify a current scene type of the current driving environment based at least on the sensor data, and obtain weight values assigned by the scene recognition model to various sensor data that match the current scene type; The result generating unit is used to generate a fusion perception result of the driving sensitive factors in the current driving environment according to various sensor data and their weight values.
13. A vehicle environment perception system, characterized in that: The system includes a scene recognition model, a fusion module and multiple candidate perception models, among which, The scene recognition model is configured to: identify a current scene type of a current driving environment based at least on input sensor data; and select at least one target perception model that matches the current scene type from the plurality of candidate perception models, and assign a weight value that matches the current scene type to target sensor data corresponding to each target perception model; Each target perception model is used to: calculate the corresponding preliminary perception results based on its corresponding target sensor data; The fusion module is used to perform weighted operations on the preliminary perception results according to corresponding weight values, and generate fused perception results for the driving sensitive factors in the current driving environment according to the operation results.
14. A vehicle comprising: a processor, a memory for storing instructions executable by the processor, and a variety of sensors; The processor implements the method according to any one of claims 1 to 11 by running the executable instructions.
15. A computer program product comprising a computer program and / or instructions, characterized in that When the computer program and / or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
V2X multi-sensor fusion method and device based on scene perception
CN115379408A
Identification method and device based on multiple modes, electronic equipment and storage medium
CN115457519A
Environment sensing method and device, storage medium and vehicle
CN117893978A
Control method of self-driving automobile and self-driving automobile
CN119018180A
Method and device for enhancing interpretability of automatic driving scene
CN119918183A
Cited By
Data fusion method and device based on multiple sensors
CN120802246A
A data fusion method and apparatus based on multiple sensors
CN120802246B