Perception method, model training method, device, equipment, medium and program product
By using multi-view image information and environmental perception model to extract features in the vehicle autonomous driving perception system, the problems of poor flexibility and low efficiency of the existing system are solved, and efficient two-dimensional and three-dimensional vehicle driving environment perception is achieved.
Patent Information
- Application Number
- CN202510092415.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The existing vehicle autonomous driving perception system has problems of poor flexibility and low efficiency, and it is impossible to perform task perception in different dimensions at the same time.
A method of perception of vehicle driving environment is adopted to realize two-dimensional and three-dimensional vehicle driving environment perception tasks by acquiring multi-view image information and extracting multi-view image features and target bird's-eye view image features based on the environment perception model.
It improves the flexibility and perceived efficiency of task perception, reduces the complexity of the perception system, and improves the accuracy of multi-view image features and target aerial view image features.
Smart Images

Figure CN120014599A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle autonomous driving, and specifically to a method for perceiving a vehicle's driving environment, a method for training an environment perception model, a device, equipment, medium, and program product. Background Art
[0002] With the vigorous development and increasing popularization of vehicle intelligence and electrification, more and more vehicles will be equipped with autonomous driving functions. Therefore, it is necessary to perceive and identify the driving environment around the vehicle. In related technologies, the perception system usually cannot perceive tasks in different dimensions at the same time, which has the problem of poor flexibility. In addition, the perception system usually contains multiple independent modules, which require each module to run separately to obtain results, and then summarize and fuse them before outputting them to downstream use, which has the problem of low efficiency. Summary of the invention
[0003] The present application provides a method for perceiving a vehicle driving environment, a method for training an environmental perception model, a device, equipment, a medium and a program product to solve the problems of poor flexibility and low efficiency of the perception system in related technologies.
[0004] In order to achieve the above-mentioned purpose, the present application provides a method for sensing a vehicle driving environment, the method comprising:
[0005] Acquire multi-view image information, and determine the image information of the target type based on the multi-view image information; wherein the target type includes at least one or more types of a forward viewing angle type, a forward narrow viewing angle type, and a periscopic viewing angle type;
[0006] Through the environment perception model, based on the image information of the target type, multi-view image features and target bird's-eye view image features corresponding to the image information of the target type are obtained; wherein the environment perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder;
[0007] Based on the multi-view image features, a two-dimensional vehicle driving environment perception task is performed to obtain a two-dimensional vehicle driving environment perception result, and based on the target bird's-eye view image features, a three-dimensional vehicle driving environment perception task is performed to obtain a three-dimensional vehicle driving environment perception result.
[0008] According to the above technical means, first, the image information of the target type is determined by the acquired multi-view image information, and then, the multi-view image features and the target bird's-eye view image features are obtained based on the image information of the target type through the environmental perception model including the multi-type feature extraction network and the bird's-eye view feature extraction network. Finally, the two-dimensional vehicle driving environment perception task can be performed based on the multi-view image features to obtain the two-dimensional vehicle driving environment perception result, and the three-dimensional vehicle driving environment perception task can be performed based on the target bird's-eye view image features to obtain the three-dimensional vehicle driving environment perception result. In this way, on the one hand, since the perception method of the present application can respectively perform the two-dimensional vehicle driving environment perception task and the three-dimensional vehicle driving environment perception task through the output multi-view image features and the target bird's-eye view image features, it can simultaneously perform task perceptions of different dimensions, thereby improving the flexibility of task perception; on the other hand, compared with the perception system of the prior art including multiple independent modules, since the perception method of the present application is implemented through the environmental perception model, the perception task can be completed through one environmental perception model, while improving the perception efficiency of the vehicle driving environment, it can also reduce the complexity of the perception system.
[0009] Furthermore, through the environmental perception model, based on the image information of the target type, multi-perspective image features and bird's-eye view image features corresponding to the image information of the target type are obtained, including: through a multi-type feature extraction network, based on the image information of the target type, multi-perspective image features corresponding to the image information of the target type are obtained; through a bird's-eye view feature extraction network, based on the multi-perspective image features, camera parameters corresponding to the multi-perspective image information, bird's-eye view image features at historical moments, and bird's-eye view parameters, target bird's-eye view image features corresponding to the image information of the target type are obtained.
[0010] According to the above-mentioned technical means, multi-view image features are obtained through a multi-type feature extraction network, and target bird's-eye view image features are obtained through a bird's-eye view feature extraction network. On the one hand, the accuracy of multi-view image features and target bird's-eye view image features is improved; on the other hand, since the environmental perception model of the present application can perform a variety of different types of perception tasks through continuous training, the scalability performance of the environmental perception model is improved.
[0011] Furthermore, through a multi-type feature extraction network, based on the image information of the target type, multi-view image features corresponding to the image information of the target type are obtained, including: through a backbone module corresponding to the target type, based on the image information of the target type, obtaining a first image feature corresponding to the image information of the target type; through a first fusion module corresponding to the target type, based on the first image feature, obtaining multi-view image features.
[0012] According to the above technical means, the first image feature is obtained through the backbone module corresponding to the target type, and then the multi-view image feature is obtained through the first fusion module corresponding to the target type. The image information of different target types can be processed in parallel through the backbone modules and the first fusion modules corresponding to different target types to obtain the multi-view image feature, which can improve the calculation speed and the accuracy of the multi-view image feature.
[0013] Furthermore, the multi-type feature extraction network includes one or more types of a first sub-network corresponding to a forward-looking perspective type, a second sub-network corresponding to a forward-looking narrow-viewing perspective type, and a third sub-network corresponding to a surrounding perspective type; through a backbone module corresponding to the target type, based on the image information of the target type, a first image feature corresponding to the image information of the target type is obtained, including at least one of the following: when the target type is a forward-looking perspective type, through the backbone module of the first sub-network, based on the image information of the forward-looking perspective type, the corresponding first image sub-feature is obtained; when the target type is a forward-looking narrow-viewing perspective type, through the backbone module of the second sub-network, based on the image information of the forward-looking narrow-viewing perspective type, the second image sub-feature corresponding to the image information of the forward-looking narrow-viewing perspective type is obtained; when the target type is a surrounding perspective type, through the backbone module of the third sub-network, based on the image information of the surrounding perspective type, the third image sub-feature corresponding to the image information of the surrounding perspective type is obtained.
[0014] According to the above technical means, among different target types, by selecting the backbone modules in the sub-networks corresponding to different target types, the image sub-features corresponding to the image information of different target types can be obtained, and multiple sub-networks can be used to process the image information in parallel to obtain the image sub-features, which can improve the calculation speed and the accuracy of the image sub-features.
[0015] Furthermore, a multi-view image feature is obtained based on the first image feature through a first fusion module corresponding to the target type, including at least one of the following: when the target type is a forward-looking perspective type, a first multi-view image feature corresponding to the first image sub-feature is obtained based on the first image sub-feature through a fusion module of the first sub-network; when the target type is a forward-looking narrow-view perspective type, a second multi-view image feature corresponding to the second image sub-feature is obtained based on the second image sub-feature through a fusion module of the second sub-network; when the target type is a peripheral perspective type, a third multi-view image feature corresponding to the third image sub-feature is obtained based on the third image sub-feature through a fusion module of the third sub-network.
[0016] According to the above technical means, among different target types, by selecting the fusion modules in the sub-networks corresponding to different target types, the multi-view image features corresponding to the image information of different target types can be obtained, and multiple sub-networks can be used to process the image sub-features in parallel to obtain the multi-view image features, which can improve the calculation speed and the accuracy of the multi-view image features.
[0017] Furthermore, a two-dimensional vehicle driving environment perception task is performed based on the multi-perspective image features to obtain a two-dimensional vehicle driving environment perception result, including: based on the two-dimensional vehicle driving environment perception task, extracting a target perspective two-dimensional image feature from the multi-perspective image features; performing a two-dimensional vehicle driving environment perception task based on the target perspective image features to obtain a two-dimensional vehicle driving environment perception result.
[0018] According to the above technical means, the two-dimensional image features of the target perspective are extracted from the multi-perspective image features, and the two-dimensional vehicle driving environment perception task is performed through the target perspective image features. The image features of the corresponding perspective can be selected for the two-dimensional vehicle driving environment perception task. Compared with using image features of all perspectives to perform the two-dimensional vehicle driving environment perception task, it can save computing resources and improve the execution speed of the perception task.
[0019] Furthermore, through a bird's-eye view feature extraction network, based on multi-view image features, camera parameters corresponding to the multi-view image information, bird's-eye view image features at historical moments, and bird's-eye view parameters, bird's-eye view image features corresponding to the image information of the target type are obtained, including: through an encoder, based on the multi-view image features, camera parameters corresponding to the multi-view image information, bird's-eye view image features at historical moments, and bird's-eye view parameters, first bird's-eye view image features corresponding to the multi-view image features are obtained; through a decoder, based on the first bird's-eye view image features, target bird's-eye view image features are obtained.
[0020] According to the above technical means, the first bird's-eye view image features corresponding to the multi-view image features are obtained through the encoder, and the target bird's-eye view image features are obtained based on the first bird's-eye view image features through the decoder, thereby improving the accuracy of the target bird's-eye view image features.
[0021] Further, determining the target type of image information based on the multi-perspective image information includes: when the multi-perspective image information is collected by a forward-looking perspective shooting device, using the forward-looking perspective type of image information as the target type of image information; when the multi-perspective image information is collected by a forward-looking narrow-viewing perspective shooting device, using the forward-looking narrow-viewing perspective type of image information as the target type of image information; when the multi-perspective image information is collected by a surrounding-viewing perspective shooting device, using the surrounding-viewing perspective type of image information as the target type of image information; wherein the surrounding-viewing perspective includes at least one of the following: left front perspective, right front perspective, left rear perspective, right rear perspective, and rearward perspective.
[0022] According to the above technical means, the image information of the target type is determined by using the image information collected by the shooting devices with different viewing angles, thereby improving the accuracy of the image information of the target type.
[0023] The present application provides a method for training an environment perception model, the method comprising:
[0024] Acquire a training data set; wherein the training data set includes first image information from multiple perspectives and an environmental perception truth value of at least one perception task corresponding to the first image information;
[0025] Obtaining, through the initial perception model, a predicted perception result of at least one perception task corresponding to the first image information based on the first image information; wherein the initial perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder;
[0026] Determine the loss based on the true value of environmental perception and the predicted perception results;
[0027] The model parameters of the initial perception model are modified based on the loss to obtain the environment perception model.
[0028] According to the above-mentioned technical means, after obtaining the predicted perception results corresponding to the first image information of multiple perspectives through the initial perception model, the total loss of the model training can be determined based on the true value of environmental perception and the predicted perception results; in this way, the initial perception model including the multi-type feature extraction network and the initial bird's-eye view feature extraction network is trained through the loss, so as to improve the prediction effect and performance of the neural network.
[0029] The present application provides a vehicle driving environment sensing device, the vehicle driving environment sensing device comprising:
[0030] A determination unit, used to obtain multi-view image information, and determine the image information of the target type based on the multi-view image information; wherein the target type includes at least one or more types of a forward viewing angle type, a forward narrow viewing angle type, and a periscopic viewing angle type;
[0031] An acquisition unit is used to acquire multi-view image features and target bird's-eye view image features corresponding to the image information of the target type based on the image information of the target type through an environment perception model; wherein the environment perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder;
[0032] The perception unit is used to perform two-dimensional vehicle driving environment perception tasks based on multi-view image features to obtain two-dimensional vehicle driving environment perception results, and to perform three-dimensional vehicle driving environment perception tasks based on target bird's-eye view image features to obtain three-dimensional vehicle driving environment perception results.
[0033] The present application provides a training device for an environment perception model, and the training device for an environment perception model includes:
[0034] A second acquisition unit is used to acquire a training data set; wherein the training data set includes multi-view first image information and an environmental perception truth value of at least one perception task corresponding to the first image information; through an initial perception model, based on the first image information, a predicted perception result of at least one perception task corresponding to the first image information is obtained; wherein the initial perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder;
[0035] A second determination unit, configured to determine the loss based on the environmental perception true value and the predicted perception result;
[0036] The training unit is used to correct the model parameters of the initial perception model based on the loss to obtain the environment perception model.
[0037] The present application provides an electronic device, including a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps in any of the above methods are implemented.
[0038] The present application provides a vehicle device, the vehicle device includes an electronic device, and the vehicle device implements the steps in any of the above methods through the electronic device.
[0039] The present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in any of the above methods are implemented.
[0040] The present application provides a computer program product, including a computer program or instructions, which implement the steps in any of the above methods when executed by a processor.
[0041] The present application provides a vehicle device, the vehicle device includes an electronic device, and the vehicle device implements the steps in any of the above methods through the electronic device.
[0042] Beneficial effects of this application:
[0043] (1) Since the perception method of the present application can respectively perform two-dimensional vehicle driving environment perception tasks and three-dimensional vehicle driving environment perception tasks through the output multi-view image features and target bird's-eye view image features, it can simultaneously perform task perceptions in different dimensions, thereby improving the flexibility of task perception;
[0044] (2) Compared with the perception system of the prior art that includes multiple independent modules, since the perception method of the present application is implemented through an environmental perception model, the perception task can be completed through one environmental perception model, which can improve the perception efficiency of the vehicle driving environment and reduce the complexity of the perception system;
[0045] (3) The accuracy of multi-view image features and target bird's-eye view image features is improved. At the same time, since the environmental perception model of the present application can perform a variety of different types of perception tasks through continuous training, the scalability of the environmental perception model is improved;
[0046] (4) By using the backbone modules corresponding to different target types and the first fusion module to process the image information of different target types in parallel to obtain multi-view image features, the calculation speed can be improved and the accuracy of the multi-view image features can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic diagram of the implementation process of a vehicle driving environment perception method provided in an embodiment of the present application Figure 1 ;
[0048] Figure 2 A schematic diagram of a collection viewing angle provided in an embodiment of the present application;
[0049] Figure 3 A schematic diagram of an environment perception model proposed in an embodiment of the present application;
[0050] Figure 4 A schematic diagram of a multi-type feature extraction network proposed in an embodiment of the present application;
[0051] Figure 5 A schematic diagram of a bird's-eye view feature extraction network proposed in an embodiment of the present application;
[0052] Figure 6 A schematic diagram of obtaining multi-view image features provided in an embodiment of the present application;
[0053] Figure 7 A schematic diagram of an implementation flow of a method for training an environment perception model provided in an embodiment of the present application;
[0054] Figure 8 A schematic diagram of the implementation process of a vehicle driving environment perception method provided in an embodiment of the present application Figure 2 ;
[0055] Fig. 9 A schematic diagram of obtaining a target bird's-eye view image feature provided by an embodiment of the present application;
[0056] Fig.10 A schematic diagram of a method for sensing a vehicle driving environment provided in an embodiment of the present application;
[0057] Fig.11 A schematic diagram of the structure of a vehicle driving environment sensing device provided in an embodiment of the present application;
[0058] Fig.12 A schematic diagram of the structure of a training device for an environmental perception model provided in an embodiment of the present application;
[0059] Fig.13 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of the present application;
[0060] Fig.14 This is a schematic diagram of the composition structure of the vehicle equipment proposed in the embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.
[0062] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0063] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0064] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0066] With the vigorous development and increasing popularity of vehicle intelligence and electrification, more and more vehicles will be equipped with autonomous driving functions. Therefore, it is necessary to first perceive and identify the driving environment around the vehicle, and then input it to the downstream modules of autonomous driving for planning, control and making correct decisions.
[0067] In related technologies, a perception system can be established through a deep neural network to perceive and identify the environment and obstacles around the vehicle based on the surrounding image data collected by the vehicle. The perception system is low-cost and can be deployed on the vehicle to perceive the real-time driving environment. However, the perception system usually contains multiple independent modules, and each module needs to run separately to obtain results, and then aggregate and fuse them before they can be output to downstream use. In addition, the perception system usually cannot perform task perceptions of different dimensions at the same time, and has a problem of poor flexibility.
[0068] The embodiment of the present application provides a method for perceiving a vehicle driving environment. First, the image information of the target type is determined by acquiring multi-view image information. Then, the multi-view image features and the target bird's-eye view image features are acquired based on the image information of the target type through an environment perception model including a multi-type feature extraction network and a bird's-eye view feature extraction network. Finally, a two-dimensional vehicle driving environment perception task can be performed based on the multi-view image features to obtain a two-dimensional vehicle driving environment perception result, and a three-dimensional vehicle driving environment perception task can be performed based on the target bird's-eye view image features to obtain a three-dimensional vehicle driving environment perception result. In this way, on the one hand, since the perception method of the present application can respectively perform the two-dimensional vehicle driving environment perception task and the three-dimensional vehicle driving environment perception task through the output multi-view image features and the target bird's-eye view image features, task perceptions of different dimensions can be performed simultaneously, thereby improving the flexibility of task perception. On the other hand, compared with the perception system of the prior art including multiple independent modules, since the perception method of the present application is implemented through an environment perception model, the perception task can be completed through one environment perception model, thereby improving the perception efficiency of the vehicle driving environment and reducing the complexity of the perception system.
[0069] Below, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application.
[0070] Figure 1 A schematic diagram of the implementation process of a vehicle driving environment perception method provided in an embodiment of the present application Figure 1 ,like Figure 1 As shown, the method includes steps 101 to 103, wherein:
[0071] Step 101, acquiring multi-view image information, and determining image information of a target type based on the multi-view image information; wherein the target type includes at least one or more types of a forward viewing angle type, a forward narrow viewing angle type, and a peripheral viewing angle type.
[0072] Here, multi-view image information refers to visual information represented in digital form and composed of pixels obtained from multiple acquisition perspectives. Multi-view image information includes image information under multiple perspectives. The acquisition perspective may include, but is not limited to, at least one of: a front perspective, a front narrow perspective, a left front perspective, a right front perspective, a left rear perspective, a right rear perspective, a rear perspective, etc.
[0073] In some embodiments, the multi-view image information may be directly acquired by the first device through a configured image acquisition device. The image acquisition device is a device with the functions of acquiring data, recording data, and transmitting data. The image acquisition device may include, but is not limited to, at least one of a camera, a radar system, and the like.
[0074] In some implementations, the image acquisition device may be set at multiple acquisition viewing angles, and then in response to the image acquisition device triggering a data acquisition function, multi-view image information may be acquired.
[0075] In some implementations, when collecting image data through a camera, the camera may be set at multiple positions to obtain multi-view image information. In implementation, one camera may correspond to one collection viewing angle.
[0076] For example, in some embodiments, the image acquisition device can be set based on different positions, so that corresponding multi-view image information can be acquired for multiple different acquisition perspectives at the same time. Figure 2 This is a schematic diagram of obtaining multi-view image information proposed in an embodiment of the present application, such as Figure 2 As shown, the image acquisition device can be 7 cameras set at 7 different positions, and then the 7 cameras are used to acquire images, so that multi-view image information of 7 acquisition perspectives at any time can be obtained, and then the acquired multi-view image information can be sent to the environment perception model, and accordingly, the environment perception model can obtain the multi-view image information. Among them, the 7 acquisition perspectives can include a front view perspective 21, a front narrow view perspective 22, a left front view perspective 23, a right front view perspective 24, a left rear view perspective 25, a right rear view perspective 26, and a rear view perspective 27.
[0077] The method for acquiring multi-view image information may include but is not limited to: reading acquired data from an image acquisition device to acquire multi-view image information, receiving acquired data sent by an image acquisition device to acquire multi-view image information, etc. For example, the image acquisition device has a function of storing data, and the execution device of the present application can acquire multi-view image information by reading data in the memory of the image acquisition device. For another example, the image acquisition device has a function of transmitting data, and by establishing a communication connection between the image acquisition device and the execution device of the present application, the execution device of the present application can receive data sent by the image acquisition device to acquire multi-view image information.
[0078] In some implementations, since different acquisition perspectives have different importance, different acquisition perspectives can be divided according to their importance, so as to obtain different types, that is, to divide target types. Among them, the target type can include at least one or more types of the forward viewing perspective type, the forward narrow viewing perspective type, and the surrounding viewing perspective type.
[0079] During implementation, it is assumed that the multiple perspectives include a forward perspective, a forward narrow perspective, and a peripheral perspective. Since the forward perspective and the forward narrow perspective contain more important information, the forward perspective can be divided into a forward perspective type, and the forward narrow perspective can be divided into a forward narrow perspective type. At the same time, other acquisition perspectives except the forward perspective and the forward narrow perspective can be used as peripheral perspectives and divided into peripheral perspective types accordingly.
[0080] In some implementations, a first correspondence between a capture perspective and a target type may be established. After acquiring multi-perspective image information, image information of the target type may be determined based on the capture perspective of the multi-perspective image information and the first correspondence.
[0081] In some implementations, when acquiring multi-view image information through a multi-view camera, the exposure time of the camera at each view needs to be synchronized as much as possible to synchronize the acquisition time of the multi-view image information.
[0082] In some embodiments, when the multi-perspective image information is collected by a forward-viewing perspective shooting device, the image information of the forward-viewing perspective type is used as the image information of the target type. When the multi-perspective image information is collected by a forward-viewing narrow-viewing perspective shooting device, the image information of the forward-viewing narrow-viewing perspective type is used as the image information of the target type. When the multi-perspective image information is collected by a surrounding-viewing perspective shooting device, the image information of the surrounding-viewing perspective type is used as the image information of the target type. The surrounding-viewing perspective includes at least one of the following: left front perspective, right front perspective, left rear perspective, right rear perspective, and rear-view perspective.
[0083] Step 102, through the environmental perception model, based on the image information of the target type, obtain multi-view image features and target bird's-eye view image features corresponding to the image information of the target type; wherein the environmental perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder.
[0084] Here, the environment perception model is a model with data processing capability. In some embodiments, the image information of the target type is input into the environment perception model, and the environment perception model outputs multi-view image features and target bird's-eye view image features corresponding to the image information of the target type.
[0085] In an embodiment of the present application, the environmental perception model may include a multi-type feature extraction network and a bird's-eye view feature extraction network. Among them, the multi-type feature extraction network is used to extract image features of different types of input image information. The bird's-eye view feature extraction network can be used to combine the acquisition parameters (such as camera parameters) corresponding to the acquisition of multi-view image information to perform spatial feature conversion and bird's-eye view (Bird-Eye-View, BEV) feature fusion on the input image features, thereby obtaining the corresponding bird's-eye view image features.
[0086] For example, in some embodiments, Figure 3 This is a schematic diagram of the environment perception model proposed in the embodiment of the present application, such as Figure 3 As shown, the environment perception model 30 may include a multi-type feature extraction network 31 and a bird's-eye view feature extraction network 32. The multi-type feature extraction network 32 may extract image features based on the input image information of different target types through different backbone modules corresponding to the target types, and the bird's-eye view feature extraction network 32 may be used to perform spatial feature conversion and BEV feature fusion on the input image features, thereby obtaining corresponding bird's-eye view image features.
[0087] For example, in some embodiments, Figure 4 A schematic diagram of a multi-type feature extraction network proposed in an embodiment of the present application, such as Figure 4 As shown, the multi-type feature extraction network 40 includes a first sub-network 41, a second sub-network 42 and a third sub-network 43 corresponding to different target types, wherein the first sub-network 41, the second sub-network 42 and the third sub-network 43 respectively include a backbone module and a first fusion module. The image information of the target type is input into the sub-network corresponding to the image information of the target type, and the multi-view image features corresponding to the image information of the target type can be obtained through the backbone module and the first fusion module of the sub-network.
[0088] For example, in some embodiments, Figure 5 This is a schematic diagram of a bird's-eye view feature extraction network proposed in an embodiment of the present application, such as Figure 5 As shown, the bird's-eye view feature extraction network 50 includes an encoder 51 and a decoder 52, wherein the spatial feature conversion and BEV feature fusion are performed on the input multi-view image features in combination with the acquisition parameters corresponding to the multi-view image information, so as to obtain the corresponding target bird's-eye view image features.
[0089] Multi-view image features refer to information obtained by extracting features from image data at multiple acquisition perspectives. Multi-view image features may include, but are not limited to, at least one of color features, texture features, shape features, etc. In some embodiments, the multi-view image feature may be at least one. Different acquisition perspectives have their own corresponding multi-view image features.
[0090] The target bird's-eye view image feature refers to a feature having multi-dimensional information after the perspective conversion process. In some embodiments, the target bird's-eye view image feature may be at least one. Different acquisition perspectives have their own corresponding target bird's-eye view image features.
[0091] In some implementations, the target bird's-eye view image features are features in vector form.
[0092] In some embodiments, the target type image information is input into a multi-type feature extraction network to obtain multi-view image features corresponding to the target type image information. The multi-view image features and acquisition parameters (e.g., camera parameters) corresponding to the target type image information are input into a bird's-eye view feature extraction network to obtain target bird's-eye view image features corresponding to the target type image information.
[0093] In some embodiments, the multi-type feature extraction network includes different sub-networks corresponding to different target types. The sub-networks include a backbone module and a first fusion module, the backbone module is used to extract semantic information of multi-view image information. The first fusion module is used to fuse the semantic information of multi-view image information and output multi-view image features.
[0094] In some implementations, the bird's eye view feature extraction network includes an encoder and a decoder.
[0095] In some embodiments, a multi-type feature extraction network can be used to obtain multi-perspective image features corresponding to the image information of the target type based on the image information of the target type, and a bird's-eye view feature extraction network can be used to obtain target bird's-eye view image features corresponding to the image information of the target type based on the multi-perspective image features, camera parameters corresponding to the multi-perspective image information, bird's-eye view image features at historical moments, and bird's-eye view parameters.
[0096] Step 103, performing a two-dimensional vehicle driving environment perception task based on the multi-view image features to obtain a two-dimensional vehicle driving environment perception result, and performing a three-dimensional vehicle driving environment perception task based on the target bird's-eye view image features to obtain a three-dimensional vehicle driving environment perception result.
[0097] Here, the two-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: lane line detection task, license plate detection task, face detection task, etc. The two-dimensional vehicle driving environment perception result is the result obtained by performing the two-dimensional vehicle driving environment perception task. The two-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: pixel point category, license plate presence, face presence, etc.
[0098] The three-dimensional vehicle driving environment perception task may include, but is not limited to: at least one of: vehicle detection task, pedestrian detection task, cyclist detection task, lane line detection task, drivable area detection task, etc. The three-dimensional vehicle driving environment perception result includes, but is not limited to: at least one of: three-dimensional detection frame information, drivable area information, lane line category, etc. The three-dimensional detection frame information may include, but is not limited to: at least one of: the center point of the three-dimensional detection frame, the length, width and height of the three-dimensional detection frame, the orientation angle of the three-dimensional detection frame, etc. The lane line category may include, but is not limited to: one of: dotted line, solid line, fishbone line, etc.
[0099] In some embodiments, the lane detection task is a semantic segmentation task, and the final class can be obtained by classifying each pixel in the image. The pixel category can be at least one of background, single line, double line, deceleration lane line, curb, wide line, etc. The license plate detection task is to detect whether a license plate appears in the image and the position of the license plate in the image. The face detection task is to detect whether a face exists in the image and the position of the face in the image.
[0100] In some implementations, the license plate detection task and the face detection task can be used to implement data privacy protection functions.
[0101] In some implementations, different vehicle driving environment perception tasks may correspond to task heads with different functions. Task heads corresponding to two-dimensional vehicle driving environment perception tasks and three-dimensional vehicle driving environment perception tasks may be selected, and then two-dimensional vehicle driving environment perception tasks may be performed based on multi-view image features to obtain two-dimensional vehicle driving environment perception results, and three-dimensional vehicle driving environment perception tasks may be performed based on target bird's-eye view image features to obtain three-dimensional vehicle driving environment perception results.
[0102] In some embodiments, the task head corresponding to the two-dimensional vehicle driving environment perception task may include, but is not limited to: at least one of a lane line detection task head, a license plate detection task head, a face detection task head, etc. The task head corresponding to the three-dimensional vehicle driving environment perception task may include, but is not limited to: at least one of a vehicle detection task head, a pedestrian detection task head, a cyclist detection task head, a lane line detection task head, a drivable area detection task head, etc.
[0103] In some embodiments, the three-dimensional detection frame can be a seven-dimensional tensor [x, y, z, l, w, h, yaw], where x, y, z respectively represent the three-dimensional spatial coordinates of the center point of the three-dimensional detection frame, l, w, h are the length, width and height of the three-dimensional detection frame, and yaw represents the angle between the direction of the three-dimensional detection frame and the x-axis of the vehicle's own coordinate system.
[0104] In some embodiments, when a new vehicle driving environment perception task is needed, the vehicle driving environment perception task can be performed through the vehicle driving environment perception task head based on multi-perspective image features and / or target bird's-eye view image features, so that the vehicle driving environment perception method of the present application is scalable and can be further expanded to perform other vehicle driving environment perception tasks.
[0105] In some embodiments, based on the two-dimensional vehicle driving environment perception task, the target perspective two-dimensional image features can be extracted from the multi-perspective image features, and then the two-dimensional vehicle driving environment perception task can be performed based on the target perspective two-dimensional image features to obtain the two-dimensional vehicle driving environment perception results.
[0106] In some embodiments, based on the collection time of multi-view image information, a corresponding perception result can be obtained for at least one perception task through an environment perception model. In other words, the image information and the perception result correspond to each other based on the collection time of the image information.
[0107] Exemplarily, in some embodiments, assuming that the acquisition time of multi-perspective image information is t, and at least one perception task includes a vehicle detection task, a pedestrian detection task, a cyclist detection task, and a lane line detection task, then the corresponding perception results obtained through the environmental perception model for at least one perception task can be that there is a vehicle at time t, there are no pedestrians at time t, there is no cyclist at time t, and the lane line is a solid line at time t.
[0108] In the embodiment of the present application, first, the image information of the target type is determined by the acquired multi-view image information, then, the multi-view image features and the target bird's-eye view image features are acquired based on the image information of the target type through the environment perception model including the multi-type feature extraction network and the bird's-eye view feature extraction network, and finally, the two-dimensional vehicle driving environment perception task can be performed based on the multi-view image features to obtain the two-dimensional vehicle driving environment perception result, and the three-dimensional vehicle driving environment perception task can be performed based on the target bird's-eye view image features to obtain the three-dimensional vehicle driving environment perception result. In this way, on the one hand, since the perception method of the present application can respectively perform the two-dimensional vehicle driving environment perception task and the three-dimensional vehicle driving environment perception task through the output multi-view image features and the target bird's-eye view image features, it can simultaneously perform task perceptions of different dimensions, thereby improving the flexibility of task perception; on the other hand, compared with the perception system of the prior art including multiple independent modules, since the perception method of the present application is implemented through the environment perception model, the perception task can be completed through one environment perception model, while improving the perception efficiency of the vehicle driving environment, it can also reduce the complexity of the perception system.
[0109] In some implementations, step 102 includes step 11 and step 12, that is, when the multi-view image features and the bird's-eye view image features corresponding to the image information of the target type are obtained based on the image information of the target type through the environment perception model, step 11 and step 12 may be performed, wherein:
[0110] Step 11, through a multi-type feature extraction network, based on the image information of the target type, obtain multi-view image features corresponding to the image information of the target type.
[0111] Here, the multi-type feature extraction network can extract features from the input data. The multi-type feature extraction network includes different sub-networks corresponding to different target types, wherein the sub-network includes a backbone module and a first fusion module.
[0112] In some implementations, the backbone module may be used to extract semantic information of the multi-view image information. The first fusion module is used to fuse the semantic information of the multi-view image information to obtain multi-view image features.
[0113] In some embodiments, since the multi-type feature extraction network includes different sub-networks corresponding to different target types, and the sub-networks include a backbone module and a first fusion module, the first image feature corresponding to the image information of the target type can be obtained based on the image information of the target type through the backbone module corresponding to the target type, and then the multi-view image feature can be obtained based on the first image feature through the first fusion module corresponding to the target type.
[0114] Step 12, through the bird's-eye view feature extraction network, based on the multi-view image features, the camera parameters corresponding to the multi-view image information, the bird's-eye view image features at historical moments, and the bird's-eye view parameters, obtain the target bird's-eye view image features corresponding to the image information of the target type.
[0115] Here, the bird's-eye view feature extraction network can perform perspective conversion and fusion processing on the input data. The bird's-eye view feature extraction network can include an encoder and a decoder.
[0116] Camera parameters are parameters that describe the performance and functions of the camera that captures the viewing angle. The camera parameters corresponding to the multi-view image information may include, but are not limited to: at least one of shutter speed, sensitivity, white balance, exposure compensation, focus mode, depth of field, focal length, aperture, etc. Bird's-eye view parameters are grid-shaped learnable parameters that can query and aggregate features in multi-view images through an attention mechanism.
[0117] The bird's-eye view image features of historical moments refer to the bird's-eye view image features acquired at historical moments. In some embodiments, the bird's-eye view image features of historical moments may be the bird's-eye view image features acquired at the last moment, or may be the bird's-eye view image features acquired over a period of history, or may be all the bird's-eye view image features acquired historically.
[0118] In some embodiments, the bird's-eye view feature extraction network can be supported by a deep learning model (Transformer_) based on the self-attention mechanism. Transformer can include a temporal self-attention layer and a spatial cross-attention layer. These two modules use a deformable attention module to convert the global attention mechanism of Transformer into a local attention mechanism, thereby effectively reducing the training time and improving the convergence speed of Transformer.
[0119] In some implementations, the Spatial CrossAttention layer is used to extract spatial features. The Temporal Self Attention layer is used to perform feature fusion. In implementation, the multi-view image features can be first extracted through the SpatialCrossAttention layer, and then the bird's-eye view image features of the historical moments are fused through the Temporal Self Attention layer to finally obtain the target bird's-eye view image features corresponding to the image information of the target type.
[0120] In some embodiments, the encoder can perform perspective conversion and fusion processing on the input data, and the decoder can perform format conversion on the data, so that when performing a three-dimensional vehicle driving environment perception task, the task head can process the format-converted data.
[0121] In some embodiments, a first bird's-eye view image feature corresponding to the multi-view image feature can be obtained through an encoder based on the multi-view image feature, the camera parameters corresponding to the multi-view image information, the bird's-eye view image features at historical moments, and the bird's-eye view parameters, and then a target bird's-eye view image feature can be obtained through a decoder based on the first bird's-eye view image feature.
[0122] In the implementation manner of the present application, multi-view image features are obtained through a multi-type feature extraction network, and target bird's-eye view image features are obtained through a bird's-eye view feature extraction network. On the one hand, the accuracy of multi-view image features and target bird's-eye view image features is improved; on the other hand, since the environmental perception model of the present application can perform a variety of different types of perception tasks through continuous training, the scalability performance of the environmental perception model is improved.
[0123] In some embodiments, step 11 includes step 111 and step 112, that is, when the multi-view image features corresponding to the image information of the target type are obtained based on the image information of the target type through the multi-type feature extraction network, step 111 and step 112 may be performed, wherein:
[0124] Step 111 , obtaining a first image feature corresponding to the image information of the target type based on the image information of the target type through a backbone module corresponding to the target type.
[0125] Here, the backbone module can be any suitable network, for example, ResNet, U-Net, etc. The first image feature includes at least one scale image feature. The first image feature can include but is not limited to: at least one of color feature, texture feature, shape feature, etc.
[0126] In some implementations, different target types may correspond to different backbone modules or the same backbone module. The capacities of the backbone modules corresponding to different target types may be the same or different.
[0127] In some implementations, the image information of the target type may be downsampled by a backbone module of the target type to obtain a first image feature of at least one scale.
[0128] In some embodiments, when the target type is a forward-looking perspective type, the first image sub-feature corresponding to the forward-looking perspective type image information can be obtained based on the image information of the forward-looking perspective type through the backbone module of the first sub-network.
[0129] In some embodiments, when the target type is a forward-looking narrow viewing angle type, the second image sub-feature corresponding to the forward-looking narrow viewing angle type image information can be obtained based on the image information of the forward-looking narrow viewing angle type through the backbone module of the second sub-network.
[0130] In some embodiments, when the target type is a panoramic perspective type, the third image sub-feature corresponding to the panoramic perspective type image information can be obtained based on the image information of the panoramic perspective type through the backbone module of the third sub-network.
[0131] Step 112: acquiring multi-view image features based on the first image features through a first fusion module corresponding to the target type.
[0132] Here, the first fusion module may be any suitable network, for example, Feature Pyramid Networks (FPN), Pyramid Pooling Module (PPM), etc. The multi-view image features may include but are not limited to: at least one of color features, texture features, shape features, etc.
[0133] In some implementations, different target types may correspond to different first fusion modules or may correspond to the same first fusion module. The capacities of the first fusion modules corresponding to different target types may be the same or different.
[0134] In some embodiments, at least one scale of first image features may be selected from the first image features, and the at least one scale of first image features may be input into a first fusion module corresponding to the target type to obtain multi-view image features. For example, the image information of the target type may be downsampled five times by a backbone module of the target type to obtain first image features of five scales, and then the last three layers of first image features may be selected from the first image features of the five scales, and the last three layers of first image features may be input into a first fusion module corresponding to the target type to obtain multi-view image features.
[0135] In some embodiments, when the target type is a front-view perspective type, a first multi-view image feature corresponding to the first image sub-feature can be obtained based on the first image sub-feature through a fusion module of the first sub-network.
[0136] In some embodiments, when the target type is a forward-looking narrow-viewing type, a second multi-viewing image feature corresponding to the second image sub-feature can be obtained based on the second image sub-feature through a fusion module of the second sub-network.
[0137] In some embodiments, when the target type is a panoramic viewing angle type, a third multi-view image feature corresponding to the third image sub-feature can be obtained based on the third image sub-feature through a fusion module of the third sub-network.
[0138] In an embodiment of the present application, a first image feature is obtained through a backbone module corresponding to the target type, and then a multi-view image feature is obtained through a first fusion module corresponding to the target type. The image information of different target types can be processed in parallel by backbone modules and first fusion modules corresponding to different target types to obtain multi-view image features, which can improve the calculation speed and the accuracy of the multi-view image features.
[0139] In some embodiments, the multi-type feature extraction network includes one or more types of a first subnetwork corresponding to a forward-looking perspective type, a second subnetwork corresponding to a forward-looking narrow-viewing perspective type, and a third subnetwork corresponding to a peripheral-viewing perspective type; step 111 includes at least one of steps 111a to 111c, that is, when a first image feature corresponding to the image information of the target type is obtained based on the image information of the target type through a trunk module corresponding to the target type, at least one of steps 111a to 111c may be performed, wherein:
[0140] Step 111a, when the target type is a forward-looking perspective type, a first image sub-feature corresponding to the forward-looking perspective type image information is obtained based on the image information of the forward-looking perspective type through a backbone module of the first sub-network.
[0141] Here, the backbone module of the first sub-network can be any suitable network, for example, ResNet, U-Net, etc. The first image sub-feature is obtained by processing the front-view type image information. The first image sub-feature may include but is not limited to: at least one of color features, texture features, shape features, etc.
[0142] In some implementations, the first sub-network, the second sub-network, and the third sub-network may be the same network or different networks.
[0143] In some implementations, the first sub-network includes a backbone module and a fusion module. The second sub-network includes a backbone module and a fusion module. The third sub-network includes a backbone module and a fusion module.
[0144] In some implementations, the front-view perspective type image information may be downsampled by a backbone module of the first sub-network to obtain a first image sub-feature of at least one scale.
[0145] Step 111b, when the target type is a forward-looking narrow-viewing angle type, a second image sub-feature corresponding to the forward-looking narrow-viewing angle type image information is obtained based on the image information of the forward-looking narrow-viewing angle type through a backbone module of the second sub-network.
[0146] Here, the backbone module of the second sub-network can be any suitable network, for example, ResNet, U-Net, etc. The second image sub-feature is obtained by processing the front-view narrow-view type image information. The second image sub-feature may include but is not limited to: at least one of color features, texture features, shape features, etc.
[0147] In some implementations, the backbone module of the second sub-network may be the same as or different from the backbone module of the first sub-network.
[0148] In some implementations, the front-view narrow-viewing angle type image information may be downsampled by a backbone module of the second sub-network to obtain a second image sub-feature of at least one scale.
[0149] Step 111c, when the target type is a panoramic viewing angle type, a third image sub-feature corresponding to the image information of the panoramic viewing angle type is obtained based on the image information of the panoramic viewing angle type through the backbone module of the third sub-network.
[0150] Here, the backbone module of the third sub-network can be any suitable network, for example, ResNet, U-Net, etc. The third image sub-feature is obtained by processing the image information of the panoramic viewing angle type. The third image sub-feature may include but is not limited to: at least one of color features, texture features, shape features, etc.
[0151] In some implementations, the backbone module of the third sub-network, the backbone module of the second sub-network, and the backbone module of the first sub-network may be the same or different.
[0152] In some implementations, the peripheral viewing angle type image information may be downsampled by a backbone module of the third sub-network to obtain a third image sub-feature of at least one scale.
[0153] In an embodiment of the present application, among different target types, image sub-features corresponding to image information of different target types are obtained by selecting the backbone modules in the sub-networks corresponding to the different target types. Multiple sub-networks can be used to process image information in parallel to obtain image sub-features, which can improve the calculation speed while also improving the accuracy of the image sub-features.
[0154] In some embodiments, step 112 includes at least one of steps 112a to 112c, that is, when acquiring multi-view image features based on the first image features through the first fusion module corresponding to the target type, at least one of steps 112a to 112c may be performed, wherein:
[0155] Step 112a, when the target type is a front-view perspective type, a first multi-view image feature corresponding to the first image sub-feature is obtained based on the first image sub-feature through a fusion module of the first sub-network.
[0156] Here, the fusion module of the first sub-network may be any suitable network, for example, FPN, PPM, etc. The first multi-view image features may include but are not limited to: at least one of color features, texture features, shape features, and the like.
[0157] In some embodiments, at least one scale of the first image sub-features may be selected from the first image sub-features, and the at least one scale of the first image sub-features may be input into a fusion module of the first sub-network of the first sub-network to obtain a first multi-view image feature.
[0158] Step 112b, when the target type is a front-view narrow-viewing type, a second multi-viewing image feature corresponding to the second image sub-feature is obtained based on the second image sub-feature through a fusion module of the second sub-network.
[0159] Here, the fusion module of the second sub-network may be any suitable network, for example, FPN, PPM, etc. The second multi-view image features may include but are not limited to: at least one of color features, texture features, shape features, and the like.
[0160] In some implementations, the fusion module of the second sub-network may be the same as or different from the fusion module of the first sub-network.
[0161] In some embodiments, at least one scale of the second image sub-features may be selected from the second image sub-features, and the at least one scale of the second image sub-features may be input into a fusion module of the second sub-network of the second sub-network to obtain a second multi-view image feature.
[0162] Step 112c, when the target type is a panoramic viewing angle type, a third multi-view image feature corresponding to the third image sub-feature is obtained based on the third image sub-feature through a fusion module of the third sub-network.
[0163] Here, the fusion module of the third sub-network may be any suitable network, for example, FPN, PPM, etc. The third multi-view image features may include but are not limited to: at least one of color features, texture features, shape features, and the like.
[0164] In some implementations, the fusion module of the third sub-network, the fusion module of the second sub-network, and the fusion module of the first sub-network may be the same or different.
[0165] In some embodiments, at least one scale of the third image sub-feature can be selected from the third image sub-features, and the at least one scale of the third image sub-feature can be input into a fusion module of the third sub-network of the third sub-network to obtain a third multi-view image feature.
[0166] Figure 6 A schematic diagram of obtaining multi-view image features provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the multi-view image feature 610 includes a first multi-view image feature, a second angle image feature, and a third multi-view image feature. The forward-viewing perspective type image information 61 is input into the backbone module 62 of the first sub-network to obtain the first image sub-feature, and then the first image sub-feature is input into the fusion module 63 of the first sub-network to obtain the first multi-view image feature. The forward-viewing narrow-viewing perspective type image information 64 is input into the backbone module 65 of the second sub-network to obtain the second image sub-feature, and then the second image sub-feature is input into the fusion module 66 of the second sub-network to obtain the second multi-view image feature. The circumferential-viewing perspective type image information 37 is input into the backbone module 68 of the third sub-network to obtain the third image sub-feature, and then the third image sub-feature is input into the fusion module 69 of the third sub-network to obtain the third multi-view image feature.
[0167] In an embodiment of the present application, among different target types, multi-view image features corresponding to image information of different target types are obtained by selecting fusion modules in sub-networks corresponding to different target types. Multiple sub-networks can be used to process image sub-features in parallel to obtain multi-view image features, which can improve the calculation speed while also improving the accuracy of multi-view image features.
[0168] In some implementations, the “performing a two-dimensional vehicle driving environment perception task based on multi-view image features to obtain a two-dimensional vehicle driving environment perception result” in step 103 includes steps 131 and 132, that is, when performing a two-dimensional vehicle driving environment perception task based on multi-view image features to obtain a two-dimensional vehicle driving environment perception result, steps 131 and 132 may be performed, wherein:
[0169] Step 131 , based on the two-dimensional vehicle driving environment perception task, extract the target perspective two-dimensional image features from the multi-perspective image features.
[0170] Here, since the two-dimensional vehicle driving environment perception task is mainly a lane line detection task, it is necessary to extract the target perspective two-dimensional image features from the multi-perspective image features containing multiple acquisition perspectives based on the two-dimensional vehicle driving environment perception task.
[0171] In some implementations, the target viewing angle may include a partial viewing angle in the acquisition viewing angle, or may include all partial viewing angles in the acquisition viewing angle.
[0172] In some embodiments, the target perspective may include, but is not limited to, at least one of a left front perspective, a right front perspective, a left rear perspective, a right rear perspective, and the like.
[0173] In some embodiments, since the multi-perspective image features corresponding to each acquisition perspective are independent of each other, after determining the target perspective based on the two-dimensional vehicle driving environment perception task, the target perspective two-dimensional image features can be extracted from the multi-perspective image features according to the target perspective.
[0174] Step 132, performing a two-dimensional vehicle driving environment perception task based on the target perspective image features to obtain a two-dimensional vehicle driving environment perception result.
[0175] Here, the two-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: lane line detection task, license plate detection task, face detection task, etc. The two-dimensional vehicle driving environment perception result is the result obtained by performing the two-dimensional vehicle driving environment perception task. The two-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: pixel point category, license plate presence, face presence, etc.
[0176] In some implementations, different two-dimensional vehicle driving environment perception tasks may correspond to task heads with different functions. A task head corresponding to the two-dimensional vehicle driving environment perception task may be selected, and then the two-dimensional vehicle driving environment perception task is performed based on the target perspective image features to obtain a two-dimensional vehicle driving environment perception result. The task head corresponding to the two-dimensional vehicle driving environment perception task may include, but is not limited to: at least one of a lane line detection task head, a license plate detection task head, and a face detection task head.
[0177] In an implementation manner of the present application, two-dimensional image features of a target perspective are extracted from multi-perspective image features, and a two-dimensional vehicle driving environment perception task is performed through the target perspective image features. Image features of corresponding perspectives can be selected for the two-dimensional vehicle driving environment perception task. Compared with using image features of all perspectives to perform the two-dimensional vehicle driving environment perception task, it can save computing resources and at the same time improve the execution speed of the perception task.
[0178] In some embodiments, step 12 includes step 121 and step 122, that is, when the bird's-eye view image features corresponding to the image information of the target type are obtained based on the multi-view image features, the camera parameters corresponding to the multi-view image information, the bird's-eye view image features at the historical moment, and the bird's-eye view parameters through the bird's-eye view feature extraction network, step 121 and step 122 may be performed, wherein:
[0179] Step 121 , obtaining, through an encoder, first bird's-eye view image features corresponding to the multi-view image features based on the multi-view image features, camera parameters corresponding to the multi-view image information, bird's-eye view image features at historical moments, and bird's-eye view parameters.
[0180] Here, the camera parameters corresponding to the multi-view image information may include but are not limited to: at least one of: shutter speed, sensitivity, white balance, exposure compensation, focus mode, depth of field, focal length, aperture, etc.
[0181] The bird's-eye view image features of historical moments refer to the bird's-eye view image features acquired at historical moments. In some embodiments, the bird's-eye view image features of historical moments may be the bird's-eye view image features acquired at the previous moment, or may be the bird's-eye view image features acquired within a period of history, or may be all the bird's-eye view image features acquired historically.
[0182] In some embodiments, multi-view image features, camera parameters corresponding to the multi-view image information, bird's-eye view image features at historical moments, and bird's-eye view parameters are input into an encoder. In the encoder, spatial features and temporal features in the bird's-eye view image features at historical moments are searched through the bird's-eye view parameters. Then, the multi-view image features, camera parameters corresponding to the multi-view image information, and spatial features and temporal features in the bird's-eye view image features at historical moments are fused to obtain a first bird's-eye view image feature corresponding to the multi-view image features.
[0183] In some implementations, when obtaining the first bird's-eye view image features corresponding to the multi-view image features in the encoder, the encoder may include a Temporal Self Attention layer and a Spatial CrossAttention layer.
[0184] Step 122 : obtaining target bird's-eye view image features based on the first bird's-eye view image features through a decoder.
[0185] Here, the first bird's-eye view image feature is input to the decoder, and the decoder processes the first bird's-eye view image feature and outputs the target bird's-eye view image feature.
[0186] In some embodiments, the decoder may be used to process the first bird's-eye view image feature, convert the first bird's-eye view image feature into a vector form, and use the first bird's-eye view image feature in the vector form as the target bird's-eye view image feature.
[0187] In some embodiments, the decoder may include a feedforward network layer, a residual connection, and a layer normalization layer. The feedforward network layer, the residual connection, and the layer normalization layer may perform a linear transformation and an addition operation on the first bird's-eye view image feature, so that the first bird's-eye view image feature may be converted into a vector form to obtain the target bird's-eye view image feature.
[0188] In some embodiments, the number of residual connections and layer normalization layers may be at least one.
[0189] In an embodiment of the present application, a first bird's-eye view image feature corresponding to a multi-view image feature is obtained through an encoder, and a target bird's-eye view image feature is obtained based on the first bird's-eye view image feature through a decoder, thereby improving the accuracy of the target bird's-eye view image feature.
[0190] In some implementations, the “determining the image information of the target type based on the multi-view image information” in step 101 includes steps 1011 to 1013, that is, when determining the image information of the target type based on the multi-view image information, steps 1011 to 1013 may be performed, wherein:
[0191] Step 1011, when the multi-view image information is collected by a forward-view shooting device, the forward-view type image information is used as the target type image information.
[0192] Here, the front view shooting device refers to a shooting device set at the front view position. The front view shooting device may include but is not limited to: at least one of a camera, a camera, etc.
[0193] In some embodiments, the forward-looking perspective camera device has a corresponding field of view angle, and the corresponding field of view angle of the forward-looking perspective camera device can be any suitable size, for example, 120°, 100°, etc.
[0194] Step 1012: When the multi-view image information is collected by a forward-looking narrow-viewing angle shooting device, the forward-looking narrow-viewing angle image information is used as the target type image information.
[0195] Here, the forward narrow viewing angle shooting device refers to a shooting device arranged at a forward narrow viewing angle position. The forward narrow viewing angle shooting device may include but is not limited to: at least one of a camera, a still camera, and the like.
[0196] In some embodiments, the forward-looking narrow-viewing angle shooting device and the forward-looking angle shooting device may be the same device or different devices.
[0197] In some embodiments, the forward-looking narrow-viewing angle shooting device has a corresponding field of view angle, and the corresponding field of view angle of the forward-looking narrow-viewing angle shooting device can be any suitable size, for example, 30°, 45°, etc.
[0198] In some embodiments, the field of view angle corresponding to the forward-looking narrow-viewing angle shooting device is smaller than the field of view angle corresponding to the forward-looking angle shooting device.
[0199] Step 1013, when the multi-perspective image information is collected by a panoramic perspective shooting device, the panoramic perspective type image information is used as the target type image information; wherein the panoramic perspective includes at least one of the following: left front perspective, right front perspective, left rear perspective, right rear perspective, and rear perspective.
[0200] Here, the panoramic viewing angle shooting device refers to a shooting device set at the panoramic viewing angle position. The panoramic viewing angle shooting device may include but is not limited to: at least one of a camera, a camera, etc. The panoramic viewing angle includes at least one of the following: left front viewing angle, right front viewing angle, left rear viewing angle, right rear viewing angle, and rear viewing angle.
[0201] In some embodiments, the forward narrow-viewing angle shooting device, the forward viewing angle shooting device, and the surrounding viewing angle shooting device may be the same device or different devices.
[0202] In some embodiments, the panoramic viewing angle shooting device has a corresponding field of view angle, and the corresponding field of view angle of the panoramic viewing angle shooting device can be any suitable size, for example, 120°, 90°, etc.
[0203] In some embodiments, the field of view angle corresponding to the panoramic viewing angle shooting device may be the same as the field of view angle corresponding to the forward viewing angle shooting device.
[0204] In the implementation manner of the present application, image information of the target type is determined through image information collected by shooting devices with different viewing angles, thereby improving the accuracy of the image information of the target type.
[0205] Based on the above embodiments, the present application provides a method for training an environment perception model. Figure 7 A schematic diagram of the implementation flow of a training method for an environment perception model provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the method includes steps 201 to 204, wherein:
[0206] Step 201, obtaining a training data set; wherein the training data set includes first image information from multiple perspectives and an environmental perception truth value of at least one perception task corresponding to the first image information.
[0207] Here, a training data set is obtained so that the training data set can be used for subsequent model training processing.
[0208] In some embodiments, the training data set may include first image information from multiple perspectives, and may also include an environmental perception truth value of at least one perception task corresponding to the first image information.
[0209] It can be understood that in the embodiments of the present application, the first image information and the environmental perception true value included in the training data set correspond to each other based on the acquisition time of the first image information. That is, for the same acquisition time, it includes first image information with multiple perspectives and the corresponding environmental perception true value for at least one perception task.
[0210] It is understandable that, in the embodiment of the present application, the multi-view first image information may be an image acquired through a plurality of different acquisition viewing angles, wherein the plurality of different acquisition viewing angles may include but are not limited to: a front viewing angle, a front narrow viewing angle, a left front viewing angle, a right front viewing angle, a left rear viewing angle, a right rear viewing angle, a rear viewing angle, etc.
[0211] That is to say, in the embodiments of the present application, there is no specific limitation on the number of acquisition viewing angles.
[0212] In some implementations, the multi-view first image information may be directly acquired by the first device through an image acquisition device configured therein, or may be acquired by other image acquisition devices and then sent to the first device.
[0213] In some embodiments, the perception task can be understood as a perception task of vehicle driving, and the perception task can be used to determine and monitor the vehicle driving environment. Among them, the environmental perception truth value of at least one perception task can represent the actual situation of the real vehicle driving environment.
[0214] Exemplarily, in some embodiments, the perception task may include, but is not limited to, at least one two-dimensional vehicle driving perception task and / or at least one three-dimensional vehicle driving perception task, which is not specifically limited in this application.
[0215] It is understandable that in the embodiments of the present application, the two-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: lane line detection task, license plate detection task, face detection task, etc. The two-dimensional vehicle driving environment perception result is the result obtained by performing the two-dimensional vehicle driving environment perception task. The two-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: pixel point category, license plate presence, face presence, etc.
[0216] It is understandable that in the embodiments of the present application, the three-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: vehicle detection task, pedestrian detection task, cyclist detection task, lane line detection task, drivable area detection task, etc. The three-dimensional vehicle driving environment perception result is the result obtained by executing the three-dimensional vehicle driving environment perception task. The three-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: three-dimensional detection frame information, drivable area information, lane line category, etc. The three-dimensional detection frame information may include but is not limited to: at least one of: the center point of the three-dimensional detection frame, the length, width and height of the three-dimensional detection frame, the orientation angle of the three-dimensional detection frame, etc. The lane line category may include but is not limited to: one of: dotted line, solid line, fishbone line, etc.
[0217] In some implementations, based on the collection time of the multi-view first image information, there is an environment perception true value corresponding to at least one perception task. That is, the first image information and the environment perception true value correspond to each other based on the collection time of the first image information.
[0218] Exemplarily, in some embodiments, assuming that the collection time of the first image information of multiple perspectives is t, and at least one perception task includes a vehicle detection task, a pedestrian detection task, a cyclist detection task, and a lane line detection task, then, at time t, the environmental perception true value of at least one perception task corresponding to the first image information can be that there is a vehicle at time t, there are no pedestrians at time t, there are no cyclists at time t, and the lane line is a solid line at time t.
[0219] That is to say, in an embodiment of the present application, the training data included in the training data set is based on image information and environmental perception true values corresponding to the image acquisition moment, wherein the image information may be a multi-perspective multi-frame image corresponding to the image acquisition moment, and the environmental perception true value may be at least one environmental perception result of at least one perception task corresponding to the image acquisition moment.
[0220] Step 202, obtaining a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through an initial perception model; wherein the initial perception model includes an initial multi-type feature extraction network and an initial bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder.
[0221] Here, after acquiring the initial perception model, the initial perception model can be further used to obtain a predicted perception result of at least one perception task corresponding to the first image information based on the first image information.
[0222] In some embodiments, when obtaining a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through an initial perception model, multiple types of image information can be first determined based on the first image information of multiple perspectives; then, multi-perspective image features and multi-perspective bird's-eye view image features can be obtained based on the multiple types of image information through the initial perception model; finally, at least one perception task can be performed based on the multi-perspective image features and the multi-perspective bird's-eye view image features to obtain a predicted perception result of the at least one perception task.
[0223] In some embodiments, the multiple types may include at least a forward-looking viewing angle type, a forward-looking narrow-viewing viewing angle type, and a peripheral-viewing viewing angle type.
[0224] In some implementations, since different acquisition perspectives have different importance, different acquisition perspectives can be divided according to their importance, so as to obtain different types, that is, to divide into multiple types. Among them, the multiple types can at least include a forward-looking perspective type, a forward-looking narrow-viewing perspective type, a circumferential-viewing perspective type, etc.
[0225] In some embodiments, the initial perception model can be used to perform at least one perception task based on the collected image information to perceive and predict the vehicle driving environment, and finally obtain a predicted perception result of the vehicle driving environment corresponding to the image information.
[0226] In some embodiments, the initial perception model may include a multi-type feature extraction network and a bird's-eye view image feature extraction network. The multi-type feature extraction network is used to extract image features from different types of input image information. The bird's-eye view image feature extraction network can be used to perform spatial feature conversion and BEV feature fusion on the input image features in combination with acquisition parameters (e.g., camera parameters) corresponding to the first image information acquired from multiple perspectives, thereby obtaining corresponding bird's-eye view image features.
[0227] In some embodiments, the predicted perception result of at least one perception task can be used to predict the situation of the vehicle driving environment. The predicted perception result of at least one perception task determined based on the first image information corresponds to the environmental perception true value of at least one perception task corresponding to the first image information in the training data set, that is, the first image information, the environmental perception true value, and the predicted perception result correspond to each other based on the acquisition time of the first image information.
[0228] Exemplarily, in some embodiments, assuming that the collection time of the first image information of multiple perspectives is t, and at least one perception task includes a vehicle detection task, a pedestrian detection task, a cyclist detection task, and a lane line detection task, then, at time t, the predicted perception results of at least one perception task corresponding to the first image information may be that there is no vehicle at time t, there is no pedestrian at time t, there is a cyclist at time t, and the lane line at time t is a dotted line.
[0229] That is to say, in an embodiment of the present application, the predicted perception result of at least one perception task obtained by the initial perception model may be different from the environmental perception true value of the corresponding at least one perception task. Therefore, the initial perception model can be further trained and optimized in combination with the predicted perception result and the environmental perception true value to improve the performance of the model.
[0230] Step 203, determining the loss based on the true value of environmental perception and the predicted perception result.
[0231] Here, after obtaining a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through the initial perception model, the loss can be further determined based on the environmental perception true value and the predicted perception result.
[0232] In some embodiments, the loss is a function that measures the gap between the model prediction result and the actual result. The loss may include but is not limited to the following:
[0233] Mean Squared Error (MSE): Also known as L2 loss, it measures the average of the squares of the differences between the predicted value and the true value. It is sensitive to outliers because outliers can significantly increase the loss value.
[0234] Mean Absolute Error (MAE): Also known as L1 loss, it measures the average of the absolute values of the difference between the predicted value and the true value. It is less sensitive to outliers because outliers do not cause the loss value to increase dramatically.
[0235] Huber loss: It combines the advantages of MSE and MAE. MSE is used when the error is small, and MAE is used when the error is large to improve the robustness to outliers.
[0236] Step 204: correct the model parameters of the initial perception model based on the loss to obtain an environment perception model.
[0237] Here, after the loss is determined based on the true value of environmental perception and the predicted perception result, the model parameters of the initial perception model can be further corrected based on the loss to finally obtain the environmental perception model.
[0238] It can be understood that in the embodiments of the present application, after obtaining the quantified loss of the gap between the model prediction results and the actual results, based on the loss, the model parameters can be adjusted through the optimization algorithm to reduce this gap, that is, the purpose of the correction is to adjust the model parameters through the optimization algorithm to minimize the loss function.
[0239] Exemplarily, in some embodiments, the optimization algorithms used when modifying the model parameters may include but are not limited to the following:
[0240] Gradient descent method: By calculating the gradient of the loss function with respect to the model parameters and updating the parameters in the opposite direction of the gradient, the loss value is gradually reduced.
[0241] Stochastic gradient descent: Only one sample or a small batch of samples is used to calculate the gradient and update the parameters in each iteration. This method is suitable for large-scale data sets, but may lead to slow convergence and unstable parameter updates.
[0242] Momentum method: A momentum term is introduced based on the gradient descent method to accelerate convergence and reduce oscillation.
[0243] Adaptive learning rate methods: such as Adam, RMSprop, etc. These methods dynamically adjust the learning rate according to historical gradient information to improve optimization efficiency and effect.
[0244] Furthermore, in an embodiment of the present application, after the model parameters of the initial perception model are corrected based on the loss to obtain the environmental perception model, at least one perception task can be performed through the environmental perception model based on the collected multi-perspective image information to obtain corresponding perception results.
[0245] In an embodiment of the present application, after obtaining the predicted perception results corresponding to the first image information of multiple perspectives through the initial perception model, the total loss of the model training can be determined based on the environmental perception true value and the predicted perception results; in this way, the initial perception model including the multi-type feature extraction network and the initial bird's-eye view feature extraction network is trained through the loss, so as to improve the prediction effect and performance of the neural network.
[0246] The following describes the application of the vehicle driving environment perception method provided in the embodiment of the present application in actual scenarios.
[0247] With the vigorous development and increasing popularity of vehicle intelligence and electrification, more and more vehicles will be equipped with autonomous driving functions. Therefore, it is necessary to first perceive and identify the driving environment around the vehicle, and then input it to the downstream modules of autonomous driving for planning, control and making correct decisions.
[0248] In related technologies, a perception system can be established through a deep neural network to perceive and identify the environment and obstacles around the vehicle based on the surrounding image data collected by the vehicle. The perception system is low-cost and can be deployed on the vehicle to perceive the real-time driving environment. However, the perception system usually contains multiple independent modules, and each module needs to run separately to obtain results, and then aggregate and fuse them before they can be output to downstream use. In addition, the perception system usually cannot perform task perceptions of different dimensions at the same time, and has a problem of poor flexibility.
[0249] The embodiment of the present application provides a method for perceiving a vehicle driving environment. First, the image information of the target type is determined by acquiring multi-view image information. Then, the multi-view image features and the target bird's-eye view image features are acquired based on the image information of the target type through an environment perception model including a multi-type feature extraction network and a bird's-eye view feature extraction network. Finally, a two-dimensional vehicle driving environment perception task can be performed based on the multi-view image features to obtain a two-dimensional vehicle driving environment perception result, and a three-dimensional vehicle driving environment perception task can be performed based on the target bird's-eye view image features to obtain a three-dimensional vehicle driving environment perception result. In this way, on the one hand, since the perception method of the present application can respectively perform the two-dimensional vehicle driving environment perception task and the three-dimensional vehicle driving environment perception task through the output multi-view image features and the target bird's-eye view image features, task perceptions of different dimensions can be performed simultaneously, thereby improving the flexibility of task perception. On the other hand, compared with the perception system of the prior art including multiple independent modules, since the perception method of the present application is implemented through an environment perception model, the perception task can be completed through one environment perception model, thereby improving the perception efficiency of the vehicle driving environment and reducing the complexity of the perception system.
[0250] Figure 8 A schematic diagram of the implementation process of a vehicle driving environment perception method provided in an embodiment of the present application Figure 2 ,like Figure 8 As shown, the method includes steps 301 to 307, wherein:
[0251] Step 301, acquiring multi-view image information, and determining image information of a target type based on the multi-view image information;
[0252] Step 302, obtaining a first image feature corresponding to the image information of the target type based on the image information of the target type through a backbone module corresponding to the target type;
[0253] Step 303, obtaining multi-view image features based on the first image features through a first fusion module corresponding to the target type;
[0254] Step 304, performing a two-dimensional vehicle driving environment perception task based on the multi-view image features to obtain a two-dimensional vehicle driving environment perception result;
[0255] Step 305, obtaining, through an encoder, first bird's-eye view image features corresponding to the multi-view image features based on the multi-view image features, camera parameters corresponding to the multi-view image information, bird's-eye view image features at historical moments, and bird's-eye view parameters;
[0256] Step 306, obtaining target bird's-eye view image features based on the first bird's-eye view image features through a decoder;
[0257] Step 307 , performing a three-dimensional vehicle driving environment perception task based on the target bird's-eye view image features to obtain a three-dimensional vehicle driving environment perception result.
[0258] Fig. 9 A schematic diagram of obtaining a target bird's-eye view image feature provided by an embodiment of the present application, such as Fig. 9 As shown, the bird's-eye view image features 91 and bird's-eye view parameters 92 at the historical moment are input into the spatiotemporal self-attention layer 93, and the bird's-eye view image features 91 and bird's-eye view parameters 92 at the historical moment, the camera parameters 95 corresponding to the multi-view image information, and the multi-view image features 96 processed by the spatiotemporal self-attention layer 93 and the first residual connection and layer normalization layer 94 in sequence are input into the spatial cross-attention layer 97, and processed by the spatial cross-attention layer 97, the second residual connection and layer normalization layer 98, the feedforward network layer 99, and the second residual connection and layer normalization layer 910 in sequence to obtain the target bird's-eye view image features 911, and then the target bird's-eye view image features 911 are input into the task head 912 to perform the vehicle driving environment perception task.
[0259] It should be noted that Fig. 9 The architecture shown is neither an encoder nor a decoder, but a network consisting of important functional layers selected from the bird's-eye view feature extraction network, through Fig. 9 The architecture shown can obtain the target bird's-eye view image features 911 based on the multi-view image features 96, the camera parameters 95 corresponding to the multi-view image information, the bird's-eye view image features 91 at the historical moment, and the bird's-eye view parameters 92.
[0260] Fig.10A schematic diagram of a method for sensing a vehicle driving environment provided in an embodiment of the present application is shown in FIG. Fig.10 As shown, first, after collecting the image information of the target type through multiple acquisition perspectives, the image information 101 of the forward-looking perspective type is input into the trunk module 1021 of the first sub-network in the trunk module 102, and sequentially passes through the trunk module 1021 of the first sub-network and the fusion module 1031 of the first sub-network in the first fusion module 103 to obtain the first image sub-feature corresponding to the forward-looking perspective type image information 101. The image information 104 of the forward-looking narrow-viewing perspective type is input into the trunk module 1022 of the second sub-network in the trunk module 102, and sequentially passes through the trunk module 1022 of the second sub-network and the fusion module 1032 of the second sub-network in the second fusion module 103 to obtain the second image sub-feature corresponding to the image information 104 of the forward-looking narrow-viewing perspective type. The image information 105 of the panoramic view type is input into the backbone module 1023 of the third sub-network in the backbone module 102, and the third image sub-feature corresponding to the image information 105 of the panoramic view type is obtained through the backbone module 1023 of the third sub-network and the fusion module 1033 of the third sub-network in the second fusion module 103. Then, the two-dimensional vehicle driving environment perception task 110 is performed through the multi-view image features including the first image sub-feature, the second image sub-feature and the third image sub-feature, wherein the two-dimensional vehicle driving environment perception task 110 includes: a lane line detection task 1101, a license plate detection task 1102 and a face detection task 1103. At the same time, the multi-view image features, the bird's-eye view parameters 105, the camera parameters and the bird's-eye view image features at the historical moment are input into the encoder 106, and are processed by the encoder 106 and the decoder 107 in turn to obtain the target bird's-eye view image features 108 in the form of vectors, wherein the number of the target bird's-eye view image features 108 is at least one. Finally, the three-dimensional vehicle driving environment perception task 109 is performed through the target bird's-eye view image features 108, wherein the three-dimensional vehicle driving environment perception task 109 includes: vehicle detection task 1091, pedestrian detection task 1092, cyclist detection task 1093, lane line detection task 1094, and drivable area detection task 1095.
[0261] Based on the above embodiments, the present application also provides a device for sensing a vehicle driving environment. Fig.11 A schematic diagram of the structure of a vehicle driving environment sensing device provided in an embodiment of the present application is shown in FIG. Fig.11 As shown, the vehicle driving environment perception device 1100 includes a first determination unit 1101, a first acquisition unit 1102 and a perception unit 1103, wherein:
[0262] A first acquisition unit 1102 is used to acquire multi-view image information;
[0263] The first determining unit 1101 is used to determine the image information of the target type based on the multi-view image information; wherein the target type includes at least one or more types of the forward viewing angle type, the forward narrow viewing angle type, and the surrounding viewing angle type;
[0264] The first acquisition unit 1101 is further used to acquire multi-view image features and target bird's-eye view image features corresponding to the image information of the target type based on the image information of the target type through the environment perception model; wherein the environment perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder;
[0265] The perception unit 1103 is used to perform a two-dimensional vehicle driving environment perception task based on multi-view image features to obtain a two-dimensional vehicle driving environment perception result, and to perform a three-dimensional vehicle driving environment perception task based on the target bird's-eye view image features to obtain a three-dimensional vehicle driving environment perception result.
[0266] In some embodiments, the first acquisition unit 1102 is further used to acquire multi-perspective image features corresponding to the image information of the target type based on the image information of the target type through a multi-type feature extraction network; and to acquire target bird's-eye view image features corresponding to the image information of the target type based on the multi-perspective image features, camera parameters corresponding to the multi-perspective image information, bird's-eye view image features at historical moments, and bird's-eye view parameters through a bird's-eye view feature extraction network.
[0267] In some embodiments, the first acquisition unit 1102 is also used to acquire a first image feature corresponding to the image information of the target type based on the image information of the target type through a trunk module corresponding to the target type; and acquire a multi-view image feature based on the first image feature through a first fusion module corresponding to the target type.
[0268] In some embodiments, the multi-type feature extraction network includes one or more types of a first sub-network corresponding to the forward-looking perspective type, a second sub-network corresponding to the forward-looking narrow perspective type, and a third sub-network corresponding to the surrounding perspective type; the first acquisition unit 1102 is also used to, when the target type is the forward-looking perspective type, acquire, based on the image information of the forward-looking perspective type, a first image sub-feature corresponding to the image information of the forward-looking perspective type through the backbone module of the first sub-network; when the target type is the forward-looking narrow perspective type, acquire, based on the image information of the forward-looking narrow perspective type, a second image sub-feature corresponding to the image information of the forward-looking narrow perspective type through the backbone module of the second sub-network; when the target type is the surrounding perspective type, acquire, based on the image information of the surrounding perspective type, a third image sub-feature corresponding to the image information of the surrounding perspective type through the backbone module of the third sub-network.
[0269] In some embodiments, the first acquisition unit 1102 is also used to, when the target type is a forward looking perspective type, acquire, based on the image information of the forward looking perspective type, a first image sub-feature corresponding to the image information of the forward looking perspective type through a backbone module of the first sub-network; when the target type is a forward narrow looking perspective type, acquire, based on the image information of the forward narrow looking perspective type, a second image sub-feature corresponding to the image information of the forward narrow looking perspective type through a backbone module of the second sub-network; when the target type is a peripheral looking perspective type, acquire, based on the image information of the peripheral looking perspective type, a third image sub-feature corresponding to the image information of the peripheral looking perspective type through a backbone module of the third sub-network.
[0270] In some embodiments, the perception unit 1103 is also used to extract target perspective two-dimensional image features from multi-perspective image features based on a two-dimensional vehicle driving environment perception task; perform the two-dimensional vehicle driving environment perception task based on the target perspective image features to obtain a two-dimensional vehicle driving environment perception result.
[0271] In some embodiments, the first acquisition unit 1102 is further used to obtain, through an encoder, first bird's-eye view image features corresponding to the multi-view image features based on the multi-view image features, camera parameters corresponding to the multi-view image information, bird's-eye view image features at historical moments, and bird's-eye view parameters; and obtain, through a decoder, target bird's-eye view image features based on the first bird's-eye view image features.
[0272] In some embodiments, the first determination unit 1101 is further used to use the forward perspective type image information as the target type image information when the multi-perspective image information is collected by a forward perspective shooting device; to use the forward narrow perspective type image information as the target type image information when the multi-perspective image information is collected by a forward narrow perspective shooting device; and to use the peripheral perspective type image information as the target type image information when the multi-perspective image information is collected by a peripheral perspective shooting device; wherein the peripheral perspective includes at least one of the following: left front perspective, right front perspective, left rear perspective, right rear perspective, and rearward perspective.
[0273] Based on the above embodiments, the present application also provides a training device for an environment perception model. Fig.12 A schematic diagram of the structure of a training device for an environmental perception model provided in an embodiment of the present application is shown in FIG. Fig.12 As shown, the training device 1200 of the environment perception model includes a second determination unit 1201, a second acquisition unit 1202 and a training unit 1203, wherein:
[0274] The second acquisition unit 1202 is used to acquire a training data set; wherein the training data set includes multi-view first image information and an environmental perception truth value corresponding to at least one perception task of the first image information; through an initial perception model, based on the first image information, a predicted perception result of at least one perception task corresponding to the first image information is obtained; wherein the initial perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder;
[0275] The second determining unit 1201 is used to determine the loss based on the environmental perception true value and the predicted perception result;
[0276] The training unit 1203 is used to correct the model parameters of the initial perception model based on the loss to obtain the environment perception model.
[0277] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0278] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0279] The present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the above method is implemented when the processor executes the computer program.
[0280] The present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. The computer-readable storage medium can be transient or non-transient.
[0281] The present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements some or all of the steps in the above method. The computer program product can be implemented specifically by hardware, software, or a combination thereof. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium, and in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0282] It should be noted that Fig.13 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Fig.13 As shown, the electronic device 130 proposed in the embodiment of the present application may include a processor 1301 , a memory 1302 , a communication interface 1303 , and a bus 1304 for connecting the processor 1301 , the memory 1302 and the communication interface 1303 .
[0283] In the embodiment of the present application, the above-mentioned processor 1301 can be at least one of an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a digital signal processor (Digital Signal Processor, DSP), a digital signal processing device (Digital Signal Processing Device, DSPD), a programmable logic device (ProgRAMmable Logic Device, PLD), a field programmable gate array (Field ProgRAMmable Gate Array, FPGA), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other, and the embodiment of the present application is not specifically limited. The computer device 130 can also include a memory 1302, which can be connected to the processor 1301, wherein the memory 1302 is used to store executable program code, the program code includes computer operation instructions, and the memory 1302 may include a high-speed RAM memory, and may also include a non-volatile memory, for example, at least two disk memories.
[0284] In the embodiment of the present application, the bus 1304 is used to connect the communication interface 1303, the processor 1301 and the memory 1302, as well as the mutual communication between these devices.
[0285] In practical applications, the memory 1302 may be a volatile memory, such as a random access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above types of memories, and provide instructions and data to the processor 1301.
[0286] Further, in an embodiment of the present application, the processor 1301 is used to obtain multi-perspective image information, and determine the image information of the target type based on the multi-perspective image information; wherein the target type includes at least one or more types of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a surrounding-view perspective type; through an environmental perception model, based on the image information of the target type, multi-perspective image features and target bird's-eye view image features corresponding to the image information of the target type are obtained; wherein the environmental perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, and the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder; based on the multi-perspective image features, a two-dimensional vehicle driving environment perception task is performed to obtain a two-dimensional vehicle driving environment perception result, and based on the target bird's-eye view image features, a three-dimensional vehicle driving environment perception task is performed to obtain a three-dimensional vehicle driving environment perception result.
[0287] Further, in an embodiment of the present application, processor 1301 is used to obtain a training data set; wherein the training data set includes multi-perspective first image information and an environmental perception true value corresponding to at least one perception task of the first image information; through an initial perception model, based on the first image information, a predicted perception result of at least one perception task corresponding to the first image information is obtained; wherein the initial perception model includes an initial multi-type feature extraction network and an initial bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, and the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder; based on the environmental perception true value and the predicted perception result, a loss is determined; based on the loss, the model parameters of the initial perception model are corrected to obtain an environmental perception model.
[0288] In the embodiments of the present application, further, Fig.14 This is a schematic diagram of the structure of the vehicle equipment proposed in the embodiment of the present application, such as Fig.14 As shown, the vehicle equipment 140 proposed in the embodiment of the present application may include an electronic device 130 .
[0289] The vehicle device 140 implements the method proposed in the above embodiment through the electronic device 130 .
[0290] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0291] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the serial number of each step / process mentioned above does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The serial numbers of the embodiments of the present application mentioned above are for description only and do not represent the advantages and disadvantages of the embodiments.
[0292] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0293] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the above-mentioned units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0294] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0295] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0296] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, read-only memories, magnetic disks or optical disks.
[0297] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0298] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.
Claims
1. A method for sensing a vehicle driving environment, characterized in that: The method comprises: Acquire multi-view image information, and determine image information of a target type based on the multi-view image information; wherein the target type includes at least one or more types of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a peripheral-view perspective type; Through the environment perception model, based on the image information of the target type, multi-view image features and target bird's-eye view image features corresponding to the image information of the target type are obtained; wherein the environment perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder; Based on the multi-view image features, a two-dimensional vehicle driving environment perception task is performed to obtain a two-dimensional vehicle driving environment perception result, and based on the target bird's-eye view image features, a three-dimensional vehicle driving environment perception task is performed to obtain a three-dimensional vehicle driving environment perception result.
2. The sensing method according to claim 1, characterized in that: The step of obtaining, through the environment perception model and based on the image information of the target type, multi-view image features and bird's-eye view image features corresponding to the image information of the target type includes: Acquiring multi-view image features corresponding to the image information of the target type based on the image information of the target type through the multi-type feature extraction network; Through the bird's-eye view feature extraction network, based on the multi-view image features, the camera parameters corresponding to the multi-view image information, the bird's-eye view image features at historical moments and the bird's-eye view parameters, the target bird's-eye view image features corresponding to the image information of the target type are obtained.
3. The sensing method according to claim 2, characterized in that: The step of obtaining multi-view image features corresponding to the image information of the target type based on the image information of the target type through the multi-type feature extraction network includes: Acquire, by the trunk module corresponding to the target type, a first image feature corresponding to the image information of the target type based on the image information of the target type; The multi-view image features are acquired based on the first image features through the first fusion module corresponding to the target type.
4. The sensing method according to claim 2, characterized in that: The multi-type feature extraction network includes one or more types of a first sub-network corresponding to the forward-looking viewing angle type, a second sub-network corresponding to the forward-looking narrow-viewing viewing angle type, and a third sub-network corresponding to the surrounding viewing angle type; The acquiring, by the trunk module corresponding to the target type and based on the image information of the target type, a first image feature corresponding to the image information of the target type comprises at least one of the following: In the case where the target type is a forward-looking perspective type, obtaining, through the backbone module of the first sub-network, a first image sub-feature corresponding to the forward-looking perspective type image information based on the forward-looking perspective type image information; In the case where the target type is a forward-looking narrow-viewing angle type, obtaining, through the backbone module of the second sub-network, a second image sub-feature corresponding to the forward-looking narrow-viewing angle type image information based on the forward-looking narrow-viewing angle type image information; In the case where the target type is a panoramic viewing angle type, a third image sub-feature corresponding to the image information of the panoramic viewing angle type is obtained based on the image information of the panoramic viewing angle type through the backbone module of the third sub-network.
5. The sensing method according to claim 4, characterized in that: The acquiring the multi-view image feature based on the first image feature by using the first fusion module corresponding to the target type includes at least one of the following: In a case where the target type is a front-view perspective type, obtaining, through a fusion module of the first sub-network, a first multi-view image feature corresponding to the first image sub-feature based on the first image sub-feature; When the target type is a forward narrow viewing angle type, obtaining, through a fusion module of the second sub-network, a second multi-view image feature corresponding to the second image sub-feature based on the second image sub-feature; In the case where the target type is a panoramic viewing angle type, a third multi-view image feature corresponding to the third image sub-feature is obtained based on the third image sub-feature through a fusion module of the third sub-network.
6. The sensing method according to claim 1, characterized in that: The performing of the two-dimensional vehicle driving environment perception task based on the multi-view image features to obtain the two-dimensional vehicle driving environment perception result includes: Based on the two-dimensional vehicle driving environment perception task, extracting target perspective two-dimensional image features from the multi-perspective image features; The two-dimensional vehicle driving environment perception task is performed based on the target perspective image features to obtain the two-dimensional vehicle driving environment perception result.
7. The sensing method according to claim 2, characterized in that: The step of obtaining the bird's-eye view image features corresponding to the image information of the target type through the bird's-eye view feature extraction network based on the multi-view image features, the camera parameters corresponding to the multi-view image information, the bird's-eye view image features at historical moments, and the bird's-eye view parameters includes: Obtaining, by the encoder, a first bird's-eye view image feature corresponding to the multi-view image feature based on the multi-view image feature, the camera parameter corresponding to the multi-view image information, the bird's-eye view image feature at the historical moment, and the bird's-eye view parameter; The target bird's-eye view image feature is acquired through the decoder based on the first bird's-eye view image feature.
8. The sensing method according to any one of claims 1 to 7, characterized in that: The determining of the image information of the target type based on the multi-view image information includes: In the case where the multi-view image information is collected by a forward-view shooting device, the image information of the forward-view type is used as the image information of the target type; In the case where the multi-view image information is collected by a forward-looking narrow-viewing angle shooting device, the image information of the forward-looking narrow-viewing angle type is used as the image information of the target type; When the multi-perspective image information is collected by a panoramic perspective shooting device, the image information of the panoramic perspective type is used as the image information of the target type; wherein the panoramic perspective includes at least one of the following: left front perspective, right front perspective, left rear perspective, right rear perspective, and rear perspective.
9. A method for training an environment perception model, characterized in that: The method comprises: Acquire a training data set; wherein the training data set includes multi-perspective first image information and an environmental perception truth value of at least one perception task corresponding to the first image information; Obtaining, by means of an initial perception model, a predicted perception result of at least one perception task corresponding to the first image information based on the first image information; wherein the initial perception model comprises a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network comprises different sub-networks corresponding to different target types, the sub-networks comprising a backbone module and a first fusion module; the bird's-eye view feature extraction network comprises an encoder and a decoder; Determining a loss based on the environmental perception true value and the predicted perception result; The model parameters of the initial perception model are modified based on the loss to obtain the environment perception model.
10. A vehicle driving environment sensing device, characterized in that: The vehicle driving environment sensing device comprises: A first acquisition unit, used to acquire multi-view image information; A first determining unit is used to determine image information of a target type based on the multi-view image information; wherein the target type includes at least one or more types of a forward viewing angle type, a forward narrow viewing angle type, and a peripheral viewing angle type; The first acquisition unit is further used to acquire, through an environment perception model and based on the image information of the target type, multi-view image features and target bird's-eye view image features corresponding to the image information of the target type; wherein the environment perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder; The perception unit is used to perform a two-dimensional vehicle driving environment perception task based on the multi-view image features to obtain a two-dimensional vehicle driving environment perception result, and to perform a three-dimensional vehicle driving environment perception task based on the target bird's-eye view image features to obtain a three-dimensional vehicle driving environment perception result.
11. A training device for an environmental perception model, characterized in that: The training device of the environmental perception model comprises: A second acquisition unit is used to acquire a training data set; wherein the training data set includes multi-view first image information and an environmental perception truth value corresponding to at least one perception task of the first image information; through an initial perception model, based on the first image information, a predicted perception result corresponding to at least one perception task of the first image information is obtained; wherein the initial perception model includes a multi-type feature extraction network and a bird's-eye view feature extraction network, the multi-type feature extraction network includes different sub-networks corresponding to different target types, the sub-networks include a backbone module and a first fusion module; the bird's-eye view feature extraction network includes an encoder and a decoder; A second determining unit, configured to determine a loss based on the environmental perception true value and the predicted perception result; A training unit is used to modify the model parameters of the initial perception model based on the loss to obtain the environment perception model.
12. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, the steps in the method according to any one of claims 1 to 8 or 9 are implemented.
13. A vehicle device, characterized in that: The vehicle equipment includes an electronic device, and the vehicle equipment implements the steps of the method according to any one of claims 1 to 8 or 9 through the electronic device.
14. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps in the method described in any one of claims 1-8 or 9 are implemented.
15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method of any one of claims 1 to 8 or 9 are implemented.