Model training method, perception method, device, equipment, medium and program product
Through the multi-level weighted loss correction initial perception model, the problem that neural networks in the prior art are difficult to achieve optimality for each perception task under ideal circumstances is solved, and the prediction effect and performance of vehicle driving environment perception is improved.
Patent Information
- Application Number
- CN202510094948.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
Existing neural network-based vehicle driving environment perception methods cannot achieve optimality for each perception task under ideal circumstances, resulting in a degradation of prediction effects and performance.
By obtaining image information from multiple perspectives and corresponding environment perception truth values, the initial perception model is used to predict the perception result, and multi-level weighted losses are determined based on the subtask weight of the perceptual task, and the model parameters are corrected to obtain the environment perception model.
Improve the prediction effect and performance of neural networks to ensure accurate perceptual results are obtained when sensing the vehicle driving environment based on neural networks.
Smart Images

Figure CN120014411A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle autonomous driving, and specifically to a method and device for training an environmental perception model, a method and device for perceiving a vehicle's driving environment, equipment, storage media, and program products. Background Art
[0002] With the vigorous development and increasing popularity of automobile intelligence and electrification, more and more cars will be equipped with autonomous driving functions. In the process of autonomous driving, it is crucial to accurately perceive and identify the environment and obstacles around the vehicle.
[0003] Currently, vehicles can be used to collect surrounding image data and combine it with deep neural networks to perceive and identify the environment and obstacles around the vehicle.
[0004] However, the common perception task recognition method based on neural network cannot guarantee that each perception task can achieve the optimal state under ideal conditions, which reduces the prediction effect and performance of the neural network, and thus cannot obtain accurate perception results when the vehicle driving environment is perceived based on the neural network. Summary of the invention
[0005] The present application provides a training method and device for an environmental perception model, a perception method and device for a vehicle driving environment, equipment, storage medium, and program product, which can realize multi-level weighted loss of perception tasks, improve the prediction effect and performance of neural networks, and thus obtain accurate perception results when perceiving the vehicle driving environment based on neural networks.
[0006] In order to achieve the above objectives, the present application provides a method for training an environment perception model, the method comprising:
[0007] Acquire a training data set; wherein the training data set includes first image information from multiple perspectives and an environmental perception truth value of at least one perception task corresponding to the first image information;
[0008] Obtaining, by means of the initial perception model, a predicted perception result of at least one perception task corresponding to the first image information based on the first image information;
[0009] Based on the true value of environmental perception and the predicted perception result, the loss is determined according to the weights of the subtasks of the perception task;
[0010] The model parameters of the initial perception model are modified based on the loss to obtain the environment perception model.
[0011] According to the above-mentioned technical means, after obtaining the predicted perception results corresponding to the first image information of multiple perspectives through the initial perception model, the total loss of the model training can be determined based on the weights of the subtasks of the perception task based on the true value of environmental perception and the predicted perception results; in this way, the loss can be weighted by selecting weights suitable for the subtasks according to the distribution differences of the subtasks of each perception task and the differences in optimization difficulty, that is, by using different weights to balance between subtasks, multi-level training losses can be achieved, thereby improving the prediction effect and performance of the neural network.
[0012] Furthermore, based on the true value of environmental perception and the predicted perception result, the loss is determined according to the weights of the subtasks of the perception task, including:
[0013] Based on the environmental perception true value and the predicted perception result, and according to the weight of each subtask of the first perception task, determining a first loss parameter; wherein the first perception task is any one of the at least one perception task;
[0014] Iterate over at least one perception task, and determine a loss based on a first loss parameter of each perception task.
[0015] According to the above technical means, the loss of each subtask can be determined separately, and then the loss of each perception task can be determined by combining the weights of the subtasks of each perception task; in this way, by using different weights between subtasks to achieve balance, multi-level training loss can be achieved, thereby improving the prediction effect and performance of the neural network.
[0016] Further, based on the environmental perception true value and the predicted perception result, according to the weight of each subtask of the first perception task, a first loss parameter is determined, including:
[0017] Determine a first loss of the first subtask based on the environmental perception true value, the predicted perception result and the loss function of the first subtask; wherein the first subtask is any subtask of the first perception task;
[0018] Determine a second loss for the first subtask based on the weight of the first subtask and the first loss;
[0019] Traverse each subtask of the first perception task, and determine the first loss parameter corresponding to the first perception task based on the second loss of each subtask.
[0020] According to the above technical means, the loss of each subtask can be determined based on the true value of environmental perception and the predicted perception result, and then the loss of each perception task can be determined in combination with the weight of the subtask of each perception task; in this way, by using different weights between subtasks to achieve balance, multi-level training loss can be achieved, thereby improving the prediction effect and performance of the neural network.
[0021] Furthermore, based on the true value of environmental perception and the predicted perception result, the loss is determined according to the weights of the subtasks of the perception task, including:
[0022] Based on the environmental perception true value and the predicted perception result, a first loss parameter is determined according to the weight of each subtask of the first perception task and a first loss adjustment parameter corresponding to the weight of each subtask; wherein the first perception task is any one of the at least one perception task;
[0023] Iterate over at least one perception task, and determine a loss based on a first loss parameter of each perception task.
[0024] According to the above technical means, when determining the loss of each perception task in combination with the weight of each subtask of the perception task, a first loss adjustment parameter corresponding to the weight of each subtask can be introduced to adaptively adjust the weight of each subtask. In this way, multi-level adaptive training loss can be achieved, thereby improving the prediction effect and performance of the neural network.
[0025] Further, based on the environmental perception true value and the predicted perception result, according to the weight of each subtask of the first perception task and the first loss adjustment parameter corresponding to the weight of each subtask, determining the first loss parameter includes:
[0026] Determine a first loss of the first subtask based on the environmental perception true value, the predicted perception result and the loss function of the first subtask; wherein the first subtask is any subtask of the first perception task;
[0027] Determine a second loss of the first subtask based on a first loss adjustment parameter corresponding to the first subtask, a weight of the first subtask, and the first loss;
[0028] Traverse each subtask of the first perception task, and determine the first loss parameter corresponding to the first perception task based on the second loss of each subtask.
[0029] According to the above-mentioned technical means, after determining the loss of each subtask based on the true value of environmental perception and the predicted perception result, in the process of determining the loss of each perception task in combination with the weight of the subtask of each perception task, a first loss adjustment parameter corresponding to the weight of each subtask can be introduced to adaptively adjust the weight of each subtask, that is, adaptively adjust the corresponding second loss adjustment parameter for different weights to avoid bias towards some perception tasks during the training process, thereby ensuring that each perception task can achieve the optimal state under ideal conditions.
[0030] Furthermore, based on the true value of environmental perception and the predicted perception result, the loss is determined according to the weights of the subtasks of the perception task, including:
[0031] Determine the target subtask among all subtasks of the perception task based on the indicator parameters of the subtask; wherein the indicator parameters of the subtask are used to characterize the importance of the subtask;
[0032] Based on the true value of environmental perception and the predicted perception result, a first loss parameter is determined according to the weight of the target subtask;
[0033] Based on a first loss parameter of a second perception task, a loss is determined; wherein the second perception task is one or more perception tasks in at least one perception task.
[0034] According to the above technical means, the target subtask can be selected from all subtasks based on the indicator parameters that characterize the importance of the subtask, and then the overall loss can be calculated according to the weight of the target subtask. This can reduce the amount of calculation in the training process and improve the training efficiency while ensuring that each perception task can achieve the optimal state under ideal conditions.
[0035] Further, traversing at least one perception task, and determining a loss based on a first loss parameter of each perception task, including:
[0036] Traversing at least one perception task, and determining a second loss parameter based on a weight of each perception task and the first loss parameter;
[0037] Based on the second loss parameter, a loss is determined.
[0038] According to the above technical means, after determining the loss of each perception task by combining the weights of the subtasks of each perception task, the total loss can be further determined by combining the weights of each perception task; in this way, a balance can be achieved by using different weights between perception tasks, combining multi-task training loss and multi-level training loss, thereby improving the prediction effect and performance of the neural network.
[0039] Further, traversing at least one perception task, and determining a loss based on a first loss parameter of each perception task, including:
[0040] Traversing at least one perceptual task, determining a second loss parameter based on a weight of each perceptual task, a second loss adjustment parameter corresponding to the weight of each perceptual task, and the first loss parameter;
[0041] Based on the second loss parameter, a loss is determined.
[0042] According to the above technical means, when the total loss is determined in combination with the weight of each perception task, a second loss adjustment parameter corresponding to the weight of each perception task can be introduced to adaptively adjust the weight of each perception task. In this way, multi-task and multi-level adaptive training losses can be achieved, thereby improving the prediction effect and performance of the neural network.
[0043] Further, based on the second loss parameter, determining the loss includes:
[0044] Determining a first operand value based on a first loss adjustment parameter corresponding to each subtask of each perception task;
[0045] Determining a second operand value based on a second loss adjustment parameter corresponding to each perception task;
[0046] Based on the second loss parameter, the first operand value, and the second operand value, a loss is determined.
[0047] According to the above-mentioned technical means, during the model training process, a first loss adjustment parameter corresponding to the weight of each subtask can be introduced to adaptively adjust the weight of each subtask, and a second loss adjustment parameter corresponding to the weight of each perception task can be introduced to adaptively adjust the weight of each perception task. That is, the corresponding loss adjustment parameters are used for different weights for adaptive adjustment, thereby ensuring that each perception task can achieve the optimal state under ideal conditions.
[0048] Further, obtaining, by means of the initial perception model, a predicted perception result of at least one perception task corresponding to the first image information based on the first image information, includes:
[0049] Determine multiple types of image information based on the first image information of multiple perspectives; wherein the multiple types include at least a forward-looking perspective type, a forward-looking narrow-view perspective type, and a peripheral-view perspective type;
[0050] Through the initial perception model, based on multiple types of image information, multi-perspective image features and multi-perspective bird's-eye view image features are obtained;
[0051] At least one perception task is performed based on multi-perspective image features and / or multi-perspective bird's-eye view image features to obtain a predicted perception result of the at least one perception task.
[0052] According to the above technical means, during the model training process, the multi-perspective image features and the multi-perspective bird's-eye view image features corresponding to the multi-perspective first image information can be output based on the initial perception model, and then the predicted perception result of at least one perception task is obtained based on the multi-perspective image features and / or the multi-perspective bird's-eye view image features. In this way, the extraction of multi-perspective image features and multi-perspective bird's-eye view image features can be completed through the neural network, so as to perform subsequent loss determination, thereby improving the prediction effect and performance of the neural network.
[0053] The present application provides a method for sensing a vehicle driving environment, the method comprising:
[0054] Acquire multi-view image information, and determine the target type image information based on the multi-view image information; wherein the target type includes at least one or more types of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a circumferential-view perspective type;
[0055] Through the environmental perception model, based on the image information of the target type, the perception result of at least one perception task corresponding to the multi-perspective image information is obtained; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task.
[0056] According to the above technical means, the vehicle driving environment is perceived through the environmental perception model, and the perception result of at least one perception task corresponding to the multi-view image information is obtained. Among them, the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task. Therefore, the environmental perception model can ensure that each perception task can achieve the optimal under ideal conditions, with good prediction effect and performance, and thus can obtain accurate perception results when the vehicle driving environment is perceived based on the neural network.
[0057] Further, obtaining a perception result of at least one perception task corresponding to the multi-view image information based on the image information of the target type through the environment perception model includes:
[0058] Through the environment perception model, based on the image information of the target type, the multi-view image features and the multi-view bird's-eye view image features are obtained;
[0059] At least one perception task is performed based on multi-perspective image features and / or multi-perspective bird's-eye view image features to obtain a perception result of the at least one perception task.
[0060] According to the above technical means, in the model reasoning process, the multi-perspective image features and multi-perspective bird's-eye view image features corresponding to the multi-perspective image information can be output based on the environmental perception model, and then the perception result of at least one perception task is obtained based on the multi-perspective image features and / or multi-perspective bird's-eye view image features. In this way, the extraction of multi-perspective image features and multi-perspective bird's-eye view image features can be completed through a neural network, thereby obtaining accurate perception results.
[0061] The present application provides a training device for an environment perception model, and the training device for an environment perception model includes:
[0062] A first acquisition unit is configured to acquire a training data set, wherein the training data set includes first image information of multiple perspectives and an environmental perception truth value of at least one perception task corresponding to the first image information; and obtain a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through an initial perception model;
[0063] A first determination unit, configured to determine a loss based on a true value of environmental perception and a predicted perception result and according to a weight of a subtask of the perception task;
[0064] The training unit is used to correct the model parameters of the initial perception model based on the loss to obtain the environment perception model.
[0065] The present application provides a vehicle driving environment sensing device, the vehicle driving environment sensing device comprising:
[0066] A second acquisition unit, used to acquire multi-view image information;
[0067] A second determination unit is used to determine the image information of the target type based on the image information of multiple perspectives; wherein the target type includes at least one or more types of a forward-looking perspective type, a forward-looking narrow-viewing perspective type, and a surrounding-viewing perspective type;
[0068] The second acquisition unit is also used to obtain the perception results of at least one perception task corresponding to the multi-perspective image information based on the image information of the target type through the environmental perception model; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task.
[0069] The present application provides an electronic device, including a processor and a memory, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps in any of the above methods are implemented.
[0070] The present application provides a vehicle device, the vehicle device includes an electronic device, and the vehicle device implements the steps in any of the above methods through the electronic device.
[0071] The present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in any of the above methods are implemented.
[0072] The present application provides a computer program product, including a computer program or instructions, which implement the steps in any of the above methods when executed by a processor.
[0073] Beneficial effects of this application:
[0074] (1) After obtaining the predicted perception results corresponding to the first image information of multiple perspectives through the initial perception model, the total loss of the model training can be determined based on the environmental perception true value and the predicted perception results and the weights of the subtasks of the perception task; in this way, the weights suitable for the subtasks can be selected to weight the losses according to the distribution differences and optimization difficulty differences of the subtasks of each perception task, that is, by using different weights between subtasks to achieve balance, multi-level training losses can be achieved, thereby improving the prediction effect and performance of the neural network;
[0075] (2) After determining the loss of each perceptual task by combining the weights of the subtasks of each perceptual task, the total loss can be further determined by combining the weights of each perceptual task; thus, the prediction effect and performance of the neural network can be improved by using different weights to balance between perceptual tasks and combining multi-task training loss and multi-level training loss;
[0076] (3) A first loss adjustment parameter corresponding to the weight of each subtask can be introduced to adaptively adjust the weight of each subtask, and a second loss adjustment parameter corresponding to the weight of each perception task can be introduced to adaptively adjust the weight of each perception task, that is, corresponding loss adjustment parameters are used for different weights to perform adaptive adjustment, so as to ensure that each perception task can achieve the optimal under ideal conditions;
[0077] (4) The vehicle driving environment is perceived through the environmental perception model, and the perception result of at least one perception task corresponding to the multi-view image information is obtained. The environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task. Therefore, the environmental perception model can ensure that each perception task can achieve the optimal state under ideal conditions, with good prediction effect and performance, and thus accurate perception results can be obtained when the vehicle driving environment is perceived based on the neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1A schematic diagram of a process flow for implementing a training method for an environment perception model proposed in an embodiment of the present application;
[0079] Figure 2 A schematic diagram of obtaining the first image information proposed in an embodiment of the present application;
[0080] Figure 3 A schematic diagram of obtaining the first image information proposed in an embodiment of the present application;
[0081] Figure 4 A schematic diagram of an initial perception model proposed in an embodiment of the present application;
[0082] Figure 5 A schematic diagram of a process flow for realizing a method for perceiving a vehicle driving environment proposed in an embodiment of the present application;
[0083] Figure 6 A schematic diagram of obtaining multi-view image information proposed in an embodiment of the present application;
[0084] Figure 7 A schematic diagram of obtaining multi-view image information proposed in an embodiment of the present application;
[0085] Figure 8 A schematic diagram of an environment perception model proposed in an embodiment of the present application;
[0086] Fig. 9 A schematic diagram of a BEV multi-task neural network perception system proposed in an embodiment of the present application;
[0087] Fig.10 A schematic diagram of a perception method based on a BEV multi-task neural network perception system proposed in an embodiment of the present application;
[0088] Fig.11 A schematic diagram of a multi-level adaptive loss training method proposed in an embodiment of the present application;
[0089] Fig.12 A schematic diagram of the structure of the training device for the environment perception model proposed in the embodiment of the present application;
[0090] Fig.13 A schematic diagram of the structure of a vehicle driving environment sensing device according to an embodiment of the present application;
[0091] Fig.14 A schematic diagram of the structure of an electronic device according to an embodiment of the present application;
[0092] Fig.15 This is a schematic diagram of the composition structure of the vehicle equipment proposed in the embodiment of the present application. DETAILED DESCRIPTION
[0093] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.
[0094] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0095] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0096] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0097] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0098] With the vigorous development and increasing popularity of automobile intelligence and electrification, more and more cars will be equipped with autonomous driving functions. It is extremely important to accurately perceive and identify the environment and obstacles around the vehicle, and then input them to the downstream modules of autonomous driving for planning, control and making correct decisions.
[0099] With the increasing development and maturity of deep learning, vehicles can be used to collect surrounding image data and combine them with deep neural networks to perceive and identify the environment and obstacles around the vehicle. This is a feasible solution and can be effectively deployed on vehicles for real-time perception, realizing the important perception module in this complex system at a relatively low cost.
[0100] Since the autonomous driving perception system is very complex and needs to perceive a variety of different elements such as lane lines, curbs, ground arrows, intersections, water barriers, divergence and confluence points, the geometric features of different elements to be perceived correspond to different perception tasks. However, it is difficult for neural networks to perceive all elements in one module, so designers will design different perception tasks into different modules. At the same time, each module may be divided into different sub-modules. For example, the lane line recognition task not only detects the lane line but also recognizes the width, direction, and category of the lane line.
[0101] Traditional multi-task neural networks design a loss function for different tasks and assign a loss weight, then weight the results of each perception task loss function to obtain the final multi-task loss. Due to differences in task data distribution and optimization difficulty, a given loss weight may cause the system to favor a certain task or subtask, and then cannot guarantee that each perception task can achieve the optimal result under ideal conditions. In this way, it is very important to balance the task effects of different perception modules so that each perception task and subtask can achieve expected performance.
[0102] In other words, the common perception task recognition method based on neural network cannot guarantee that each perception task can achieve the optimal state under ideal conditions, which reduces the prediction effect and performance of the neural network, and thus cannot obtain accurate perception results when the vehicle driving environment perception is based on the neural network.
[0103] In order to solve the above problems, the present application provides a training method and device for an environmental perception model, a perception method and device for a vehicle driving environment, a device, a storage medium, and a program product. The training device for the environmental perception model obtains a training data set; wherein the training data set includes multi-perspective first image information and an environmental perception truth value corresponding to at least one perception task of the first image information; through an initial perception model, based on the first image information, a predicted perception result of at least one perception task corresponding to the first image information is obtained; based on the environmental perception truth value and the predicted perception result, a loss is determined according to the weight of the subtask of the perception task; based on the loss, the model parameters of the initial perception model are corrected to obtain an environmental perception model. The perception device for the vehicle driving environment obtains multi-perspective image information, and determines the image information of the target type based on the multi-perspective image information; wherein the target type includes at least one or more types of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a surrounding-view perspective type; through the environmental perception model, based on the image information of the target type, a perception result of at least one perception task corresponding to the multi-perspective image information is obtained; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; and the loss is determined based on the weight of the subtask of each perception task. It can be seen that in this application, on the one hand, after obtaining the predicted perception results corresponding to the first image information of multiple perspectives through the initial perception model, the total loss of model training can be determined based on the environmental perception true value and the predicted perception results, according to the weights of the subtasks of the perception task; in this way, the weights suitable for the subtasks can be selected for weighting the losses according to the distribution differences of the subtasks of each perception task and the differences in the optimization difficulty, that is, by using different weights between the subtasks to balance, multi-level training losses are achieved, thereby improving the prediction effect and performance of the neural network; on the other hand, the vehicle driving environment is perceived through the environmental perception model, and the perception results of at least one perception task corresponding to the multi-perspective image information are obtained. Among them, the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task, therefore, the environmental perception model can ensure that each perception task can achieve the optimal under ideal conditions, with better prediction effect and performance, and thus accurate perception results can be obtained when the vehicle driving environment is perceived based on the neural network.
[0104] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0105] An embodiment of the present application provides a method for training an environmental perception model. The method for training an environmental perception model can be applied to a training device or electronic device for an environmental perception model, and can also be applied to any terminal that includes a training device or electronic device for an environmental perception model.
[0106] Below, taking the training device of the environment perception model as an example, the training method of the environment perception model proposed in the embodiment of the present application is exemplified.
[0107] Furthermore, in the embodiments of the present application, Figure 1 The following is a flow chart of the training method for the environment perception model proposed in the embodiment of the present application. Figure 1 As shown, the training method of the environment perception model may include the following steps:
[0108] Step 101: Acquire a training data set; wherein the training data set includes first image information from multiple perspectives and an environmental perception truth value of at least one perception task corresponding to the first image information.
[0109] In an embodiment of the present application, a training data set is obtained so that the training data set can be used for subsequent model training processing.
[0110] In an embodiment of the present application, the training data set may include first image information from multiple perspectives, and may also include an environmental perception truth value of at least one perception task corresponding to the first image information.
[0111] It can be understood that in the embodiments of the present application, the first image information and the environmental perception true value included in the training data set correspond to each other based on the acquisition time of the first image information. That is, for the same acquisition time, it includes first image information with multiple perspectives and the corresponding environmental perception true value for at least one perception task.
[0112] It is understandable that, in the embodiment of the present application, the multi-view first image information may be an image acquired through a plurality of different acquisition viewing angles, wherein the plurality of different acquisition viewing angles may include but are not limited to: a front viewing angle, a front narrow viewing angle, a left front viewing angle, a right front viewing angle, a left rear viewing angle, a right rear viewing angle, a rear viewing angle, etc.
[0113] That is to say, in the embodiments of the present application, there is no specific limitation on the number of acquisition viewing angles.
[0114] In an embodiment of the present application, the multi-perspective first image information may be directly acquired by the environmental perception model training device through a configured image acquisition device, or may be acquired by other image acquisition devices and then sent to the environmental perception model training device.
[0115] For example, in some embodiments, the training device of the environment perception model can configure different image acquisition devices based on different positions, so that corresponding multiple first image information can be acquired for multiple different acquisition perspectives at the same time. Figure 2This is a schematic diagram of obtaining the first image information proposed in the embodiment of the present application, such as Figure 2 As shown, the training device of the environment perception model can set 6 cameras at 6 different positions respectively, and then use the set 6 cameras to collect images, so as to obtain image information of 6 collection perspectives at any time, that is, obtain the first image information of multiple perspectives. Among them, the 6 collection perspectives can include a front view perspective 21, a left front view perspective 22, a right front view perspective 23, a left rear view perspective 24, a right rear view perspective 25, and a rear view perspective 26.
[0116] For example, in some embodiments, the image acquisition device may be set based on different positions, so that a plurality of first image information corresponding to a plurality of different acquisition viewing angles may be acquired at the same time. Figure 3 This is a schematic diagram of obtaining the first image information proposed in the embodiment of the present application, such as Figure 3 As shown, the image acquisition device can be 7 cameras set at 7 different positions, and then the 7 cameras are used to perform image acquisition, so that image information of 7 acquisition perspectives at any time can be obtained, and then the acquired image information can be sent to the training device of the environment perception model, and accordingly, the training device of the environment perception model can obtain the first image information of multiple perspectives. Among them, the 7 acquisition perspectives can include a front view perspective 31, a front narrow view perspective 32, a left front view perspective 33, a right front view perspective 34, a left rear view perspective 35, a right rear view perspective 36, and a rear view perspective 37.
[0117] Further, in the embodiments of the present application, the perception task can be understood as a perception task of vehicle driving, and the perception task can be used to determine and monitor the vehicle driving environment. Among them, the environmental perception truth value of at least one perception task can represent the actual situation of the real vehicle driving environment.
[0118] Exemplarily, in some embodiments, the perception task may include, but is not limited to, at least one two-dimensional vehicle driving perception task and / or at least one three-dimensional vehicle driving perception task, which is not specifically limited in this application.
[0119] It is understandable that in the embodiments of the present application, the two-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: lane line detection task, license plate detection task, face detection task, etc. The two-dimensional vehicle driving environment perception result is the result obtained by performing the two-dimensional vehicle driving environment perception task. The two-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: pixel point category, license plate presence, face presence, etc.
[0120] It is understandable that in the embodiments of the present application, the three-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: vehicle detection task, pedestrian detection task, cyclist detection task, lane line detection task, drivable area detection task, etc. The three-dimensional vehicle driving environment perception result is the result obtained by executing the three-dimensional vehicle driving environment perception task. The three-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: three-dimensional detection frame information, drivable area information, lane line category, etc. The three-dimensional detection frame information may include but is not limited to: at least one of: the center point of the three-dimensional detection frame, the length, width and height of the three-dimensional detection frame, the orientation angle of the three-dimensional detection frame, etc. The lane line category may include but is not limited to: one of: dotted line, solid line, fishbone line, etc.
[0121] In an embodiment of the present application, based on the collection time of the multi-view first image information, there is an environment perception true value corresponding to at least one perception task. That is, the first image information and the environment perception true value correspond to each other based on the collection time of the first image information.
[0122] Exemplarily, in some embodiments, assuming that the collection time of the first image information of multiple perspectives is t, and at least one perception task includes a vehicle detection task, a pedestrian detection task, a cyclist detection task, and a lane line detection task, then, at time t, the environmental perception true value of at least one perception task corresponding to the first image information can be that there is a vehicle at time t, there are no pedestrians at time t, there are no cyclists at time t, and the lane line is a solid line at time t.
[0123] That is to say, in an embodiment of the present application, the training data included in the training data set is based on image information and environmental perception true values corresponding to the image acquisition moment, wherein the image information may be a multi-perspective multi-frame image corresponding to the image acquisition moment, and the environmental perception true value may be at least one environmental perception result of at least one perception task corresponding to the image acquisition moment.
[0124] Step 102: Obtain a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through the initial perception model.
[0125] In an embodiment of the present application, after acquiring the initial perception model, the initial perception model can be further used to obtain a predicted perception result of at least one perception task corresponding to the first image information based on the first image information.
[0126] It can be understood that in an embodiment of the present application, when obtaining a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through an initial perception model, multiple types of image information can be first determined based on the multi-perspective first image information; then, multi-perspective image features and / or multi-perspective bird's-eye view image features can be obtained based on the multiple types of image information through the initial perception model; finally, at least one perception task can be performed based on the multi-perspective image features and / or multi-perspective bird's-eye view image features to obtain a predicted perception result of at least one perception task.
[0127] In an embodiment of the present application, the multiple types may include at least a forward-looking viewing angle type, a forward-looking narrow-viewing viewing angle type, and a peripheral-viewing viewing angle type.
[0128] In some implementations, since different acquisition perspectives have different importance, different acquisition perspectives can be divided according to their importance, so as to obtain different types, that is, to divide into multiple types. Among them, the multiple types can at least include a forward-looking perspective type, a forward-looking narrow-viewing perspective type, a circumferential-viewing perspective type, etc.
[0129] It can be understood that in the embodiments of the present application, it is assumed that the multiple perspectives include a forward perspective, a forward narrow perspective, a left front perspective, a right front perspective, a left rear perspective, a right rear perspective, and a rear perspective. Since the forward perspective and the forward narrow perspective contain more important information, the forward perspective can be divided into a forward perspective type, and the forward narrow perspective can be divided into a forward narrow perspective type. At the same time, other acquisition perspectives except the forward perspective and the forward narrow perspective can be used as peripheral perspectives and divided into corresponding peripheral perspective types.
[0130] It can be understood that in an embodiment of the present application, after determining multiple types of image information based on the first image information of multiple perspectives, the multiple types of image information can be further input into the initial perception model as input information to output multi-perspective image features and / or multi-perspective bird's-eye view image features.
[0131] Furthermore, in an embodiment of the present application, the initial perception model can be used to perform at least one perception task based on the collected image information to perceive and predict the vehicle driving environment, and finally obtain a predicted perception result of the vehicle driving environment corresponding to the image information.
[0132] In an embodiment of the present application, the initial perception model may include a multi-type feature extraction network and a bird's-eye view feature extraction network. Among them, the multi-type feature extraction network is used to extract image features of different types of input image information. The bird's-eye view feature extraction network can be used to combine the acquisition parameters (such as camera parameters) corresponding to the first image information acquired from multiple perspectives to perform spatial feature conversion and bird's-eye view (BEV) feature fusion on the input image features, thereby obtaining the corresponding bird's-eye view image features.
[0133] For example, in some embodiments, Figure 4 This is a schematic diagram of the initial perception model proposed in the embodiment of the present application, such as Figure 4 As shown, the initial perception model may include a multi-type feature extraction network and a bird's-eye view feature extraction network. The multi-type feature extraction network can extract image features based on the input image information of different target types through different backbone modules corresponding to the target types, and the bird's-eye view feature extraction network can be used to perform spatial feature conversion and BEV feature fusion on the input image features, thereby obtaining the corresponding bird's-eye view image features.
[0134] Furthermore, in an embodiment of the present application, after obtaining multi-perspective image features and / or multi-perspective bird's-eye view image features through an initial perception model, at least one perception task can be finally performed based on the multi-perspective image features and / or multi-perspective bird's-eye view image features to obtain a predicted perception result of at least one perception task.
[0135] It can be understood that in the embodiments of the present application, the predicted perception result of at least one perception task can be used to predict the situation of the vehicle driving environment. Among them, the predicted perception result of at least one perception task determined based on the first image information corresponds to the environmental perception true value of at least one perception task corresponding to the first image information in the training data set, that is, the first image information, the environmental perception true value and the predicted perception result correspond to each other based on the acquisition time of the first image information.
[0136] Exemplarily, in some embodiments, assuming that the collection time of the first image information of multiple perspectives is t, and at least one perception task includes a vehicle detection task, a pedestrian detection task, a cyclist detection task, and a lane line detection task, then, at time t, the predicted perception results of at least one perception task corresponding to the first image information may be that there is no vehicle at time t, there is no pedestrian at time t, there is a cyclist at time t, and the lane line at time t is a dotted line.
[0137] That is to say, in an embodiment of the present application, the predicted perception result of at least one perception task obtained by the initial perception model may be different from the environmental perception true value of the corresponding at least one perception task. Therefore, the initial perception model can be further trained and optimized in combination with the predicted perception result and the environmental perception true value to improve the performance of the model.
[0138] Step 103: Based on the true value of environmental perception and the predicted perception result, the loss is determined according to the weights of the subtasks of the perception task.
[0139] In an embodiment of the present application, after obtaining the predicted perception result of at least one perception task corresponding to the first image information based on the first image information through the initial perception model, the loss can be further determined based on the weights of the subtasks of the perception task based on the environmental perception true value and the predicted perception result.
[0140] It is understandable that in the embodiment of the present application, in order to train and optimize the initial perception model, the training device of the environment perception model can calculate the loss based on the one-to-one correspondence between the predicted perception result and the true value of the environment perception. The loss can be determined in combination with the weight of the subtask of each perception task.
[0141] Furthermore, in an embodiment of the present application, for each subtask of at least one perception task, taking into account that the optimization difficulty between different subtasks may also be different, different weights may be selected to balance between the subtasks, and then different weights corresponding to different subtasks may be used to determine the final loss used to train the model.
[0142] In an embodiment of the present application, each perception task may correspond to one or more subtasks, where the training difficulty and objectives of different subtasks may be different due to differences in data distribution and optimization difficulty. Therefore, different loss weights may be selected for different subtasks to avoid favoring some of the subtasks during the training process, thereby ensuring that each subtask can achieve the optimal state under ideal conditions.
[0143] It can be understood that in the embodiments of the present application, the training device of the environmental perception model can pre-set corresponding weights for each subtask in each perception task, and the weights can be understood as loss weights in the model training process.
[0144] It is understandable that in the embodiments of the present application, the loss determined by the weights of the subtasks of the perception task based on the true value of environmental perception and the predicted perception result can be understood as the overall loss of model training. In other words, the total loss corresponding to the training of the initial perception model can be calculated by the weights (loss weights) corresponding to the subtasks of each perception task.
[0145] For example, in some embodiments, the process of calculating the total loss of the initial perception model by the weights corresponding to the subtasks of each perception task can be expressed as formula (1):
[0146]
[0147] Among them, w ts represents the weight corresponding to the subtask s of the perception task t, t s represents the number of subtasks corresponding to the perception task t, L(W) represents the total loss, and L ts (W) represents the loss corresponding to the subtask s of the perceptual task t, and n represents the number of perceptual tasks.
[0148] Furthermore, in an embodiment of the present application, when determining the loss based on the environmental perception true value and the predicted perception result and according to the weights of the subtasks of the perception task, the first loss parameter can be determined based on the environmental perception true value and the predicted perception result and according to the weight of each subtask of the first perception task; then at least one perception task can be traversed to determine the loss based on the first loss parameter of each perception task.
[0149] It is understandable that in the embodiment of the present application, the first perception task is any one of the at least one perception task. That is, for any perception task, the first loss parameter corresponding to the perception task can be calculated based on the weight corresponding to each subtask of the perception task, for example Then, each perception task can be traversed to determine the first loss parameter corresponding to each perception task, and then determine the final loss.
[0150] In an embodiment of the present application, the number of perception tasks t and the number of subtasks ts may both be integers greater than 0, and the present application does not specifically limit the values of t and ts.
[0151] Further, in an embodiment of the present application, when determining the first loss parameter based on the environmental perception true value and the predicted perception result and according to the weight of each subtask of the first perception task, the first loss of the first subtask can be determined based on the environmental perception true value, the predicted perception result and the loss function of the first subtask; then the second loss of the first subtask can be determined based on the weight of the first subtask and the first loss; finally, each subtask of the first perception task can be traversed, and the first loss parameter corresponding to the first perception task can be determined based on the second loss of each subtask.
[0152] It can be understood that in the embodiment of the present application, the first subtask is any subtask of the first perception task. That is to say, for any subtask, the corresponding first loss can be calculated based on the environmental perception true value and the predicted perception result through the corresponding loss function, for example, the first loss L of subtask s of perception task t is ts (W), and then combined with the weights corresponding to the subtasks, such as the weight w of the subtask s of the perception task t ts , determine the corresponding weighted second loss, for example, the second loss w of subtask s of perception task t ts L ts (W), and finally traverse each subtask, determine the second loss of each subtask, and then determine the first loss parameter of each perception task.
[0153] Exemplarily, in some embodiments, for a perception task, the loss corresponding to the subtask can be determined based on the loss function corresponding to the subtask, wherein different subtasks correspond to different losses. For example, the subtasks of the perception task of lane line recognition correspond to roadside loss, lane line loss, and other losses; the subtasks of the perception task of intersection recognition correspond to divergence point loss, confluence point loss, and other losses; the subtasks of the perception task of ground arrow recognition correspond to direction loss, width loss, and other losses; the subtasks of the perception task of static obstacles correspond to cone barrel loss, water barrier loss, and other losses; the subtasks of the perception task of intersection recognition correspond to straight arrow loss, U-turn arrow loss, and other losses; the subtasks of the perception task of road sign recognition correspond to speed bump loss, stop line loss, and other losses.
[0154] Furthermore, in an embodiment of the present application, when determining the loss based on the environmental perception true value and the predicted perception result and according to the weight of the subtask of the perception task, a first loss adjustment parameter corresponding to the weight of each subtask may be introduced to adaptively adjust the weight of each subtask. Specifically, the first loss parameter may be determined based on the environmental perception true value and the predicted perception result, according to the weight of each subtask of the first perception task and the first loss adjustment parameter corresponding to the weight of each subtask; then, at least one perception task may be traversed to determine the loss based on the first loss parameter of each perception task.
[0155] For example, in some embodiments, the process of calculating the total loss of the initial perception model by using the weight corresponding to each subtask of the perception task and the corresponding first loss adjustment parameter can be expressed as formula (2):
[0156]
[0157] Among them, w ts represents the weight corresponding to the subtask s of the perception task t, ts represents the number of subtasks corresponding to the perception task t, L(W) represents the total loss, and L ts (W) represents the loss corresponding to the subtask s of the perception task t, σ ts represents the first loss adjustment parameter corresponding to the subtask s of the perception task t, and n represents the number of perception tasks.
[0158] It can be understood that in the embodiments of the present application, for any perception task, the first loss parameter corresponding to the perception task can be calculated based on the weight of each subtask of the perception task and the adjustment item corresponding to the weight of the subtask, that is, the first loss adjustment parameter, for example Then, each perception task can be traversed to determine the first loss parameter corresponding to each perception task, and then determine the final loss.
[0159] It can be understood that in the embodiment of the present application, the first loss adjustment parameter is used to adaptively adjust the weight of the subtask. The value range of the first loss adjustment parameter can be (0, 1), and for different subtasks, the corresponding value of the first loss adjustment parameter can be different, that is, the first loss adjustment parameter is an adaptive adjustment item corresponding to the weight of the subtask.
[0160] Further, in an embodiment of the present application, when determining the first loss parameter based on the environmental perception true value and the predicted perception result, according to the weight of each subtask of the first perception task and the first loss adjustment parameter corresponding to the weight of each subtask, the first loss of the first subtask can be determined based on the environmental perception true value, the predicted perception result and the loss function of the first subtask; then, the second loss of the first subtask can be determined based on the first loss adjustment parameter corresponding to the first subtask, the weight of the first subtask and the first loss; finally, each subtask of the first perception task is traversed, and the first loss parameter corresponding to the first perception task is determined based on the second loss of each subtask.
[0161] It can be understood that in the embodiment of the present application, for any subtask, the corresponding first loss can be calculated based on the environmental perception true value and the predicted perception result through the corresponding loss function, for example, the first loss L of the subtask s of the perception task t is ts (W), and then combine the weight corresponding to the subtask and the first loss adjustment parameter corresponding to the weight, for example, the weight w of the subtask s of the perception task t ts and the corresponding first loss adjustment parameter σ ts , determine the corresponding weighted second loss, for example, the second loss of subtask s of perception task t Finally, each subtask is traversed, the second loss of each subtask is determined, and then the first loss parameter of each perception task is determined.
[0162] Further, in an embodiment of the present application, when traversing at least one perception task and determining the loss based on a first loss parameter of each perception task, at least one perception task can be traversed, and a second loss parameter can be determined based on the weight of each perception task and the first loss parameter; and then the loss can be determined based on the second loss parameter.
[0163] That is to say, in an embodiment of the present application, when determining the total loss of the model training using the first loss parameter corresponding to each perception task, different weights can be introduced for different perception tasks to perform weighted operations on the first loss parameter, thereby avoiding bias towards some of the perception tasks during the training process, thereby ensuring that each perception task can achieve the optimal result under ideal conditions.
[0164] For example, in some embodiments, the process of introducing the weights corresponding to the perception task to calculate the total loss can be expressed as formula (3):
[0165]
[0166] Among them, w t represents the weight corresponding to the perception task t, w ts represents the weight corresponding to the subtask s of the perception task t, t s represents the number of subtasks corresponding to the perception task t, L(W) represents the total loss, and L ts (W) represents the loss corresponding to the subtask s of the perception task t, σ ts represents the first loss adjustment parameter corresponding to the subtask s of the perception task t, and n represents the number of perception tasks.
[0167] Further, in an embodiment of the present application, when determining the loss based on the second loss parameter, the first loss adjustment parameter corresponding to each subtask of each perception task may be selected to determine the first operand value, for example The loss may then be determined based on the second loss parameter and the first operand value.
[0168] Further, in an embodiment of the present application, when traversing at least one perception task and determining the loss based on a first loss parameter of each perception task, traversing at least one perception task, determining a second loss parameter based on the weight of each perception task, a second loss adjustment parameter corresponding to the weight of each perception task, and the first loss parameter; and then determining the loss based on the second loss parameter.
[0169] That is to say, in an embodiment of the present application, when the first loss parameter corresponding to each perception task is used to determine the total loss of the model training, in the process of introducing different weights to perform weighted operation on the first loss parameter, the corresponding second loss adjustment parameter can also be used for adaptive adjustment according to different weights to avoid bias towards some perception tasks during the training process, thereby ensuring that each perception task can achieve the optimal state under ideal conditions.
[0170] For example, in some embodiments, the process of introducing the weight corresponding to the perception task and the corresponding second loss adjustment parameter to calculate the total loss can be expressed as formula (4):
[0171]
[0172] Among them, w t represents the weight corresponding to the perception task t, w ts represents the weight corresponding to the subtask s of the perception task t, t s represents the number of subtasks corresponding to the perception task t, L(W) represents the total loss, and L ts (W) represents the loss corresponding to the subtask s of the perception task t, σ ts represents the first loss adjustment parameter corresponding to the subtask s of the perception task t, σ t represents the second loss adjustment parameter corresponding to the perception task t, and n represents the number of perception tasks.
[0173] It can be understood that in the embodiment of the present application, the second loss adjustment parameter is used to adaptively adjust the weight of the perception task. The value range of the second loss adjustment parameter can be (0, 1), and for different perception tasks, the corresponding value of the second loss adjustment parameter can be different, that is, the second loss adjustment parameter is an adaptive adjustment item corresponding to the weight of the perception task.
[0174] Further, in an embodiment of the present application, when determining the loss based on the second loss parameter, the first loss adjustment parameter corresponding to each subtask of each perception task may be selected to determine the first operand value, for example At the same time, a second loss adjustment parameter corresponding to each perception task may be selected to determine the second operand value, for example The loss may ultimately be determined based on the second loss parameter, the first operand value, and the second operand value.
[0175] It can be seen that in an embodiment of the present application, in the process of determining the total loss of the initial perception model training based on the environmental perception true value and the predicted perception result, the weight of the subtask of the perception task can be introduced to weight the loss corresponding to the subtask. At the same time, the adjustment item (first loss adjustment parameter) can be used when weighting the loss corresponding to the subtask to further adaptively adjust the weight of the subtask. In this way, optimization can be performed on the basis of multi-level weighted losses, and the weights of each subtask can be adaptively modified through the proposed multi-level adaptive weighted loss. Furthermore, the weight of the perception task can be introduced to weight the loss corresponding to the perception task. At the same time, the adjustment item (second loss adjustment parameter) can be used when weighting the loss corresponding to the perception task to further adaptively adjust the weight of the perception task. In this way, optimization can be performed on the basis of multi-task weighted losses, and the weights of each perception task can be adaptively modified through the proposed multi-task adaptive weighted losses.
[0176] Furthermore, in an embodiment of the present application, when determining the loss based on the environmental perception true value and the predicted perception result and according to the weights of the subtasks of the perception task, the target subtask can be first determined among all the subtasks of the perception task based on the indicator parameters of the subtask; then, based on the environmental perception true value and the predicted perception result and according to the weight of the target subtask, a first loss parameter can be determined; finally, the loss can be determined based on the first loss parameter of the second perception task; wherein the second perception task is one or more perception tasks among at least one perception task.
[0177] In an embodiment of the present application, the indicator parameter of the subtask may be used to characterize the importance of the subtask, that is, different subtasks of the same perception task may have different corresponding importance.
[0178] Accordingly, in an embodiment of the present application, when determining a target subtask among all subtasks of a perception task based on the indicator parameters of the subtask, all subtasks of the perception task can be screened according to the indicator parameters of the subtask, and one or more subtasks with higher importance can be selected as the target subtasks of the perception task.
[0179] In an embodiment of the present application, the first loss parameter corresponding to the perception task can be determined based on the true value of environmental perception and the predicted perception result, according to the weight of the target subtask; then each second perception task can be traversed, and the loss can be determined based on the first loss parameter of each second perception task.
[0180] It is understandable that in the embodiment of the present application, the second perception task is one or more perception tasks in at least one perception task. That is to say, for one or more perception tasks in all perception tasks, the first loss parameters corresponding to the one or more perception tasks can be calculated based on the weights corresponding to the target subtasks of the perception tasks, and then the final loss can be determined.
[0181] It can be seen that in an embodiment of the present application, when determining the overall loss, some or all of the perception tasks can be selected to determine the loss, that is, it is not necessary to traverse each perception task, but one or more perception tasks can be selected to calculate the overall loss; at the same time, for the perception task used to calculate the loss, the first loss parameter of the perception task can also be determined based on the weights of some or all of the subtasks of the perception task, that is, it is not necessary to traverse each subtask, but one or more subtasks can be selected to calculate the first loss parameter of the perception task.
[0182] It is understood that in the embodiments of the present application, the loss used to correct the model parameters generally involves quantifying the gap between the model prediction result and the actual result. When calculating the loss of the perception task and / or the subtask of the perception task, the choice of the loss function is not specifically limited, wherein the loss functions selected for different perception tasks and / or subtasks of the perception task may be the same or different, and the present application does not specifically limit it.
[0183] It is understood that in the embodiments of the present application, the loss function is a function that measures the gap between the model prediction result and the actual result. Among them, the loss function may include but is not limited to the following:
[0184] Mean Squared Error (MSE): Also known as L2 loss, it measures the average of the squares of the differences between the predicted value and the true value. It is sensitive to outliers because outliers can significantly increase the loss value.
[0185] Mean Absolute Error (MAE): Also known as L1 loss, it measures the average of the absolute values of the difference between the predicted value and the true value. It is less sensitive to outliers because outliers do not cause the loss value to increase dramatically.
[0186] Huber loss: It combines the advantages of MSE and MAE. MSE is used when the error is small, and MAE is used when the error is large to improve the robustness to outliers.
[0187] Step 104: Correct the model parameters of the initial perception model based on the loss to obtain an environment perception model.
[0188] In an embodiment of the present application, after determining the loss based on the environmental perception true value and the predicted perception result and according to the weights of the subtasks of the perception task, the model parameters of the initial perception model can be further corrected based on the loss to finally obtain the environmental perception model.
[0189] It can be understood that in the embodiments of the present application, after obtaining the quantified loss of the gap between the model prediction results and the actual results, based on the loss, the model parameters can be adjusted through the optimization algorithm to reduce this gap, that is, the purpose of the correction is to adjust the model parameters through the optimization algorithm to minimize the loss function.
[0190] Exemplarily, in some embodiments, the optimization algorithms used when modifying the model parameters may include but are not limited to the following:
[0191] Gradient descent method: By calculating the gradient of the loss function with respect to the model parameters and updating the parameters in the opposite direction of the gradient, the loss value is gradually reduced.
[0192] Stochastic gradient descent: Only one sample or a small batch of samples is used to calculate the gradient and update the parameters in each iteration. This method is suitable for large-scale data sets, but may lead to slow convergence and unstable parameter updates.
[0193] Momentum method: A momentum term is introduced based on the gradient descent method to accelerate convergence and reduce oscillation.
[0194] Adaptive learning rate methods: such as Adam, RMSprop, etc. These methods dynamically adjust the learning rate according to historical gradient information to improve optimization efficiency and effect.
[0195] Furthermore, in an embodiment of the present application, after the model parameters of the initial perception model are corrected based on the loss to obtain the environmental perception model, at least one perception task can be performed through the environmental perception model based on the collected multi-perspective image information to obtain corresponding perception results.
[0196] To sum up, in the training method of an environmental perception model proposed in an embodiment of the present application, during the training process of the environmental perception model, taking into account the differences in data distribution and optimization difficulty between different subtasks corresponding to the perception task, the loss of the corrected model parameters is determined in combination with the weights of the subtasks, that is, for different subtasks, different weights are used to balance the subtasks, thereby achieving subtask weight optimization and forming a multi-level weighted loss, so that the final total loss can ensure that each subtask can be optimized to an ideal state, thereby obtaining an environmental perception model with better prediction effect based on loss correction training.
[0197] The present application provides a training method for an environmental perception model, wherein the training device of the environmental perception model obtains a training data set; wherein the training data set includes multi-perspective first image information and the environmental perception true value of at least one perception task corresponding to the first image information; through the initial perception model, based on the first image information, a predicted perception result of at least one perception task corresponding to the first image information is obtained; based on the environmental perception true value and the predicted perception result, according to the weight of the subtask of the perception task, the loss is determined; based on the loss, the model parameters of the initial perception model are corrected to obtain the environmental perception model. It can be seen that in the present application, after the predicted perception result corresponding to the multi-perspective first image information is obtained through the initial perception model, the total loss of the model training can be determined based on the environmental perception true value and the predicted perception result, according to the weight of the subtask of the perception task; in this way, the weight suitable for the subtask can be selected for the distribution difference of the subtask of each perception task and the optimization difficulty difference to weight the loss, that is, by using different weights to balance between subtasks, a multi-level training loss is achieved, thereby improving the prediction effect and performance of the neural network.
[0198] Based on the above embodiments, another embodiment of the present application provides a method for sensing a vehicle driving environment. The method for sensing a vehicle driving environment can be applied to a sensing device or electronic device for a vehicle driving environment, and can also be applied to any terminal that includes a sensing device or electronic device for a vehicle driving environment.
[0199] Below, taking the vehicle driving environment perception device as an example, the vehicle driving environment perception method proposed in the embodiment of the present application is exemplarily described.
[0200] Furthermore, in the embodiments of the present application, Figure 5 The following is a flow chart of the method for realizing the vehicle driving environment perception proposed in the embodiment of the present application. Figure 5 As shown, the method for sensing the vehicle driving environment may include the following steps:
[0201] Step 201: Acquire multi-view image information, and determine target type image information based on the multi-view image information; wherein the target type includes at least one or more of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a peripheral-view perspective type.
[0202] In an embodiment of the present application, a perception device of a vehicle driving environment may first acquire multi-perspective image information, and then determine image information of a corresponding target type based on the multi-perspective image information.
[0203] It is understood that in the embodiment of the present application, the multi-view image information may be an image acquired through a plurality of different acquisition viewing angles. The plurality of different acquisition viewing angles may include, but are not limited to, a front viewing angle, a front narrow viewing angle, a left front viewing angle, a right front viewing angle, a left rear viewing angle, a right rear viewing angle, a rear viewing angle, etc.
[0204] That is to say, in the embodiments of the present application, there is no specific limitation on the number of acquisition viewing angles.
[0205] In an embodiment of the present application, the multi-perspective image information may be directly acquired by a perception device of the vehicle's driving environment through a configured image acquisition device, or may be acquired by other image acquisition devices and then sent to the perception device of the vehicle's driving environment.
[0206] For example, in some embodiments, the sensing device of the vehicle driving environment can configure different image acquisition devices based on different positions, so that corresponding multiple image information can be collected from multiple different acquisition perspectives at the same time. Figure 6 This is a schematic diagram of obtaining multi-view image information proposed in an embodiment of the present application, such as Figure 6 As shown, the sensing device of the vehicle driving environment can be a vehicle device, and the vehicle device can be provided with 5 cameras at 5 different positions, and then the 5 cameras are used to collect images, so that the image information of 5 collected viewing angles at any time can be obtained, that is, the first image information of multiple viewing angles can be obtained. Among them, the 5 collected viewing angles can include a front viewing angle 61, a left front viewing angle 62, a right front viewing angle 63, a left rear viewing angle 64, and a right rear viewing angle 65.
[0207] For example, in some embodiments, the sensing device of the vehicle driving environment can configure different image acquisition devices based on different positions, so that corresponding multiple image information can be collected from multiple different acquisition perspectives at the same time. Figure 7 This is a schematic diagram of obtaining multi-view image information proposed in an embodiment of the present application, such as Figure 7 As shown, the sensing device of the vehicle driving environment can be a vehicle device, and the vehicle device can be provided with 7 cameras at 7 different positions, and then the 7 cameras are used to collect images, so that image information of 7 collection angles at any time can be obtained, that is, image information of multiple angles can be obtained. Among them, the 7 collection angles can include a front view angle 71, a front narrow view angle 72, a left front view angle 73, a right front view angle 74, a left rear view angle 75, a right rear view angle 76, and a rear view angle 77.
[0208] It is understandable that in the embodiment of the present application, after acquiring the multi-view image information, the image information of the corresponding target type can be further determined based on the multi-view image information. The target type can include at least one or more of the forward viewing angle type, the forward narrow viewing angle type, and the surrounding viewing angle type.
[0209] In some implementations, since different acquisition perspectives have different importance, the different acquisition perspectives may be divided according to their importance, thereby obtaining different types, that is, dividing the target types.
[0210] It can be understood that in the embodiments of the present application, it is assumed that the multiple perspectives include a forward perspective, a forward narrow perspective, a left front perspective, a right front perspective, a left rear perspective, a right rear perspective, and a rear perspective. Since the forward perspective and the forward narrow perspective contain more important information, the forward perspective can be divided into a forward perspective type, and the forward narrow perspective can be divided into a forward narrow perspective type. At the same time, other acquisition perspectives except the forward perspective and the forward narrow perspective can be used as peripheral perspectives and divided into corresponding peripheral perspective types.
[0211] Step 202: Obtain, through the environmental perception model, a perception result of at least one perception task corresponding to the multi-perspective image information based on the image information of the target type; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; and the loss is determined based on the weight of the subtask of each perception task.
[0212] In an embodiment of the present application, after acquiring multi-perspective image information and determining the image information of the target type based on the multi-perspective image information, the perception device of the vehicle's driving environment can obtain the perception results of at least one perception task corresponding to the multi-perspective image information based on the image information of the target type through the environmental perception model.
[0213] Furthermore, in the embodiments of the present application, the perception task may be understood as a perception task of vehicle driving, and the perception task may be used to determine and monitor the vehicle driving environment.
[0214] Accordingly, in some embodiments, the perception result of at least one perception task can characterize the predicted condition of the vehicle driving environment, that is, the perception result can be used to predict the condition of the vehicle driving environment.
[0215] Exemplarily, in some embodiments, the perception task may include, but is not limited to, at least one two-dimensional vehicle driving perception task and / or at least one three-dimensional vehicle driving perception task, which is not specifically limited in this application.
[0216] It is understandable that in the embodiments of the present application, the two-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: lane line detection task, license plate detection task, face detection task, etc. The two-dimensional vehicle driving environment perception result is the result obtained by performing the two-dimensional vehicle driving environment perception task. The two-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: pixel point category, license plate presence, face presence, etc.
[0217] It is understandable that in the embodiments of the present application, the three-dimensional vehicle driving environment perception task may include but is not limited to: at least one of: vehicle detection task, pedestrian detection task, cyclist detection task, lane line detection task, drivable area detection task, etc. The three-dimensional vehicle driving environment perception result is the result obtained by executing the three-dimensional vehicle driving environment perception task. The three-dimensional vehicle driving environment perception result may include but is not limited to: at least one of: three-dimensional detection frame information, drivable area information, lane line category, etc. The three-dimensional detection frame information may include but is not limited to: at least one of: the center point of the three-dimensional detection frame, the length, width and height of the three-dimensional detection frame, the orientation angle of the three-dimensional detection frame, etc. The lane line category may include but is not limited to: one of: dotted line, solid line, fishbone line, etc.
[0218] In an embodiment of the present application, based on the acquisition time of multi-view image information, a corresponding perception result can be obtained for at least one perception task through an environment perception model. In other words, the image information and the perception result correspond to each other based on the acquisition time of the image information.
[0219] Exemplarily, in some embodiments, assuming that the acquisition time of multi-perspective image information is t, and at least one perception task includes a vehicle detection task, a pedestrian detection task, a cyclist detection task, and a lane line detection task, then the corresponding perception results obtained through the environmental perception model for at least one perception task can be that there is a vehicle at time t, there are no pedestrians at time t, there is no cyclist at time t, and the lane line is a solid line at time t.
[0220] Furthermore, in an embodiment of the present application, when obtaining the perception result of at least one perception task corresponding to multi-perspective image information based on the image information of the target type through the environmental perception model, the multi-perspective image features and / or multi-perspective bird's-eye view image features can be first obtained through the environmental perception model based on the image information of the target type; and then at least one perception task is performed based on the multi-perspective image features and / or multi-perspective bird's-eye view image features to obtain the perception result of at least one perception task.
[0221] It can be understood that in an embodiment of the present application, after determining the image information of the target type based on the multi-perspective image information, the image information of the target type can be further input as input information into the environmental perception model to output multi-perspective image features and / or multi-perspective bird's-eye view image features.
[0222] Furthermore, in an embodiment of the present application, the environmental perception model can be used to perform at least one perception task based on the collected image information to perceive and predict the vehicle driving environment, and finally obtain a predicted perception result of the vehicle driving environment corresponding to the image information.
[0223] In an embodiment of the present application, the environmental perception model may include a multi-type feature extraction network and a bird's-eye view feature extraction network. Among them, the multi-type feature extraction network is used to extract image features of different types of input image information. The bird's-eye view feature extraction network can be used to combine the acquisition parameters (such as camera parameters) corresponding to the first image information acquired from multiple perspectives to perform spatial feature conversion and BEV feature fusion on the input image features, thereby obtaining the corresponding bird's-eye view image features.
[0224] For example, in some embodiments, Figure 8 This is a schematic diagram of the environment perception model proposed in the embodiment of the present application, such as Figure 8 As shown, the environment perception model may include a multi-type feature extraction network and a bird's-eye view feature extraction network. The multi-type feature extraction network can extract image features based on the input image information of different target types through different backbone modules corresponding to the target types, and the bird's-eye view feature extraction network can be used to perform spatial feature conversion and BEV feature fusion on the input image features to obtain the corresponding bird's-eye view image features.
[0225] It can be understood that in an embodiment of the present application, after obtaining multi-perspective image features and / or multi-perspective bird's-eye view image features through an initial perception model, at least one perception task can ultimately be performed based on the multi-perspective image features and / or multi-perspective bird's-eye view image features to obtain a perception result of at least one perception task.
[0226] Further, in an embodiment of the present application, the environment perception model used to predict the vehicle driving environment can be obtained by modifying the model parameters of the initial perception model based on the loss, wherein the loss is determined based on the weight of the subtask of each perception task.
[0227] It is understandable that in an embodiment of the present application, in order to train and optimize the initial perception model, the loss can be calculated based on the one-to-one correspondence between the predicted perception results and the true value of environmental perception. Among them, the loss can be determined in combination with the weight of the subtask of each perception task. This is because for each subtask of the perception task in at least one perception task, considering that the optimization difficulty between different subtasks will also be different, it is possible to choose to use different weights between subtasks for balance, and then different weights corresponding to different subtasks can be used to determine the loss ultimately used for training the model.
[0228] It can be understood that in the embodiments of the present application, each perception task may correspond to one or more subtasks, where the training difficulty and objectives of different subtasks may be different due to differences in data distribution and optimization difficulty. Therefore, different loss weights can be selected for different subtasks to avoid bias towards some of the subtasks during the training process, thereby ensuring that each subtask can achieve the optimal state under ideal conditions.
[0229] It can be understood that in the embodiments of the present application, a corresponding weight can be set in advance for each subtask in each perception task, and the weight can be understood as the loss weight in the model training process.
[0230] That is to say, in an embodiment of the present application, the total loss corresponding to the training of the initial perception model can be calculated by the weights (loss weights) corresponding to the subtasks of each perception task.
[0231] Exemplarily, in some embodiments, the process of calculating the total loss of the initial perception model by using the weights corresponding to the subtasks of each perception task can be expressed as formula (1).
[0232] Further, in an embodiment of the present application, when determining the loss based on the environmental perception true value and the predicted perception result according to the weights of the subtasks of the perception task, for any perception task, the first loss parameter corresponding to the perception task can be calculated based on the weight corresponding to each subtask of the perception task, for example Then, each perception task can be traversed to determine the first loss parameter corresponding to each perception task, and then determine the final loss.
[0233] In an embodiment of the present application, the number of perception tasks t and the number of subtasks ts may both be integers greater than 0, and the present application does not specifically limit the values of t and ts.
[0234] Further, in an embodiment of the present application, when determining the first loss parameter based on the environmental perception true value and the predicted perception result and according to the weight of each subtask of the first perception task, for any subtask, the corresponding first loss can be obtained by first calculating the corresponding loss function based on the environmental perception true value and the predicted perception result. For example, the first loss L of the subtask s of the perception task t is ts (W), and then combined with the weights corresponding to the subtasks, such as the weight w of the subtask s of the perception task t ts , determine the corresponding weighted second loss, for example, the second loss w of subtask s of perception task t ts L ts (W), and finally traverse each subtask, determine the second loss of each subtask, and then determine the first loss parameter of each perception task.
[0235] Exemplarily, in some embodiments, for a perception task, the loss corresponding to the subtask can be determined based on the loss function corresponding to the subtask, wherein different subtasks correspond to different losses. For example, the subtasks of the perception task of lane line recognition correspond to roadside loss, lane line loss, and other losses; the subtasks of the perception task of intersection recognition correspond to divergence point loss, confluence point loss, and other losses; the subtasks of the perception task of ground arrow recognition correspond to direction loss, width loss, and other losses; the subtasks of the perception task of static obstacles correspond to cone barrel loss, water barrier loss, and other losses; the subtasks of the perception task of intersection recognition correspond to straight arrow loss, U-turn arrow loss, and other losses; the subtasks of the perception task of road sign recognition correspond to speed bump loss, stop line loss, and other losses.
[0236] Furthermore, in an embodiment of the present application, when determining the loss based on the environmental perception true value and the predicted perception result and according to the weight of the subtask of the perception task, a first loss adjustment parameter corresponding to the weight of each subtask may be introduced to adaptively adjust the weight of each subtask. Specifically, the first loss parameter may be determined based on the environmental perception true value and the predicted perception result, according to the weight of each subtask of the first perception task and the first loss adjustment parameter corresponding to the weight of each subtask; then, at least one perception task may be traversed to determine the loss based on the first loss parameter of each perception task.
[0237] Exemplarily, in some embodiments, the process of calculating the total loss of the initial perception model by using the weights corresponding to the subtasks of each perception task and the corresponding first loss adjustment parameter can be expressed as formula (2).
[0238] It can be understood that in the embodiments of the present application, for any perception task, the first loss parameter corresponding to the perception task can be calculated based on the weight of each subtask of the perception task and the adjustment item corresponding to the weight of the subtask, that is, the first loss adjustment parameter, for example Then, each perception task can be traversed to determine the first loss parameter corresponding to each perception task, and then determine the final loss.
[0239] It can be understood that in the embodiment of the present application, the first loss adjustment parameter is used to adaptively adjust the weight of the subtask. The value range of the first loss adjustment parameter can be (0, 1), and for different subtasks, the corresponding value of the first loss adjustment parameter can be different, that is, the first loss adjustment parameter is an adaptive adjustment item corresponding to the weight of the subtask.
[0240] Further, in an embodiment of the present application, when determining the first loss parameter based on the environmental perception true value and the predicted perception result, according to the weight of each subtask of the first perception task and the first loss adjustment parameter corresponding to the weight of each subtask, for any subtask, the corresponding first loss can be first calculated based on the environmental perception true value and the predicted perception result through the corresponding loss function, for example, the first loss L of the subtask s of the perception task t is ts (W), and then combine the weight corresponding to the subtask and the first loss adjustment parameter corresponding to the weight, for example, the weight w of the subtask s of the perception task t ts and the corresponding first loss adjustment parameter σ ts , determine the corresponding weighted second loss, for example, the second loss of subtask s of perception task t Finally, each subtask is traversed, the second loss of each subtask is determined, and then the first loss parameter of each perception task is determined.
[0241] Furthermore, in an embodiment of the present application, when determining the total loss of the model training using the first loss parameter corresponding to each perception task, different weights can be introduced for different perception tasks to perform weighted operations on the first loss parameter, thereby avoiding bias towards some of the perception tasks during the training process, thereby ensuring that each perception task can achieve the optimal result under ideal conditions.
[0242] Exemplarily, in some embodiments, the process of introducing the weights corresponding to the perception task to calculate the total loss can be expressed as formula (3).
[0243] Further, in an embodiment of the present application, when determining the loss based on the second loss parameter, the first loss adjustment parameter corresponding to each subtask of each perception task may be selected to determine the first operand value, for example The loss may then be determined based on the second loss parameter and the first operand value.
[0244] Furthermore, in an embodiment of the present application, when determining the total loss of model training using the first loss parameter corresponding to each perception task, in the process of introducing different weights to perform weighted operations on the first loss parameter, the corresponding second loss adjustment parameter can be used for adaptive adjustments according to different weights to avoid bias towards some perception tasks during the training process, thereby ensuring that each perception task can achieve the optimal state under ideal conditions.
[0245] Exemplarily, in some embodiments, the process of introducing the weight corresponding to the perception task and the corresponding second loss adjustment parameter to calculate the total loss can be expressed as formula (4).
[0246] It can be understood that in the embodiment of the present application, the second loss adjustment parameter is used to adaptively adjust the weight of the perception task. The value range of the second loss adjustment parameter can be (0, 1), and for different perception tasks, the corresponding value of the second loss adjustment parameter can be different, that is, the second loss adjustment parameter is an adaptive adjustment item corresponding to the weight of the perception task.
[0247] Further, in an embodiment of the present application, when determining the loss based on the second loss parameter, the first loss adjustment parameter corresponding to each subtask of each perception task may be selected to determine the first operand value, for example At the same time, a second loss adjustment parameter corresponding to each perception task may be selected to determine the second operand value, for example The loss may ultimately be determined based on the second loss parameter, the first operand value, and the second operand value.
[0248] It can be seen that in an embodiment of the present application, during the training process of the environmental perception model, the weights of the subtasks of the perception task can be introduced to weight the losses corresponding to the subtasks. At the same time, the adjustment items (first loss adjustment parameters) can be used when weighting the losses corresponding to the subtasks to further adaptively adjust the weights of the subtasks. In this way, optimization can be performed on the basis of multi-level weighted losses, and the weights of each subtask can be adaptively modified through the proposed multi-level adaptive weighted losses. Furthermore, the weights of the perception tasks can be introduced to weight the losses corresponding to the perception tasks. At the same time, the adjustment items (second loss adjustment parameters) can be used when weighting the losses corresponding to the perception tasks to further adaptively adjust the weights of the perception tasks. In this way, optimization can be performed on the basis of multi-task weighted losses, and the weights of each perception task can be adaptively modified through the proposed multi-task adaptive weighted losses.
[0249] The present application provides a method for perceiving a vehicle driving environment, wherein a perception device for the vehicle driving environment acquires image information from multiple perspectives, and determines image information of a target type based on the image information from multiple perspectives; wherein the target type includes at least one or more of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a circumferential perspective type; through an environmental perception model, based on the image information of the target type, a perception result of at least one perception task corresponding to the image information from multiple perspectives is acquired; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; and the loss is determined based on the weight of the subtask of each perception task. It can be seen that in the present application, the vehicle driving environment is perceived through the environmental perception model, and the perception result of at least one perception task corresponding to the image information from multiple perspectives is acquired. wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; and the loss is determined based on the weight of the subtask of each perception task, therefore, the environmental perception model can ensure that each perception task can achieve the optimal under ideal conditions, and has a good prediction effect and performance, and thus can obtain accurate perception results when the vehicle driving environment is perceived based on a neural network.
[0250] Based on the above embodiments, another embodiment of the present application proposes a multi-level adaptive training loss and method for a vehicle environment perception system, which may include a multi-level adaptive loss training loss and method for a vehicle perception system for a BEV neural network, which can solve the current problems of balancing different tasks and optimizing effects when using the BEV model for multi-task perception, thereby ensuring the performance of the perception system in a comprehensive autonomous driving scenario.
[0251] The multi-level adaptive training loss and method for the vehicle environment perception system proposed in the embodiment of the present application can be applied to neural network models with perception prediction functions, including but not limited to BEV multi-task neural network perception systems. That is, the present application does not limit the specific type and results of the neural network model, and is not coupled with any specific perception system.
[0252] For example, in some embodiments, Fig. 9 This is a schematic diagram of the BEV multi-task neural network perception system proposed in the embodiment of the present application, such as Fig. 9 As shown, the BEV multi-task neural network perception system may include a multi-type feature extraction network and a bird's-eye view feature extraction network, wherein the multi-type feature extraction network may perform image feature extraction based on the input image information of different target types through different backbone modules corresponding to the target types (e.g. backbone module 1, backbone module 2, etc.), and the bird's-eye view feature extraction network may be used to perform spatial feature conversion and BEV feature fusion on the input image features, thereby obtaining corresponding bird's-eye view image features.
[0253] For example, in some embodiments, Fig.10 This is a schematic diagram of a perception method based on a BEV multi-task neural network perception system proposed in an embodiment of the present application, such as Fig.10 As shown, assuming that the acquired multi-perspective image information includes image information of six acquisition perspectives, namely, the front view perspective, the left front view perspective, the right front view perspective, the left rear view perspective, the right rear view perspective, and the rear view perspective, then two target types of image information can be determined accordingly, that is, the front view perspective is divided into the front view perspective type to obtain the image information of the front view perspective type, and the other acquisition perspectives except the front view perspective can be used as the surrounding view perspective, and the corresponding image information is divided into the surrounding view perspective type. The multi-type feature extraction network can perform image feature extraction based on the input image information of different target types through different backbone modules corresponding to the target type. For example, the multi-type feature extraction network performs forward feature extraction and surrounding feature extraction based on the image information of the forward view perspective type, and then inputs the obtained multi-perspective image features into the bird's-eye view feature extraction network for spatial feature conversion and BEV feature fusion, and performs processing of at least one subsequent perception task based on the obtained multi-perspective image features and / or multi-perspective bird's-eye view image features. Among them, at least one perception task may include but is not limited to BEV lane line recognition, intersection recognition, ground arrow recognition, intersection recognition, road sign recognition, static obstacles and other key environmental elements for recognition. For example, 6 perception tasks can correspond to 6 key task identifiers respectively. Among them, each perception task head is independently responsible for the perception and recognition of a type of task.
[0254] That is to say, in an embodiment of the present application, first, multiple cameras (cameras) with different viewing angles can be used to capture images (image information from multiple viewing angles). Taking the images captured by cameras in six directions, namely, the front view, rear view, left front view, left rear view, right front view, and right rear view of the vehicle as an example, the image information from these six captured viewing angles can be divided to obtain image data of two different target types. For image data of two different target types, two backbone neural networks can be used to extract features from the images respectively, wherein the image information of the rear view, left front view, left rear view, right front view, and right rear view corresponds to the image information of the peripheral vision type, and a neural network (backbone neural network) can be shared. At the same time, for the image data of the forward vision type, an independent feature extraction network (backbone neural network) is used to extract features. After extracting the image features from different viewing angles (image features from multiple viewing angles), subsequent feature extraction of the bird's-eye view image can be performed. Specifically, the internal and external parameters (acquisition parameters) corresponding to the camera of each acquisition perspective can be used to calculate the homography transformation matrix from each camera perspective to the BEV perspective, and then the homography matrix is used to convert the features to the BEV perspective and the BEV features of different perspectives are fused by average pooling to output the bird's-eye view image features of multiple perspectives. Finally, the spatially fused BEV features (bird's-eye view image features of multiple perspectives) and / or the image features of multiple perspectives can be sent to the environmental task perception identifier for perception and recognition of specific task targets (at least one perception task). Among them, the environmental perception system can use different key task identifiers to respectively identify different key environmental elements such as BEV lane lines, road arrows, intersections, road signs, divergence / merging points, static obstacles, etc., that is, each perception task head is independently responsible for the perception and recognition of a type of task.
[0255] Furthermore, in the embodiments of the present application, Fig.11 A schematic diagram of the multi-level adaptive loss training method proposed in the embodiment of the present application is shown in FIG. Fig.11As shown, it is assumed that the BEV multi-task perception system includes perception tasks such as lane line recognition, intersection recognition, ground arrow recognition, intersection recognition, road sign recognition, and static obstacle recognition. For each perception task, there may be different subtasks. Accordingly, the losses of different subtasks may be determined separately during the training process. For example, the subtasks of the perception task of lane line recognition correspond to curb loss, lane line loss, and other losses; the subtasks of the perception task of intersection recognition correspond to divergence point loss, merging point loss, and other losses; the subtasks of the perception task of ground arrow recognition correspond to direction loss, width loss, and other losses; the subtasks of the perception task of static obstacles correspond to cone loss, water barrier loss, and other losses; the subtasks of the perception task of intersection recognition correspond to straight arrow loss, U-turn arrow loss, and other losses; the subtasks of the perception task of road sign recognition correspond to speed bump loss, stop line loss, and other losses.
[0256] Accordingly, in the embodiment of the present application, when determining the loss L of each subtask of the perception task respectively, ts (W) After that, the adaptive term of the subtask can be combined, that is, the first loss adjustment parameter σ corresponding to the subtask ts , and the weight of the subtask w ts , determine the loss parameters corresponding to the perception task, such as Then, the adaptive term of the perceptual task can be combined, that is, the second loss adjustment parameter σ corresponding to the perceptual task t , and the weight of the perception task w t , determine the final overall loss L(W).
[0257] For example, in some embodiments, assuming that the BEV multi-task perception system includes n key tasks, that is, the number of perception tasks is n, then the loss function can be expressed as formula (5):
[0258]
[0259] Among them, L(W) can represent the overall multi-task loss, which can be composed of the loss L of each perception task t (W), weighted summation is performed, and n = 6 represents the number of perception tasks.
[0260] Considering the different complexity of different recognition tasks, relatively complex tasks are relatively difficult to optimize. For example, the lane line recognition task involves the recognition of position, category, topology information, width, direction and other information, which requires the loss function to pay more attention to it. Relatively simple tasks, such as ground arrow recognition, do not need to be assigned too much weight. Therefore, it is necessary to weight the loss of each perception task. In this way, the multi-task weighted loss can be expressed as formula (6):
[0261]
[0262] Among them, w t It can represent the weight assigned to each perception task. The goal of multi-task optimization is to ensure that each perception task can be optimized to an ideal state by setting different weights wt.
[0263] Since each perceptual task identifier corresponds to a different subtask, the optimization difficulty between different subtasks will also be different, so it is also necessary to use different weights between subtasks to balance. Therefore, the embodiment of the present application considers the weight optimization of subtasks on the basis of traditional multi-task weighted loss, forming a multi-level weighted loss, which is further expressed as formula (7):
[0264]
[0265] Among them, w ts represents the weight corresponding to the subtask s of the perception task t, t s represents the number of subtasks corresponding to the perception task t, L(W) represents the total loss, and L ts (W) represents the loss corresponding to the subtask s of the perception task t.
[0266] Multi-level adaptive weighted loss, directly using the multi-level weighted loss, the weight parameters that need to be manually adjusted are linearly related to the number of tasks and the subtasks contained in each perception task. Adjusting each weight parameter to be more reasonable by conducting multiple experiments requires a large amount of resource overhead, and it is difficult to obtain a set of ideal weight parameters. Therefore, the embodiment of the present application is optimized on the basis of the multi-level weighted loss, and a multi-level adaptive weighted loss is proposed, so that the weights of each perception task and its subtasks can be adaptively modified. Specifically, the embodiment of the present application adds a multi-level adjustment item on the basis of the weighted loss, and this adjustment item is used to guide the adjustment of the initial weight.
[0267] For example, in some embodiments, adding an adjustment term (second loss adjustment parameter) for the perception task to guide the adjustment of the initial weight can be expressed as formula (8):
[0268]
[0269] Among them, σ t Represents the second loss adjustment parameter corresponding to the perception task t.
[0270] Considering each perception task and its subtasks, the multi-level adaptive weighted loss can finally be expressed as formula (4)
[0271] It is understandable that, in the embodiment of the present application, the first loss adjustment parameter is used to adaptively adjust the weight of the subtask, and for different subtasks, the corresponding value of the first loss adjustment parameter may be different, that is, the first loss adjustment parameter is an adaptive adjustment item corresponding to the weight of the subtask. The second loss adjustment parameter is used to adaptively adjust the weight of the perception task, and for different perception tasks, the corresponding value of the second loss adjustment parameter may be different, that is, the second loss adjustment parameter is an adaptive adjustment item corresponding to the weight of the perception task.
[0272] It can be understood that, in the embodiment of the present application, the value range of the first loss adjustment parameter and the second loss adjustment parameter can be (0, 1).
[0273] It can be seen that in the embodiment of the present application, in the process of model training, several groups of weights can be set according to experience. After obtaining the reasoning results of each perception task, the single-task adaptive loss value with the adjustment item added is calculated according to the corresponding true value (generally the result of manual annotation) and the loss function of the single-task submodule, and then the adjustment item is added to each single-task adaptive loss value and weighted to obtain the multi-task adaptive loss value of the entire system. The network weights are updated using the AdamW optimization algorithm according to the final loss, and finally the optimized multi-task model is obtained, that is, the final result of this system.
[0274] Furthermore, in an embodiment of the present application, after completing model training based on the proposed multi-level adaptive weighted loss and obtaining the environmental perception model, the environmental perception model can be used to perform at least one perception task based on the collected multi-perspective image information to obtain corresponding perception results.
[0275] Exemplarily, in some embodiments, the images of 6 different perspectives (multi-perspective image information) collected by the camera are first input into the environmental perception model, and the extracted multi-perspective image features are obtained through a multi-type feature extraction network. The multi-perspective image features are input into the bird's-eye view feature extraction network, projected into the BEV feature space through the homography matrix, and the features in the BEV space are fused through the maximum pooling to obtain the BEV features, that is, the multi-perspective bird's-eye view image features. Finally, the BEV features (and / or multi-perspective image features) can be used as shared features and sent to the respective designated heads for perception and recognition of each task. Among them, the environmental perception tasks specifically include lane line recognition, intersection recognition, ground arrow recognition, road sign recognition, and intersection recognition. Among them, lane line recognition is divided into the roadside and the lane line itself, the intersection includes the starting and ending points of the divergence and merging, the U-turn point, and the virtual-real transformation point, the ground arrow includes straight, U-turn, left turn, and right turn, the road sign includes the stop line, zebra crossing, longitudinal speed bump and transverse speed bump, and the static obstacle detection includes cone barrel and water horse detection.
[0276] To sum up, in the training method of an environmental perception model proposed in an embodiment of the present application, during the training process of the environmental perception model, taking into account the differences in data distribution and optimization difficulty between different subtasks corresponding to the perception task, the loss of the corrected model parameters is determined in combination with the weights of the subtasks, that is, for different subtasks, different weights are used to balance the subtasks, thereby achieving subtask weight optimization and forming a multi-level weighted loss, so that the final total loss can ensure that each subtask can be optimized to an ideal state, thereby obtaining an environmental perception model with better prediction effect based on loss correction training.
[0277] The present application provides a training method for an environmental perception model and a method for perceiving a vehicle driving environment. On the one hand, after obtaining the predicted perception results corresponding to the first image information of multiple perspectives through the initial perception model, the total loss of the model training can be determined based on the environmental perception true value and the predicted perception results, according to the weights of the subtasks of the perception task; in this way, the weights suitable for the subtasks can be selected for weighting the loss according to the distribution differences of the subtasks of each perception task and the differences in the optimization difficulty, that is, by using different weights between the subtasks to balance, a multi-level training loss is achieved, thereby improving the prediction effect and performance of the neural network; on the other hand, the vehicle driving environment is perceived through the environmental perception model, and the perception results of at least one perception task corresponding to the image information of multiple perspectives are obtained. Among them, the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task, therefore, the environmental perception model can ensure that each perception task can achieve the optimal under ideal conditions, with better prediction effect and performance, and thus accurate perception results can be obtained when the vehicle driving environment is perceived based on the neural network.
[0278] Based on the above embodiment, in another embodiment of the present application, Fig.12 A schematic diagram of the structure of the training device for the environment perception model proposed in the embodiment of the present application is shown in FIG. Fig.12 As shown, the training device 110 of the environment perception model proposed in the embodiment of the present application may include:
[0279] The first acquisition unit 1101 is used to acquire a training data set; wherein the training data set includes multi-view first image information and an environmental perception truth value of at least one perception task corresponding to the first image information; and obtain a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through an initial perception model;
[0280] A first determining unit 1102 is used to determine the loss based on the environmental perception true value and the predicted perception result and according to the weights of the subtasks of the perception task;
[0281] The training unit 1103 is used to correct the model parameters of the initial perception model based on the loss to obtain the environment perception model.
[0282] In the embodiments of the present application, further, Fig.13 This is a schematic diagram of the structure of the vehicle driving environment sensing device proposed in the embodiment of the present application, such as Fig.13 As shown, the vehicle driving environment perception device 120 proposed in the embodiment of the present application may include:
[0283] The second acquisition unit 1201 is used to acquire multi-view image information;
[0284] The second determining unit 1202 is used to determine the image information of the target type based on the image information of multiple perspectives; wherein the target type includes at least one or more types of the forward viewing perspective type, the forward narrow viewing perspective type, and the surrounding viewing perspective type;
[0285] The second acquisition unit 1201 is also used to obtain the perception results of at least one perception task corresponding to multi-perspective image information based on the image information of the target type through the environmental perception model; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task.
[0286] In the embodiments of the present application, further, Fig.14 This is a schematic diagram of the structure of the electronic device proposed in the embodiment of the present application, such as Fig.14 As shown, the electronic device 130 proposed in the embodiment of the present application may include a processor 1301 , a memory 1302 , a communication interface 1303 , and a bus 1304 for connecting the processor 1301 , the memory 1302 and the communication interface 1303 .
[0287] In the embodiment of the present application, the above-mentioned processor 1301 can be at least one of an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a digital signal processor (Digital Signal Processor, DSP), a digital signal processing device (Digital Signal Processing Device, DSPD), a programmable logic device (ProgRAMmable Logic Device, PLD), a field programmable gate array (Field ProgRAMmable Gate Array, FPGA), a central processing unit (Central Processing Unit, CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other, and the embodiment of the present application is not specifically limited. The computer device 130 can also include a memory 1302, which can be connected to the processor 1301, wherein the memory 1302 is used to store executable program code, the program code includes computer operation instructions, and the memory 1302 may include a high-speed RAM memory, and may also include a non-volatile memory, for example, at least two disk memories.
[0288] In the embodiment of the present application, the bus 1304 is used to connect the communication interface 1303, the processor 1301 and the memory 1302, as well as the mutual communication between these devices.
[0289] In practical applications, the memory 1302 may be a volatile memory, such as a random access memory (RAM); or a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above types of memories, and provide instructions and data to the processor 1301.
[0290] Further, in an embodiment of the present application, processor 1301 is used to obtain a training data set; wherein the training data set includes multi-perspective first image information and an environmental perception true value of at least one perception task corresponding to the first image information; through an initial perception model, based on the first image information, a predicted perception result of at least one perception task corresponding to the first image information is obtained; based on the environmental perception true value and the predicted perception result, a loss is determined according to the weights of the subtasks of the perception task; and model parameters of the initial perception model are corrected based on the loss to obtain an environmental perception model.
[0291] Furthermore, in an embodiment of the present application, the processor 1301 is used to obtain multi-perspective image information, and determine the image information of the target type based on the multi-perspective image information; wherein the target type includes at least one or more of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a peripheral-view perspective type; through an environmental perception model, based on the image information of the target type, a perception result of at least one perception task corresponding to the multi-perspective image information is obtained; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; and the loss is determined based on the weight of the subtask of each perception task.
[0292] In the embodiments of the present application, further, Fig.15 This is a schematic diagram of the structure of the vehicle equipment proposed in the embodiment of the present application, such as Fig.15 As shown, the vehicle equipment 140 proposed in the embodiment of the present application may include an electronic device 130 .
[0293] The vehicle device 140 implements the method proposed in the above embodiment through the electronic device 130 .
[0294] In addition, each functional module in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or software functional modules.
[0295] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0296] An embodiment of the present application provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the method proposed in the above embodiment is implemented.
[0297] Specifically, the program instructions corresponding to the training method of an environment perception model in this embodiment can be stored on a storage medium such as a CD, a hard disk, a USB flash drive, etc. When the program instructions corresponding to the training method of an environment perception model in the storage medium are read or executed by an electronic device, the following steps are included:
[0298] Acquire a training data set; wherein the training data set includes first image information from multiple perspectives and an environmental perception truth value of at least one perception task corresponding to the first image information;
[0299] Obtaining, by means of the initial perception model, a predicted perception result of at least one perception task corresponding to the first image information based on the first image information;
[0300] Based on the true value of environmental perception and the predicted perception result, the loss is determined according to the weights of the subtasks of the perception task;
[0301] The model parameters of the initial perception model are modified based on the loss to obtain the environment perception model.
[0302] Specifically, the program instructions corresponding to the method for sensing a vehicle driving environment in this embodiment may be stored on a storage medium such as a CD, a hard disk, or a USB flash drive. When the program instructions corresponding to the method for sensing a vehicle driving environment in the storage medium are read or executed by an electronic device, the following steps are included:
[0303] Acquire multi-view image information, and determine the target type image information based on the multi-view image information; wherein the target type includes at least one or more types of a forward-looking perspective type, a forward-looking narrow-view perspective type, and a circumferential-view perspective type;
[0304] Through the environmental perception model, based on the image information of the target type, the perception result of at least one perception task corresponding to the multi-perspective image information is obtained; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; the loss is determined based on the weight of the subtask of each perception task.
[0305] The embodiment of the present application also provides a computer program product.
[0306] In some embodiments, the computer program product may include a computer program or instructions.
[0307] In some embodiments, the computer program product can be applied to the computer device in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the computer device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0308] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0309] The present application is described with reference to implementation flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0310] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which is implemented in the implementation flow diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0311] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing the steps in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0312] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. A training method for an environment perception model, characterized in that: The method comprises: Acquire a training data set; wherein the training data set includes multi-perspective first image information and an environmental perception truth value of at least one perception task corresponding to the first image information; Obtaining, by means of an initial perception model, a predicted perception result of at least one perception task corresponding to the first image information based on the first image information; Based on the environmental perception true value and the predicted perception result, and according to the weights of the subtasks of the perception task, determining the loss; The model parameters of the initial perception model are modified based on the loss to obtain the environment perception model.
2. The method according to claim 1, characterized in that The determining of the loss based on the environmental perception true value and the predicted perception result and according to the weight of the subtask of the perception task includes: Based on the environmental perception true value and the predicted perception result, determining a first loss parameter according to the weight of each of the subtasks of the first perception task; wherein the first perception task is any one of the at least one perception task; The at least one perception task is traversed, and the loss is determined based on the first loss parameter of each perception task.
3. The method according to claim 2, characterized in that The determining of the first loss parameter based on the environmental perception true value and the predicted perception result and according to the weight of each of the subtasks of the first perception task includes: Determine a first loss of the first subtask based on the environmental perception true value, the predicted perception result and the loss function of the first subtask; wherein the first subtask is any subtask of the first perception task; Determining a second loss of the first subtask based on the weight of the first subtask and the first loss; Traverse each of the subtasks of the first perception task, and determine the first loss parameter corresponding to the first perception task based on the second loss of each of the subtasks.
4. The method according to claim 1, characterized in that: The determining of the loss based on the environmental perception true value and the predicted perception result and according to the weight of the subtask of the perception task includes: Based on the environmental perception true value and the predicted perception result, a first loss parameter is determined according to the weight of each of the subtasks of the first perception task and a first loss adjustment parameter corresponding to the weight of each of the subtasks; wherein the first perception task is any one of the at least one perception task; The at least one perception task is traversed, and the loss is determined based on the first loss parameter of each perception task.
5. The method according to claim 4, characterized in that The determining of the first loss parameter based on the environmental perception true value and the predicted perception result according to the weight of each of the subtasks of the first perception task and the first loss adjustment parameter corresponding to the weight of each of the subtasks includes: Determine a first loss of the first subtask based on the environmental perception true value, the predicted perception result and the loss function of the first subtask; wherein the first subtask is any subtask of the first perception task; determining a second loss of the first subtask based on the first loss adjustment parameter corresponding to the first subtask, the weight of the first subtask, and the first loss; Traverse each of the subtasks of the first perception task, and determine the first loss parameter corresponding to the first perception task based on the second loss of each of the subtasks.
6. The method according to claim 1, characterized in that The determining of the loss based on the environmental perception true value and the predicted perception result and according to the weight of the subtask of the perception task includes: Determining a target subtask among all subtasks of the perception task based on an indicator parameter of the subtask; wherein the indicator parameter of the subtask is used to characterize the importance of the subtask; Based on the environmental perception true value and the predicted perception result, and according to the weight of the target subtask, determining a first loss parameter; The loss is determined based on the first loss parameter of a second perception task; wherein the second perception task is one or more perception tasks among the at least one perception task.
7. The method according to any one of claims 2 to 5, characterized in that: The traversing the at least one perception task and determining the loss based on the first loss parameter of each of the perception tasks includes: Traversing the at least one perception task, and determining a second loss parameter based on a weight of each of the perception tasks and the first loss parameter; Based on the second loss parameter, the loss is determined.
8. The method according to any one of claims 2 to 5, characterized in that: The traversing the at least one perception task and determining the loss based on the first loss parameter of each of the perception tasks includes: Traversing the at least one perception task, and determining a second loss parameter based on a weight of each of the perception tasks, a second loss adjustment parameter corresponding to the weight of each of the perception tasks, and the first loss parameter; Based on the second loss parameter, the loss is determined.
9. The method according to claim 8, characterized in that The determining the loss based on the second loss parameter comprises: Determining a first operand value based on the first loss adjustment parameter corresponding to each of the subtasks of each of the perception tasks; Determining a second operand value based on the second loss adjustment parameter corresponding to each of the perception tasks; The loss is determined based on the second loss parameter, the first operand value, and the second operand value.
10. The method according to any one of claims 1 to 6, characterized in that: The obtaining, based on the first image information and using the initial perception model, a predicted perception result of at least one perception task corresponding to the first image information includes: Determine multiple types of image information based on the first image information of multiple perspectives; wherein the multiple types include at least a forward-looking perspective type, a forward-looking narrow-view perspective type, and a peripheral-view perspective type; Acquire multi-perspective image features and multi-perspective bird's-eye view image features based on the multiple types of image information through the initial perception model; At least one perception task is performed based on the multi-perspective image features and / or the multi-perspective bird's-eye view image features to obtain a predicted perception result of the at least one perception task.
11. A method for sensing a vehicle driving environment, characterized in that: The method comprises: Acquire multi-view image information, and determine the image information of the target type based on the multi-view image information; wherein the target type includes at least one or more types of a forward-looking viewing angle type, a forward-looking narrow-view viewing angle type, and a surrounding viewing angle type; Through the environmental perception model, based on the image information of the target type, the perception results of at least one perception task corresponding to the multi-perspective image information are obtained; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; and the loss is determined based on the weights of the subtasks of the perception task.
12. The method according to claim 11, characterized in that The acquiring, by using the environment perception model and based on the image information of the target type, a perception result of at least one perception task corresponding to the multi-perspective image information comprises: Acquire multi-perspective image features and multi-perspective bird's-eye view image features based on the image information of the target type through the environment perception model; At least one perception task is performed based on the multi-perspective image features and / or the multi-perspective bird's-eye view image features to obtain a perception result of the at least one perception task.
13. A training device for an environmental perception model, characterized in that: The training device of the environmental perception model comprises: A first acquisition unit is configured to acquire a training data set, wherein the training data set includes multi-view first image information and an environmental perception truth value of at least one perception task corresponding to the first image information; and obtain a predicted perception result of at least one perception task corresponding to the first image information based on the first image information through an initial perception model; A first determining unit, configured to determine a loss based on the environmental perception true value and the predicted perception result and according to the weights of the subtasks of the perception task; A training unit is used to modify the model parameters of the initial perception model based on the loss to obtain the environment perception model.
14. A vehicle driving environment sensing device, characterized in that: The vehicle driving environment sensing device comprises: A second acquisition unit, used to acquire multi-view image information; A second determining unit is used to determine the image information of the target type based on the image information of the multiple perspectives; wherein the target type includes at least one or more types of a forward-looking perspective type, a forward-looking narrow-viewing perspective type, and a surrounding-viewing perspective type; The second acquisition unit is also used to obtain the perception results of at least one perception task corresponding to the multi-perspective image information based on the image information of the target type through the environmental perception model; wherein the environmental perception model is obtained by correcting the model parameters of the initial perception model based on the loss; and the loss is determined based on the weights of the subtasks of the perception task.
15. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program executable on the processor, and when the processor executes the computer program, the steps in the method according to any one of claims 1 to 9 or 10 to 11 are implemented.
16. A vehicle device, characterized in that: The vehicle equipment includes an electronic device, and the vehicle equipment implements the steps in the method of any one of claims 1-10 or 11-12 through the electronic device.
17. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps in the method described in any one of claims 1-10 or 11-12 are implemented.
18. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method of any one of claims 1-10 or 11-12 are implemented.