Beverage evaporation model training method, device, system, vehicle and readable storage medium

By combining multi-view image encoding features and loss function weighted training in the BEV model, the BEV perception model is optimized, solving the problem that the backbone network weights are limited to the ImageNet domain, and achieving efficient domain transfer and improved perception prediction performance of the BEV model.

CN116524324BActive Publication Date: 2026-02-06CHONGQING CHANGAN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310424714.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2026-02-06
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

During training, the backbone network weights of existing BEV models are limited to the initial ImageNet domain, making it difficult to fully realize the potential of multi-task perception for BEVs, resulting in poor perception and prediction performance.

Method used

By acquiring multiple 2D image encoding features from different perspectives of the target scene, processing them using a convolutional backbone network, and then inputting them into the bird's-eye view BEV perception model and a monocular auxiliary network, the BEV perception model is optimized by combining the first loss function and the second loss function for weighted training, forming a joint training to achieve efficient domain transfer.

Benefits of technology

It significantly improves the perception and prediction performance of the BEV model, enhances training efficiency and the accuracy of perception results, and avoids the impact of loss functions of irrelevant task head types on joint training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524324B_ABST
    Figure CN116524324B_ABST
Patent Text Reader

Abstract

The application discloses a BEV model training method, device and system, a vehicle and a readable storage medium, relates to the technical field of vehicles, and aims to improve the perception and prediction performance of a BEV model. The method comprises the following steps: acquiring a plurality of two-dimensional image coding features of different perspectives of a target scene; inputting the plurality of two-dimensional image coding features into a bird's-eye view (BEV) perception model to obtain a first perception result; the first perception result comprises any one of a first three-dimensional target, a BEV road and a lane line in the target scene; inputting the plurality of two-dimensional image coding features into a monocular auxiliary network to obtain a second perception result; the second perception result comprises any one of a second three-dimensional target in the target scene and a semantic segmentation result; performing weighted training on a first loss function and a second loss function to obtain a target network weight; and optimizing the BEV perception model by using the target network weight, and determining the optimized BEV perception model as a target BEV model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicles, in particular to a BEV model training method, device, system, vehicle and readable storage medium. BACKGROUND

[0002] With the development of automobile intelligence, vehicle automatic driving technology emerges as the times require, and visual perception is an important way to realize vehicle intelligent driving environment detection. With its rich information, low cost and other outstanding advantages, it has become one of the main technical means of vehicle external perception.

[0003] For example, the bird's-eye-view (BEV) perception based on the fusion of multi-camera visual angle image features around the vehicle body is an important pre-fusion perception method to realize the all-around information perception of the external environment of the vehicle. Compared with monocular visual perception, it has the advantage of cross-view multi-view feature fusion, which can effectively improve the 360° all-around environmental perception performance of the vehicle and reduce the difficulty of post-processing.

[0004] Although the training of the existing BEV model uses driving scene data sets, the image coding network is always at the front of the overall information flow, and the end-to-end back propagation mechanism has limited impact on the domain migration of the overall backbone network weight. The coding network weight is often limited to local optimization within the initial domain range of ImageNet, which is difficult to fully exploit the feature extraction potential of the backbone network for BEV multi-task perception, and also leads to poor perception and prediction performance of the existing BEV models. SUMMARY

[0005] One of the purposes of the present application is to provide a BEV model training method, device, system, vehicle and readable storage medium to improve the perception and prediction performance of the BEV model.

[0006] According to the first aspect of the present application, a BEV model training method is provided, which comprises: acquiring a plurality of two-dimensional image coding features of different visual angles of a target scene; inputting the plurality of two-dimensional image coding features into a bird's-eye-view (BEV) perception model to obtain a first perception result; the first perception result includes any one of a first three-dimensional target, a BEV road and a lane line in the target scene; inputting the plurality of two-dimensional image coding features into a monocular auxiliary network to obtain a second perception result; the second perception result includes any one of a second three-dimensional target in the target scene and a semantic segmentation result; weighting training of a first loss function and a second loss function to obtain a target network weight; the first loss function is a loss function corresponding to the BEV perception model, and the second loss function is a loss function corresponding to the monocular auxiliary network; using the target network weight to optimize the BEV perception model, and determining the optimized BEV perception model as a target BEV model.

[0007] According to the above technical means, after obtaining the multiple two-dimensional image coding features of different perspectives of the target scene, the multiple two-dimensional image coding features are input into the bird's eye view (BEV) perception model to obtain a first perception result; the multiple two-dimensional image coding features are input into the monocular auxiliary network to obtain a second perception result; the first perception result includes any one of a first three-dimensional target, a BEV road, and a lane line in the target scene, and the second perception result includes any one of a second three-dimensional target in the target scene and a semantic segmentation result. In this way, multi-dimensional perception in the target scene can be realized based on multiple models. Further, the target network weight is obtained by weighted training of the first loss function and the second loss function, and the BEV perception model is optimized using the target network weight, and the optimized BEV perception model is determined as a target BEV model. The first loss function is a loss function corresponding to the BEV perception model, and the second loss function is a loss function corresponding to the monocular auxiliary network. In this way, the monocular auxiliary network can be jointly trained with the BEV perception model to optimize the BEV perception model to obtain the target BEV model, realize efficient domain migration of the backbone coding network, and significantly improve the perception prediction performance of the BEV model without affecting the structure and inference efficiency of the BEV model.

[0008] Further, obtaining the multiple two-dimensional image coding features of different perspectives of the target scene includes: obtaining multiple two-dimensional images of different perspectives of the target scene; and processing the multiple two-dimensional images using a convolutional backbone network to obtain the multiple two-dimensional image coding features.

[0009] According to the above technical means, the multiple two-dimensional image coding features are determined by the convolutional backbone network, so that the bird's eye view (BEV) perception model and the monocular auxiliary network can determine the first perception result and the second perception result according to the two-dimensional image coding features.

[0010] Further, the target network weight is obtained by weighted training of the first loss function and the second loss function, including: determining a target loss function, the target loss function being a function obtained by weighted fusion of the first loss function and the second loss function; and performing joint network backpropagation training using multiple sets of sample data and the target loss function until a preset training end condition is reached to obtain a loss target network weight; wherein each set of sample data includes a sample image of a sample scene, a first sample perception result, and a second sample perception result.

[0011] According to the above technical means, the target network weight is obtained by performing joint network backpropagation training using multiple sets of sample data and the target loss function, so that the BEV perception model can be optimized by the target network weight to improve the perception prediction performance of the BEV model.

[0012] Further, in a case where the first perception result only includes the first three-dimensional target, the second loss function is only related to the second three-dimensional target; in a case where the first perception result only includes the BEV road and / or the lane line, the second loss function is only related to the semantic segmentation result.

[0013] According to the above technical means, by setting the task head type related to the second loss function to be the same as the perception type of the first perception result, it can be ensured that the second loss function and the first loss function can be jointly trained smoothly, and the influence of the irrelevant task head type loss function on the joint training is avoided, thereby improving the training efficiency and the accuracy of the perception result of the BEV perception model.

[0014] In a second aspect, a BEV model training apparatus is provided, which includes an acquisition unit and a processing unit. The acquisition unit is configured to acquire a plurality of two-dimensional image encoding features of different perspectives of a target scene. The processing unit is configured to input the plurality of two-dimensional image encoding features into a bird's eye view (BEV) perception model to obtain a first perception result. The first perception result includes any one of a first three-dimensional target, a BEV road, and a lane line in the target scene. The processing unit is further configured to input the plurality of two-dimensional image encoding features into a monocular auxiliary network to obtain a second perception result. The second perception result includes any one of a second three-dimensional target in the target scene and a semantic segmentation result. The processing unit is further configured to perform weighted training on a first loss function and a second loss function to obtain a target network weight. The first loss function is a loss function corresponding to the BEV perception model, and the second loss function is a loss function corresponding to the monocular auxiliary network. The processing unit is further configured to optimize the BEV perception model using the target network weight, and determine an optimized BEV perception model as a target BEV model.

[0015] Further, the acquisition unit is specifically configured to acquire a plurality of two-dimensional images of different perspectives of a target scene, and process the plurality of two-dimensional images using a convolution backbone network to obtain a plurality of two-dimensional image encoding features.

[0016] Further, the processing unit is specifically configured to determine a target loss function, the target loss function being a function obtained by weighted fusion of the first loss function and the second loss function; perform joint network backpropagation training using a plurality of sets of sample data and the target loss function until a preset training end condition is reached to obtain a loss target network weight; and each set of sample data includes a sample image of a sample scene, a first sample perception result, and a second sample perception result.

[0017] Further, in a case where the first perception result only includes the first three-dimensional target, the second loss function is only related to the second three-dimensional target; in a case where the first perception result only includes the BEV road and / or the lane line, the second loss function is only related to the semantic segmentation result.

[0018] In a third aspect, a BEV model training system is provided, which comprises a training device configured to perform the method of the first aspect or any possible design of the first aspect.

[0019] In a fourth aspect, a BEV model training device is provided, which comprises a processor, a memory for storing processor-executable instructions, and the processor is configured to execute the instructions to perform the functions performed in the first aspect or any possible design of the first aspect.

[0020] In a fifth aspect, a vehicle is provided, which comprises the BEV model training system provided in the first aspect.

[0021] In a sixth aspect, a BEV model training device is provided, which can implement the functions performed by the BEV model training device in the above aspects or any possible design of the BEV model training device, and the functions can be implemented by hardware, for example, in a possible design, the BEV model training device can comprise a processor and a communication interface, and the processor can be configured to support the BEV model training device to implement the functions involved in the above first aspect or any possible design of the first aspect.

[0022] In another possible design, the BEV model training device can further comprise a memory for storing computer-executable instructions and data necessary for the BEV model training device. When the BEV model training device is running, the processor executes the computer-executable instructions stored in the memory to make the BEV model training device perform the BEV model training method of the above first aspect or any possible design of the first aspect.

[0023] In a seventh aspect, a computer-readable storage medium is provided, which can be a readable non-volatile storage medium, and the computer-readable storage medium stores computer instructions or programs, which, when running on a computer, can make the computer perform the BEV model training method of the above first aspect or any possible design of the above aspect.

[0024] In an eighth aspect, a computer program product comprising instructions is provided, which, when running on a computer, can make the computer perform the BEV model training method of the above first aspect or any possible design of the above aspect.

[0025] Advantages of the present application:

[0026] (1) After acquiring multiple 2D image coding features from different perspectives of the target scene, the multiple 2D image coding features are input into the bird's-eye view BEV perception model to obtain the first perception result; the multiple 2D image coding features are input into the monocular auxiliary network to obtain the second perception result; the first perception result includes any one of the first 3D target, BEV road and lane lines in the target scene, and the second perception result includes any one of the second 3D target and semantic segmentation results in the target scene; thus, multi-dimensional perception of the target scene can be achieved based on multiple models. Further, the target network weights are obtained by weighted training of the first loss function and the second loss function; and the BEV perception model is optimized using the target network weights, and the optimized BEV perception model is determined as the target BEV model. The first loss function is the loss function corresponding to the BEV perception model, and the second loss function is the loss function corresponding to the monocular auxiliary network. Thus, the monocular auxiliary network can form a joint training with the BEV perception model, optimize the BEV perception model to obtain the target BEV model, realize efficient domain transfer of the backbone coding network, and significantly improve the perception and prediction performance of the BEV model without affecting the structure and inference efficiency of the BEV model.

[0027] (2) By using a convolutional backbone network to determine multiple two-dimensional image coding features, the bird's-eye view BEV perception model and monocular auxiliary network can determine the first perception result and the second perception result based on the two-dimensional image coding features.

[0028] (3) By using multiple sets of sample data and the target loss function to perform joint network backpropagation training, the target network weights are obtained. In this way, the BEV perception model can be optimized through the target network weights to improve the perception and prediction performance of the BEV model.

[0029] (4) By setting the task head type associated with the second loss function to be the same as the perception type of the first perception result, it is possible to ensure that the second loss function and the first loss function can be jointly trained smoothly, and avoid the influence of unrelated task head type loss functions on joint training, thereby improving training efficiency and the accuracy of the perception results of the BEV perception model.

[0030] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0031] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application, and do not constitute an undue limitation of this application.

[0032] Figure 1A structural schematic diagram of a BEV model training system provided by an embodiment of the present application is provided.

[0033] Figure 2 A structural schematic diagram of a BEV model training device provided by an embodiment of the present application is provided.

[0034] Figure 3 A flowchart of a BEV model training method provided by an embodiment of the present application is provided.

[0035] Figure 4 A flowchart of a BEV model training method provided by an embodiment of the present application is provided.

[0036] Figure 5 A flowchart of a BEV model training method provided by an embodiment of the present application is provided.

[0037] Figure 6 A structural schematic diagram of a BEV model training device provided by an embodiment of the present application is provided. DETAILED DESCRIPTION

[0038] In order to make the ordinary person skilled in the art better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings.

[0039] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the embodiments of the present application, as detailed in the appended claims.

[0040] It should also be understood that the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or components.

[0041] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concepts of the present application in a schematic manner, and only show the components related to the present application in the diagrams, not the number, shape and size of the components when actually implemented. The actual implementation of each component may be a random change in shape, number and proportion, and the layout pattern of the components may be more complex.

[0042] With the development of automobile intelligence, vehicle automatic driving technology emerges as the times require, and visual perception is an important way to realize vehicle intelligent driving environment detection. With its rich information, low cost and other outstanding advantages, it has become one of the main technical means of vehicle external perception.

[0043] For example, the bird's-eye-view (BEV) perception based on the fusion of multi-camera visual image features around the vehicle body is an important pre-fusion perception method to realize the all-around information perception of the external environment of the vehicle. Compared with monocular visual perception, it has the advantage of cross-view multi-view feature fusion, which can effectively improve the 360° all-around environmental perception performance of the vehicle and reduce the difficulty of post-processing.

[0044] The existing multi-camera visual BEV perception algorithm model mainly includes two typical information processing methods based on LSS architecture and Transformer architecture.

[0045] The LSS architecture is based on physical optical perspective transformation as the theoretical basis, combined with the internal and external parameters of the camera, to map the coding features of each camera visual image to the BEV space or plane, realize the BEV feature construction, and complete the result output through the decoding network connection of each task head. This method has strong physical interpretability, clear modularization, and is easy to engineer, but it is very sensitive to camera labeling accuracy.

[0046] The Transformer architecture is based on the principle of cross-attention, which calculates the correlation between each camera image pixel coding feature and BEV space voxel or plane pixel, realizes the corresponding mapping of image features to BEV features. A large number of existing literature experiments show that compared with the more traditional LSS architecture, the Transformer architecture often obtains better prediction results and has strong tolerance to camera labeling errors, but its shortcomings are also very prominent, i.e. high computational complexity, which is not conducive to engineering transformation.

[0047] In fact, whether it is LSS architecture or Transformer architecture, the front part generally needs to use a traditional two-dimensional (2D) convolution backbone network to encode the features of each camera input image to extract image features, and then input them to the subsequent feature conversion link for further processing. The training of such encoding backbone network often follows the conventional processing method, i.e. using the network weights pre-trained on large-scale classification datasets such as ImageNet as the initial weights of BEV perception for training. However, the ImageNet dataset mainly focuses on the classification task of natural scene images, which has obvious source domain differences with driving scenes centered on roads, pedestrians, vehicles, traffic signs, etc. On the other hand, such classification features also have target domain differences with the required BEV perception tasks centered on three-dimensional (3D) target detection.

[0048] Although the training of existing BEV models all uses a driving scene dataset, because the image encoding network is always at the front of the overall information flow, the end-to-end back propagation mechanism has limited impact on the domain migration of the overall backbone network weight, often causing the encoding network weight to be limited to local optimization within the initial domain range of ImageNet, making it difficult to fully exploit the feature extraction potential of the backbone network oriented to BEV multi-task perception, and also leading to poor perception and prediction performance of existing BEV models.

[0049] For example, in related technologies, multi-angle camera images can be encoded by an encoding network to extract visual features, and then the visual features are projected into the ego-vehicle coordinate system voxel features in combination with the internal and external parameter information of each camera, and are input into a 3D target detection head, a drivable area segmentation head, and a lane line segmentation head, respectively, to output the inference results of 3D target detection, drivable area BEV segmentation, and lane line BEV segmentation. This method is a typical BEV computing architecture, and the process is simple and clear, but as mentioned above, the training of the backbone encoding network weight often only uses the end loss back propagation optimization of the BEV architecture, which ultimately easily leads to limited prediction performance.

[0050] For another example, in related technologies, 3D spatial position alignment can also be performed on multi-frame image 3D coordinates, and then the 2D image features of each of the multi-frame images and the aligned 3D spatial coordinates are fused to obtain 2D image features containing 3D spatial position information, and then semantic segmentation and 3D target detection are performed in combination with the embedding features of multiple fixed position points in the BEV space and the embedding features of multiple pre-trained position points in the 3D space. Although this method strengthens the embedding and use of 3D spatial position information, the overall network structure is complex, the training and inference calculation efficiency is low, and the performance bottleneck of the aforementioned backbone feature encoding network still exists.

[0051] In view of this, the embodiments of the present application provide a BEV model training method, which applies a training device in a BEV model training system of a vehicle. The BEV model training system further comprises an external relay, an external power supply interface, an internal relay, an internal power supply interface, and an EDS. The external relay is connected to the EDS and the external power supply interface, and the internal relay is connected to the EDS and the internal power supply interface. The method comprises: determining that the internal power supply interface has a power demand; controlling the internal relay to be in an open state when the external power supply interface is in a charging state; controlling the internal relay and the external relay to be in a conductive state and controlling the EDS to supply power when the external power supply interface is in a discharging state; and controlling the internal relay to be in a conductive state and controlling the external relay to be in an open state and controlling the EDS to supply power when the external power supply interface is neither in a charging state nor in a discharging state.

[0052] The method provided by the embodiment of the present application is described in detail below with reference to the accompanying drawings of the specification.

[0053] It should be noted that the BEV model training system described in the embodiments of the present application is for more clearly illustrating the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. It is known to those skilled in the art that, with the evolution of the BEV model training system and the appearance of other BEV model training systems, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0054] The BEV model training system provided by the embodiments of the present application can be applied to a vehicle. The vehicle can be any type of vehicle. For example, the vehicle can be a fuel vehicle, a hybrid vehicle, a new energy vehicle, etc., and the embodiments of the present application do not limit the specific technology, specific quantity and specific device form used by the vehicle.

[0055] The vehicle includes an external power supply interface and an internal power supply interface. The vehicle can supply power to an external power consumption device through the external power supply interface. For example, the external power supply interface can discharge power to the power consumption device in the form of universal serial bus (USB) power supply or socket interface power supply. At the same time, the external power supply interface can also charge the vehicle in the form of inserting a charging device, for example, a direct current charging gun or an alternating current charging gun. The vehicle can also supply power to an internal power consumption device through the internal power supply interface.

[0056] Figure 1 A schematic diagram of a BEV model training system 10 provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the BEV model training system 10 can include a 2D convolution backbone network 11 (also referred to as a 2D visual feature encoding backbone network), a BEV perception model 12 (also referred to as a BEV perception model), and a monocular auxiliary network 13. Figure 1

[0057] The 2D convolution backbone network 11, the BEV perception model 12, and the monocular auxiliary network 13 are connected. For example, they can be wirelessly connected.

[0058] ​The 2D convolution backbone network 11 is configured to receive multiple two-dimensional images of different perspectives of a target scene, process the multiple two-dimensional images, obtain multiple two-dimensional image encoding features, and send the multiple two-dimensional image encoding features to the BEV perception model 12 and the monocular auxiliary network 13. The BEV perception model 12 is configured to receive the multiple two-dimensional image encoding features, process the multiple two-dimensional image encoding features, and obtain a first three-dimensional target, a BEV road, and a lane line in the target scene. The monocular auxiliary network 13 is configured to receive the multiple two-dimensional image encoding features, process the multiple two-dimensional image encoding features, and obtain a second three-dimensional target and a semantic segmentation result in the target scene.

[0059] It should be noted that, Figure 1 The example framework is only for illustration, Figure 1 The names of the various modules included in the example are not limited, and in addition to Figure 1 the functional modules shown, other modules can also be included, which are not limited by the embodiments of the present application.

[0060] In a specific implementation, Figure 2 The training device in the example can adopt Figure 2 the component structure shown, or include Figure 2 the components shown. Figure 2 A structural schematic diagram of a BEV model training device 200 provided by an embodiment of the present application is shown. The BEV model training device 200 can be a training device in a BEV model training system, or the BEV model training device 200 can be a chip or a system on a chip in the training device. As shown in Figure 2 The BEV model training device 200 includes a processor 201, a communication interface 202, and a communication line 203.

[0061] Further, the BEV model training device 200 can also include a memory 204. The processor 201, the memory 204, and the communication interface 202 can be connected through the communication line 203.

[0062] The processor 201 can be a CPU, a general processor, a network processor (NP), a digital signal processing (DSP), a microprocessor, a micro-training device, a programmable logic device (PLD), or any combination thereof. The processor 201 can also be other devices with processing functions, such as a circuit, a device, or a software module, which are not limited.

[0063] Communication interface 202 is used to communicate with other devices or other communication networks. Communication interface 202 can be a module, circuit, communication interface, or any device capable of enabling communication.

[0064] Communication line 203 is used to transmit information between the components included in the BEV model training device 200.

[0065] Memory 204 is used to store instructions executable by processor 201. These instructions may be computer programs.

[0066] The memory 204 can be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions; it can also be a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions; it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, etc., without limitation.

[0067] It should be noted that the memory 204 can exist independently of the processor 201 or can be integrated with the processor 201. The memory 204 can be used to store instructions, program code, or some data. The memory 204 can be located inside or outside the BEV model training device 200, without limitation. The processor 201 is used to execute the instructions stored in the memory 204 to implement the BEV model training method provided in the following embodiments of this application.

[0068] In one example, processor 201 may include one or more CPUs, for example, Figure 2 CPU0 and CPU1 in the CPU.

[0069] As an optional implementation, the BEV model training device 200 includes multiple processors, for example, besides Figure 2 In addition to processor 201, it may also include processor 205.

[0070] It should be pointed out that, Figure 2 The composition shown does not constitute a basis for the interpretation of this invention. Figure 1 The limitations of each device in the process, except Figure 2 In addition to the components shown,Figure 1 The various training devices in the process may include more Figure 2 More or fewer components, or combinations of certain components, or different arrangements of components.

[0071] In this embodiment of the application, the chip system may be composed of chips or may include chips and other discrete devices.

[0072] Furthermore, the actions, terms, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are merely examples, and other names may be used in specific implementations without limitation.

[0073] The following is combined Figure 2 The BEV model training system shown herein describes the BEV model training method provided in the embodiments of this application.

[0074] This application uses an example of application to a training device for illustration. Figure 3 As shown, the method includes the following steps S301-S305:

[0075] S301, The training device acquires multiple two-dimensional image coding features from different perspectives of the target scene.

[0076] The target scene can be a road scene, such as an urban road scene. Different viewpoints can be set as needed, for example, they can include the viewpoint directly in front of the vehicle, the viewpoint directly behind the vehicle, the viewpoint to the left front of the vehicle, the viewpoint to the right front of the vehicle, the viewpoint to the left of the vehicle, and the viewpoint to the right of the vehicle. When there are N different viewpoints, the multiple two-dimensional image encoding features can be Fi (i = 1, 2, ..., N).

[0077] As one possible implementation, the training device can acquire multiple two-dimensional image coding features from different perspectives of the target scene through a data storage device connected to the training device.

[0078] As another possible implementation, the training device can obtain multiple two-dimensional image coding features from different perspectives of the target scene through a convolutional backbone network connected to the training device.

[0079] S302. The training device inputs multiple two-dimensional image encoding features into the bird's-eye view BEV perception model to obtain the first perception result.

[0080] The first perception result includes any one of the following: a first three-dimensional target in the target scene, the BEV road, and lane lines. The first three-dimensional target may include cross-view three-dimensional targets and / or non-cross-view three-dimensional targets.

[0081] For example, the three-dimensional target can be a vehicle, a pedestrian, etc. The cross-view three-dimensional target includes a first-view three-dimensional target and a second-view three-dimensional target. The first-view three-dimensional target and the second-view three-dimensional target are different components of the same three-dimensional target, the first-view three-dimensional target is located in the first two-dimensional image coding feature, and the second-view three-dimensional target is located in the second two-dimensional image coding feature. The first two-dimensional image coding feature is any one of the plurality of two-dimensional image coding features, and the second two-dimensional image coding feature is different from the first two-dimensional image coding feature.

[0082] As a further possible implementation, the training apparatus can obtain the first perception result through a plurality of task heads in the BEV perception model.

[0083] The plurality of task heads in the BEV perception model can include a cross-view three-dimensional target detection head, a BEV road segmentation head, and a BEV lane line segmentation head. The cross-view three-dimensional target detection head can be used to detect a first three-dimensional target in the two-dimensional image coding feature. The BEV road segmentation head can be used to detect a BEV road in the two-dimensional image coding feature. The BEV lane line segmentation head can be used to detect a lane line in the two-dimensional image coding feature.

[0084] It should be noted that the BEV perception model Model BEV Various BEV perception algorithms can be generally included but not limited to LSS or Transformer architecture. For example, it can be BEVFormer series, BEVDet series, PETR series, BEVDepth series, etc. Specifically, it can be selected according to the demand of computing resource and time efficiency.

[0085] S303, the training apparatus inputs the plurality of two-dimensional image coding features into the monocular auxiliary network to obtain a second perception result.

[0086] The second perception result includes a second three-dimensional target in the target scene and any one of the semantic segmentation results.

[0087] As a possible implementation, the training apparatus can obtain the second perception result through a plurality of task heads in the monocular auxiliary network.

[0088] For example, the plurality of task heads in the monocular auxiliary network can include a monocular three-dimensional target detection head (Model 3D ), a monocular semantic segmentation head (Model Seg ). The monocular three-dimensional target detection head can be used to detect a second three-dimensional target in the two-dimensional image coding feature. The monocular semantic segmentation head can perform semantic segmentation on the two-dimensional image coding feature.

[0089] Model 3DVarious monocular vision 3D target detection algorithm models can be used, such as MonoRCNN, FCOS3D, DD3D, PGD, and the like. Seg Various scene semantic segmentation algorithm models can be used in general, such as FCN, U-Net, SegNet, DeepLab, and the like. Specific selection can be made according to the computing resource and time efficiency requirements.

[0090] In S304, the training device performs weighted training on the first loss function and the second loss function to obtain a target network weight.

[0091] The first loss function is a loss function corresponding to the BEV perception model, and the second loss function is a loss function corresponding to the monocular auxiliary network. For example, the first loss function can be Loss BEV , and the second loss function can be Loss MON .

[0092] As a possible implementation, after determining the first loss function and the second loss function, the training device can perform training according to the first loss function and the second loss function, adjust the network weight until a preset training end condition is reached, and take the network weight at the end of training as the target network weight.

[0093] It should be noted that, in the case where the first perception result only includes the first three-dimensional target, the second loss function is only related to the second three-dimensional target; in the case where the first perception result only includes the BEV road and / or lane line, the second loss function is only related to the semantic segmentation result.

[0094] It can be understood that the task head type related to the second loss function is the same as the perception type of the first perception result, so that the joint training of the second loss function and the first loss function can be smoothly performed, and the influence of non-related task head type loss functions on the joint training is avoided, thereby improving the training efficiency and the accuracy of the perception result of the BEV perception model.

[0095] In S305, the training device optimizes the BEV perception model using the target network weight, and determines the optimized BEV perception model as a target BEV model.

[0096] As a possible implementation, the training device can modify the network weight in the BEV perception model to the target network weight to optimize the BEV perception model, and determine the optimized BEV perception model as the target BEV model.

[0097] Based on the technical solutions provided in the present application, after obtaining the multiple two-dimensional image coding features of different perspectives of the target scene, the multiple two-dimensional image coding features are input into a bird's eye view (BEV) perception model to obtain a first perception result; the multiple two-dimensional image coding features are input into a monocular auxiliary network to obtain a second perception result; the first perception result includes any one of a first three-dimensional target in the target scene, a BEV road, and a lane line, and the second perception result includes any one of a second three-dimensional target in the target scene and a semantic segmentation result. In this way, multi-dimensional perception in the target scene can be realized based on multiple models. Further, by performing weighted training on a first loss function and a second loss function, a target network weight is obtained; and the target network weight is used to optimize the BEV perception model, and the optimized BEV perception model is determined as a target BEV model. The first loss function is a loss function corresponding to the BEV perception model, and the second loss function is a loss function corresponding to the monocular auxiliary network. In this way, the monocular auxiliary network can be jointly trained with the BEV perception model to optimize the BEV perception model to obtain the target BEV model, efficiently perform domain migration on the backbone coding network, and significantly improve the perception prediction performance of the BEV model without affecting the structure and inference efficiency of the BEV model.

[0098] In some embodiments, as shown in FIG. 1, to obtain the multiple two-dimensional image coding features, S301 in the BEV model training method of the present application can further include S401-S402. Figure 4

[0099] S401, the training device obtains multiple two-dimensional images of different perspectives of the target scene.

[0100] The multiple two-dimensional images can be 2D visual images obtained by vehicle-mounted cameras of different perspectives. The multiple two-dimensional images can be I i (i = 1, 2,..., N).

[0101] As a possible implementation, the training device can be connected with multiple perspective sensors arranged at different positions of the vehicle, and the multiple perspective sensors are used to capture the multiple two-dimensional images of different perspectives of the target scene. The multiple perspective sensors of the vehicle can send the multiple two-dimensional images of different perspectives of the target scene to the training device based on a preset frequency, and then the training device obtains the multiple two-dimensional images of different perspectives of the target scene.

[0102] ​As another possible implementation manner, after receiving the BEV model training instruction (for example, the BEV model training instruction sent by the user to the electronic device), the training device can send a request message for requesting to obtain the plurality of two-dimensional images of different perspectives of the target scene to a server storing the plurality of two-dimensional images of different perspectives of the target scene. Correspondingly, after receiving the request message, the server can obtain the plurality of two-dimensional images of different perspectives of the target scene from a database or other storage platform in response to the request message, and send the plurality of two-dimensional images of different perspectives of the target scene to the training device. Correspondingly, the training device can obtain the plurality of two-dimensional images of different perspectives of the target scene.

[0103] S402, the training device processes the plurality of two-dimensional images by using the convolution backbone network to obtain a plurality of two-dimensional image encoding features.

[0104] The type of the convolution backbone network can be set as needed. For example, it can be a series of convolutional neural network structures such as ResNet, EfficientNet, VoVNetV2, SwinTransformer, and ConvNeXt. The specific selection can be based on the demand for computing resources and time efficiency.

[0105] As a possible implementation manner, the training device can share weights for the plurality of two-dimensional images by using the 2D convolution backbone network, and respectively perform convolution processing on the plurality of two-dimensional images after sharing the weights to obtain the plurality of two-dimensional image encoding features.

[0106] It can be understood that by using the convolution backbone network to determine the plurality of two-dimensional image encoding features, the bird's eye view BEV perception model and the monocular auxiliary network can determine the first perception result and the second perception result based on the two-dimensional image encoding features.

[0107] In some embodiments, as shown in Figure 5 To obtain the target network weight, S304 in the BEV model training method of the present application can further include the following S501-S502.

[0108] S501, the training device determines a target loss function.

[0109] The target loss function is a function obtained by weighted fusion of the first loss function and the second loss function.

[0110] As a possible implementation manner, the training device can determine a weighting parameter of the first loss function and the second loss function, and perform weighted fusion on the first loss function and the second loss function based on the weighting parameter to obtain the target loss function.

[0111] For example, the target loss function can be Loss = a-Loss BEV + b-Loss MON .

[0112] wherein a > 0, b > 0, and a + b = 1.

[0113] S502, the training device performs joint network back propagation training by using multiple sets of sample data and the target loss function until a preset training end condition is reached, and obtains optimal network weights.

[0114] Each set of sample data includes a sample image of a sample scene, a first sample perception result, and a second sample perception result.

[0115] The purpose of back propagation is to continuously correct the weight parameters and bias parameters of the BEV perception model when the actual output and the perception result in the sample data differ greatly, so that the output of the BEV perception model is consistent with the perception result in the sample data.

[0116] As a possible implementation, the training device can adjust the network weights through a back propagation algorithm, and continuously iterate the above process until the loss function reaches a preset value (0.1) or the number of iterations reaches a preset number (500,000 times), to obtain optimal network weights, and use the optimal network weights as target network weights.

[0117] It can be understood that by using multiple sets of sample data and the target loss function to perform joint network back propagation training, the target network weights are obtained, so that the BEV perception model can be optimized by using the target network weights, and the BEV model perception prediction performance can be improved.

[0118] The various schemes in the above embodiments of the present application can be combined under the premise of no contradiction.

[0119] The embodiments of the present application can divide the function modules or function units of the BEV model training device or the training device according to the above method examples, for example, each function module or function unit can be divided according to each function, or two or more functions can be integrated in one processing module. The above integrated module can be realized in the form of hardware or in the form of software function module or function unit. In the embodiments of the present application, the division of the module or unit is illustrative, and is only a logical function division. In actual implementation, there can be another division method.

[0120] In the case of dividing each function module according to each function, Figure 6A structural diagram of a BEV model training apparatus 600 is shown, which can be a training apparatus or a chip applied in the training apparatus. The BEV model training apparatus 600 can be used to perform the functions of the training apparatus involved in the above embodiments. Figure 6 The BEV model training apparatus 600 shown can include an acquisition unit 601 and a processing unit 602. The acquisition unit 601 is configured to acquire a plurality of two-dimensional image encoding features of different perspectives of a target scene. The processing unit 602 is configured to input the plurality of two-dimensional image encoding features into a bird's eye view (BEV) perception model to obtain a first perception result. The first perception result includes any one of a first three-dimensional target, a BEV road, and a lane line in the target scene. The processing unit 602 is further configured to input the plurality of two-dimensional image encoding features into a monocular auxiliary network to obtain a second perception result. The second perception result includes any one of a second three-dimensional target in the target scene and a semantic segmentation result. The processing unit 602 is further configured to perform weighted training on a first loss function and a second loss function to obtain a target network weight. The first loss function is a loss function corresponding to the BEV perception model, and the second loss function is a loss function corresponding to the monocular auxiliary network. The processing unit 602 is further configured to optimize the BEV perception model by using the target network weight, and determine an optimized BEV perception model as a target BEV model.

[0121] Further, the acquisition unit 601 is specifically configured to acquire a plurality of two-dimensional images of different perspectives of the target scene, and process the plurality of two-dimensional images by using a convolution backbone network to obtain a plurality of two-dimensional image encoding features.

[0122] Further, the processing unit 602 is specifically configured to determine a target loss function, which is a function obtained by weighted fusion of the first loss function and the second loss function. The processing unit 602 is further configured to perform joint network backpropagation training by using a plurality of sets of sample data and the target loss function until a preset training end condition is reached to obtain a loss target network weight. Each set of sample data includes a sample image of a sample scene, a first sample perception result, and a second sample perception result.

[0123] Further, in a case where the first perception result only includes the first three-dimensional target, the second loss function is only related to the second three-dimensional target. In a case where the first perception result only includes the BEV road and / or the lane line, the second loss function is only related to the semantic segmentation result.

[0124] The embodiments of the present application further provide a computer-readable storage medium. All or part of the processes in the above method embodiments can be instructed by a computer program to relevant hardware to complete, and the program can be stored in the computer-readable storage medium. When the program is executed, the program can include the processes of the above method embodiments. The computer-readable storage medium can be an internal storage unit of the BEV model training apparatus or the training apparatus (including the data sending end and / or the data receiving end) of any of the preceding embodiments, such as a hard disk or a memory of the BEV model training apparatus. The computer-readable storage medium can also be an external storage device of the BEV model training apparatus, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Further, the computer-readable storage medium can include both the internal storage unit and the external storage device of the BEV model training apparatus. The computer-readable storage medium is used to store the computer program and other programs and data required by the BEV model training apparatus. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0125] The embodiments of the present application further provide a vehicle comprising the BEV model training system, the training apparatus or the BEV model training apparatus involved in the above method embodiments.

[0126] In addition, the actions, terms and the like involved between the embodiments of the present application can be mutually referred to and are not limited. The message name or parameter name in the message between the devices in the embodiments of the present application is only an example, and other names can also be used in specific implementation, which is not limited.

[0127] It should be noted that the terms "first" and "second" and the like in the specification, claims and drawings of the present application are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed or can optionally include other steps or units inherent to the process, method, product or device.

[0128] It should be understood that, in the present application, "at least one" refers to one or more, "multiple" refers to two or more, "at least two" refers to two or three and three or more, and "and / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0130] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between devices or units, which can be electrical, mechanical or other forms.

[0131] The units described as separate components can or can not be physically separated, and the components shown as units can be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0132] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0133] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium, includes a plurality of instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, and various storage program codes.

[0134] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for training a BEV model, characterized in that, The method comprises: obtaining a plurality of two-dimensional image encoding features of different perspectives of a target scene; the different perspectives include one or more of a front perspective, a rear perspective, a left front perspective, a right front perspective, a left perspective and a right perspective of a vehicle; inputting the plurality of two-dimensional image encoding features into a bird's eye view (BEV) perception model to obtain a first perception result; the first perception result includes any one of a first three-dimensional target, a BEV road and a lane line in the target scene; inputting the plurality of two-dimensional image encoding features into a monocular auxiliary network to obtain a second perception result; the second perception result includes a second three-dimensional target in the target scene and a semantic segmentation result; determining a target loss function, which is a function obtained by weighted fusion of a first loss function and a second loss function; the first loss function is a loss function corresponding to the BEV perception model, and the second loss function is a loss function corresponding to the monocular auxiliary network; performing joint network back propagation training on a plurality of sets of sample data and the target loss function until a preset training end condition is reached to obtain loss target network weights; wherein each set of sample data includes a sample image of a sample scene, a first sample perception result and a second sample perception result; optimizing the BEV perception model using the target network weights, and determining the optimized BEV perception model as a target BEV model.

2. The method of claim 1, wherein, The method comprises: obtaining a plurality of two-dimensional images of different perspectives of a target scene; processing the plurality of two-dimensional images using a convolutional backbone network to obtain a plurality of two-dimensional image encoding features.

3. The method according to claim 1 or 2, characterized in that, In a case where the first perception result only includes the first three-dimensional target, the second loss function is only related to the second three-dimensional target; in a case where the first perception result only includes the BEV road and / or the lane line, the second loss function is only related to the semantic segmentation result.

4. A BEV model training apparatus, characterized by, The device comprises an acquisition unit and a processing unit. The acquisition unit is configured to obtain a plurality of two-dimensional image encoding features of different perspectives of a target scene; the different perspectives include one or more of a front perspective, a rear perspective, a left front perspective, a right front perspective, a left perspective and a right perspective of a vehicle. The processing unit is configured to input the plurality of two-dimensional image encoding features into a bird's eye view (BEV) perception model to obtain a first perception result; the first perception result includes any one of a first three-dimensional target, a BEV road and a lane line in the target scene. The processing unit is further configured to input the plurality of two-dimensional image encoding features into a monocular auxiliary network to obtain a second perception result; the second perception result includes a second three-dimensional target in the target scene and a semantic segmentation result. The processing unit is further configured to: determine a target loss function, the target loss function being a function of a weighted fusion of a first loss function and a second loss function; the first loss function being a loss function corresponding to the BEV perception model, and the second loss function being a loss function corresponding to the monocular auxiliary network; perform joint network backpropagation training using multiple sets of sample data and the target loss function until a preset training end condition is reached, to obtain loss target network weights; wherein each set of sample data includes a sample image of a sample scene, a first sample perception result, and a second sample perception result; The processing unit is further configured to optimize the BEV perception model using the target network weights, and determine the optimized BEV perception model as a target BEV model.

5. The apparatus of claim 4, wherein, The acquisition unit is specifically configured to: acquire multiple two-dimensional images of different perspectives of a target scene; use a convolutional backbone network to process the multiple two-dimensional images to obtain multiple two-dimensional image encoding features.

6. The apparatus of claim 4 or 5, wherein, In the case where the first perception result only includes the first three-dimensional target, the second loss function is only related to the second three-dimensional target; in the case where the first perception result only includes the BEV road and / or the lane line, the second loss function is only related to the semantic segmentation result.

7. A BEV model training system, comprising: The BEV model training system includes a training device configured to perform the method of any one of claims 1 to 3.

8. A BEV model training apparatus, comprising: comprise: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method of any one of claims 1 to 3.

9. A vehicle characterized by comprising: The BEV model training system of claim 7.

10. A computer-readable storage medium, characterized in that, When the computer-executable instructions stored in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can perform the method of any one of claims 1 to 3.

Citation Information

Patent Citations

  • Monocular 3D detection frame prediction method and device

    CN115661779A