Perception model generation method, perception method, device, equipment, vehicle and medium

By generating a perception model of encoding and decoding modules for each driving scenario, the problem of insufficient perception accuracy and efficiency in different scenarios is solved, dynamic adjustment and efficient adaptation of the perception model are achieved, and driving safety is improved.

CN120220095APending Publication Date: 2025-06-27XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311801534.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to meet the diverse needs of perception models in different driving scenarios, resulting in insufficient perception accuracy and efficiency, which affects driving safety.

Method used

By generating an initial perception model, the model includes a coding module and a decoding module corresponding to each driving scenario, encode and decode according to the perception configuration information of different driving scenarios, and dynamically adjust the perception configuration information to adapt to the perception needs of different scenarios.

Benefits of technology

It realizes the perception accuracy and efficiency of adapting to different driving scenarios without increasing vehicle computing resources, and improves driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220095A_ABST
    Figure CN120220095A_ABST
Patent Text Reader

Abstract

The invention relates to a perception model generation method and device, a perception method and device, equipment, a vehicle and a medium. The perception model generation method comprises the following steps: generating an initial perception model according to perception configuration information corresponding to a plurality of driving scenes, the initial perception model at least comprising a coding module corresponding to each driving scene and a decoding module corresponding to a plurality of driving scenes, each coding module can perform coding according to the sensing configuration information corresponding to the driving scene; obtaining training sample data corresponding to each driving scene, wherein the training sample data comprises a sample scene image in the driving scene and a perception task labeling result corresponding to the sample scene image; and training the initial perception model according to the training sample data corresponding to each driving scene to obtain a target perception model. On the premise that vehicle computing resources are not increased, the requirements of different driving scenes for sensing precision can be met, the sensing accuracy and efficiency are improved, and then the driving safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of vehicles, and in particular, to a method for generating a perception model, a perception method, a device, a device, a vehicle, and a medium. Background Art

[0002] With the continuous development of technology, autonomous driving technology is increasingly applied to vehicles. Among them, the perception model plays an important role in the process of path planning and vehicle behavior control for vehicle autonomous driving. Specifically, vehicle autonomous driving needs to be completed based on the perception data of the vehicle's surrounding environment. The perception data is obtained by the perception model performing tasks such as obstacle target detection, obstacle tracking, obstacle trajectory prediction, drivable area recognition, lane line recognition, traffic sign recognition, traffic light recognition, etc.

[0003] Generally, in different driving scenarios, factors such as the directions, distances, and speeds that the autonomous driving system or the driver needs to pay attention to are different, that is, different driving scenarios have different requirements for the perception model. Exemplarily, for a high-speed driving scenario, the vehicle speed is usually relatively fast, and a large braking distance is required. Therefore, in a high-speed driving scenario, the perception model needs to have the ability to perceive front and rear at ultra-long distances. For an urban driving scenario, when the vehicle is driving at an intersection, it needs to pay attention to a large number of lateral traffic participants. Therefore, in an urban driving scenario, the perception model needs to have a strong ability to perceive lateral obstacles. For a parking driving scenario, the vehicle needs to accurately know its distance from other vehicles and its relative position to the parking space. Therefore, in a parking driving scenario, the perception model needs to have a more refined perception ability. To ensure driving safety, the perception model needs to meet the perception requirements in different driving scenarios. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a method for generating a perception model, a perception method, a device, a device, a vehicle, and a medium.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for generating a perception model, including:

[0006] Generating an initial perception model according to the perception configuration information corresponding to each of a plurality of driving scenarios, where the initial perception model includes at least an encoding module corresponding to each of the driving scenarios and a decoding module corresponding to the plurality of driving scenarios, and each encoding module can perform encoding according to the perception configuration information corresponding to the driving scenario;

[0007] Obtaining the training sample data corresponding to each of the driving scenarios, where the training sample data includes the sample scenario image in the driving scenario and the perception task annotation result corresponding to the sample scenario image;

[0008] Train the initial perception model according to the training sample data corresponding to each driving scenario to obtain a target perception model.

[0009] Optionally, the encoding module corresponding to each driving scenario is used to generate perception task features in the driving scenario according to the sample scenario image in the driving scenario and the perception configuration information corresponding to the driving scenario;

[0010] The decoding module is used to perform a decoding operation on the perception task features to obtain task query features, and the task query features are used to generate perception task prediction results.

[0011] Optionally, the decoding module is used to perform a token serialization process on the perception task features, encode each token using a positional encoding function to obtain encoded features, and decode the encoded features to obtain task query features.

[0012] Optionally, the training the initial perception model according to the training sample data corresponding to each driving scenario to obtain a target perception model includes:

[0013] For each driving scenario, input the sample scenario image corresponding to the driving scenario into the initial perception model to obtain the perception task prediction result of the driving scenario output by the initial perception model;

[0014] Optimize the initial perception model according to the perception task prediction result and the perception task annotation result of each driving scenario to obtain a target perception model.

[0015] Optionally, the inputting the sample scenario image corresponding to the driving scenario into the initial perception model to obtain the perception task prediction result of the driving scenario output by the initial perception model includes:

[0016] Extract image features from the sample scenario image corresponding to the driving scenario, and input the image features into the encoding module corresponding to the driving scenario to obtain the perception task features in the driving scenario generated by the encoding module based on the perception configuration information corresponding to the driving scenario;

[0017] Input the perception task features into the decoding module to obtain the task query features generated by the decoding module;

[0018] Generate the perception task prediction result of the driving scenario according to the task query features.

[0019] Optionally, the perception configuration information includes a perception range and / or a perception granularity.

[0020] According to a second aspect of the embodiments of the present disclosure, a perception method is provided. The method includes:

[0021] Determine a target driving scenario to be perceived;

[0022] According to the target driving scenario, determine a target encoding module corresponding to the target driving scenario in the target perception model, where the target perception model is generated according to the perception model generation method described in the first aspect of the embodiments of the present disclosure;

[0023] Input the current scenario image of the vehicle into the target encoding module and the decoding module in sequence to obtain a perception result.

[0024] Optionally, the determining the target driving scenario to be perceived includes:

[0025] Determine the target driving scenario to be perceived according to the driving scenario input by the user; or

[0026] Determine the target driving scenario to be perceived according to the current scenario image of the vehicle.

[0027] According to a third aspect of the embodiments of the present disclosure, a perception model generation device is provided. The perception model generation device includes:

[0028] A generation module configured to generate an initial perception model according to the perception configuration information corresponding to multiple driving scenarios. The initial perception model includes at least an encoding module corresponding to each driving scenario and a decoding module corresponding to the multiple driving scenarios. Each encoding module can perform encoding according to the perception configuration information corresponding to the driving scenario;

[0029] An acquisition module configured to acquire the training sample data corresponding to each driving scenario, where the training sample data includes the sample scenario image in the driving scenario and the perception task annotation result corresponding to the sample scenario image;

[0030] A training module configured to train the initial perception model according to the training sample data corresponding to each driving scenario to obtain a target perception model.

[0031] According to a fourth aspect of the embodiments of the present disclosure, a perception device is provided. The perception device includes:

[0032] A first determination module configured to determine a target driving scenario to be perceived;

[0033] A second determination module configured to determine a target encoding module corresponding to the target driving scenario in the target perception model according to the target driving scenario, where the target perception model is generated according to the perception model generation method described in the first aspect of the embodiments of the present disclosure;

[0034] An input module, configured to sequentially input current scene images of a vehicle into the target encoding module and the decoding module to obtain a perception result.

[0035] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0036] A processor;

[0037] A memory for storing processor-executable instructions;

[0038] Wherein, when the processor is configured to execute the executable instructions, the steps of the perception model generation method described in the first aspect of the embodiments of the present disclosure are implemented.

[0039] According to a sixth aspect of the embodiments of the present disclosure, there is provided a vehicle, including:

[0040] A target perception model, which is generated according to the perception model generation method described in the first aspect of the embodiments of the present disclosure;

[0041] A processor;

[0042] A memory for storing processor-executable instructions;

[0043] Wherein, when the processor is configured to execute the executable instructions, the steps of the perception method described in the second aspect of the embodiments of the present disclosure are implemented.

[0044] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the perception model generation method described in the first aspect of the embodiments of the present disclosure are implemented, or the steps of the perception method described in the second aspect of the embodiments of the present disclosure are implemented.

[0045] With the above technical solutions, since the initial perception model includes an encoding module corresponding to each driving scene and a decoding module corresponding to multiple driving scenes, the model structure of the target perception model obtained based on the initial perception model is small. Subsequently, when using the target perception model for perception, the consumption of vehicle computing resources can be reduced. In addition, since the generated target perception model includes an encoding module corresponding to each driving scene, during the perception process, the target perception model can select the corresponding encoding module according to the driving scene, and thus can dynamically adjust the perception configuration information to adapt to the requirements of different driving scenes for perception accuracy, improve the accuracy and efficiency of perception, and further improve driving safety.

[0046] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0047] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0048] Figure 1 It is a flowchart of a method for generating a perception model shown according to an exemplary embodiment.

[0049] Figure 2 It is a schematic structural diagram of a decoding module shown according to an exemplary embodiment.

[0050] Figure 3 It is a schematic diagram of a training process shown according to an exemplary embodiment.

[0051] Figure 4 It is a flowchart of a perception method shown according to an exemplary embodiment.

[0052] Figure 5 It is a block diagram of a device for generating a perception model shown according to an exemplary embodiment.

[0053] Figure 6 It is a block diagram of a perception device shown according to an exemplary embodiment.

[0054] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment.

[0055] Figure 8 It is a block diagram of a vehicle shown according to an exemplary embodiment. Detailed Embodiments

[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0057] It should be noted that all actions of obtaining signals, information, or data in this application are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.

[0058] In the related art, in order to meet the perception requirements in different driving scenarios, in one approach, a perception model corresponding to each driving scenario is generated, and in each driving scenario, the perception model corresponding to the driving scenario is used for perception. In this way, multiple perception models need to be deployed in the vehicle, occupying a large amount of vehicle computing resources. In another approach, a perception model is arranged in the vehicle. Regardless of the driving scenario, the perception model is used for perception without distinction. That is to say, the vehicle uses the same perception algorithm for perception regardless of the driving scenario. In this way, since the perception model cannot adapt to the requirements of different driving scenarios for perception accuracy, the accuracy and efficiency of perception are low, affecting driving safety.

[0059] In view of this, the present disclosure provides a perception model generation method, a perception method, a device, a device, a vehicle and a medium, so that the generated perception model can adapt to the requirements of different driving scenarios for perception accuracy without increasing the vehicle computing resources, improve the accuracy and efficiency of perception, and thus improve driving safety.

[0060] Figure 1 It is a flowchart of a perception model generation method shown according to an exemplary embodiment. As Figure 1 shown, the perception model generation method may include the following steps.

[0061] In step S11, an initial perception model is generated according to the perception configuration information corresponding to multiple driving scenarios.

[0062] Among them, the initial perception model at least includes an encoding module corresponding to each driving scenario and a decoding module corresponding to multiple driving scenarios. Each encoding module can perform encoding according to the perception configuration information corresponding to the driving scenario.

[0063] In the present disclosure, the perception configuration information corresponding to each driving scenario can be set in advance. The perception configuration information is related to the accuracy of the perception result predicted by the perception model. Since the accuracy requirements for the predicted perception result are different in different driving scenarios, the perception configuration information corresponding to different driving scenarios is different.

[0064] The perception configuration information may include a perception range and / or a perception granularity, and the perception range may include a front-back range and a left-right range. Exemplarily, the perception configuration information includes a perception range and a perception granularity. Assume that multiple driving scenarios include a highway driving scenario, an urban driving scenario, and a parking driving scenario. As described in the background art, in the highway driving scenario, the perception model needs to have the ability to perceive over a very long distance in the front-back direction. In the urban driving scenario, the perception model needs to have a strong ability to perceive lateral obstacles. In the parking driving scenario, the perception model needs to have a more refined perception ability. Therefore, the front-back range in the perception configuration information corresponding to the highway driving scenario should be greater than the front-back ranges in the perception configuration information corresponding to the urban driving scenario and the parking driving scenario. The left-right range in the perception configuration information corresponding to the urban driving scenario should be greater than the left-right ranges in the perception configuration information corresponding to the highway driving scenario and the parking driving scenario. The perception granularity in the perception configuration information corresponding to the parking driving scenario should be smaller than the perception granularities in the perception configuration information corresponding to the highway driving scenario and the urban driving scenario.

[0065] For example, the perception configuration information corresponding to the highway driving scenario may be: the left-right perception range is 40 m, the front-back range is 160 m, and the perception granularity is 0.8 m / grid. The perception configuration information corresponding to the urban driving scenario may be: the left-right perception range is 80 m, the front-back range is 80 m, and the perception granularity is 0.8 m / grid. The perception configuration information corresponding to the parking driving scenario may be: the left-right perception range is 32 m, the front-back range is 32 m, and the perception granularity is 0.4 m / grid.

[0066] It should be understood that during the process of the perception model predicting the perception result, the space is divided into grids to obtain multiple grids, and the perception result is predicted based on the feature information in the divided grids. Among them, a perception granularity of 0.8 m / grid means that the length and width of the grid obtained by dividing the space into grids are 0.8 m. A perception granularity of 0.4 m / grid means that the length and width of the grid obtained by dividing the space into grids are 0.4 m.

[0067] In the present disclosure, the number of encoding modules included in the generated initial perception model is the same as the number of driving scenarios in the multiple driving scenarios. Exemplarily, assume that the perception configuration information corresponding to three driving scenarios is set in advance, then the generated initial perception model includes three encoding modules. For each driving scenario, the perception configuration information corresponding to the driving scenario can be written into an encoding module to obtain the encoding module corresponding to the driving scenario. In this way, when predicting the perception result, the encoding module corresponding to the driving scenario can perform encoding according to the perception configuration information. For example, the space is divided into grids according to the perception configuration information to obtain multiple grids. In addition, the initial perception model includes, in addition to multiple encoding modules, a decoding module.

[0068] In step S12, training sample data corresponding to each driving scenario is obtained.

[0069] The training sample data may include a sample scenario image in the driving scenario and an annotation result of a perception task corresponding to the sample scenario image. Among them, a relatively mature annotation method can be used to annotate the sample scenario image to obtain the annotation result of the perception task.

[0070] In step S13, based on the training sample data corresponding to each driving scenario, the initial perception model is trained to obtain a target perception model.

[0071] It should be understood that since the initial perception model includes an encoding module corresponding to each driving scenario and a decoding module corresponding to multiple driving scenarios, the target perception model obtained by training the initial perception model also includes an encoding module corresponding to each driving scenario and a decoding module corresponding to multiple driving scenarios.

[0072] In this way, since the initial perception model includes an encoding module corresponding to each driving scenario and a decoding module corresponding to multiple driving scenarios, the model structure of the target perception model obtained based on the initial perception model is relatively small. Subsequently, when using the target perception model for perception, the consumption of vehicle computing resources can be reduced. In addition, since the generated target perception model includes an encoding module corresponding to each driving scenario, during the perception process, the target perception model can select the corresponding encoding module according to the driving scenario, and then can dynamically adjust the perception configuration information to adapt to the requirements of different driving scenarios for perception accuracy, improve the accuracy and efficiency of perception, and thus improve driving safety.

[0073] In the present disclosure, the generated target perception model can be applied to tasks such as bird's-eye view perception tasks, obstacle perception information, lane line perception tasks, drivable area perception tasks, etc. Therefore, the target perception model generated in the present disclosure has strong scalability.

[0074] In addition, the encoding module corresponding to each driving scenario is used to generate perception task features in the driving scenario according to the sample scenario image in the driving scenario and the perception configuration information corresponding to the driving scenario. Among them, taking the perception task as a bird's-eye view perception task as an example, the perception task feature may be a bird's-eye view feature.

[0075] The decoding module is used to perform a decoding operation on the perception task feature to obtain a task query feature, and the task query feature is used to generate a perception task prediction result.

[0076] Exemplarily, the decoding module is used to perform a token serialization process on the perception task feature, encode each token using a position encoding function to obtain an encoded feature, and decode the encoded feature to obtain a task query feature.

[0077] In the present disclosure, the initial perception model includes a plurality of encoding modules and a decoding module. Since the perception configuration information corresponding to each encoding module is different, the sizes of the perception task features generated by the encoding modules are different. In order to enable the decoding module to adapt to the different-sized perception task features output by the encoding modules corresponding to different driving scenarios, the encoding module can adopt a Transformer network to perform token serialization processing on the perception task features to obtain a plurality of tokens, and encode each token using a position encoding function to obtain encoded features. Then, the encoding module decodes the encoded features to obtain task query features.

[0078] Figure 2 It is a schematic structural diagram of a decoding module shown according to an exemplary embodiment. As Figure 2 shown, assume that the perception task features output by the encoding module corresponding to the urban driving scenario are shown as rectangle 1 in the figure. In this rectangle 1, there are a plurality of grids divided according to the perception configuration information corresponding to the urban driving scenario. The decoding module performs token serialization processing on the perception task features to obtain a plurality of tokens, and performs position encoding on each token to obtain encoded features. In Figure 2 it, the encoded features include serialized tokens and position encodings for indicating the position information corresponding to each token. The decoding module decodes according to the initialized task query feature task query (Q) and the encoded features to obtain the task query feature task query that extracts the perception task features in the urban driving scenario. Among them, the task query feature task query is used to generate a perception task prediction result.

[0079] After generating the initial perception model, the initial perception model is trained to obtain a target perception model.

[0080] In one embodiment, step S13 of training the initial perception model according to the training sample data corresponding to each driving scenario to obtain a target perception model may include: for each driving scenario, inputting the sample scenario image corresponding to the driving scenario into the initial perception model to obtain the perception task prediction result of the driving scenario output by the initial perception model; and optimizing the initial perception model according to the perception task prediction result and the perception task annotation result of each driving scenario to obtain the target perception model.

[0081] Exemplarily, image features are extracted from a sample scene image corresponding to a driving scene, and the image features are input into an encoding module corresponding to the driving scene to obtain perception task features in the driving scene generated by the encoding module based on the perception configuration information corresponding to the driving scene; the perception task features are input into a decoding module to obtain task query features generated by the decoding module; according to the task query features, a perception task prediction result for the driving scene is generated.

[0082] Figure 3 is a schematic diagram of a training process shown according to an exemplary embodiment. As Figure 3 shown, image features are extracted from a sample scene image corresponding to a high-speed driving scene, hereinafter referred to as first image features, and the first image features are input into a high-speed encoding module corresponding to the high-speed driving scene to obtain perception task features in the high-speed driving scene generated by the high-speed encoding module based on the perception configuration information corresponding to the high-speed driving scene as Figure 3 shown by rectangle 2 in. Thereafter, the perception task features in the high-speed driving scene are input into a decoding module to obtain task query features in the high-speed driving scene output by the decoding module. Finally, the task query features are passed through a fully connected network to obtain a perception task prediction result for the high-speed driving scene.

[0083] Image features are extracted from a sample scene image corresponding to an urban driving scene, hereinafter referred to as second image features, and the second image features are input into an urban encoding module to obtain perception task features in the urban driving scene generated by the urban encoding module based on the perception configuration information corresponding to the urban driving scene as Figure 3 shown by rectangle 3 in. Thereafter, the perception task features in the urban driving scene are input into a decoding module to obtain task query features in the urban driving scene output by the decoding module. Finally, the task query features are passed through a fully connected network to obtain a perception task prediction result for the urban driving scene.

[0084] Image features are extracted from a sample scene image corresponding to a parking driving scene, hereinafter referred to as third image features, and the third image features are input into a parking encoding module to obtain perception task features in the parking driving scene generated by the parking encoding module based on the perception configuration information corresponding to the parking driving scene as Figure 3 shown by rectangle 4 in. Thereafter, the perception task features in the parking driving scene are input into a decoding module to obtain task query features in the parking driving scene output by the decoding module. Finally, the task query features are passed through a fully connected network to obtain a perception task prediction result for the parking driving scene.

[0085] Among them, as Figure 3 shown, an image feature extraction module can be used to extract image features from the sample scene image.

[0086] After obtaining the perception task prediction results for each driving scenario in the above manner, optimize the initial perception model according to the perception task prediction results and perception task annotation results for each driving scenario to obtain the target perception model.

[0087] In the present disclosure, in accordance with Figure 3 the manner shown, perform forward propagation calculation on the model using the sample scenario images corresponding to multiple driving scenarios to obtain the perception task prediction results for each driving scenario. Then, perform gradient backpropagation to optimize the model. Exemplarily, for each driving scenario, calculate the prediction error according to the perception task prediction result and perception task annotation result of this driving scenario. Based on the prediction errors in each driving scenario comprehensively, optimize the initial perception model, that is, determine the perception-related parameters in the initial perception model to obtain the target perception model.

[0088] In this way, adopting the above training process, comprehensively utilize the perception task prediction results and perception task annotation results of each driving scenario to optimize the initial perception model and obtain the target perception model. In this way, the perception ability of the target perception model for all driving scenarios is enhanced.

[0089] After obtaining the target perception model in the above manner, the target perception model can be deployed on the vehicle to perform perception tasks using the target perception model during the vehicle driving process to ensure driving safety.

[0090] Based on the same inventive concept, the present disclosure also provides a perception method. Figure 4 It is a flowchart of a perception method shown according to an exemplary embodiment. As Figure 4 shown, the perception method may include the following steps.

[0091] In step S41, determine the target driving scenario to be perceived.

[0092] In one implementation, the target driving scenario to be perceived can be determined according to the driving scenario input by the user. For example, the user can input the driving scenario according to their own needs and determine this driving scenario as the target driving scenario to be perceived.

[0093] In another implementation, the target driving scenario to be perceived can be determined using a scenario state machine. Exemplarily, according to the current scenario image of the vehicle, determine the target driving scenario to be perceived. For example, the scenario state machine obtains the current scenario image of the vehicle, performs scenario recognition on this current scenario image, obtains the current driving scenario of the vehicle, and determines this driving scenario as the target driving scenario to be perceived.

[0094] In step S42, determine the target encoding module corresponding to the target driving scenario in the target perception model.

[0095] Among them, the target perception model is generated according to the perception model generation method provided by the present disclosure.

[0096] Exemplarily, referring to Figure 3 , if the target driving scenario is a highway driving scenario, the determined target encoding module is a highway encoding module; if the target driving scenario is an urban driving scenario, the determined target encoding module is an urban encoding module; if the target driving scenario is a parking driving scenario, the determined target encoding module is a parking encoding module.

[0097] In step S43, the current scene image of the vehicle is sequentially input into the target encoding module and the decoding module to obtain a perception result.

[0098] In this way, without increasing the computing resources of the vehicle, the target encoding module corresponding to the target driving scenario to be perceived can be selected, and then the perception configuration information can be dynamically adjusted to adapt to the requirements of different driving scenarios for perception accuracy, improve the accuracy and efficiency of perception, and further improve driving safety.

[0099] Based on the same inventive concept, the present disclosure also provides a perception model generation device. Figure 5 It is a block diagram of a perception model generation device shown according to an exemplary embodiment. As Figure 5 shown, the perception model generation device 500 may include:

[0100] A generation module 501, configured to generate an initial perception model according to the perception configuration information corresponding to multiple driving scenarios, where the initial perception model at least includes an encoding module corresponding to each of the driving scenarios and a decoding module corresponding to the multiple driving scenarios, and each encoding module can perform encoding according to the perception configuration information corresponding to the driving scenario;

[0101] An acquisition module 502, configured to acquire the training sample data corresponding to each of the driving scenarios, where the training sample data includes a sample scene image in the driving scenario and a perception task annotation result corresponding to the sample scene image;

[0102] A training module 503, configured to train the initial perception model according to the training sample data corresponding to each of the driving scenarios to obtain a target perception model.

[0103] Optionally, each encoding module corresponding to each of the driving scenarios is used to generate a perception task feature in the driving scenario according to the sample scene image in the driving scenario and the perception configuration information corresponding to the driving scenario;

[0104] The decoding module is used to perform a decoding operation on the perception task feature to obtain a task query feature, and the task query feature is used to generate a perception task prediction result.

[0105] Optionally, the decoding module is used to perform a token serialization process on the perception task feature, encode each token using a positional encoding function to obtain an encoded feature, and decode the encoded feature to obtain a task query feature.

[0106] Optionally, the training module 503 may include:

[0107] A first input sub-module, configured to input the sample scene image corresponding to each driving scene into the initial perception model for each driving scene, and obtain the perception task prediction result of the driving scene output by the initial perception model;

[0108] An optimization sub-module, configured to optimize the initial perception model according to the perception task prediction result and the perception task annotation result of each driving scene to obtain a target perception model.

[0109] Optionally, the first input sub-module may include:

[0110] An extraction sub-module, configured to extract an image feature from the sample scene image corresponding to the driving scene, and input the image feature into the encoding module corresponding to the driving scene to obtain the perception task feature of the driving scene generated by the encoding module based on the perception configuration information corresponding to the driving scene;

[0111] A second input sub-module, configured to input the perception task feature into the decoding module to obtain the task query feature generated by the decoding module;

[0112] A generation sub-module, configured to generate the perception task prediction result of the driving scene according to the task query feature.

[0113] Optionally, the perception configuration information includes a perception range and / or a perception granularity.

[0114] Based on the same inventive concept, the present disclosure also provides a perception device. Figure 6 It is a block diagram of a perception device shown according to an exemplary embodiment. As Figure 6 shown, the perception device 600 may include:

[0115] A first determination module 601, configured to determine a target driving scene to be perceived;

[0116] A second determination module 602, configured to determine a target encoding module corresponding to the target driving scenario in the target perception model according to the target driving scenario, where the target perception model is generated according to the perception model generation method provided by the present disclosure;

[0117] An input module 603, configured to sequentially input a current scene image of the vehicle into the target encoding module and the decoding module to obtain a perception result.

[0118] Optionally, the first determination module 601 is configured to:

[0119] Determine a target driving scenario to be perceived according to a driving scenario input by a user; or

[0120] Determine a target driving scenario to be perceived according to a current scene image of the vehicle.

[0121] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0122] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the perception model generation method provided by the present disclosure are implemented.

[0123] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the perception method provided by the present disclosure are implemented.

[0124] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0125] Refer to Figure 7 , the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output interface 712, a sensor component 714, and a communication component 716.

[0126] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-described perception model generation method. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.

[0127] The memory 704 is configured to store various types of data to support the operation of the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, and the like. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.

[0128] The power component 706 provides power to various components of the electronic device 700. The power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.

[0129] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0130] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.

[0131] The input / output interface 712 provides an interface between the processing component 702 and a peripheral interface module, and the peripheral interface module may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0132] The sensor component 714 includes one or more sensors for providing an assessment of various aspects of the state of the electronic device 700. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and a change in the temperature of the electronic device 700. The sensor component 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 may further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0133] The communication component 716 is configured to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0134] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described perception model generation method.

[0135] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, and the above instructions can be executed by a processor 720 of the electronic device 700 to complete the above-described perception model generation method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0136] Figure 8 is a block diagram of a vehicle shown according to an exemplary embodiment. For example, the vehicle 800 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 800 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.

[0137] Referring to Figure 8 , the vehicle 800 may include various subsystems. For example, the infotainment system 810, the perception system 820, the decision control system 830, the drive system 840, and the computing platform 850. Among them, the vehicle 800 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 800 may be interconnected by wired or wireless means. In addition, a target perception model is deployed on the vehicle, and the target perception model is generated according to the perception model generation method provided by the present disclosure.

[0138] In some embodiments, the infotainment system 810 may include a communication system, an entertainment system, and a navigation system, etc.

[0139] The perception system 820 may include several sensors for sensing information about the environment around the vehicle 800. For example, the perception system 820 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), a lidar, a millimeter wave radar, an ultrasonic radar, and a camera device.

[0140] The decision control system 830 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.

[0141] The drive system 840 may include components that provide motive power for the vehicle 800. In one embodiment, the drive system 840 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.

[0142] Some or all functions of the vehicle 800 are controlled by the computing platform 850. The computing platform 850 may include at least one processor 851 and a memory 852, and the processor 851 may execute instructions 853 stored in the memory 852.

[0143] The processor 851 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0144] The memory 852 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0145] In addition to the instructions 853, the memory 852 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 852 can be used by the computing platform 850.

[0146] In an embodiment of the present disclosure, the processor 851 may execute the instructions 853 to complete all or part of the steps of the above-described perception method.

[0147] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-described perception model generation method or perception method when executed by the programmable device.

[0148] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0149] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A method for generating a perception model, characterized in that, Including: Generate an initial perception model according to the perception configuration information corresponding to each of multiple driving scenarios. The initial perception model at least includes an encoding module corresponding to each of the driving scenarios and a decoding module corresponding to the multiple driving scenarios. Each encoding module can perform encoding according to the perception configuration information corresponding to the driving scenario; Obtain the training sample data corresponding to each of the driving scenarios. The training sample data includes the sample scenario image in the driving scenario and the perception task annotation result corresponding to the sample scenario image; Train the initial perception model according to the training sample data corresponding to each of the driving scenarios to obtain a target perception model.

2. The method for generating a perception model according to claim 1, wherein: Each encoding module corresponding to each of the driving scenarios is used to generate a perception task feature in the driving scenario according to the sample scenario image in the driving scenario and the perception configuration information corresponding to the driving scenario; The decoding module is used to perform a decoding operation on the perception task feature to obtain a task query feature, and the task query feature is used to generate a perception task prediction result.

3. The method for generating a perception model according to claim 2, wherein: The decoding module is used to perform a token serialization process on the perception task feature, encode each token using a position encoding function to obtain an encoded feature, and decode the encoded feature to obtain a task query feature.

4. The method for generating a perception model according to claim 1, wherein, The training the initial perception model according to the training sample data corresponding to each of the driving scenarios to obtain a target perception model includes: For each of the driving scenarios, input the sample scenario image corresponding to the driving scenario into the initial perception model to obtain the perception task prediction result of the driving scenario output by the initial perception model; Optimize the initial perception model according to the perception task prediction result and the perception task annotation result of each of the driving scenarios to obtain a target perception model.

5. The method for generating a perception model according to claim 4, wherein The inputting the sample scenario image corresponding to the driving scenario into the initial perception model to obtain the perception task prediction result of the driving scenario output by the initial perception model includes: Extract image features from the sample scenario image corresponding to the driving scenario, and input the image features into the encoding module corresponding to the driving scenario to obtain the perception task feature in the driving scenario generated by the encoding module based on the perception configuration information corresponding to the driving scenario; Input the perception task feature into the decoding module to obtain the task query feature generated by the decoding module; Generate the perception task prediction result of the driving scenario according to the task query feature.

6. The method for generating a perception model according to any one of claims 1-5, characterized in that, The perception configuration information includes a perception range and / or a perception granularity.

7. A perception method, characterized in that, The method includes: Determine a target driving scenario to be perceived; Determine a target encoding module corresponding to the target driving scenario in the target perception model according to the target driving scenario. The target perception model is generated according to the method for generating a perception model according to any one of claims 1-6. Input the current scene image of the vehicle into the target encoding module and decoding module in sequence to obtain a perception result.

8. The sensing method according to claim 7, wherein The determination of the target driving scene to be perceived includes: Determine the target driving scene to be perceived according to the driving scene input by the user; or Determine the target driving scene to be perceived according to the current scene image of the vehicle.

9. A perception model generation device, characterized in that, The perception model generation device includes: A generation module configured to generate an initial perception model according to the perception configuration information corresponding to multiple driving scenes. The initial perception model includes at least an encoding module corresponding to each driving scene and a decoding module corresponding to the multiple driving scenes. Each encoding module can perform encoding according to the perception configuration information corresponding to the driving scene; An acquisition module configured to acquire the training sample data corresponding to each driving scene. The training sample data includes the sample scene image in the driving scene and the perception task annotation result corresponding to the sample scene image; A training module configured to train the initial perception model according to the training sample data corresponding to each driving scene to obtain a target perception model.

10. A sensing device, characterized in that, The perception device includes: A first determination module configured to determine the target driving scene to be perceived; A second determination module configured to determine the target encoding module corresponding to the target driving scene in the target perception model according to the target driving scene. The target perception model is generated according to the perception model generation method described in any one of claims 1-6; An input module configured to input the current scene image of the vehicle into the target encoding module and decoding module in sequence to obtain a perception result.

11. An electronic device, characterized in that, Includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the steps of the perception model generation method described in any one of claims 1-6 when executing the executable instructions.

12. A vehicle, characterized in that, Includes: A target perception model, the target perception model is generated according to the perception model generation method described in any one of claims 1-6; A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the steps of the perception method described in claim 7 or 8 when executing the executable instructions.

13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the perception model generation method described in any one of claims 1-6 are implemented, or the steps of the perception method described in claim 7 or 8 are implemented.