Method for determining the detection result state of an object, model training method, and apparatus
The method improves the identification of perception anomalies in autonomous driving by using a pre-trained model based on object and map information, enhancing generalization and efficiency in anomaly detection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING HORIZON ROBOTICS TECH RES & DEV CO LTD
- Filing Date
- 2024-03-15
- Publication Date
- 2026-06-02
AI Technical Summary
Existing perception systems in autonomous driving and driving assistance face limitations in identifying objects with perception anomalies, leading to low development efficiency, poor generalization ability, and limited coverage of abnormal situations, which affect trajectory prediction and vehicle control.
A method and apparatus for determining the perception result state of an object using a pre-trained sensing result state detection model based on the target object's trajectory information, surrounding objects' trajectory information, and map element information, improving generalization ability and efficiency.
Enhances the ability to identify objects with perception anomalies, covers a wider range of anomaly situations, and improves processing efficiency compared to artificially designed rules.
Smart Images

Figure 2026517577000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to artificial intelligence technology, and particularly to a method for determining the perception result state of an object, a model training method, and an apparatus.
Background Art
[0002] In fields such as autonomous driving and driving assistance, the stability of perception has a great impact on the effect of trajectory prediction and affects the planning and control of vehicles. Currently, due to the limitations of perception capabilities, it is still impossible to completely avoid the occurrence of perception anomalies. Therefore, in subsequent trajectory prediction, it is usually necessary to identify objects (or obstacles) with perception anomalies to prevent or reduce interference with trajectory prediction, vehicle planning, and control, and avoid giving users a bad experience. Identifying objects with perception anomalies based on artificially designed identification rules is likely to cause problems such as low development efficiency, poor generalization ability, and limited coverage of perception anomaly situations.
Summary of the Invention
Problems to be Solved by the Invention
[0003] Embodiments of the present disclosure provide a method for determining the perception result state of an object, a model training method, and an apparatus, which are helpful to improve the generalization ability of identification for objects with perception anomalies and improve work efficiency and processing efficiency.
Means for Solving the Problems
[0004] According to a first embodiment of the embodiments of the present disclosure, a method for determining the sensing result state of an object is provided, comprising the steps of: acquiring sensing results and map element information corresponding to each of at least one time frame in which the sensing results include state information in a first coordinate system of at least one object; determining first trajectory information corresponding to a target object among the objects and second trajectory information of other objects around the target object based on the sensing results corresponding to each of the time frames; determining a detection result corresponding to the target object using a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information and the map element information; and determining the sensing result state of the target object based on the detection result.
[0005] A second aspect of the embodiments of the present disclosure provides a method for training a sensing result state detection model, comprising the steps of: acquiring first training trajectory information corresponding to each of the training target objects in at least one training target object, second training trajectory information for other surrounding training targets corresponding to each of the training target objects, and map element information corresponding to each of the training target objects; and training a pre-constructed sensing result state detection network based on each of the first training trajectory information, each of the second training trajectory information, and each of the map element information to obtain a trained sensing result state detection model.
[0006] According to a third aspect of the embodiments of the present disclosure, an object sensing result state determination device is provided, comprising: a first acquisition module used to acquire sensing results and map element information corresponding to each of the time frames in at least one time frame, wherein the sensing results include state information in a first coordinate system of at least one object; a first processing module used to determine first trajectory information corresponding to a target object among the objects and second trajectory information of other objects around the target object based on the sensing results corresponding to each of the time frames; a second processing module used to determine a detection result corresponding to the target object using a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information and the map element information; and a third processing module used to determine the sensing result state of the target object based on the detection result.
[0007] A training device for a sensing result state detection model is provided, which includes: a second acquisition module used to acquire first training trajectory information corresponding to each of the training target objects in at least one training target object, second training trajectory information of other surrounding training objects corresponding to each of the training target objects, and map element information corresponding to each of the training target objects; and a fourth processing module used to train a pre-constructed sensing result state detection network based on each of the first training trajectory information, each of the second training trajectory information, and each of the map element information, in order to obtain a trained sensing result state detection model.
[0008] A fifth embodiment of the embodiments of the present disclosure is provided, which stores a computer-readable storage medium that stores a computer program for performing a method for determining the sensing result state of an object as described in any of the above embodiments of the present disclosure, or a computer program for performing a method for training a sensing result state detection model as described in any of the above embodiments of the present disclosure.
[0009] According to a sixth embodiment of the embodiments of the present disclosure, an electronic device is provided, comprising a processor and a memory for storing instructions that the processor can execute, wherein the processor reads and executes the executable instructions from the memory to implement a method for determining the sensing result state of an object as described in any of the embodiments of the present disclosure, or a method for training a sensing result state detection model as described in any of the embodiments of the present disclosure.
[0010] An embodiment of a seventh aspect of the present disclosure provides a computer program product in which, when executed by a processor, instructions in the computer program product either perform a method for determining the sensing result state of an object as described in any of the above embodiments of the present disclosure, or a method for training a sensing result state detection model as described in any of the above embodiments of the present disclosure. [Effects of the Invention]
[0011] According to the object sensing result state determination method, model training method, and apparatus provided in the above embodiments of this disclosure, the sensing result state of a target object can be determined by utilizing a pre-trained sensing result state detection model based on the target object's own trajectory information, the trajectory information of other objects in its vicinity, and the surrounding map element information. This improves the generalization ability for identifying objects with sensing anomalies, helps to cover a wider range of sensing anomaly situations, and improves work efficiency and processing efficiency compared to artificially designing rules.
[0012] The technical proposal of this disclosure will be described in more detail below with reference to the drawings and examples. [Brief explanation of the drawing]
[0013] [Figure 1] This is one exemplary application scenario of the method for determining the sensing result state of an object provided in this disclosure. [Figure 2] This is a flowchart of a method for determining the sensing result state of an object, as provided by one exemplary embodiment of the present disclosure. [Figure 3] This is a flowchart of a method for determining the sensing result state of an object, as provided in another exemplary embodiment of the present disclosure. [Figure 4] This is a schematic diagram of the target coordinate system provided by one exemplary embodiment of the present disclosure. [Figure 5] This is a schematic diagram of the trajectory of the first trajectory information of a target object provided in one exemplary embodiment of the present disclosure. [Figure 6] This is a flowchart of a method for determining the sensing result state of an object, as provided in one further exemplary embodiment of the present disclosure. [Figure 7] This is a schematic diagram of an attention network provided by one exemplary embodiment of the present disclosure. [Figure 8] This is a flowchart of a method for determining the sensing result state of an object, as provided by another exemplary embodiment of the present disclosure. [Figure 9] This is a flowchart of a training method for a sensing result state detection model provided in one exemplary embodiment of the present disclosure. [Figure 10] This is a flowchart of a training method for a sensing result state detection model provided in another exemplary embodiment of the present disclosure. [Figure 11] This is a schematic diagram of a multitasking network provided by one exemplary embodiment of the present disclosure. [Figure 12] This is a flowchart of a training method for a sensing result state detection model provided in one further exemplary embodiment of the present disclosure. [Figure 13]It is a schematic structural diagram of a determination device for the sensed result state of an object provided by one exemplary embodiment of the present disclosure. [Figure 14] It is a schematic structural diagram of a determination device for the sensed result state of an object provided by another exemplary embodiment of the present disclosure. [Figure 15] It is a schematic structural diagram of a training device for a sensed result state detection model provided by one exemplary embodiment of the present disclosure. [Figure 16] It is a schematic structural diagram of one application embodiment of an electronic device of the present disclosure.
Embodiments for Carrying Out the Invention
[0014] Hereinafter, in order to explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail while referring to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all of the embodiments, and the present disclosure is not limited to the exemplary embodiments.
[0015] Unless otherwise specifically described, the relative arrangements, mathematical formulas and numerical values of the members and steps described in these embodiments do not limit the scope of the present disclosure.
[0016] [Summary of the Present Disclosure] In the process of realizing the present disclosure, the inventors have discovered that in the fields of autonomous driving, driving assistance, etc., the stability of perception has a great impact on the effect of trajectory prediction and affects the planning and control of vehicles. At present, due to the limitations of perception capabilities, it is still impossible to completely avoid the occurrence of abnormal perception situations. Therefore, in subsequent trajectory predictions, it is usually necessary to identify abnormally perceived objects (or obstacles), thereby preventing or reducing interference with trajectory prediction, vehicle planning, and control, and avoiding giving users a bad experience. When identifying an abnormally perceived object based on an artificially designed identification rule and obtaining information about the abnormally perceived object, the identification rule can be related to many things such as whether the yaw angle of the object jumps, whether the type of the object is stable, whether the object drifts horizontally, etc. Therefore, designing the rule artificially has low development efficiency, poor generalization ability, and is likely to lead to limited coverage of abnormal perception situations.
[0017] [Exemplary Overview] FIG. 1 is an exemplary application scenario of a method for determining the perception result state of an object provided by the present disclosure.
[0018] In application scenarios such as autonomous driving and driver assistance, the object sensing result state determination method disclosed herein can be used to train a pre-constructed sensing result state detection network based on first training trajectory information corresponding to each training target object in at least one training target object, second training trajectory information of other surrounding training targets, and map element information of the area where the training target object is located, thereby obtaining a trained sensing result state detection model. The sensing result state detection model can be deployed to an object sensing result state determination device on a vehicle. During the vehicle's driving process, sensing results and map element information corresponding to each time frame in at least one time frame can be acquired. The sensing results may include state information of the sensing object in a first coordinate system. The first coordinate system may be a world coordinate system or a vehicle coordinate system of the corresponding time frame, and is not specifically limited. The world coordinate system and the vehicle coordinate system can be converted to each other depending on the vehicle's position and orientation. Based on the sensing results corresponding to each time frame, first trajectory information corresponding to the target object within each object, and second trajectory information for other objects surrounding the target object can be determined. Based on the first trajectory information, second trajectory information, and map element information, the detection result of the target object can be determined using a pre-trained sensing result state detection model, and furthermore, the sensing result state of the target object can be determined based on the detection result. The sensing result state can include normal and abnormal states. The first trajectory information can include at least one piece of information such as the position, orientation, velocity, and acceleration of the target object in each time frame. Similarly, the second trajectory information can include at least one piece of information such as the position, orientation (direction or angle), velocity, and acceleration of other objects in each time frame. Specifically, this can be set according to actual needs. Map element information can include static element information such as lane markings and road curbs.
[0019] The embodiments of this disclosure can determine the accuracy of the perceived state of an object based on a perceived state detection model, improve the generalization ability of identifying objects with perceived anomalies, help cover a wider range of perceived anomaly situations, and improve work efficiency and processing efficiency compared to artificially designing rules.
[0020] [Example Method] Figure 2 is a flowchart of a method for determining the sensing result state of an object provided by one exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, such as automotive computing platforms, and as shown in Figure 2, the method of the embodiment of the present disclosure may include the following steps.
[0021] In step 201, sensing results and map element information corresponding to each time frame within at least one time frame are acquired, and the sensing results include state information of at least one object in a first coordinate system.
[0022] Here, at least one time frame can be set according to actual needs. For example, in the process of a vehicle's journey, at least one time frame may include the current time frame and at least one historical time frame. The first coordinate system may be a world coordinate system, a map coordinate system, or the vehicle coordinate system of the corresponding time frame. Specifically, it can be set according to actual needs. The object may include dynamic objects such as other vehicles and pedestrians around the vehicle obtained through sensing, and may also include the vehicle itself. Map element information may include static element information such as lane markings and road curbs. Map element information may be a set of map element information corresponding to each time frame, or a set of map element information corresponding to all time frames within at least one time frame. Specifically, it can be set according to actual needs. The object state information may include at least one of the following pieces of information obtained through sensing: position, orientation, velocity, angular velocity, acceleration, type, etc. The type may refer to the classification of the object obtained through sensing. For example, the object may be a vehicle, pedestrian, or other obstacle.
[0023] In one selective embodiment, the sensing result can be obtained based on data collected by at least one sensor among sensors such as a camera, laser radar, millimeter-wave radar, and ultrasonic radar, and a corresponding sensing algorithm or sensing model. For example, images collected by a camera on a vehicle can be processed based on a pre-trained target detection model to obtain the type and bounding box of each object in the image, point clouds corresponding to each object can be obtained based on the laser radar, and the radar point clouds and image detection results can be integrated to determine the state of each object in a first coordinate system. The specific method for acquiring the sensing result can be set according to the actual needs.
[0024] In one selective embodiment, map element information can be obtained by extracting it from a map of a corresponding region, which can be determined based on the vehicle's position or orientation corresponding to each time frame. The map describes relevant information for various elements, such as lane markings and road curbs. Based on the vehicle's position or orientation on the map corresponding to each time frame, a preset range around the vehicle can be determined, and information on relevant elements within that preset range on the map can be used as map element information.
[0025] In one optional example, step 201 may be performed by the processor calling a corresponding instruction stored in memory, or by a first acquisition module executed by the processor.
[0026] In step 202, based on the sensing results corresponding to each time frame, the first trajectory information corresponding to the target object within each object, and the second trajectory information of other objects around the target object are determined.
[0027] Here, the target object can be determined based on the object in the sensing result corresponding to the current time frame, or it can be determined by integrating sensing results from multiple time frames. For example, the target object may be any one of the objects in the sensing result of the current time frame, and the number of target objects can be set according to the actual needs. For example, each object in the sensing result corresponding to the current time frame can be a target object, and specifically, it can be set according to the actual needs. The first trajectory information may include the position and orientation of the target object in each time frame, and may also include velocity, acceleration, etc., and specifically, it can be set according to the actual needs.
[0028] In one selective embodiment, objects in the sensing results corresponding to each time frame are identified based on preset identification rules, the sensing state of a first object that satisfies the preset identification rules is determined to be an abnormal state, the first object is filtered from each sensing result to obtain filtered sensing results, and further, based on the filtered sensing results corresponding to each time frame, first trajectory information corresponding to the target object among the objects, and second trajectory information of other objects around the target object can be determined. The preset identification rules can be set according to actual needs and may include rules such as whether the yaw angle of an object jumps, whether the type of object is stable, and whether the object drifts laterally. Objects that are clearly abnormal are identified first based on the preset identification rules, and then subsequent processing is performed on the remaining objects to identify their sensing state.
[0029] In one selective embodiment, the number of other objects surrounding the target object may be zero, one, or more, and is not specifically limited. In practical applications, the input format for the second trajectory information of other objects can be set so that it can be adapted to a sensing result state detection model. For example, the input format can be set to the second trajectory information of a first number (e.g., 16) of other objects. If the number of other objects surrounding the target object does not meet the first number, or if there are no other objects around the target object, empty values (e.g., 0) can be filled to represent the portion of the second trajectory information of other objects that do not actually exist. For example, if there are actually 5 other objects around the target object, in an input format representing the second trajectory information of 16 other objects, the second trajectory information of 5 other objects has valid data, and the second trajectory information of the remaining 11 other objects is filled with empty values, thereby ensuring that the second trajectory information of other objects satisfies the input requirements of the model.
[0030] In one optional example, step 202 may be performed by the processor calling a corresponding instruction stored in memory, or by a first processing module executed by the processor.
[0031] In step 203, based on the first trajectory information, the second trajectory information, and map element information, the detection result corresponding to the target object is determined using a pre-trained sensing result state detection model.
[0032] Here, the specific structure of the sensing result state detection model can be set according to the actual needs, for example, a detection model based on a convolutional neural network and an attention network can be employed, and this embodiment is not limited thereto. The detection result corresponding to the target object can include the sensing result state probability of the target object predicted and obtained by the model, and the sensing result state probability can include at least one of a first probability that the sensing result state is a first state and a second probability that the sensing result state is a second state. The first state may be a normal state, the second state may be an abnormal state, or the first state may be an abnormal state and the second state may be a normal state, and specifically, this can be set according to the actual needs. The representation forms of the first and second states can also be set according to the actual needs, for example, the first state may be represented by 1 and the second state by 0.
[0033] In one optional example, step 203 may be performed by the processor calling a corresponding instruction stored in memory, or by a second processing module executed by the processor.
[0034] In step 204, the detection status of the target object is determined based on the detection results.
[0035] In one selective embodiment, the perceived state of a target object can be determined based on the perceived state probability included in the detection result and the corresponding probability threshold. The probability threshold can be set according to the actual needs.
[0036] In one optional example, step 204 may be performed by the processor calling a corresponding instruction stored in memory, or by a third processing module executed by the processor.
[0037] The method for determining the sensing state of an object provided in this embodiment can determine the sensing state of a target object by utilizing a pre-trained sensing state detection model based on the target object's own trajectory information, the trajectory information of other objects in its vicinity, and the surrounding map element information. This improves the generalization ability for identifying objects with sensing anomalies, helps cover a wider range of sensing anomaly situations, and improves work efficiency and processing efficiency compared to artificially designing rules.
[0038] Figure 3 is a flowchart of a method for determining the sensing result state of an object, as provided in another exemplary embodiment of the present disclosure.
[0039] In one selective embodiment, at least one time frame may include a current time frame and at least one historical time frame, and the state information may include the position and orientation of the object.
[0040] Here, position can include the spatial coordinates of the object in the first coordinate system, and orientation can include the attitude or angle of the object in the first coordinate system, for example, the forward direction of a vehicle.
[0041] In one selective embodiment, determining the first trajectory information corresponding to the target object among the objects, based on the sensing results corresponding to each time frame in step 202, may include the following steps:
[0042] Step 2021 determines a target coordinate system with the target object as the origin, based on the target object's current position and orientation in the current time frame.
[0043] Here, the target coordinate system may be a coordinate system in which the position of the target object in the current time frame is the origin, the orientation of the target object in the current time frame is the first coordinate axis, and the direction perpendicular to the orientation of the target object is the second coordinate axis. For example, if the object is a vehicle, the target coordinate system may be the vehicle coordinate system of that vehicle, with the center of the rear axle of the vehicle as the origin, the length direction of the vehicle as the first coordinate axis, and the width direction of the vehicle as the second coordinate axis.
[0044] Exemplary, Figure 4 is a schematic diagram of a target coordinate system provided by one exemplary embodiment of the present disclosure, where x0y represents the target coordinate system of the target object.
[0045] In one optional example, step 2021 may be performed by the processor calling a corresponding instruction stored in memory, or by a first decision unit executed by the processor.
[0046] In step 2022, for any one historical time frame, the state information of the target object in that historical time frame is transformed into the target coordinate system based on the transformation relationship between the first coordinate system and the target coordinate system, and the first state information in the target coordinate system corresponding to that historical time frame is obtained.
[0047] Here, if the first coordinate system is the world coordinate system or the map coordinate system, the transformation relationship between the first coordinate system and the target coordinate system can be determined based on the rotation and translation of the target coordinate system relative to the first coordinate system. If the first coordinate system is the vehicle coordinate system of the corresponding time frame, the transformation relationship between the first coordinate system and the target coordinate system can be determined based on the transformation relationship between the vehicle coordinate system and the world coordinate system, and the transformation relationship between the world coordinate system and the target coordinate system. Based on the transformation relationship, the state information of any one of the historical time frames of the target object can be transformed into the target coordinate system, and the first state information in the target coordinate system corresponding to that historical time frame can be obtained. The specific transformation principle will not be explained. The first state information may include the position and orientation in the target coordinate system. Based on this, the state information in the target coordinate system corresponding to each time frame of the target object can be obtained.
[0048] In one optional example, step 2022 may be performed by the processor calling a corresponding instruction stored in memory, or by a first translation unit executed by the processor.
[0049] In step 2023, the first trajectory information corresponding to the target object is determined based on the first state information corresponding to each historical time frame and the origin state information corresponding to the current time frame.
[0050] Here, the origin state information may be the first state information of the target object in the target coordinate system. The current time frame and each time frame within each historical time frame can correspond to one state of the target object, and based on this, the first trajectory information corresponding to the target object can be determined.
[0051] Exemplary, Figure 5 is a schematic diagram of the trajectory of first trajectory information of a target object provided by one exemplary embodiment of the present disclosure. In this example, at least one time frame may include the current time frame and five historical time frames to determine the first trajectory information of the target object in the target coordinate system.
[0052] In one optional example, step 2023 may be performed by the processor calling a corresponding instruction stored in memory, or by a second decision unit executed by the processor.
[0053] This embodiment converts the state information of the target object in the first coordinate system for each time frame of the target object into a target coordinate system with the target object as the origin. This makes it easier to determine the relative relationship between the state of the target object in its historical time frame and the state of the current time frame, as well as the relative relationship between the target object and other objects and map elements in the current time frame, thereby providing more effective data for determining the sensing result state of the target object corresponding to the current time frame.
[0054] In one selective embodiment, determining second trajectory information of other objects around the target object based on the sensing results corresponding to each time frame in step 202 may include the following steps:
[0055] In step 2024, for any one time frame, the state information of other objects in that time frame is transformed into the target coordinate system based on the transformation relationship between the first coordinate system and the target coordinate system, and the second state information of the other objects in that time frame is obtained.
[0056] Here, the transformation of the state information of other objects at each time frame into the target coordinate system is similar to that of the target object mentioned above, and therefore will not be explained here.
[0057] In one optional example, step 2024 may be performed by the processor calling a corresponding instruction stored in memory, or by a second translation unit executed by the processor.
[0058] In step 2025, for any one other object, the second trajectory information corresponding to that other object is determined based on the second state information of that other object in each time frame.
[0059] The determination of the second trajectory information is similar to that of the first trajectory information described above, so we will omit the explanation here.
[0060] In one selective embodiment, if at least one other object exists around the target object, a second trajectory information corresponding to each of those other objects can be determined for each other object using the method described above.
[0061] In one selective embodiment, if no other objects exist around the target object, it is not necessary to perform steps 2024 to 2025, and it can be directly determined that the second trajectory information of each other object input into the model is a filled empty value.
[0062] In one optional example, step 2025 may be performed by the processor calling a corresponding instruction stored in memory, or by a third decision unit executed by the processor.
[0063] This embodiment helps to further improve the accuracy of prediction results by assisting in model prediction of the perceived state of the target object based on the relative position and relative orientation between the other object and the target object, by converting the state information of other objects in each time frame to the target coordinate system of the target object.
[0064] Figure 6 is a flowchart of a method for determining the sensing result state of an object, provided by one further exemplary embodiment of the present disclosure.
[0065] In one selective embodiment, the method of the embodiment of the present disclosure may further include step 2026 of converting map element information to a target coordinate system based on a transformation relationship between the map coordinate system and the target coordinate system corresponding to the map element information, thereby obtaining target map element information.
[0066] Here, the principle for determining the transformation relationship between the map coordinate system and the target coordinate system is similar to that of the first coordinate system described above, and therefore will not be explained here.
[0067] In one optional example, step 2026 may be performed by the processor calling a corresponding instruction stored in memory, or by a third translation unit executed by the processor.
[0068] Step 203, which determines a detection result corresponding to a target object using a pre-trained sensing result state detection model based on first trajectory information, second trajectory information, and map element information, may include step 203a, which determines a detection result corresponding to a target object using a pre-trained sensing result state detection model based on first trajectory information, second trajectory information, and target map element information.
[0069] In one selective embodiment, the sensing result state detection model can be obtained by training it during the training process based on a first training trajectory information of a training target object in a unified coordinate system, a second training trajectory information of other training objects surrounding the training target object, and map element information. For specific training processes, please refer to the corresponding training method examples that follow.
[0070] In one optional example, step 203a may be performed by the processor calling a corresponding instruction stored in memory, or by a second processing module executed by the processor.
[0071] This embodiment unifies the map element information, first trajectory information, and second trajectory information into the same coordinate system by converting the map element information to the target coordinate system, thereby further improving the accuracy of the representation of the relative relationships between the target object, other objects, and map elements, and thereby improving the accuracy of the sensing result state.
[0072] In one selective embodiment, step 203 may specifically include the following steps:
[0073] In step 2031, the first feature extraction network in the sensing result state detection model is used to extract features from the first trajectory information and obtain the first feature.
[0074] Here, the first feature extraction network can be any feasible feature extraction network, such as a series of feature extraction networks based on convolutional neural networks, specifically a feature extraction network based on VGGNet (Visual Geometry Group Net), a feature extraction network based on ResNet (Residual Network), etc., and is not specifically limited.
[0075] In one optional example, step 2031 may be performed by the processor calling a corresponding instruction stored in memory, or by a first processing unit executed by the processor.
[0076] In step 2032, the second feature extraction network in the sensing result state detection model is used to extract features from the second trajectory information and obtain the second feature corresponding to the second trajectory information.
[0077] Here, the second feature extraction network can be any feasible feature extraction network, such as a feature extraction network based on VGGNet (Visual Geometry Group Net) or a feature extraction network based on ResNet (Residual Network), and is not specifically limited.
[0078] In one optional example, step 2032 may be performed by the processor calling a corresponding instruction stored in memory, or by a second processing unit executed by the processor.
[0079] In step 2033, a third feature extraction network in the sensing result state detection model is used to extract features from map element information and obtain static features.
[0080] Here, the third feature extraction network can be any feasible feature extraction network, such as a feature extraction network based on VGGNet (Visual Geometry Group Net) or a feature extraction network based on ResNet (Residual Network), and is not specifically limited to such networks.
[0081] In one selective embodiment, the network structures of the first feature extraction network, the second feature extraction network, and the third feature extraction network can be set to be the same or different depending on the actual needs, and the embodiments of this disclosure are not limited thereto.
[0082] In one selective embodiment, the first feature extraction network and the second feature extraction network may be the same feature extraction network.
[0083] Steps 2031, 2032, and 2033 do not distinguish in terms of priority.
[0084] In one optional example, step 2033 may be performed by the processor calling a corresponding instruction stored in memory, or by a third processing unit executed by the processor.
[0085] In step 2034, the attention network in the sensing result state detection model is used to process the first feature, second feature, and static feature to obtain the attention result.
[0086] Here, the specific structure of the attention network can be set according to the actual needs, and the attention network is used for cross-attention between the first feature, the second feature, and static features, thereby obtaining feature correlations between the first feature, the second feature, and static features.
[0087] Exemplary, an attention network may include, but is not specifically limited to, a self-attention layer, a cross-attention layer, and other related layers, such as an add-and-normalize layer. The attention network may be a single-head attention network or a multi-head attention network, but is not specifically limited to that.
[0088] Exemplary, Figure 7 is a schematic diagram of the structure of an attention network provided by one exemplary embodiment of the present disclosure. A first feature can be mapped to a first query vector Q1, a first key vector K1, and a first value vector V1 by a preset mapping relationship, and a first intermediate feature can be obtained through a self-attention layer, an add and normalize layer (Add&Norm), the first intermediate feature can be mapped to a second query vector Q2, the static feature can be mapped to a second key vector K2 and a second value vector V2, the second feature can be mapped to a third key vector K3 and a third value vector V3, cross-attention can be performed on the second query vector Q2 and the second key vector K2 and second value vector V2 to obtain a first attention result, cross-attention can be performed on the second query vector Q2 and the third key vector K3 and third value vector V3 to obtain a second attention result, and the first attention result, the second attention result, and the first intermediate feature are the attention results of the attention network. This is merely one illustrative structure and can be set to other possible structures depending on the actual needs in actual applications, and is not limited to the above structure.
[0089] In one optional example, step 2034 may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth processing unit executed by the processor.
[0090] In step 2035, the attention result is processed using the feature fusion network in the sensing result state detection model to obtain fused features.
[0091] Here, the feature fusion network may be a fully connected layer, and it achieves feature fusion of the attention results.
[0092] For example, the first attention result, the second attention result, and the first intermediate feature obtained in the above example can be combined using Concat to obtain a combined feature.
[0093] In one optional example, step 2035 may be performed by the processor calling a corresponding instruction stored in memory, or by a fifth processing unit executed by the processor.
[0094] In step 2036, the detection head network in the sensing result state detection model is used to process the fused features and obtain the detection result corresponding to the target object.
[0095] Here, the detection head network may be a multilayer perceptron (MLP), which is used to map fused features to the detection results.
[0096] In one optional example, step 2036 may be performed by the processor calling a corresponding instruction stored in memory, or by a sixth processing unit executed by the processor.
[0097] This embodiment realizes spatial correlation between the first feature of the first trajectory information of a target object, the second feature of the second trajectory information of other surrounding objects, and the static features of map element information based on an attention network. This is used to predict the perceived state of the target object and helps to improve the accuracy of the perceived state.
[0098] In one selective embodiment, the detection result may include at least one of a first probability that the perceived state of the target object is a first state and a second probability that it is a second state.
[0099] In one selective embodiment, step 204, which determines the sensing result state of a target object based on the detection result, may include step 2041, which determines the sensing result state of a target object based on at least one of a first probability and a second probability included in the detection result and a corresponding probability threshold.
[0100] Here, the probability threshold can be set according to the actual needs. For example, the probability threshold may include a first threshold corresponding to the first probability and a second threshold corresponding to the second probability, and the first and second thresholds were used as mapping thresholds between the first and second probabilities and the perceived result state, respectively.
[0101] For example, if the first state is a normal state, and the first probability is greater than the first threshold, then the perceived state is determined to be a normal state.
[0102] For example, if the second state is an abnormal state, and the second probability is greater than the second threshold, then the perceived state is determined to be an abnormal state.
[0103] In one optional example, step 2041 may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth decision unit executed by the processor.
[0104] Figure 8 is a flowchart of a method for determining the sensing result state of an object provided by another exemplary embodiment of the present disclosure.
[0105] In one selective embodiment, the method of the embodiment of the present disclosure may further include step 301: identifying an object in the sensing result corresponding to each time frame based on a preset identification rule; determining the sensing result state of a first object that satisfies the preset identification rule as an abnormal state; and filtering the first object from the sensing result corresponding to each time frame to obtain the filtered sensing result.
[0106] Here, preset identification rules can be set according to actual needs. For example, preset identification rules may include rules such as whether the yaw angle of an object jumps, whether the type of object is stable, and whether the object drifts laterally. If the yaw angle of an object jumps, it can be determined that the perceived state of the object is abnormal. If the type of an object jumps, it can be determined that the perceived state of the object is abnormal. If the object drifts laterally, it can be determined that the perceived state of the object is abnormal. A jump in yaw angle can refer to the change in yaw angle in adjacent time frames being greater than the change threshold. An unstable (jumping) object type can refer to the same object corresponding to at least two types in at least one time frame. For example, an object's type in time frame t-2 might be a vehicle, and its type in time frame t might be a pedestrian. Lateral drift of an object can refer to the change in the object's lateral coordinates exceeding the corresponding threshold. The lateral coordinates may be the lateral coordinates in the vehicle coordinate system of the object in the corresponding time frame. Based on preset identification rules, some objects with abnormal detection behavior are identified first, and the detection status of the remaining objects is identified through subsequent processing.
[0107] In one optional example, step 301 may be performed by the processor calling a corresponding instruction stored in memory, or by a preprocessing module executed by the processor.
[0108] Step 202, which determines first trajectory information corresponding to the target object within each object and second trajectory information of other objects around the target object based on the sensing results corresponding to each time frame, may include step 202a, which determines first trajectory information corresponding to the target object within each object and second trajectory information of other objects around the target object based on the filtered sensing results corresponding to each time frame.
[0109] The specific procedures for this step are similar to those described in step 202 above, and therefore will not be explained here.
[0110] In one optional example, step 202a may be performed by the processor calling a corresponding instruction stored in memory, or by a first processing module executed by the processor.
[0111] This embodiment identifies objects in the sensing result based on preset identification rules and identifies objects that are clearly abnormal in sensing. On the one hand, this reduces the amount of data required for subsequent model predictions and improves processing efficiency. On the other hand, it reduces the interference that these clearly abnormal objects have on model inference and further improves the accuracy of the sensing result state.
[0112] Each of the embodiments described herein can be implemented independently or combined in any combination, provided they do not conflict, and can be specifically configured according to actual needs, and the embodiments described herein are not limited thereto.
[0113] The method for determining the sensing result state of any one object provided in the embodiments of this disclosure can be executed by any device having appropriate data processing capabilities, which includes, but is not limited to, terminal devices and servers. Alternatively, the method for determining the sensing result state of any one object provided in the embodiments of this disclosure can be executed by a processor, for example, by calling a corresponding instruction stored in memory to execute the method for determining the sensing result state of any one object described in the embodiments of this disclosure. Further explanation is omitted below.
[0114] Figure 9 is a flowchart of a training method for a sensing result state detection model provided by one exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, specifically electronic devices such as servers and terminal devices. As shown in Figure 9, the method of the embodiment of the present disclosure may include the following steps:
[0115] Step 401 obtains first training trajectory information corresponding to each training target object in at least one training target object, second training trajectory information for other surrounding training targets corresponding to each training target object, and map element information corresponding to each training target object.
[0116] Here, the training target object can be determined based on pre-collected sensing result data, and the sensing result data may include sensing results from at least one time frame. The first training trajectory information is similar to the first trajectory information described above, and the second training trajectory information is similar to the second trajectory information described above, and their explanation is omitted here.
[0117] In one optional example, step 401 may be performed by the processor calling a corresponding instruction stored in memory, or by a second acquisition module executed by the processor.
[0118] In step 402, the pre-constructed sensing result state detection network is trained based on each first training trajectory information, each second training trajectory information, and each map element information to obtain a trained sensing result state detection model.
[0119] Here, the network structure of the sensing result state detection network can be described by referring to the sensing result state detection model of the embodiment described above, and will not be explained here.
[0120] In one selective embodiment, the state label data required for the model training process can be determined based on the sensing results of the target object in the future time frame in the sensing results data, thereby enabling the model training to integrate the object's historical state, current state, and future state, further improving the model's performance. Here, for the future time frame of the target object, any one of the multiple time frames in the sensing results data can be used as the current time frame. In this case, the time frame before the current time frame is the historical time frame, and the time frame after the current time frame is the future time frame.
[0121] In one selective embodiment, the sensing result state detection model can be obtained by training the sensing result state detection network in conjunction with at least one of a trajectory prediction head network and a trajectory planning head network based on imitation learning, thereby improving the performance of the sensing result state detection model.
[0122] In one optional example, step 402 may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth processing module executed by the processor.
[0123] The training method for a sensing result state detection model provided by the embodiments of this disclosure trains the sensing result state detection model using first training trajectory information of a training target object, second training trajectory information of other training objects surrounding the training target object, and map element information to obtain a detection model that can effectively detect the sensing result state of an object. This method is used to identify objects with sensing anomalies, helps to improve the generalization ability of identifying objects with sensing anomalies, covers a wider range of sensing anomaly situations, and helps to improve work efficiency and processing efficiency compared to artificially designing rules.
[0124] Figure 10 is a flowchart of a training method for a sensing result state detection model provided in another exemplary embodiment of the present disclosure.
[0125] In one selective embodiment, obtaining first training trajectory information corresponding to each training target object in at least one training target object in step 401, and second training trajectory information for other surrounding training targets corresponding to each training target object, may include the following steps:
[0126] Step 4011 involves obtaining sensing results corresponding to each time frame in at least one time frame, and the sensing results include state information of at least one object in a first coordinate system.
[0127] Here, at least one time frame may be a time frame within any historical time period. For example, sensing result data for a certain time period during the actual driving process of a vehicle can be obtained as sensing results corresponding to each time frame within at least one time frame, and can be specifically set according to actual needs. The first coordinate system and object state information can be found in the embodiments described above and will not be explained here.
[0128] In step 4012, one of the time frames in each time frame is designated as the current time frame, and for any object in the current time frame, a first sensing result state corresponding to that object is determined based on the object's state information in the current time frame and the previous time frame, as well as the preset identification rules.
[0129] Here, the specific details of the preset identification rules can be found in the previously mentioned section. Based on the state information of the object in the current time frame and the previous time frame, the specific conditions of the object, such as its position and orientation, can be determined for each time frame. The position and orientation of the object in each time frame are then compared with the preset identification rules to determine the perceived state of the object, which is set as the first perceived state. For example, if it is determined that the object is drifting laterally based on its position and orientation in each time frame, then the first perceived state of the object can be determined to be an abnormal state.
[0130] In step 4013, in response to the first sensing result state corresponding to the object being the first state, the object is designated as the training target object.
[0131] Here, the first state may be a normal state, and if the first perceived result state corresponding to the object is a normal state, the object can be used as a training target object, thus avoiding the unfavorable interference of clearly abnormal objects from being used as training target objects in model training and inference.
[0132] In one selective embodiment, the second state may be set as the normal state, in which case step 4013 should set the object as the training target object in response to the first sensing result state corresponding to the object being the second state.
[0133] In step 4014, the first training trajectory information corresponding to the training target object is determined based on the state information of the training target object in the current time frame and the previous time frame.
[0134] Here, the determination of the first training trajectory information can be done by referring to the first trajectory information mentioned above, and therefore the explanation is omitted here.
[0135] In step 4015, other objects surrounding the training target object in the current time frame and the previous time frame are designated as other training objects.
[0136] In step 4016, based on the state information of each other training object in the current time frame and the previous time frame, the second training trajectory information corresponding to each other training object is determined.
[0137] Here, the determination of the second training trajectory information can be done by referring to the aforementioned second trajectory information, and therefore the explanation is omitted here.
[0138] Step 401 may further include step 4017, which obtains map element information corresponding to at least one time frame or map element information corresponding to each training target object.
[0139] Here, the same set of map element information can be used for each time frame in at least one time frame; that is, the map element information includes map element information within the domain range of each time frame, and the map element information corresponding to each training target object is the same set of map element information. It is also possible to associate one set of map element information with each training target object, and such map element information includes, but is not specifically limited to, map element information within the range in which the target object is involved.
[0140] In one optional example, steps 4011 through 4017 above may be performed by the processor calling corresponding instructions stored in memory, or by a second acquisition module executed by the processor.
[0141] This embodiment uses previously collected sensing result data to determine training target objects, which are then used to train a sensing result state detection model, helping to improve model performance.
[0142] In one selective embodiment, the method of the embodiment of the present disclosure may further include the following steps:
[0143] Step 5011 determines the second perceived result state of the training target object in a subsequent time frame.
[0144] Here, the later time frame may include at least one future time frame relative to the current time frame, and the second perceived result state may be determined based on preset identification rules. For example, the second perceived result state of the training target object is determined based on whether or not a situation such as lateral drift, type instability, or yaw angle jump occurs in each of the at least one future time frame. For example, if the target object drifts laterally in a future time frame, the second perceived result state may be determined to be an abnormal state. If no situation matching the preset identification rule occurs, the second perceived result state may be determined to be a normal state.
[0145] In step 5012, a state label corresponding to the training target object is determined based on the second sensing result state.
[0146] Here, different states of the second sensing result state can be represented by different labels. For example, if the second sensing result state is a normal state, the corresponding state label is 1 or 0, and similarly, if the second sensing result state is an abnormal state, the corresponding state label is 0 or 1.
[0147] In step 5013, state label data is determined based on the state labels corresponding to each training target object.
[0148] Here, multiple training target objects can determine training sample data for multiple training samples, and the state labels corresponding to each training sample form state label data corresponding to the training sample data.
[0149] In one optional example, steps 5011 to 5013 above may be performed by the processor calling corresponding instructions stored in memory, or by a second acquisition module executed by the processor.
[0150] In this embodiment, the state label of the training target object is determined based on the sensing result of the training target object's future time frame, and used for model training. Since the state label represents the actual state of the training target object in its future time frame, the historical state, current state, and future state of the training target object can be integrated to further improve model performance.
[0151] Step 402 may specifically include training a pre-built sensing result state detection network based on each first training trajectory information, each second training trajectory information, each map element information, and state label data, in order to obtain a trained sensing result state detection model.
[0152] In one selective embodiment, the method of the embodiment of the present disclosure may further include the following steps:
[0153] In step 5021, a trajectory label corresponding to the training target object is determined based on the second sensing result state and the state information of the training target object in subsequent time frames.
[0154] Here, the trajectory label may include an output label and a corresponding trajectory type, the trajectory type may be at least one of the preset types, the preset type may refer to a type of driving trajectory that is possible to occur within a predetermined future time (e.g., 6 seconds) of the vehicle based on the vehicle's speed, straight-line driving, turning, etc., and when trajectory prediction is performed, each preset type may correspond to a set of anchor trajectory points, each set of anchor trajectory points may include a preset number (e.g., 12) of coordinate points, the coordinate points may be coordinate points in a coordinate system with the object whose trajectory is to be predicted (e.g., a training target object) as the origin, the coordinate points represent possible trajectories that may occur in the future of the object, and the future trajectory of the object is determined by predicting the probability and offset value of each preset type of the object, for example, the sum of the anchor trajectory points and offset value corresponding to at least one preset type with a relatively high probability may be the future trajectory of the object. During the training process, the trajectory type of the trajectory label can be determined based on the future time frame of the training target object, to which the actual future trajectory of the training target object belongs. The trajectory label may further include a non-output label. Here, the output label indicates that the sensing result state of the training target object is normal and that a predicted trajectory can be output, while the non-output label indicates that the sensing result state of the training target object is abnormal and that trajectory prediction will not be performed. The representation form of the trajectory label can be set according to the actual needs, and this embodiment is not limited thereto.
[0155] In step 5022, trajectory label data is determined based on the trajectory labels corresponding to each training target object.
[0156] In one optional example, steps 5021 and 5022 may be performed by the processor calling corresponding instructions stored in memory, or by a second acquisition module executed by the processor.
[0157] Step 402, which involves training a pre-constructed sensing result state detection network based on each first training trajectory information, each second training trajectory information, and each map element information to obtain a trained sensing result state detection model, may include step 402a, which involves co-training a pre-constructed multitask network, including a sensing result state detection network and a trajectory prediction head network, based on each first training trajectory information, each second training trajectory information, each map element information, state label data, and trajectory label data to obtain a trained sensing result state detection model.
[0158] Here, the multitasking network may include a sensing result state detection network and a trajectory prediction network, and the sensing result state detection network and the trajectory prediction network may share the front part of the network, with different head networks performing different tasks.
[0159] In one selective embodiment, Figure 11 is a schematic diagram of a multitasking network provided by one exemplary embodiment of the present disclosure. As shown in Figure 11, the multitasking network may include a first feature extraction network, a second feature extraction network, a third feature extraction network, an attention network, a feature fusion network, a detection head network for a sensing result state detection network, and a trajectory prediction head network for a trajectory prediction network.Here, both the first and second feature extraction networks can be used to extract trajectory information, and the first and second feature extraction networks can employ the same feature extraction network. During the training process, the first feature extraction network can perform feature extraction on the first training trajectory information of the training target object to obtain the first training features, the second feature extraction network can perform feature extraction on the second training trajectory information of other training objects surrounding the training target object to obtain the second training features, and the third feature extraction network can perform feature extraction on map element information to obtain static training features. The first training features, second training features, and static training features obtained through extraction can be used in an attention network to perform attention operations and obtain training attention results. The training attention results can then be used in a feature fusion network to perform feature fusion and obtain training fusion features. A characteristic can be obtained, the training fusion feature can obtain training state detection results by the detection head network, the training fusion feature can obtain training trajectory prediction results by the trajectory prediction head network, the network loss can be determined based on the training sensing result state detection results and corresponding state label data, and the training trajectory prediction results and corresponding trajectory label data, the network parameters of the multitask network can be updated based on the network loss, and the updated multitask network can be obtained, if the updated multitask network satisfies the training termination condition, training is terminated, otherwise the updated multitask network is repeatedly updated by the above process until the updated multitask network satisfies the training termination condition, and a multitask model can be obtained, thereafter the portion of the multitask model excluding the trajectory prediction head network can be made into a trained sensing result state detection model. The portion excluding the detection head network can also be made into a trajectory prediction model.In the model application process, the specific application of each part of the network can be found in the previously mentioned examples, and will not be explained here.
[0160] In one optional example, step 402a may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth processing module executed by the processor.
[0161] This embodiment determines the trajectory label of the training target object based on the sensing results of future frames of the training target object, and further improves the performance of the sensing result state detection model by jointly training the sensing result state detection network and the trajectory prediction network.
[0162] Figure 12 is a flowchart of a training method for a sensing result state detection model provided in one further exemplary embodiment of the present disclosure.
[0163] In one selective embodiment, step 402, which involves training a pre-built sensing result state detection network based on each first training trajectory information, each second training trajectory information, and each map element information to obtain a trained sensing result state detection model, may include the following steps:
[0164] In step 4021, the multitasking network is determined based on the sensing result state detection network and the trajectory prediction head network.
[0165] Here, the specific network structure of the multitasking network can be configured according to the actual needs.
[0166] In step 4022, the training state detection result and training trajectory prediction result are determined using a multitasking network based on each first training trajectory information, each second training trajectory information, and each map element information.
[0167] The specific reasoning process for this step can be found in the previously mentioned content and will not be explained here.
[0168] In step 4023, the first loss is determined based on the training state detection results and the corresponding state label data.
[0169] Here, the first loss can be determined based on a first preset loss function, which can be any feasible loss function such as the mean squared error loss function, the mean absolute error loss function, or the cross-entropy loss function.
[0170] In step 4024, the second loss is determined based on the training trajectory prediction results and trajectory label data.
[0171] Here, the second loss can be determined based on a second preset loss function, which can be any feasible loss function such as the mean squared error loss function, mean absolute error loss function, or cross-entropy loss function. Specifically, it can be set according to the actual needs.
[0172] In step 4025, the total loss is determined based on the first and second losses.
[0173] Here, the total loss can be obtained by weighting the first and second losses with preset weights, the specific preset weights can be set according to actual needs, and the embodiments of this disclosure are not limited thereto.
[0174] In step 4026, in response to the overall loss not meeting the training termination condition, the multitask network is updated based on the overall loss, and the updated multitask network is obtained.
[0175] Here, the training termination condition may include at least one of the following conditions: model convergence and the number of iterations reaching a preset threshold. Updates to a multitask network can be implemented based on any feasible gradient descent algorithm, such as the stochastic gradient descent algorithm or the stochastic average gradient descent algorithm.
[0176] In step 4027, the updated multitasking network is set as the multitasking network, and step 4022 is repeated and executed.
[0177] In step 4028, a trained multitask model is obtained in response to the overall loss meeting the training termination condition.
[0178] In step 4029, the sensing result state detection network portion of the multitasking model is defined as the sensing result state detection model.
[0179] In one optional example, steps 4021 to 4029 above may be performed by the processor calling corresponding instructions stored in memory, or by a fourth processing module executed by the processor.
[0180] In this embodiment, by jointly training the sensing result state detection network and the trajectory prediction network, the losses from the two tasks can be integrated to update the network parameters, thereby further improving model performance.
[0181] The embodiments of this disclosure can capture potential features of anomaly-perceiving obstacles, such as location distribution, surrounding obstacle conditions, and occlusion conditions, through model training, which helps improve processing efficiency and the generalization performance of identifying anomaly-perceiving objects. Furthermore, by making full use of future trajectories during the training process, possible relationships between the current state of the obstacle and future trajectories can be considered, further improving the accuracy of the perceived state compared to making judgments based solely on historical trajectories. In addition, the model obtained through training can be used to identify anomaly-perceiving obstacles, and in actual applications, it can reduce delays caused by various rule-based decisions, resulting in faster identification speeds and higher efficiency.
[0182] Each of the embodiments described herein can be implemented independently or combined in any combination, provided they do not conflict, and can be specifically configured according to actual needs, and the embodiments described herein are not limited thereto.
[0183] Any training method for a sensing result state detection model provided in the embodiments of this disclosure can be performed by any device having appropriate data processing capabilities, including, but not limited to, terminal devices and servers. Alternatively, any training method for a sensing result state detection model provided in the embodiments of this disclosure can be performed by a processor, for example, by calling a corresponding instruction stored in memory to perform any training method for a sensing result state detection model described in the embodiments of this disclosure. Further explanation is omitted below.
[0184] As those skilled in the art will understand, all or some of the steps to implement the above-described method embodiment can be completed by hardware related to program instructions, the aforementioned program can be stored in a computer-readable storage medium, and when the program is executed, the steps of the above-described method embodiment are performed, the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks or optical disks.
[0185] [Example device] Figure 13 is a schematic diagram of an object sensing result state determination device provided in one exemplary embodiment of the present disclosure. The device of this embodiment can be used to implement an embodiment of the corresponding object sensing result state determination method of the present disclosure. The device shown in Figure 13 may include a first acquisition module 601, a first processing module 602, a second processing module 603, and a third processing module 604.
[0186] The first acquisition module 601 can be used to acquire sensing results and map element information corresponding to each of the time frames in at least one time frame, the sensing results including state information of at least one object in a first coordinate system.
[0187] The first processing module 602 can be used to determine first trajectory information corresponding to a target object among the objects, and second trajectory information of other objects around the target object, based on the sensing results corresponding to each of the time frames.
[0188] The second processing module 603 can be used to determine the detection result corresponding to the target object by utilizing a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information, and the map element information.
[0189] The third processing module 604 can be used to determine the sensing result state of the target object based on the detection result.
[0190] Figure 14 is a schematic diagram of a device for determining the sensing result state of an object, provided in another exemplary embodiment of the present disclosure.
[0191] In one selective embodiment, at least one time frame includes a current time frame and at least one historical time frame, the state information includes the position and orientation of the object, and the first processing module 602 may include a first decision unit 6021, a first conversion unit 6022, and a second decision unit 6023.
[0192] The first decision unit 6021 can be used to determine a target coordinate system with the target object as the origin, based on the position and orientation of the target object in the current time frame.
[0193] The first transformation unit 6022 can be used to transform the state information of a target object in any one historical time frame into the target coordinate system based on the transformation relationship between the first coordinate system and the target coordinate system, and to obtain the first state information in the target coordinate system corresponding to that historical time frame.
[0194] The second decision unit 6023 can be used to determine the first trajectory information corresponding to the target object based on the first state information corresponding to each historical time frame and the origin state information corresponding to the current time frame, the origin state information being the first state information of the target object in the target coordinate system.
[0195] In one selective embodiment, the first processing module 602 may further include a second conversion unit 6024 and a third decision unit 6025.
[0196] The second transformation unit 6024 can be used to transform the state information of another object in any one time frame into the target coordinate system based on the transformation relationship between the first coordinate system and the target coordinate system, and to obtain the second state information of the other object in that time frame.
[0197] The third decision unit 6025 can be used to determine a second trajectory information corresponding to any one other object, based on the second state information of that other object in each time frame.
[0198] In one selective embodiment, the first processing module 602 may further include a third transformation unit 6026 used to transform map element information into a target coordinate system and obtain target map element information, based on a transformation relationship between the map coordinate system and the target coordinate system corresponding to the map element information.
[0199] Specifically, the second processing module 603 can be used to determine the detection result corresponding to the target object by utilizing a pre-trained sensing result state detection model based on the first trajectory information, each second trajectory information, and target map element information.
[0200] In one selective embodiment, the second processing module 603 may include a first processing unit 6031, a second processing unit 6032, a third processing unit 6033, a fourth processing unit 6034, a fifth processing unit 6035, and a sixth processing unit 6036.
[0201] The first processing unit 6031 can be used to extract features from the first trajectory information and obtain the first feature by utilizing the first feature extraction network in the sensing result state detection model.
[0202] The second processing unit 6032 can be used to extract features from the second trajectory information using the second feature extraction network in the sensing result state detection model, and to obtain a second feature corresponding to the second trajectory information.
[0203] The third processing unit 6033 can be used to extract features from map element information and obtain static features by utilizing the third feature extraction network in the sensing result state detection model.
[0204] The fourth processing unit 6034 can be used to process the first feature, second feature, and static feature using the attention network in the sensing result state detection model to obtain an attention result.
[0205] The fifth processing unit 6035 can be used to process the attention result and obtain fused features by utilizing the feature fusion network in the sensing result state detection model.
[0206] The sixth processing unit 6036 can be used to process fused features using the detection head network in the sensing result state detection model to obtain detection results corresponding to the target object.
[0207] In one selective embodiment, the detection result may include at least one of a first probability that the perceived state of the target object is a first state and a second probability that it is a second state, and the third processing module 604 may include a fourth determination unit 6041 that can be used to determine the perceived state of the target object based on at least one of the first and second probabilities and a corresponding probability threshold.
[0208] In one selective embodiment, the apparatus of the embodiment of the present disclosure may further include a preprocessing module 610 used to identify objects in the sensing results corresponding to each time frame based on a preset identification rule, determine the sensing result state of a first object that satisfies the preset identification rule as an abnormal state, and filter the first object from the sensing results corresponding to each time frame to obtain the filtered sensing results.
[0209] Specifically, the first processing module 602 can be used to determine first trajectory information corresponding to the target object within each object, and second trajectory information of other objects around the target object, based on the filtered sensing results corresponding to each time frame.
[0210] Each of the embodiments described herein may be implemented individually or combined in any combination as long as it does not conflict, and can be specifically configured according to actual needs, and the embodiments described herein are not limited to these.
[0211] Figure 15 is a schematic diagram of a training apparatus for a sensing result state detection model provided in one exemplary embodiment of the present disclosure. The apparatus of this embodiment can be used to implement an embodiment of the corresponding sensing result state detection model training method of the present disclosure. The apparatus shown in Figure 15 may include a second acquisition module 701 and a fourth processing module 702.
[0212] The second acquisition module 701 can be used to acquire first training trajectory information corresponding to each training target object in at least one training target object, second training trajectory information of other surrounding training targets corresponding to each training target object, and map element information corresponding to each training target object.
[0213] The fourth processing module 702 can be used to train a pre-constructed sensing result state detection network based on each first training trajectory information, each second training trajectory information, and each map element information, in order to obtain a trained sensing result state detection model.
[0214] In one selective embodiment, the second acquisition module 701 can be used to acquire sensing results corresponding to each time frame in at least one time frame, the sensing results may include state information in a first coordinate system of at least one object, any one time frame in each time frame may be the current time frame, and for any one object in the current time frame, it can be used to determine a first sensing result state corresponding to the object based on the object's state information in the current time frame and previous time frame and a preset identification rule, in response to the first sensing result state corresponding to the object being a first state, it can be used to designate the object as a training target object, it can be used to determine first training trajectory information corresponding to the training target object based on the state information corresponding to the training target object in its current time frame and previous time frame, it can be used to designate other objects in the surrounding area of the training target object in its current time frame and previous time frame as other training objects, and it can be used to determine second training trajectory information corresponding to each other training object based on the state information of each other training object in its current time frame and previous time frame.
[0215] In one selective embodiment, the second acquisition module 701 can be used to determine a second perceived result state of the training target object in a subsequent time frame, to determine a state label corresponding to the training target object based on the second perceived result state, and to determine state label data based on the state label corresponding to each training target object.
[0216] The fourth processing module 702 can be used to train a pre-built sensing result state detection network based on each first training trajectory information, each second training trajectory information, each map element information, and state label data, in order to obtain a trained sensing result state detection model.
[0217] In one selective embodiment, the second acquisition module 701 can be used to determine a trajectory label corresponding to the training target object based on the second sensing result state and the state information of the training target object in a subsequent time frame, and can be used to determine trajectory label data based on the trajectory label corresponding to each training target object.
[0218] The fourth processing module 702 can be used to perform joint training on a pre-built multitask network, including a sensing result state detection network and a trajectory prediction head network, based on each first training trajectory information, each second training trajectory information, each map element information, state label data, and trajectory label data, in order to obtain a trained sensing result state detection model.
[0219] In one selective embodiment, the fourth processing module 702 can be used to determine a multitask network based on a sensing result state detection network and a trajectory prediction head network; it can be used to determine training state detection results and training trajectory prediction results using the multitask network based on each first training trajectory information, each second training trajectory information and each map element information; it can be used to determine a first loss based on the training state detection results and corresponding state label data; it can be used to determine a second loss based on the training trajectory prediction results and trajectory label data; and it can be used to determine an overall loss based on the first and second losses. In response to the failure to meet the training termination conditions, the multitask network can be updated based on the total loss to obtain the updated multitask network. The updated multitask network can then be used to repeatedly perform the step of determining the training state detection result and the training trajectory prediction result using the multitask network based on each first training trajectory information, each second training trajectory information, and each map element information. In response to the total loss meeting the training termination conditions, a trained multitask model can be obtained. The sensing result state detection network portion of the multitask model can then be used to obtain the sensing result state detection model.
[0220] Each of the embodiments described herein may be implemented individually or combined in any combination, provided they do not conflict, and can be specifically configured according to actual needs, and the embodiments described herein are not limited thereto.
[0221] Beneficial technical effects corresponding to exemplary embodiments of this apparatus can be found by referring to the corresponding beneficial technical effects of the exemplary method portion described above, and are therefore omitted from this description.
[0222] [Example electronic device] Figure 16 is a structural diagram of one electronic device provided by an embodiment of the present disclosure, which includes at least one processor 11 and memory 12.
[0223] The processor 11 may be a central processing unit (CPU) or another form of processing unit having data processing capability and / or instruction execution capability, and can control other components in the electronic device 10 to perform a desired function.
[0224] Memory 12 may include one or more computer program products, including various forms of computer-readable storage media such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or fast cache memory (cache). Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored in the computer-readable storage media, and the processor 11 can execute one or more computer program instructions to realize the methods of each embodiment of the present disclosure described above and / or other desired functions.
[0225] In one example, the electronic device 10 may further include an input device 13 and an output device 14, and these components are connected to each other via a bus system and / or other forms of connection mechanisms (not shown).
[0226] The input device 13 may further include, for example, a keyboard, a mouse, and the like.
[0227] The output device 14 can output various types of information to the outside, and may include, for example, a display, speaker, printer, communication network, and remote output devices connected thereto.
[0228] Naturally, for the sake of simplification, Figure 16 shows only some of the components of the electronic device 10 relevant to this disclosure, and components such as buses and input / output interfaces are omitted. Beyond this, the electronic device 10 may include any other appropriate components depending on the specific application.
[0229] [Exemplary computer program products and computer-readable storage media] ] In addition to the methods and apparatus described above, embodiments of the present disclosure may also provide a computer program product that, when executed by a processor, includes computer program instructions that cause the processor to perform steps of the various embodiments of the present disclosure described in the “Exemplary Methods” section above.
[0230] Computer program products can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and common procedural programming languages such as the C language or similar programming languages, to create program code for performing the operations of the embodiments of this disclosure. The program code may run entirely on a user computing device, partially on a user computing device, as a standalone software package, partially on a user computing device, partially on a remote computing device, or entirely on a remote computing device or server.
[0231] In addition, embodiments of the present disclosure may further be a computer-readable storage medium storing computer program instructions that, when executed by a processor, cause the processor to perform the steps of the various embodiments of the present disclosure described in the “Exemplary Methods” section above.
[0232] Computer-readable storage media may employ any combination of one or more readable media. The readable media may be readable signal media or readable storage media. Readable storage media may include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any combination thereof. More specific examples (non-exclusive list) of readable storage media include electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0233] While the basic principles of this disclosure have been explained above with reference to specific examples, the advantages, advantages, and effects mentioned in this disclosure are merely illustrative and not limiting, and it is not assumed that these advantages, advantages, and effects must be present in each example of this disclosure. Furthermore, the specific details disclosed above are merely illustrative and intended to facilitate understanding, and are not limiting, nor do they imply that this disclosure must be implemented in the specific details described above.
[0234] Those skilled in the art can make various modifications and variations to this disclosure without departing from the spirit and scope of the application. Thus, if such modifications and variations of the application fall within the scope of the claims of this disclosure and the equivalent art, this disclosure is intended to include such modifications and variations.
[0235] [Cross-reference of related applications] This disclosure claims priority to the Chinese patent application filed with the China National Intellectual Property Administration on March 27, 2023, application number CN202310309049.4, with the title of the invention "Method for determining the sensing result state of an object, model training method and apparatus," all of which are incorporated herein by reference.
Claims
1. A step of acquiring sensing results and map element information corresponding to each of the time frames in at least one time frame, wherein the sensing results include state information of at least one object in a first coordinate system; The steps include determining, based on the sensing results corresponding to each of the aforementioned time frames, first trajectory information corresponding to a target object among the aforementioned objects, and second trajectory information of other objects surrounding the target object, The steps include determining a detection result corresponding to the target object using a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information, and the map element information, A method for determining the sensing result state of an object, comprising the step of determining the sensing result state of the target object based on the detection result.
2. The step of determining the detection result corresponding to the target object using a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information, and the map element information is as follows: The first step is to extract features from the first trajectory information using the first feature extraction network in the aforementioned sensing result state detection model to obtain the first feature, The steps include: using the second feature extraction network in the aforementioned sensing result state detection model to extract features from the second trajectory information and obtain a second feature corresponding to the second trajectory information; The steps include: using the third feature extraction network in the aforementioned sensing result state detection model to extract features from the map element information and obtain static features; The steps include: using the attention network in the sensing result state detection model to process the first feature, the second feature, and the static feature to obtain an attention result; The steps include: using the feature fusion network in the aforementioned sensing result state detection model to process the attention result and obtain a fused feature; The method according to claim 1, comprising the step of using the detection head network in the sensing result state detection model to process the fused features and obtain the detection result corresponding to the target object.
3. The at least one time frame includes the current time frame and at least one historical time frame, and the state information includes the position and orientation of the object. Determining first trajectory information corresponding to the target object among the objects based on the sensing results corresponding to each of the time frames is: A step of determining a target coordinate system with the target object as the origin, based on the position and orientation of the target object in the current time frame. For any one of the historical time frames, the state information of the target object in that historical time frame is transformed into the target coordinate system based on the transformation relationship between the first coordinate system and the target coordinate system, and the first state information in the target coordinate system corresponding to that historical time frame is obtained. The method according to claim 1, comprising the step of determining the first trajectory information corresponding to the target object based on the first state information corresponding to each of the historical time frames and the origin state information corresponding to the current time frame, wherein the origin state information is the first state information of the target object in the target coordinate system.
4. Determining second trajectory information of other objects around the target object based on the sensing results corresponding to each of the aforementioned time frames is: For any one of the aforementioned time frames, the state information of the other object in that time frame is transformed into the target coordinate system based on the transformation relationship between the first coordinate system and the target coordinate system, and a second state information of the other object in that time frame is obtained. The method according to claim 3, comprising the step of determining a second trajectory information corresponding to any one of the other objects based on the second state information of the other object in each of the time frames.
5. The process further includes the step of converting the map element information to the target coordinate system based on the transformation relationship between the map coordinate system corresponding to the map element information and the target coordinate system, and obtaining target map element information. The step of determining the detection result corresponding to the target object using a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information, and the map element information is as follows: The method according to claim 3, further comprising the step of determining a detection result corresponding to the target object using a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information, and the target map element information.
6. The detection result includes at least one of a first probability that the perceived state of the target object is a first state and a second probability that it is a second state. The step of determining the sensing result state of the target object based on the detection result is: The method according to any one of claims 1 to 5, comprising the step of determining the sensing result state of the target object based on at least one of the first probability and the second probability and a corresponding probability threshold.
7. The steps include: identifying an object in the sensing result corresponding to each of the time frames based on a preset identification rule, and determining that the sensing result state of a first object that satisfies the preset identification rule is an abnormal state; The method further includes the step of filtering the first object from the sensing results corresponding to each of the time frames to obtain the filtered sensing results, The step of determining, based on the sensing results corresponding to each of the aforementioned time frames, first trajectory information corresponding to a target object among the aforementioned objects, and second trajectory information of other objects surrounding the target object, is as follows: The method according to any one of claims 1 to 5, comprising the step of determining first trajectory information corresponding to a target object among the objects, and second trajectory information of other objects around the target object, based on the filtered sensing results corresponding to each of the time frames.
8. Steps include obtaining first training trajectory information corresponding to each of the training target objects in at least one training target object, second training trajectory information of other surrounding training targets corresponding to each of the training target objects, and map element information corresponding to each of the training target objects, A method for training a sensing result state detection model, comprising the steps of: training a pre-constructed sensing result state detection network based on each of the first training trajectory information, each of the second training trajectory information, and each of the map element information to obtain a trained sensing result state detection model.
9. The step of training a pre-constructed sensing result state detection network based on each of the first training trajectory information, each of the second training trajectory information, and each of the map element information, in order to obtain a trained sensing result state detection model, is as follows: The steps include determining a multitasking network based on the aforementioned sensing result state detection network and trajectory prediction head network, A step of determining a training state detection result and a training trajectory prediction result using the multitask network based on each of the first training trajectory information, each of the second training trajectory information, and each of the map element information, The steps include determining a first loss based on the training state detection result and the corresponding state label data, The steps include determining a second loss based on the aforementioned training trajectory prediction results and trajectory label data, A step of determining the total loss based on the first loss and the second loss, In response to the fact that the overall loss does not meet the training termination condition, the multitask network is updated based on the overall loss, and the updated multitask network is obtained. The updated multitasking network is defined as the multitasking network, and the step of repeatedly executing the step of determining the training state detection result and the training trajectory prediction result using the multitasking network based on each of the first training trajectory information, each of the second training trajectory information, and each of the map element information, The steps include obtaining a trained multitask model in response to the total loss meeting the training termination condition, and The method according to claim 8, comprising the step of making the sensing result state detection network portion in the multitasking model the sensing result state detection model.
10. Acquiring first training trajectory information corresponding to each training target object in at least one training target object, and second training trajectory information for other surrounding training targets corresponding to each training target object, A step of obtaining a sensing result corresponding to each of the time frames in at least one time frame, wherein the sensing result includes state information of at least one object in a first coordinate system. The steps include: setting one of the time frames in each of the aforementioned time frames as the current time frame; determining a first sensing result state corresponding to one of the objects in the current time frame based on the object's state information in the current time frame and the previous time frame, and a preset identification rule; In response to the first sensing result state corresponding to the object being the first state, the step of designating the object as the training target object, A step of determining the first training trajectory information corresponding to the training target object based on the state information of the training target object in the current time frame and the previous time frame, The steps include: making other objects surrounding the training target object in the current time frame and the previous time frame into other training objects; The method according to claim 9, comprising the step of determining a second training trajectory information corresponding to each of the other training objects based on the state information of the current time frame and the previous time frame of each of the other training objects.
11. The steps include determining the second perception result state of the training target object in a subsequent time frame, The steps include determining a state label corresponding to the training target object based on the second sensing result state, The steps include determining the state label data based on the state label corresponding to each of the training target objects, A step of determining a trajectory label corresponding to the training target object based on the second sensing result state and the state information of the training target object in the subsequent time frame, The method according to claim 10, further comprising the step of determining the trajectory label data based on the trajectory label corresponding to each of the training target objects.
12. A first acquisition module used to acquire sensing results and map element information corresponding to each of at least one time frames, wherein the sensing results include state information of at least one object in a first coordinate system, A first processing module used to determine, based on the sensing results corresponding to each of the aforementioned time frames, first trajectory information corresponding to a target object among the aforementioned objects, and second trajectory information of other objects surrounding the target object, A second processing module used to determine the detection result corresponding to the target object, using a pre-trained sensing result state detection model based on the first trajectory information, the second trajectory information, and the map element information; A device for determining the sensing result state of an object, comprising: a third processing module used to determine the sensing result state of the target object based on the detection result.
13. A second acquisition module used to acquire first training trajectory information corresponding to each of at least one training target object, second training trajectory information of other surrounding training objects corresponding to each of the training target objects, and map element information corresponding to each of the training target objects, A training device for a sensing result state detection model, comprising: a fourth processing module used to train a pre-constructed sensing result state detection network based on each of the first training trajectory information, each of the second training trajectory information, and each of the map element information, in order to obtain a trained sensing result state detection model.
14. A computer-readable storage medium storing a computer program for performing a method for determining the sensing result state of an object according to any one of claims 1 to 7, or a method for training a sensing result state detection model according to any one of claims 8 to 11.
15. Processor and The processor includes a memory for storing executable instructions, Electronic device wherein the processor reads and executes the executable instructions from the memory to realize a method for determining the sensing result state of an object according to any one of claims 1 to 7, or a method for training a sensing result state detection model according to any one of claims 8 to 11.