Object Detection Method, Vehicle Control Method, Device, Vehicle, Medium and Chip
By using environmental data at multiple moments in the autonomous driving vehicle for two-stage optimization and solution, the initial position and speed information of the target object are obtained, the problem of jumping in the target object status information is solved, and the accuracy of decision-making planning and control of the autonomous driving vehicle is improved.
Patent Information
- Application Number
- CN202411612966.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-12
AI Technical Summary
In the prior art, there are problems of inaccuracy and jump in the prediction results of target object detection models of autonomous driving vehicles, which affect the accuracy of decision planning and control.
By obtaining the initial position information and velocity information of the target object based on environmental data at multiple moments, the two-stage optimization solution method is used to determine the target state information of the target object, including position, speed and heading angle, to ensure that the speed information complies with the uniform motion constraints and avoid jumping of the state information.
It improves the accuracy of decision-making planning and control of autonomous vehicles, ensures that the target status information is smoother and more accurate in the timing dimension, and improves the stability of the driving experience.
Smart Images

Figure CN119239656B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving technology, and more particularly, to an object detection method, a vehicle control method, a device, a vehicle, a medium, and a chip. Background Art
[0002] In the field of autonomous driving, the position, speed, and other state information of target objects located in the vehicle's surrounding environment can be detected based on the environmental data around the vehicle. The state information of these target objects can be used for the decision-making planning and control of the vehicle.
[0003] In the related art, the environmental data around the vehicle can be analyzed and processed based on a target object detection model to obtain the state information of the target object. However, the results predicted by the model have problems of inaccuracy and jump, which affect the decision-making planning and control of autonomous vehicles. Summary of the Invention
[0004] In view of this, an embodiment of the present disclosure proposes a new technical solution for object detection.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an object detection method, the method including:
[0006] Based on the environmental data at multiple moments, obtain the initial position information of the target object at each of the multiple moments; wherein, the environmental data includes real-scan data of the external environment of the vehicle, and the target object is a dynamic object in the external environment of the vehicle;
[0007] Based on a preset first detection target, determine the initial velocity information of the target object at each moment according to the initial position information of the target object at multiple moments, so that the initial velocity information conforms to the first detection target; wherein, the initial velocity information includes a first velocity component of the target object along a first coordinate axis direction and a second velocity component of the target object along a second coordinate axis direction, and the first detection target includes that the velocity difference of the target object along the same coordinate axis direction at adjacent moments is less than or equal to a preset velocity threshold;
[0008] Based on the initial position information and the initial velocity information at the multiple moments, determine the target state information of the target object at each moment; wherein, the target state information includes the target position information, the target velocity information, and the target heading angle information of the target object.
[0009] Optionally, the initial position information includes the center point coordinates of the target object, and the center point coordinates include a first coordinate of the center point of the target object on the first coordinate axis and a second coordinate of the center point of the target object on the second coordinate axis;
[0010] The first detection target includes a constraint term determined based on the first coordinate, second coordinate, first velocity component, and second velocity component of the target object at different times.
[0011] Optionally, the constraint term of the first detection target includes at least one of the following:
[0012] The difference between the first coordinate difference of the center point of the target object between the first time and the second time and the first moving distance is less than or equal to a first preset threshold; wherein, the first time is the time earlier than the second time and adjacent to the second time among the multiple times, the first coordinate difference is the difference between the first coordinate of the center point at the first time and the first coordinate at the second time, and the first moving distance is the distance calculated based on the first velocity component and the time difference between the first time and the second time;
[0013] The difference between the second coordinate difference of the center point of the target object between the first time and the second time and the second moving distance is less than or equal to a second preset threshold; wherein, the second coordinate difference is the difference between the second coordinate of the center point at the first time and the second coordinate at the second time, and the second moving distance is the distance calculated based on the second velocity component and the time difference between the first time and the second time;
[0014] The difference between the first velocity component of the target object at the first time and the first velocity component of the target object at the second time is less than or equal to a third preset threshold;
[0015] The difference between the second velocity component of the target object at the first time and the second velocity component of the target object at the second time is less than or equal to a fourth preset threshold.
[0016] Optionally, the initial position information is the three-dimensional position information of the target object, the center point coordinates further include the third coordinate of the center point of the target object on the third coordinate axis, and the initial velocity information further includes the third velocity component of the target object along the direction of the third coordinate axis;
[0017] The constraint term of the first detection target further includes at least one of the following:
[0018] The difference between the third coordinate difference of the center point between the first time and the second time and the third moving distance is less than or equal to a fifth preset threshold; wherein, the third coordinate difference is the difference between the third coordinate of the center point at the first time and the third coordinate at the second time, and the second moving distance is the distance calculated based on the third velocity component and the time difference between the first time and the second time;
[0019] The difference between the third velocity component of the target object at the first moment and the third velocity component of the target object at the second moment is less than or equal to a sixth preset threshold.
[0020] Optionally, determining the target state information of the target object at each moment according to the initial position information and initial velocity information at the multiple moments includes:
[0021] Based on a preset second detection target, determining the target state information of the target object at each moment according to the initial position information and initial velocity information at the multiple moments, so that the target state information conforms to the second detection target;
[0022] Wherein, the second detection target includes constraint conditions determined based on state quantities of the target object, and the state quantities include at least one of the center point coordinates, size, heading angle, velocity, acceleration, and angular velocity of the target object.
[0023] Optionally,
[0024] The second detection target includes: the difference between the predicted target velocity and the initial velocity of the target object at each moment is less than or equal to a preset difference, and the initial velocity is a velocity value calculated according to the initial velocity information.
[0025] Optionally, the second detection target includes constraint conditions corresponding to multiple parts of the target object, where:
[0026] The constraint conditions for different parts of the target object correspond to different confidence levels, and the confidence level of the constraint condition for each part is determined according to the number of position points corresponding to the part in the environmental data; the confidence level is used to represent the weight of the constraint condition for the part among the constraint conditions for multiple parts.
[0027] Optionally,
[0028] The multiple parts of the target object include the head and tail of the target object;
[0029] The constraint condition for each part of the target object is a constraint condition determined based on the center point coordinate of the part, the center point coordinate of the target object, the size of the target object, and the heading angle of the target object.
[0030] Optionally, the second detection target includes at least one of the following:
[0031] The difference in acceleration of the target object at adjacent moments is less than or equal to a preset acceleration threshold;
[0032] The difference in angular velocity of the target object at adjacent moments is less than or equal to a preset angular velocity threshold.
[0033] Optionally, obtaining the initial position information of the target object at each of the multiple moments according to the environmental data at the multiple moments includes:
[0034] Based on the target object detection model, determining at least one candidate object at each moment according to the environmental data and the attribute information of the candidate object at that moment; wherein, the attribute information includes the position information of the candidate object, and the size information and / or heading angle information of the candidate object;
[0035] Performing matching between candidate objects according to the similarity of the attribute information between candidate objects at different moments, determining the attribute information of the same candidate object at each moment, taking the candidate object as the target object, and taking the position information in the attribute information of the candidate object as the initial position information of the target object at each moment.
[0036] According to a second aspect of the embodiments of the present disclosure, a vehicle control method is provided, and the method includes:
[0037] Determining a target object in the external environment where the vehicle is located;
[0038] Controlling the vehicle to travel according to the target state information of the target object;
[0039] Wherein, the target state information is the state information of the target object at each moment determined according to the initial position information and the initial speed information of the target object at multiple moments, the initial position information is the position information of the target object obtained according to the environmental data at multiple moments, the initial speed information is the initial speed information of the target object at each moment determined based on a preset first detection target and according to the initial position information; the environmental data includes real-scan data of the external environment of the vehicle, the target object is a dynamic object in the external environment of the vehicle, the initial speed information includes a first speed component of the target object along a first coordinate axis direction and a second speed component of the target object along a second coordinate axis direction, the first detection target includes that the speed difference of the target object along the same coordinate axis direction at adjacent moments is less than or equal to a preset speed threshold, and the target state information includes the target position information, target speed information and target heading angle information of the target object.
[0040] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory and a processor, the memory is used for storing computer instructions, and the processor is used for calling the computer instructions from the memory to execute the method according to any one of the first aspect and / or the second aspect.
[0041] According to a fourth aspect of the embodiments of the present disclosure, a vehicle is provided, including a memory and a processor. The memory is used to store computer instructions, and the processor is used to call the computer instructions from the memory to execute the method described in any one of the first aspect and / or the second aspect.
[0042] According to a fifth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the first aspect and / or the second aspect is implemented.
[0043] According to a sixth aspect of the embodiments of the present disclosure, a chip is provided, including a processing unit, and the processing unit is configured to execute the method described in any one of the first aspect and / or the second aspect.
[0044] Based on the target detection method provided by the embodiments of the present disclosure, based on the initial position information at multiple moments, through two-stage optimization and solution, the determined target state information is more accurate and smooth in the time series dimension, avoiding the problem of state information jumps of the target object, thereby improving the accuracy of decision-making planning and control of autonomous vehicles.
[0045] Through the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings, other features and advantages of the present disclosure will become clear. Description of the Drawings
[0046] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present disclosure, and together with the description are used to explain the principles of the present disclosure.
[0047] Figure 1 It is a schematic diagram of an intelligent networked system to which the method provided by the embodiments of the present disclosure can be applied.
[0048] Figure 2 It is according to Figure 1 The structural schematic diagram of a vehicle provided by the shown embodiment.
[0049] Figure 3 It is a schematic flowchart of a target detection method provided by the embodiments of the present disclosure.
[0050] Figure 4 It is a schematic flowchart of a target detection method provided by the embodiments of the present disclosure.
[0051] Figure 5 It is a schematic flowchart of a vehicle control method provided by the embodiments of the present disclosure.
[0052] Figure 6 It is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. Detailed Implementation Modes
[0053] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0054] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present disclosure or its application or use.
[0055] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the above technologies, methods, and devices should be regarded as part of the specification.
[0056] In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0057] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0058] First, the application scenarios of the embodiments of the present disclosure will be described.
[0059] Figure 1 is a schematic diagram of an intelligent networked system 100 to which the method provided by the embodiments of the present disclosure can be applied. As Figure 1 shown, the intelligent networked system 100 may include: a vehicle 101, a server 102, and a user terminal 103.
[0060] In some examples, the vehicle 101 may be a vehicle with an autonomous driving function. Among them, autonomous driving is also known as driverless or intelligent driving. A vehicle with an autonomous driving function can perform driving tasks such as environmental perception, decision-making and planning, and control execution. The levels of autonomous driving can refer to the automotive intelligent grading standard formulated by the Society of Automotive Engineers (SAE). For example, the L0 level is manual driving, L1 is assisted driving, L2 is partial autonomous driving, L3 is conditional autonomous driving, L4 is highly autonomous driving, and L5 is fully autonomous driving. The above classification methods for the levels of autonomous driving are only for illustration, and the present disclosure embodiments do not limit the classification criteria and levels of autonomous driving.
[0061] In some examples, the server 102 can be a single server or a distributed server cluster composed of multiple servers, and its deployment method can include a local server or a cloud server. The server 102 can communicate with the vehicle 101 and / or the user terminal 103 based on a communication network, and provide various services for the vehicle 101 and / or the user terminal 103. For example, the server can receive the perception data sent by the vehicle and provide services such as high-precision maps, data analysis, and decision-making planning for the vehicle. Also, for example, the server can receive query instructions or control instructions sent by the user terminal and provide corresponding services for the user.
[0062] In some examples, the user terminal 103 can be any form of electronic device that provides services for users, such as a personal computer, a laptop, a smart tablet, a smart phone, a smart wearable device, etc. The user can interact with the vehicle or the server through the human-machine interaction terminal configured on the vehicle 101, or can also interact with the vehicle or the server through the user terminal 103. For example, query the status and / or parameters of the vehicle through the user terminal, or control the vehicle to execute a set task and / or modify configuration parameters, etc.; wherein, the user terminal runs an application program based on the intelligent networked system to realize the interaction with the vehicle or the server. The application program can be a local application, a web application or a small program, etc., which is not limited here.
[0063] In some examples, the above application program running on the user terminal can provide authentication or authorization services for users. Users who have successfully authenticated and been granted corresponding permissions can query and / or control the vehicle within the granted permissions.
[0064] The vehicle 101, the server 102 and the user terminal 103 can communicate through the communication link provided by the communication network 104. The communication network 104 can include one or more networks of any type. For example, the communication network 104 can include the Internet, a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a public switched telephone network (PSTN), a satellite communication network, Wi-Fi, 2G, 3G, 4G, 5G, 6G, NB-IoT, eMTC, infrared, Bluetooth, NFC and other networks that provide communication, or a combination of the above multiple networks. The communication networks between the vehicle 101 and the server 102, between the user terminal 103 and the server 102, and between the user terminal 103 and the vehicle 101 can be the same or different.
[0065] It should be noted that Figure 1The structure of the intelligent networked system 100 shown is only schematic. The intelligent networked system in the embodiments of the present disclosure is not limited to the above structure and may include more or fewer devices as needed, or the devices may be combined or split. For example, the intelligent networked system may not include a user terminal and / or a server; for another example, the user terminal and the server may be deployed in combination.
[0066] Figure 2 is provided according to Figure 1 the schematic diagram of a vehicle 101 shown in the embodiment. As Figure 2 shown, the vehicle 101 may include a sensing component 1011, a computing platform 1012, an execution component 1013, etc. Among them, the sensing component 1011, the computing platform 1012, and the execution component 1013 may be connected by a bus or other means.
[0067] In some examples, the sensing component 1011 may be used to collect information about the vehicle itself or the outside. The sensing component 1011 may include at least one of a vision sensing unit, a radar, a positioning and navigation unit, an inertial measurement unit (IMU), or other sensing units. Among them, the vision sensor unit may include one or more cameras, the radar may include at least one of a lidar, a millimeter-wave radar, an ultrasonic radar, or other radars, and the positioning and navigation unit may include at least one of a GPS system, a Beidou system, or other global positioning systems.
[0068] In some examples, the computing platform 1012 may include a device with computing capabilities for processing the sensed information collected by the sensing component 1011 to obtain control information and sending corresponding control instructions to the execution component 1013, so that the execution component 1013 performs corresponding actions, thereby realizing the control of the vehicle 101. Exemplarily, the computing platform 1012 may perform actions such as simultaneous localization and mapping (SLAM), path planning, and behavior decision-making on the vehicle, thereby realizing autonomous control of the vehicle. The computing platform 1012 may include at least one processor and at least one memory. Each processor may execute the instructions stored in the memory alone or jointly to implement the method provided in the embodiments of the present disclosure. The processors in the embodiments of the present disclosure may include at least one of a central processing unit (CPU), a graphic process unit (GPU), a neural-network processing unit (NPU), a tensor processing unit (TPU), a data processing unit (DPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a system on chip (SOC), an application specific integrated circuit (ASIC), a microcontroller unit (MCU), or other processors. The memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. In addition to storing instructions, the memory may also store data, such as high-precision maps, path information, the position, direction, speed, etc. of the vehicle. The data stored in the memory may be acquired and used by the processor.
[0069] In some examples, the computing platform of the vehicle may execute computing tasks independently or communicate with a server to complete computing tasks. For example, the computing platform of the vehicle may cooperate with the server to complete corresponding computing tasks.
[0070] The computing platform 1012 can be set in the vehicle 101, and part or all of the computing platform 1012 can also be set in the server corresponding to the vehicle. For example, functions with higher real-time requirements in the computing platform 1012 are set in the vehicle, and functions with lower real-time requirements are set in the server corresponding to the vehicle.
[0071] In some examples, the execution component 1013 is used to perform corresponding actions based on the control of the computing platform 1012, so that the vehicle 101 completes the moving task. The execution component 1013 can include, for example, a power component, a braking component, a transmission component, a steering component, etc.
[0072] It should be noted that Figure 2 The structure of the vehicle 101 shown in is only schematic. The vehicle in the embodiments of the present disclosure is not limited to the above structure, and may include more or fewer components according to needs, or the devices may be combined or split. For example, the vehicle may not include the above-mentioned computing platform. For another example, the vehicle may further include a communication component, an interface component, a multimedia component, an input component, an output component, etc.
[0073] The embodiments of the present disclosure can be applied to the autonomous driving scenario. In the autonomous driving scenario, it is necessary to detect the state information such as the position and speed of the target object in the surrounding environment of the vehicle to assist in decision-making planning and control. The target object can be a dynamic object such as another vehicle or a pedestrian.
[0074] In the related art, the environmental data around the vehicle can be analyzed and processed based on the target object detection model to obtain the state information of the target object. However, the results obtained by model prediction have problems of inaccuracy and jump, which affect the decision-making planning and control of autonomous driving vehicles. For example, due to the limited sensing range of the vehicle's sensors, if the sensors do not fully observe the target object (such as another vehicle) or the quality of the sensor data obtained is crossed, it will cause large jumps in the state information of the detected target object (such as position, physical size, speed, etc.), which will bring wrong signals to the path decision-making planning of the host vehicle and thus affect the normal and smooth driving of the vehicle.
[0075] In view of the problems in the related art, the embodiments of the present disclosure provide a target detection method, which can make the determined target state information more accurate and smooth in the time series dimension through two-stage optimization and solution based on the initial position information at multiple moments, avoiding the problem of jump in the state information of the target object, thereby improving the accuracy of decision-making planning and control of autonomous driving vehicles.
[0076] Figure 3 is a schematic flow chart of a target detection method provided by the embodiments of the present disclosure. The target detection method can be executed by Figure 1 the vehicle and / or server shown in. AsFigure 3 As shown in Figure 3 , the object detection method of this embodiment may include the following steps S310 to S330.
[0077] Step S310: Obtain the initial position information of the target object at each of multiple moments according to the environmental data at multiple moments.
[0078] In some examples, the environmental data may include real-scene scan data of the external environment of the vehicle. Here, the "real scene" in the real-scene scan data refers to the scene in the real world or the actual world, that is, the external environment as the data collection object exists in the external environment in the real world; and the "scan" in the real-scene scan data is intended to reflect that the sensor has scanned the external environment when collecting data for the external environment, rather than restricting the specific method of the sensor collecting data. Among them, the real-scene scan data may include at least one of the image data collected by the camera for the external environment of the vehicle and the point cloud data collected by the radar for the external environment of the vehicle. Exemplarily, the environmental data may include point cloud data, image data, or a combination of both, and the environmental data may also be fusion data obtained by feature fusion based on point cloud data and image data.
[0079] In some examples, the above multiple moments may be a continuous time series. For example, the multiple moments may be a periodic series with a set time interval, and the set time interval may be any pre-set time such as 5 milliseconds, 20 milliseconds, or 100 milliseconds. For example, based on the set time interval, the real-scene scan data of the external environment of the vehicle may be periodically obtained through the sensor as the environmental data. In this way, continuous environmental data in the time dimension can be obtained.
[0080] In some examples, the above target object may be a dynamic object in the external environment of the vehicle. Exemplarily, the target object may include dynamic objects such as other vehicles (vehicles other than the host vehicle), pedestrians, animals, etc.
[0081] Exemplarily, the environmental data may be input into a pre-generated target object detection model to detect the target object and the initial position information of the target object at each moment. If the environmental data is point cloud data, the target object detection model may be a point cloud detection model; if the environmental data is image data, the target object detection model may be an image detection model.
[0082] Step S320: Based on a preset first detection target, determine the initial velocity information of the target object at each moment according to the initial position information of the target object at multiple moments, so that the initial velocity information conforms to the first detection target.
[0083] In some examples, the initial velocity information may include a first velocity component of the target object along the first coordinate axis direction and a second velocity component of the target object along the second coordinate axis direction. The first coordinate axis and the second coordinate axis may be perpendicular to each other. For example, the first coordinate axis is the x-axis, and the first velocity component is used to indicate the velocity of the target object along the x-axis direction; the second coordinate axis is the y-axis, and the second velocity component is used to indicate the velocity of the target object along the y-axis direction.
[0084] In some examples, the first detection target may include that the velocity difference of the target object along the same coordinate axis direction at adjacent moments is less than or equal to a preset velocity threshold. The preset velocity threshold may be a threshold set in advance, such as 0 or a value close to 0. It should be noted that the difference in the embodiments of the present disclosure may be the absolute value of the difference. For example, the velocity difference may be the absolute value of the velocity difference. Optionally, the first detection target may be referred to as a uniform motion constraint or a constant velocity constraint.
[0085] In some examples, step S320 may solve for the optimal solution of the velocity information that meets the first detection target through linear optimization or non-linear optimization methods, and use it as the initial velocity information. The initial velocity information may also be referred to as the initial velocity value.
[0086] Step S330: Determine the target state information of the target object at each moment according to the initial position information and the initial velocity information at multiple moments.
[0087] In some examples, the target state information may include the target position information, the target velocity information, and the target heading angle information of the target object.
[0088] In some examples, the target state information may include at least one of the target position information, the target velocity information, the target heading angle information, the target acceleration information, the target angular velocity information, and the target size information of the target object.
[0089] The target state information of the target object may also be referred to as geometric attributes or attributes.
[0090] It should be noted that the target position information and the initial position information may be the same or different. For example, the above steps can be used to optimize and solve the initial position information to obtain new target position information. Optionally, the target velocity information may include the velocity scalar of the target object.
[0091] In some examples, step S330 can obtain the optimal solution of the state information that meets the detection objective through linear optimization or non-linear optimization, and use it as the target state information. In this way, through step S320 and step S330, two-stage optimization and solution for target object detection are achieved. In the first stage (step S320), initial velocity information (quasi-true value of velocity) is obtained based on the first detection objective. In the second stage (step S330), the initial velocity information can be used as the initial value and the observation item to determine the target state information of the target object (such as target position information, target velocity information, and target heading angle information), so that a more stable and accurate target state information can be obtained.
[0092] In some examples, the above-mentioned multiple moments can be a continuous time series. Based on the environmental data of the target object in the continuous time series, the temporal change situation of the target object can be obtained, so that the target state information of the detected target object is smoother in the temporal dimension.
[0093] Using the method of the above steps S310 to S330 provides a new target detection method. According to the environmental data at multiple moments, the initial position information of the target object at each of the multiple moments is obtained; based on a preset first detection objective, according to the initial position information of the target object at multiple moments, the initial velocity information of the target object at each moment is determined so that the initial velocity information meets the preset first detection objective; then, according to the initial position information and the initial velocity information at multiple moments, the target state information of the target object at each moment is determined; wherein, the target state information includes the target position information, target velocity information, and target heading angle information of the target object. In this way, based on the initial position information at multiple moments, through two-stage optimization and solution, the determined target state information can be more accurate and smoother in the temporal dimension, avoiding the problem of state information jump of the target object, thereby improving the accuracy of decision-making planning and control of autonomous vehicles.
[0094] In some embodiments of the present disclosure, the initial position information of the target object determined in step S310 above can be information in the form of a detection box (BOX). The detection box can be an area surrounding the above-mentioned target object, and the position information (i.e., the initial position information) of the target object can be represented by the detection box. Optionally, the detection box can also represent the size (i.e., the initial size information) and the heading angle (i.e., the initial heading angle information) of the target object. The size can include the length, width, and height of the detection box.
[0095] Exemplarily, the detection box can be a two-dimensional detection box or a three-dimensional detection box (3DBOX). For example, the detection box can be a rectangular box or a cuboid. Taking the three-dimensional detection box as an example, the attribute information of the three-dimensional detection box can include the center point coordinates (x, y, z) of the target object, the size of the target object (length l, width w, height h), the heading angle (i.e., the direction angle) of the target object, the category of the target object (such as vehicle, pedestrian, bicycle, etc.), and the object identifier (ID) of the target object, etc., one or more of the information. Among them, the center point coordinates of the target object determined based on the detection box can be used as the above-mentioned initial position information. Similarly, the size information of the target object determined based on the detection box can also be referred to as the initial size information of the target object, and the heading angle of the target object determined based on the detection box can also be referred to as the initial heading angle information of the target object.
[0096] In some embodiments, the above-mentioned initial position information may include the center point coordinates of the target object, and the center point coordinates include the first coordinate of the center point of the target object on the first coordinate axis and the second coordinate of the second coordinate axis. Exemplarily, the first coordinate axis is the x-axis, and the second coordinate axis is the y-axis. That is, the center point coordinates may include two-dimensional coordinates (x, y).
[0097] In some examples, the first detection target in the above step S320 may include a constraint term determined based on the first coordinate, the second coordinate, the first velocity component, and the second velocity component of the target object at different times.
[0098] In this way, based on the constraint term determined by the above first coordinate, second coordinate, first velocity component, and second velocity component, the constraint and range of the solution can be defined by the constraint term, so that the obtained initial velocity information satisfies the first detection target, that is, the initial velocity information at multiple times is more accurate and smooth in the time sequence dimension, thereby avoiding velocity jumps at adjacent times.
[0099] Exemplarily, the constraint terms of the above first detection target may include at least one of the following first constraint term to fourth constraint term:
[0100] First constraint term: The difference between the first coordinate difference of the center point of the target object between the first moment and the second moment and the first moving distance is less than or equal to the first preset threshold; wherein, the first moment is the moment earlier than the second moment and adjacent to the second moment among multiple moments. For example, the first moment is the (i - 1)-th moment, the second moment is the i-th moment, the first coordinate difference is the difference between the first coordinate of the center point at the first moment and the first coordinate at the second moment, the first moving distance is the distance calculated based on the first velocity component and the time difference between the first moment and the second moment, and the first preset threshold may be 0 or a preset value close to 0.
[0101] Taking the first preset threshold as 0 as an example, the first constraint term can be expressed by the following formula (1):
[0102]
[0103] Wherein, represents the first coordinate of the center point of the target object at the i-th moment (such as the second moment above), represents the first coordinate of the center point of the target object at the (i - 1)-th moment (such as the first moment above), represents the first velocity component of the target object on the first coordinate axis (x-axis) at the i-th moment, and Δt represents the time difference between the i-th moment and the (i - 1)-th moment.
[0104] Second constraint term: The difference between the second coordinate difference of the center point of the target object between the first moment and the second moment and the second moving distance is less than or equal to the second preset threshold; wherein, the second coordinate difference is the difference between the second coordinate of the center point at the first moment and the second coordinate at the second moment, and the second moving distance is the distance calculated based on the second velocity component and the time difference between the first moment and the second moment, and the second preset threshold can be 0 or a preset value close to 0.
[0105] Taking the second preset threshold as 0 as an example, the second constraint term can be expressed by the following formula (2):
[0106]
[0107] Wherein, represents the second coordinate of the center point of the target object at the i-th moment (such as the second moment above), represents the second coordinate of the center point of the target object at the (i - 1)-th moment (such as the first moment above), represents the second velocity component of the target object on the second coordinate axis (x-axis) at the i-th moment, and Δt represents the time difference between the i-th moment and the (i - 1)-th moment.
[0108] Third constraint term: The difference between the first velocity component of the target object at the first moment and the first velocity component of the target object at the second moment is less than or equal to the third preset threshold, and the third preset threshold can be 0 or a preset value close to 0.
[0109] Taking the third preset threshold as 0 as an example, the third constraint term can be expressed by the following formula (3):
[0110]
[0111] Wherein, represents the first velocity component of the target object at the i-th moment (such as the second moment above), Represents the first velocity component of the target object at the (i - 1)-th moment (e.g., the first moment mentioned above).
[0112] Fourth constraint term: The difference between the second velocity component of the target object at the first moment and the second velocity component of the target object at the second moment is less than or equal to a fourth preset threshold, and the fourth preset threshold can be 0 or a preset value close to 0.
[0113] Taking the fourth preset threshold as 0 as an example, this fourth constraint term can be represented by the following formula (4):
[0114]
[0115] Wherein, Represents the second velocity component of the target object at the i-th moment (e.g., the second moment mentioned above), Represents the second velocity component of the target object at the (i - 1)-th moment (e.g., the first moment mentioned above).
[0116] It should be noted that the above first to fourth constraint terms are constructed alone or in any combination to obtain the first detection target. By way of example, the first detection target may include the first constraint term and the third constraint term, or, the second constraint term and the fourth constraint term, or, the first constraint term and the second constraint term, or, the third constraint term and the fourth constraint term, or, the whole of the first to fourth constraint terms.
[0117] In this way, based on the above first to fourth constraint terms, the target object can be constrained to move at a uniform speed on the x-axis and / or y-axis, thereby avoiding jumps in speed at adjacent moments, making the initial velocity information at multiple moments more accurate and smoother in the time series dimension.
[0118] In some examples, the above initial position information may be the three-dimensional position information of the target object, the above center point coordinates may further include the third coordinate of the center point of the target object on the third coordinate axis, and the above initial velocity information may further include the third velocity component of the target object on the third coordinate axis. The third coordinate axis may be perpendicular to the above first coordinate axis and the second coordinate axis pairwise to form a three-dimensional coordinate axis. For example, the first coordinate axis may be the x-axis, the second coordinate axis may be the y-axis, and the third coordinate axis may be the z-axis. That is, the center point coordinates may be three-dimensional coordinates (x, y, z), and the above initial velocity information may include the velocity components of the target object in the three directions of the x-axis, y-axis, and z-axis.
[0119] In some examples, the constraint terms of the first detection target may further include at least one of the following fifth to sixth constraint terms:
[0120] The fifth constraint item: The difference between the third coordinate difference of the center point between the first moment and the second moment and the third movement distance is less than or equal to a fifth preset threshold; wherein, the third coordinate difference is the difference between the third coordinate of the center point at the first moment and the third coordinate at the second moment, the second movement distance is the distance calculated based on the third velocity component and the time difference between the first moment and the second moment, and the first preset threshold can be 0 or a preset value close to 0.
[0121] Taking the fifth preset threshold as 0 as an example, the fifth constraint item can be expressed by the following formula (5):
[0122]
[0123] Wherein, represents the third coordinate of the center point of the target object at the i-th moment (such as the second moment above), represents the third coordinate of the center point of the target object at the i-1-th moment (such as the first moment above), represents the third velocity component of the target object on the third coordinate axis (z-axis) at the i-th moment, and Δt represents the time difference between the i-th moment and the i-1-th moment.
[0124] The sixth constraint item: The difference between the third velocity component of the target object at the first moment and the third velocity component of the target object at the second moment is less than or equal to a sixth preset threshold, and the sixth preset threshold can be 0 or a preset value close to 0.
[0125] Taking the sixth preset threshold as 0 as an example, the sixth constraint item can be expressed by the following formula (6):
[0126]
[0127] Wherein, represents the third velocity component of the target object at the i-th moment (such as the second moment above), represents the third velocity component of the target object at the i-1-th moment (such as the first moment above).
[0128] In this way, based on the above fifth to sixth constraint items, the uniform motion of the target object on the z-axis can be constrained, thereby avoiding the jump of the speed at adjacent moments, and making the initial speed information at multiple moments more accurate and smooth in the time sequence dimension.
[0129] In some embodiments of the present disclosure, the above step S330 can be implemented based on the following method:
[0130] Based on a preset second detection target, determine the target state information of the target object at each moment according to the initial position information and initial speed information at multiple moments, so that the target state information conforms to the second detection target.
[0131] In some examples, the second detection target may include a constraint condition determined based on a state quantity of the target object. The state quantity may include at least one of the center point coordinates, size, heading angle, speed, acceleration, and angular velocity of the target object. The size may include at least one of the length, width, and height of the target object. The state quantity may also be referred to as an optimization state quantity, that is, the optimization state quantity set in the optimization model corresponding to the second detection target. The optimization model may be a multi-constraint optimization model.
[0132] Exemplarily, the state quantity X may be expressed as the following formula (7):
[0133]
[0134] Wherein, X represents the state quantity, x represents the first coordinate of the center point of the target object on the first coordinate axis (x-axis), y represents the second coordinate of the center point of the target object on the second coordinate axis (y-axis), θ represents the heading angle of the target object, v represents the speed of the target object, Ω represents the angular velocity of the target object, and a represents the acceleration of the target object.
[0135] It should be noted that the above formula (7) is only an example. The state quantity may include more or fewer parameters. For example, it may also include at least one of the size of the target object: length l, width w, and height h.
[0136] In some examples, the above second detection target may include at least one of the following:
[0137] The acceleration difference of the target object at adjacent moments is less than or equal to a preset acceleration threshold (for example, 0);
[0138] The angular velocity difference of the target object at adjacent moments is less than or equal to a preset angular velocity threshold (for example, 0).
[0139] Exemplarily, constant acceleration and constant angular velocity (constant turn rate) may be used as the basic constraints of the second detection target to determine the constraint corresponding to the second detection target.
[0140] Taking the preset acceleration threshold and the preset angular velocity threshold both being 0 as an example, the second detection target may include a constraint term represented by the following formula (8):
[0141]
[0142] Wherein, the (i - 1)-th moment is the previous moment adjacent to the i-th moment among a plurality of moments, and x i represents the first coordinate of the center point of the target object at the i-th moment (for example, the second moment), y i represents the second coordinate of the center point of the target object at the i-th moment, θi Denotes the heading angle of the target object at the \(i\)-th moment, \(v\) i Denotes the speed of the target object at the \(i\)-th moment, \(\Omega\) i Denotes the angular velocity of the target object at the \(i\)-th moment, \(a\) i Denotes the acceleration of the target object at the \(i\)-th moment, \(x\) i-1 Denotes the first coordinate of the center point of the target object at the \((i - 1)\)-th moment (e.g., the coordinate on the first coordinate axis \(x\)-axis), \(y\) i-1 Denotes the second coordinate of the center point of the target object at the \((i - 1)\)-th moment (e.g., the coordinate on the second coordinate axis \(y\)-axis), \(\theta\) i-1 Denotes the heading angle of the target object at the \((i - 1)\)-th moment, \(v\) i-1 Denotes the speed of the target object at the \((i - 1)\)-th moment, \(\Omega\) i-1 Denotes the angular velocity of the target object at the \((i - 1)\)-th moment, \(a\) i-1 Denotes the acceleration of the target object at the \((i - 1)\)-th moment, \(\delta t\) represents the time difference between the \(i\)-th moment and the \((i - 1)\)-th moment, Denotes the displacement of the target object along the \(x\)-axis within the time \(\delta t\), Denotes the displacement of the target object along the \(y\)-axis within the time \(\delta t\), \(\Omega\) i-1 \(\cdot\delta t\) represents the angular change amount of the target object within the time \(\delta t\), \(a\) i \(\cdot\delta t\) represents the speed change amount of the target object within the time \(\delta t\).
[0143] In some examples, the above second detection target may include: the difference between the predicted target speed and the initial speed of the target object at each moment is less than or equal to a preset difference, and the initial speed is a speed value calculated based on the initial speed information.
[0144] Taking the preset difference as 0 as an example, the second detection target may include the constraint term represented by the following formula (9):
[0145]
[0146] Among them, the subscript \(i\) may represent any moment among multiple moments, Denotes the first speed component of the target object on the first coordinate axis \(x\)-axis at the \(i\)-th moment, that is Denotes the first speed component in the initial speed information determined in step S320, Denotes the second speed component of the target object on the second coordinate axis \(y\)-axis at the \(i\)-th moment, that is Denotes the second speed component in the initial speed information determined in step S320, \(v\) i Denotes the target speed scalar of the target object at the \(i\)-th moment, that is, the scalar representation of the target speed information of the optimization target in step S330.
[0147] In this way, using the initial speed information determined in step S302 to constrain the target speed of the vehicle not only enables the optimized speed to converge quickly, but also, in combination with other constraint terms, makes other optimization results in the target state information more accurate.
[0148] In some examples, the above-mentioned second detection target may further include constraint conditions corresponding to multiple parts of the target object, where: the constraint conditions for different parts of the target object correspond to different confidence levels, and the confidence level of the constraint condition for each part is determined according to the number of position points corresponding to that part in the environmental data; the confidence level is used to represent the weight of the constraint condition of the part among the multiple part constraint conditions. Optionally, the confidence level may be any value between 0 and 1, and the greater the confidence level, the greater the weight.
[0149] In some examples, the more the number of position points of a certain part, the higher the confidence level of that part, that is, the higher the weight.
[0150] In some examples, the multiple parts of the above-mentioned target object include the head and tail of the target object, and the constraint conditions for each part of the target object may be constraint conditions determined based on the center point coordinates of the part, the center point coordinates of the target object, the size of the target object, and the heading angle of the target object.
[0151] Exemplarily, the constraint conditions corresponding to the head and tail of the target object included in the second detection target may be represented by the following formulas (10) and (11):
[0152]
[0153] where the subscript i may represent any moment among multiple moments, x i represents the first coordinate (x-axis coordinate) of the center point of the target object at the i-th moment, y i represents the second coordinate (y-axis coordinate) of the center point of the target object at the i-th moment, l i represents the length of the target object at the i-th moment, θ i represents the heading angle of the target object at the i-th moment, and represent the observed values of the center point of the head of the target object (i.e., the x-axis coordinate and y-axis coordinate of the center point of the head of the detection box detected based on the target object detection model), and represent the observed values of the center point of the tail of the target object (the x-axis coordinate and y-axis coordinate of the center point of the tail of the detection box detected based on the target object detection model).
[0154] The above and It can be expressed by the following formulas (12) and (13):
[0155]
[0156] Wherein, represents the x-axis coordinate of the center point of the head of the target object at the i-th moment, represents the x-axis coordinate of the center point of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the center point x-axis coordinate), represents the y-axis coordinate of the center point of the head of the target object at the i-th moment, represents the y-axis coordinate of the center point of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the center point y-axis coordinate), represents the length in the initial size of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the physical size length), represents the initial heading angle of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the heading angle).
[0157] The above and can be expressed by the following formulas (14) and (15):
[0158]
[0159] Wherein, represents the x-axis coordinate of the center point of the tail of the target object at the i-th moment, represents the x-axis coordinate of the center point of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the center point x-axis coordinate), represents the y-axis coordinate of the center point of the tail of the target object at the i-th moment, represents the y-axis coordinate of the center point of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the center point y-axis coordinate), represents the length in the initial size of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the physical size length), represents the initial heading angle of the target object detected based on the target object detection model at the i-th moment (i.e., the observed value of the heading angle).
[0160] Based on the center point of the head, the number of position points corresponding to the head can be determined, and then the confidence level of the constraint condition corresponding to the head can be determined; similarly, based on the center point of the tail, the number of position points corresponding to the tail can be determined, and then the confidence level of the constraint condition corresponding to the tail can be determined.
[0161] Exemplarily, based on the detection box of the target object (including initial position information, initial size information, and initial heading angle information), multiple position points located within the detection box can be screened out from the point set of the original environmental data (such as the point set of point cloud data or the pixel set of image data) as the first point set P i 。
[0162] According to the coordinates of the above-mentioned head center point, position points whose distance from the head is less than or equal to a specific distance are screened out from the first point set as the position point set corresponding to the head, thereby obtaining the number of position points corresponding to the head. This specific distance can be the width w of the target object or any other arbitrarily set value.
[0163] Furthermore, the confidence level of the constraint condition corresponding to the head can be determined based on the number of position points corresponding to the head. For example, the confidence level of the head can be calculated through the following formula (16):
[0164]
[0165] wherein, represents the confidence level of the constraint condition corresponding to the head, F i represents the position point set corresponding to the head, |F i | represents the number of position points corresponding to the head, M is a preset parameter, used to indicate that if the number of position points corresponding to the head is greater than or equal to M, the confidence level of the head is the maximum value 1, and M can be any positive integer set in advance, such as 20 or 100.
[0166] According to the coordinates of the above-mentioned tail center point, position points whose distance from the tail center point is less than or equal to a specific distance are screened out from the first point set as the position point set corresponding to the tail, thereby obtaining the number of position points corresponding to the tail. This specific distance can be the width w of the target object.
[0167] Furthermore, the confidence level of the constraint condition corresponding to the tail can be determined based on the number of position points corresponding to the tail. For example, the confidence level of the tail can be calculated through the following formula (17):
[0168]
[0169] wherein, represents the confidence level of the constraint condition corresponding to the tail, B i represents the position point set corresponding to the tail, |B i | represents the number of position points corresponding to the tail, M is a preset parameter, used to indicate that if the number of position points corresponding to the tail is greater than or equal to M, the confidence level of the tail is the maximum value 1, and M can be any positive integer set in advance, such as 20 or 100.
[0170] In this way, when constructing the constraint conditions corresponding to the second detection target, different confidence levels can be set for different parts of the target object, which can further improve the accuracy of the target state information obtained based on the constraint conditions. Taking the target object as another vehicle and the environmental data as point cloud data as an example, it is illustrated as follows: When another vehicle is directly in front of the host vehicle, only the tail of the other vehicle can be observed and the point cloud is also distributed at the tail of the other vehicle. At this time, the detection box (such as 3DBOX) detected based on the target object detection model (such as a deep learning model) fits very accurately to the tail of the other vehicle, but the fitting accuracy of its physical size length and the center point of the head of the other vehicle is very low. On the contrary, if the other vehicle is directly behind the host vehicle, only the head of the other vehicle can be observed and the point cloud is also distributed at the head of the other vehicle. At this time, the detection box detected based on the target object detection model fits very accurately to the head of the other vehicle, but the fitting accuracy of its physical size length and the center point of the tail of the other vehicle is very low. Therefore, when constructing the constraint, it is necessary to consider its observation covariance at the same time. By setting the confidence levels of different parts, the entire target detection result can be made smoother and the detection accuracy can be higher.
[0171] In some examples, the above-mentioned second detection target may further include an observation constraint term based on the physical size length. For example, this observation constraint term based on the physical size length can be represented by the following formula (18):
[0172]
[0173] where l i represents the target size length (i.e., the state quantity of the physical size length) in the target state information of the target object, represents the length in the initial size of the target object detected based on the target object detection model (i.e., the observed quantity of the physical size length).
[0174] In some examples, the above-mentioned second detection target may further include an observation constraint term based on the heading angle. For example, this observation constraint term based on the physical size length can be represented by the following formula (19):
[0175]
[0176] where θ i represents the target heading angle (i.e., the state quantity of the heading angle) in the target state information of the target object, represents the initial heading angle of the target object detected based on the target object detection model (i.e., the observed quantity of the heading angle).
[0177] In this way, the accuracy of the detection result can be further increased and the precision of the target detection result can be improved.
[0178] In some examples, in the above step S330, the one or more constraint items included in the second detection target can be added to the same non-linear optimization problem, and then a preset non-linear optimization algorithm (such as a method based on Gauss-Newton iteration) is used to iteratively update and solve the target state information of the target object (such as one or more of the center point position, size, speed, acceleration, and angular velocity of the target object), so that the state information of the target object that is smoother and more accurate can be calculated, improving the accuracy and smoothness of target detection.
[0179] In some embodiments of the present disclosure, the above step S310 may include the following steps S3101 and S3102.
[0180] Step S3101, based on the target object detection model, determine at least one candidate object at each moment and the attribute information of the candidate object at that moment according to the environmental data at multiple moments.
[0181] Wherein, the attribute information includes the position information of the candidate object, and the size information and / or heading angle information of the candidate object.
[0182] In some examples, the position information may be position information based on the world coordinate system.
[0183] Exemplarily, based on the target object detection model, the position information of at least one candidate object in the vehicle coordinate system at each moment can be determined; according to the attitude information of the vehicle, the position information in the vehicle coordinate system is converted into the position information in the world coordinate system, so that the position information in the world coordinate system is used as the position information in the attribute information of the target object. Among them:
[0184] The vehicle coordinate system can also be referred to as the vehicle's own coordinate system or the body coordinate system, which is a coordinate system established with the vehicle as the reference point. The origin of the vehicle coordinate system can be the center point of the vehicle, and the directions of the coordinate axes can be determined according to the moving direction of the vehicle. For example, the vehicle coordinate system includes an x-axis, a y-axis, and a z-axis. The x-axis can point to the forward direction of the vehicle, the y-axis can point to the left side of the vehicle, and the z-axis can be perpendicular to the ground where the vehicle is located and point upward of the vehicle. Through this vehicle coordinate system, the movement of the vehicle relative to itself and the state of the vehicle's surrounding environment can be described, such as the speed, acceleration, and steering angle of the vehicle.
[0185] The world coordinate system can also be called the global coordinate system, which can establish a global reference framework to describe the global position and orientation of the vehicle in the map or the entire environment. The origin of this world coordinate system can be a selected fixed point, such as a certain geographical location or a specific reference point on the map, and the directions of the coordinate axes can also be pre-set directions, such as the absolute directions of north, south, east, and west based on the Earth. For example, the x-axis of the world coordinate system can point to the due north direction on the Earth's surface, the y-axis can point to the due east direction on the Earth's surface, and the z-axis can be perpendicular to the Earth's surface and point above the ground.
[0186] During the driving process of the vehicle, since the origin of the vehicle coordinate system may be different at different times, after converting the position information of the target object from the vehicle coordinate system to the world coordinate system, it is beneficial to align the position information at different times to the same coordinate origin to facilitate the processing of multiple position information.
[0187] In some examples, the pose information of the vehicle can include the position and heading angle of the vehicle. Optionally, the pose information of the vehicle can be obtained based on SLAM (Simultaneous Localization and Mapping), for example, self-vehicle localization can be achieved based on SLAM, so as to obtain the position and orientation angle of the self-vehicle and obtain the pose information of the self-vehicle.
[0188] Step S3102, match the candidate objects according to the similarity of the attribute information between the candidate objects at different times, determine the attribute information of the same candidate object at each time, and use this candidate object as the target object, and use the position information in the attribute information of this candidate object as the initial position information of the target object at each time.
[0189] Exemplarily, the environmental data at each time can be called a frame of environmental data, the attribute information of the candidate object can be represented in the form of a detection box, and the similarity of the attribute information between the candidate objects at different times can be the intersection over union. In this way, the detection boxes of the candidate objects in each frame of environmental data can be matched by calculating the intersection over union (IOU) of the detection boxes pairwise to obtain the detection box sequence of each candidate object at each time, so as to obtain the initial position information of the target object (i.e., the candidate object) at each time.
[0190] Taking the environmental data at two consecutive moments as an example, three candidate objects A1, B1, and C1 are detected in the environmental data at the first moment, and two candidate objects A2 and B2 are detected in the environmental data at the second moment. The intersection over union (IoU) between each pair of candidate objects is calculated based on their attribute information (such as detection boxes) to determine which candidate objects in the environmental data at these two moments represent the same object. For example, by comparing their IoU, it can be determined that A1 and A2 represent the same object, and B1 and B2 represent another object. Thus, the attribute information of target objects A and B at the two moments can be obtained, and the initial position information of target objects A and B at the two moments can also be obtained.
[0191] In this way, based on the environmental data at multiple moments, the initial position information of the target object at each moment among multiple moments can be accurately obtained.
[0192] In some embodiments of the present disclosure, the above environmental data can be obtained in any of the following ways:
[0193] In some examples, the above environmental data can be collected by the vehicle through its own sensors. For example, the environmental data can include point cloud data collected based on the vehicle's radar (such as lidar), image data collected based on the vehicle's vision acquisition unit (such as a camera), or fusion data obtained by feature fusion based on the point cloud data and the image data.
[0194] In other examples, the environmental data can also be collected by other sensors outside the vehicle. For example, the environmental data can include image data of the vehicle's surrounding environment collected based on sensing devices (such as cameras) installed on the road.
[0195] In still other examples, the environmental data can also be jointly collected by the vehicle's own sensors and other sensors outside the vehicle. For example, the environmental data can include point cloud data collected through the vehicle's radar, image data of the vehicle's surrounding environment collected through sensing devices installed on the road, or fusion data obtained by feature fusion based on the point cloud data and the image data.
[0196] In this way, the environmental data can be obtained in multiple ways to improve the accuracy of target detection.
[0197] It should be noted that the above environmental data can also be other data collected by other means, and the specific categories and acquisition methods of the environmental data in the embodiments of the present disclosure are not limited.
[0198] Figure 4 is a schematic flowchart of another target detection method provided by the embodiments of the present disclosure. This target detection method can be performed by Figure 1executed by the vehicle and / or server shown. As Figure 4 As shown, the object detection method of this embodiment may include the following steps S410 to S460.
[0199] Step S410: Determine the detection frame at each moment according to the environmental data at multiple moments.
[0200] Exemplarily, based on an object detection model (such as a point cloud detection model for deep learning), at least one detection frame at each moment can be detected in the environmental data at multiple moments collected by the vehicle. The detection frame can enclose a candidate object. For example, the detection frame can include the candidate object and the attribute information of the candidate object at this moment, that is, the attribute information of the candidate object can be realized based on the detection frame.
[0201] Step S420: Determine the pose information of the vehicle at each moment among multiple moments.
[0202] Exemplarily, the pose of the ego-vehicle movement at the timestamp of each frame of environmental data (i.e., each moment) can be calculated through SLAM technology as the pose information of the vehicle. The pose information can include the position and heading angle of the vehicle, and the position can be the position coordinates of the vehicle in the world coordinate system.
[0203] Step S430: According to the vehicle body pose, convert the coordinates of the detection frame into the coordinates in the world coordinate system.
[0204] Step S440: Perform matching and tracking on the detection frames at multiple moments to determine the detection frame of the target object at each moment, and the detection frame includes the initial position information of the target object.
[0205] Exemplarily, the detection frames at multiple moments can be matched in the world coordinate system by calculating the intersection over union (IOU) of the detection frames pairwise to obtain the detection frames of each target object at each moment.
[0206] In some examples, the candidates can be matched according to the similarity of the attribute information between the candidate objects at different moments, the attribute information of the same candidate object at each moment can be determined, and the candidate object is used as the target object, and the position information in the attribute information of the candidate object is used as the initial position information of the target object at each moment. The detection frame can include the initial position information of the target object.
[0207] Step S450: Based on a preset first detection target, determine the initial velocity information of the target object at each moment according to the initial position information of the target object at multiple moments, so that the initial velocity information meets the first detection target.
[0208] Among them, the specific implementation of step S450 can refer to the description of step S320 in the foregoing embodiments of the present disclosure, and will not be elaborated here.
[0209] Step S460: Determine the target state information of the target object at each moment according to the initial position information and initial velocity information at multiple moments.
[0210] In some examples, the target state information includes at least one of the target position information, target velocity information, target heading angle information, target acceleration information, target angular velocity information, and target size information of the target object.
[0211] Among them, the specific implementation of step S460 can refer to the description of step S330 in the foregoing embodiments of the present disclosure, and will not be elaborated here.
[0212] Using the above method, by using the detection result based on the target object detection model (such as the initial position information) and the vehicle body pose, and combining with the kinematic model,
[0213] Two-stage optimization is adopted for solution, that is, in the first stage of optimization, the initial velocity information (the quasi-true value of the velocity) of the target object is obtained. This initial velocity information can be used as the initial value and the observation term for the second stage of optimization. Then, in the second stage, based on the constant acceleration constraint and / or the constant angular velocity constraint in the kinematic model, the center points of different parts of the target object are used as the observation terms, and different confidence levels are calculated respectively. And by fusing the heading angle constraint and the physical size length constraint, the overall optimization of each state information of the target object is carried out, and more stable and accurate target state information is output, so that the process of the host vehicle planning and controlling based on the target state information of the target object is also more stable, and the user obtains a more comfortable experience.
[0214] The above two-stage optimization target detection method in the embodiments of the present disclosure has smaller error and higher accuracy than the single-stage optimization (only considering the constant velocity model) target detection method in the related art. Specifically, under the same test conditions (the vehicle and road conditions are the same), the error comparison between the two methods is shown in the following table:
[0215]
[0216] In the above table, the absolute error of longitudinal speed, the absolute error of lateral speed, the absolute error of longitudinal position, the absolute error of lateral position, and the absolute error of heading angle are all the differences between the detection results obtained based on the corresponding object detection method and the actual results. Among them: the longitudinal speed may be the first speed component of the target speed of the target object along the first coordinate axis, the lateral speed may be the second speed component of the target speed of the target object along the second coordinate axis, the longitudinal position may be the first coordinate of the target position of the target object on the first coordinate axis, the lateral position may be the second coordinate of the target position of the target object on the second coordinate axis, and the heading angle may be the target heading angle of the target object.
[0217] In some embodiments of the present disclosure, the object detection method can be used in data annotation. For example, a plurality of environmental data can be obtained by a data acquisition vehicle, and based on the object detection method, the environmental data can be annotated to determine the target object and the target state information of the target object therein. The annotated data can be used to train the model of an autonomous vehicle to improve the accuracy of the vehicle's detection of the target object.
[0218] In some other embodiments of the present disclosure, the object detection method can be used in vehicle control. For example, the target state information of the target object can be determined based on the object detection method, and the vehicle can be controlled to travel based on the target state information.
[0219] Figure 5 is a schematic flowchart of a vehicle control method provided by an embodiment of the present disclosure. The vehicle control method can be executed by Figure 1 the vehicle and / or server shown. As Figure 5 shown, the vehicle control method of this embodiment may include:
[0220] Step S510, determining a target object in the external environment where the vehicle is located.
[0221] Step S520, controlling the vehicle to travel according to the target state information of the target object.
[0222] Among them, the target state information is the state information of the target object at each moment determined based on the initial position information and initial velocity information of the target object at multiple moments. The initial position information is the position information of the target object obtained based on the environmental data at multiple moments, and the initial velocity information is the initial velocity information of the target object at each moment determined based on a preset first detection target and according to the initial position information. The environmental data includes real-scan data of the external environment of the vehicle, the target object is a dynamic object in the external environment of the vehicle, the initial velocity information includes a first velocity component of the target object along the first coordinate axis direction and a second velocity component of the target object along the second coordinate axis direction, the first detection target includes that the velocity difference of the target object along the same coordinate axis direction at adjacent moments is less than or equal to a preset velocity threshold, and the target state information includes the target position information, target velocity information, and target heading angle information of the target object.
[0223] The acquisition method of the target state information of the target object in this embodiment can refer to the description in the foregoing embodiments of the present disclosure, and will not be elaborated here.
[0224] In this way, since the detection result of the target object is smoother and more accurate, the process of planning and controlling the vehicle is more stable, providing a more comfortable experience for the user.
[0225] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 6 shown, the electronic device 1000 may include a memory 1010 and a processor 1020. The memory 1010 may be used to store computer instructions, and the processor 1020 may be used to call the computer instructions from the memory 1010 to execute all or part of the steps of any method in the foregoing embodiments of the present disclosure. Among them, the processor may be one or more, and the one or more processors may execute instructions alone or jointly. The memory may also be one or more, and the one or more memories may store the above computer instructions alone or jointly.
[0226] An embodiment of the present disclosure also provides a vehicle, which may include a memory and a processor. The memory may be used to store computer instructions, and the processor may be used to call the computer instructions from the memory to execute all or part of the steps of any method in the foregoing embodiments of the present disclosure. Among them, the processor may be one or more, and the one or more processors may execute instructions alone or jointly. The memory may also be one or more, and the one or more memories may store the above computer instructions alone or jointly.
[0227] The vehicle in the foregoing embodiments of the present disclosure may be an electric vehicle, a hybrid vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle may be an autonomous vehicle or a non-autonomous vehicle. By way of example, the vehicle provided in this embodiment may be Figure 1 or Figure 2 the vehicle shown.
[0228] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, all or part of the steps of any one of the methods in the foregoing embodiments of the present disclosure are implemented. Optionally, the computer-readable storage medium may be a non-transitory storage medium, but is not limited thereto, and it may also be a transitory storage medium.
[0229] The embodiments of the present disclosure also provide a chip, which may include a processing unit, and the processing unit may be used to execute all or part of the steps of any one of the methods in the foregoing embodiments of the present disclosure. The chip may be in the form of an application specific integrated circuit (ASIC), a system on chip (SOC), a field programmable gate array (FPGA), etc., and this embodiment does not limit this.
[0230] The embodiments of the present disclosure also provide a computer program product, which may include a computer program, and when the computer program is executed by a processor, any one of the methods in the foregoing embodiments of the present disclosure may be implemented.
[0231] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for causing a processor to implement any one of the methods in the foregoing embodiments of the present disclosure are uploaded.
[0232] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0233] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0234] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, which may include object - oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0235] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0236] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0237] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0238] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It should be noted that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.
[0239] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skill in the technical field to understand the embodiments disclosed herein. The scope of the present disclosure is defined by the appended claims.
Claims
1. A target detection method, characterized in that, The method includes: Obtaining initial position information of a target object at each of the multiple moments according to environmental data at the multiple moments; wherein, the environmental data includes real-scene scanning data of the external environment of the vehicle, and the target object is a dynamic object in the external environment of the vehicle; Based on a preset first detection target, determining initial velocity information of the target object at each moment according to the initial position information of the target object at the multiple moments, so that the initial velocity information conforms to the first detection target; Based on a preset second detection target, determining target state information of the target object at each moment according to the initial position information and the initial velocity information at the multiple moments, so that the target state information conforms to the second detection target; wherein, the target state information includes target position information, target velocity information and target heading angle information of the target object, the target position information is obtained according to the initial position information, the second detection target includes constraint conditions corresponding to multiple parts of the target object, and the constraint conditions of different parts of the target object correspond to different confidence levels, and the confidence level of the constraint condition of each part is determined according to the number of position points corresponding to the part in the environmental data; the confidence level is used to represent the weight of the constraint condition of the part among the constraint conditions of multiple parts.
2. The method according to claim 1, characterized in that, The initial position information includes the center point coordinates of the target object, and the center point coordinates include a first coordinate of the center point of the target object on a first coordinate axis and a second coordinate on a second coordinate axis; The first detection target includes constraint terms determined based on the first coordinate, the second coordinate, the first velocity component and the second velocity component of the target object at different moments.
3. The method according to claim 2, characterized in that, The constraint terms of the first detection target include at least one of the following: The difference between the first coordinate difference of the center point of the target object between the first moment and the second moment and the first movement distance is less than or equal to a first preset threshold; wherein, the first moment is the moment earlier than and adjacent to the second moment among the multiple moments, the first coordinate difference is the difference between the first coordinate of the center point at the first moment and the first coordinate at the second moment, and the first movement distance is the distance calculated based on the first velocity component and the time difference between the first moment and the second moment; The difference between the second coordinate difference of the center point of the target object between the first moment and the second moment and the second movement distance is less than or equal to a second preset threshold; wherein, the second coordinate difference is the difference between the second coordinate of the center point at the first moment and the second coordinate at the second moment, and the second movement distance is the distance calculated based on the second velocity component and the time difference between the first moment and the second moment; The difference between the first velocity component of the target object at the first moment and the first velocity component of the target object at the second moment is less than or equal to a third preset threshold; The difference between the second velocity component of the target object at the first moment and the second velocity component of the target object at the second moment is less than or equal to a fourth preset threshold.
4. The method according to claim 3, characterized in that, The initial position information is the three-dimensional position information of the target object. The center point coordinates further include the third coordinate of the center point of the target object on the third coordinate axis. The initial velocity information further includes the third velocity component of the target object along the direction of the third coordinate axis. The constraint terms of the first detection target further include at least one of the following: The difference between the third coordinate difference of the center point between the first moment and the second moment and the third moving distance is less than or equal to a fifth preset threshold. Wherein, the third coordinate difference is the difference between the third coordinate of the center point at the first moment and the third coordinate at the second moment, and the second moving distance is the distance calculated based on the third velocity component and the time difference between the first moment and the second moment. The difference between the third velocity component of the target object at the first moment and the third velocity component of the target object at the second moment is less than or equal to a sixth preset threshold.
5. The method according to claim 1, wherein The second detection target further includes a constraint condition determined based on the state quantity of the target object. The state quantity includes at least one of the center point coordinates, size, heading angle, velocity, acceleration, and angular velocity of the target object.
6. The method according to claim 5, wherein The second detection target further includes: the difference between the target velocity predicted for each moment of the target object and the initial velocity is less than or equal to a preset difference, and the initial velocity is the velocity value calculated based on the initial velocity information.
7. The method according to claim 1, wherein The multiple parts of the target object include the head and tail of the target object. The constraint condition for each part of the target object is a constraint condition determined based on the center point coordinates of the part, the center point coordinates of the target object, the size of the target object, and the heading angle of the target object.
8. The method according to claim 5, wherein The second detection target further includes at least one of the following: The acceleration difference of the target object at adjacent moments is less than or equal to a preset acceleration threshold. The angular velocity difference of the target object at adjacent moments is less than or equal to a preset angular velocity threshold.
9. The method according to any one of claims 1 to 8, characterized in that, The obtaining of the initial position information of the target object at each of the multiple moments according to the environmental data of the multiple moments includes: Based on a target object detection model, determining at least one candidate object at each moment and the attribute information of the candidate object at that moment from the environmental data. Wherein, the attribute information includes the position information of the candidate object, and the size information and / or heading angle information of the candidate object. Performing matching between candidate objects according to the similarity of the attribute information between candidate objects at different moments, determining the attribute information of the same candidate object at each moment, and using the candidate object as the target object and the position information in the attribute information of the candidate object as the initial position information of the target object at each moment.
10. A vehicle control method, characterized in that, The method includes: Determining a target object in the external environment where the vehicle is located. Controlling the vehicle to travel according to the target state information of the target object. Among them, the target state information is the state information of the target object at each moment determined based on a preset second detection target and according to the initial position information and initial velocity information of the target object at multiple moments. The initial position information is the position information of the target object obtained based on environmental data at multiple moments. The initial velocity information is the initial velocity information of the target object at each moment determined based on a preset first detection target and according to the initial position information. The environmental data includes real-scene scanning data of the external environment of the vehicle. The target object is a dynamic object in the external environment of the vehicle. The target state information includes the target position information, target velocity information, and target heading angle information of the target object. The target position information is obtained based on the initial position information. The second detection target includes constraint conditions corresponding to multiple parts of the target object. The constraint conditions for different parts of the target object correspond to different confidence levels. The confidence level of the constraint conditions for each part is determined based on the number of position points corresponding to the part in the environmental data. The confidence level is used to represent the weight of the constraint condition of the part among the multiple part constraint conditions.
11. An electronic device, characterized in that, It includes a memory and a processor. The memory is used to store computer instructions, and the processor is used to call the computer instructions from the memory to execute the method according to any one of claims 1 to 10.
12. A vehicle, characterized in that, It includes a memory and a processor. The memory is used to store computer instructions, and the processor is used to call the computer instructions from the memory to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 10.
14. A chip, characterized in that, It includes a processing unit, and the processing unit is used to execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Estimation method, device and equipment of target motion state quantity, and storage medium
CN116466707A
Speed determination method and device based on laser radar, equipment and storage medium
CN116540252A