Track information determination method and device, equipment and storage medium

By generating and filtering trajectory information in different spatial dimensions, problems such as false detection in obstacle trajectory determination are solved, improving the accuracy of trajectory and the performance of the autonomous driving system.

CN119975405APending Publication Date: 2025-05-13ZHEJIANG LEAPMOTOR TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510033080.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, there are problems such as misdetection in determining obstacle trajectory, resulting in low accuracy of obstacle trajectory.

Method used

By obtaining several frames of the target vehicle, trajectory information (two-dimensional and three-dimensional) of different spatial dimensions are generated, and the three-dimensional trajectory information is used to filter the three-dimensional trajectory information to obtain more accurate target trajectory information.

Benefits of technology

It improves the accuracy of obstacle trajectory, reduces the false detection rate, and enhances the performance of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975405A_ABST
    Figure CN119975405A_ABST
Patent Text Reader

Abstract

The invention discloses a trajectory information determination method and device, equipment and a storage medium. The trajectory information determination method comprises the following steps: acquiring a plurality of frames of images acquired by at least one image acquisition device in a target vehicle for the environment where the target vehicle is located at the current moment; based on the plurality of frames of images, first track information and second track information of a target object are generated respectively, the spatial dimensions of the first track information and the second track information are different, and the target object is an obstacle in the environment; and filtering the second track information by using the first track information to obtain target track information of the target object. According to the scheme, the accuracy of obtaining the target trajectory information can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device, equipment and storage medium for determining trajectory information. Background Art

[0002] The autonomous driving function is one of the core functions of smart cars. Among them, the realization of the autonomous driving function can rely on the perception module as the upstream of the function realization, and the quality of its perception directly affects the performance of the entire autonomous driving. In some application scenarios, the trajectory determination of the target object is particularly important in the realization of the autonomous driving function. The target object can be an obstacle. The trajectory determination of the obstacle depends on the obstacle detection and ranging method. The traditional method is to directly perform target detection on the collected image to obtain the trajectory of the obstacle. However, the existing target detection process may have a certain degree of false detection and other problems, which will result in a low accuracy of the determined obstacle trajectory.

[0003] In view of the existing technical defects, how to provide an effective solution for determining trajectory information is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention

[0004] The present application at least provides a trajectory information determination method, apparatus, device and storage medium.

[0005] The present application provides a method for determining trajectory information, comprising: obtaining a plurality of frames of images of an environment in which the target vehicle is located by at least one image acquisition device in a target vehicle at a current moment; generating first trajectory information and second trajectory information of a target object based on the plurality of frames of images, respectively, wherein the first trajectory information and the second trajectory information have different spatial dimensions, and the target object is an obstacle in the environment; and filtering the second trajectory information using the first trajectory information to obtain target trajectory information of the target object.

[0006] In some embodiments, the first trajectory information includes at least one two-dimensional detection frame related to the target object, and the second trajectory information includes at least one three-dimensional detection frame related to the target object. The second trajectory information is filtered using the first trajectory information to obtain target trajectory information of the target object, including: performing the following steps for each three-dimensional detection frame: projecting the three-dimensional detection frame to the spatial dimension where the first trajectory information is located to obtain a projection frame corresponding to the three-dimensional detection frame; obtaining a target matching result of the projection frame according to the degree of overlap between the projection frame and each two-dimensional detection frame, the target matching result including a projection frame matching success or a projection frame matching failure; in response to the target matching result being a projection frame matching success, using the second trajectory information to which the projection frame belongs as the target trajectory information of the target object; or, in response to the target matching result being a projection frame matching failure, deleting the second trajectory information to which the projection frame belongs.

[0007] In some embodiments, each image acquisition device has a different acquisition angle, and the three-dimensional detection frame is projected to the spatial dimension where the first trajectory information is located to obtain a projection frame corresponding to the three-dimensional detection frame, including: obtaining calibration parameters corresponding to each image acquisition device, each calibration parameter characterizing the relative relationship between the image coordinate system and the world coordinate system of the image acquired by each image acquisition device; determining the projection corner points at each acquisition angle according to the initial corner points corresponding to the three-dimensional detection frame and each calibration parameter; and determining the projection frame based on the projection corner points at each acquisition angle.

[0008] In some embodiments, a projection frame is determined based on the projection corner points at each acquisition angle, including: selecting target projection corner points from the projection corner points at each acquisition angle, at least some of the target projection corner points are within a target image range, and the target image range is an image range corresponding to the acquisition angle to which the target projection corner points belong; in response to all of the target projection corner points being within the target image range, using a projection rectangular frame corresponding to the target projection corner point as the projection frame; or, in response to at least one of the target projection corner points not being within the target image range, determining a new projection rectangular frame based on the projection corner points not being within the target image range, and using the new projection rectangular frame as the projection frame.

[0009] In some embodiments, based on several frames of images, first trajectory information and second trajectory information of the target object are generated respectively, including: performing first detection processing and second detection processing on each image respectively to obtain a first initial target sequence and a second initial target sequence corresponding to each image, the spatial dimension of each first initial target sequence is two-dimensional, and the spatial dimension of each second initial target sequence is three-dimensional; according to the difference between a preset target sequence and the first initial target sequence and / or the second target sequence, determining a first target sequence corresponding to the first initial target sequence and / or a second target sequence corresponding to the second initial target sequence, the preset target sequence being a target sequence of several frames of historical images collected at a preset moment before the current moment; using the first initial target sequence or the first target sequence as the first trajectory information; using the second initial target sequence or the second target sequence as the second trajectory information.

[0010] In some embodiments, the preset target sequence includes a first preset target sequence with a two-dimensional spatial dimension, the first preset target sequence includes a plurality of first preset detection frames, the first initial target sequence includes a first initial detection frame for each target object, and the first target sequence corresponding to the first initial target sequence and / or the second target sequence are determined based on the difference between the preset target sequence and the first initial target sequence and / or the second target sequence, including: determining a first matching result based on the degree of overlap between each first preset detection frame and each first initial detection frame, the first matching result including at least one of successful matching of at least one first initial detection frame, failed matching of at least one first initial detection frame, and failed matching of at least one first preset detection frame; and updating the first initial target sequence based on the first matching result to obtain a first target sequence corresponding to the first initial target sequence.

[0011] In some embodiments, the preset target sequence includes a second preset target sequence with a two-dimensional spatial dimension, the second preset target sequence includes a plurality of second preset detection frames, the second initial target sequence includes a second initial detection frame for each target object, and the first target sequence corresponding to the first initial target sequence and / or the second target sequence are determined based on the difference between the preset target sequence and the first initial target sequence and / or the second target sequence, including: determining a second matching result based on the spatial distance between each second preset detection frame and each second initial detection frame, the second matching result including at least one of successful matching of at least one second initial detection frame, failed matching of at least one second initial detection frame, and failed matching of at least one second preset detection frame; and updating the second initial target sequence based on the second matching result to obtain a second target sequence corresponding to the second initial target sequence.

[0012] The present application provides a trajectory information determination device, including: an acquisition module, a generation module and a filtering module; the acquisition module is used to acquire a plurality of frames of images acquired by at least one image acquisition device in a target vehicle at a current moment of an environment where the target vehicle is located; the generation module is used to generate first trajectory information and second trajectory information of a target object based on the plurality of frames of images, respectively, the first trajectory information and the second trajectory information have different spatial dimensions, and the target object is an obstacle in the environment; the filtering module is used to filter the second trajectory information using the first trajectory information to obtain the target trajectory information of the target object.

[0013] The present application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-mentioned trajectory information determination method.

[0014] The present application provides a computer-readable storage medium on which program instructions are stored. When the program instructions are executed by a processor, the above-mentioned trajectory information determination method is implemented.

[0015] The above scheme, based on collecting several frames of images of the environment in which the target vehicle is located, generates first trajectory information and second trajectory information of the target object respectively. The first trajectory information and the second trajectory information have different spatial dimension information. The target object is an obstacle in the environment. The second trajectory information is filtered using the first trajectory information with different spatial dimension information to obtain the target trajectory information of the target object, so that the filtered target trajectory information is more accurate.

[0016] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.

[0018] Figure 1 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 1 ;

[0019] Figure 2 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 2 ;

[0020] Figure 3 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 3 ;

[0021] Figure 4 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 4 ;

[0022] Figure 5 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 5 ;

[0023] Figure 6 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 6 ;

[0024] Figure 7 It is a structural diagram of an embodiment of a trajectory information determination device of the present application;

[0025] Figure 8 It is a structural schematic diagram of an embodiment of the electronic device of the present application;

[0026] Fig. 9It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0027] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.

[0028] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0029] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the objects associated before and after are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of, for example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C.

[0030] The present application provides some trajectory information determination methods and trajectory information determination devices. The application scenarios of the trajectory information determination method include but are not limited to determining the trajectory information of a target object, wherein the target object may be an obstacle. The execution subject of the trajectory information determination method may be a trajectory information determination device. For example, the trajectory information determination device may be arranged in a terminal device or a server or other processing device, wherein the terminal device may be a device for trajectory information determination, a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, etc. In some possible implementations, the trajectory information determination method may be implemented by a processor calling computer-readable instructions stored in a memory.

[0031] See also Figure 1 , Figure 1 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 1 Specifically, the trajectory information determination method may include the following steps:

[0032] Step S11: obtaining a plurality of frames of images captured by at least one image acquisition device in the target vehicle of the environment where the target vehicle is located at the current moment.

[0033] The target vehicle may be a vehicle that needs to determine the trajectory information of the environment in which it is located. At least one may refer to one or more. Among them, the image acquisition device is a device arranged outside the target vehicle and capable of acquiring images of the environment in which the target vehicle is located. Different image acquisition devices are arranged at different positions outside the body of the target vehicle. It can be understood that the images acquired by different image acquisition devices are images acquired from different perspectives. Several frames of images are images acquired by each image acquisition device at the current moment respectively of the environment in which the target vehicle is located. Among them, the number of several frames of images is less than or equal to the number of the above-mentioned at least one image acquisition device. The environment in which the target vehicle is located may refer to the surrounding physical space in which the target vehicle is located. The environment in which the target vehicle is located includes at least one target object. Each target object may refer to an obstacle in the environment in which the target vehicle is located. The obstacle may be at least one of a tree, a building, a construction area, other vehicles, pedestrians, and a bicycle. Other vehicles are vehicles other than the target vehicle. It can be understood that the present application does not limit the type or number of target objects. The present application takes the target object as other vehicles in the environment in comparison with the target vehicle as an example.

[0034] Exemplarily, the above step S11 may be that at the current moment, the vehicle system corresponding to the target vehicle may trigger each image acquisition device to perform image acquisition as needed. For example, in the automatic driving system of the target vehicle, each image acquisition device may continue to work and capture the image of the target environment in real time. For example, the target vehicle triggers each image acquisition device to acquire images in a specific driving state (such as vehicle starting, lane changing, parking, driving, etc.).

[0035] Step S12: Based on a plurality of frame images, first trajectory information and second trajectory information of the target object are generated respectively.

[0036] The first trajectory information and the second trajectory information have different spatial dimensions, and the target object is an obstacle in the environment. The target object is an obstacle in the environment. Specifically, the target object of the present application includes but is not limited to a certain type of obstacle in the environment where the target vehicle is located. The present application takes the target object as other vehicles in the environment compared to the target vehicle as an example. The first trajectory information of the target object may refer to the trajectory information of each target of the obstacle in each frame image in the first spatial dimension. The second trajectory information of the target object may refer to the trajectory information of each target of the obstacle in each frame image in the second spatial dimension. Among them, the first spatial dimension is lower than the second spatial dimension. For example, each target of the obstacle in each frame image may be each vehicle among other vehicles. It can be understood that the captured image contains multiple image elements, and when the image element is an obstacle, the trajectory information of the element is generated. That is, each target of the obstacle may be each image element corresponding to the obstacle.

[0037] For example, an image captured by a certain image acquisition device includes image elements A, B, and C. Image element A is a tree. Image element B is another vehicle. Image element C is a pedestrian. In some application scenarios, when the obstacle is another vehicle, each target of the obstacle includes image element B. In other application scenarios, when the obstacle is another vehicle and a pedestrian, each target of the obstacle includes image element B and image element C.

[0038] Exemplarily, the first trajectory information of the target object includes the size and position of each target of the obstacle in each frame image in the first spatial dimension. Exemplarily, the second trajectory information of the target object includes the size, position, speed, and heading angle of each target of the obstacle in each frame image in the second spatial dimension.

[0039] The above step S12 may include: performing the following steps for each frame of image: performing a first detection process on the image to obtain the trajectory information of each target of the obstacle in the frame of image in the first spatial dimension. In some application scenarios, the first detection process may include but is not limited to performing target detection process related to the first spatial dimension on each frame of image. In other application scenarios, the first detection process may also be to first input each frame of image into a preset feature extraction module for feature extraction to obtain the extracted features of each frame of image. Then, the extracted features of each frame of image are subjected to target detection process related to the first spatial dimension to obtain the trajectory information of each target of the obstacle in each frame of image in the first spatial dimension.

[0040] Specifically, the above step S12 may include: performing a first detection process on each image to obtain the trajectory information of each target of the obstacle in each frame image in the second spatial dimension. In some application scenarios, the second detection process may include but is not limited to performing target detection process related to the second spatial dimension on each frame image. In other application scenarios, the second detection process may also be to first input each frame image into a preset feature extraction module for feature extraction to obtain the extracted features of each frame image. Then, the extracted features of each frame image are subjected to target detection process related to the second spatial dimension to obtain the trajectory information of each target of the obstacle in each frame image in the second spatial dimension.

[0041] Exemplarily, the preset feature extraction module is used to extract features from each input frame image to obtain extracted features of each frame image. Among them, a preset feature extraction network may be provided on the preset feature extraction module. Specifically, the preset feature extraction network may be a feature extraction network such as a convolutional neural network, a recurrent neural network, a long short-term memory network, a gated recurrent unit, a BiGRU neural network, an attention mechanism, etc. Exemplarily, the preset feature extraction network in each prediction module may be a long short-term memory network (LSTM, Long Short-Term Memory) in a recurrent neural network.

[0042] Step S13: Filter the second trajectory information using the first trajectory information to obtain target trajectory information of the target object.

[0043] The target trajectory information of the target object may be trajectory information of at least part of the objects in each target of the obstacle in each frame image in the second space dimension.

[0044] In some application scenarios, the above step S13 may be to determine whether the size difference between the first size information and the second size information meets the filtering condition. The filtering condition may include a size difference threshold. The second trajectory information corresponding to the second size information that meets the filtering condition is retained, and the retained second trajectory information is used as the target trajectory information of the target object. The second trajectory information corresponding to the second size information that does not meet the filtering condition is deleted. Among them, the first size information is used to indicate the size of each target of the obstacle in each frame image contained in the first trajectory information in the first spatial dimension. The second size information is used to indicate the size of each target of the obstacle in each frame image contained in the second trajectory information in the second spatial dimension. The size difference may be the difference between the size of a specific target of the obstacle in each frame image after the first size information and the second size information are converted to the same spatial dimension.

[0045] In some other application scenarios, the above step S13 may be to determine whether the position offset between the first position information and the second position information meets the filtering condition. The filtering condition may include a position offset threshold. The second trajectory information corresponding to the second position information that meets the filtering condition is retained, and the retained second trajectory information is used as the target trajectory information of the target object. The second trajectory information corresponding to the second position information that does not meet the filtering condition is deleted. Among them, the first position information is used to indicate the position of each target of the obstacle in each frame image contained in the first trajectory information in the first spatial dimension. The second position information is used to indicate the position of each target of the obstacle in each frame image contained in the second trajectory information in the second spatial dimension. The position offset may be the difference between the position of a specific target of the obstacle in each frame image after the first position information and the second position information are converted to the same spatial dimension.

[0046] The above scheme, based on collecting several frames of images of the environment in which the target vehicle is located, generates first trajectory information and second trajectory information of the target object respectively. The first trajectory information and the second trajectory information have different spatial dimension information. The target object is an obstacle in the environment. The second trajectory information is filtered using the first trajectory information with different spatial dimension information to obtain the target trajectory information of the target object, so that the filtered target trajectory information is more accurate.

[0047] See also Figure 2 , Figure 2 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 2 .

[0048] In some embodiments, the above step S12 may include the following steps: Step S21: Perform a first detection process and a second detection process on each image respectively to obtain a first initial target sequence and a second initial target sequence corresponding to each image. The spatial dimension of each first initial target sequence is two-dimensional, and the spatial dimension of each second initial target sequence is three-dimensional. Step S22: Determine the first target sequence corresponding to the first initial target sequence and / or the second target sequence corresponding to the second initial target sequence based on the difference between the preset target sequence and the first initial target sequence and / or the second target sequence. The preset target sequence is a target sequence of several frames of historical images collected at a preset time before the current moment. Step S23: Use the first initial target sequence or the first target sequence as the first trajectory information. Step S24: Use the second initial target sequence or the second target sequence as the second trajectory information.

[0049] In some application scenarios, the first detection process includes at least target detection processing related to the two-dimensional space dimension. The above-mentioned first space dimension may be a two-dimensional space dimension. Exemplarily, the target detection process related to the two-dimensional space dimension may be a 2D detection process. Wherein, each frame image corresponds to a first initial target sequence. Each first initial target sequence includes the first detection results of each target of the obstacle in the frame image after the corresponding image is subjected to the first detection process. The first detection result of each target may include at least the detection box of each target in the first space dimension.

[0050] In some other application scenarios, the second detection process includes at least target detection processing related to the three-dimensional space dimension. The above-mentioned second space dimension may be a three-dimensional space dimension. Exemplarily, the target detection process related to the three-dimensional space dimension may be a 3D detection process. Among them, each frame image corresponds to a second initial target sequence. The second initial target sequence includes the second detection results of each target of the obstacle in each frame image after all images are processed by the second detection process. The second detection result of each target may include at least the detection box of each target in the second space dimension.

[0051] For example, the number of at least one image acquisition device in the present application may be 6. The layout of the 6 image acquisition devices can cover a 360° viewing angle of the environment in which the target vehicle is located. At least part of the acquisition field of view of each image acquisition device overlaps.

[0052] Exemplarily, the first detection process in the above step S21 may be to input each frame of image into a 2D perception network respectively, and obtain a BBox containing all dynamic obstacles on each frame of image. The input of the 2D perception network is a single image, and the output of the 2D perception network is a BBox containing all dynamic obstacles on the image. It is understandable that the 2D perception network may be a network module or algorithm module that has been pre-trained and is capable of performing the first detection process. Among them, images captured by different image acquisition devices may be input into different 2D perception networks, or may be input into the same 2D perception network. Different 2D perception networks may have the same structure but different parameter quantities, or may have different structures and parameter quantities. It is understandable that the image captured by each image acquisition device can correspond to a 2D perception network.

[0053] The first detection result of each target, that is, the output information corresponding to the above BBox in each frame of the image, may include at least one of the following: the category of the obstacle, the whole vehicle detection frame rectangle, the head and tail detection frame rectangle, the preset part frame rectangle, the tire line, the confidence level and other information. Exemplarily, the category of the obstacle may be a car, a person, a rider, a two-wheeled vehicle, a three-wheeled vehicle, etc. The whole vehicle detection frame rectangle, the head and tail detection frame rectangle and the preset part frame rectangle may all refer to a rectangular detection frame obtained by the left side of the upper left pixel point + width + height of at least part of the image area of ​​the target object. It can be understood that the image area of ​​the target object included in the whole vehicle detection frame rectangle, the head and tail detection frame rectangle and the preset part frame rectangle is different. For example, the image area of ​​the target object included in the whole vehicle detection frame rectangle is at least the entire image area to which the target object belongs. For example, the image area of ​​the target object included in the head and tail detection frame rectangle is at least the head image area and the tail image area to which the target object belongs. For example, the image area of ​​the target object included in the preset part detection frame rectangle is at least the image area of ​​the preset part to which the target object belongs. The preset part may be the tire part of the target object. It is understandable that for the specific target of the obstacle in each frame image, the number of tire parts may be different, that is, the number of preset part frame rectangles may be different. For example, in the case where the category of the obstacle is a large truck, there will be multiple tires, and the number of preset part frame rectangles is multiple. In the case where the category of the obstacle is a four-wheeled or two-wheeled vehicle, the number of preset part frame rectangles is 2. In some application scenarios, when the specific target of the obstacle in each frame image is in the edge area of ​​the image, the number of preset part frame rectangles may be 1. The above confidence level can be the credibility of the output information corresponding to the output BBox of the 2D perception network. It is understandable that the above-mentioned first initial target sequence includes the first detection result of the target of the obstacle in the environment where the target vehicle is located corresponding to a frame image collected at the current moment.

[0054] Exemplarily, the second detection process in the above step S21 may be to input all frame images into the 3D perception network separately to obtain the BBox of the dynamic obstacle in the world system. The input of the 3D perception network is a single or multiple images, and the output of the 3D perception network is the BBox of the dynamic obstacle in the world system. In some application scenarios, when the input is multiple images, the 3D perception network can output targets with a more complete perspective, such as front view, rear view, and side view. In other application scenarios, when the input of the 3D perception network is a single image, generally a forward-looking image, the 3D perception network can also output the target object, but it can only output the target object in the forward-looking field of view. Compared with the case of multiple images, the perception range of a single image is relatively small.

[0055] The second detection result of each of the above targets, that is, the output information of the BBox of the dynamic obstacle in the world system, may include at least one of the following: the category of the obstacle, the 3D position of the center point of the bottom edge, the length, width and height, the heading angle, and the degree of confidence. Exemplarily, the category of the obstacle may be a car, a person, a cyclist, a two-wheeled vehicle, a three-wheeled vehicle, etc. The 3D position of the center point of the bottom edge may include the coordinate value of the x-axis, the coordinate value of the y-axis, and the coordinate value of the z-axis of the center point of the bottom edge of the target of the obstacle in the world coordinate system. The length, width and height may refer to the physical size of the target of the obstacle, which can characterize the size of the target of the obstacle. The heading angle may indicate the orientation of the target of the obstacle, that is, the angle between the main axis of the target and the reference direction. The reference direction may be the north direction or the forward direction of the target vehicle. The above degree of confidence may be the credibility of the output information corresponding to the output BBox of the 3D perception network. It can be understood that the above second initial target sequence includes the second detection result of the target of the obstacle in the environment where the target vehicle is located at the current moment.

[0056] For example, the present application takes the example that a 2D perception network can train 6 perception reasoning networks corresponding to 6 image acquisition devices. The layout settings of the 6 image acquisition devices on the target vehicle can be front view, left front, left rear, rear, right rear, and right front. The 3D perception network is a separate perception network, and the input is 6 images collected by 6 image acquisition devices at the same time.

[0057] On the vehicle side and in the local offline environment, the operation is that the six image acquisition devices are used to acquire multiple groups of images in real time, and the images contained in each group of images need to ensure that the timestamps are aligned. The time of the six images contained in each group of images is consistent. Each image group is sent to six 2D perception networks respectively, and is sent to the 3D perception network at the same time. The above step S11 can obtain the current image group acquired at the current moment. The images contained in the current image group are the images corresponding to the current moment. The above step S21 can be to input each image of the current image group into different 2D perception networks for the first detection processing to obtain the first initial target sequence corresponding to each image. And, each image of the current image group is input into the same 3D perception network for the second detection processing to obtain the second initial target sequence corresponding to each image.

[0058] It is understandable that the first initial target sequence may include the track information of at least one target belonging to an obstacle in two-dimensional space, and the second initial target sequence may include the track information of at least one target belonging to an obstacle in three-dimensional space.

[0059] The preset target sequence is a target sequence of several frames of historical images collected at a preset moment before the current moment. The preset moment may be a moment before the current moment, or any moment before the current moment. In other application scenarios, the acquisition frequencies of the image acquisition devices are the same, and the preset moment may be the moment corresponding to the last frame of image collected by the image acquisition device. It can be understood that several frames of historical images may be the last frames of image collected by each image acquisition device at the preset moment. In some application fields, the preset target sequence, that is, the target sequence of several frames of historical images, may be historical target trajectory information obtained by filtering the second historical trajectory information using the first historical trajectory information generated by each frame of historical image. The first historical trajectory information may be a first initial target sequence corresponding to the preset moment obtained by performing a first detection process on each frame of historical image. The second historical trajectory information may be a second initial target sequence corresponding to the preset moment obtained by performing a second detection process on each frame of historical image. In other application scenarios, the preset target sequence includes a first preset target sequence corresponding to the first historical trajectory information and a second preset target sequence corresponding to the second historical trajectory information.

[0060] In some application scenarios, the process of obtaining the first target sequence corresponding to the first initial target sequence may be to perform any one or more of overlapping area filtering, confidence filtering, and frame supplementation processing on the same target in the first initial target sequence and the preset target sequence to obtain a new first initial sequence, and use the new first initial sequence as the first target sequence. Wherein, overlapping area filtering may be when the overlapping area between the first initial target sequence and the same target in the preset target sequence is greater than the area threshold, the trajectory information of the target in the first initial target sequence is deleted to obtain a new first initial target sequence. Confidence filtering may be when the confidence level of the target in the first initial target sequence is lower than the confidence threshold, the trajectory information of the target in the first initial target sequence is deleted to obtain a new first initial target sequence. Frame supplementation processing may be when any target exists in the preset target sequence but does not exist in the first initial sequence, the trajectory information of any target in the preset target sequence is added to the first initial sequence to obtain a new first initial target sequence.

[0061] In other application scenarios, the process of obtaining the second target sequence corresponding to the second initial target sequence may be to perform any one or more of overlapping area filtering, confidence filtering, and frame supplementation processing on the same target in the second initial target sequence and the preset target sequence to obtain a new second initial sequence, and use the new second initial sequence as the second target sequence. Wherein, overlapping area filtering may be when the overlapping area between the second initial target sequence and the same target in the preset target sequence is greater than the area threshold, the trajectory information of the target in the second initial target sequence is deleted to obtain a new second initial target sequence. Confidence filtering may be when the confidence level of the target in the second initial target sequence is lower than the confidence threshold, the trajectory information of the target in the second initial target sequence is deleted to obtain a new second initial target sequence. Frame supplementation processing may be when any target exists in the preset target sequence but does not exist in the second initial sequence, the trajectory information of any target in the preset target sequence is added to the second initial sequence to obtain a new second initial target sequence.

[0062] The first trajectory information is used to represent the trajectory information corresponding to at least one target obtained by performing the first detection process on each frame image. The second trajectory information is used to represent the trajectory information corresponding to at least one target obtained by performing the second detection process on each frame image.

[0063] In some application scenarios, the above step S22 may be to determine the first target sequence corresponding to the first initial target sequence according to the difference between the first preset target sequence and the first initial target sequence in the preset target sequence. And, directly determine the second initial target sequence as the second target sequence. The above step S23 may be to use the first target sequence as the first trajectory information. The above step S24 may be to use the second initial target sequence as the second trajectory information.

[0064] In other application scenarios, the above step S22 may be to determine the second target sequence corresponding to the second initial target sequence according to the difference between the second preset target sequence and the second initial target sequence in the preset target sequence. And, directly determine the first initial target sequence as the first target sequence. The above step S23 may be to use the first initial target sequence as the first trajectory information. The above step S24 may be to use the second target sequence as the second trajectory information.

[0065] In other application scenarios, the above step S22 may be to determine the first target sequence corresponding to the first initial target sequence according to the difference between the first preset target sequence and the first initial target sequence in the preset target sequence. And, according to the difference between the second preset target sequence and the second initial target sequence in the preset target sequence, determine the second target sequence corresponding to the second initial target sequence. The above step S23 may be to use the first target sequence as the first trajectory information. The above step S24 may be to use the second target sequence as the second trajectory information.

[0066] See also Figure 3 , Figure 3 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 3 .

[0067] In some embodiments, the preset target sequence includes a first preset target sequence with a two-dimensional spatial dimension, the first preset target sequence includes a plurality of first preset detection frames, and the first initial target sequence includes a first initial detection frame for each target object. The above step S22 may include the following steps: Step S31: Determine the first matching result according to the degree of overlap between each first preset detection frame and each first initial detection frame. The first matching result includes at least one of a successful match of at least one first initial detection frame, a failed match of at least one first initial detection frame, and a failed match of at least one first preset detection frame. Step S32: Update the first initial target sequence according to the first matching result to obtain a first target sequence corresponding to the first initial target sequence.

[0068] Each first preset detection frame in the first preset target sequence corresponds to a detection frame of a historical target. Each first initial detection frame in the first initial target sequence corresponds to a detection frame of a current target. The current target may be one of the target objects. The target objects may belong to the same type of obstacle or different types of obstacles.

[0069] The first matching result includes the matching results of each first preset detection frame and each first initial detection frame. Successful matching of the first initial detection frame is used to indicate that the match between the first initial detection frame and the first preset detection frame in the matching pair is successful, and a matching pair exists. Failed matching of the first initial detection frame is used to indicate that there is no matching pair between the first initial detection frame and each first preset detection frame. Failed matching of the first preset detection frame is used to indicate that the match between the first initial detection frame and the first preset detection frame in the matching pair fails, or there is no matching pair between the first preset detection frame and each first initial detection frame. Specifically, the first matching result includes any one, any two, or all three of successful matching of at least one first initial detection frame, failed matching of at least one first initial detection frame, and failed matching of at least one first preset detection frame.

[0070] The above step S31 may include: the degree of overlap between each first preset detection frame and each first initial detection frame may be an IoU value between each first preset detection frame and each first initial detection frame. The IoU values ​​between each first preset detection frame and each first initial detection frame are calculated respectively. The larger the IoU, the higher the degree of overlap between the first preset detection frame and the first initial detection frame, and the higher the similarity. The IoU value represents the degree of overlap between any two first preset detection frames and the first initial detection frame, and the IoU value is constructed as a cost matrix, in which the row represents the detection frame of the historical target and the column represents the detection frame of the current target. The cost matrix is ​​solved using the Hungarian algorithm to find the matching pair with the largest total IoU. The Hungarian algorithm assigns a first initial detection frame to each first preset detection frame so that the total matching cost IoU value is maximized, or the minimized cost 1-IoU is minimized. According to the relationship between the IoU value and the first threshold, the above first matching result is determined.

[0071] In some application scenarios, the above step S31 may be that for a matching pair whose IoU value is higher than the first threshold, the first initial detection frame of the matching pair is matched successfully. The matching pair includes the first initial detection frame and the first preset detection frame. The above step S32 may be that the index of each target in the first initial target sequence is retained to the first target sequence. For each initial detection frame, in response to the first matching result including the first initial detection frame matching successfully, the first initial detection frame is added to the index of the corresponding target in the first target sequence.

[0072] In some other application scenarios, the above step S31 may be that if a first preset detection frame does not find a matching first initial detection frame, or a matching pair with an IoU value lower than a first threshold, the first preset detection frame is considered to have failed to match. The above step S32 may be that the index of each target in the first initial target sequence is retained in the first target sequence. For each preset detection frame, in response to the first matching result including the failure to match the first preset detection frame, it is necessary to process the tracking loss of the first preset detection frame, that is, to add the first preset detection frame to the index of the corresponding target in the first target sequence.

[0073] In other application scenarios, the above step S31 may be that if the first initial detection frame and any first preset detection frame cannot match, then it is considered that the first initial detection frame fails to match, and the first initial detection frame cannot find any first preset detection frame as a matching pair, and the first initial detection frame may be a new target. The above step S32 can retain the index of each target in the first initial target sequence to the first target sequence. For each initial detection frame, in response to the first matching result including the failure of matching the first initial detection frame, the target corresponding to the first initial detection frame is used as a newly added target. The index of the newly added target is added to the first target sequence. Then the first initial detection frame is directly added to the index of the newly added target corresponding to the first target sequence.

[0074] In other application scenarios, each image in the above step S21 corresponds to a first initial target sequence. The first preset target sequence includes a plurality of first preset detection frames. According to the historical images collected by different image acquisition devices, the plurality of first preset detection frames are divided to obtain at least one first preset detection frame group. The historical targets corresponding to the first preset detection frames in each first preset detection frame group belong to the historical images collected by the same image acquisition device. The above step of determining the first target sequence corresponding to the first initial target sequence according to the difference between the first preset target sequence and the first initial target sequence in the preset target sequence can be, for each first initial target sequence, the degree of overlap between the first preset detection frames in the first target preset detection frame group and the first initial detection frames in the first initial target sequence, determining the above first matching result, and executing the above step S32. The first target preset detection frame group is one of the above first preset detection frame groups, and the image acquisition device corresponding to the first target preset detection frame group is the same as the image acquisition device to which the image corresponding to the first initial target sequence belongs.

[0075] It can be considered that the first target sequence obtained by updating the first initial target sequence using the first matching result determined by the degree of overlap between each first preset detection frame and each first initial detection frame can make the first target sequence determined more accurately.

[0076] See also Figure 4 , Figure 4 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 4 .

[0077] In some embodiments, the preset target sequence includes a second preset target sequence with a two-dimensional spatial dimension, the second preset target sequence includes a plurality of second preset detection frames, and the second initial target sequence includes a second initial detection frame of each target object. The above step S22 may include the following steps: Step S41: Determine the second matching result based on the spatial distance between each second preset detection frame and each second initial detection frame. The second matching result includes at least one of at least one second initial detection frame successfully matching, at least one second initial detection frame failing to match, and at least one second preset detection frame failing to match. Step S42: Update the second initial target sequence based on the second matching result to obtain a second target sequence corresponding to the second initial target sequence.

[0078] Each second preset detection frame in the second preset target sequence corresponds to a detection frame of a historical target. Each second initial detection frame in the second initial target sequence corresponds to a detection frame of a current target. The current target may be one of the target objects. The target objects may belong to the same type of obstacle or different types of obstacles.

[0079] The second matching result includes the matching results of each second preset detection frame and each second initial detection frame. The successful matching of the second initial detection frame is used to indicate that the match between the second initial detection frame and the second preset detection frame in the matching pair is successful, and a matching pair exists. The failed matching of the second initial detection frame is used to indicate that there is no matching pair between the second initial detection frame and each second preset detection frame. The failed matching of the second preset detection frame is used to indicate that the match between the second initial detection frame and the second preset detection frame in the matching pair fails, or there is no matching pair between the second preset detection frame and each second initial detection frame. Specifically, the second matching result includes any one, any two, or all three of the successful matching of at least one second initial detection frame, the failed matching of at least one second initial detection frame, and the failed matching of at least one second preset detection frame.

[0080] The above step S41 may include: the spatial distance between each second preset detection frame and each second initial detection frame may be an IoU value between each second preset detection frame and each second initial detection frame. The IoU values ​​between each second preset detection frame and each second initial detection frame are calculated respectively. The larger the IoU, the smaller the spatial distance between the second preset detection frame and the second initial detection frame, and the higher the similarity. The IoU value represents the spatial distance between any two second preset detection frames and the second initial detection frame, and the IoU value is constructed as a cost matrix, where the row represents the detection frame of the historical target and the column represents the detection frame of the current target. The cost matrix is ​​solved using the Hungarian algorithm to find the matching pair with the largest total IoU. The Hungarian algorithm assigns a second initial detection frame to each second preset detection frame so that the total matching cost IoU value is maximized, or the minimization cost 1-IoU is minimized. According to the relationship between the IoU value and the second threshold, the above second matching result is determined.

[0081] In some application scenarios, the above step S41 may be that for a matching pair whose IoU value is higher than the second threshold, the second initial detection frame of the matching pair is matched successfully. The matching pair includes the second initial detection frame and the second preset detection frame. The above step S42 may be that the index of each target in the second initial target sequence is retained to the second target sequence. For each initial detection frame, in response to the second matching result including the second initial detection frame matching successfully, the second initial detection frame is added to the index of the corresponding target in the second target sequence.

[0082] In other application scenarios, the above step S41 may be that if a second preset detection frame does not find a matching second initial detection frame, or a matching pair with an IoU value lower than a second threshold, the second preset detection frame is considered to have failed to match. The above step S42 may be that the index of each target in the second initial target sequence is retained in the second target sequence. For each preset detection frame, in response to the second matching result including the failure to match the second preset detection frame, it is necessary to process the second preset detection frame for tracking loss, that is, to add the second preset detection frame to the index of the corresponding target in the second target sequence.

[0083] In other application scenarios, the above step S41 may be that if the second initial detection frame and any second preset detection frame cannot match, then it is considered that the second initial detection frame fails to match, and the second initial detection frame cannot find any second preset detection frame as a matching pair, and the second initial detection frame may be a new target. The above step S42 can retain the index of each target in the second initial target sequence to the second target sequence. For each initial detection frame, in response to the second matching result including the failure of matching the second initial detection frame, the target corresponding to the second initial detection frame is used as a newly added target. The index of the newly added target is added to the second target sequence. Then the second initial detection frame is directly added to the index of the newly added target corresponding to the second target sequence.

[0084] In some other application scenarios, each image in the above step S21 corresponds to a second initial target sequence. The second preset target sequence includes a plurality of second preset detection frames. According to the historical images collected by different image acquisition devices, the plurality of second preset detection frames are divided to obtain at least one second preset detection frame group. The historical targets corresponding to each second preset detection frame in each second preset detection frame group belong to the historical images collected by the same image acquisition device. The above step of determining the second target sequence corresponding to the second initial target sequence according to the difference between the second preset target sequence and the second initial target sequence in the preset target sequence can be, for each second initial target sequence, the spatial distance between each second preset detection frame in the second target preset detection frame group and each second initial detection frame in the second initial target sequence, determining the above second matching result, and executing the above step S42. The second target preset detection frame group is one of the above second preset detection frame groups, and the image acquisition device corresponding to the second target preset detection frame group is the same as the image acquisition device to which the image corresponding to the second initial target sequence belongs.

[0085] It can be considered that the second target sequence obtained by updating the second initial target sequence using the second matching result determined by the spatial distance between each second preset detection frame and each second initial detection frame can make the second target sequence determined more accurately.

[0086] Exemplarily, the above step S22 may be a post-processing method of treating the output of the perception network as sensor data and performing multi-target tracking + fusion on the sensor data. The output of the perception network includes the output of 6 2D perception networks and 1 3D perception network. The output of the perception network is subjected to multi-target tracking respectively, and then fused. The fusion adopts a two-stage method, that is, first fusion with 3D and then fusion with 2D.

[0087] Specifically, the above-mentioned determination of the first target sequence corresponding to the first initial target sequence according to the difference between the first preset target sequence and the first initial target sequence in the preset target sequence may include: inputting the first preset target sequence and the first initial target sequence in the preset target sequence into the 2D MOT module to obtain the first target sequence output by the 2D MOT module. Among them, the 2D MOT module can be an algorithm module provided with non-maximum suppression and multi-target tracking. Among them, non-maximum suppression (nms) compares the similarity of two targets to be compared, and deletes the target with smaller attributes when the similarity is large. The attribute selected in the 2D MOT module is the confidence, and the set similarity is the 2D rectangular intersection ratio, that is, the degree of overlap between the detection boxes in the two target sequences. Exemplarily, for each BBox output by the 2D perception network, the re-detection box is filtered by nms. Generally, when the overlapping area of ​​the two 2D boxes is greater than 0.6, the target with lower confidence is filtered out (deleted). For each BBox output by the 2D perception network, continuous tracking between frames is performed through multi-target tracking. If the target loses one or several frames, frame supplementation is performed, and frame supplementation flag information is provided for the target that needs frame supplementation for use during fusion. The frame supplementation flag information includes but is not limited to the indication information of the target that needs frame supplementation in the first initial target sequence and the indication information of the target that can be used as the frame supplementation object in the first preset target sequence.

[0088] Specifically, the above-mentioned determination of the second target sequence corresponding to the second initial target sequence according to the difference between the second preset target sequence and the second initial target sequence in the preset target sequence may include: inputting the second preset target sequence and the second initial target sequence in the preset target sequence into the 3D MOT module to obtain the second target sequence output by the 3D MOT module. Among them, the 3D MOT module can be an algorithm module provided with non-maximum suppression and multi-target tracking. Among them, non-maximum suppression (nms) compares the similarity of two targets to be compared, deletes the target with smaller attributes when the similarity is large, and the attribute selected in the 3D MOT module is the confidence, and the set similarity is the center point distance, that is, the spatial distance between the detection frames in the two target sequences. Exemplarily, for the BBox output by the 3D perception network, nms and multi-target tracking are performed, the re-detection frame is filtered by nms, and continuous tracking between frames is performed by multi-target tracking, and the target loses one or several frames for frame supplementation processing. Frame supplementation processing requires providing frame supplementation flag information. The frame supplement flag information includes, but is not limited to, indication information for targets in the first initial target sequence that need frame supplementation, and indication information for targets in the first preset target sequence that can be used as frame supplementation objects. For targets in the second target sequence output by the 3D MOT module, speed calculations are performed between multiple frames, and the position and speed are smoothed according to the Kalman filter to prevent abnormalities. 3D position and speed are the attributes that obstacles ultimately output to downstream modules for use. The 3D perception network generally only outputs the position of a single-frame target, and the speed needs to be fitted and solved. When smoothing the speed, the position is also smoothed using the Kalman filter. Even if there is a position jump in a single-frame 3D perception network, the result of the final tracking will not show a particularly large position jump.

[0089] The first trajectory information includes the position and size of each target in the first initial target sequence or the first target sequence in the first spatial dimension. The second trajectory information includes the position, size, speed and yaw angle of each target in the second initial target sequence or the second target sequence in the second spatial dimension. In some application scenarios, the above step S23 may be to use the first target sequence as the first trajectory information. The above step S24 may be to use the second initial target sequence as the second trajectory information. In other application scenarios, the above step S23 may be to use the first initial target sequence as the first trajectory information. The above step S24 may be to use the second target sequence as the second trajectory information. In other application scenarios, the above step S23 may be to use the first target sequence as the first trajectory information. The above step S24 may be to use the second target sequence as the second trajectory information.

[0090] It can be considered that using the first initial target sequence or the first target sequence as the first trajectory information and using the second initial target sequence or the second target sequence as the second trajectory information can improve the accuracy of the first trajectory information and the second trajectory information.

[0091] See also Figure 5 , Figure 5 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 5 .

[0092] In some embodiments, the first trajectory information includes at least one two-dimensional detection frame related to the target object, and the second trajectory information includes at least one three-dimensional detection frame related to the target object. The above step S13 may include the following steps: performing the following steps on each three-dimensional detection frame: Figure 5 The following steps are shown: Step S51: Project the three-dimensional detection frame to the spatial dimension where the first trajectory information is located to obtain the projection frame corresponding to the three-dimensional detection frame. Step S52: Obtain the target matching result of the projection frame according to the degree of overlap between the projection frame and each two-dimensional detection frame. The target matching result includes a successful projection frame match or a failed projection frame match. Step S53: In response to the target matching result being a successful projection frame match, the second trajectory information belonging to the projection frame is used as the target trajectory information of the target object. Or, Step S54: In response to the target matching result being a failed projection frame match, the second trajectory information belonging to the projection frame is deleted.

[0093] In some application scenarios, at least one two-dimensional detection frame associated with the target object may refer to at least one two-dimensional detection frame associated with the obstacle. When the obstacles are of one or more categories, each two-dimensional detection frame is a detection frame of a target belonging to any one category of obstacles in the first spatial dimension. In some application scenarios, at least one three-dimensional detection frame associated with the target object may refer to at least one three-dimensional detection frame associated with the obstacle. When the obstacles are of one or more categories, each three-dimensional detection frame is a detection frame of a target belonging to any one category of obstacles in the first spatial dimension.

[0094] The above step S51 may be to project the three-dimensional detection frame to the first spatial dimension where the first trajectory information is located, obtain a projected rectangular frame corresponding to the three-dimensional detection frame, and directly use the projected rectangular frame corresponding to the three-dimensional detection frame as the projection frame.

[0095] See also Figure 6 , Figure 6 This is a schematic diagram of the process of an embodiment of the method for determining trajectory information of the present application. Figure 6 .

[0096] In some embodiments, the acquisition angles of the image acquisition devices are different, and the above step S51 may include the following steps: Step S61: Acquire the calibration parameters corresponding to each image acquisition device. Each calibration parameter represents the relative relationship between the image coordinate system and the world coordinate system of the image acquired by each image acquisition device. Step S62: Determine the projection corner points at each acquisition angle based on the initial corner points corresponding to the three-dimensional detection frame and each calibration parameter. Step S63: Determine the projection frame based on the projection corner points at each acquisition angle.

[0097] The acquisition angle of the image acquisition device may refer to the viewing angle of the image acquisition device. The layout of each image acquisition device on the target vehicle is different, and the acquisition angle of each image acquisition device is different. The calibration parameters corresponding to each image acquisition device may be camera calibration parameters. The image coordinate system may refer to a two-dimensional plane coordinate system. Each calibration parameter includes internal parameters and external parameters. Among them, the internal parameters can characterize the imaging characteristics of any image acquisition device, including but not limited to the internal parameter matrix of the image acquisition device and the parameters that can constitute the internal parameter matrix. The external parameters can characterize the position and posture of any image acquisition device in the world coordinate system, including but not limited to the external parameter matrix of the image acquisition device and the parameters that can constitute the external parameter matrix. Exemplarily, the external parameters include but are not limited to the external parameter rotation matrix and translation matrix, the rotation matrix and translation matrix of the world system to the camera system. Each calibration parameter may also include the corner points detected and calibrated in each image collected by each image acquisition device. Among them, the corner points detected and calibrated in each image can correspond to a target image range. The target image range is the range where the image collected by the image acquisition device is located.

[0098] The initial corner points corresponding to each three-dimensional detection frame are several corner points on each three-dimensional detection frame, for example, 8 corner points. The projection corner points at each acquisition angle are used to represent the corner points of each three-dimensional detection frame after being projected to the first spatial dimension using the calibration parameters corresponding to each acquisition angle. For example, the number of each image acquisition device is 6. The number of calibration parameters is 6. For each three-dimensional detection frame, there are 6 groups of projection corner points at each acquisition angle. The image acquired at each acquisition angle corresponds to a target image range. For each three-dimensional detection frame, only some of the projection corner points in the 6 groups of projection corner points are within the target image range corresponding to each acquisition angle. In some application scenarios, for each three-dimensional detection frame, when the projection corner points belonging to acquisition angle A are compared with the target image range B corresponding to the image acquired at acquisition angle B, the projection corner points belonging to acquisition angle A may not be within the target image range B. In other application scenarios, when the projection corner points belonging to acquisition angle A are compared with the target image range A corresponding to the image acquired at acquisition angle A, some of the projection corner points belonging to acquisition angle A may not be within the target image range A.

[0099] In some application scenarios, the above step S63 may be that for each 3D detection frame, in response to at least part of the projection corner points belonging to the same acquisition angle being within the target image range corresponding to the acquisition angle, all the projection corner points under the acquisition angle are used as the target projection corner points of the image corresponding to the target image range. The projection rectangular frame corresponding to the target projection corner point is directly used as the above projection frame.

[0100] In other application scenarios, the above step S63 may be to delete all projection corner points under the acquisition angle for each 3D detection frame in response to all projection corner points under the same acquisition angle not being within the target image range corresponding to the acquisition angle.

[0101] In some embodiments, determining a projection frame based on the projection corner points at each acquisition angle includes: selecting a target projection corner point from the projection corner points at each acquisition angle. At least some of the target projection corner points are within a target image range, and the target image range is an image range corresponding to the acquisition angle to which the target projection corner point belongs. In response to all of the target projection corner points being within the target image range, a projection rectangular frame corresponding to the target projection corner point is used as the projection frame. Or, in response to at least one of the target projection corner points not being within the target image range, a new projection rectangular frame is determined based on the projection corner points not within the target image range, and the new projection rectangular frame is used as the projection frame.

[0102] The above-mentioned step of selecting target projection corner points from the projection corner points under each acquisition angle includes: for each three-dimensional detection frame, in response to at least some of the projection corner points belonging to the same acquisition angle being within the target image range corresponding to the acquisition angle, all the projection corner points under the acquisition angle are used as the target projection corner points of the image corresponding to the target image range.

[0103] After the step of selecting the target projection corner point from the projection corner points under each acquisition angle, it is determined whether all the projection corner points in the target projection corner points are within the target image range. In the case that all the projection corner points in the target projection corner points are within the target image range, the projection rectangular frame corresponding to the target projection corner point is used as the projection frame. In the case that at least one of the target projection corner points is not within the target image range, a new projection rectangular frame is determined based on the projection corner points that are not within the target image range. And the new projection rectangular frame is used as the projection frame. The step of determining a new projection rectangular frame based on the projection corner points that are not within the target image range can be that for each projection corner point that is not within the target image range, a reference projection corner point corresponding to the projection corner point that is not within the target image range is determined. The reference projection corner point is a projection corner point that is adjacent to the projection corner point that is not within the target image range and is within the target image range. The intersection point between the target line and the image boundary of the target image range is used as the new corner point corresponding to the projection corner point that is not within the target image range. Among them, the target line is the line connecting the projection corner point that is not within the target image range and the reference projection corner point. A new projection rectangular frame is determined based on the new corner points and the projection corner points of the target projection corner points that are within the target image range.

[0104] The projection frames belonging to the same acquisition angle are regarded as a projection target sequence. Each projection target sequence includes the index of the projection target and the projection frames under the index of the projection target.

[0105] The above step S52 may be to determine the target matching result of each projection frame in each projection target sequence according to the degree of overlap between each projection target sequence and the first target sequence or the first initial target sequence in the first trajectory information. Specifically, the target matching result of each projection frame in each projection target sequence is determined according to the degree of overlap between each projection frame in each projection target sequence and each two-dimensional detection frame in the first target sequence or the first initial target sequence in the first trajectory information.

[0106] The target matching result includes the matching results of each projection detection frame and each two-dimensional detection frame. The projection detection frame matching success is used to indicate that the matching between the two-dimensional detection frame and the projection detection frame in the matching pair is successful, and a matching pair exists. The projection detection frame matching failure is used to indicate that there is no matching pair between the two-dimensional detection frame and each projection detection frame. Specifically, the target matching result includes any one of at least one projection detection frame matching success, at least one projection detection frame matching failure, or both.

[0107] The above step S31 may include: the degree of overlap between each projected detection frame and each two-dimensional detection frame may be an IoU value between each projected detection frame and each two-dimensional detection frame. The IoU values ​​between each projected detection frame and each two-dimensional detection frame are calculated respectively. The larger the IoU, the higher the degree of overlap between the projected detection frame and the two-dimensional detection frame, and the higher the similarity. The IoU value represents the degree of overlap between any two projected detection frames and the two-dimensional detection frames, and the IoU value is constructed as a cost matrix, where the rows represent the detection frames of the historical targets and the columns represent the detection frames of the current targets. The cost matrix is ​​solved using the Hungarian algorithm to find the matching pair with the largest total IoU. The Hungarian algorithm assigns a two-dimensional detection frame to each projected detection frame so that the total matching cost IoU value is maximized, or the minimized cost 1-IoU is minimized. According to the relationship between the IoU value and the first threshold, the above target matching result is determined.

[0108] In some application scenarios, the above step S52 may be that for a matching pair whose IoU value is higher than the first threshold, the projection detection box of the matching pair is matched successfully. The matching pair includes a two-dimensional detection box and a projection detection box. The above step S53 may be that the second trajectory information to which the projection box belongs is used as the target trajectory information of the target object.

[0109] In other application scenarios, the above step S52 may be that if a certain projection detection frame does not find a matching two-dimensional detection frame, or the matching pair with an IoU value lower than the first threshold, the projection detection frame is considered to have failed to match. The above step S53 may be that the second track information to which the projection frame belongs is deleted.

[0110] Exemplarily, the above step S42 may include: setting a virtual output track sequence (the sequence is empty initially), first matching it with the output target of the 3D perception network, obtaining the track according to the matching result, and generating the initialization (first frame) track for the output of the unmatched 3D perception network. The output of the matched 3D perception network corresponds to the output virtual track, and subsequent matching is performed. The unmatched output virtual track results are deleted, and then the second target sequence is determined. The second target sequence is used as the second trajectory information.

[0111] Corresponding to each virtual track (including the initial frame) that matches the 3D perception result, that is, the second track information, the 3D perception information contained in the second track information, that is, each three-dimensional detection frame, is projected onto each 2D image using the camera calibration parameters to obtain a 2D rectangular frame on each image, that is, the above-mentioned projection frame. Before projecting each corner point of the three-dimensional detection frame to obtain the projection corner point, the above-mentioned calibration parameters are obtained. Exemplarily, the process of projecting each three-dimensional detection frame to obtain the projection frame corresponding to the three-dimensional detection frame satisfies the preset projection relationship. Please refer to formula (1) for the preset projection relationship:

[0112]

[0113] Among them, K can represent the camera intrinsic parameter matrix, which is R 3×3 and t 3×1 The camera extrinsic rotation matrix and translation matrix t can be represented respectively 3×1 . and They can represent the rotation matrix and translation matrix of the world system to the camera system respectively. p(x,y,z) is used to represent the coordinate information of a 3D point, that is, any corner point of the 3D detection box. P(u,v) is used to represent the coordinate information of the 2D pixel point after projection, that is, the projection corner point. c It can represent the normalization coefficient. Where u and v are the coordinate information of the projection corner point. 1×3 The matrix [0,0,0] can be represented.

[0114] Exemplarily, for each three-dimensional detection frame in the second trajectory information, the attributes are the center point and the length, width and height, and 8 corner points are obtained. The 8 corner points are projected to the 2D image to obtain 8 2D points, and the minimum rectangular envelope is obtained for these 8 points. In the projection process, there will be a situation where the projection is outside the image. At this time, the rectangular frame is processed so that it falls within the image range, and two points are selected to fall inside and outside the image respectively. The line connecting it and the edge of the image is obtained, and the point is used to replace the point outside the image, and then the minimum rectangular envelope is solved. For 6 images, each image has a minimum rectangular envelope obtained by projecting each three-dimensional detection frame in the second trajectory information onto the 2D image. And the minimum rectangular envelope is used as the projection frame corresponding to the three-dimensional detection frame. And each three-dimensional detection frame in the second trajectory information is used as a virtual track, and the 3D information of each three-dimensional detection frame in the second trajectory information and the projection frame corresponding to each three-dimensional detection frame. The 3D information of each three-dimensional detection frame includes the size, position, speed, and yaw angle of the three-dimensional detection frame.

[0115] The above step S52 can be a virtual track with 3D information, which is then matched with the 2D results after tracking on the 6 images. The virtual track with 3D information can refer to the projection frame corresponding to each three-dimensional detection frame. Specifically, the target matching result is determined based on the degree of overlap between the projection frame corresponding to each three-dimensional detection frame and the two-dimensional detection frame at the same acquisition angle. Bipartite graph matching is performed using the two-dimensional detection frame at the same acquisition angle and the projection frame corresponding to the three-dimensional detection frame. The Hungarian matching algorithm is used, and the iou of the detection frame is set as the cost matrix to match the target matching result. For example, the same acquisition angle is the acquisition angle of the image acquisition device corresponding to the front view.

[0116] The target matching result can be a 2D and virtual track matching result, a single 2D result, or a single virtual track result. Different target matching results have different output strategies.

[0117] The 2D and virtual track matching result may refer to the above projection frame matching success. Single 2D result and single virtual track result may refer to the above projection frame matching failure. Single 2D result may refer to the case where there is only a target corresponding to the 2D detection frame. Single virtual track result may refer to the case where there is only a target corresponding to the projection detection frame.

[0118] In some application scenarios, if the target matching result is a successful projection frame match, the target is a target with a high degree of confidence, and the 3D attribute information (size, position, speed, yaw angle) in the second trajectory information corresponding to the target is used to update the virtual track information and use it as the target track information. In other application scenarios, for a single 2D result or a single virtual track result, that is, a projection frame match failure, the target corresponding to the projection frame is determined as a false detection target, and the 3D attribute information in the second trajectory information corresponding to the false detection target is not output.

[0119] It can be considered that the target trajectory information determined by the first matching result, the second matching result and the target matching result is the more accurate trajectory information in the second trajectory information. The target trajectory information is output as the final target, and the output attributes are information such as size, position, speed, and yaw angle. The purpose of the above step S13 filtering is to further screen the real target and further prevent the second trajectory information from being misdetected, that is, the target tracked by the 2D perception network and the 3D perception network at the same time is a real and credible target, which can improve the accuracy of the target trajectory information.

[0120] In some embodiments, after the above step S13, the target trajectory information is input into the target function module to determine the output result of the target function module.

[0121] The target function module has a target function, which includes but is not limited to planning and controlling the target vehicle, dynamic obstacle avoidance, parking planning and other functions. The output result of the target function module can be an output result related to the target function. Exemplarily, the target function is mainly to perceive obstacles for the target vehicle. The application scenarios of the target trajectory information of the present application include but are not limited to automatic cruise, memory driving, AEB emergency braking, obstacle perception of parking, memory parking and dynamic obstacle avoidance.

[0122] The above scheme, based on collecting several frames of images of the environment in which the target vehicle is located, generates first trajectory information and second trajectory information of the target object respectively. The first trajectory information and the second trajectory information have different spatial dimension information. The target object is an obstacle in the environment. The second trajectory information is filtered using the first trajectory information with different spatial dimension information to obtain the target trajectory information of the target object, so that the filtered target trajectory information is more accurate.

[0123] See also Figure 7 , Figure 7 It is a structural diagram of an embodiment of the trajectory information determination device of the present application. The trajectory information determination device 70 includes an acquisition module 71, a generation module 72 and a filtering module 73. The acquisition module 71 is used to acquire a plurality of frames of images of the environment in which the target vehicle is located at the current moment acquired by at least one image acquisition device in the target vehicle; the generation module 72 is used to generate the first trajectory information and the second trajectory information of the target object based on the plurality of frames of images, respectively, the first trajectory information and the second trajectory information have different spatial dimensions, and the target object is an obstacle in the environment; the filtering module 73 is used to filter the second trajectory information using the first trajectory information to obtain the target trajectory information of the target object.

[0124] In the above scheme, after obtaining the working status information of the target object at multiple times, supplementary information is determined based on the working status information, and the supplementary information at least includes candidate temperatures obtained by determining the trajectory information of the target object under different prediction modes, which can carry the temperature indication information of the target object under different prediction modes. Subsequently, the feature information of the initial features obtained by combining the working status information and the supplementary information is richer, and the target trajectory information determination model is used to predict and process the initial features with richer feature information, so that the trajectory information determination result is more accurate.

[0125] For the functions performed by each module, please refer to the trajectory information determination method, which will not be repeated here.

[0126] See also Figure 8 , Figure 8 80 is a schematic diagram of the structure of an embodiment of an electronic device of the present application. The electronic device 80 includes a memory 81 and a processor 82, and the processor 82 is used to execute program instructions stored in the memory 81 to implement the steps in the above-mentioned trajectory information determination method embodiment. In a specific implementation scenario, the electronic device 80 may include but is not limited to: a microcomputer, a server, and in addition, the electronic device 80 may also include a mobile device such as a laptop computer and a tablet computer, which is not limited here.

[0127] Specifically, the processor 82 is used to control itself and the memory 81 to implement the steps in the above-mentioned trajectory information determination method embodiment. The processor 82 can also be called a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with signal processing capabilities. The processor 82 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 82 can be implemented by an integrated circuit chip.

[0128] In the above scheme, after obtaining the working status information of the target object at multiple times, supplementary information is determined based on the working status information, and the supplementary information at least includes candidate temperatures obtained by determining the trajectory information of the target object under different prediction modes, which can carry the temperature indication information of the target object under different prediction modes. Subsequently, the feature information of the initial features obtained by combining the working status information and the supplementary information is richer, and the target trajectory information determination model is used to predict and process the initial features with richer feature information, so that the trajectory information determination result is more accurate.

[0129] See also Fig. 9 , Fig. 9 The computer-readable storage medium 90 stores program instructions 901, which implement the steps in any of the above-mentioned trajectory information determination method embodiments when executed by a processor.

[0130] In the above scheme, after obtaining the working status information of the target object at multiple times, supplementary information is determined based on the working status information, and the supplementary information at least includes candidate temperatures obtained by determining the trajectory information of the target object under different prediction modes, which can carry the temperature indication information of the target object under different prediction modes. Subsequently, the feature information of the initial features obtained by combining the working status information and the supplementary information is richer, and the target trajectory information determination model is used to predict and process the initial features with richer feature information, so that the trajectory information determination result is more accurate.

[0131] In some embodiments, the functions or modules included in the system provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0132] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0133] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0134] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0135] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.

Claims

1. A method for determining trajectory information, characterized in that: The method comprises: Acquire a plurality of frames of images of the environment where the target vehicle is located captured by at least one image acquisition device in the target vehicle at the current moment; Based on the plurality of frames of images, first trajectory information and second trajectory information of a target object are respectively generated, wherein the first trajectory information and the second trajectory information have different spatial dimensions, and the target object is an obstacle in the environment; The second trajectory information is filtered using the first trajectory information to obtain target trajectory information of the target object.

2. The method according to claim 1, characterized in that The first trajectory information includes at least one two-dimensional detection box related to the target object, the second trajectory information includes at least one three-dimensional detection box related to the target object, and the filtering of the second trajectory information by using the first trajectory information to obtain the target trajectory information of the target object includes: For each of the three-dimensional detection frames, the following steps are performed: Projecting the three-dimensional detection frame to the spatial dimension where the first trajectory information is located to obtain a projection frame corresponding to the three-dimensional detection frame; Obtaining a target matching result of the projection frame according to the degree of overlap between the projection frame and each of the two-dimensional detection frames, wherein the target matching result includes a successful matching of the projection frame or a failed matching of the projection frame; In response to the target matching result being that the projection frame matches successfully, taking the second track information to which the projection frame belongs as the target track information of the target object; or, In response to the target matching result being that the projection frame fails to match, the second track information to which the projection frame belongs is deleted.

3. The method according to claim 2, characterized in that The image acquisition devices have different acquisition angles, and the three-dimensional detection frame is projected to the space dimension where the first trajectory information is located to obtain a projection frame corresponding to the three-dimensional detection frame, including: Acquire calibration parameters corresponding to each of the image acquisition devices, each of the calibration parameters representing a relative relationship between an image coordinate system and a world coordinate system of an image acquired by each of the image acquisition devices; Determining the projection corner points at each acquisition angle according to the initial corner points corresponding to the three-dimensional detection frame and the calibration parameters; The projection frame is determined based on the projection corner points at each of the acquisition angles.

4. The method according to claim 3, characterized in that The determining the projection frame based on the projection corner points at the acquisition angles includes: Selecting a target projection corner point from the projection corner points at each of the acquisition angles, wherein at least some of the target projection corner points are within a target image range, and the target image range is an image range corresponding to the acquisition angle to which the target projection corner point belongs; In response to all the projection corner points of the target projection corner points being within the target image range, taking the projection rectangular frame corresponding to the target projection corner points as the projection frame; or, In response to at least one of the target projection corner points not being within the target image range, a new projection rectangular frame is determined based on the projection corner points not being within the target image range, and the new projection rectangular frame is used as the projection frame.

5. The method according to any one of claims 1 to 4, characterized in that: The generating first trajectory information and second trajectory information of the target object based on the plurality of frame images respectively comprises: Performing a first detection process and a second detection process on each of the images respectively, to obtain a first initial target sequence and a second initial target sequence corresponding to each of the images, wherein the space dimension of each of the first initial target sequences is two-dimensional, and the space dimension of each of the second initial target sequences is three-dimensional; Determine, according to a difference between a preset target sequence and the first initial target sequence and / or the second target sequence, a first target sequence corresponding to the first initial target sequence and / or a second target sequence corresponding to the second initial target sequence, wherein the preset target sequence is a target sequence of several frames of historical images collected at a preset time before a current time; Using the first initial target sequence or the first target sequence as the first trajectory information; The second initial target sequence or the second target sequence is used as the second trajectory information.

6. The method according to claim 5, characterized in that The preset target sequence includes a first preset target sequence with a two-dimensional spatial dimension, the first preset target sequence includes a plurality of first preset detection frames, the first initial target sequence includes a first initial detection frame of each target object, and determining, according to a difference between the preset target sequence and the first initial target sequence and / or the second target sequence, a first target sequence corresponding to the first initial target sequence and / or a second target sequence corresponding to the second initial target sequence, comprises: Determining a first matching result according to the overlap degree between each of the first preset detection frames and each of the first initial detection frames, the first matching result including at least one of a successful match of at least one of the first initial detection frames, a failed match of at least one of the first initial detection frames, and a failed match of at least one of the first preset detection frames; The first initial target sequence is updated according to the first matching result to obtain a first target sequence corresponding to the first initial target sequence.

7. The method according to claim 5, characterized in that The preset target sequence includes a second preset target sequence with a two-dimensional spatial dimension, the second preset target sequence includes a plurality of second preset detection frames, the second initial target sequence includes a second initial detection frame of each target object, and determining a first target sequence corresponding to the first initial target sequence and / or a second target sequence corresponding to the second initial target sequence according to a difference between the preset target sequence and the first initial target sequence and / or the second target sequence, comprises: determining a second matching result according to a spatial distance between each of the second preset detection frames and each of the second initial detection frames, the second matching result including at least one of a successful match of at least one of the second initial detection frames, a failed match of at least one of the second initial detection frames, and a failed match of at least one of the second preset detection frames; The second initial target sequence is updated according to the second matching result to obtain a second target sequence corresponding to the second initial target sequence.

8. A trajectory information determination device, characterized in that: include: An acquisition module, used to acquire a number of frames of images captured by at least one image acquisition device in the target vehicle at the current moment of the environment where the target vehicle is located; A generating module, configured to generate first trajectory information and second trajectory information of a target object based on the plurality of frame images, respectively, wherein the first trajectory information and the second trajectory information have different spatial dimensions, and the target object is an obstacle in the environment; The filtering module is used to filter the second trajectory information by using the first trajectory information to obtain the target trajectory information of the target object.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 7.