Space grid corresponding motion state prediction method and device, and autonomous vehicle
By combining current-time images with historical motion state information through iterative processing, the motion state of the surrounding spatial grid of autonomous vehicles is predicted, solving the problem of high cost of radar point cloud data acquisition and achieving a balance between accuracy and cost.
Patent Information
- Application Number
- CN202411787039.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-12-05
AI Technical Summary
Existing autonomous vehicles use radar to collect point cloud data for obstacle perception and prediction, which results in high costs. How can we reduce costs while ensuring accuracy?
By iteratively processing based on current-time images and historical motion state prediction information, the motion state of the spatial grid within the current space is predicted. This reduces costs by using image sensors to collect current-time images and combining them with historical motion state prediction information.
It achieves accurate prediction of the motion state of the surrounding spatial grid at the current moment, reducing costs.
Smart Images

Figure CN119810790B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of autonomous driving, in particular to the technical field of image processing, environment perception, etc. BACKGROUND
[0002] In the field of autonomous driving, the accuracy of the perception of surrounding obstacles by an autonomous vehicle plays a crucial role in the safety of the autonomous vehicle driving. In the prior art, the autonomous vehicle usually uses point cloud data collected by a radar to perceive or predict obstacles in the surrounding space, so as to determine the relevant features or relevant information of the surrounding space. However, this approach requires the use of a laser radar, resulting in high cost. Therefore, how to reduce the cost while ensuring the accurate prediction of the relevant features or relevant information of the surrounding space by the autonomous vehicle has become a technical problem to be solved. SUMMARY
[0003] The present disclosure provides a method and device for predicting the motion state of a spatial grid, an apparatus, a storage medium, and an autonomous vehicle.
[0004] According to an aspect of the present disclosure, a method for predicting the motion state of a spatial grid is provided, comprising:
[0005] determining feature vectors of a plurality of spatial grids in a first space based on a plurality of images of the first space collected at a current time;
[0006] performing multiple iterations of processing the feature vectors of the plurality of spatial grids in the first space and motion state prediction information corresponding to a plurality of spatial grids in a second space to obtain motion state prediction information corresponding to the plurality of spatial grids in the first space, wherein the motion state prediction information corresponding to the plurality of spatial grids in the second space is obtained at a historical time.
[0007] According to an aspect of the present disclosure, a device for predicting the motion state of a spatial grid is provided, comprising:
[0008] a feature extraction module configured to determine feature vectors of a plurality of spatial grids in a first space based on a plurality of images of the first space collected at a current time;
[0009] a prediction module configured to perform multiple iterations of processing the feature vectors of the plurality of spatial grids in the first space and motion state prediction information corresponding to a plurality of spatial grids in a second space to obtain motion state prediction information corresponding to the plurality of spatial grids in the first space, wherein the motion state prediction information corresponding to the plurality of spatial grids in the second space is obtained at a historical time.
[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising:
[0011] at least one processor; and
[0012] a memory communicatively connected with the at least one processor; wherein
[0013] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.
[0014] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method according to any of the embodiments of the present disclosure.
[0015] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.
[0016] According to another aspect of the present disclosure, an autonomous vehicle is provided, comprising the electronic device described above.
[0017] By adopting the method provided in the embodiment, the motion state of the plurality of spatial grids in the first space at the current time is predicted by using the plurality of images at the current time and the motion state prediction information at the historical time, so that the motion state of each spatial grid in the current space can be predicted by using only the images collected at the current time and in combination with the motion state prediction information at the historical time. Since the prediction result at the historical time is combined in the prediction process, the accuracy of the prediction of the motion state of the surrounding spatial grid at the current time can be ensured, and since only the images at the current time need to be collected to realize the prediction of the motion state of each spatial grid in the current space, the cost is reduced.
[0018] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:
[0020] Figure 1 is a flowchart of a method for predicting the motion state corresponding to the spatial grid according to an embodiment of the present disclosure;
[0021] Figure 2 is a flowchart of a method for predicting the motion state corresponding to the spatial grid according to another embodiment of the present disclosure;
[0022] Figure 3 is a schematic block diagram of a device for predicting motion state of a spatial grid according to an embodiment of the present disclosure;
[0023] Figure 4 is a schematic block diagram of a device for predicting motion state of a spatial grid according to another embodiment of the present disclosure;
[0024] Figure 5 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, in which various details of embodiments of the present disclosure are set forth to assist in the understanding of the present disclosure. It should be appreciated that the embodiments described herein are merely exemplary and are not intended to limit the scope of the present disclosure. It will be readily apparent to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted herein.
[0026] Figure 1 is a schematic flowchart of a method for predicting motion state of a spatial grid according to an embodiment of the present disclosure, comprising:
[0027] S110, determining feature vectors of a plurality of spatial grids in a first space based on a plurality of images of the first space collected at a current time.
[0028] S120, performing multiple iterations of processing on the feature vectors of the plurality of spatial grids in the first space and motion state prediction information corresponding to a plurality of spatial grids in a second space to obtain motion state prediction information corresponding to the plurality of spatial grids in the first space, wherein the motion state prediction information corresponding to the plurality of spatial grids in the second space is obtained at a historical time.
[0029] The above method for predicting motion state of a spatial grid can be implemented by an electronic device. The electronic device can be a terminal or a server. Exemplarily, the electronic device can be a vehicle-mounted terminal or other terminal device having computing capability and capable of communicating with a vehicle. Exemplarily, the server can be a single server capable of communicating with a vehicle or one or more servers in a server cluster. It should be understood that the above are only exemplary descriptions of execution subjects capable of executing the method for predicting motion state of a spatial grid provided by the present embodiment, and in actual processing, the electronic device capable of executing the method for predicting motion state of a spatial grid provided by the present embodiment is not limited to the types of devices mentioned in the above examples.
[0030] By adopting the method provided in the embodiment, the motion state of the plurality of spatial grids in the first space at the current time is predicted by the plurality of images at the current time and the motion state prediction information at the historical time, so that the motion state of each spatial grid in the current space can be predicted by using only the images collected at the current time and in combination with the motion state prediction information at the historical time. Since the prediction result at the historical time is combined in the prediction process, the accuracy of the prediction of the motion state of the surrounding spatial grid at the current time can be ensured, and since the prediction of the motion state of each spatial grid in the current space can be realized only by collecting the images at the current time, the cost is reduced.
[0031] In an embodiment, the manner of acquiring the plurality of images of the first space collected at the current time can include: acquiring the plurality of images of the first space collected by the image sensor at the current time.
[0032] It should be noted that the electronic device side executing the method for predicting the motion state of the spatial grid can pre-set the acquisition interval between the acquisition time, and the specific duration of the acquisition interval can be configured according to actual conditions, such as 0.1 seconds, or 1 second, or longer or shorter, which is not limited here.
[0033] The current time can refer to the current acquisition time, that is, the electronic device can control the image sensor to collect, and the image sensor will perform collection once every acquisition interval. Specifically, acquiring the plurality of images of the first space collected by the image sensor at the current time can refer to that the electronic device controls the image sensor to collect at the current time, and acquires the plurality of images of the first space collected by the image sensor at the current time.
[0034] The first space can refer to the space where the vehicle is currently located, or the space where the vehicle is located at the current time.
[0035] The range of the first space can be related to the pre-set acquisition range and the sensing range of the image sensor. The pre-set acquisition range can be configured according to actual conditions, such as the pre-set acquisition range can be a range of 120 meters vertically and 120 meters horizontally and 3 meters vertically with the image sensor as the origin.
[0036] The image sensor can be any kind of device or sensor with image collection function arranged on the vehicle.
[0037] The number of image sensors can be one or more.
[0038] Exemplarily, the number of image sensors can be multiple, and the collection angles of different image sensors in the multiple image sensors can be different. For example, the multiple image sensors can include multiple cameras or multiple camera heads arranged on the vehicle, and the collection angles of different cameras or different camera heads are different.
[0039] In this case, the multiple images of the first space can be two-dimensional images of at least part of the range in the first space collected by each image sensor in the multiple image sensors. Since different image sensors in the multiple image sensors can have different collection angles, each image sensor can only collect two-dimensional images in part of the range of the first space where the vehicle is currently located, different image sensors can collect two-dimensional images in different ranges in the first space where the vehicle is currently located, and the two-dimensional images collected by different image sensors can have partially overlapping ranges, but all image sensors can collect all two-dimensional images of all ranges in the first space where the vehicle is currently located, that is, the collection of two-dimensional images collected by all image sensors can cover all ranges of the first space. How all image sensors can collect all two-dimensional images covering all ranges of the first space is related to the setting of the collection angle of each image sensor, which is not limited here.
[0040] Exemplarily, the number of image sensors can be one, for example, the image sensor can include a surround view camera arranged on the vehicle. In this case, the multiple images of the first space can be multiple two-dimensional images in the surround view of the first space collected by the image sensor. It should be understood that each two-dimensional image can be an image of part of the range of the first space where the vehicle is located, and different two-dimensional images contain images of different ranges of the first space, but all two-dimensional images contain all ranges of the first space.
[0041] In an embodiment, the feature vectors of the multiple spatial grids in the first space are determined based on the multiple images of the first space collected at the current time, including: performing feature extraction on each pixel point in each image in the multiple images of the first space collected at the current time to obtain a feature vector of each pixel point in each image; and obtaining a feature vector of each spatial grid in the multiple spatial grids in the first space based on the feature vector of each pixel point in each image. The feature vector can be a semantic vector, which is used to describe shapes, textures, colors, etc., which are not listed one by one here.
[0042] The feature extraction of each pixel point in each image of the plurality of images of the first space collected at the current moment can include: performing feature extraction on each pixel point in each image of the plurality of images of the first space collected at the current moment through a feature extraction network to obtain a feature vector of each pixel point in each image. The feature extraction network can be a ResNet (Residual Network), and specifically, the ResNet can include one of ResNet50 and Resnet101.
[0043] For example, the feature extraction of each pixel point in each image of the plurality of images of the first space collected at the current moment through the feature extraction network to obtain a feature vector of each pixel point in each image can include: inputting the mth image of the first space collected at the current moment into the ResNet to obtain a feature vector of each pixel point in the mth image output by the ResNet, where m is an integer greater than or equal to 1, and the mth image is any one of the plurality of images. Since the processing of each image is the same as that of the mth image, it will not be described one by one.
[0044] The feature vector of each spatial grid in the first space can be obtained based on the feature vector of each pixel point in each image. The plurality of spatial grids in the first space can be generated based on the first space range and a spatial grid parameter. The feature vector of one or more pixel points corresponding to each spatial grid in the first space can be obtained by mapping the feature vector of each pixel point in each image to the corresponding spatial grid in the first space. The feature vector of each spatial grid in the first space can be obtained based on the feature vector of one or more pixel points corresponding to each spatial grid in the first space.
[0045] The spatial grid parameter can include at least one of a size of the spatial grid and a number of the spatial grid. The size of the spatial grid and the number of the spatial grid can be set according to actual conditions, and the size of the spatial grid can also be referred to as the resolution of the spatial grid.
[0046] The mapping of the feature vector of each pixel point in each image to the corresponding spatial grid in the first space to obtain the feature vector of one or more pixel points corresponding to each spatial grid in the first space can include: determining the corresponding spatial grid of each pixel point in the first space based on the correspondence between the coordinates of each image and the coordinates of the first space; and mapping the feature vector of each pixel point in each image to the corresponding spatial grid of each pixel point in the first space to obtain the feature vector of one or more pixel points corresponding to each spatial grid in the first space. The coordinates of each image can be coordinates in an image coordinate system, and the specific setting of the image coordinate system is not limited in the embodiment. The coordinates of the first space can be coordinates in a spatial coordinate system, and the spatial coordinate system can be a coordinate system with the vehicle as the origin, the horizontal plane as the xy plane, and the vertical line of the horizontal plane as the z axis. It should be noted that this is only an exemplary description of the spatial coordinate system, and the spatial coordinate system can also be set in other ways in actual processing, which is not limited or exhaustive. The correspondence between the coordinates of each image and the coordinates of the first space can refer to the correspondence between the coordinates of each image in the image coordinate system and the coordinates of the first space in the spatial coordinate system. The generation and specific content of the correspondence are not limited in the embodiment.
[0047] Taking the zth pixel point in the kth image (k is a positive integer) in the plurality of images as an example, the mapping of the feature vector of each pixel point in each image to the corresponding spatial grid in the first space to obtain the feature vector of one or more pixel points corresponding to each spatial grid in the first space can include: mapping the feature vector of the zth pixel point in the kth image to the nth spatial grid corresponding to the zth pixel point in the first space to obtain the feature vector of one pixel point corresponding to the nth spatial grid, and n is a positive integer.
[0048] The obtaining of the feature vector of each spatial grid in the first space based on the feature vector of one or more pixel points corresponding to each spatial grid in the first space can include: calculating the feature vector of one or more pixel points corresponding to each spatial grid in the first space based on a specified calculation method to obtain the feature vector of each spatial grid in the first space. The specified calculation method can be one of the following: average calculation, maximum calculation, etc. It should be understood that this is only an exemplary description, and the specified calculation method can also include other related statistical methods in actual processing, which is not limited or exhaustive.
[0049] In an example, the generating the plurality of spatial grids in the first space based on the first spatial range and the spatial grid parameter, and mapping the feature vector of each pixel in each image to the corresponding spatial grid in the first space to obtain the feature vector of one or more pixels corresponding to each spatial grid in the first space can be implemented by a spatial generation network. Specifically, the feature vector of each pixel in each image can be input into the spatial generation network to obtain the feature vector of each spatial grid in the plurality of spatial grids in the first space output by the spatial generation network. The spatial generation network can be a Cam2voxel (Camera 2 voxel) network.
[0050] In an example, the mapping the feature vector of each pixel in each image to the corresponding spatial grid in the first space to obtain the feature vector of one or more pixels corresponding to each spatial grid in the first space can be implemented by a first mapping model. Specifically, the feature vector of each pixel in each image can be input into the first mapping model to obtain the feature vector of one or more pixels corresponding to each spatial grid in the first space output by the first mapping model. The first mapping model can be trained based on a Transformer model, and the specific training method of the Transformer model is not limited in the present application.
[0051] In an example, the mapping the feature vector of each pixel in each image to the corresponding spatial grid in the first space to obtain the feature vector of one or more pixels corresponding to each spatial grid in the first space; and obtaining the feature vector of each spatial grid in the first space based on the feature vector of one or more pixels corresponding to each spatial grid in the first space can be implemented by a second mapping model. Specifically, the feature vector of each pixel in each image can be input into the second mapping model to obtain the feature vector of each spatial grid in the first space output by the second mapping model. The second mapping model can be trained based on a Transformer model, and the specific training method of the Transformer model is not limited in the present application.
[0052] In an embodiment, the plurality of times of iteration processing of the feature vectors of the plurality of spatial grids in the first space and the corresponding motion state prediction information of the plurality of spatial grids in the second space to obtain the corresponding motion state prediction information of the plurality of spatial grids in the first space comprises: in the i th iteration processing of the plurality of times of iteration processing, generating the to-be-processed information of the i th iteration processing based on the feature vectors of the plurality of spatial grids in the first space and the corresponding motion state prediction information of the plurality of spatial grids in the second space, wherein i is a positive integer; inputting the to-be-processed information of the i th iteration processing and the i th first condition vector into a target model to obtain the i th motion state prediction information of the plurality of spatial grids in the first space output by the target model; and in the case that the i th iteration processing satisfies a constraint condition, taking the i th motion state prediction information of the plurality of spatial grids in the first space as the corresponding motion state prediction information of the plurality of spatial grids in the first space.
[0053] The motion state comprises at least one of the following: velocity, orientation, acceleration, and the like, and the i th motion state prediction information comprises at least one of the following: i th predicted velocity, i th predicted orientation, i th predicted acceleration, and the like, which are not listed one by one here.
[0054] The i th motion state prediction information of any one of the spatial grids in the first space can be used to represent at least one of the following: i th predicted velocity, i th predicted orientation, i th predicted acceleration, and the like, of the occupant object corresponding to the spatial grid in the first space. The occupant object can comprise one of the following: obstacle vehicle, pedestrian, green belt, building, and the like.
[0055] It should be understood that the occupant objects corresponding to different spatial grids in the first space can be the same or different, and it is not limited here whether different spatial grids correspond to the same occupant object. For each spatial grid, at least one of the i th predicted velocity, i th predicted orientation, i th predicted acceleration, and the like, of the corresponding occupant object can be determined independently.
[0056] The number of the second spaces can be one or more.
[0057] If the number of the second spaces is one, the second space can refer to the space in which the vehicle is located at a historical time. The historical time can refer to the last collection time adjacent to the current time. The second space can be the same as or at least partially different from the first space. Specifically, the plurality of spatial grids of the second space can be the same as the plurality of spatial grids of the first space; or the plurality of spatial grids of the second space can be partially the same as and at least partially different from the plurality of spatial grids of the first space.
[0058] If the number of the second spaces is multiple, each of the multiple second spaces refers to a space in which the vehicle is located at each historical moment, wherein different second spaces correspond to different historical moments, that is, the multiple second spaces correspond to multiple historical moments, the last historical moment in the multiple historical moments is the last collection moment adjacent to the current moment, and the interval between two adjacent historical moments in the multiple historical moments is equal to the collection interval. The collection interval is described above in the same manner and will not be repeated.
[0059] In the multiple second spaces, different second spaces can be the same or at least partially different, specifically, multiple space grids of different second spaces can be the same; or multiple space grids of different second spaces can be partially the same and at least partially different. In addition, any one of the multiple second spaces can be the same or at least partially different from the first space, specifically, multiple space grids of any one of the second spaces can be the same as multiple space grids of the first space; or multiple space grids of any one of the second spaces can be partially the same and at least partially different from multiple space grids of the first space.
[0060] The motion state prediction information corresponding to any one of the space grids in any one of the second spaces can be used to represent at least one of a predicted speed, a predicted direction, a predicted acceleration, etc. of the occupying object corresponding to the space grid in the second space.
[0061] Optionally, the generating of the to-be-processed information of the i th iteration processing based on the feature vectors of the multiple space grids in the first space and the motion state prediction information corresponding to the multiple space grids in the second space can include one of the following: taking the feature vectors of the multiple space grids in the first space and the motion state prediction information corresponding to the multiple space grids in the second space as the to-be-processed information of the i th iteration processing; and taking the feature vectors of the multiple space grids in the first space, the motion state prediction information corresponding to the multiple space grids in the second space, and the feature vectors of the multiple space grids in the second space as the to-be-processed information of the i th iteration processing.
[0062] The target model can be obtained by training a preset model, and the preset model can be a diffusion model. In this embodiment, the target model is used to perform a denoising process on the to-be-processed content, which can also be referred to as a reverse denoising process. The reverse denoising process can make the to-be-processed content clearer or more accurate, and then predict prediction information related to the to-be-processed content based on the clearer or more accurate to-be-processed content.
[0063] The inputting the information to be processed in the i-th iteration processing and the i-th first condition vector into a target model to obtain i-th motion state prediction information corresponding to a plurality of spatial grids in the first space output by the target model can include: inputting the information to be processed in the i-th iteration processing and the i-th first condition vector into a target model to obtain i-th prediction information corresponding to a plurality of spatial grids in the first space output by the target model, wherein the i-th prediction information corresponding to a plurality of spatial grids in the first space includes i-th motion state prediction information corresponding to a plurality of spatial grids in the first space.
[0064] In a case where i is equal to 1, the i-th first condition vector is a random noise vector; in a case where i is greater than 1, the i-th first condition vector is i-1-th motion state prediction information corresponding to a plurality of spatial grids in the first space. The random noise vector can be obtained according to Gaussian operation coding, and the specific obtaining manner is not limited by the present application.
[0065] Optionally, the generating the information to be processed in the i-th iteration processing based on the feature vector of the plurality of spatial grids in the first space and the motion state prediction information corresponding to a plurality of spatial grids in the second space can include one of: taking the feature vector of the plurality of spatial grids in the first space, the motion state prediction information corresponding to a plurality of spatial grids in the second space, and the occupancy state prediction information corresponding to a plurality of spatial grids in the second space as the information to be processed in the i-th iteration processing; and taking the feature vector of the plurality of spatial grids in the first space, the motion state prediction information corresponding to a plurality of spatial grids in the second space, the occupancy state prediction information corresponding to a plurality of spatial grids in the second space, and the feature vector of the plurality of spatial grids in the second space as the information to be processed in the i-th iteration processing.
[0066] The inputting the information to be processed in the i-th iteration processing and the i-th first condition vector into a target model to obtain i-th motion state prediction information corresponding to a plurality of spatial grids in the first space output by the target model can include: inputting the information to be processed in the i-th iteration processing and the i-th first condition vector into a target model to obtain i-th prediction information corresponding to a plurality of spatial grids in the first space output by the target model, wherein the i-th prediction information corresponding to a plurality of spatial grids in the first space includes i-th motion state prediction information corresponding to a plurality of spatial grids in the first space, i-th occupancy state prediction information corresponding to a plurality of spatial grids in the first space.
[0067] In a case where i is equal to 1, the i th first conditional vector is a random noise vector; in a case where i is greater than 1, the i th first conditional vector is i-1 th motion state prediction information corresponding to a plurality of spatial grids in the first space and i-1 th occupancy state prediction information corresponding to the plurality of spatial grids in the first space.
[0068] Taking a g th spatial grid (g is a positive integer) in the plurality of spatial grids in the first space as an example, the i th occupancy state prediction information corresponding to the plurality of spatial grids in the first space can include: prediction information about whether the g th spatial grid in the first space is occupied in the i th time.
[0069] In a case where the prediction information about whether the g th spatial grid in the first space is occupied in the i th time is occupation, the i th occupancy state prediction information corresponding to the g th spatial grid in the first space can further include: occupancy attribute corresponding to the g th spatial grid in the first space. The occupancy attribute includes at least one of the following: texture, shape, color, occupancy object, and the like, which are not listed one by one here.
[0070] In a case where the i th iteration processing meets the constraint condition, the i th motion state prediction information corresponding to the plurality of spatial grids in the first space is taken as the motion state prediction information corresponding to the plurality of spatial grids in the first space, including: judging whether the i th iteration processing meets the constraint condition, and in a case where the i th iteration processing meets the constraint condition, taking the i th motion state prediction information corresponding to the plurality of spatial grids in the first space as the motion state prediction information corresponding to the plurality of spatial grids in the first space.
[0071] The manner of determining that the i th iteration processing meets the constraint condition includes at least one of the following: in a case where the i th iteration processing is the last iteration processing, it is determined that the i th iteration processing meets the constraint condition; and in a case where the i th occupancy state prediction information corresponding to the plurality of spatial grids in the first space meets the quality requirement, it is determined that the i th iteration processing meets the constraint condition.
[0072] In a case where the i th iteration processing is the last iteration processing, it is determined that the i th iteration processing meets the constraint condition; and in a case where the i th occupancy state prediction information corresponding to the plurality of spatial grids in the first space meets the quality requirement, it is determined that the i th iteration processing meets the constraint condition.
[0073] The manner of determining whether the i-th occupation state prediction information corresponding to the plurality of spatial grids in the first space satisfies the quality requirement comprises: in a case where a related parameter of the i-th occupation state prediction information corresponding to the plurality of spatial grids in the first space satisfies a preset condition, determining that the i-th occupation state prediction information corresponding to the plurality of spatial grids in the first space satisfies the quality requirement. The related parameter comprises at least one of a mean square error, a peak signal-to-noise ratio, brightness, contrast, and the like.
[0074] In addition, the method can further comprise: in a case where the i-th iteration processing does not satisfy the constraint condition, performing an (i+1)-th iteration processing of the plurality of iteration processings. The (i+1)-th iteration processing of the plurality of iteration processings has the same manner as the i-th iteration processing of the plurality of iteration processings, which will not be described herein.
[0075] It should be noted that, after obtaining the motion state prediction information corresponding to the plurality of spatial grids in the first space, the motion state prediction information corresponding to each spatial grid in the first space can be directly used for prediction processing at a next time.
[0076] Alternatively, after obtaining the motion state prediction information corresponding to the plurality of spatial grids in the first space, the following processing can be performed: determining one or more spatial grids corresponding to each occupied object in the first space based on the occupied object corresponding to each spatial grid in the plurality of spatial grids in the first space, obtaining motion state prediction information of each occupied object based on the motion state prediction information corresponding to each spatial grid corresponding to each occupied object, and updating the motion state prediction information corresponding to each spatial grid in the first space based on the motion state prediction information of each occupied object.
[0077] Taking any one occupied object as an example, obtaining the motion state prediction information of the object based on the motion state prediction information corresponding to each spatial grid corresponding to each object can comprise at least one of the following: obtaining a predicted velocity of the occupied object based on a predicted velocity in the motion state prediction information corresponding to each spatial grid corresponding to the occupied object; obtaining a predicted orientation of the occupied object based on a predicted orientation in the motion state prediction information corresponding to each spatial grid corresponding to the occupied object; and obtaining a predicted acceleration of the occupied object based on a predicted acceleration in the motion state prediction information corresponding to each spatial grid corresponding to the occupied object. The specific manner of obtaining the predicted velocity, the predicted orientation, and the predicted acceleration of the occupied object can be configured according to actual conditions, such as calculating an average value, or taking a maximum value, or taking a minimum value, or the like, which will not be limited or enumerated herein.
[0078] Still taking any one occupying object as an example, updating the motion state prediction information corresponding to each spatial grid in the first space based on the motion state prediction information of each occupying object can be that the motion state prediction information corresponding to each spatial grid in the first space corresponding to the occupying object is set to the motion state prediction information of the occupying object.
[0079] In this way, the feature vectors of the plurality of spatial grids in the first space at the current time and the motion state prediction information of the plurality of spatial grids in the second space obtained at the historical time can be iteratively processed by the target model multiple times, and finally the motion state prediction information of the plurality of spatial grids in the first space at the current time is obtained. In this way, the accuracy and efficiency of obtaining the motion state prediction information corresponding to the plurality of spatial grids in the first space can be improved by the target model executing multiple iterations. Moreover, the output result of the last iteration is referred to in each iteration, so that the accuracy of obtaining the motion state prediction information corresponding to the plurality of spatial grids in the first space can be further improved.
[0080] In an embodiment, in the case where there is no historical time (i.e. the first acquisition time in the current driving task is executed by the target vehicle), the method further comprises: setting an initial speed as the motion state prediction information of the plurality of spatial grids in the first space at the current time. Wherein, the initial speed can be set according to actual conditions, for example, the initial speed can be 0 or set to empty.
[0081] In combination Figure 2 The above embodiments are exemplarily described, including:
[0082] S201, obtaining a plurality of images of the first space acquired by the image sensor at the current time (i.e. a plurality of two-dimensional images in the surround view of the first space acquired by the surround camera).
[0083] S202, performing feature extraction on each pixel point in each image of the plurality of images of the first space acquired at the current time to obtain a feature vector of each pixel point in each image.
[0084] S203, inputting the feature vector of each pixel point in each image of the plurality of images of the first space into a space generation network (i.e. Cam2voxel) to obtain the feature vector of each spatial grid in the plurality of spatial grids in the first space output by the space generation network.
[0085] S204, performing multiple iterations on the feature vectors of the plurality of spatial grids in the first space and the motion state prediction information corresponding to the plurality of spatial grids in the second space to obtain the motion state prediction information corresponding to the plurality of spatial grids in the first space.
[0086] The specific process of the plurality of iteration processes includes:
[0087] S2041, in the i-th iteration process of the plurality of iteration processes, based on the feature vectors of the plurality of spatial grids in the first space, the motion state prediction information corresponding to the plurality of spatial grids in the second space, generate the to-be-processed information of the i-th iteration process; input the to-be-processed information of the i-th iteration process and the i-th first condition vector into the target model to obtain the i-th motion state prediction information corresponding to the plurality of spatial grids in the first space output by the target model. Wherein, in the case that i is equal to 1, the i-th first condition vector is a random noise vector; in the case that i is greater than 1, the i-th first condition vector is the (i-1)-th motion state prediction information corresponding to the plurality of spatial grids in the first space.
[0088] S2042, judge whether the i-th iteration process meets the constraint condition, in the case that the constraint condition is met, the i-th motion state prediction information corresponding to the plurality of spatial grids in the first space is taken as the motion state prediction information corresponding to the plurality of spatial grids in the first space, and the process is ended; in the case that the constraint condition is not met, return to S2041 to execute the (i+1)-th iteration process of the plurality of iteration processes.
[0089] In an embodiment, the method further comprises: based on the sample and the sample label, the preset model is trained for a plurality of times to obtain the target model, wherein the sample includes the feature vectors of the plurality of spatial grids in the third space at the first time and the motion state information corresponding to the plurality of spatial grids in the fourth space at the second time, and the first time is later than the second time.
[0090] The preset model can be a diffusion model.
[0091] Wherein, the acquisition method of the feature vectors of the plurality of spatial grids in the third space at the first time is similar to the acquisition method of the feature vectors of the plurality of spatial grids in the first space, which will not be repeated here. In addition, the feature vectors of the plurality of spatial grids in the third space at the first time can also be acquired by other methods, which are not limited by the present application.
[0092] The first time can refer to the first acquisition time, wherein the related description of the acquisition time is the same as that in the above embodiment, which will not be repeated here.
[0093] The third space can refer to a space where the vehicle is located at the first time, or a space where the vehicle is located at the first time. The range of the third space can be related to a pre-set collection range and a sensing range of the image sensor. The pre-set collection range can be configured according to actual conditions, and the specific configuration manner can be set according to actual conditions. The application is not limited.
[0094] The second time can be one or more.
[0095] If the second time is one, the second time refers to the last collection time adjacent to the first time. The fourth space at the second time can refer to a space where the vehicle is located at the second time. The fourth space can be the same as or at least partially different from the third space. Specifically, the plurality of space grids of the fourth space can be the same as the plurality of space grids of the third space. Alternatively, the plurality of space grids of the fourth space can be partially the same and at least partially different from the plurality of space grids of the third space.
[0096] If the second time is more than one, the second time refers to each historical collection time before the first time. The fourth space at the second time can refer to a space where the vehicle is located at each historical collection time. Different fourth spaces correspond to different historical times, that is, a plurality of fourth spaces correspond to a plurality of historical times. The last historical time in the plurality of historical times is the last collection time adjacent to the first time. The interval between two adjacent historical times in the plurality of historical times is equal to the collection interval. The collection interval is the same as the aforementioned embodiment, and will not be repeated.
[0097] The different fourth spaces in the plurality of fourth spaces can be the same or at least partially different. Specifically, the plurality of space grids of different fourth spaces can be the same. Alternatively, the plurality of space grids of different fourth spaces can be partially the same and at least partially different. In addition, any one of the plurality of fourth spaces can be the same as or at least partially different from the third space. Specifically, the plurality of space grids of any one of the fourth spaces can be the same as the plurality of space grids of the third space. Alternatively, the plurality of space grids of any one of the fourth spaces can be partially the same and at least partially different from the plurality of space grids of the third space.
[0098] The acquisition manner of the motion state information corresponding to the plurality of space grids in the fourth space at the second time is not limited by the application. For example, the motion state information corresponding to the plurality of space grids in the fourth space at the second time can be obtained based on point cloud data collected by a point cloud collection device on the vehicle at the second time, or obtained by other speed detection devices. Here, it is not limited or exhausted. The point cloud collection device can be one of a laser radar, a millimeter wave radar, and the like.
[0099] Thus, the pre-set model is iteratively trained based on the feature vectors of the plurality of spatial grids in the third space at the first time point and the motion state information corresponding to the plurality of spatial grids in the fourth space at the second time point and the sample label, to obtain the target model. In this way, the pre-set model can be trained multiple times, so that the trained pre-set model can predict the motion state information corresponding to the plurality of spatial grids in the space according to the image information, thereby ensuring the accuracy of the prediction of the motion state of the spatial grid by the trained target model.
[0100] In an implementation, the training of the pre-set model based on the sample and the sample label to obtain the target model comprises: in the jth training of the multiple training, inputting the sample and the jth second conditional vector into the pre-set model to obtain the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time point output by the pre-set model, wherein j is a positive integer; determining a jth loss based on the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time point and the sample label; training the pre-set model based on the jth loss to obtain the pre-set model after the jth training; and in the case that the jth training is the last training of the multiple training, taking the pre-set model after the jth training as the target model.
[0101] Optionally, the inputting of the sample and the jth second conditional vector into the pre-set model to obtain the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time point output by the pre-set model comprises: inputting the sample and the jth second conditional vector into the pre-set model to obtain the jth prediction information corresponding to the plurality of spatial grids in the third space at the first time point output by the pre-set model, wherein the jth prediction information corresponding to the plurality of spatial grids in the third space at the first time point comprises the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time point.
[0102] The sample comprises: the feature vectors of the plurality of spatial grids in the third space at the first time point and the motion state information corresponding to the plurality of spatial grids in the fourth space at the second time point.
[0103] Alternatively, the sample comprises: the feature vectors of the plurality of spatial grids in the third space at the first time point, the motion state information corresponding to the plurality of spatial grids in the fourth space at the second time point, and the feature vectors of the plurality of spatial grids in the fourth space at the second time point. The application does not limit the manner of obtaining the feature vectors of the plurality of spatial grids in the fourth space at the second time point.
[0104] The sample label can comprise a motion state sample label. The motion state sample label can comprise a real motion state corresponding to each of a plurality of spatial grids in the third space at the first time. The real motion state corresponding to each of the plurality of spatial grids can comprise at least one of a real position, a real velocity, and a real orientation corresponding to each of the plurality of spatial grids. The real motion state corresponding to each of the plurality of spatial grids can be obtained in any manner.
[0105] In a case where j is equal to 1, the jth second conditional vector is obtained by adding a random noise encoding vector to the sample label. In a case where j is greater than 1, the jth second conditional vector is obtained by adding a random noise vector to a (j-1)th second conditional vector.
[0106] In a case where the sample label is added with a random noise encoding vector to obtain the jth second conditional vector, the following operations can be performed: in a case where the motion state sample label comprises a real motion state corresponding to each of a plurality of spatial grids in the third space, the random noise encoding vector is added to the real motion state corresponding to each of the plurality of spatial grids to obtain the jth second conditional vector; or in a case where the motion state sample label comprises a real motion state corresponding to each of a plurality of spatial grids in the third space, the random noise encoding vector is added to the real motion state corresponding to part of the plurality of spatial grids to obtain the jth second conditional vector.
[0107] In a case where the (j-1)th second conditional vector is added with a random noise vector to obtain the jth second conditional vector, the operation is similar to the above operation, and thus will not be described herein.
[0108] The determining of the jth loss based on the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time and the sample label comprises determining the jth loss based on the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time and the motion state sample label. The jth loss can be obtained by using a cross-entropy loss function, which is not limited in this embodiment.
[0109] Optionally, the inputting of the sample and the jth second conditional vector into the preset model to obtain the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time output by the preset model comprises inputting the sample and the jth second conditional vector into the preset model to obtain jth prediction information corresponding to the plurality of spatial grids in the third space at the first time output by the preset model, wherein the jth prediction information corresponding to the plurality of spatial grids in the third space at the first time comprises the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time and jth occupancy state prediction information corresponding to the plurality of spatial grids in the third space at the first time.
[0110] The sample includes: feature vectors of a plurality of spatial grids in the third space at the first time, motion state information corresponding to the plurality of spatial grids in the fourth space at the second time, and occupancy state information corresponding to the plurality of spatial grids in the fourth space at the second time. The occupancy state information corresponding to the plurality of spatial grids in the fourth space at the second time can be obtained based on point cloud data collected by a point cloud collection device on the vehicle at the second time, or obtained by other occupancy state detection devices, which is not limited or exhausted herein.
[0111] Alternatively, the sample can include: feature vectors of a plurality of spatial grids in the third space at the first time, motion state information corresponding to the plurality of spatial grids in the fourth space at the second time, occupancy state information corresponding to the plurality of spatial grids in the fourth space at the second time, and feature vectors of the plurality of spatial grids in the fourth space at the second time.
[0112] The sample label can include: a motion state sample label and an occupancy state sample label.
[0113] The motion state sample label can be a true motion state corresponding to each spatial grid of the plurality of spatial grids in the third space. The true motion state corresponding to each spatial grid of the plurality of spatial grids in the third space can be obtained in the same manner as the above examples, which will not be repeated here.
[0114] The occupancy state sample label can be a true occupancy state corresponding to each spatial grid of the plurality of spatial grids in the third space.
[0115] The manner of obtaining the true occupancy state corresponding to each spatial grid of the plurality of spatial grids in the third space can include: obtaining point cloud information in the third space; and obtaining the true occupancy state corresponding to each spatial grid of the plurality of spatial grids in the third space based on the point cloud information in the third space.
[0116] The obtaining of the real occupancy state corresponding to each of the plurality of spatial grids in the third space based on the point cloud information in the third space can include: mapping each point cloud feature in the point cloud information in the third space to a corresponding spatial grid in the third space to obtain one or more point cloud features corresponding to each of the plurality of spatial grids in the third space; and obtaining the real occupancy state corresponding to each of the plurality of spatial grids in the third space based on the one or more point cloud features corresponding to each of the plurality of spatial grids in the third space. Any one point cloud feature can be used to represent geometric features, color features, texture features, and the like, which are not listed here.
[0117] The mapping of each point cloud feature in the point cloud information in the third space to a corresponding spatial grid in the third space to obtain one or more point cloud features corresponding to each of the plurality of spatial grids in the third space can include: mapping each point cloud feature in the point cloud information in the third space to a corresponding spatial grid in the third space based on a correspondence between the coordinates of the point cloud information in the third space and the coordinates of the third space to obtain one or more point cloud features corresponding to each of the plurality of spatial grids in the third space. The coordinates of the point cloud information in the third space can be coordinates in a radar coordinate system, and the specific setting manner of the radar coordinate system is not limited in the embodiment. The determination manner of the coordinates of the third space is the same as the determination manner of the coordinates of the first space, which is not described here. The correspondence between the coordinates of the point cloud information in the third space and the coordinates of the third space can refer to the correspondence between each coordinate of the third space in the radar coordinate system and each coordinate of the third space in the spatial coordinate system, and the generation and specific content of the correspondence are not limited in the embodiment.
[0118] Taking the yth spatial grid (y is a positive integer) of the plurality of spatial grids in the third space as an example, the obtaining of the real occupancy state corresponding to each of the plurality of spatial grids in the third space based on the one or more point cloud features corresponding to each of the plurality of spatial grids in the third space can include: selecting a same and most numerous point cloud feature from the one or more point cloud features corresponding to the yth spatial grid in the third space as the real occupancy state corresponding to the yth spatial grid.
[0119] The real occupancy state can be referred to as a semantic label truth value, and the real occupancy state corresponding to any spatial grid can be represented in a one-hot coding (or one-bit effective coding) manner. For example, assuming that there are three optional occupancy states in total, and the real occupancy state corresponding to a spatial grid is the optional occupancy state 2, the one-hot coding of the real occupancy state corresponding to the spatial grid can be represented as {0, 1, 0}, where the first bit coding corresponds to the optional occupancy state 1, the second bit coding corresponds to the optional occupancy state 2, and the third bit coding corresponds to the optional occupancy state 3. If the second bit coding takes the value 1, it indicates that the real occupancy state corresponding to the spatial grid is the optional occupancy state 2.
[0120] In a case where j is equal to 1, the jth second conditional vector is obtained by adding a random noise coding vector to the sample label; in a case where j is greater than 1, the jth second conditional vector is obtained by adding a random noise vector to the (j-1)th second conditional vector.
[0121] In a case where the sample label is added with a random noise coding vector to obtain the jth second conditional vector, it can include: adding a random noise coding vector to the motion state sample label to obtain a jth motion state conditional vector; adding a random noise coding vector to the occupancy state sample label to obtain a jth occupancy state conditional vector; and taking the jth motion state conditional vector and the jth occupancy state conditional vector as the jth second conditional vector. The specific processing manner of adding a random noise coding vector to the motion state sample label to obtain the jth motion state conditional vector is similar to the foregoing embodiments, and will not be repeated.
[0122] In a case where the occupancy state sample label is added with a random noise coding vector to obtain the jth occupancy state conditional vector, it can include one of the following: adding a random noise coding vector to the one-hot coding corresponding to each spatial grid in the plurality of spatial grids in the third space of the occupancy state sample label to obtain the jth occupancy state conditional vector; and adding a random noise coding vector to the one-hot coding corresponding to part of the spatial grids in the plurality of spatial grids in the third space of the occupancy state sample label to obtain the jth occupancy state conditional vector. For example, assuming that there are three optional occupancy states in total, and the real occupancy state corresponding to a spatial grid is the optional occupancy state 2, the one-hot coding of the real occupancy state corresponding to the spatial grid can be represented as {0, 1, 0}, and after adding a random noise coding vector to the one-hot coding, the jth occupancy state conditional vector can be obtained as {-0.5, 0.6, 0.1}.
[0123] The jth second condition vector is obtained by adding a random noise vector to the (j-1)th second condition vector, including: adding a random noise encoding vector to the (j-1)th motion state condition vector to obtain the jth motion state condition vector; adding a random noise encoding vector to the (j-1)th occupancy state condition vector to obtain the jth occupancy state condition vector; and taking the jth motion state condition vector and the jth occupancy state condition vector as the jth second condition vector. The process of adding a random noise encoding vector to the (j-1)th motion state condition vector to obtain the jth motion state condition vector is similar to the foregoing example; the process of adding a random noise encoding vector to the (j-1)th occupancy state condition vector to obtain the jth occupancy state condition vector is also similar to the foregoing example, and will not be described in detail.
[0124] The jth loss is determined based on the jth motion state prediction information corresponding to the plurality of spatial grids in the third space and the sample label at the first time point, including: determining a jth first sub-loss based on the jth motion state prediction information corresponding to the plurality of spatial grids in the third space and the motion state sample label; determining a jth second sub-loss based on the jth occupancy state prediction information corresponding to the plurality of spatial grids in the third space and the occupancy state sample label; and determining the jth loss based on the jth first sub-loss and the jth second sub-loss. The first sub-loss and the second sub-loss can be cross-entropy losses.
[0125] For example, the jth loss is determined based on the jth first sub-loss and the jth second sub-loss, which can be: taking the sum of the jth first sub-loss and the jth second sub-loss as the jth loss. It should be pointed out that the jth loss can also be calculated in other ways, such as calculating the average value, the maximum value, etc. of the jth first sub-loss and the jth second sub-loss, which will not be listed here.
[0126] The preset model is trained based on the jth loss to obtain the jth trained preset model, which can include: adjusting the parameters of the preset model based on the jth loss; and obtaining the jth trained preset model when a convergence condition is met. The convergence condition can be at least one of: the jth loss is less than or equal to a threshold value; the number of cycles of training reaches a preset number, etc. The threshold value can be set according to actual conditions, which is not limited in the present application.
[0127] The preset model after the jth training is taken as the target model in a case that the jth training is the last training in the multiple trainings, including: judging whether the jth training is the last training in the multiple trainings, and taking the preset model after the jth training as the target model in a case that the jth training is the last training in the multiple trainings.
[0128] The jth training being the last training can be determined according to a preset training number. For example, in a case that the number of the jth training reaches the preset training number, it can be determined that the jth training is the last training, otherwise, it is determined that the jth training is not the last training. The preset training number can be configured according to actual conditions, for example, it can be 8, or 6, or larger or smaller, which is not limited in the embodiment.
[0129] In addition, it can also include: in a case that the jth training is not the last training in the multiple trainings, performing the j+1th training in the multiple trainings. The j+1th training in the multiple trainings is the same as the jth training in the multiple trainings, which is not repeated here.
[0130] It should be understood that the foregoing embodiment provides a plurality of trainings on the preset model based on the sample and the sample label to obtain the target model. The processing can be performed by an electronic device executing the spatial grid corresponding motion state prediction method provided in the present application. Alternatively, the plurality of trainings on the preset model based on the sample and the sample label to obtain the target model can also be performed by other devices or equipment. The way of training the target model by other devices or equipment is the same as the embodiment, which is not repeated. In the case of training the target model by other devices or equipment, the electronic device can receive and save the target model from the other devices or equipment in advance.
[0131] In this way, the preset model is trained multiple times, and the preset model after the jth training is taken as the target model in a case that the jth training meets the convergence condition and the jth training is the last training in the multiple trainings. In this way, the target model can predict the motion state information corresponding to multiple spatial grids in the space according to the image information, which can ensure the accuracy of the trained target model in predicting the motion state of the spatial grid. In addition, the second condition vector (i.e. the sample label) gradually increases the noise coding in different input times of the preset model, which can make the preset model learn the relationship between the sample label and the sample of different clarity, and then accurately predict the input information of different clarity during model prediction, thereby improving the accuracy of model prediction.
[0132] Figure 3 A schematic block diagram of a spatial grid-corresponding motion state prediction device according to an embodiment of this disclosure is shown. Figure 3 As shown, it includes:
[0133] The feature extraction module 301 is used to determine the feature vectors of multiple spatial grids in the first space based on multiple images of the first space acquired at the current time.
[0134] The prediction module 302 is used to perform multiple iterative processing on the feature vectors of multiple spatial grids in the first space and the motion state prediction information corresponding to multiple spatial grids in the second space to obtain the motion state prediction information corresponding to multiple spatial grids in the first space, wherein the motion state prediction information corresponding to multiple spatial grids in the second space is obtained at historical time.
[0135] The prediction module is configured to, in the i-th iteration of the multiple iterations, generate the information to be processed for the i-th iteration based on the feature vectors of multiple spatial grids in the first space and the motion state prediction information corresponding to multiple spatial grids in the second space, where i is a positive integer; input the information to be processed for the i-th iteration and the i-th first condition vector into the target model to obtain the i-th motion state prediction information corresponding to multiple spatial grids in the first space output by the target model; and, if the i-th iteration satisfies the constraint conditions, use the i-th motion state prediction information corresponding to multiple spatial grids in the first space as the motion state prediction information corresponding to multiple spatial grids in the first space.
[0136] Wherein, when i equals 1, the i-th first condition vector is a random noise vector; when i is greater than 1, the i-th first condition vector is the (i-1)-th motion state prediction information corresponding to multiple spatial grids in the first space.
[0137] like Figure 4 The device further includes:
[0138] The model training module 401 is used to train the preset model multiple times based on samples and sample labels to obtain the target model. The samples include feature vectors of multiple spatial grids in the third space at the first time and motion state information corresponding to multiple spatial grids in the fourth space at the second time. The first time is later than the second time.
[0139] The model training module is configured to, in the jth training of the multiple training, input a sample and a jth second conditional vector into a preset model to obtain jth motion state prediction information corresponding to a plurality of spatial grids in a third space at the first moment output by the preset model, where j is a positive integer; determine a jth loss based on the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first moment and the sample label; train the preset model based on the jth loss to obtain the preset model after the jth training; and in a case where the jth training is the last training of the multiple training, take the preset model after the jth training as the target model.
[0140] In a case where j is equal to 1, the jth second conditional vector is obtained by adding a random noise encoding vector to the sample label; and in a case where j is greater than 1, the jth second conditional vector is obtained by adding a random noise vector to a (j-1)th second conditional vector.
[0141] The specific functions and examples of the modules and sub-modules of the apparatuses in the embodiments of the present disclosure are described above in the corresponding steps of the method embodiments, and will not be described here again.
[0142] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0143] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0144] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0145] As Figure 5As shown, the electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded into a random access memory (RAM) 503 from a storage unit 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0146] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, and the like; an output unit 507, such as various types of displays, a speaker, and the like; a storage unit 508, such as a magnetic disk, an optical disk, and the like; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0147] The computing unit 501 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the above-described methods can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, at least one step of the above-described methods can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the above-described methods by any other appropriate means (e.g., by means of firmware).
[0148] According to another aspect of the present disclosure, there is provided an autonomous vehicle comprising the above-described electronic device.
[0149] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0150] Program code for carrying out methods of the present disclosure can be written in any combination of at least one programming language. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.
[0151] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include at least one line of electrical wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0152] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0153] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0154] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established by computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers combined with a blockchain.
[0155] It should be understood that various forms of flow shown above can be used, with steps reordered, added, or removed. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without limitation herein, so long as the desired results of the technology disclosed in the present disclosure are achieved.
[0156] The specific embodiments described above are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the principles of the present disclosure. Any further modifications, equivalents, and / or alternatives of the present disclosure as described within the principles of the present disclosure are intended to fall within the scope of the present disclosure.
Claims
1. A method for predicting motion state of spatial grids, comprising: determining feature vectors of a plurality of spatial grids in a first space based on a plurality of images of the first space collected at a current time; performing a plurality of iterations of processing the feature vectors of the plurality of spatial grids in the first space and motion state prediction information corresponding to a plurality of spatial grids in a second space to obtain the motion state prediction information corresponding to the plurality of spatial grids in the first space, wherein the motion state prediction information corresponding to the plurality of spatial grids in the second space is obtained at a historical time; and wherein the performing the plurality of iterations of processing the feature vectors of the plurality of spatial grids in the first space and the motion state prediction information corresponding to the plurality of spatial grids in the second space to obtain the motion state prediction information corresponding to the plurality of spatial grids in the first space comprises, in an i-th iteration of the plurality of iterations, generating to-be-processed information of the i-th iteration based on the feature vectors of the plurality of spatial grids in the first space and the motion state prediction information corresponding to the plurality of spatial grids in the second space, wherein i is a positive integer; inputting the to-be-processed information of the i-th iteration and an i-th first condition vector into a target model to obtain i-th motion state prediction information corresponding to the plurality of spatial grids in the first space output by the target model; and in a case where the i-th iteration satisfies a constraint condition, taking the i-th motion state prediction information corresponding to the plurality of spatial grids in the first space as the motion state prediction information corresponding to the plurality of spatial grids in the first space.
2. The method of claim 1, wherein in a case where i is equal to 1, the i-th first condition vector is a random noise vector; and in a case where i is greater than 1, the i-th first condition vector is (i-1)-th motion state prediction information corresponding to the plurality of spatial grids in the first space.
3. The method of claim 1 or 2, further comprising: training a preset model a plurality of times based on samples and sample labels to obtain the target model, wherein the samples comprise feature vectors of a plurality of spatial grids in a third space at a first time and motion state information corresponding to a plurality of spatial grids in a fourth space at a second time, and the first time is later than the second time. The training the preset model a plurality of times based on the samples and the sample labels to obtain the target model comprises: in a j-th training of the plurality of trainings, inputting a sample and a j-th second condition vector into the preset model to obtain j-th motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time output by the preset model, wherein j is a positive integer; determining a j-th loss based on the j-th motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time and the sample label; training the preset model based on the j-th loss to obtain the preset model after the j-th training; and wherein the j is greater than 1. 4. The method of claim 3, wherein, In a case that the jth training is the last training in the plurality of trainings, the preset model after the jth training is taken as the target model.
5. The method of claim 4, wherein, In a case that j is equal to 1, the jth second condition vector is obtained by adding a random noise coding vector to the sample label; In a case that j is greater than 1, the jth second condition vector is obtained by adding a random noise vector to a (j-1)th second condition vector.
6. An apparatus for predicting motion states of spatial grids, comprising: a feature extraction module configured to determine feature vectors of a plurality of spatial grids in a first space based on a plurality of images of the first space captured at a current time; a prediction module configured to perform a plurality of iterations on the feature vectors of the plurality of spatial grids in the first space and motion state prediction information corresponding to a plurality of spatial grids in a second space to obtain the motion state prediction information corresponding to the plurality of spatial grids in the first space, wherein the motion state prediction information corresponding to the plurality of spatial grids in the second space is obtained at a historical time; the prediction module is configured to, in an ith iteration of the plurality of iterations, generate to-be-processed information of the ith iteration based on the feature vectors of the plurality of spatial grids in the first space and the motion state prediction information corresponding to the plurality of spatial grids in the second space, wherein i is a positive integer; input the to-be-processed information of the ith iteration and an ith first condition vector into a target model to obtain ith motion state prediction information of the plurality of spatial grids in the first space output by the target model; and in a case that the ith iteration meets a constraint condition, take the ith motion state prediction information of the plurality of spatial grids in the first space as the motion state prediction information corresponding to the plurality of spatial grids in the first space.
7. The apparatus of claim 6, wherein, in a case that i is equal to 1, the ith first condition vector is a random noise vector; in a case that i is greater than 1, the ith first condition vector is (i-1)th motion state prediction information corresponding to the plurality of spatial grids in the first space.
8. The apparatus of claim 6 or 7, further comprising: a model training module configured to perform a plurality of trainings on a preset model based on a sample and a sample label to obtain the target model, wherein the sample comprises feature vectors of a plurality of spatial grids in a third space at a first time and motion state information corresponding to a plurality of spatial grids in a fourth space at a second time, and the first time is later than the second time.
9. The apparatus of claim 8, wherein, The model training module is configured to, in the jth training of the multiple training, input a sample and a jth second conditional vector into a preset model to obtain jth motion state prediction information corresponding to a plurality of spatial grids in a third space at a first time output by the preset model, where j is a positive integer; determine a jth loss based on the jth motion state prediction information corresponding to the plurality of spatial grids in the third space at the first time and a sample label; train the preset model based on the jth loss to obtain the preset model after the jth training; and in a case where the jth training is the last training of the multiple training, use the preset model after the jth training as the target model.
10. The apparatus of claim 9, wherein, in a case where j is equal to 1, the jth second conditional vector is obtained by adding a random noise encoding vector to the sample label; in a case where j is greater than 1, the jth second conditional vector is obtained by adding a random noise vector to a (j-1)th second conditional vector.
11. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.
12. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-5.
13. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-5.
14. An autonomous vehicle comprising the electronic device of claim 11.
Citation Information
Patent Citations
Vehicle track prediction method based on graph convolutional neural network in network connection environment
CN116933005A
Automatic driving model, automatic driving method and device based on state node prediction
CN118551806A