Scene information prediction method and device, equipment and storage medium
By constructing a scene prediction model that combines historical voxel representation and time control information, the problem of flexibility and high training cost of scenario information prediction methods in the prior art is solved, efficient and accurate prediction of future scenarios is achieved, and the security and adaptability of target objects in complex environments is improved.
Patent Information
- Application Number
- CN202510473177.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
The existing scenario information prediction methods are difficult to accurately model the impact of various environmental factors on motion trends, the prediction flexibility is insufficient, and the training cost is high, making it difficult to generalize to different environments.
By obtaining the historical time period image of the movable target object, constructing historical voxel representations, and combining time control information, using the scene prediction model to generate target voxel representations of the target time period, achieving efficient prediction of future scenarios.
It improves the flexibility and generalization ability of scenario information prediction, can more accurately perceive and understand the future environment, optimize independent decision-making, and improve the security and adaptability of target objects in complex dynamic scenarios.
Smart Images

Figure CN120388188A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to a method, apparatus, device, and storage medium for predicting scene information. Background Art
[0002] Trajectory prediction technology mainly involves reasonably inferring the motion trends of surrounding dynamic objects in a complex environment to assist in the decision-making of path planning for a movable target object. Trajectory prediction generally relies on key technologies such as environmental perception, historical motion data analysis, and future scene information prediction. Trajectory prediction methods may face challenges when dealing with complex scenarios. For example, it is difficult for scene information prediction methods to accurately model the influence of multiple environmental factors on motion trends. Therefore, how to improve the stability and generalization ability of future scene information prediction has become an important problem to be solved urgently. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for predicting scene information is provided. The method includes: obtaining a plurality of images of a scene where a movable target object is located in a historical time period; determining a historical voxel representation of the scene in the historical time period based on the plurality of images; and using a scene prediction model to generate a target voxel representation of the scene in a target time period based on the historical voxel representation and time control information indicating the target time period, where the target voxel representation is used to generate a trajectory of the target object in the target time period.
[0004] In a second aspect of the present disclosure, a device for predicting scene information is provided. The device includes: an obtaining module configured to obtain a plurality of images of a scene where a movable target object is located in a historical time period; a historical voxel representation determination module configured to determine a historical voxel representation of the scene in the historical time period based on the plurality of images; and a target voxel representation generation module configured to generate a target voxel representation of the scene in a target time period based on the historical voxel representation and time control information indicating the target time period, where the target voxel representation is used to generate a trajectory of the target object in the target time period.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory, where the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor. When the instructions are executed by the at least one processor, the device is caused to execute the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, there is provided a computer program product. The computer program product includes computer-executable instructions which, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this content part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0010] Figure 1 A schematic diagram showing an example environment in which the embodiments of the present disclosure can be implemented;
[0011] Figure 2 A flowchart showing an example process of scene information prediction according to some embodiments of the present disclosure;
[0012] Figure 3 A schematic diagram showing an example architecture for generating a target voxel representation according to some embodiments of the present disclosure;
[0013] Figure 4 A schematic diagram showing an example architecture for training a scene prediction model according to some embodiments of the present disclosure;
[0014] Figure 5 A schematic diagram showing an example of determining a corresponding depth estimation map according to some embodiments of the present disclosure;
[0015] Figure 6 A schematic structural block diagram showing a scene information prediction device according to some embodiments of the present disclosure;
[0016] Figure 7 A block diagram showing an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0018] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.
[0019] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this article, the "model" can also be referred to as a "machine learning model", a "machine learning network", a "neural network", or a "network", and these terms are used interchangeably herein.
[0020] In this article, unless explicitly stated, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.
[0021] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0022] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained through appropriate means according to the relevant laws and regulations.
[0023] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by it will require obtaining and using the user's personal information, so that the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solution of the present disclosure according to the prompt message.
[0024] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0025] It is understandable that the above-mentioned notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0026] Figure 1 FIG. shows a schematic diagram of an exemplary environment in which embodiments of the present disclosure can be implemented. Referring to Figure 1 , in the environment 100, one or more movable objects are deployed, for example, the first robot 110-1, the second robot 110-2, the third robot 110-3, and the like. The first robot 110-1, the second robot 110-2, and the third robot 110-3 can also be collectively or individually referred to as the robot 110. A control device can be deployed in each robot 110 to control the operation of the robot. For example, as Figure 1 shown, control devices 120-1, 120-2, and 120-3 are respectively deployed at the first robot 110-1, the second robot 110-2, and the third robot 110-3, which can also be collectively or individually referred to as the control device 120. In some embodiments, only one robot may be deployed in the environment 100, and only the movement path of this one robot needs to be planned. In some embodiments, multiple robots may be deployed in the environment 100. In this case, path planning can be performed for multiple robots in the environment 100 simultaneously, or path planning can be performed only for any one robot in the environment 100.
[0027] In the embodiments of the present disclosure, the robot 110 can be used for various suitable purposes. For example, it can be a transportation robot for delivering goods, and the control device 120 can be used to plan the movement path of the robot 110. The robot 110 can perform, for example, a package delivery service and a warehouse management service. As an example, the robot 110 can work in an indoor environment, such as a warehousing environment.
[0028] In Figure 1 the environment 100, the control device 120 can be any type of microcontroller, programmable logic controller, industrial control computer, single-board computer, field programmable gate array, digital signal processor, multi-core processor, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the control device 120 can control the robot 110 based on control instructions generated by itself.
[0029] In some embodiments, the operation of the robot 110 can be controlled via the server 130. The server 130 can receive data related to various tasks and remotely control the operation of the robot 110. The control device 120 collects data of the robot 110 and communicates with the server 130 via the network 140 to send the data to the server 130. The server 130 determines control instructions for the robot based on the acquired data and sends the control instructions to the control device 120. The control device 120 controls the robot 110 based on the received control instructions. The server 130 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, and the like.
[0030] In some embodiments, the robot 110 can be connected to and communicate with the server 130 via the network 140. The robot 110 can send data related to various tasks executed by the robot 110 to the server 130 in real time.
[0031] It should be understood that the structures and functions of the various elements in the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure. In addition, although a robot is shown in Figure 1 , embodiments of the present disclosure can also be applied to other types of movable objects.
[0032] As briefly mentioned above, with the development of intelligent perception, intelligent control, and machine learning technologies, trajectory prediction has become a key task for movable objects to achieve efficient decision-making and safe planning in complex dynamic environments. For example, in the fields of autonomous driving, robot perception, etc., scene information prediction, as an important link between environmental perception and path decision-making, directly affects the safety and reliability of trajectory prediction. The core task of scene information prediction is to predict the information of the environment where the movable object will be located in the future time period based on historical information, so that the autonomous driving system or robot can plan a reasonable path in advance. Furthermore, potential collisions can be avoided and driving efficiency can be optimized. However, traditional scene information prediction methods still face many challenges. Several typical challenges are described below.
[0033] A typical challenge involves the flexibility of prediction. For example, the future time period predicted by traditional scene information prediction methods is uncontrollable. This limitation may lead to inaccurate decisions during application processes such as autonomous driving and intelligent path planning. For example, in autonomous driving, it may be necessary to control the evolution of scene information at different time points to make more accurate path adjustments. The predictions of traditional scene information prediction methods are difficult to meet this requirement.
[0034] Another typical challenge involves the issue of training costs. Specifically, traditionally, most of the scenario information prediction methods need to obtain the scenario information in the future time period for supervised learning and require a large amount of three-dimensional (3D) semantic annotations. However, the cost of manually annotating 3D data is extremely high. In addition, the distribution of 3D data varies greatly in different environments (such as urban roads, highways, closed campuses, warehouses), resulting in difficulty in generalizing the training data. This further increases the difficulty of obtaining training data.
[0035] In view of this, according to an embodiment of the present disclosure, an improved solution for scenario information prediction is proposed to at least partially solve one or more of the above problems. In this solution, first, multiple images of the scenario where the movable target object is located in the historical time period are obtained. Further, based on the multiple images, the historical voxel representation of the scenario in the historical time period is determined. Finally, a scenario prediction model can be used to generate the target voxel representation of the scenario in the target time period based on the historical voxel representation and the time control information indicating the target time period. The target voxel representation can be used to generate the trajectory of the target object in the target time period.
[0036] In an embodiment of the present disclosure, by obtaining multiple frames of images in the historical time period, constructing the historical voxel representation, and combining the time control mechanism to predict the voxel representation in the future target time period. In this way, an efficient prediction of future scenario information is achieved. In this way, based on the historical scenario information, the voxel representation of the environment at any future time point can be flexibly predicted, providing reliable input data for trajectory prediction and motion path planning. In this way, the future environment can be more accurately perceived and understood, autonomous decision-making can be optimized, and the safety and adaptability of the target object in a complex dynamic scenario can be improved.
[0037] Some example embodiments of the present disclosure will be described in detail below with reference to the examples of the drawings.
[0038] Figure 2 The flowchart of an example process 200 for scenario information prediction according to some embodiments of the present disclosure is shown. The process 200 will be mainly described below with reference to Figure 1 the example environment, but it should be understood that this is only exemplary and is not intended to impose any limitations. The embodiments of the present disclosure can also be applied to other types of environments. In addition, the process 200 will be described with respect to the server 130, but this is only exemplary. The process 200 can be executed by the remote server of the target object (for example, the server 130), or can be executed by the control device of the target object itself (for example, the control device 120), or can be completed by the cooperation of the remote server and the control device.
[0039] At block 210, the server 130 obtains multiple images of the scene where the movable target object is located in a historical time period. In an embodiment of the present disclosure, the target object may include various suitable types of movable objects, such as, but not limited to, an autonomous driving system (such as an autonomous vehicle), a drone, a robot, etc. Such objects can autonomously move towards a set target in the environment where they are located. For example, the target object may be one or more of the first robot 110-1, the second robot 110-2, and the third robot 110-3 described above (collectively or individually referred to as the robot 110). Further, for safe and efficient movement, it is necessary to sense and predict the environment where the target object is located. Based on the results of the sensing and prediction, the next action of the target object can be determined (such as determining the traveling direction, planning a reasonable path, etc.).
[0040] In some embodiments, the scene where the target object is located may refer to the surrounding environment that the target object has experienced within a certain time range, which may include static objects, dynamic objects, etc. Static objects may include, for example, roads, obstacles, buildings, shelves, etc., which are stationary within a relatively long time period. Dynamic objects may include, for example, people, animals, other movable objects, etc., which may move in real time.
[0041] During the movement of the target object, information about the environment around it in a certain historical time period can be collected by an image acquisition device (such as a camera, a vision sensor), etc., for further constructing a three-dimensional understanding or predicting the future scene. Such an image acquisition device may be set in the surrounding environment or may be set on the target object. For example, the server 130 may obtain T frame images of one or more perspectives collected by a camera on the target object within a historical time period. These images record the environmental changes experienced by the target object and provide basic data for subsequent future scene prediction.
[0042] Continuing to refer to Figure 2 , at block 220, the server 130 determines a historical voxel representation of the scene in the historical time period based on the multiple images. Exemplarily, after receiving the multiple images, the server 130 may convert the multiple images into a three-dimensional voxel representation (3D voxel representation) through a three-dimensional encoder (3D encoder). The three-dimensional encoder can be used to extract the spatial structure and geometric information contained in the images and model the scene in the three-dimensional voxel space. For example, the voxel representation may be a regular three-dimensional grid structure, and each cell of which can be used to represent whether the corresponding position in space is occupied, the density value of the voxel, a semantic label, or other attributes.
[0043] In some embodiments, the historical voxel representation may include a set of voxel representations at multiple historical moments. The set includes a number of voxel representations corresponding to multiple frames of images or multiple historical moments. A certain voxel representation in the historical voxel representation may correspond to the three-dimensional spatial features formed after the image collected at a certain moment is processed by the encoder.
[0044] As an example, the historical time period may represent a continuous time interval, such as a number of seconds in the past. During this time period, the server 130 may continuously collect image frames at a fixed frequency (such as several frames per second), and each frame of image corresponds to a specific moment within this time period. The multiple moments included in the historical time period can be understood as the set of the acquisition time points of these image frames. For example, within the past 1 second, there is one frame of image corresponding to every 0.2 seconds, for a total of 5 frames. Each frame represents a specific past moment (e.g., t0 = 0s, t1 = 0.2s, t2 = 0.4s,...). Further, the server 130 may construct corresponding voxel representations based on the image frames corresponding to each moment, that is, each frame of image generates a three-dimensional scene expression at the corresponding moment after three-dimensional modeling. The voxel representation of each frame reflects the three-dimensional spatial environment in which the target object is located at the corresponding moment. Therefore, the multi-frame voxel representations can be used to describe the continuous three-dimensional modeling results of the scene where the target object is located within the historical time period.
[0045] Alternatively or additionally, the multiple images within the historical time period are not limited to an image sequence continuously acquired at a fixed time interval. In other words, the multiple images may include images acquired in a non-equidistant manner, such as images acquired at critical moments manually set according to application requirements, or images acquired triggered by detected events. In this case, the server 130 can also construct corresponding multi-frame voxel representations based on these non-continuous or non-equidistant images, so as to support subsequent scene information prediction and trajectory planning.
[0046] Return reference Figure 2 , at block 230, the server 130 uses the scene prediction model to generate a target voxel representation of the scene in the target time period based on the historical voxel representation and the time control information indicating the target time period. The target voxel representation is used to generate the trajectory of the target object in the target time period.
[0047] Traditionally, the scene information prediction method lacks the ability to perform directional prediction for any target time point, thus limiting the adaptability to the changing timing requirements in complex dynamic environments. To improve the prediction flexibility and accuracy, time control information can be introduced when performing scene prediction, and this information can be used to indicate the target time period or time point for model prediction.
[0048] In some embodiments, the time control information may include forms such as absolute time representing the target time, relative timestamps, time codes, etc. After receiving the time control information, the scene prediction model may use it together with the historical voxel representation as input, and generate a target voxel representation corresponding to the target time period through a time perception mechanism. With the time control information, the scene prediction model can achieve modeling of the scene state at any target time point, thereby significantly improving the temporal resolution and flexibility of the prediction. In this way, more refined and dynamic trajectory planning and decision control can be supported.
[0049] Figure 3 FIG. shows a schematic diagram of an exemplary architecture 300 for generating a target voxel representation according to some embodiments of the present disclosure. In this example, the scene prediction model may include a feature extraction network 311, a feed-forward neural network 312, a control module 313, a fusion module 314, a convolutional layer 315, etc. However, this is only exemplary, and the scene prediction model may include other model layers. As Figure 3 shown, the server 130 may use the feature extraction network 311 included in the scene prediction model to perform feature extraction on the historical voxel representation 301 to generate corresponding scene features 302-1, 302-2, and 302-3 at multiple scales of the scene, which are also collectively or individually referred to as scene features 302. For example, the server 130 may input the voxel representations of T frames in the historical time period into the feature extraction network 311 (such as a residual network) for feature extraction. The size of the historical voxel representation is T×C×X×Y×Z, where C is the number of channels, and X, Y, and Z respectively correspond to the length, width, and depth of the three-dimensional voxel data. For example, the feature extraction network 311 may perform convolution on the historical voxel representation data in the three-dimensional space to extract deep semantic features in the space and obtain scene features 302 at multiple scales of the scene. The multi-scale scene features 302 may correspond to the structural information at different spatial scales in the voxel representation. For example, local details are extracted at lower levels, and the overall structure is extracted at higher levels.
[0050] Furthermore, to achieve accurate prediction of the scene state within a future target time period, time control information 303 can be introduced as a prediction condition. The feed-forward neural network 312 included in the scene prediction model can perform feature extraction on the time control information 303 to obtain time control features at multiple scales that match the above-mentioned multi-scale scene features 302. The time control information 303 can include the time step (timestep) representing the target prediction time period. For example, predicting N frames means predicting the scenes of the subsequent N time steps. Alternatively or additionally, the specific time control information can be determined based on user input. The introduction of this time control feature enables the model to not only rely on the historical spatial scene during prediction but also clarify that the target is to predict at "which second in the future", avoiding the uncertainty in the time dimension when simply predicting scene information based on the historical voxel representation.
[0051] As Figure 3 shown, the control module 313 in the scene prediction model can further determine the target voxel representation 305 based on the corresponding scene features 302 at multiple scales and the corresponding time control features at multiple scales.
[0052] As an example, the multi-scale scene features 302 and the time control features can be input into the control module 313. In the control module 313, regulatory parameters are applied to the scene features 302 through operations such as layer normalization, scaling, and offset. Further, for each scale among the multiple scales, the control module 313 applies the attention mechanism to the scene features at that scale by taking the time control features at that scale as a condition. Subsequently, operations such as layer normalization and scaling and offset based on the conditional regulatory parameters are performed to obtain the regulated scene features 304-1, 304-2, and 304-3 at that scale, which are also collectively or individually referred to as the regulated scene features 304.
[0053] In some embodiments, the server 130 can determine the target voxel representation 305 by fusing the regulated scene features 304 determined for each scale respectively. Exemplarily, for each scale among the multiple scales, the control module 313 can take the time control features at that scale as a regulatory condition and perform regulatory processing on the scene features at that scale based on the attention mechanism, thereby generating the regulated multi-scale scene features 304. By fusing the regulated scene features at each scale (for example, performing operations such as splicing, weighting, or convolutional summation through the fusion module 314 and the convolutional layer 315), the target voxel data with a size of N×C×X×Y×Z can be finally obtained, where N is the number of frames of the target voxel representation indicated by the time control information 303, C is the number of channels, and X, Y, and Z respectively correspond to the length, width, and depth of the three-dimensional voxel data.
[0054] It should be noted that the number of scales of the above scene feature 302 and the specific composition of each module and network structure for operations such as feature extraction, attention regulation, and feature fusion included in the scene prediction model are only examples. Embodiments of the present disclosure can be flexibly adjusted according to specific application requirements, data scale, model accuracy requirements, and other factors. For example, the scene prediction model can include more or fewer feature extraction modules 311, feed-forward neural networks 312, fusion modules 314, convolutional layers 315, etc., or be replaced with other modules or network structures with equivalent functions, so as to optimize aspects such as scene prediction accuracy, computational efficiency, and model capacity.
[0055] In some embodiments, the foregoing process of generating the target voxel representation based on the historical voxel representation and the time control information is based on the architecture of the scene prediction model. This process can be executed in the inference phase or in the training phase of the scene prediction model. Some embodiments are described below for the training of the scene prediction model.
[0056] In some embodiments, first, the server 130 can generate the target voxel representation of the target time period based on the input historical voxel representation and the corresponding time control information. The server 130 can determine the depth estimation information for the target time period based on the target voxel representation, and the depth estimation information represents the estimated depth of the elements in the scene during the target time period. This is a process of mapping the three-dimensional voxel representation to a depth map in a two-dimensional perspective, representing the estimated distance from the observation point to the object surface when passing through the scene in each pixel direction.
[0057] Figure 4 A schematic diagram of an example architecture 400 for training a scene prediction model according to some embodiments of the present disclosure is shown. As Figure 4 shown, the server 130 can generate the target voxel representation 402 of the target time period based on the input historical voxel representation 401 and the corresponding time control information, using the scene prediction model 411 to be trained. The target voxel representation 402 includes the corresponding voxel representations at multiple moments within the target time period.
[0058] In some examples, each moment can be mapped to a rendering frame. Each frame can also correspond to a moment. In other words, if the target time period is divided into N discrete moments (such as t1, t2,..., t N )), then the server 130 can generate the corresponding voxel representations at these N moments respectively as its three-dimensional semantic basis. Therefore, each frame of the target voxel representation can correspond to the visualization or rendering performance of each moment of the target time period, and each moment is the positioning of the frame on the time axis.
[0059] Further, the depth information determination module 412 may determine depth estimation information 403 for a target time period based on the target voxel representation 402. The depth estimation information 403 may include a plurality of estimated depth maps respectively corresponding to the plurality of moments. Each estimated depth map represents the spatial depth estimation value of each pixel point in a frame of image corresponding to a moment. For each moment among the plurality of moments, the depth information determination module 412 may convert the voxel representation corresponding to this moment into a signed distance field representation (Signed Distance Field, SDF) corresponding to this moment. Based on the signed distance field representation corresponding to this moment, an estimated depth map corresponding to this moment is determined.
[0060] For each moment within the target time period, Figure 5 FIG. 500 is a schematic diagram showing an example of determining a corresponding estimated depth map according to some embodiments of the present disclosure. As Figure 5 shown, a feed-forward neural network 511 (such as a multi-layer perceptron) may convert the target voxel representation 402 into a corresponding signed distance field representation 501. The signed distance field can model the distance of any point in three-dimensional space relative to the scene surface, where a positive value indicates that the point is outside the scene surface, a negative value indicates that the point is inside, and a zero value indicates that the point is on the surface. Then, the sampling module 512 may sample the three-dimensional points 502 in the scene. For example, starting from the camera center o, a ray is emitted in the direction r of each pixel in the image, and D points are uniformly sampled on each ray, which can be represented by the following formula:
[0061] {P i = o + t i r i | i = 1,... D, t i <t t+1} (1)
[0062] where t i is the depth value of the i-th sampling point on the ray, and P i is the corresponding 3D spatial position.
[0063] Subsequently, the server 130 may use the sampling function 513 to find each sampled three-dimensional point 502, that is, P i corresponding SDF value 503, that is, S i . Further, the server 130 may further use the volume rendering module 514 to penetrate the SDF field along the pixel direction from the simulated viewing perspective, calculate the intersection position with the surface, and thus obtain the estimated depth map 504 at this moment. For example, the server 130 processes these SDF values to estimate the opacity αj of each sampling point, that is, the possibility that this point is the object surface, which can be represented by the following formula:
[0064]
[0065] Among them, σ(s) is an activation function (such as the sigmoid function), which is used to smoothly map the SDF value to the interval (0, 1) to express the "surface possibility" of a certain point in space, and can be expressed by the following formula:
[0066]
[0067] Among them, β is a learnable parameter used to control the sharpness of the surface edge. In order to integrate the contributions of all sampling points into a final depth estimate di, a volumetric rendering method is adopted, and it is carried out through the following weighted summation formula:
[0068]
[0069] Among them, the weight wj represents the "surface contribution degree" of each sampling point, which is determined by the transparency of this point and can be expressed by the following formula:
[0070]
[0071] Among them, Tj is the cumulative transparency of all points before the current point and can be expressed by the following formula:
[0072]
[0073] This ensures that if a surface has been encountered at the previous sampling points, the weights of the subsequent points will be suppressed. After the rays of all pixels complete the above process, the depth values di at each pixel position can be aggregated into a complete estimated depth map 504, and this depth map can be expressed by the following formula:
[0074]
[0075] Among them, D render is the estimated depth map 504 at this moment, and R represents the two-dimensional array corresponding to this depth map, and its dimensions are the height (H) and width (W) of the image.
[0076] Continue to refer to Figure 4, the server 130 can obtain the depth tag information 407 for the target time period. The depth tag information 407 represents the true depth or pseudo-true depth of the elements in the scene during the target time period, and thus can be used as tags to train the scene prediction model 411. The depth tag information 407 includes a plurality of true depth maps corresponding to multiple moments within the target time period. The server 130 first obtains a plurality of images 404 of the scene at multiple moments within the target time period and radar data 406 collected at multiple moments in the scene.
[0077] For each of the multiple moments, the depth estimation model 414 can determine a relative depth map 405 corresponding to that moment. The relative depth map 405 indicates the depth relationship of each pixel point in the image at that moment relative to a reference point. Further, the calibration module 415 can calibrate the relative depth 405 corresponding to that moment based on the radar data 406 at that moment to obtain the true depth map at that moment.
[0078] As an example, the server 130 first collects a plurality of image data 404 of the scene within the target time period (e.g., from an RGB camera) and radar data 406 at multiple moments (e.g., lidar point cloud). Then, for each of the multiple moments, the server 130 can generate a relative depth map 405 based on the image data at that moment through a pre-trained depth estimation model 414 such as a Monocular Depth Estimation (MDE). The relative depth map 405 represents the relative depth relationship of each pixel in the image relative to a certain reference pixel or region, and it still lacks a true scale value. Therefore, the server 130 can perform scale calibration and geometric alignment on the above relative depth map 405 based on the radar data 406 collected at that moment. For example, the server 130 can use methods such as Random Sample Consensus (RANSAC) for calibration to make the estimated relative depth in the image have a unified scale, thereby determining the pseudo-true depth that can be used for supervision corresponding to the image at that moment.
[0079] In some embodiments, the service 130 can update the scene prediction model based on the depth estimation information and the depth tag information. For example, the service 130 can construct a loss function 413 based on the depth estimation information 403 and the depth tag information 407. Based on the loss function 413, the scene prediction model 411 is updated. Calculate a supervision signal (such as a depth loss function), and accordingly update the parameters of the scene prediction model to optimize the model's ability to model the spatial structure in the future time period. Through this training mechanism, the scene prediction model 411 can, in the absence of explicit three-dimensional annotation data, rely on two-dimensional depth supervision to complete the learning of the modeling ability of the target voxel representation with a specified time series, thereby improving the prediction accuracy of future scene changes.
[0080] As an example, the loss function 413 can be expressed by the following formula:
[0081]
[0082] where N is the total number of frames represented by the target voxel. represents the estimated depth map corresponding to the predicted voxel representation of the i-th frame. D i represents the ground truth depth map corresponding to the i-th frame image. The loss function 413 can be defined as the average depth error of all frames. The server 130 can perform backpropagation based on the loss function 413 to update and optimize the parameters in the scene prediction model 411. Through the above method, a method for effectively training a scene prediction model without relying on real 3D data annotation is realized. The core lies in introducing depth information as an intermediate supervision signal, bypassing direct 3D annotation, and indirectly learning the 3D structure through supervision in the image space. In addition, this training mechanism can make full use of low-cost image and radar data, greatly improving the convenience and scalability of training data construction.
[0083] In short, although the network model in the present disclosure still aims to generate voxel representations of several future frames during the inference stage, its training process does not use traditional 3D annotation ground truth data as supervision. Through the above-mentioned deep supervision learning method, the target voxel representation is mapped to the depth map, and then the ground truth depth map is used as the supervision signal to realize an end-to-end training process starting from image / radar data. This mechanism significantly reduces the training threshold while ensuring good prediction performance.
[0084] Through the description of various embodiments of the present disclosure, it can be more clearly understood that in the embodiments of the present disclosure, on the one hand, a time control mechanism is introduced, and the time control information is input into the model control module as conditional information. In this way, the scene prediction model can flexibly predict the target voxel representation corresponding to any time point, thus significantly improving the flexibility and generalization ability of temporal modeling. On the other hand, a mechanism based on deep supervision learning is adopted, and the pseudo-ground truth depth obtained by calibrating the image depth estimation result and radar data is used as the supervision signal. In this way, high-quality target voxel representation modeling and model training can be realized without expensive 3D annotation data, significantly reducing data dependence and training costs.
[0085] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 6FIG. 600 is a schematic structural block diagram of a scene information prediction device according to some embodiments of the present disclosure. The device 600 may be implemented as or included in the server 130. Each module / component in the device 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0086] As shown in Figure 6 , the device 600 includes an acquisition module 610, a historical voxel representation determination module 620, and a target voxel representation generation module 630. In some embodiments, the acquisition module 610 is configured to acquire a plurality of images of a scene where a movable target object is located during a historical time period; the historical voxel representation determination module 620 is configured to determine a historical voxel representation of the scene during the historical time period based on the plurality of images; and the target voxel representation generation module 630 is configured to generate a target voxel representation of the scene during a target time period based on the historical voxel representation and time control information indicating the target time period, where the target voxel representation is used to generate a trajectory of the target object during the target time period.
[0087] In some embodiments, the target voxel representation generation module 630 is further configured to perform feature extraction on the historical voxel representation by using a feature extraction network included in the scene prediction model to generate corresponding scene features of the scene at multiple scales; perform feature extraction on the time control information by using a feedforward neural network included in the scene prediction model to obtain corresponding time control features at multiple scales; and determine the target voxel representation based on the corresponding scene features at multiple scales and the corresponding time control features at multiple scales.
[0088] In some embodiments, the target voxel representation generation module 630 is further configured to, for each scale in the multiple scales, apply an attention mechanism to the scene features at that scale by using the time control features at that scale as a condition to obtain the regulated scene features at that scale; and determine the target voxel representation by fusing the regulated scene features determined for the multiple scales respectively.
[0089] In some embodiments, the device 600 further includes a training module, which is configured to determine depth estimation information for the target time period based on the target voxel representation, where the depth estimation information represents the estimated depth of elements in the scene during the target time period; acquire depth label information for the target time period, where the depth label information represents the true depth of elements in the scene during the target time period; and update the scene prediction model based on the depth estimation information and the depth label information.
[0090] In some embodiments, where the target voxel representation includes corresponding voxel representations for multiple moments within a target time period, the depth estimation information includes multiple estimated depth maps respectively corresponding to the multiple moments, and the multiple estimated depth maps are determined by: for each of the multiple moments, converting the voxel representation corresponding to that moment into a signed distance field representation corresponding to that moment; and based on the signed distance field representation corresponding to that moment, determining the estimated depth map corresponding to that moment.
[0091] In some embodiments, the depth label information includes multiple ground truth depth maps respectively corresponding to multiple moments within a target time period, and a given ground truth depth map among the multiple ground truth depth maps is determined by: acquiring multiple images of a scene at multiple moments within the target time period and radar data acquired at multiple moments in the scene; for each of the multiple moments, using a depth estimation model to determine a relative depth map corresponding to that moment, the relative depth map indicating the depth relationship of each pixel point in the image at that moment relative to a reference point; and using the radar data at that moment to calibrate the relative depth map corresponding to that moment to obtain the ground truth depth map at that moment.
[0092] In some embodiments, the training module is further configured to construct a loss function based on the difference between the depth estimation information and the depth label information; and update the scene prediction model based on the loss function.
[0093] In some embodiments, the time control information is determined based on user input.
[0094] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. The electronic device 700 may be used, for example, to implement the control device 120 as shown in Figure 1 It should be understood that Figure 7 The electronic device 700 shown is merely exemplary and should not constitute any limitation to the functions and scope of the embodiments described herein.
[0095] Referring to Figure 7 , the electronic device 700 is in the form of a general-purpose electronic device. The components of the electronic device 700 may include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 700.
[0096] The electronic device 700 generally includes multiple computer storage media. Such media can be any available media accessible to the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be removable or non-removable media and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be capable of storing information and / or data and can be accessed within the electronic device 700.
[0097] The electronic device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 7 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules that are configured to execute the various methods or actions of the various embodiments of the present disclosure.
[0098] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented in a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the electronic device 700 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0099] The input device 760 can be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 760 can be one or more output devices, such as a display, speaker, printer, etc. The electronic device 700 can also communicate with one or more external devices (not shown) as needed via the communication unit 740, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 700, or communicate with any device that enables the electronic device 700 to communicate with one or more other electronic devices (such as a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0100] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0101] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0102] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause a computer, a programmable data processing device, and / or other devices to operate in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0103] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operation steps are executed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0105] The implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The determination of the terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art in the field to understand the various implementation manners disclosed herein.
Claims
1. A method for predicting scene information, comprising: Obtaining a plurality of images of a scene where a movable target object is located in a historical time period; Determining a historical voxel representation of the scene in the historical time period based on the plurality of images; And Using a scene prediction model, based on the historical voxel representation and time control information indicating a target time period, generating a target voxel representation of the scene in the target time period, where the target voxel representation is used to generate a trajectory of the target object in the target time period.
2. The method according to claim 1, wherein generating the target voxel representation in the target time period includes: Performing feature extraction on the historical voxel representation by using a feature extraction network included in the scene prediction model to generate corresponding scene features of the scene at multiple scales; Performing feature extraction on the time control information by using a feedforward neural network included in the scene prediction model to obtain corresponding time control features at the multiple scales; And Based on the corresponding scene features at the multiple scales and the corresponding time control features at the multiple scales, determining the target voxel representation.
3. The method according to claim 2, wherein determining the target voxel representation includes: For each of the multiple scales, by taking the time control feature at this scale as a condition, applying an attention mechanism to the scene feature at this scale to obtain a regulated scene feature at this scale; And By fusing the regulated scene features respectively determined for the multiple scales, determining the target voxel representation.
4. The method according to claim 1, wherein the method is executed during the training of the scene prediction model, and the method further includes: Based on the target voxel representation, determining depth estimation information for the target time period, where the depth estimation information represents the estimated depth of elements in the scene within the target time period; Obtaining depth label information for the target time period, where the depth label information represents the true depth of the elements in the scene within the target time period; And Based on the depth estimation information and the depth label information, updating the scene prediction model.
5. The method according to claim 4, wherein the target voxel representation includes corresponding voxel representations of multiple moments within the target time period, the depth estimation information includes multiple estimated depth maps respectively corresponding to the multiple moments, and the multiple estimated depth maps are determined by the following method: For each of the multiple moments, Converting the voxel representation corresponding to this moment into a signed distance field representation corresponding to this moment; and Based on the signed distance field representation corresponding to this moment, determining an estimated depth map corresponding to this moment.
6. The method according to claim 4, wherein the depth label information includes multiple true depth maps respectively corresponding to multiple moments within the target time period, and a given true depth map among the multiple true depth maps is determined by the following method: Obtain multiple images of the scene at multiple moments within the target time period and radar data collected at the multiple moments in the scene; For each of the multiple moments, Use a depth estimation model to determine a relative depth map corresponding to that moment, where the relative depth map indicates the depth relationship of each pixel point in the image at that moment relative to a reference point; and Use the radar data at that moment to calibrate the relative depth map corresponding to that moment to obtain a ground truth depth map at that moment.
7. The method according to claim 4, wherein updating the scene prediction model based on the depth estimation information and the depth label information includes: Construct a loss function based on the difference between the depth estimation information and the depth label information; And Update the scene prediction model based on the loss function.
8. The method according to claim 1, wherein the time control information is determined based on user input.
9. A scene information prediction device, comprising: An acquisition module configured to acquire multiple images of a scene where a movable target object is located in a historical time period; A historical voxel representation determination module configured to determine a historical voxel representation of the scene in the historical time period based on the multiple images; And A target voxel representation generation module configured to generate a target voxel representation of the scene in the target time period based on the historical voxel representation and time control information indicating the target time period, where the target voxel representation is used to generate a trajectory of the target object in the target time period.
10. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 8.
12. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions when executed by a processor implement the method according to any one of claims 1 to 8.