Scene modeling method and device, vehicle, equipment, medium and chip

By extracting scene features from multiple perspectives and predicting future scene features, the problem of low efficiency and insufficient accuracy in scene modeling is solved, achieving efficient and accurate multi-dimensional scene modeling, and improving the adaptation of downstream tasks and the accuracy of vehicle testing.

CN121861202APending Publication Date: 2026-04-14XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAOMI EV TECH CO LTD
Filing Date
2025-12-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing technologies, scene modeling algorithms are inefficient at processing large amounts of modeling data, resulting in poor modeling efficiency and insufficient modeling accuracy and completeness.

Method used

By extracting scene features from multiple perspectives based on the current time step and updating historical scene features, the scene features of future time steps are predicted. Multidimensional scene modeling is then performed by combining the features of the first and second target scenes. Techniques such as feature refinement, fusion, and coordinate system transformation are employed to reduce the computational load of the algorithm and improve the modeling accuracy and completeness.

Benefits of technology

It improved the efficiency and accuracy of scene modeling, optimized the modeling method, and enhanced the adaptability to downstream tasks and the accuracy of vehicle performance testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121861202A_ABST
    Figure CN121861202A_ABST
Patent Text Reader

Abstract

The invention provides a scene modeling method and device, a vehicle, equipment, a medium and a chip, and the method comprises the steps: obtaining a multi-view scene feature of an image corresponding to a current time step based on a to-be-processed image collected at the current time step; the historical scene features are updated based on the multi-view scene features to obtain first target scene features corresponding to a first time step sequence, second target scene features corresponding to a second time step sequence are predicted and determined, and the second time step sequence is a future time step sequence of the first time step sequence; and based on the first target scene feature and the second target scene feature, modeling the target scene to obtain a target multi-dimensional scene of the target scene. According to the invention, the operand of a multi-dimensional modeling algorithm is reduced, the modeling efficiency of the target scene is improved, and the accuracy of displaying the change condition of the target scene by the target multi-dimensional scene modeling is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to artificial intelligence technology applied to the vehicle field, and particularly to a scene modeling method, apparatus, vehicle, equipment, medium, and chip. Background Technology

[0002] With the development of technology, more and more tasks need to rely on modeling of real-world scenarios. Among the related technologies, large amounts of collected modeling data can be processed to construct corresponding 3D models. However, the modeling algorithms are complex and the modeling efficiency is poor. Summary of the Invention

[0003] This disclosure aims to at least partially address one of the technical problems in the related art.

[0004] Therefore, the first aspect of this disclosure proposes a scene modeling method.

[0005] The second aspect of this disclosure proposes a scene modeling apparatus.

[0006] The third aspect of this disclosure proposes a vehicle.

[0007] The fourth aspect of this disclosure provides for an electronic device.

[0008] The fifth aspect of this disclosure provides for a computer-readable storage medium.

[0009] The sixth aspect of this disclosure proposes a chip.

[0010] The first aspect of this disclosure proposes a scene modeling method, comprising: obtaining multi-view scene features of the image corresponding to the current time step based on the image acquired at the current time step to be processed; updating historical scene features based on the multi-view scene features to obtain a first target scene feature corresponding to a first time step sequence, wherein the historical scene features are obtained based on images acquired at processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene; predicting and determining a second target scene feature corresponding to a second time step sequence based on the first target scene feature, wherein the second time step sequence is a future time step sequence of the first time step sequence; and modeling the target scene based on the first target scene feature and the second target scene feature to obtain a target multi-dimensional scene of the target scene.

[0011] The scene modeling method proposed in this disclosure, compared with related technologies that require data processing of large amounts of modeling data to achieve scene modeling, extracts modeling features from multi-view images at each time step in the first time step sequence one by one, reducing the computational load of multi-dimensional modeling algorithms and improving the modeling efficiency of target scenes. Based on the first target scene features, it predicts the second target scene features corresponding to the second time step sequence, improving the prediction accuracy of the second target scene features. Based on the first and second target scene features, it performs multi-dimensional scene modeling, improving the completeness and accuracy of the target multi-dimensional scene modeling, improving the accuracy of the target multi-dimensional scene modeling in displaying changes in the target scene, optimizing the modeling method and the modeling effect of target multi-dimensional scene modeling. In the scenario of executing downstream tasks based on target multi-dimensional scene modeling, it improves the adaptability of the task decision of the downstream task to the target scene. In the scenario of vehicle performance testing based on target multi-dimensional scene modeling, it improves the testing accuracy and precision of vehicle performance testing, and optimizes the task execution quality of the downstream tasks corresponding to target multi-dimensional scene modeling.

[0012] The scene modeling method proposed in the first aspect of this disclosure also includes the following technical features: According to one embodiment of this disclosure, the acquired image is a multi-view image. The step of obtaining the multi-view scene features of the image corresponding to the current time step based on the image acquired at the current time step includes: extracting single-view image features corresponding to each single-view image from the multi-view acquired image at the current time step; determining newly added single-view features in each single-view image based on the historical scene features and the single-view image features; and determining the multi-view scene features corresponding to the current time step based on the newly added single-view features.

[0013] According to one embodiment of this disclosure, determining the newly added single-view features in each single-view image based on the historical scene features and the single-view image features includes: acquiring the feature acquisition pose of unit scene features in the historical scene features, wherein the feature acquisition pose is obtained based on the device pose of the image acquisition device corresponding to the unit scene feature; for any single-view image, determining the approximate acquisition pose corresponding to the single-view acquisition pose of the single-view image from the feature acquisition pose of the unit scene features, to determine the single-view associated features, wherein the single-view associated features are determined based on unit scene features within the feature range of the approximate acquisition pose, and the feature range is determined based on the frustum range of the approximate acquisition pose in the target scene; filtering the single-view image features based on the single-view associated features to determine the filtered newly added single-view features.

[0014] By approximating the pose to obtain single-view associated features, the correlation between single-view associated features and single-view image features is improved, thereby increasing the accuracy of single-view associated feature acquisition.

[0015] According to one embodiment of this disclosure, for any single-view image, determining the approximate acquisition pose corresponding to the single-view acquisition pose of each single-view image from the feature acquisition pose of the unit scene features, so as to determine the single-view associated features of each single-view image, wherein the single-view associated features are determined based on the unit scene features within the feature range of the approximate acquisition pose, and the feature range is determined based on the view frustum range of the approximate acquisition pose in the target scene, includes: for any single-view image, determining associated unit scene features from the unit scene features within the feature range based on the feature distance between the position of the unit scene features within the feature range and the position of the approximate acquisition pose; and performing a local coordinate system transformation on the associated unit scene features based on the point cloud data of the associated unit scene features to transform the associated unit scene features to the coordinate system corresponding to the approximate acquisition pose, so as to determine the single-view associated features.

[0016] Based on feature distance, associated unit scene features are determined from unit scene features within the visible range, and corresponding single-view associated features are determined based on coordinate system transformation. This realizes the view integration of associated unit scene features acquired from multiple views, providing an integration foundation for subsequent integration of single-view associated features and single-view image features.

[0017] According to one embodiment of this disclosure, determining the multi-view scene features corresponding to the current time step based on the newly added features of each single-view image includes: for any single-view image, refining the single-view related features of the single-view image to determine the refined single-view related features, and fusing the refined single-view related features with the newly added single-view features of the single-view image to determine the single-view scene features of the fused single-view image, wherein the feature refinement is based on at least one of feature dimensionality reduction, feature selection, feature encoding, feature optimization, and feature abstraction; and determining the multi-view scene features corresponding to the current time step based on the single-view scene features.

[0018] Single-view scene features are determined based on refined single-view associated features and newly added single-view features, which improves the signal-to-noise ratio of single-view scene features and optimizes the feature quality of single-view scene features.

[0019] According to one embodiment of this disclosure, updating historical scene features based on the scene features to obtain a first target scene feature corresponding to a first time step sequence, wherein the historical scene features are obtained based on the acquired images of processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene, includes: in response to the current time step being the last time step in the first time step sequence, updating the historical scene features based on the multi-view scene features corresponding to the current time step, and determining the first target scene feature.

[0020] When the current time step to be processed is identified as the last time step, the first target scene feature is determined based on the multi-view scene features, which improves the accuracy of the task completion status determination of the first target scene feature extraction, and thus improves the completeness of the first target scene feature extraction.

[0021] According to one embodiment of this disclosure, the method further includes: responding to the current time step being a non-last time step in the first time step sequence, updating the historical scene features based on the multi-view scene features corresponding to the current time step, and determining the updated historical scene features; returning to continue processing the next acquired image corresponding to the next time step to be processed in the first time step sequence, until the last time step is reached, and determining the first target scene features based on the multi-view scene features corresponding to the last time step and the updated historical scene features corresponding to the last time step.

[0022] When it is identified that the current time step to be processed is not the last time step, the next multi-view image corresponding to the next time step is processed until the multi-view image of the last time step is obtained, so as to determine the historical scene features of the target. This improves the accuracy of the task completion status judgment of the first target scene feature extraction task, and thus improves the completeness of the first target scene feature extraction.

[0023] According to one embodiment of this disclosure, the step of predicting and determining the second target scene features corresponding to the second time step sequence based on the first target scene features, wherein the second time step sequence is a future time step sequence of the first time step sequence, includes: predicting the lifecycle of the target scene based on the first target scene features to obtain the predicted second time step sequence, wherein the second time step sequence is one of the remaining lifecycle of the target scene after the first time step sequence, a future time step sequence of the first time step sequence, and the next adjacent time step of the current time step; and predicting the change trend of the target scene within the second time step sequence based on the scene change trend corresponding to the first target scene features to determine the second target scene features.

[0024] Predicting the second target scene features based on the first target scene features improves the prediction accuracy of the second target scene features.

[0025] According to one embodiment of this disclosure, the step of modeling the target scene based on the first target scene features and the second target scene features to obtain a target multidimensional scene includes: determining single-view detection boxes in the multi-view images of each time step in the first time step sequence and the second time step sequence, so as to obtain the scene features within the target boxes of each single-view detection box based on the first target scene features and the second target scene features; performing Gaussian decoding on the scene features within the target boxes of each single-view detection box at each time step to determine the single-view model fragments corresponding to the single-view detection boxes, and constructing scene models at each time step based on each single-view model fragment; determining the target multidimensional scene model based on the scene models at each time step, wherein the target multidimensional scene model is obtained by connecting and rendering the scene models of any two adjacent time steps.

[0026] Based on scene modeling at each time step, the target multidimensional scene modeling is determined, realizing the modeling of the target scene in both spatial and temporal dimensions. This improves the accuracy of the target multidimensional scene modeling in describing the target scene. By integrating and splicing the regional models of the defined area to construct the scene model, the modeling accuracy of the scene modeling is improved. By connecting and rendering the scene models of two adjacent time steps, the modeling quality of the target multidimensional scene modeling is improved, and its visual display effect is optimized.

[0027] According to one embodiment of this disclosure, the method further includes: obtaining the first target scene features and the second target scene features through a trained target scene modeling model, so as to output the target multi-dimensional scene model of the target scene.

[0028] By constructing a multi-dimensional target scene model using a trained target scene modeling model, the modeling accuracy and efficiency of multi-dimensional target scene modeling are improved.

[0029] According to one embodiment of this disclosure, obtaining the target scene modeling model includes: acquiring sample multi-view images at sample time steps in a sample time step sequence; determining training samples based on each sample single-view image, the sample single-view image acquisition pose of each sample single-view image, and the sample object detection box of each sample single-view image in the sample multi-view images; extracting scene features from the sample multi-view images at each sample time step sequentially using the scene modeling model to be trained, to determine the output multi-dimensional scene modeling of the scene modeling model to be trained; training the scene modeling model to be trained based on the output multi-dimensional scene modeling and the label multi-dimensional scene modeling corresponding to the training samples, and determining the trained target scene modeling model.

[0030] Training samples are constructed based on multi-view images of samples at each time step in the sample time step sequence. This enables the scene modeling model to learn the ability to perform multi-dimensional modeling based on spatial and temporal dimensions, reducing the amount of computational computation and complexity of a single operation of the scene modeling model, and reducing the possibility of abnormalities in the model's operation process due to the processing of large amounts of data.

[0031] According to one embodiment of this disclosure, the step of training a scene modeling model to be trained based on the output multidimensional scene modeling and the label multidimensional scene modeling corresponding to the training samples, and determining the trained target scene modeling model, includes: determining the photometric reconstruction loss, perception loss, depth loss, optical flow regularization loss, lifetime regularization loss, sky depth loss, sky transparency loss, and object distribution loss corresponding to the scene modeling model based on the output multidimensional scene modeling and the label multidimensional scene modeling; weighting the photometric reconstruction loss, perception loss, depth loss, optical flow regularization loss, lifetime regularization loss, sky depth loss, sky transparency loss, and object distribution loss to determine the training loss of the scene modeling model; and iteratively optimizing the parameters of the scene modeling model based on the training loss to determine the trained target modeling model.

[0032] Based on various loss mechanisms, a multi-objective optimization strategy is constructed to optimize the scene modeling model, thereby improving the training effect, performance, and robustness of the scene modeling model.

[0033] A second aspect of this disclosure proposes a scene modeling apparatus, comprising: an extraction module for obtaining multi-view scene features of an image corresponding to the current time step based on an image acquired at the current time step to be processed; an acquisition module for updating historical scene features based on the multi-view scene features to obtain first target scene features corresponding to a first time step sequence, wherein the historical scene features are obtained based on images acquired at processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene; a prediction module for predicting and determining second target scene features corresponding to a second time step sequence based on the first target scene features, wherein the second time step sequence is a future time step sequence of the first time step sequence; and a modeling module for modeling a target scene based on the first target scene features and the second target scene features to obtain a target multi-dimensional scene of the target scene.

[0034] This disclosure provides a third aspect of a vehicle for implementing the scene modeling method as described in the first aspect above.

[0035] This disclosure provides a fourth aspect of an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute instructions to implement the scene modeling method as described in the first aspect above.

[0036] The fifth aspect of this disclosure provides a computer-readable storage medium that, when executed by a processor of an electronic device, enables the electronic device to perform the scene modeling method as described in the first aspect above.

[0037] A sixth aspect of this disclosure provides a chip including one or more interface circuits and one or more processors; the interface circuits are configured to receive signals and send the signals to the processors, the signals including computer instructions stored in a memory, which, when executed by the processors, cause the chip to perform the scene modeling method as described in the first aspect above.

[0038] It should be understood that the description herein is not intended to identify key or essential features of the embodiments thereof, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0039] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which: Figure 1 This is a flowchart illustrating a scene modeling method according to an embodiment of the present disclosure; Figure 2This is a flowchart illustrating a scene modeling method according to another embodiment of the present disclosure; Figure 3 This is a flowchart illustrating a scene modeling method according to another embodiment of the present disclosure; Figure 4 This is a flowchart illustrating a scene modeling method according to another embodiment of the present disclosure; Figure 5 This is a flowchart illustrating a method for training a scene modeling model according to an embodiment of the present disclosure. Figure 6 This is a schematic diagram of the structure of a scene modeling apparatus according to an embodiment of the present disclosure; Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure; Figure 8 This is a schematic diagram of the structure of a chip according to an embodiment of the present disclosure. Detailed Implementation

[0040] Embodiments of this disclosure are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.

[0041] The following description, with reference to the accompanying drawings, outlines a scene modeling method, apparatus, vehicle, device, medium, and chip according to embodiments of this disclosure.

[0042] Figure 1 This is a flowchart illustrating a scene modeling method according to an embodiment of the present disclosure, as shown below. Figure 1 As shown, the method includes: S101, Based on the image acquired at the current time step to be processed, obtain the multi-view scene features of the image corresponding to the current time step.

[0043] In this embodiment of the disclosure, the scene for which model construction is required can be determined as the target scene. Within a set time range, multiple single-view images of the target scene can be acquired based on a preset image acquisition time interval, thereby obtaining multiple single-view acquired images of the target scene in each time period within the set time range.

[0044] In this process, the time range can be divided into time steps based on a preset image acquisition time interval using a set time step division method to determine multiple time steps after division. Then, the multiple time steps are constructed into a sequence based on the order from early to late to obtain the corresponding time step sequence. This time step sequence can be determined as the first time step sequence.

[0045] In this scenario, based on the image acquisition time of each of the multiple single-view acquired images in each time period, the time steps of the multiple single-view acquired images can be classified to determine the time step to which each of the multiple single-view acquired images belongs.

[0046] In this scenario, based on any time step, multiple single-view acquired images belonging to that time step can be obtained, and then a multi-view image at that time step can be obtained based on these multiple single-view acquired images as the image acquired at that time step.

[0047] In this embodiment of the disclosure, features required for modeling can be extracted from the images acquired at each of the multiple time steps included in the first time step sequence based on the order from earliest to latest. Among the multiple time steps in the first time step sequence that have not yet undergone feature extraction, the time step with the earliest time dimension can be determined as the current time step to be processed.

[0048] In this scenario, the features required for modeling can be extracted from each single-view image included in the multi-view image corresponding to the image acquired at the current time step, and the extracted features can be determined as the scene features of each single-view image. In this scenario, the scene features of each of the multiple single-view images can be integrated to obtain the scene features corresponding to the multi-view image acquired at the current time step, which can be used as the multi-view scene features of the image corresponding to the current time step.

[0049] It should be noted that, regarding the acquisition of multi-view images, image acquisition devices can be set up at different image acquisition locations in the target scene. Single-view images of corresponding perspectives can be acquired through multiple image acquisition devices, thereby obtaining multi-view images at each time step.

[0050] The image acquisition device can be a camera, which can be deployed in multiple directions around the target scene; no specific limitations are made here.

[0051] S102, update the historical scene features based on multi-view scene features to obtain the first target scene features corresponding to the first time step sequence. The historical scene features are obtained based on the images collected in the processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene.

[0052] In this embodiment of the disclosure, the time step to which the multi-view image that has undergone feature extraction processing in the first time step sequence belongs can be determined as the processed time step in the first time step sequence. In this scenario, the features extracted from the multi-view image corresponding to the processed time step that are required for modeling can be determined as the extracted historical scene features.

[0053] It should be noted that in scenarios where multiple time steps have been processed as the first time step sequence, the historical scene features can include the modeling features of any modeling object in the target scene at different time steps.

[0054] In this scenario, the extracted historical scene features can be updated based on the multi-view scene features corresponding to the current time step. The scene features required for modeling the target scene can be obtained based on the updated scene features. The scene features can be understood as the modeling features of the target scene extracted from the images collected at each time step in the first time step sequence. The scene features can be determined as the first target scene features corresponding to the first time step sequence.

[0055] S103, based on the features of the first target scene, predict and determine the features of the second target scene corresponding to the second time step sequence, wherein the second time step sequence is the future time step sequence of the first time step sequence.

[0056] When modeling a target scene, it may not be possible to acquire images of the target scene for the entire modeling cycle. In this case, the changes that the target scene may undergo in the future time range can be predicted based on the images of the target scene acquired at each time step in the first time step sequence. Specifically, the future time range can be divided into time steps based on the division interval of the first time step sequence, and the sequence of time steps obtained from this division can be determined as the future time step sequence corresponding to the first time step sequence, which is the second time step sequence.

[0057] The modeling cycle of the target scene can be the scene lifecycle of the target scene or other time periods; no specific limitation is made here.

[0058] In this scenario, based on the scene change trend prediction method in related technologies, the first target scene features extracted from the first time step sequence can be used to predict the possible changes of the target scene at each time step in the second time step sequence. Then, based on the prediction results, the scene features corresponding to the second time step sequence can be extracted as the second target scene features corresponding to the second time step sequence.

[0059] Among them, the second target scene feature is the modeling feature used when modeling the target scene in the time dimension corresponding to the second time step sequence.

[0060] S104. Based on the features of the first target scene and the features of the second target scene, the target scene is modeled to obtain the target multi-dimensional scene of the target scene.

[0061] In this embodiment of the disclosure, based on the first target scene features and the second target scene features, the features required for modeling the target scene within the modeling cycle can be obtained. It should be noted that the features include the modeling features of the target scene in the spatial dimension, as well as the modeling features of the target scene in the time dimension within the modeling cycle corresponding to the first time step sequence and the second time step sequence.

[0062] In this scenario, multidimensional modeling methods from relevant technologies can be used to perform multidimensional modeling based on this feature, thereby obtaining the multidimensional model corresponding to the target scene, which serves as the target multidimensional scene model corresponding to the target scene.

[0063] It should be noted that the target multi-dimensional scene modeling proposed in this embodiment includes the display of any object, which will present different display results as time changes.

[0064] The scene modeling method proposed in this disclosure, compared with related technologies that require data processing of large amounts of modeling data to achieve scene modeling, extracts modeling features from multi-view images at each time step in the first time step sequence one by one, reducing the computational load of multi-dimensional modeling algorithms and improving the modeling efficiency of target scenes. Based on the first target scene features, it predicts the second target scene features corresponding to the second time step sequence, improving the prediction accuracy of the second target scene features. Based on the first and second target scene features, it performs multi-dimensional scene modeling, improving the completeness and accuracy of the target multi-dimensional scene modeling, improving the accuracy of the target multi-dimensional scene modeling in displaying changes in the target scene, optimizing the modeling method and the modeling effect of target multi-dimensional scene modeling. In the scenario of executing downstream tasks based on target multi-dimensional scene modeling, it improves the adaptability of the task decision of the downstream task to the target scene. In the scenario of vehicle performance testing based on target multi-dimensional scene modeling, it improves the testing accuracy and precision of vehicle performance testing, and optimizes the task execution quality of the downstream tasks corresponding to target multi-dimensional scene modeling.

[0065] In the above embodiments, the acquisition of the target multi-dimensional scene model can also be combined with Figure 2 To understand further, Figure 2 This is a flowchart illustrating a scene modeling method according to another embodiment of the present disclosure, as shown below. Figure 2 As shown, the method includes: S201, extract the single-view image features corresponding to each single-view image from the multi-view acquired images at the current time step.

[0066] In this embodiment of the disclosure, for the multi-view acquired images acquired at the current time step, any single-view image can be obtained, and the image feature acquisition algorithm in the related technology can be used to process the single-view image to extract the features in the single-view image as the single-view image features of the single-view image.

[0067] S202, Based on historical scene features and single-view image features, determine the newly added single-view features in each single-view image.

[0068] In some possible implementations, the feature acquisition pose of unit scene features included in the historical scene features is obtained, wherein the feature acquisition pose is obtained based on the device pose of the image acquisition device corresponding to the unit scene features.

[0069] In this embodiment of the disclosure, each unit feature (token) included in the extracted historical scene features can be determined as a unit scene feature in the extracted historical scene features. In this scenario, a single-view image of any unit scene feature can be obtained, and the device pose of the image acquisition device when acquiring the single-view image can be obtained. Based on the device pose, the acquisition pose of the unit scene feature extracted in the single-view image can be determined.

[0070] The acquisition pose can be determined as the feature acquisition pose of the scene features of the unit.

[0071] It should be noted that the image acquisition device mentioned above can be a camera or other devices, and no specific limitation is made here.

[0072] In some possible implementations, for any single-view image, an approximate acquisition pose corresponding to the single-view acquisition pose of the single-view image is determined from the feature acquisition pose of the unit scene features, so as to determine the single-view associated features. The single-view associated features are determined based on the unit scene features within the feature range of the approximate acquisition pose, and the feature range is determined based on the view frustum range of the approximate acquisition pose in the target scene.

[0073] In this embodiment of the disclosure, the pose of the image acquisition device that acquires any single-view image can be obtained. This pose can be determined as the single-view acquisition pose of the single-view image. In this scenario, based on the single-view acquisition pose, an approximate pose that meets the set conditions can be determined from the feature acquisition pose of the unit scene features, and used as the approximate acquisition pose corresponding to the single-view acquisition pose.

[0074] In some possible implementations, if the pose deviation between the feature acquisition pose of the unit scene features and the single-view acquisition pose of the single-view image falls within a set pose deviation range, the feature acquisition pose is determined to be an approximate acquisition pose of the single-view acquisition pose.

[0075] In this embodiment of the disclosure, the feature acquisition pose of any unit scene feature and the single-view acquisition pose corresponding to the single-view image can be processed by the pose deviation algorithm in the related technology. The deviation value between the feature acquisition pose and the single-view acquisition pose is obtained based on the result of the algorithm processing, and is used as the pose deviation between the two.

[0076] In this scenario, the pose deviation can be compared with the set pose deviation range. When the pose deviation falls within the pose deviation range, it can be determined that the feature acquisition pose is the approximate acquisition pose of the single-view acquisition pose.

[0077] In some possible implementations, single-view associated features are determined based on unit scene features within the feature range of the approximate acquisition pose, wherein the feature range is determined based on the visible range of the approximate acquisition pose in the target scene, and the visible range includes at least the visual cone range corresponding to the approximate acquisition pose.

[0078] In this embodiment of the present disclosure, the range of the visual cone corresponding to the approximate acquisition pose can be determined as the visible range of the approximate acquisition pose. In this scenario, some features that are related to the single-view image features can be selected from the unit scene features included in the visible range of the approximate acquisition pose. These features can be determined as the single-view related features of the single-view image in the extracted historical scene features.

[0079] In some possible implementations, the correlation degree algorithm in related technologies can be used to process the unit scene features and single-view image features included in the visible range to obtain the correlation degree parameter between each unit scene feature and the single-view image feature, and then the corresponding single-view related features can be determined from the unit scene features included in the visible range based on the correlation degree parameter.

[0080] In other possible implementations, for any single-view image, associated unit scene features are determined from the unit scene features within the feature range based on the feature distance between the unit scene features within the feature range and the position of the approximate acquisition pose.

[0081] In this embodiment of the disclosure, single-view associated features can also be determined based on the distance relationship between each unit scene feature and the approximate acquisition pose. In this case, the feature distance algorithm in the related technology can be used to process the position features corresponding to the unit scene features and the approximate acquisition pose to obtain the feature distance between the unit scene features and the position corresponding to the approximate acquisition pose.

[0082] In this scenario, a set number of unit scene features can be determined based on the feature distances corresponding to each unit scene feature to obtain the corresponding single-view associated features. Specifically, the unit scene features within the feature range can be sorted in ascending order of feature distance values ​​to obtain sorted unit scene features. Starting from the first unit scene feature in the sorted unit scene features, a set number of unit scene features can be determined as associated unit scene features.

[0083] As an example, such as Figure 3 As shown, for Figure 3 The approximate acquisition pose shown includes unit scene features k1, k2, k3, ..., kK, kK+1, ..., kK+n within its feature range. These unit scene features k1, k2, k3, ..., kK, kK+1, ..., kK+n can be sorted in ascending order of distance value. Then, the distance can be determined from the sorted k1, k2, k3, ..., kK, kK+1, ..., kK+n. Figure 3 The approximate acquisition pose corresponds to the nearest K unit scene features k1, k2, k3, ..., kK, and k1, k2, k3, ..., kK are identified as the associated unit scene features of the single-view image in the extracted historical scene features.

[0084] In some possible implementations, the scene features of the associated units are transformed into a local coordinate system based on the point cloud data of the scene features of the associated units, so as to transform the scene features of the associated units into the corresponding coordinate system of the approximate acquisition pose, so as to determine the single-view associated features.

[0085] In this embodiment of the disclosure, the scene features of the associated units obtained by approximate acquisition pose may not be acquired based on the same single viewpoint. In this scenario, it is necessary to integrate the scene features of the associated units. Specifically, point cloud data of the scene features of the associated units can be obtained to obtain the coordinate information of the scene features of the associated units in the corresponding point cloud coordinate system.

[0086] In this scenario, the point cloud coordinate information of the associated unit scene features can be processed by the coordinate system transformation algorithm in the relevant technology to transform the coordinate information of the associated unit scene features to the coordinate system corresponding to the approximate acquisition pose, and the associated unit scene features after coordinate system transformation are determined as single-view associated features.

[0087] The coordinate system corresponding to the approximate acquisition pose can be the camera coordinate system of the image acquisition device corresponding to the approximate acquisition pose, or it can be other coordinate systems; no specific limitation is made here.

[0088] In some possible implementations, single-view image features are filtered based on single-view associated features to determine the newly added single-view features after filtering.

[0089] In this embodiment of the disclosure, there may be some overlap between the single-view correlation feature and the single-view image feature. The overlapping feature can be identified as the correlation unit feature of the single-view correlation feature in the single-view image feature.

[0090] In this scenario, the associated unit features can be filtered out from the single-view image features to obtain the remaining features after filtering. These remaining features are the features that need to be added to the extracted historical scene features from the single-view image features, which are the single-view newly added features carried in the single-view image features.

[0091] S203, based on the newly added features of each single view, determine the multi-view scene features corresponding to the current time step.

[0092] In some possible implementations, for any single-view image, the single-view associated features of the single-view image are refined to determine the refined single-view associated features, and the refined single-view associated features are fused with the newly added single-view features of the single-view image to determine the fused single-view scene features.

[0093] Feature refinement is based on at least one of the following: feature dimensionality reduction, feature selection, feature encoding, feature optimization, and feature abstraction.

[0094] In this embodiment of the disclosure, the single-view associated features can be refined based on the feature refinement method in the related technology to obtain the refined single-view associated features.

[0095] It should be noted that feature refinement methods may include at least one of the following: feature dimensionality reduction methods, feature selection methods, feature encoding methods, feature optimization methods, and feature abstraction methods in related technologies. They may also include other feature processing methods that can achieve feature refinement, without specific limitations here.

[0096] In this embodiment of the disclosure, the refined single-view associated features and the newly added single-view features can be processed by the feature fusion algorithm in the related technology, and then the feature fusion between the refined single-view associated features and the newly added single-view features can be realized based on the algorithm processing, and the fused features are determined as single-view scene features.

[0097] In some possible implementations, the multi-view scene features corresponding to the current time step are determined based on the single-view scene features of each single-view image.

[0098] In this embodiment of the disclosure, for the multi-view image corresponding to the current time step to be processed, after obtaining the single-view scene features of each single-view image included therein, the single-view scene features can be integrated, and the multi-view scene features corresponding to the multi-view image can be obtained based on the integrated features, which are used as the multi-view scene features corresponding to the current time step.

[0099] S204, update the historical scene features based on multi-view scene features to obtain the first target scene features corresponding to the first time step sequence. The historical scene features are obtained based on the images collected in the processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene.

[0100] In some possible implementations, in response to the current time step being the last time step in the first time step sequence, the historical scene features are updated based on the multi-view scene features corresponding to the current time step to determine the first target scene features.

[0101] In this embodiment of the disclosure, for the current time step in which feature extraction is performed, it can be identified whether the time step is the last time step in the first time step sequence. In the scenario where the time step is the last time step, it can be determined that after extracting the multi-view scene features of the multi-view image corresponding to the time step, all the multi-view images collected for the target scene that needs to be modeled have been subjected to scene feature extraction.

[0102] In this scenario, scene features can be integrated and updated based on the extracted historical scene features and the multi-view scene features extracted from the multi-view images at the last time step. The single-view image features corresponding to each other in the extracted historical scene features are deleted, and the currently extracted multi-view scene features are merged into the deleted historical scene features. This achieves the integration and update between the multi-view scene features and the extracted historical scene features, and then the updated historical scene features are determined as the first target scene features required for multi-dimensional modeling of the target scene.

[0103] This can be understood as follows: by using the refined single-view associated features included in the multi-view scene features, the unrefined single-view associated features in the extracted historical scene features can be overwritten and updated, and the newly added single-view features included in the multi-view scene features can be merged into the historical scene features. This achieves the feature increment from the distinguishing features carried in the current multi-view image to the extracted historical scene features, thereby obtaining the first target scene features.

[0104] In other possible implementations, in response to the current time step being a non-last time step in the first time step sequence, the historical scene features are updated based on the multi-view scene features, the updated historical scene features are determined, and the process returns to continue processing the next acquired image corresponding to the next time step to be processed in the first time step sequence until the last time step is reached. Based on the multi-view scene features corresponding to the last time step and the updated historical scene features corresponding to the last time step, the first target scene features are determined.

[0105] In this embodiment of the disclosure, if the current processing time step is not the last time step in the first time step sequence, it can be determined that after extracting the multi-view scene features of the multi-view image corresponding to the current time step, it is still necessary to continue to extract scene features from the multi-view images of the remaining unprocessed time steps included in the first time step sequence.

[0106] In this scenario, the extracted historical scene features can be updated based on the multi-view scene features extracted from the multi-view image at the current time step, resulting in updated historical scene features. The earliest time step from the remaining unprocessed time steps is determined as the next time step to be processed. Scene features are then extracted from the next acquired image corresponding to the next time step, and the updated historical scene features are updated based on the extracted new multi-view scene features, until the multi-view scene features corresponding to the last time step are extracted. Based on the multi-view scene features corresponding to the last time step and the historical scene features extracted from all time steps in the first time step sequence excluding the last time step, the first target scene features required for target scene modeling can be obtained.

[0107] S205, based on the features of the first target scene, predict and determine the features of the second target scene corresponding to the second time step sequence, wherein the second time step sequence is the future time step sequence of the first time step sequence.

[0108] In some possible implementations, based on the characteristics of the first target scene, a lifecycle prediction is performed on the target scene to obtain a predicted second time step sequence, wherein the second time step sequence is the remaining lifecycle of the target scene after the first time step sequence, the future time step sequence of the first time step sequence, and the next adjacent time step of the current time step.

[0109] In this embodiment of the disclosure, when the modeling period of the target scene is its life cycle, the life cycle of the target scene can be predicted based on the first target scene features. That is, the life evolution process and time of the target scene can be predicted based on the first target scene features, thereby obtaining the predicted life cycle of the target scene.

[0110] In this scenario, the time period corresponding to the first time step sequence for which image acquisition has already taken place can be removed from the predicted lifetime, and the remaining lifetime after removal can be determined as the time range to which the second time step sequence for which image acquisition has not taken place belongs.

[0111] In this embodiment of the disclosure, the second time step sequence can also be the next adjacent time step of the current time step. It can be understood that the scene features that are integrated with the multi-view scene features extracted at the current time step and the extracted historical scene features can be used to predict the possible changes of the target scene in the next adjacent time step. In this scenario, the second time step sequence can be determined as the next adjacent time step corresponding to the current time step.

[0112] In some possible implementations, the change trend of the target scene within the second time step sequence is predicted based on the scene change trend corresponding to the first target scene features, so as to determine the second target scene features.

[0113] In this embodiment of the disclosure, the first target scene features can be processed by a scene change trend prediction algorithm in the related technology. Then, based on the result of the algorithm processing, the scene changes that the target scene may produce in the second time step sequence can be obtained. Then, from the predicted scene changes corresponding to the second time step sequence, the scene modeling features required for modeling the target scene in the time dimension corresponding to the second time step sequence can be obtained as the second target scene features corresponding to the second time step sequence.

[0114] S206. Based on the features of the first target scene and the features of the second target scene, the target scene is modeled to obtain the target multi-dimensional scene of the target scene.

[0115] In some possible implementations, single-view detection boxes are determined in the multi-view images at each time step in the first time step sequence and the second time step sequence, so as to obtain the scene features within the target box of each single-view detection box based on the first target scene features and the second target scene features.

[0116] In this embodiment of the disclosure, for any time step included in the first time step sequence and the second time step sequence, each single-view image included in the multi-view image acquired at that time step is provided with a corresponding detection box, wherein the detection box can define the modeling object in the single-view image.

[0117] In this scenario, the detection box in the single-view image can be identified as the single-view detection box.

[0118] In this embodiment of the disclosure, each unit scene feature included in the first target scene feature and the second target scene feature corresponding to the target scene has its own single-view detection box. In this scenario, the detection boxes of each unit scene feature included in the first target scene feature and the second target scene feature can be divided based on the single-view detection boxes corresponding to each time step in the first time step sequence and the second time step sequence, so as to determine the single-view detection box to which each unit scene feature belongs from all the single-view detection boxes.

[0119] In this scenario, all the unit scene features included in each single-view detection box corresponding to the first time step sequence and the second time step sequence can be obtained. Among them, all the unit scene features included in any single-view detection box can be determined as the scene features within the target box corresponding to that single-view detection box.

[0120] In some possible implementations, Gaussian decoding is performed on the scene features within the target box of the single-view detection box at each time step to determine the single-view model fragments corresponding to the single-view detection box, and scene modeling at each time step is constructed based on each single-view model fragment.

[0121] Specifically, for any time step, Gaussian decoding is performed on the scene features within the target box of the single-view detection box at that time step to determine the single-view model fragment corresponding to the single-view detection box.

[0122] In this embodiment of the disclosure, Gaussian decoding can be performed on the scene features within the target box of the single-view detection box at any time step based on the Gaussian decoding method in the related technology to obtain the corresponding Gaussian distribution. Based on the rendering of the Gaussian distribution, model fragments of the region bounded by the single-view detection box can be constructed, and the model fragments can be identified as the single-view model fragments corresponding to the single-view detection box.

[0123] In some possible implementations, there are multiple single-view detection boxes at each time step, and multiple single-view region detection boxes corresponding to the defined region are determined from the multiple single-view detection boxes based on the defined region of any single-view detection box.

[0124] In this embodiment of the disclosure, there are multiple single-view images at each time step in the time step sequence. In this scenario, there are multiple single-view detection boxes at each time step. In this scenario, for any single-view detection box, the region inside the box of the single-view detection box can be obtained, and this region can be determined as the bounding region of the single-view detection box. In this scenario, the associated detection box of the bounding region can be determined from the multiple single-view detection boxes included in the time step, and used as multiple single-view region detection boxes corresponding to the bounding region.

[0125] Among them, the multiple single-view region detection boxes corresponding to the bounding area mentioned above can be understood as different single-view regions of the same area in the target scene.

[0126] In some possible implementations, the region model of the bounding region is determined based on the single-view model fragments of each of the multiple single-view region detection boxes.

[0127] In this embodiment of the disclosure, the model fragments of each single-view model of multiple single-view region detection boxes can be spliced ​​together based on the three-dimensional model splicing method in the related technology to obtain the spliced ​​multi-view model, which can then be determined as the region model of the bounded area.

[0128] In some possible implementations, the regional models at each time step are integrated and stitched together to construct the scene model at that time step.

[0129] In this embodiment of the disclosure, for any time step, the region models obtained from the detection boxes of each single-view region in the multi-view image corresponding to that time step can be integrated and stitched together to obtain a model based on the multi-view image of the target scene at that time step, which is used as the scene model corresponding to that time step.

[0130] In some possible implementations, a target multidimensional scene model is determined based on the scene modeling at each time step, wherein the target multidimensional scene model is obtained by connecting and rendering the scene models of any two adjacent time steps.

[0131] In this embodiment of the disclosure, for the target scene, the scene modeling based on each time step can obtain a model that can characterize the changes of the target scene as the time step changes. That is, the modeling includes the relevant feature information of the target scene in the spatial dimension and the relevant feature information of the target scene in the time dimension. In this scenario, the above-mentioned modeling including the spatial dimension and time dimension of the target scene can be determined as the target multidimensional scene modeling of the target scene.

[0132] In some possible implementations, the scene models of any two adjacent time steps in the first time step sequence and the second time step sequence are connected and rendered to integrate the scene models in the time dimension and determine the integrated target multi-dimensional scene model.

[0133] In this embodiment of the disclosure, the scene models of any two adjacent time steps can be connected, integrated and rendered, thereby realizing the integration of the scene models of any two adjacent time steps in the time dimension. The model presented after the integration of the two scene models can show the scene changes of the target scene within the time range corresponding to the two time steps.

[0134] In this scenario, the scene modeling of each time step in the time step sequence can be integrated and rendered based on the connection between the scene models of any two adjacent time steps, thereby achieving the integration of scene models in the time dimension and obtaining the target multi-dimensional scene model corresponding to the target scene.

[0135] The scene modeling method proposed in this disclosure, compared to related technologies that require data processing of large amounts of modeling data to achieve scene modeling, extracts modeling features step-by-step in the first and second time-step sequences, reducing the computational load of multi-dimensional modeling algorithms and improving the modeling efficiency of the target scene. It is based on newly added single-view features that differ from extracted historical scene features in single-view image features, and integrates these newly added single-view features with the extracted historical scene features, reducing the computational load of feature integration algorithms and the computational load of feature representation acquisition required for multi-dimensional scene modeling, thereby reducing the complexity of the multi-dimensional scene modeling algorithm. The single-view associated features are refined, and the extracted historical scene features are updated based on the refined single-view associated features and newly added single-view features, which improves the feature quality of the target historical scene features and thus improves the modeling quality of the target 3D scene modeling. Based on the first target scene features, the second target scene features corresponding to the second time step sequence are predicted, which improves the prediction accuracy of the second target scene features. Multi-dimensional scene modeling is performed based on the first and second target scene features, which improves the completeness and accuracy of the target multi-dimensional scene modeling, improves the accuracy of the target multi-dimensional scene modeling in displaying the changes in the target scene, and optimizes the modeling method and modeling effect of the target multi-dimensional scene modeling.

[0136] In the above embodiments, the acquisition of the target multi-dimensional scene model can also be combined with Figure 4 understand, Figure 4 This is a flowchart illustrating a scene modeling method according to another embodiment of the present disclosure, as shown below. Figure 4 As shown, the method includes: S401 obtains the first target scene features and the second target scene features through the trained target scene modeling model, so as to output the target scene multidimensional scene model.

[0137] In this embodiment of the disclosure, the target scene can also be modeled in multiple dimensions by relying on the model, so as to obtain the target multi-dimensional scene model of the target scene proposed in the above embodiment.

[0138] The training of the target scene model can be understood in conjunction with the following: In some possible implementations, sample multi-view images at sample time steps in the sample time step sequence are acquired, and training samples are determined based on the sample single-view images in the sample multi-view images, the sample single-view image acquisition pose of each sample single-view image, and the sample object detection box of each sample single-view image.

[0139] In this embodiment of the disclosure, the acquisition of the target multi-dimensional scene modeling proposed in the above embodiments can be achieved through a model, wherein the model used to construct the multi-dimensional scene modeling can be determined as a scene modeling model.

[0140] In some possible implementations, before constructing scene modeling based on the scene modeling model, the model can be trained, and the time step sequence corresponding to the sample data used during training can be determined as the sample time step sequence. In this scenario, multi-view images at the sample time steps in the sample time step sequence can be obtained and determined as the sample multi-view images required for model training.

[0141] Among them, the multi-view sample image includes each single-view image, which can be identified as each sample single-view image included in the multi-view sample image. The image information of the sample single-view image carries the acquisition pose when the image was acquired, which can be identified as the sample single-view image acquisition pose corresponding to the sample single-view image. The sample single-view image also carries a detection box that defines the modeling object, which can be identified as the sample object detection box in the sample single-view image.

[0142] In this scenario, samples can be constructed based on the sample construction method in related technologies, based on the multi-view images of samples at the sample time step, the single-view images of the sample single-view images included in the multi-view images of samples, and the sample object detection boxes in the single-view images of samples, and the constructed samples are determined as training samples used during model training.

[0143] In some possible implementations, scene features are extracted sequentially from the multi-view images of samples at each time step using the scene modeling model to be trained, so as to determine the output multi-dimensional scene modeling of the scene modeling model to be trained.

[0144] In this embodiment of the disclosure, the model to be trained can be determined as the scene modeling model to be trained. In this scenario, training samples can be input into the scene modeling model to be trained, and the scene modeling model can extract the single-view scene features required for modeling the multi-view images of the training samples, including the single-view images of the samples.

[0145] In other words, through the scene modeling model, the multi-view images of samples at each time step in the sample time step sequence can be processed sequentially. The order can be determined based on the time sequence from early to late. In this scenario, the scene modeling model can extract scene features from each single-view image of the multi-view image of any sample time step to obtain the single-view scene features corresponding to each single-view image of the multi-view image.

[0146] The process of extracting scene features from single-view images of samples can be combined with the above. Figures 1 to 3 The relevant content in the embodiments is understood and will not be repeated here.

[0147] In this scenario, the scene modeling model to be trained can extract the single-view scene features from the single-view images of each sample multi-view image, obtain the historical scene features required for scene modeling of the scene corresponding to the multi-view images of each sample under the time step sequence of the sample, and then construct the corresponding multi-dimensional scene model, which is used as the output multi-dimensional scene model of the scene model to be trained.

[0148] It should be noted that the scene modeling model can be built based on the neural network model (Transformer) in related technologies, or it can be built based on other models; no specific limitation is made here.

[0149] In some possible implementations, the scene modeling model to be trained is trained based on the output multidimensional scene modeling and the label multidimensional scene modeling corresponding to the training samples, and the trained target scene modeling model is determined.

[0150] In this embodiment of the disclosure, the model parameters of the scene modeling model can be iteratively optimized based on the output multidimensional scene modeling of the scene modeling model to be trained and the label multidimensional scene modeling corresponding to the training samples, using the model training method in the related art, so as to realize the model training of the scene modeling model.

[0151] Among them, the training loss of the scene modeling model can be determined based on the output multidimensional scene modeling and the label multidimensional scene modeling, and the parameters of the scene modeling model can be iterated based on the training loss.

[0152] In some possible implementations, based on output multidimensional scene modeling and label multidimensional scene modeling, the photometric reconstruction loss, perception loss, depth loss, optical flow regularization loss, lifetime regularization loss, sky depth loss, sky transparency loss, and object distribution loss corresponding to the scene modeling model are determined. The photometric reconstruction loss, perception loss, depth loss, optical flow regularization loss, lifetime regularization loss, sky depth loss, sky transparency loss, and object distribution loss are weighted to determine the model training loss of the scene modeling model.

[0153] As an example, the following formula can be used to understand this:

[0154] In the above formula, This represents the training loss of the scene modeling model to be trained. Indicates the loss in photometric reconstruction. Indicates perceived loss. Indicates deep loss. This represents the optical flow regularization loss. This represents the lifecycle regularization loss. Indicates the loss of sky depth. This indicates a loss of sky transparency. This represents the loss due to the distribution of objects.

[0155] as well as, The weighting coefficients representing the photometric reconstruction loss. The weighting coefficients represent the perceived loss. The weighting coefficients represent the depth loss. The weighting coefficients representing the optical flow regularization loss are: The weighting coefficients represent the lifecycle regularization loss. The weighting coefficients representing the sky depth loss. The weighting coefficients representing the loss of sky transparency. The weighting coefficients represent the loss due to object distribution.

[0156] In some possible implementations, the parameters of the scene modeling model are iteratively optimized based on the training loss to determine the trained target modeling model.

[0157] In this embodiment of the disclosure, the iterative optimization method of model parameters in related technologies can be used to iteratively optimize the parameters of the scene modeling model to be trained based on the training loss, so as to obtain the scene modeling model with adjusted parameters.

[0158] In this scenario, when the scene modeling model after parameter adjustment is identified as meeting the preset model training termination condition, the model training can be terminated, and the scene modeling model after parameter adjustment can be identified as the trained target scene modeling model.

[0159] Accordingly, when it is identified that the scene modeling model after parameter adjustment does not meet the model training termination condition, the process can return to obtain the next training sample and perform the next round of model training on the scene modeling model after parameter adjustment until the model training termination condition is met, and the trained target scene modeling model can be obtained.

[0160] The scene modeling method proposed in this disclosure constructs training samples based on multi-view images of samples at each time step in the sample time step sequence. This enables the scene modeling model to learn the ability to perform multi-dimensional modeling based on spatial and temporal dimensions. By extracting scene features from the multi-view images of samples at each time step sequentially frame by frame, compared to related technologies that input a large amount of data required for modeling into the model, the algorithm load and computational complexity of the scene modeling model are reduced. This reduces the possibility of abnormalities in the model's computation process due to large amounts of data processing, improves the model's stability and robustness, and thus improves the modeling efficiency and accuracy of multi-dimensional scene modeling. It also optimizes the model performance and improves the model's generalization ability. By constructing a target multi-dimensional scene model using the trained target scene modeling model, the modeling accuracy and speed of target multi-dimensional scene modeling are improved, and the modeling quality of target multi-dimensional scene modeling is optimized.

[0161] In the above embodiments, the training of the scene modeling model can also be combined with... Figure 5 To understand further, Figure 5 This is a flowchart illustrating a method for training a scene modeling model according to an embodiment of the present disclosure.

[0162] like Figure 5 As shown, Figure 5 The sequence consisting of time step t-2, time step t-1, and time step t is determined as the sample time step sequence, as shown below. Figure 5 As shown, the single-view images of each sample included in the multi-view image t-2 corresponding to time step t-2 are input to... Figure 5 The update module t-2 shown extracts features from each single-view image included in the multi-view image t-2, thereby extracting the historical scene features t-2 carried in the multi-view image t-2.

[0163] like Figure 5As shown, after extracting the historical scene feature t-2, the single-view images of each sample included in the multi-view image t-1 corresponding to the sample time step t-1 can be input into the database. Figure 5 The update module t-1 shown extracts features from each single-view image included in the multi-view image t-1, thereby extracting the scene features carried in the multi-view image t-1.

[0164] In this scenario, historical scene feature t-2 can be considered as extracted historical scene features in the current scene. The associated scene features extracted from the multi-view image t-1 can be determined from historical scene feature t-2, and based on... Figure 5 The filter 1 shown refines the associated scene features. Based on the refined associated scene features and the scene features carried in the sample multi-view image t-1 extracted by the update module t-1, the following is obtained: Figure 5 The historical scene features t-1 are shown.

[0165] like Figure 5 As shown, after extracting the historical scene feature t-1, through... Figure 5 The update module t shown is paired with Figure 5 The scene features of the sample multi-view image t corresponding to the sample time step t are extracted, and the associated scene features of the sample multi-view image t in the historical scene features t-1 are refined by filter 2. Then, the refined associated scene features and the scene features extracted by the sample multi-view image t extracted by the update module t are integrated to obtain the historical scene features t extracted from the sample multi-view images of each sample time step in the sample time step sequence.

[0166] In this scenario, the scene modeling model to be trained can also be used to predict the scene features of the future sample time step sequence corresponding to the sample time step sequence based on the historical scene features t, thereby obtaining the predicted scene features corresponding to the future sample time step sequence.

[0167] like Figure 5 As shown, the scene modeling model deploys... Figure 5 The Gaussian decoding layer shown is used to perform Gaussian decoding on historical scene features t and predicted scene features, thereby constructing a multi-dimensional scene model of the corresponding scene.

[0168] In this scenario, the multidimensional scene model can be used as the output multidimensional scene model of the scene modeling model. Based on the output multidimensional scene model and the corresponding label multidimensional scene model, the corresponding training loss can be obtained to train and optimize the model, resulting in the trained target scene modeling model.

[0169] The training method for the scene modeling model proposed in this disclosure constructs training samples based on multi-view images of samples at each time step in the sample time step sequence. This enables the scene modeling model to learn the ability to perform multi-dimensional modeling based on spatial and temporal dimensions. By sequentially extracting scene features from the multi-view images of samples at each time step frame by frame, compared to related technologies that input all the large amounts of data required for modeling into the model, the algorithm load and computational complexity of the scene modeling model are reduced. This reduces the possibility of abnormalities in the model's computation process due to the processing of large amounts of data, improves the model's stability and robustness, and thus improves the modeling efficiency and accuracy of multi-dimensional scene modeling, optimizes the model performance of the scene modeling model, and enhances the model's generalization ability.

[0170] Corresponding to the scene modeling methods proposed in the above embodiments, an embodiment of this disclosure also proposes a scene modeling device. Since the scene modeling device proposed in this disclosure corresponds to the scene modeling methods proposed in the above embodiments, the implementation methods of the above-mentioned scene modeling methods are also applicable to the scene modeling device proposed in this disclosure, and will not be described in detail in the following embodiments.

[0171] Figure 6 This is a schematic diagram of the structure of a scene modeling apparatus according to an embodiment of the present disclosure, as shown below. Figure 6 As shown, the scene modeling device 600 includes an extraction module 61, an acquisition module 62, a prediction module 63, and a modeling module 64, wherein: Extraction module 61 is used to obtain multi-view scene features of the image corresponding to the current time step based on the image acquired at the current time step to be processed; The acquisition module 62 is used to update the historical scene features based on the multi-view scene features to obtain the first target scene features corresponding to the first time step sequence. The historical scene features are obtained based on the images acquired in the processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene. Prediction module 63 is used to predict and determine the second target scene features corresponding to the second time step sequence based on the first target scene features, wherein the second time step sequence is the future time step sequence of the first time step sequence; Modeling module 64 is used to model the target scene based on the first target scene features and the second target scene features to obtain the target multi-dimensional scene of the target scene.

[0172] In this embodiment of the disclosure, the extraction module 61 is further configured to: extract single-view image features corresponding to each single-view image from the multi-view acquisition images acquired at the current time step; determine the newly added single-view features in each single-view image based on the historical scene features and the single-view image features; and determine the multi-view scene features corresponding to the current time step based on the newly added single-view features.

[0173] In this embodiment of the disclosure, the extraction module 61 is further configured to: acquire the feature acquisition pose of unit scene features in historical scene features, wherein the feature acquisition pose is obtained based on the device pose of the image acquisition device corresponding to the unit scene feature; for any single-view image, determine the approximate acquisition pose corresponding to the single-view acquisition pose of the single-view image from the feature acquisition pose of the unit scene features, so as to determine the single-view associated features, wherein the single-view associated features are determined based on the unit scene features within the feature range of the approximate acquisition pose, and the feature range is determined based on the frustum range of the approximate acquisition pose in the target scene; filter the single-view image features based on the single-view associated features to determine the filtered single-view newly added features.

[0174] In this embodiment of the present disclosure, the extraction module 61 is further configured to: for any single-view image, determine associated unit scene features from the unit scene features within the feature range based on the feature distance between the unit scene features within the feature range and the position of the approximate acquisition pose; and perform local coordinate system transformation on the associated unit scene features based on the point cloud data of the associated unit scene features to transform the associated unit scene features to the corresponding coordinate system of the approximate acquisition pose in order to determine the single-view associated features.

[0175] In this embodiment of the disclosure, the extraction module 61 is further configured to: for any single-view image, refine the single-view related features of the single-view image, determine the refined single-view related features, and fuse the refined single-view related features with the newly added single-view features of the single-view image to determine the single-view scene features of the fused single-view image, wherein the feature refinement is based on at least one of feature dimensionality reduction, feature selection, feature encoding, feature optimization, and feature abstraction; and determine the multi-view scene features corresponding to the current time step based on each single-view scene feature.

[0176] In this embodiment of the disclosure, the acquisition module 62 is further configured to: in response to the current time step being the last time step in the first time step sequence, update the historical scene features based on the multi-view scene features corresponding to the current time step, and determine the first target scene features.

[0177] In this embodiment of the disclosure, the acquisition module 62 is further configured to: respond to the current time step being a non-last time step in the first time step sequence, update the historical scene features based on the multi-view scene features corresponding to the current time step, and determine the updated historical scene features; return to continue processing the next acquired image corresponding to the next time step to be processed in the first time step sequence, until the last time step is reached, and determine the first target scene features based on the multi-view scene features corresponding to the last time step and the updated historical scene features corresponding to the last time step.

[0178] In this embodiment of the disclosure, the prediction module 63 is further configured to: perform lifecycle prediction on the target scene based on the first target scene features to obtain a predicted second time step sequence, wherein the second time step sequence is one of the future time step sequence of the target scene after the first time sequence and the next adjacent time step of the current time step; and predict the change trend of the target scene in the second time step sequence based on the scene change trend corresponding to the first target scene features to determine the second target scene features.

[0179] In this embodiment of the disclosure, the modeling module 64 is further configured to: determine single-view detection boxes in the multi-view images of each time step in the first time step sequence and the second time step sequence, so as to obtain the scene features within the target box of each single-view detection box based on the first target scene features and the second target scene features; perform Gaussian decoding on the scene features within the target box of each single-view detection box at each time step to determine the single-view model fragments corresponding to the single-view detection boxes, and construct scene models at each time step based on each single-view model fragment; and determine the target multi-dimensional scene model based on the scene models at each time step, wherein the target multi-dimensional scene model is obtained by connecting and rendering the scene models of any two adjacent time steps.

[0180] In this embodiment of the disclosure, the modeling module 64 is further configured to: obtain first target scene features and second target scene features through a trained target scene modeling model, so as to output a target multi-dimensional scene model of the target scene.

[0181] In this embodiment of the present disclosure, the modeling module 64 is further configured to: acquire sample multi-view images at sample time steps in the sample time step sequence; determine training samples based on the sample single-view images, the sample single-view image acquisition poses of each sample single-view image, and the sample object detection boxes of each sample single-view image; extract scene features from the sample multi-view images at each sample time step sequentially through the scene modeling model to be trained, so as to determine the output multi-dimensional scene modeling of the scene modeling model to be trained; train the scene modeling model to be trained based on the output multi-dimensional scene modeling and the label multi-dimensional scene modeling corresponding to the training samples, and determine the trained target scene modeling model.

[0182] In this embodiment of the disclosure, the modeling module 64 is further configured to: determine the photometric reconstruction loss, perception loss, depth loss, optical flow regularization loss, lifetime regularization loss, sky depth loss, sky transparency loss, and object distribution loss corresponding to the scene modeling model based on the output multidimensional scene modeling and the label multidimensional scene modeling; weight the photometric reconstruction loss, perception loss, depth loss, optical flow regularization loss, lifetime regularization loss, sky depth loss, sky transparency loss, and object distribution loss to determine the training loss of the scene modeling model; and iteratively optimize the parameters of the scene modeling model based on the training loss to determine the trained target modeling model.

[0183] The scene modeling device proposed in this disclosure, compared to related technologies that require data processing of large amounts of modeling data to achieve scene modeling, extracts modeling features step-by-step in the first and second time-step sequences, reducing the computational load of multi-dimensional modeling algorithms and improving the modeling efficiency of the target scene. It adds new single-view features based on differences between single-view image features and extracted historical scene features, and integrates these new single-view features with the extracted historical scene features, reducing the computational load of feature integration algorithms and the computational load of feature representation acquisition required for multi-dimensional scene modeling, thereby reducing the complexity of the multi-dimensional scene modeling algorithm. The single-view associated features are refined, and the extracted historical scene features are updated based on the refined single-view associated features and newly added single-view features, which improves the feature quality of the target historical scene features and thus improves the modeling quality of the target 3D scene modeling. Based on the first target scene features, the second target scene features corresponding to the second time step sequence are predicted, which improves the prediction accuracy of the second target scene features. Multi-dimensional scene modeling is performed based on the first and second target scene features, which improves the completeness and accuracy of the target multi-dimensional scene modeling, improves the accuracy of the target multi-dimensional scene modeling in displaying the changes in the target scene, and optimizes the modeling method and modeling effect of the target multi-dimensional scene modeling.

[0184] To achieve the above embodiments, this disclosure also provides a vehicle for implementing the scene modeling method proposed in the above embodiments.

[0185] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. For example, the electronic device 700 may be a vehicle, mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0186] Reference Figure 7The electronic device 700 may include one or more of the following components: processing component 702, memory 704, power component 706, multimedia component 708, audio component 710, input / output (I / O) interface 712, sensor component 714, and communication component 716.

[0187] Processing component 702 typically controls the overall operation of electronic device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the scene modeling method described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.

[0188] Memory 704 is configured to store various types of data to support the operation of electronic device 700. Examples of this data include instructions for any application or method operating on electronic device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0189] Power component 706 provides power to various components of electronic device 700. Power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 700.

[0190] Multimedia component 708 includes a screen that provides an output interface between electronic device 700 and user. In some embodiments, the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen may be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When electronic device 700 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0191] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when electronic device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.

[0192] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0193] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of electronic device 700. For example, sensor assembly 714 may detect the on / off state of electronic device 700, the relative positioning of components such as the display and keypad of electronic device 700, changes in position of electronic device 700 or a component of electronic device 700, the presence or absence of user contact with electronic device 700, orientation or acceleration / deceleration of electronic device 700, and temperature changes of electronic device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include an optical sensor, such as a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0194] Communication component 716 is configured to facilitate wired or wireless communication between electronic device 700 and other devices. Electronic device 700 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra-Wideband (UWB), Bluetooth, and other technologies.

[0195] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described scene modeling method.

[0196] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of an electronic device 700 to complete the scene modeling method described above. For example, the non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0197] To implement the above embodiments, this disclosure also proposes a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the steps of the scene modeling method provided in this disclosure.

[0198] Alternatively, the computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0199] To implement the above embodiments, this disclosure also proposes a chip including an interface circuit and a processing circuit coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is configured to implement the steps of the scene modeling method provided in this disclosure.

[0200] Figure 8 This is a schematic diagram of the structure of a chip according to an embodiment of this disclosure. See also... Figure 8 The diagram shown is a schematic representation of the structure of chip 800, but is not limited to this.

[0201] Chip 800 includes processing circuitry 801, which is configured to execute any of the above scene modeling methods.

[0202] In some embodiments, the chip 800 further includes one or more interface circuits 802. Optionally, the interface circuit 802 is connected to the memory 803, and the interface circuit 802 can be used to receive signals from the memory 803 or other devices, and the interface circuit 802 can be used to send signals to the memory 803 or other devices. For example, the interface circuit 802 can read instructions stored in the memory 803 and send the instructions to the processing circuit 801.

[0203] In some embodiments, the interface circuit 802 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 801 performs other steps.

[0204] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.

[0205] In some embodiments, chip 800 further includes one or more memories 803 for storing instructions. Optionally, all or part of the memories 803 may be located outside of chip 800.

[0206] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program, which, when executed by a processor, implements the steps of the scene modeling method provided in this disclosure.

[0207] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0208] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A scene modeling method, characterized in that, The method includes: Based on the image acquired at the current time step to be processed, the multi-view scene features of the image corresponding to the current time step are obtained; The historical scene features are updated based on the multi-view scene features to obtain the first target scene features corresponding to the first time step sequence. The historical scene features are obtained based on the images collected in the processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene. Based on the first target scene features, predict and determine the second target scene features corresponding to the second time step sequence, wherein the second time step sequence is a future time step sequence of the first time step sequence; Based on the first target scene features and the second target scene features, the target scene is modeled to obtain the target multi-dimensional scene of the target scene.

2. The method according to claim 1, characterized in that, The acquired images are multi-view images. Based on the images acquired at the current time step to be processed, the multi-view scene features of the image corresponding to the current time step are obtained, including: Extract single-view image features corresponding to each single-view image from the multi-view images acquired at the current time step; Based on the historical scene features and the single-view image features, determine the newly added single-view features in each single-view image; Based on the newly added features of each single view, the multi-view scene features corresponding to the current time step are determined.

3. The method according to claim 2, characterized in that, The step of determining newly added single-view features in each single-view image based on the historical scene features and the single-view image features includes: The feature acquisition pose of a unit scene feature in the historical scene features is obtained, wherein the feature acquisition pose is obtained based on the device pose of the image acquisition device corresponding to the unit scene feature; For any single-view image, from the feature acquisition pose of the unit scene features, determine the approximate acquisition pose corresponding to the single-view acquisition pose of the single-view image, so as to determine the single-view associated features, wherein the single-view associated features are determined based on the unit scene features within the feature range of the approximate acquisition pose, and the feature range is determined based on the visual frustum range of the approximate acquisition pose in the target scene. Based on the single-view correlation features, the single-view image features are filtered to determine the newly added single-view features after filtering.

4. The method according to claim 3, characterized in that, For any single-view image, from the feature acquisition pose of the unit scene features, determine the approximate acquisition pose corresponding to the single-view acquisition pose of each single-view image, so as to determine the single-view associated features of each single-view image. The single-view associated features are determined based on the unit scene features within the feature range of the approximate acquisition pose, and the feature range is determined based on the frustum range of the approximate acquisition pose in the target scene, including: For any single-view image, based on the feature distance between the unit scene features within the feature range and the position of the approximate acquisition pose, the associated unit scene features are determined from the unit scene features within the feature range; Based on the point cloud data of the scene features of the associated unit, a local coordinate system transformation is performed on the scene features of the associated unit to transform them into the corresponding coordinate system of the approximate acquisition pose, so as to determine the single-view associated features.

5. The method according to claim 2, characterized in that, The step of determining the multi-view scene features corresponding to the current time step based on the newly added features from each single viewpoint includes: For any single-view image, the single-view related features of the single-view image are refined to determine the refined single-view related features. The refined single-view related features are then fused with the newly added single-view features of the single-view image to determine the single-view scene features of the fused single-view image. The feature refinement is based on at least one of feature dimensionality reduction, feature selection, feature encoding, feature optimization, and feature abstraction. Based on the features of each single-view scene, the multi-view scene features corresponding to the current time step are determined.

6. The method according to claim 1, characterized in that, The step of updating historical scene features based on the scene features to obtain first target scene features corresponding to the first time step sequence, wherein the historical scene features are obtained based on the acquired images of the processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene, includes: In response to the fact that the current time step is the last time step in the first time step sequence, the historical scene features are updated based on the multi-view scene features corresponding to the current time step to determine the first target scene features.

7. The method according to claim 6, characterized in that, The method further includes: In response to the fact that the current time step is not the last time step in the first time step sequence, the historical scene features are updated based on the multi-view scene features corresponding to the current time step, and the updated historical scene features are determined. Return to continue processing the next acquired image corresponding to the next time step to be processed in the first time step sequence, until the last time step is reached. Based on the multi-view scene features corresponding to the last time step and the updated historical scene features corresponding to the last time step, determine the first target scene features.

8. The method according to claim 1, characterized in that, The step of predicting and determining the second target scene features corresponding to the second time step sequence based on the first target scene features, wherein the second time step sequence is a future time step sequence of the first time step sequence, includes: Based on the first target scene features, the life cycle of the target scene is predicted to obtain the predicted second time step sequence, wherein the second time step sequence is the remaining life cycle of the target scene after the first time sequence, the second time step sequence is a future time step sequence of the first time step sequence, and the next adjacent time step of the current time step; Based on the scene change trend corresponding to the first target scene feature, the change trend of the target scene in the second time step sequence is predicted to determine the second target scene feature.

9. The method according to claim 1, characterized in that, The step of modeling the target scene based on the first target scene features and the second target scene features to obtain the target multi-dimensional scene of the target scene includes: Determine the single-view detection box in the multi-view image of each time step in the first time step sequence and the second time step sequence, so as to obtain the scene features within the target box of each single-view detection box based on the first target scene features and the second target scene features; Gaussian decoding is performed on the scene features within the target boxes of the single-view detection boxes at each time step to determine the single-view model fragments corresponding to the single-view detection boxes, and scene modeling at each time step is constructed based on each single-view model fragment. Based on the scene modeling at each time step, the target multi-dimensional scene modeling is determined, wherein the target multi-dimensional scene modeling is obtained by connecting and rendering the scene models of any two adjacent time steps.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: The first target scene features and the second target scene features are obtained through the trained target scene modeling model to output the target scene multidimensional scene model.

11. The method according to claim 10, characterized in that, Obtaining the target scene modeling model includes: Acquire multi-view images of samples at sample time steps in the sample time step sequence, and determine training samples based on the single-view images of each sample in the multi-view images, the single-view image acquisition pose of each sample single-view image, and the object detection box of each sample single-view image. Using the scene modeling model to be trained, scene features are extracted from the multi-view images of the samples at each time step in sequence to determine the output multi-dimensional scene modeling of the scene model to be trained. Based on the output multidimensional scene modeling and the label multidimensional scene modeling corresponding to the training samples, the scene modeling model to be trained is trained to determine the trained target scene modeling model.

12. The method according to claim 11, characterized in that, The step of training the scene modeling model to be trained based on the output multidimensional scene modeling and the label multidimensional scene modeling corresponding to the training samples, and determining the trained target scene modeling model, includes: Based on the output multidimensional scene modeling and the label multidimensional scene modeling, the photometric reconstruction loss, perception loss, depth loss, optical flow regularization loss, lifetime regularization loss, sky depth loss, sky transparency loss, and object distribution loss corresponding to the scene modeling model are determined. The training loss of the scene modeling model is determined by weighting the photometric reconstruction loss, the perception loss, the depth loss, the optical flow regularization loss, the lifetime regularization loss, the sky depth loss, the sky transparency loss, and the object distribution loss. The parameters of the scene modeling model are iteratively optimized based on the training loss to determine the trained target modeling model.

13. A scene modeling device, characterized in that, The device includes: The extraction module is used to obtain multi-view scene features of the image corresponding to the current time step based on the image acquired at the current time step to be processed; The acquisition module is used to update the historical scene features based on the multi-view scene features to obtain the first target scene features corresponding to the first time step sequence, wherein the historical scene features are obtained based on the images acquired in the processed time steps in the first time step sequence, and the first time step sequence is obtained based on the image acquisition time of the target scene. The prediction module is used to predict and determine the second target scene features corresponding to the second time step sequence based on the first target scene features, wherein the second time step sequence is a future time step sequence of the first time step sequence; The modeling module is used to model the target scene based on the first target scene features and the second target scene features to obtain the target multi-dimensional scene of the target scene.

14. The apparatus according to claim 13, characterized in that, The extraction module is also used for: Extract single-view image features corresponding to each single-view image from the multi-view images acquired at the current time step; Based on the historical scene features and the single-view image features, determine the newly added single-view features in each single-view image; Based on the newly added features of each single view, the multi-view scene features corresponding to the current time step are determined.

15. The apparatus according to claim 13, characterized in that, The prediction module is also used for: Based on the first target scene features, the life cycle of the target scene is predicted to obtain the predicted second time step sequence, wherein the second time step sequence is the remaining life cycle of the target scene after the first time sequence, the second time step sequence is a future time step sequence of the first time step sequence, and the next adjacent time step of the current time step; Based on the scene change trend corresponding to the first target scene feature, the change trend of the target scene in the second time step sequence is predicted to determine the second target scene feature.

16. The apparatus according to claim 13, characterized in that, The modeling module is also used for: Determine the single-view detection box in the multi-view image of each time step in the first time step sequence and the second time step sequence, so as to obtain the scene features within the target box of each single-view detection box based on the first target scene features and the second target scene features; Gaussian decoding is performed on the scene features within the target boxes of the single-view detection boxes at each time step to determine the single-view model fragments corresponding to the single-view detection boxes, and scene modeling at each time step is constructed based on each single-view model fragment. Based on the scene modeling at each time step, the target multi-dimensional scene modeling is determined, wherein the target multi-dimensional scene modeling is obtained by connecting and rendering the scene models of any two adjacent time steps.

17. The apparatus according to any one of claims 13-16, characterized in that, The modeling module is also used for: The first target scene features and the second target scene features are obtained through the trained target scene modeling model to output the target scene multidimensional scene model.

18. A vehicle, characterized in that, The vehicle is used to implement the method as described in any one of claims 1-12.

19. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute instructions to implement the method as described in any one of claims 1-12.

20. A computer-readable storage medium, wherein instructions in the computer-readable storage medium, when executed by a processor of an electronic device, enable the electronic device to perform the method as described in any one of claims 1-12.

21. A chip, characterized in that, It includes one or more interface circuits and one or more processors; the interface circuits are used to receive signals and send the signals to the processors, the signals including computer instructions stored in a memory, which, when executed by the processor, cause the chip to perform the steps of any one of claims 1-12.