Three-dimensional model construction method and device and electronic equipment

By combining feature extraction and visual tracking algorithms to process on-board video data, the problem of low accuracy in three-dimensional environment reconstruction is solved, and the construction of high-precision three-dimensional models is realized, which is suitable for diverse application scenarios and improves user experience.

CN120411342APending Publication Date: 2025-08-01GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264716.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, in the process of reconstructing a three-dimensional environment based on vehicle video data, there are problems such as low accuracy of the three-dimensional model and poor user experience, especially in ordinary mobile MR devices, which consume huge resources and single application scenarios.

Method used

By combining feature extraction and visual tracking algorithms, the on-board video data are processed, the image feature point trajectory is obtained, and the position and attitude are determined by combining interchangeable image file information and multi-sensor fusion algorithm, and a high-precision three-dimensional model is finally built.

Benefits of technology

It improves the authenticity and accuracy of the three-dimensional model, reduces data acquisition costs, expands the scope of application, and is suitable for diverse scenarios such as smart cities, real-life game scenes and remote tourism excursions, improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411342A_ABST
    Figure CN120411342A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional model construction method and apparatus, and an electronic device. The method comprises the steps of obtaining a target vehicle-mounted network video; performing feature extraction on each frame of image in the target vehicle-mounted network video to obtain an image feature corresponding to each frame of image; performing feature point tracking on the target vehicle-mounted network video according to each frame of image and the image feature corresponding to each frame of image, and determining feature point tracks of the same feature point in the target vehicle-mounted network video in different frames of images; according to exchangeable image file information and the feature point track corresponding to the target vehicle-mounted network video, determining a position posture corresponding to each frame of image in the target vehicle-mounted network video, the exchangeable image file information comprising attribute information and shooting data of the target vehicle-mounted network video; and obtaining the three-dimensional model based on the target vehicle-mounted network video and the position attitude corresponding to each frame of image in the target vehicle-mounted network video, thereby improving the precision of the three-dimensional model constructed based on the vehicle-mounted video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more particularly, to a method and apparatus for constructing a three-dimensional model and an electronic device. Background Art

[0002] With the development of automotive technology, the application of setting cameras on vehicles to capture images of the vehicle's environment has become increasingly widespread. In related technologies, a Mixed Reality (MR) technology for three-dimensional environment reconstruction based on in-vehicle video data has been proposed. Based on this, there are challenges in accurately constructing a three-dimensional model during the process of three-dimensional environment reconstruction based on in-vehicle video data. Summary of the Invention

[0003] In view of this, embodiments of this application propose a method and apparatus for constructing a three-dimensional model and an electronic device to address the above problems.

[0004] In a first aspect, an embodiment of this application provides a method for constructing a three-dimensional model, the method including: obtaining a target in-vehicle network video; extracting features from each frame of the target in-vehicle network video to obtain image features corresponding to each frame of the image; tracking feature points of the target in-vehicle network video according to each frame of the image and the image features corresponding to each frame of the image to determine a feature point trajectory of the same feature point in different frames of the target in-vehicle network video; determining a position and attitude corresponding to each frame of the target in-vehicle network video according to the exchangeable image file information corresponding to the target in-vehicle network video and the feature point trajectory, where the exchangeable image file information includes attribute information and shooting data of the target in-vehicle network video; and obtaining a three-dimensional model based on the target in-vehicle network video and the position and attitude corresponding to each frame of the target in-vehicle network video.

[0005] In a second aspect, an embodiment of the present application provides a device for constructing a three-dimensional model. The device includes: a network video acquisition module, a feature extraction module, a feature point trajectory determination module, a position and attitude determination module, and a three-dimensional model acquisition module. Among them, the network video acquisition module is used to acquire a target vehicle-mounted network video; the feature extraction module is used to perform feature extraction on each frame image in the target vehicle-mounted network video to obtain image features corresponding to each frame image; the feature point trajectory determination module is used to perform feature point tracking on the target vehicle-mounted network video according to each frame image and the image features corresponding to each frame image to determine the feature point trajectory of the same feature point in different frame images in the target vehicle-mounted network video; the position and attitude determination module is used to determine the position and attitude corresponding to each frame image in the target vehicle-mounted network video according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectory, where the exchangeable image file information includes the attribute information and shooting data of the target vehicle-mounted network video; the three-dimensional model acquisition module is used to acquire a three-dimensional model based on the target vehicle-mounted network video and the position and attitude corresponding to each frame image in the target vehicle-mounted network video.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor. The memory is coupled to the processor, and the memory stores instructions. When the instructions are executed by the processor, the processor executes the above method.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, and the program code can be called by a processor to execute the above method.

[0008] In the solution of the present application, by acquiring a target vehicle-mounted network video, performing feature extraction on each frame image in the target vehicle-mounted network video to obtain image features corresponding to each frame image, performing feature point tracking on the target vehicle-mounted network video according to each frame image and the image features corresponding to each frame image to determine the feature point trajectory of the same feature point in different frame images in the target vehicle-mounted network video, determining the position and attitude corresponding to each frame image in the target vehicle-mounted network video according to the exchangeable image file information corresponding to the target vehicle-mounted network video, which includes the attribute information and shooting data of the target vehicle-mounted network video, and the feature point trajectory, and acquiring a three-dimensional model based on the target vehicle-mounted network video and the position and attitude corresponding to each frame image in the target vehicle-mounted network video, the vehicle-mounted video data is processed by combining feature extraction and visual tracking algorithms, improving the authenticity of the constructed three-dimensional model, and a three-dimensional model is obtained by combining the results of visual data and shooting parameter fusion, improving the accuracy of the absolute position and true scale of the constructed three-dimensional model and improving the accuracy of the three-dimensional model. Description of the Drawings

[0009] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0010] Figure 1 It shows a schematic flowchart of a method for constructing a three-dimensional model provided by an embodiment of the present application;

[0011] Figure 2 It shows a schematic flowchart of a method for obtaining the position and attitude corresponding to an image provided by an embodiment of the present application;

[0012] Figure 3 It shows a schematic flowchart of a method for constructing a three-dimensional model provided by an embodiment of the present application;

[0013] Figure 4 It shows a schematic flowchart of a method for constructing a three-dimensional model provided by an embodiment of the present application;

[0014] Figure 5 It shows a block diagram of modules of a device for constructing a three-dimensional model provided by an embodiment of the present application;

[0015] Figure 6 It shows a block diagram of an electronic device for executing the method for constructing a three-dimensional model according to an embodiment of the present application;

[0016] Figure 7 It shows a storage unit for storing or carrying program codes for implementing the method for constructing a three-dimensional model according to an embodiment of the present application. Detailed implementation manners

[0017] To enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application.

[0018] To better understand the solutions of the embodiments of the present application, the following first explains the technical terms used in the embodiments of the present application.

[0019] The Exchangeable Image File Format (EXIF) is a file format customized specifically for digital camera photos, used to record the attribute information and shooting data of digital photos. EXIF information mainly includes the following categories: camera setting information (such as ISO, white balance, saturation, sharpness, etc.), shooting records (such as shutter speed, aperture size, focal length, shooting time, etc.), image information (such as manufacturer, resolution, etc.), thumbnail (such as width and height of the thumbnail, etc.), GPS information (such as longitude, latitude, altitude, etc. at the time of shooting).

[0020] Optical flow method is an important concept in the field of computer vision for detecting the movement of objects in the visual field; it describes the movement of observed objects, surfaces or edges caused by the movement relative to the observer.

[0021] The six degrees of freedom (6DOF) position and orientation of an image refer to the ability to determine the position and direction of an object in three-dimensional space. 6DOF includes three translational degrees of freedom and three rotational degrees of freedom; the three translational degrees of freedom describe the position of the object in three-dimensional space, usually represented by a three-dimensional vector (X, Y, Z) to represent the coordinates of the object in the world coordinate system or camera coordinate system; the three rotational degrees of freedom describe the orientation of the object in three-dimensional space, usually represented by a rotation matrix or quaternion to represent the rotation of the object relative to a certain reference coordinate system.

[0022] The COCO dataset (Common Objects in Context) is a large-scale computer vision dataset that is widely used in tasks such as object detection, instance segmentation, and keypoint detection, and is a dataset with the highest generality and the widest scope.

[0023] Gaussian Splatting is a three-dimensional data format for real-time radiance field rendering, mainly used to generate high-quality three-dimensional images and videos. It represents the objects in a three-dimensional scene as a collection of Gaussian kernels, and each Gaussian kernel represents a part of the object surface, thus achieving a high-fidelity rendering effect.

[0024] MR technology, Mixed Reality, is the further development of virtual reality technology (VR) and augmented reality technology (AR). It presents virtual scene information in the real scene, constructs an interactive feedback information loop, and enhances the realism of the user experience.

[0025] The following elaborates in detail on the implementation details of the technical solution of the embodiments of this application:

[0026] In the related art, for the development of an autonomous driving system for vehicles, a solution has been proposed to use the Neural Radiance Fields (NERF) algorithm to map and render the collected vehicle data, achieving the integration of vehicle data. However, the collection of vehicle data requires a dedicated collection vehicle and other high-cost hardware to provide the position and attitude of the images, which will result in a high cost for three-dimensional environment reconstruction based on vehicle data and cannot be easily extended to a larger range. In addition, this solution is targeted at the autonomous driving scenario of vehicles, suffering from the problem of a single application scenario. At the same time, the algorithm framework of NERF used therein cannot be used on ordinary mobile MR devices, resulting in huge resource consumption for three-dimensional environment construction. Based on this, in the process of applying the three-dimensional environment reconstructed based on vehicle data to MR technology, there are problems of low accuracy of the three-dimensional environment and low user experience of vehicle-mounted MR. Therefore, in the related art, there is a challenge in improving the accuracy of the three-dimensional model in the process of three-dimensional environment reconstruction based on vehicle video data.

[0027] In view of the above problems, through long-term research, the inventors have proposed a method, an apparatus, and an electronic device for constructing a three-dimensional model provided in the embodiments of the present application. By combining feature extraction and visual tracking algorithms to process vehicle video data, the authenticity of the constructed three-dimensional model is improved, and a three-dimensional model is obtained by combining the results of visual data and shooting parameter fusion, improving the accuracy of the absolute position and true scale of the constructed three-dimensional model and enhancing the precision of the three-dimensional model. Among them, the specific method for constructing the three-dimensional model will be described in detail in the subsequent embodiments.

[0028] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of a method for constructing a three-dimensional model provided in an embodiment of the present application. In a specific embodiment, the method for constructing the three-dimensional model can be applied to a three-dimensional model construction apparatus 200 as shown in Figure 5 and an electronic device 100 configured with the three-dimensional model construction apparatus 200 ( Figure 6 ). Hereinafter, taking the electronic device as an example, the specific process of this embodiment will be described. Of course, it can be understood that the electronic device to which this embodiment is applied may include devices such as a desktop computer, a laptop computer, a vehicle-mounted terminal, and a vehicle-mounted large screen, which are not limited herein. Next, the process shown in Figure 1 will be elaborated in detail. The method for constructing the three-dimensional model may specifically include the following steps:

[0029] Step S110: Obtain a target vehicle network video.

[0030] In some embodiments, the electronic device may obtain the target vehicle network video from an associated cloud or the electronic device itself. Among them, the target vehicle network video may be a video collected by cameras installed at different positions of the vehicle. Optionally, the electronic device may collect vehicle video data through a search engine, a video sharing platform, or a dedicated dataset library, and obtain the target vehicle network video from the collected vehicle video data; thus, three-dimensional reconstruction can be performed using the network video data, and a wider and more diverse set of vehicle scene data can be obtained without relying on specific devices or sensors, reducing the cost and difficulty of data collection.

[0031] Optionally, a target vehicle network video acquisition script may be pre-stored in the electronic device, and the electronic device may obtain the target vehicle network video based on this script; optionally, the electronic device may also obtain the target vehicle network video from an associated cloud or the electronic device itself through a web crawler. Among them, the number of target vehicle network videos obtained by the electronic device may be one or more, which is not limited herein.

[0032] Step S120: Extract features from each frame image in the target vehicle network video to obtain the image features corresponding to each frame image.

[0033] In some embodiments, after the electronic device obtains the target vehicle network video, it may extract features from each frame image in the target vehicle network video to obtain the image features corresponding to each frame image. Among them, a deep learning network for extracting features from vehicle videos may be pre-set in the electronic device. Among them, this deep learning network may be used to extract general features, abstract features, low-level features, high-level features, etc. of each frame image in the video; correspondingly, the electronic device may determine the features extracted by the deep learning network as the image features. Among them, the image features of each frame image may include basic visual elements such as corners and edges of the image, and may also include elements such as the texture of the image and the objects in the image.

[0034] Among them, this deep learning network may be composed of a convolutional neural network or other deep learning architectures (such as autoencoders or generative adversarial networks, etc.). Optionally, this deep learning network may be obtained by the electronic device from an associated cloud or the electronic device itself, or may be obtained by the electronic device based on third-party experimental data. Exemplarily, the electronic device may obtain the COCO dataset from an associated cloud or the electronic device itself, and may perform network training based on this COCO dataset to obtain a deep learning network for extracting image features from vehicle videos.

[0035] Among them, after the electronic device obtains the target vehicle network video, it may input this target vehicle network video into this deep learning network, and obtain the image features corresponding to each frame image in the target vehicle network video output by this deep learning network.

[0036] Step S130: According to the respective frame images and the image features corresponding to the respective frame images, perform feature point tracking on the target vehicle-mounted network video to determine the feature point trajectories of the same feature point in different frame images of the target vehicle-mounted network video.

[0037] In some embodiments, after the electronic device obtains the image features corresponding to the respective frame images in the target vehicle-mounted network video, it may perform feature point tracking on the target vehicle-mounted network video according to the respective frame images and the image features corresponding to the respective frame images, and determine the feature point trajectories of the same feature point in different frame images of the target vehicle-mounted network video. Among them, the electronic device may determine the feature points (such as corner points, edge points, spots, etc.) in the respective frame images based on the image features corresponding to the respective frame images, and may obtain the motion information of the feature points in the respective frame images in the target vehicle-mounted network video based on a visual tracking algorithm, and may process the motion information of the same feature point in different frame images of the target vehicle-mounted network video in chronological order to obtain the feature point trajectories of the same feature point in different frame images of the target vehicle-mounted network video; thus, by combining deep learning and a visual tracking algorithm, high-precision three-dimensional reconstruction of vehicle-mounted video data is achieved, a more realistic, accurate, and high-quality three-dimensional model is generated, and the effect of applying the three-dimensional model is improved.

[0038] Among them, the motion information of the feature point in different frame images may include the speed of the pixel point represented by the feature point. For example, the motion direction and speed magnitude of the pixel point on the image plane may be represented by a two-dimensional vector (u, v), where u represents the speed in the horizontal direction and v represents the speed in the vertical direction.

[0039] Step S140: According to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectories, determine the position and attitude corresponding to each frame image in the target vehicle-mounted network video, where the exchangeable image file information includes the attribute information and shooting data of the target vehicle-mounted network video.

[0040] In some embodiments, considering that in the process of obtaining a standard video file, the EXIF information corresponding to the standard video file may also be obtained; correspondingly, in the process of the electronic device obtaining the target vehicle-mounted network video, the EXIF information corresponding to the target vehicle-mounted network video may also be obtained. Among them, the EXIF information corresponding to the target vehicle-mounted network video may include the attribute information and shooting data of the target vehicle-mounted network video, such as shooting time, shooting focal length, inertial measurement unit IMU information, global positioning system GPS information, etc.

[0041] After the electronic device obtains the feature point trajectories of the same feature point in different frame images of the target vehicle-mounted network video, it can determine the position and orientation corresponding to each frame image in the target vehicle-mounted network video according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectories.

[0042] Optionally, a multi-sensor fusion algorithm can be preset in the electronic device; among them, the electronic device can fuse the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectories based on the multi-sensor fusion algorithm to obtain the position and orientation corresponding to each frame image in the target vehicle-mounted network video. It can be understood that the GPS information contains the absolute position. Based on this, the multi-sensor fusion algorithm can estimate the absolute position and true scale information of the subsequent constructed three-dimensional model, thereby improving the accuracy and precision of the three-dimensional reconstruction.

[0043] Among them, the position and orientation of the image can include the spatial position and direction of the object in the image. For example, it can include the two-dimensional pose information of the object in the image (such as the 4-degree-of-freedom position and orientation of the image, etc.), the three-dimensional pose information of the object in the image (such as the 6-degree-of-freedom position and orientation of the image, etc.), etc. In this embodiment, considering ensuring the accuracy and authenticity of the model in the three-dimensional space, the position and orientation of the image can be the 6-degree-of-freedom position and orientation of the image.

[0044] Step S150: Obtain a three-dimensional model based on the target vehicle-mounted network video and the position and orientation corresponding to each frame image in the target vehicle-mounted network video.

[0045] In some embodiments, after the electronic device obtains the position and orientation corresponding to each frame image in the target vehicle-mounted network video, it can obtain a three-dimensional model based on the target vehicle-mounted network video and the position and orientation corresponding to each frame image in the target vehicle-mounted network video. Among them, the electronic device can fuse each frame image in the target vehicle-mounted network video and the position and orientation corresponding to each frame image to obtain a map and the position and orientation corresponding to the map. Among them, the electronic device can perform real-time rendering and mapping based on the deep learning neural network rendering algorithm, the map, and the position and orientation corresponding to the map to obtain a three-dimensional model. The deep learning neural network rendering algorithm can be preset in the electronic device, or can be obtained from an associated cloud or electronic device, which is not limited herein.

[0046] In some embodiments, considering the situation of overlapping areas in the obtained map, in order to avoid mapping errors, a loop detection algorithm and a global optimization algorithm can be preset in the electronic device. Among them, the electronic device can identify the closed-loop information in the map based on the loop detection algorithm, and can optimize the topological structure and geometric shape in the map based on the closed-loop information and the global optimization algorithm, so as to improve the accuracy of the obtained map and the position and attitude corresponding to the map. Accordingly, the electronic device can perform real-time rendering mapping based on the deep learning neural network rendering algorithm, the optimized map, and the position and attitude corresponding to the optimized map to obtain a three-dimensional model, so that the user can freely move and three-dimensionally view in the entire map with six degrees of freedom, providing the user with a more immersive experience and also improving the visualization effect and interactivity of the map.

[0047] The method for constructing a three-dimensional model provided by an embodiment of the present application obtains a target vehicle-mounted network video, extracts features from each frame of image in the target vehicle-mounted network video to obtain the image features corresponding to each frame of image, and tracks feature points of the target vehicle-mounted network video according to each frame of image and the image features corresponding to each frame of image to determine the feature point trajectories of the same feature point in different frames of image in the target vehicle-mounted network video, and determines the position and attitude corresponding to each frame of image in the target vehicle-mounted network video according to the exchangeable image file information including the attribute information of the target vehicle-mounted network video and the shooting data and the feature point trajectories corresponding to the target vehicle-mounted network video, and obtains a three-dimensional model based on the target vehicle-mounted network video and the position and attitude corresponding to each frame of image in the target vehicle-mounted network video, so as to process the vehicle-mounted video data by combining feature extraction and visual tracking algorithms, improve the authenticity of the constructed three-dimensional model, and obtain a three-dimensional model by combining the results of visual data and shooting parameter fusion, improve the accuracy of the absolute position and true scale of the constructed three-dimensional model, and improve the accuracy of the three-dimensional model.

[0048] In some embodiments, during the process of the electronic device obtaining the target vehicle-mounted network video, at least one vehicle-mounted network video can be obtained, and the vehicle-mounted network video that meets the preset conditions in the at least one vehicle-mounted network video can be determined as the target vehicle-mounted network video.

[0049] Among them, a preset condition can be pre-set in the electronic device. The preset condition can be set independently by the user or obtained through third-party experimental data, which is not limited herein. Among them, the preset condition can include at least one of the following: the video is a video of continuous frames, the resolution of the video is greater than a preset resolution, the video includes a moving object, and the duration of the video is greater than a preset duration. The video being a video of continuous frames can be understood as the video being a video of continuous frames without including editing. The preset resolution can be pre-set in the electronic device, can be set independently by the user, or can be obtained through third-party experimental data. Exemplarily, the preset resolution is set by the user to 2K, 3K, etc. The video including a moving object can be understood as that the video must contain a moving object, that is, a video in which all objects in the video are stationary does not meet the preset condition. The preset duration can be pre-set in the electronic device, can be set independently by the user, or can be obtained through third-party experimental data. Exemplarily, the preset duration is set by the user to 5 min, 6 min, etc.

[0050] Exemplarily, the electronic device can obtain a list of at least one in-vehicle network video by crawling from each video sharing platform in the manner of keyword search, and can screen the in-vehicle network videos that meet the preset conditions from the list of the at least one in-vehicle network video as the target in-vehicle network videos, so as to screen videos with higher quality and stronger diversity for 3D model reconstruction, improving the accuracy and precision of 3D reconstruction.

[0051] It can be understood that the amount of in-vehicle video data on the network is very large, the cost of data acquisition is very low, and it can be easily extended to a very large range. Based on this, in this embodiment, the electronic device uses the existing in-vehicle video data source on the network for 3D reconstruction, breaking through the bottleneck of traditional in-vehicle scenarios that require special equipment and high-cost acquisition; through automatic collection and screening of network videos, a more extensive and flexible data acquisition method is realized, which is applicable to the extended application of large-scale scenarios, such as, including but not limited to: the construction of smart cities, the construction and use of real-world game scenarios, the scenarios of remote cultural tourism tours, etc., improving the application scope of in-vehicle MR.

[0052] In some embodiments, considering ensuring the accuracy and authenticity of the constructed three-dimensional model in the three-dimensional space, the position and attitude corresponding to each frame image in the target vehicle-mounted network video may include the six-degree-of-freedom position and attitude corresponding to each frame image. Among them, when the electronic device determines the position and attitude corresponding to each frame image in the target vehicle-mounted network video according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectory of the same feature point in different frame images of the target vehicle-mounted network video, if it is determined that the shooting data includes sensor information, then the six-degree-of-freedom position and attitude corresponding to each frame image in the target vehicle-mounted network video can be determined according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectory. Exemplarily, if the electronic device determines that the shooting data includes at least one of the information of a vision sensor, an inertial measurement unit, and a global positioning system, etc., then the six-degree-of-freedom position and attitude corresponding to each frame image in the target vehicle-mounted network video can be determined according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectory.

[0053] Among them, the sensor information may include the information of a vision sensor (such as image information, etc.), the information of an inertial measurement unit (such as IMU data, etc.), and the information of a global positioning system (such as GPS position information, etc.).

[0054] Among them, considering that the data acquisition times of different sensors are different, in this embodiment, when the electronic device determines the position and attitude corresponding to each frame image in the target vehicle-mounted network video according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectory of the same feature point in different frame images of the target vehicle-mounted network video, if it is determined that the shooting data includes at least one of the information of a vision sensor, an inertial measurement unit, and a global positioning system, etc., then the time of collecting information by different sensors can be synchronized and calibrated. Exemplarily, if the electronic device determines that the EXIF information corresponding to the target vehicle-mounted network video contains timestamp information, then it may not be necessary to calibrate the time additionally. If it is determined that the EXIF information does not contain timestamp information, then the time of collecting information by different sensors can be synchronized and calibrated based on the timestamp calibration algorithm.

[0055] Among them, considering that each sensor has its own accuracy range (for example, GPS may have an accuracy of meters or centimeters), and each sensor also has noise. Based on this, in this embodiment, in order to improve the accuracy and precision of 3D reconstruction, the electronic device may be pre-set with the noise and uncertainty parameters of the sensors, and these parameters can be brought into the multi-sensor fusion algorithm to effectively eliminate the uncertainties and noises of different sensors. Among them, the electronic device can, based on the multi-sensor fusion algorithm, perform fusion processing on the exchangeable image file information corresponding to the target vehicle network video and the feature point trajectories of the same feature point in different frame images in the target vehicle network video to obtain the absolute position and true scale information of the 3D model. Among them, the electronic device can use a non-linear optimization method to jointly estimate the information from different sensors (for example, fuse the data collected by vision, IMU, and GPS sensors) to obtain the absolute position and true scale information of the 3D model, thereby fusing sensor data such as visual data, IMU data, and GPS information, improving the accuracy of the absolute position and true scale information of the 3D model, making the reconstruction result more reliable, and also improving the accuracy of positioning using the 3D model.

[0056] Exemplarily, please refer to Figure 2 , which shows a schematic flowchart of obtaining the position and attitude corresponding to an image provided by an embodiment of the present application. Among them, the electronic device can obtain at least one vehicle network video from an associated cloud or electronic device, and can screen and download the vehicle network video that meets the preset conditions from the at least one vehicle network video and determine it as the target vehicle network video.

[0057] Among them, after the electronic device obtains the target vehicle network video, it can perform feature extraction on each frame image in the target vehicle network video based on deep learning to obtain the image features corresponding to each frame image. Among them, the electronic device can also, based on a visual tracking algorithm, according to each frame image and the image features corresponding to each frame image, perform feature point tracking on the target vehicle network video to determine the feature point trajectories of the same feature point in different frame images in the target vehicle network video, thereby overcoming the complexity and precision challenges of map building based on network videos in a way that combines deep learning and visual tracking algorithms, obtaining high-precision modeling results in rich and variable vehicle video scenarios, greatly improving the accuracy and adaptability of 3D reconstruction, effectively overcoming the difficulties of MR applications in visual map building, providing rich content resources for MR applications, and also effectively overcoming the content gap of composing maps for MR applications through network videos.

[0058] Among them, if the electronic device determines that the shooting data in the exchangeable image file information corresponding to the target vehicle-mounted network video includes at least one of the information of the vision sensor, the inertial measurement unit, and the global positioning system, the six-degree-of-freedom position and attitude corresponding to each frame of the image in the target vehicle-mounted network video can be determined according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectory; thereby, sensor data such as vision data, IMU, and GPS are fused, improving the absolute position accuracy and true scale of the three-dimensional model, making the model suitable for high-precision positioning requirements, and this multi-modal fusion not only improves the accuracy of the three-dimensional model but also provides flexible data formats (such as three-dimensional point clouds, three-dimensional models, and neural rendering maps, etc.) for different application scenarios, providing multiple visualization methods for the three-dimensional model and improving the user experience.

[0059] Please refer to Figure 3 , Figure 3 which shows a schematic flow chart of a method for constructing a three-dimensional model provided by an embodiment of the present application. This method is applied to the above-mentioned electronic device, and the following will elaborate in detail on the Figure 3 flow shown. The method for constructing the three-dimensional model may specifically include the following steps:

[0060] Step S210: Obtain a target vehicle-mounted network video, and the number of the target vehicle-mounted network videos is multiple.

[0061] In some embodiments, the electronic device can obtain it from an associated cloud or the electronic device; among them, the number of the target vehicle-mounted network videos can be multiple, and correspondingly, the electronic device can perform three-dimensional mapping based on different videos to improve the integrity of the three-dimensional model.

[0062] Step S220: Extract features from each frame of the image in the target vehicle-mounted network video to obtain the image features corresponding to each frame of the image.

[0063] For the specific description of step S220, please refer to the description of step S120 in the previous text, and details will not be repeated here.

[0064] Step S230: Determine the target feature points corresponding to the target vehicle-mounted network video according to the image features corresponding to each frame of the image, where the target feature points include the corner points in the target vehicle-mounted network video.

[0065] In some embodiments, after the electronic device obtains the image features corresponding to each frame of the target vehicle-mounted network video, it may determine the target feature points corresponding to the target vehicle-mounted network video according to the image features corresponding to each frame of the image. Among them, the image features corresponding to each frame of the image may be general features corresponding to each frame of the image. Among them, the electronic device may determine the key feature points (such as corner points, edge points, spots, etc. of the image in the target vehicle-mounted network video) in the target vehicle-mounted network video according to the image features corresponding to each frame of the image, and may determine the key feature points as the target feature points corresponding to the target vehicle-mounted network video.

[0066] As an implementable manner, the electronic device may determine the corner points in the target vehicle-mounted network video according to the image features corresponding to each frame of the image, and determine the corner points as the target feature points corresponding to the target vehicle-mounted network video.

[0067] Step S240: In the target vehicle-mounted network video, track the target feature points to obtain the motion information of the target feature points in different frame images of the target vehicle-mounted network video.

[0068] In some embodiments, after the electronic device determines the target feature points corresponding to the target vehicle-mounted network video, it may track the target feature points in the target vehicle-mounted network video to obtain the motion information of the target feature points in different frame images of the target vehicle-mounted network video.

[0069] Among them, a deep learning model (such as a convolutional neural network, a generative adversarial network, etc.) may be pre-set in the electronic device, and the deep learning model may be used to track the target feature points in the video to obtain the motion information of the target feature points in different frame images of the video. Exemplarily, the electronic device may input the target vehicle-mounted network video and the features corresponding to each frame of the target vehicle-mounted network video into the deep learning model, and the deep learning model may determine the target feature points corresponding to the target vehicle-mounted network video and obtain the motion information of the target feature points in consecutive frame images of the target vehicle-mounted network video. [[ID=

[13] ]]

[0070] In some embodiments, the electronic device may track the target feature points in different frame images of the target vehicle-mounted network video based on the optical flow method and a tracker to obtain the motion information of the target feature points in consecutive frame images (such as the speed in the horizontal direction, the speed in the vertical direction, etc.).

[0071] Step S250: Determine the feature point trajectory of the target feature points according to the motion information.

[0072] In some embodiments, after the electronic device obtains the motion information of the target feature points corresponding to the target in-vehicle network video in consecutive frame images, it can determine the feature point trajectory of the target feature points according to the motion information. Among them, the electronic device can obtain the feature point trajectory of the target feature points based on the time sequence and the motion information of the target feature points in consecutive frame images, so as to perform three-dimensional reconstruction according to the feature point trajectory, improving the authenticity and accuracy of three-dimensional mapping. Among them, the feature point trajectory of the target feature points can also be understood as the tracking trajectory of the target feature points.

[0073] Step S260: Based on the loop detection algorithm, multiple target in-vehicle network videos, and the position and attitude corresponding to each frame image in each target in-vehicle network video, perform cross-positioning on the multiple target in-vehicle network videos to obtain the loop constraint information of the multiple target in-vehicle network videos.

[0074] In some embodiments, the number of target in-vehicle network videos obtained by the electronic device is multiple. During the process of three-dimensional reconstruction based on multiple target in-vehicle network videos, considering that there are overlapping regions between different videos, during the process of three-dimensional reconstruction based on multiple target in-vehicle network videos, the overlap and conflict between the maps will cause errors in the three-dimensional mapping results. To reduce the mapping errors in the overlapping regions between the maps of different videos and improve the stability and accuracy of the boundaries between the maps, in this embodiment, the electronic device may be pre-set with a loop detection algorithm. The electronic device can perform cross-positioning on multiple target in-vehicle network videos based on the loop detection algorithm, multiple target in-vehicle network videos, and the position and attitude corresponding to each frame image in each target in-vehicle network video, to obtain the loop constraint information of the multiple target in-vehicle network videos.

[0075] Among them, during the process of the electronic device performing loop detection, it can process two pairs of target in-vehicle network videos respectively. For example, select the detected feature points in the image from target in-vehicle network video A and locate the detected feature points in target in-vehicle network video B. If the positioning is successful, the image corresponding to the detected feature points and the position and attitude corresponding to the image can be obtained and used as the loop constraint information of target in-vehicle network video A and target in-vehicle network video B. Correspondingly, the electronic device can obtain the loop constraint information of multiple target in-vehicle network videos based on the loop detection method.

[0076] Step S270: Based on the loop constraint information, multiple target in-vehicle network videos, and the position and attitude corresponding to each frame image in the multiple target in-vehicle network videos, generate a target map and the position and attitude corresponding to the target map.

[0077] After obtaining the loop closure constraint information of multiple target in-vehicle network videos, an electronic device can generate a target map and the position and orientation corresponding to the target map based on the loop closure constraint information, the multiple target in-vehicle network videos, and the position and orientation corresponding to each frame image in the multiple target in-vehicle network videos. Among them, the electronic device can fuse the multiple target in-vehicle network videos and the position and orientation corresponding to each frame image in the multiple target in-vehicle network videos to connect and integrate the 3D reconstruction results from different in-vehicle videos and construct a consistent map. Among them, the electronic device can optimize the overlaps and conflicts in the fused map based on the loop closure constraint information to improve the consistency of the map. The overlaps in the map can be understood as the overlaps of the coverage areas of the maps corresponding to different videos, and the conflicts in the map can be understood as the loop closure constraint information in the mis-matched multiple target in-vehicle network videos.

[0078] In some embodiments, the electronic device can fuse the multiple target in-vehicle network videos and the position and orientation corresponding to each frame image in the multiple target in-vehicle network videos based on the loop closure constraint information to obtain an initial map and the position and orientation corresponding to the initial map, and can perform topological structure optimization and geometric shape optimization on the initial map and the position and orientation corresponding to the initial map based on a non-linear optimization algorithm to obtain the target map and the position and orientation corresponding to the target map, thereby ensuring the integrity and consistency of the map through loop detection and non-linear optimization and providing accurate and complete 3D environment support for users in immersive browsing and navigation applications. Among them, the topological structure optimization and geometric shape optimization can include transforming the position and orientation of the map, transforming the coordinate system corresponding to the map, etc., which are not limited herein.

[0079] Among them, it can be understood that after the electronic device identifies a closed loop in the map based on the loop detection algorithm, it can optimize the topological structure and geometric shape of the map through a non-linear optimization algorithm so that all the data in different target in-vehicle network videos can be fused together, eliminating the inconsistencies between the acquisitions of different target in-vehicle network videos and improving the accuracy of the target map and the position and orientation corresponding to the target map. Among them, through global optimization and map fusion technology, the mapping results of different target in-vehicle network videos can be connected to generate a consistent map, improving the integrity and consistency of the map and providing a better browsing and navigation experience for users.

[0080] Step S280: Render and map the target map based on the target map and the position and orientation corresponding to the target map to obtain the 3D model, where the 3D model includes at least one of the point cloud reconstruction content, model reconstruction content, and neural network rendering reconstruction content corresponding to the target map.

[0081] After obtaining the target map and the position and orientation corresponding to the target map, the electronic device can perform rendering and mapping on the target map based on the target map and the position and orientation corresponding to the target map to obtain a three-dimensional model. Among them, a rendering and mapping network (such as a convolutional neural network, a generative adversarial network, etc.) can be pre-set in the electronic device. The electronic device can input the target map and the position and orientation corresponding to the target map into the rendering and mapping network, so that the rendering and mapping network can identify the representation and visual features of the target map and perform rendering and mapping to output a three-dimensional model.

[0082] Among them, considering balancing the speed of rendering the map and the quality of rendering the map, and comprehensively considering the processing and storage capabilities of the electronic device for large-scale data, the three-dimensional model can include at least one of the point cloud reconstruction content, model reconstruction content, and neural network rendering reconstruction content corresponding to the target map. Among them, the point cloud reconstruction content can represent a set of multiple point clouds, and each point cloud can contain information such as the three-dimensional position information of the point cloud and the color of the point cloud. Among them, the model reconstruction content can represent a mesh model and can contain triangular face information and texture information of the three-dimensional model. Among them, the neural network rendering reconstruction content can represent a set of multiple point clouds (which can represent advanced point cloud reconstruction), and each point cloud can contain information such as the three-dimensional position information of the point cloud, covariance (such as a 3×3 matrix, etc.), opacity, color, optical characteristics (such as spherical harmonic function parameters, etc.).

[0083] Exemplarily, please refer to Figure 4 , which shows a schematic flow chart of a method for constructing a three-dimensional model provided by an embodiment of the present application. Among them, the electronic device can obtain multiple target vehicle-mounted network videos, and can perform feature point tracking on each target vehicle-mounted network video by combining deep learning and visual tracking algorithms to obtain the feature point trajectories corresponding to each target vehicle-mounted network video, and can fuse the exchangeable image file information corresponding to each target vehicle-mounted network video and the feature point trajectories corresponding to each target vehicle-mounted network video to obtain the position and orientation corresponding to each frame of image in each target vehicle-mounted network video, and can perform cross-positioning and global optimization based on the position and orientation corresponding to each frame of image in each target vehicle-mounted network video and multiple target vehicle-mounted network videos to obtain a target map with multi-map fusion and the position and orientation corresponding to the target map, and can input the target map and the position and orientation corresponding to the target map into the rendering and mapping network for rendering and mapping to obtain a three-dimensional model including point cloud reconstruction content, model reconstruction content, and neural network rendering reconstruction content output by the rendering and mapping network. Thus, the deep learning neural network is used to perform real-time rendering and mapping on the map, enabling users to freely move and perform three-dimensional browsing in the entire map, providing users with a more immersive experience, and also improving the visualization effect and interactivity of the map.

[0084] In some embodiments, a deep learning model may be pre - set in an electronic device. The deep learning model may include a Gaussian Splating algorithm framework. Among them, the electronic device may input a target vehicle - borne network video into the deep learning model. The deep learning model extracts features of each frame image in the target vehicle - borne network video to obtain image features corresponding to each frame image, and may track feature points of the target vehicle - borne network video according to each frame image and the image features corresponding to each frame image, determine the feature point trajectory of the same feature point in different frame images in the target vehicle - borne network video, and may determine the position and pose corresponding to each frame image in the target vehicle - borne network video according to the exchangeable image file information corresponding to the target vehicle - borne network video and the feature point trajectory, and may obtain a 3D model and output it based on the target vehicle - borne network video and the position and pose corresponding to each frame image in the target vehicle - borne network video. Among them, when the number of target vehicle - borne network videos input into the deep learning model is multiple, the deep learning model may perform cross - positioning on the multiple target vehicle - borne network videos based on a loop - closing detection algorithm, the multiple target vehicle - borne network videos, and the position and pose corresponding to each frame image in each target vehicle - borne network video to obtain loop - closing constraint information of the multiple target vehicle - borne network videos; and may generate a target map and the position and pose corresponding to the target map based on the loop - closing constraint information, the multiple target vehicle - borne network videos, and the position and pose corresponding to each frame image in the multiple target vehicle - borne network videos; and may perform rendering and mapping on the target map based on the target map and the position and pose corresponding to the target map to obtain a 3D model including point cloud reconstruction content, model reconstruction content, and neural network rendering reconstruction content, thereby introducing the Gaussian Splating algorithm framework, avoiding the problems of poor real - time performance and high resource requirements when the MR device uses the traditional neural network rendering framework, and supporting efficient real - time rendering on low - computing - power MR devices, enabling users to perform 6 - degree - of - freedom real - time mobile 3D browsing of the map on ordinary mobile MR devices, and optimizing the user experience in the MR scenario.

[0085] The method for constructing a 3D model provided by an embodiment of the present application, compared with Figure 1The construction method of the three-dimensional model shown. In this embodiment, the number of target vehicle-mounted network videos is multiple. This embodiment can also perform cross-positioning on multiple target vehicle-mounted network videos based on the loop detection algorithm, multiple target vehicle-mounted network videos, and the position and attitude corresponding to each frame image in each target vehicle-mounted network video to obtain the loop constraint information of multiple target vehicle-mounted network videos; based on the loop constraint information, multiple target vehicle-mounted network videos, and the position and attitude corresponding to each frame image in multiple target vehicle-mounted network videos, generate a target map and the position and attitude corresponding to the target map; based on the target map and the position and attitude corresponding to the target map, perform rendering and mapping on the target map to obtain a three-dimensional model, where the three-dimensional model includes at least one of the point cloud reconstruction content, model reconstruction content, and neural network rendering reconstruction content corresponding to the target map, so as to perform three-dimensional mapping based on multiple target vehicle-mounted network videos, perform cross-positioning on multiple target vehicle-mounted network videos, reduce the error of the map, improve the consistency and integrity of the generated three-dimensional model, and improve the user experience. This embodiment can also determine the target feature points corresponding to the target vehicle-mounted network video according to the image features corresponding to each frame image, where the target feature points include corner points in the target vehicle-mounted network video; in the target vehicle-mounted network video, track the target feature points to obtain the motion information of the target feature points in different frame images of the target vehicle-mounted network video; determine the feature point trajectory of the target feature points according to the motion information, so as to perform three-dimensional reconstruction by combining the results of video feature extraction and visual tracking, improve the authenticity and accuracy of the three-dimensional model, and improve the precision of the three-dimensional model.

[0086] Please refer to Figure 5 , Figure 5 which shows a module block diagram of a three-dimensional model construction device provided by an embodiment of the present application. The three-dimensional model construction device 200 is applied to the above-mentioned electronic device. The following will elaborate in detail on Figure 5 the process shown. The three-dimensional model construction device 200 includes: a network video acquisition module 210, a feature extraction module 220, a feature point trajectory determination module 230, a position and attitude determination module 240, and a three-dimensional model acquisition module 250, where:

[0087] The network video acquisition module 210 is used to acquire target vehicle-mounted network videos.

[0088] The feature extraction module 220 is used to extract features from each frame image in the target vehicle-mounted network video to obtain the image features corresponding to each frame image.

[0089] The feature point trajectory determination module 230 is used to perform feature point tracking on the target vehicle-mounted network video according to each frame image and the image features corresponding to each frame image, and determine the feature point trajectory of the same feature point in different frame images of the target vehicle-mounted network video.

[0090] A position and attitude determination module 230, configured to determine the position and attitude corresponding to each frame of image in the target vehicle-mounted network video according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectory, wherein the exchangeable image file information includes the attribute information and shooting data of the target vehicle-mounted network video.

[0091] A 3D model acquisition module 240, configured to acquire a 3D model based on the target vehicle-mounted network video and the position and attitude corresponding to each frame of image in the target vehicle-mounted network video.

[0092] Further, the number of the target vehicle-mounted network videos is multiple, and the 3D model acquisition module 240 may include: a loop detection unit, a target map generation unit, and a rendering and mapping unit, wherein:

[0093] The loop detection unit is configured to perform cross-positioning on multiple target vehicle-mounted network videos based on a loop detection algorithm, the multiple target vehicle-mounted network videos, and the position and attitude corresponding to each frame of image in each target vehicle-mounted network video, so as to obtain loop constraint information of the multiple target vehicle-mounted network videos.

[0094] The target map generation unit is configured to generate a target map and the position and attitude corresponding to the target map based on the loop constraint information, the multiple target vehicle-mounted network videos, and the position and attitude corresponding to each frame of image in the multiple target vehicle-mounted network videos.

[0095] The rendering and mapping unit is configured to perform rendering and mapping on the target map based on the target map and the position and attitude corresponding to the target map, so as to obtain the 3D model, wherein the 3D model includes at least one of point cloud reconstruction content, model reconstruction content, and neural network rendering and reconstruction content corresponding to the target map.

[0096] Further, the target map generation unit may include: a video fusion unit and a global optimization unit, wherein:

[0097] The video fusion unit is configured to fuse the multiple target vehicle-mounted network videos and the position and attitude corresponding to each frame of image in the multiple target vehicle-mounted network videos based on the loop constraint information, so as to obtain an initial map and the position and attitude corresponding to the initial map.

[0098] The global optimization unit is configured to perform topological structure optimization and geometric shape optimization on the initial map and the position and attitude corresponding to the initial map based on a non-linear optimization algorithm, so as to obtain the target map and the position and attitude corresponding to the target map.

[0099] Further, the feature point trajectory determination module 230 may include: a target feature point determination unit, a motion information acquisition unit, and a feature point trajectory determination unit, where:

[0100] The target feature point determination unit is configured to determine the target feature points corresponding to the target vehicle network video according to the image features corresponding to each frame of the image, where the target feature points include corner points in the target vehicle network video.

[0101] The motion information acquisition unit is configured to track the target feature points in the target vehicle network video to obtain the motion information of the target feature points in different frames of the image in the target vehicle network video.

[0102] The feature point trajectory determination unit is configured to determine the feature point trajectory of the target feature points according to the motion information.

[0103] Further, the motion information acquisition unit may include: a motion information acquisition subunit, where:

[0104] The motion information acquisition subunit is configured to track the target feature points in different frames of the image of the target vehicle network video based on the optical flow method and a tracker to obtain the motion information.

[0105] Further, the position and pose corresponding to each frame of the image in the target vehicle network video include the six-degree-of-freedom position and pose corresponding to each frame of the image, and the position and pose determination module 230 may include: a position and pose determination unit, where:

[0106] The position and pose determination unit is configured to, if the shooting data includes at least one of the information of a vision sensor, the information of an inertial measurement unit, and the information of a global positioning system, determine the six-degree-of-freedom position and pose corresponding to each frame of the image in the target vehicle network video according to the exchangeable image file information corresponding to the target vehicle network video and the feature point trajectory.

[0107] Further, the network video acquisition module 210 may include: a vehicle network video acquisition unit and a target vehicle network video determination unit, where:

[0108] The vehicle network video acquisition unit is configured to acquire at least one vehicle network video.

[0109] The target vehicle network video determination unit is configured to determine the vehicle network video that meets the preset conditions in the at least one vehicle network video as the target vehicle network video, where the preset conditions include at least one of the video being a video of consecutive frames, the resolution of the video being greater than a preset resolution, the video including a moving object, and the duration of the video being greater than a preset duration.

[0110] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0111] In several embodiments provided by the present application, the coupling between modules can be electrical, mechanical or other forms of coupling.

[0112] In addition, in each embodiment of the present application, each functional module can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0113] Please refer to Figure 6 , which shows a structural block diagram of an electronic device 100 provided by an embodiment of the present application. The electronic device 100 can be an electronic device with processing capabilities such as a vehicle, a tablet computer, a robot, etc. The electronic device 100 in the present application can include one or more of the following components: a processor 110, a memory 120, and one or more application programs, where one or more application programs can be stored in the memory 120 and configured to be executed by one or more processors 110, and one or more programs are configured to execute the methods described in the foregoing method embodiments.

[0114] Among them, the processor 110 may include one or more processing cores. The processor 110 connects various parts within the entire electronic device 100 through various interfaces and lines, and executes various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 110 and may be implemented separately through a communication chip.

[0115] The memory 120 may include random access memory (RAM) and may also include read-only memory. The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the electronic device 100 (such as phone book, audio and video data, chat record data, etc.).

[0116] Please refer to Figure 7 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 300, and the program code can be called by the processor to execute the method described in the above method embodiments.

[0117] The computer-readable storage medium 300 may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 300 has a storage space for program code 310 that executes any of the method steps in the above-described method. These program codes may be read out from or written into one or more computer program products. The program code 310 may be compressed in a suitable form, for example.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for constructing a three-dimensional model, characterized in that, The method includes: Obtaining a target vehicle-mounted network video; Performing feature extraction on each frame image in the target vehicle-mounted network video to obtain the image features corresponding to each frame image; Performing feature point tracking on the target vehicle-mounted network video according to each frame image and the image features corresponding to each frame image, and determining the feature point trajectories of the same feature point in different frame images in the target vehicle-mounted network video; Determining the position and attitude corresponding to each frame image in the target vehicle-mounted network video according to the exchangeable image file information corresponding to the target vehicle-mounted network video and the feature point trajectories, where the exchangeable image file information includes the attribute information and shooting data of the target vehicle-mounted network video; Obtaining a three-dimensional model based on the target vehicle-mounted network video and the position and attitude corresponding to each frame image in the target vehicle-mounted network video.

2. The method according to claim 1, characterized in that, The number of the target vehicle-mounted network videos is multiple. The obtaining a three-dimensional model based on the target vehicle-mounted network video and the position and attitude corresponding to each frame image in the target vehicle-mounted network video includes: Performing cross-positioning on multiple target vehicle-mounted network videos based on a loop detection algorithm, the multiple target vehicle-mounted network videos, and the position and attitude corresponding to each frame image in each target vehicle-mounted network video to obtain loop constraint information of the multiple target vehicle-mounted network videos; Generating a target map and the position and attitude corresponding to the target map based on the loop constraint information, the multiple target vehicle-mounted network videos, and the position and attitude corresponding to each frame image in the multiple target vehicle-mounted network videos; Performing rendering and mapping on the target map based on the target map and the position and attitude corresponding to the target map to obtain the three-dimensional model, where the three-dimensional model includes at least one of point cloud reconstruction content, model reconstruction content, and neural network rendering reconstruction content corresponding to the target map.

3. The method according to claim 2, wherein The generating a target map and the position and attitude corresponding to the target map based on the loop constraint information, the multiple target vehicle-mounted network videos, and the position and attitude corresponding to each frame image in the multiple target vehicle-mounted network videos includes: Fusing the multiple target vehicle-mounted network videos and the position and attitude corresponding to each frame image in the multiple target vehicle-mounted network videos based on the loop constraint information to obtain an initial map and the position and attitude corresponding to the initial map; Performing topological structure optimization and geometric shape optimization on the initial map and the position and attitude corresponding to the initial map based on a nonlinear optimization algorithm to obtain the target map and the position and attitude corresponding to the target map.

4. The method according to claim 1, characterized in that, The performing feature point tracking on the target vehicle-mounted network video according to each frame image and the image features corresponding to each frame image, and determining the feature point trajectories of the same feature point in different frame images in the target vehicle-mounted network video includes: Determining target feature points corresponding to the target vehicle-mounted network video according to the image features corresponding to each frame image, where the target feature points include corner points in the target vehicle-mounted network video; In the target in-vehicle network video, track the target feature points to obtain the motion information of the target feature points in different frame images of the target in-vehicle network video; Determine the feature point trajectory of the target feature points according to the motion information.

5. The method according to claim 4, wherein The step of tracking the target feature points in the target in-vehicle network video to obtain the motion information of the target feature points in different frame images of the target in-vehicle network video includes: Based on the optical flow method and a tracker, track the target feature points in different frame images of the target in-vehicle network video to obtain the motion information.

6. The method according to any one of claims 1-5, characterized in that, The position and attitude corresponding to each frame image in the target in-vehicle network video include the six-degree-of-freedom position and attitude corresponding to each frame image. The step of determining the position and attitude corresponding to each frame image in the target in-vehicle network video according to the exchangeable image file information corresponding to the target in-vehicle network video and the feature point trajectory includes: If the shooting data includes at least one of the information of a vision sensor, an inertial measurement unit, and a global positioning system, determine the six-degree-of-freedom position and attitude corresponding to each frame image in the target in-vehicle network video according to the exchangeable image file information corresponding to the target in-vehicle network video and the feature point trajectory.

7. The method according to any one of claims 1-5, characterized in that The step of obtaining the target in-vehicle network video includes: Obtain at least one in-vehicle network video; Determine the in-vehicle network video that meets the preset conditions in the at least one in-vehicle network video as the target in-vehicle network video, where the preset conditions include at least one of the video being a video of consecutive frames, the resolution of the video being greater than a preset resolution, the video including a moving object, and the duration of the video being greater than a preset duration.

8. A three-dimensional model construction device, characterized in that The device includes: A network video acquisition module, configured to obtain a target in-vehicle network video; A feature extraction module, configured to extract features from each frame image in the target in-vehicle network video to obtain the image features corresponding to each frame image; A feature point trajectory determination module, configured to perform feature point tracking on the target in-vehicle network video according to each frame image and the image features corresponding to each frame image, and determine the feature point trajectory of the same feature point in different frame images of the target in-vehicle network video; A position and attitude determination module, configured to determine the position and attitude corresponding to each frame image in the target in-vehicle network video according to the exchangeable image file information corresponding to the target in-vehicle network video and the feature point trajectory, where the exchangeable image file information includes the attribute information and shooting data of the target in-vehicle network video; A three-dimensional model acquisition module, configured to obtain a three-dimensional model based on the target in-vehicle network video and the position and attitude corresponding to each frame image in the target in-vehicle network video.

9. An electronic device, characterized in that, including: One or more processors; A memory; One or more applications, where the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, and the program code can be called by a processor to execute the method according to any one of claims 1-7.