Three-dimensional point cloud model construction method and device, equipment and storage medium

By acquiring key point feature vectors and pose data from video frames to construct a 3D point cloud model, the problem of poor surface rendering of 3D point clouds in existing technologies is solved, and better rendering results are achieved.

CN121482277APending Publication Date: 2026-02-06CHINA MOBILE (JIANGXI) VIRTUAL REALITY TECH CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511718386.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies produce poor surface rendering of 3D point clouds, failing to effectively consider the reconstruction of surface images.

Method used

By acquiring multiple video frames, identifying the feature vectors of key points, calculating the pose data of the video frames, and constructing a 3D point cloud model based on the pose data, the surface rendering effect is improved.

Benefits of technology

It improves the surface rendering effect of 3D point cloud models, making them closer to the real surface of video frames and improving rendering quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482277A_ABST
    Figure CN121482277A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional point cloud model construction method and device, equipment and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: obtaining a plurality of video frames which are video frames in a video stream of a first scene; identifying feature vectors corresponding to the plurality of first key points in each video frame; determining a plurality of second key points of each video frame, the plurality of video frames including a first video frame, feature vectors corresponding to the plurality of second key points of the first video frame being respectively matched with feature vectors corresponding to different first key points in the second video frame, the plurality of second key points of the first video frame are part of key points in the plurality of first key points of the first video frame, and the second video frame is a previous video frame of the first video frame; respectively calculating pose data corresponding to each video frame based on the plurality of second key points of the plurality of video frames; and constructing a three-dimensional point cloud model based on the pose data corresponding to the plurality of video frames. According to the invention, the three-dimensional point cloud surface rendering effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a three-dimensional point cloud model construction method and device, equipment and a storage medium. BACKGROUND

[0002] Augmented Reality (AR) technology presents virtual objects constructed by a computer and information such as text, patterns, videos and the like as a prompt in a real world around a user, and allows the user to naturally interact with the virtual world information. In the related art, the application of AR technology is usually realized by constructing a three-dimensional point cloud. However, in the related art, the construction of a three-dimensional point cloud depends on the reconstruction of a spatial geometric position, and does not consider the reconstruction of a surface image of a three-dimensional point cloud, resulting in poor three-dimensional point cloud surface rendering effect in the related art.

[0003] It can be seen that the related art has the problem of poor three-dimensional point cloud surface rendering effect. SUMMARY

[0004] The embodiments of the present application provide a three-dimensional point cloud model construction method and device, electronic equipment and a readable storage medium to solve the problem of poor three-dimensional point cloud surface rendering effect in the related art.

[0005] To solve the above problem, the present application is implemented as follows:

[0006] In a first aspect, the embodiments of the present application provide a three-dimensional point cloud model construction method, comprising:

[0007] Obtaining a plurality of video frames, the plurality of video frames being video frames in a video stream of a first scene;

[0008] Identifying feature vectors corresponding to a plurality of first key points in each video frame;

[0009] Determining a plurality of second key points of each video frame, the plurality of video frames including a first video frame, feature vectors corresponding to the plurality of second key points of the first video frame respectively matching feature vectors corresponding to different first key points in a second video frame, the plurality of second key points of the first video frame being part of the plurality of first key points of the first video frame, and the second video frame being a previous video frame of the first video frame;

[0010] Respectively calculating pose data corresponding to each video frame based on the plurality of second key points of the plurality of video frames;

[0011] Constructing a three-dimensional point cloud model based on the pose data corresponding to the plurality of video frames.

[0012] In a second aspect, an embodiment of the present application further provides a three-dimensional point cloud model construction device, comprising:

[0013] An acquisition module is configured to acquire a plurality of video frames, the plurality of video frames being video frames in a video stream of a first scene;

[0014] A recognition module is configured to recognize feature vectors corresponding to a plurality of first key points in each video frame;

[0015] A determination module is configured to determine a plurality of second key points of each video frame, the plurality of video frames including a first video frame, feature vectors corresponding to the plurality of second key points of the first video frame matching feature vectors corresponding to different first key points in a second video frame respectively, the plurality of second key points of the first video frame being part of the plurality of first key points of the first video frame, the second video frame being a previous video frame of the first video frame;

[0016] A calculation module is configured to calculate pose data corresponding to each video frame based on the plurality of second key points of the plurality of video frames respectively;

[0017] A first construction module is configured to construct a three-dimensional point cloud model based on the pose data corresponding to the plurality of video frames.

[0018] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a transceiver and a processor,

[0019] The transceiver is configured to acquire a plurality of video frames, the plurality of video frames being video frames in a video stream of a first scene;

[0020] The processor is configured to recognize feature vectors corresponding to a plurality of first key points in each video frame;

[0021] The processor is further configured to determine a plurality of second key points of each video frame, the plurality of video frames including a first video frame, feature vectors corresponding to the plurality of second key points of the first video frame matching feature vectors corresponding to different first key points in a second video frame respectively, the plurality of second key points of the first video frame being part of the plurality of first key points of the first video frame, the second video frame being a previous video frame of the first video frame;

[0022] The processor is further configured to calculate pose data corresponding to each video frame based on the plurality of second key points of the plurality of video frames respectively;

[0023] The processor is further configured to construct a three-dimensional point cloud model based on the pose data corresponding to the plurality of video frames.

[0024] In a fourth aspect, an electronic device is provided, which includes a processor, a memory, and a program stored in the memory and executable on the processor, and when the program is executed by the processor, the steps of the three-dimensional point cloud model construction method of the first aspect are implemented.

[0025] In a fifth aspect, a computer readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the three-dimensional point cloud model construction method of the first aspect are implemented.

[0026] In a sixth aspect, a computer program product is provided, which includes computer instructions, and when the computer instructions are executed by a processor, the steps of the three-dimensional point cloud model construction method of the first aspect are implemented.

[0027] In the embodiment of the present application, a plurality of video frames are obtained, the plurality of video frames being video frames in a video stream of a first scene; feature vectors corresponding to a plurality of first key points in each video frame are identified; a plurality of second key points of each video frame are determined, the plurality of video frames including a first video frame, feature vectors corresponding to the plurality of second key points of the first video frame respectively matching feature vectors corresponding to different first key points in a second video frame, the plurality of second key points of the first video frame being part of the plurality of first key points of the first video frame, the second video frame being a previous video frame of the first video frame; pose data corresponding to each video frame is respectively calculated based on the plurality of second key points of the plurality of video frames; and a three-dimensional point cloud model is constructed based on the pose data corresponding to the plurality of video frames. In this way, by calculating the pose data corresponding to the video frames in the video stream and then constructing the three-dimensional point cloud model based on the pose data corresponding to the plurality of video frames, the surface of the constructed three-dimensional point cloud model is close to the video frames, which effectively improves the rendering effect of the surface of the three-dimensional point cloud relative to the related art which relies on spatial geometric positions to construct a three-dimensional point cloud. BRIEF DESCRIPTION OF DRAWINGS

[0028] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.

[0029] Figure 1 is a flowchart of a three-dimensional point cloud model construction method provided by an embodiment of the present application;

[0030] Figure 2 is a flowchart of a three-dimensional point cloud model construction method provided by an embodiment of the present application;

[0031] Figure 3 This is a structural diagram of a deep neural network model provided in an embodiment of the present invention;

[0032] Figure 4 This is a schematic diagram of instance segmentation of a three-dimensional point cloud model provided in an embodiment of the present invention;

[0033] Figure 5 This is a schematic diagram of Example 3 provided in the embodiments of the present invention;

[0034] Figure 6 This is a schematic diagram of Example 3 after the movement provided in the embodiment of the present invention;

[0035] Figure 7 This is a schematic diagram of the replaced Example 3 provided in the embodiment of the present invention;

[0036] Figure 8 This is a structural diagram of a three-dimensional point cloud model construction device provided in an embodiment of the present invention;

[0037] Figure 9 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] Please see Figure 1 , Figure 1 This is a flowchart of a three-dimensional point cloud model construction method provided by an embodiment of the present invention, such as... Figure 1 As shown, it includes the following steps:

[0040] Step 101: Obtain multiple video frames, which are video frames in the video stream of the first scene.

[0041] The aforementioned multiple video frames are video frames captured on-site at the first scene. In some implementations, a video stream can be captured on-site at the first scene, and then frames can be extracted from the video stream to obtain multiple video frames. For example, the captured video stream S can be represented as S={f i =(C i D i )} i , where C i Represents a Red, Green, Blue (RGB) image, D iRepresents a depth image, f i This represents the i-th video frame.

[0042] Among them, frame extraction from the video stream can be performed by extracting frames from the video stream of the first scene according to a preset period, such as a preset period of 0.2s, that is, a video frame is obtained by extracting frames every 0.2s.

[0043] The first scenario mentioned above is the scenario that requires 3D point cloud modeling. This is achieved by shooting a video stream on-site in the first scenario and extracting frames from the video stream to obtain multiple video frames, which are then used to construct a 3D point cloud model of the first scenario.

[0044] Step 102: Identify the feature vectors corresponding to multiple first key points in each video frame.

[0045] The aforementioned multiple first key points are key points in a video frame, and the features of a local region in the video frame can be determined through these multiple first key points. In some embodiments, the multiple first key points may be corner points, edge points, and / or points of interest in the video frame.

[0046] The aforementioned feature vector is the feature vector of the first key point, used to characterize the features of the location of the first key point. It should be noted that the pose of the camera used to capture the video stream changes during the shooting process. At this time, the position of objects in the captured video frames will change, and the pose change of the camera can be determined by the change in the object position. Therefore, in this invention, feature vectors corresponding to multiple first key points in each video frame are identified, and the same key points in different video frames are determined by using these feature vectors, thereby realizing the determination of the pose change of the camera.

[0047] Step 103: Determine multiple second key points for each video frame. The multiple video frames include a first video frame. The feature vectors corresponding to the multiple second key points of the first video frame are respectively matched with the feature vectors corresponding to different first key points in the second video frame. The multiple second key points of the first video frame are some of the multiple first key points of the first video frame. The second video frame is the video frame preceding the first video frame.

[0048] The feature vectors of the multiple second key points are matched with the feature vectors of the first key points in the previous video frame. That is, for multiple second key points in the first video frame, there is a corresponding first feature point in the second video frame. In this way, the pose change of each video frame relative to the previous video frame is calculated by comparing the multiple second key points of each video frame with the first key points in the previous video frame.

[0049] The number of second keypoints in each video frame can be the same or different, but the number of second keypoints in each video frame must be less than or equal to the number of first keypoints in the previous video frame. For example, there may be 10 first keypoints in the second video frame and 11 first keypoints in the first video frame, but the number of second keypoints in the first video frame must be less than or equal to 10, and the feature vectors of the second keypoints must match the feature vectors of the first keypoints in the second video frame.

[0050] Step 104: Calculate the pose data corresponding to each video frame based on the multiple second key points of the multiple video frames.

[0051] It is understandable that multiple second keypoints have matching first keypoints in the previous video frame. By calculating the offset between each second keypoint and its matching first keypoint, the pose data of the shooting device can be calculated based on the offset of multiple second keypoints, and thus the pose data corresponding to each video frame can be obtained.

[0052] Step 105: Construct a three-dimensional point cloud model based on the pose data corresponding to the multiple video frames.

[0053] In this embodiment of the invention, multiple video frames are acquired, which are video frames in a video stream of a first scene; feature vectors corresponding to multiple first key points in each video frame are identified; multiple second key points in each video frame are determined, the multiple video frames include a first video frame, the feature vectors corresponding to the multiple second key points of the first video frame are respectively matched with the feature vectors corresponding to different first key points in a second video frame, the multiple second key points of the first video frame are some of the multiple first key points of the first video frame, and the second video frame is the video frame preceding the first video frame; pose data corresponding to each video frame is calculated based on the multiple second key points of the multiple video frames; a three-dimensional point cloud model is constructed based on the pose data corresponding to the multiple video frames. Thus, by calculating the pose data corresponding to the video frames in the video stream, and then constructing a three-dimensional point cloud model based on the pose data corresponding to multiple video frames, the surface of the constructed three-dimensional point cloud model is close to that of the video frames. Compared with the method of constructing three-dimensional point clouds based on spatial geometric positions in related technologies, this effectively improves the surface rendering effect of the three-dimensional point cloud.

[0054] In one embodiment, calculating the pose data corresponding to each video frame based on the plurality of second keypoints of the plurality of video frames includes:

[0055] Calculate the distance between each second keypoint in each video frame, where the distance is the distance between the second keypoint and the matched first keypoint;

[0056] The pose data corresponding to each video frame is calculated, and the pose data corresponding to the first video frame is calculated based on the distance between the plurality of second key points of the first video frame.

[0057] It should be noted that when the pose of the shooting device changes, the positions of key points in the video frame it captures will also change. Different second key points in the video frame are located in different positions. By calculating the distance of each second key point, the change of the current video frame relative to the previous video frame can be determined, that is, the pose change of the corresponding video frame, and thus the pose data corresponding to each video frame can be obtained.

[0058] The distance of each second keypoint is the distance between the second keypoint and the matched first keypoint. This distance can be the straight-line distance or the Euclidean distance between the second keypoint and the matched first keypoint.

[0059] In one embodiment, the plurality of second keypoints of the first video frame are obtained in the following manner:

[0060] Calculate the first root mean square of the feature vectors corresponding to multiple first key points of the first video frame, and calculate the second root mean square of the feature vectors corresponding to multiple first key points of the second video frame.

[0061] Calculate the transformation matrix based on the first root mean square and the second root mean square;

[0062] Based on the transformation matrix, calculate the reprojection vectors of multiple first key points of the first video frame in the second video frame;

[0063] The plurality of second key points are determined, wherein the plurality of second key points are key points in the plurality of first key points of the first video frame whose reprojection vectors respectively match the feature vectors of the first key points in the second video frame.

[0064] The first root mean square (RMS) represents the overall characteristics of multiple first key points in the first video frame, and the second RM represents the overall characteristics of multiple first key points in the second video frame. The transformation matrix is ​​calculated based on the first and second RMs, and the transformation of the first and second video frames as a whole can be determined through the transformation matrix.

[0065] In some implementations, the transformation matrix can be calculated using the following formula:

[0066] T ij (ρ)=(T j -1 ×T i (ρ);

[0067] In the formula, T ij(ρ) represents the transformation matrix, T i T represents the root mean square of the i-th video frame. j This represents the root mean square of the j-th video frame, where the i-th and j-th video frames are adjacent video frames.

[0068] Furthermore, after obtaining the transformation matrix, the reprojection vectors of multiple first key points in the first video frame in the second video frame are calculated using the transformation matrix. The reprojection vectors represent the possible positions of the first key points in the first video frame in the second video frame, and thus multiple second key points can be determined through the reprojection vectors.

[0069] Among them, multiple second key points are key points in multiple first key points of the first video frame whose reprojection vectors match the feature vectors of the first key points in the second video frame. The matching of the reprojection vector with the feature vector of the first key point can be that the similarity between the reprojection vector and the feature vector of the first key point is greater than a set threshold.

[0070] For example, the set of multiple first keypoints of the first video frame is set P. cur Let the set of multiple first keypoints of the second video frame be set Q. cur The first and second root mean squares are calculated using the two sets respectively, and then the transformation matrix is ​​calculated. The set P is then transformed using the transformation matrix. cur The reprojection vectors of multiple first keypoints in the second video frame are used to determine the set P. cur With set Q cur The key points that correspond to each other, that is, determining the set P cur Several second key points in the process.

[0071] In some implementations, this can be achieved by adjusting set P. cur and set Q cur The covariance of the points and P cur and set Q cur Conditional analysis is performed on the cross-covariance between the two frames to determine the correspondence between key points. Specifically, if any condition number (i.e., set P) is used... cur Set Q cur The covariance of the points, or P cur and set Q curIf the cross-covariance between the two is high, it is considered an unstable correspondence. If the error between the reprojection vector after transformation matrix processing and the eigenvector of the first keypoint is large (e.g., greater than the set error threshold) or the conditional analysis determines that it is unstable, then the two are considered to be mismatched. Correspondences can be removed in order of error until this situation is no longer met, or there are too few correspondences to determine the rigid transformation. The remaining part are keypoints with matching relationships.

[0072] Additionally, if the first video frame f i Second video frame f j If no effective transformation can be generated between the two video frames (i.e., no matching keypoint exists), then all correspondences between the two video frames will be discarded, and it will be assumed that there is no second keyframe corresponding to the first video frame.

[0073] Furthermore, in addition to determining whether the reprojection vector and the first keypoint vector match by their similarity, as described above, matching can also be determined by the reprojection area. Specifically, in one embodiment, determining the plurality of second keypoints includes:

[0074] Multiple intermediate key points are determined, and a first key point of the second video frame corresponding to each intermediate key point is determined. The intermediate key points are key points in the multiple first key points of the first video frame whose reprojection vectors match the feature vectors of the first key points in the second video frame.

[0075] Calculate the projected area corresponding to each intermediate key point. The plurality of intermediate key points include a first intermediate key point. The projected area corresponding to the first intermediate key point is the area of ​​a first rectangle. The distance between the diagonals of the first rectangle is the distance between the first intermediate key point and the first key point of the second video frame corresponding to the first intermediate key point.

[0076] The plurality of second key points are determined from the plurality of intermediate key points, wherein the plurality of second key points are key points whose projected area is greater than a set area threshold among the plurality of intermediate key points.

[0077] Each of the aforementioned intermediate keypoints is the keypoint projected from the first keypoint in the first video frame onto the second video frame. It should be noted that when matching the first keypoints in two video frames, if the distance (physical dimension spanned) between two matching first keypoints is too close, errors may occur (e.g., if three or more keypoints are too close together, the correspondence between pairs of keypoints cannot be determined), meaning it's impossible to determine if it's a true match. In such cases, these keypoints need to be discarded to obtain multiple second keypoints with higher accuracy.

[0078] Specifically, in this embodiment of the invention, the distance between the first intermediate key point and the first key point of the second video frame corresponding to the first intermediate key point is calculated. The area of ​​the first rectangle is calculated based on the distance. This area is the projected area. Multiple second key points are determined from multiple intermediate key points by using a preset area threshold, so that each second key point can have a certain distance from the matched first key point, thereby improving the accuracy of multiple second key points.

[0079] For example, for the first video frame f i The first intermediate key point P and the corresponding second video frame f j For each keypoint Q, project two keypoints onto a plane defined by two principal axes. The projected area is obtained by the region enclosed by the two keypoints. If the area covered by P and Q is insufficient (i.e., the projected area is less than or equal to a set area threshold), the correspondence is considered ambiguous, meaning keypoints P and Q do not correspond, and the matching relationship needs to be discarded.

[0080] In some implementations, the reprojection error of the transformation matrix between the first video frame and the second video frame can be calculated using each second keypoint and the corresponding first keypoint. The reprojection error can then be used to determine the average depth difference, normal deviation, and photometric consistency of the reprojection.

[0081] Wherein, let π be the camera intrinsic parameter of the downsampled image, starting from the first video frame f i to the second video frame f j The reprojection error can be calculated using the following formula:

[0082] ;

[0083] In the formula Er(f i ,f j ) represents the reprojection error, x, y represent the positions of the key points in the video frame, and p i,x,y q represents the feature vector of the key points in the first video frame. i,x,y T represents the feature vector of the key points in the second video frame. ij ( ) represents the transformation matrix.

[0084] Furthermore, p i,x,y =p i low (x,y), p i low ( ) indicates the spatial position of the camera.

[0085] In some implementations, for video frames that require calculation of the corresponding second keyframe, the filtered and downsampled color intensity C is... ilow and depth D i low It is cached to improve filtering efficiency. Each D i low Camera spatial position p i low and normal N i low It is also calculated and cached.

[0086] In one embodiment, the distance of each second keypoint in each video frame includes a projected distance and / or a Euclidean space distance, and calculating the distance of each second keypoint in each video frame includes:

[0087] Calculate the projection distance of each second keypoint in the first video frame, wherein the plurality of second keypoints includes a third keypoint, and the projection distance of the third keypoint is the distance between the position of the third keypoint in the first video frame and the projection position of the third keypoint in the second video frame; and / or,

[0088] The Euclidean spatial distance of each second keypoint in the first video frame is calculated, and the Euclidean spatial distance of the third keypoint is calculated based on the position of the third keypoint in the depth map corresponding to the first video frame and the preset pose data.

[0089] The aforementioned projection distance characterizes the change in distance between matched keypoints on the projection plane in adjacent video frames. It can be calculated using the position of a keypoint in the first video frame and its projected position in the second video frame. In some implementations, the projection distance of the i-th keypoint can be expressed as... The first sum of the projected distances of multiple second keypoints can be expressed as:

[0090] ;

[0091] E in the formula sparse (X) represents the first sum value, x i x represents the position of the i-th second keypoint in the first video frame. i ' indicates the position of the i-th second keypoint in the second video frame.

[0092] The aforementioned Euclidean spatial distance is used to characterize the spatial distance variation of matched keypoints in adjacent video frames. Specifically, it can be calculated from the position of the depth map corresponding to the first video frame and the preset pose data. In some embodiments, the Euclidean spatial distance can be obtained by weighting the position of the depth map corresponding to the first video frame and the preset pose data.

[0093] The second sum of the Euclidean spatial distances between multiple second keypoints can be calculated using the following formula:

[0094] E dense (X)=W photo E photo (X)+W geo E geo (X);

[0095] ;

[0096] ;

[0097] E in the formula dense (X) represents the second sum, W photo and W photo τ represents the weighting coefficients, π represents the preset pose data, and d represents the homogeneous transformation. i,k I represents the position of the i-th second keypoint in the depth map. i and I j D represents the i-th and j-th adjacent video frames, respectively. i and D j Let i and j represent the i-th and j-th depth maps, respectively.

[0098] In one embodiment, calculating the pose data corresponding to each video frame includes:

[0099] Calculate the distance energy value, which is obtained by weighting the first sum of the projected distances of the plurality of second key points with the second sum of the Euclidean spatial distances of the plurality of second key points;

[0100] Adjust the preset pose data until the distance energy value meets the preset condition, and set the preset pose data as the pose data.

[0101] The preset conditions include the following:

[0102] The distance energy value is less than a set energy threshold;

[0103] The distance energy value is the minimum value after multiple adjustments to the preset pose data.

[0104] The aforementioned distance energy value is used to characterize the pose change between the first video frame and the second video frame. It should be noted that the smaller the distance energy value, the closer the preset pose data is to the actual pose data. Therefore, in this embodiment of the invention, the preset condition is set as follows: the distance energy value is less than a set energy threshold, or the distance energy value is the minimum value after multiple adjustments of the preset pose data, thereby effectively improving the accuracy of the calculated quota pose data.

[0105] The distance energy value can be calculated using the image formula:

[0106] E align (X)=W sparse E sparse (X)+W dense E dense (X);

[0107] E in the formula align (X) represents the distance energy value, W sparse and W dense This represents the weighting coefficient.

[0108] Furthermore, the preset pose data and the actual pose data can be represented by the following data structure:

[0109] X=(R0,t0,…,R s ,t s ) T =(x0,…,x N ) T ;

[0110] Where R represents the camera rotation matrix, T represents the camera translation vector, and x0,…,x N This represents the pose parameters of different video frames.

[0111] In one embodiment, constructing a 3D point cloud model based on the pose data corresponding to the plurality of video frames includes:

[0112] Obtain the historical truncation distance field of the three-dimensional point cloud model, which is constructed based on video frames prior to the first video frame;

[0113] Calculate voxel coordinates based on the pose data corresponding to the first video frame;

[0114] The historical truncation distance field is updated based on the voxel coordinates to obtain the first truncation distance field;

[0115] The three-dimensional point cloud model is updated based on the first truncated distance field.

[0116] In this embodiment of the invention, voxel coordinates are calculated based on the pose data corresponding to the first video frame; the historical truncated distance field is updated based on the voxel coordinates to obtain a first truncated distance field; and the 3D point cloud model is updated based on the first truncated distance field. Thus, by updating the truncated distance field using the pose data corresponding to each video frame, and then updating the 3D point cloud model using the truncated distance field, the construction of the 3D point cloud model of the first scene is completed.

[0117] The first cutoff distance field can be updated through the fusion operation using the following formula:

[0118] ;

[0119] The first truncated distance field is updated through the elimination operation using the following formula:

[0120] ;

[0121] In the formula, v is the voxel coordinate, D and W represent the historical cutoff distance field before the update, and D' and W' represent the first cutoff distance field after the update.

[0122] Specifically, given a new frame of depth image and camera pose, the existing truncated range field is updated using a fusion operation; whenever the old pose is optimized and updated, the old camera pose is removed using an elimination operation for fusion update, thus completing the final 3D reconstruction.

[0123] The process of constructing the 3D point cloud model in this invention can be achieved through... Figure 2 As shown, it includes the following steps:

[0124] Capture video frames from an RGB image;

[0125] Multiple secondary key points are determined based on video frames, specifically through reprojection.

[0126] Pose parameter optimization can be obtained by calculating the distance energy value mentioned above;

[0127] Construct a 3D point cloud model.

[0128] In one embodiment, after constructing the 3D point cloud model based on the pose data corresponding to the plurality of video frames, the method further includes:

[0129] Instance segmentation is performed on the plurality of video frames to obtain first segmentation data corresponding to each instance in the plurality of instances, wherein the plurality of instances are object instances included in the plurality of video frames;

[0130] The multiple instances in the three-dimensional point cloud model are segmented to obtain a first segmented point cloud corresponding to each instance;

[0131] Project the first segmented point cloud corresponding to each of the multiple instances to obtain the second segmented data corresponding to each instance;

[0132] Construct a Gaussian 3D model for each instance, wherein the plurality of instances includes a first instance, and the Gaussian 3D model of the first instance is constructed based on the first segmentation data and the second segmentation data of the first instance.

[0133] In this embodiment of the invention, the plurality of video frames are segmented to obtain first segmentation data corresponding to each instance in the plurality of instances, wherein the plurality of instances are object instances included in the plurality of video frames; the plurality of instances in the 3D point cloud model are segmented to obtain first segmented point clouds corresponding to each instance; the first segmented point clouds corresponding to the plurality of instances are projected to obtain second segmentation data corresponding to each instance; a Gaussian 3D model of each instance is constructed, wherein the plurality of instances include a first instance, and the Gaussian 3D model of the first instance is constructed based on the first segmentation data and the second segmentation data of the first instance. Thus, by segmenting the constructed 3D point cloud model to obtain Gaussian 3D models corresponding to different instances, model replacement can be performed on the Gaussian 3D model corresponding to the instance, further improving the display effect of the 3D point cloud model.

[0134] The above-mentioned instance segmentation of the multiple video frames to obtain the first segmentation data corresponding to each instance can be achieved by a sampling image segmentation network performing semantic segmentation and instance segmentation on images of different video frames to obtain segmentation data C corresponding to each video frame. i seg The segmented data carries semantic annotation information.

[0135] Furthermore, the segmentation result of the 3D point cloud model is projected, denoted as C. j seg Using C i seg and C j seg The correspondence can be used to determine C. i seg As C j seg The constraints are thus obtained, resulting in C that is consistent across all perspectives. i seg This is to achieve instance segmentation.

[0136] In some implementations, instance segmentation of the multiple video frames can be achieved using a deep neural network combining self-attention and masked cross-attention mechanisms. Specifically, such as... Figure 3 As shown, the model includes a feature backbone, a decoder, a mask module, and a decoder layer for query refinement.

[0137] The feature backbone outputs multi-scale features F, and the decoder iteratively refines the instance query X (i.e., the first segmented point cloud after segmentation), given the point features and instance query.

[0138] The masking module predicts a semantic class and instance heatmap for each query, generating a binary instance mask B (after thresholding). At the heart of the model are instance queries, each representing an object instance in the scene, and predicting a corresponding point-level instance mask. The masking module uses refined instance queries and point features, and returns the semantic class and binary instance mask based on the dot product between the point features and the instance queries.

[0139] The instance query is iteratively refined by the decoder, using a cross-attention mechanism to process the point features extracted from the feature backbone, and a self-attention mechanism to process other instance query features. The multi-level iteration produces the final instance query result.

[0140] The above model is used to segment instances in a 3D point cloud model. The segmentation result can be as follows: Figure 4 As shown, there are three examples: Example 1, Example 2, and Example 3.

[0141] In one embodiment, constructing the Gaussian 3D model for each instance includes:

[0142] The initial model is trained based on the first segmentation data and the second segmentation data of the first instance to obtain an intermediate 3D model. The initial model is used to represent instances in a 3D scene.

[0143] A rendered image is generated based on the intermediate 3D model;

[0144] Calculate the loss value corresponding to the intermediate 3D model based on the rendered image and the first segmentation data;

[0145] If the loss value is less than a preset loss threshold, the intermediate 3D model is set as the Gaussian 3D model of the first instance.

[0146] In this embodiment of the invention, an initial model is trained based on the first and second segmentation data of the first instance to obtain an intermediate 3D model. The initial model is used to represent instances in a 3D scene. A rendered image is generated based on the intermediate 3D model. A loss value corresponding to the intermediate 3D model is calculated based on the rendered image and the first segmentation data. If the loss value is less than a preset loss threshold, the intermediate 3D model is set as a Gaussian 3D model of the first instance. In this way, a Gaussian 3D model is obtained by training using the first and second segmentation data, enabling the Gaussian 3D model to accurately represent the situation of the instance.

[0147] Among them, the loss value L gaussian It can be calculated using the following formula:

[0148] L gaussian =L color +λ 2d L 2d +λ 3d L 3d ;

[0149] In the formula, L color L represents the cross-entropy loss value between the rendered image and the first segmentation data. 2d L represents the loss value of segmentation precision when Gaussian rendering is applied to 2D. 3d Let λ represent the regularization loss value. 2d and λ 3d This represents the weighting coefficient.

[0150] Furthermore, the regularization loss value L 3d It can be calculated using the following formula:

[0151] ;

[0152] In the formula, m represents the number of sampling points, F(e j F(e') represents the feature of the rendered sample point. j ) represents the characteristics of the first segment of data at the sampling point.

[0153] In some implementations, an identity encoding vector V of length 16 can be initialized for each Gaussian. IE Gaussian representation as information for 3D Gaussian segmentation. During 3D Gaussian reconstruction, the Gaussian representation is optimized based on the loss values ​​of the rendered image and the ground image. Simultaneously, a feature classifier is trained to optimize V... IE The representation of Gaussian segmentation information. The above regularization loss value L. 3d By leveraging the consistency of 3D space, we strive to ensure that the identity codes of the k nearest 3D Gaussians are within V. IE The proximity allows for more adequate supervision of 3D Gaussians located inside or heavily occluded 3D objects, improving the quality and accuracy of Gaussian 3D models.

[0154] Furthermore, to enhance the interactivity of the reconstructed 3D model, this invention can also provide moving components and other replaceable components for different Gaussian 3D models to enable the movement or replacement of the Gaussian 3D model. For example, Figure 5 and Figure 6 As shown, the Gaussian 3D model can be moved; as Figure 5 and Figure 7 As shown, the Gaussian 3D model can be replaced, thereby effectively improving the display effect of the 3D point cloud model.

[0155] Please seeFigure 8 , Figure 8 This is a structural diagram of a three-dimensional point cloud model construction device provided in an embodiment of the present invention, as shown below. Figure 8 As shown, the 3D point cloud model construction device 800 includes:

[0156] Acquisition module 801 is used to acquire multiple video frames, wherein the multiple video frames are video frames in the video stream of the first scene;

[0157] The recognition module 802 is used to recognize the feature vectors corresponding to multiple first key points in each video frame;

[0158] The determining module 803 is used to determine multiple second key points of each video frame, the multiple video frames including a first video frame, the feature vectors corresponding to the multiple second key points of the first video frame are respectively matched with the feature vectors corresponding to different first key points in the second video frame, the multiple second key points of the first video frame are some key points among the multiple first key points of the first video frame, and the second video frame is the video frame preceding the first video frame.

[0159] The calculation module 804 is used to calculate the pose data corresponding to each video frame based on multiple second key points of the multiple video frames;

[0160] The first construction module 805 is used to construct a three-dimensional point cloud model based on the pose data corresponding to the multiple video frames.

[0161] In one embodiment, the computing module 804 includes:

[0162] The first calculation submodule is used to calculate the distance of each second keypoint in each video frame, wherein the distance is the distance between the second keypoint and the matched first keypoint;

[0163] The second calculation submodule is used to calculate the pose data corresponding to each video frame, wherein the pose data corresponding to the first video frame is calculated based on the distance between the plurality of second key points of the first video frame.

[0164] In one embodiment, the plurality of second keypoints of the first video frame are obtained in the following manner:

[0165] Calculate the first root mean square of the feature vectors corresponding to multiple first key points of the first video frame, and calculate the second root mean square of the feature vectors corresponding to multiple first key points of the second video frame.

[0166] Calculate the transformation matrix based on the first root mean square and the second root mean square;

[0167] Based on the transformation matrix, calculate the reprojection vectors of multiple first key points of the first video frame in the second video frame;

[0168] The plurality of second key points are determined, wherein the plurality of second key points are key points in the plurality of first key points of the first video frame whose reprojection vectors respectively match the feature vectors of the first key points in the second video frame.

[0169] In one embodiment, determining the plurality of second key points includes:

[0170] Multiple intermediate key points are determined, and a first key point of the second video frame corresponding to each intermediate key point is determined. The intermediate key points are key points in the multiple first key points of the first video frame whose reprojection vectors match the feature vectors of the first key points in the second video frame.

[0171] Calculate the projected area corresponding to each intermediate key point. The plurality of intermediate key points include a first intermediate key point. The projected area corresponding to the first intermediate key point is the area of ​​a first rectangle. The distance between the diagonals of the first rectangle is the distance between the first intermediate key point and the first key point of the second video frame corresponding to the first intermediate key point.

[0172] The plurality of second key points are determined from the plurality of intermediate key points, wherein the plurality of second key points are key points whose projected area is greater than a set area threshold among the plurality of intermediate key points.

[0173] In one embodiment, the distance of each second keypoint in each video frame includes a projected distance and / or a Euclidean spatial distance, and the first calculation submodule includes:

[0174] A first calculation unit is configured to calculate the projection distance of each second keypoint in the first video frame, wherein the plurality of second keypoints includes a third keypoint, and the projection distance of the third keypoint is the distance between the position of the third keypoint in the first video frame and the projection position of the third keypoint in the second video frame; and / or,

[0175] The second calculation unit is used to calculate the Euclidean space distance of each second key point in the first video frame. The Euclidean space distance of the third key point is calculated based on the position of the third key point in the depth map corresponding to the first video frame and preset pose data.

[0176] In one embodiment, the second computing submodule includes:

[0177] The third calculation unit is used to calculate the distance energy value, which is obtained by weighting the first sum of the projected distances of the plurality of second key points and the second sum of the Euclidean spatial distances of the plurality of second key points.

[0178] A setting unit is used to adjust preset pose data until the distance energy value meets preset conditions, and then set the preset pose data as the pose data.

[0179] The preset conditions include the following:

[0180] The distance energy value is less than a set energy threshold;

[0181] The distance energy value is the minimum value after multiple adjustments to the preset pose data.

[0182] In one embodiment, the first building module 805 includes:

[0183] The acquisition submodule is used to acquire the historical truncation distance field of the three-dimensional point cloud model, which is constructed based on video frames before the first video frame;

[0184] The third calculation submodule is used to calculate voxel coordinates based on the pose data corresponding to the first video frame.

[0185] The first update submodule is used to update the historical truncation distance field based on the voxel coordinates to obtain the first truncation distance field;

[0186] The second update submodule is used to update the three-dimensional point cloud model based on the first truncated distance field.

[0187] In one embodiment, the three-dimensional point cloud model construction device 800 further includes:

[0188] The first segmentation module is used to perform instance segmentation on the plurality of video frames to obtain first segmentation data corresponding to each instance in the plurality of instances, wherein the plurality of instances are object instances included in the plurality of video frames;

[0189] The second segmentation module is used to segment the multiple instances in the three-dimensional point cloud model to obtain a first segmented point cloud corresponding to each instance;

[0190] The projection module is used to project the first segmented point cloud corresponding to the plurality of instances respectively to obtain the second segmented data corresponding to each instance;

[0191] The second construction module is used to construct a Gaussian 3D model for each instance, wherein the plurality of instances includes a first instance, and the Gaussian 3D model of the first instance is constructed based on the first segmentation data and the second segmentation data of the first instance.

[0192] In one embodiment, the second building module includes:

[0193] The training submodule is used to train the initial model based on the first segmentation data and the second segmentation data of the first instance to obtain an intermediate 3D model, wherein the initial model is used to represent instances in a 3D scene.

[0194] A generation submodule is used to generate a rendered image based on the intermediate 3D model;

[0195] The fourth calculation submodule is used to calculate the loss value corresponding to the intermediate 3D model based on the rendered image and the first segmentation data;

[0196] The processing submodule is used to set the intermediate 3D model as the Gaussian 3D model of the first instance when the loss value is less than a preset loss threshold.

[0197] The three-dimensional point cloud model construction device provided in this embodiment of the invention can realize each process of each embodiment of the above-mentioned three-dimensional point cloud model construction method. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0198] It should be noted that the three-dimensional point cloud model construction device in the embodiments of the present invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0199] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the above-described functionality. Figure 1 The various processes of the three-dimensional point cloud model construction method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0200] For details, see Figure 9 As shown, this embodiment of the invention also provides an electronic device, including a bus 901, a transceiver 902, an antenna 903, a bus interface 904, a processor 905, and a memory 906.

[0201] The transceiver 902 is used to acquire multiple video frames, which are video frames in the video stream of the first scene.

[0202] The processor 905 is used to identify feature vectors corresponding to multiple first key points in each video frame;

[0203] The processor 905 is further configured to determine a plurality of second key points for each video frame, the plurality of video frames including a first video frame, wherein the feature vectors corresponding to the plurality of second key points of the first video frame are respectively matched with the feature vectors corresponding to different first key points in the second video frame, the plurality of second key points of the first video frame are some key points among the plurality of first key points of the first video frame, and the second video frame is the video frame preceding the first video frame.

[0204] The processor 905 is further configured to calculate the pose data corresponding to each video frame based on multiple second key points of the multiple video frames;

[0205] The processor 905 is also used to construct a three-dimensional point cloud model based on the pose data corresponding to the multiple video frames.

[0206] In one embodiment, calculating the pose data corresponding to each video frame based on the plurality of second keypoints of the plurality of video frames includes:

[0207] Calculate the distance between each second keypoint in each video frame, where the distance is the distance between the second keypoint and the matched first keypoint;

[0208] The pose data corresponding to each video frame is calculated, and the pose data corresponding to the first video frame is calculated based on the distance between the plurality of second key points of the first video frame.

[0209] In one embodiment, the plurality of second keypoints of the first video frame are obtained in the following manner:

[0210] Calculate the first root mean square of the feature vectors corresponding to multiple first key points of the first video frame, and calculate the second root mean square of the feature vectors corresponding to multiple first key points of the second video frame.

[0211] Calculate the transformation matrix based on the first root mean square and the second root mean square;

[0212] Based on the transformation matrix, calculate the reprojection vectors of multiple first key points of the first video frame in the second video frame;

[0213] The plurality of second key points are determined, wherein the plurality of second key points are key points in the plurality of first key points of the first video frame whose reprojection vectors respectively match the feature vectors of the first key points in the second video frame.

[0214] In one embodiment, determining the plurality of second key points includes:

[0215] Multiple intermediate key points are determined, and a first key point of the second video frame corresponding to each intermediate key point is determined. The intermediate key points are key points in the multiple first key points of the first video frame whose reprojection vectors match the feature vectors of the first key points in the second video frame.

[0216] Calculate the projected area corresponding to each intermediate key point. The plurality of intermediate key points include a first intermediate key point. The projected area corresponding to the first intermediate key point is the area of ​​a first rectangle. The distance between the diagonals of the first rectangle is the distance between the first intermediate key point and the first key point of the second video frame corresponding to the first intermediate key point.

[0217] The plurality of second key points are determined from the plurality of intermediate key points, wherein the plurality of second key points are key points whose projected area is greater than a set area threshold among the plurality of intermediate key points.

[0218] In one embodiment, the distance of each second keypoint in each video frame includes a projected distance and / or a Euclidean space distance, and calculating the distance of each second keypoint in each video frame includes:

[0219] Calculate the projection distance of each second keypoint in the first video frame, wherein the plurality of second keypoints includes a third keypoint, and the projection distance of the third keypoint is the distance between the position of the third keypoint in the first video frame and the projection position of the third keypoint in the second video frame; and / or,

[0220] The Euclidean spatial distance of each second keypoint in the first video frame is calculated, and the Euclidean spatial distance of the third keypoint is calculated based on the position of the third keypoint in the depth map corresponding to the first video frame and the preset pose data.

[0221] In one embodiment, calculating the pose data corresponding to each video frame includes:

[0222] Calculate the distance energy value, which is obtained by weighting the first sum of the projected distances of the plurality of second key points with the second sum of the Euclidean spatial distances of the plurality of second key points;

[0223] Adjust the preset pose data until the distance energy value meets the preset condition, and set the preset pose data as the pose data.

[0224] The preset conditions include the following:

[0225] The distance energy value is less than a set energy threshold;

[0226] The distance energy value is the minimum value after multiple adjustments to the preset pose data.

[0227] In one embodiment, constructing a 3D point cloud model based on the pose data corresponding to the plurality of video frames includes:

[0228] Obtain the historical truncation distance field of the three-dimensional point cloud model, which is constructed based on video frames prior to the first video frame;

[0229] Calculate voxel coordinates based on the pose data corresponding to the first video frame;

[0230] The historical truncation distance field is updated based on the voxel coordinates to obtain the first truncation distance field;

[0231] The three-dimensional point cloud model is updated based on the first truncated distance field.

[0232] In one embodiment, the processor 905 is further configured to perform instance segmentation on the plurality of video frames to obtain first segmentation data corresponding to each instance in the plurality of instances, wherein the plurality of instances are object instances included in the plurality of video frames;

[0233] The processor 905 is further configured to segment the plurality of instances in the three-dimensional point cloud model to obtain a first segmented point cloud corresponding to each instance;

[0234] The processor 905 is further configured to project the first segmentation point cloud corresponding to the plurality of instances respectively to obtain the second segmentation data corresponding to each instance;

[0235] The processor 905 is further configured to construct a Gaussian 3D model for each instance, wherein the plurality of instances includes a first instance, and the Gaussian 3D model of the first instance is constructed based on the first segmentation data and the second segmentation data of the first instance.

[0236] In one embodiment, constructing the Gaussian 3D model for each instance includes:

[0237] The initial model is trained based on the first segmentation data and the second segmentation data of the first instance to obtain an intermediate 3D model. The initial model is used to represent instances in a 3D scene.

[0238] A rendered image is generated based on the intermediate 3D model;

[0239] Calculate the loss value corresponding to the intermediate 3D model based on the rendered image and the first segmentation data;

[0240] If the loss value is less than a preset loss threshold, the intermediate 3D model is set as the Gaussian 3D model of the first instance.

[0241] existFigure 9 In this document, a bus architecture (represented by bus 901) is used. Bus 901 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 905 and memory represented by memory 906. Bus 901 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 904 provides an interface between bus 901 and transceiver 902. Transceiver 902 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 905 is transmitted over a wireless medium via antenna 903, which further receives data and transmits it to processor 905.

[0242] Processor 905 manages bus 901 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 906 can be used to store data used by processor 905 during operation.

[0243] Optionally, the processor 905 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a graphics processing unit (GPU).

[0244] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the above-described functions. Figure 1 The various processes corresponding to the embodiments of the 3D point cloud model construction method, and which achieve the same technical effect, will not be described again here to avoid repetition. The computer-readable storage medium mentioned includes, for example, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0245] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The various processes corresponding to the three-dimensional point cloud model construction method embodiments, and which can achieve the same technical effect, will not be described in detail here to avoid repetition.

[0246] In the embodiments of this invention, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing eight possibilities: A alone, B alone, C alone, both A and B present, both B and C present, both A and C present, and A, B, and C present.

[0247] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0248] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0249] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for constructing a three-dimensional point cloud model, characterized in that, include: Acquire multiple video frames, wherein the multiple video frames are video frames in the video stream of the first scene; Identify the feature vectors corresponding to multiple first keypoints in each video frame; Multiple second key points are determined for each video frame, the multiple video frames include a first video frame, the feature vectors corresponding to the multiple second key points of the first video frame are respectively matched with the feature vectors corresponding to different first key points in the second video frame, the multiple second key points of the first video frame are some key points among the multiple first key points of the first video frame, and the second video frame is the video frame preceding the first video frame. The pose data corresponding to each video frame is calculated based on the multiple second key points of the multiple video frames; A three-dimensional point cloud model is constructed based on the pose data corresponding to the multiple video frames.

2. The method as described in claim 1, characterized in that, The step of calculating the pose data corresponding to each video frame based on multiple second key points of the multiple video frames includes: Calculate the distance between each second keypoint in each video frame, where the distance is the distance between the second keypoint and the matched first keypoint; The pose data corresponding to each video frame is calculated, and the pose data corresponding to the first video frame is calculated based on the distance between the plurality of second key points of the first video frame.

3. The method as described in claim 2, characterized in that, The multiple second key points of the first video frame are obtained in the following manner: Calculate the first root mean square of the feature vectors corresponding to multiple first key points of the first video frame, and calculate the second root mean square of the feature vectors corresponding to multiple first key points of the second video frame. Calculate the transformation matrix based on the first root mean square and the second root mean square; Based on the transformation matrix, calculate the reprojection vectors of multiple first key points of the first video frame in the second video frame; The plurality of second key points are determined, wherein the plurality of second key points are key points in the plurality of first key points of the first video frame whose reprojection vectors respectively match the feature vectors of the first key points in the second video frame.

4. The method as described in claim 3, characterized in that, Determining the plurality of second key points includes: Multiple intermediate key points are determined, and a first key point of the second video frame corresponding to each intermediate key point is determined. The intermediate key points are key points in the multiple first key points of the first video frame whose reprojection vectors match the feature vectors of the first key points in the second video frame. Calculate the projected area corresponding to each intermediate key point. The plurality of intermediate key points include a first intermediate key point. The projected area corresponding to the first intermediate key point is the area of ​​a first rectangle. The distance between the diagonals of the first rectangle is the distance between the first intermediate key point and the first key point of the second video frame corresponding to the first intermediate key point. The plurality of second key points are determined from the plurality of intermediate key points, wherein the plurality of second key points are key points whose projected area is greater than a set area threshold among the plurality of intermediate key points.

5. The method according to any one of claims 2 to 4, characterized in that, The distance of each second keypoint in each video frame includes the projected distance and / or the Euclidean space distance. The calculation of the distance of each second keypoint in each video frame includes: Calculate the projection distance of each second keypoint in the first video frame, wherein the plurality of second keypoints includes a third keypoint, and the projection distance of the third keypoint is the distance between the position of the third keypoint in the first video frame and the projection position of the third keypoint in the second video frame; and / or, The Euclidean spatial distance of each second keypoint in the first video frame is calculated, and the Euclidean spatial distance of the third keypoint is calculated based on the position of the third keypoint in the depth map corresponding to the first video frame and the preset pose data.

6. The method as described in claim 5, characterized in that, The calculation of the pose data corresponding to each video frame includes: Calculate the distance energy value, which is obtained by weighting the first sum of the projected distances of the plurality of second key points with the second sum of the Euclidean spatial distances of the plurality of second key points; Adjust the preset pose data until the distance energy value meets the preset condition, and set the preset pose data as the pose data. The preset conditions include the following: The distance energy value is less than a set energy threshold; The distance energy value is the minimum value after multiple adjustments to the preset pose data.

7. The method according to any one of claims 1 to 4, characterized in that, The construction of a 3D point cloud model based on the pose data corresponding to the multiple video frames includes: Obtain the historical truncation distance field of the three-dimensional point cloud model, which is constructed based on video frames prior to the first video frame; Calculate voxel coordinates based on the pose data corresponding to the first video frame; The historical truncation distance field is updated based on the voxel coordinates to obtain the first truncation distance field; The three-dimensional point cloud model is updated based on the first truncated distance field.

8. The method according to any one of claims 1 to 4, characterized in that, After constructing the 3D point cloud model based on the pose data corresponding to the multiple video frames, the method further includes: Instance segmentation is performed on the plurality of video frames to obtain first segmentation data corresponding to each instance in the plurality of instances, wherein the plurality of instances are object instances included in the plurality of video frames; The multiple instances in the three-dimensional point cloud model are segmented to obtain a first segmented point cloud corresponding to each instance; Project the first segmented point cloud corresponding to each of the multiple instances to obtain the second segmented data corresponding to each instance; Construct a Gaussian 3D model for each instance, wherein the plurality of instances includes a first instance, and the Gaussian 3D model of the first instance is constructed based on the first segmentation data and the second segmentation data of the first instance.

9. The method as described in claim 8, characterized in that, The construction of the Gaussian 3D model for each instance includes: The initial model is trained based on the first segmentation data and the second segmentation data of the first instance to obtain an intermediate 3D model. The initial model is used to represent instances in a 3D scene. A rendered image is generated based on the intermediate 3D model; Calculate the loss value corresponding to the intermediate 3D model based on the rendered image and the first segmentation data; If the loss value is less than a preset loss threshold, the intermediate 3D model is set as the Gaussian 3D model of the first instance.

10. A three-dimensional point cloud model construction device, characterized in that, include: The acquisition module is used to acquire multiple video frames, wherein the multiple video frames are video frames in the video stream of the first scene; The recognition module is used to identify the feature vectors corresponding to multiple first keypoints in each video frame; The determining module is used to determine multiple second key points of each video frame, the multiple video frames including a first video frame, the feature vectors corresponding to the multiple second key points of the first video frame are respectively matched with the feature vectors corresponding to different first key points in the second video frame, the multiple second key points of the first video frame are some key points among the multiple first key points of the first video frame, and the second video frame is the video frame preceding the first video frame. The calculation module is used to calculate the pose data corresponding to each video frame based on multiple second key points of the multiple video frames; The first construction module is used to construct a three-dimensional point cloud model based on the pose data corresponding to the multiple video frames.

11. An electronic device, characterized in that, Including transceivers and processors, The transceiver is used to acquire multiple video frames, which are video frames in the video stream of the first scene. The processor is used to identify feature vectors corresponding to multiple first key points in each video frame; The processor is further configured to determine a plurality of second key points for each video frame, the plurality of video frames including a first video frame, wherein the feature vectors corresponding to the plurality of second key points of the first video frame are respectively matched with the feature vectors corresponding to different first key points in the second video frame, the plurality of second key points of the first video frame are some key points among the plurality of first key points of the first video frame, and the second video frame is the video frame preceding the first video frame. The processor is further configured to calculate the pose data corresponding to each video frame based on multiple second key points of the multiple video frames; The processor is also used to construct a three-dimensional point cloud model based on the pose data corresponding to the plurality of video frames.

12. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the three-dimensional point cloud model construction method as described in any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the three-dimensional point cloud model construction method as described in any one of claims 1 to 9.

14. A computer program product, characterized in that, It includes computer instructions, which, when executed by a processor, implement the steps of the three-dimensional point cloud model construction method as described in any one of claims 1 to 9.