Pose optimization method, device and equipment and readable storage medium
By using the backpropagation gradient descent algorithm and multi-scale loss function in the SLAM system, the camera pose is optimized, and the instability problem of pose estimation is solved when the camera is fast motion or scene changes are complex, and the stability and accuracy of pose estimation are improved.
Patent Information
- Application Number
- CN202510027266.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
The existing SLAM system based on Gaussian point clouds has unstable pose estimation accuracy when the camera moves rapidly or scene changes are complex, making it difficult to provide sufficiently stable gradient information.
The backpropagation gradient descent algorithm is used to optimize the camera position by setting the multi-round optimization process, using multi-scale loss functions (such as key point loss, object-level loss or scene-level loss) in the first K rounds, and using depth and color loss functions in the rear L-K rounds.
Through the use of multi-scale loss functions, gradient updates become smoother and hierarchical, avoiding unstable fluctuations in gradient updates in traditional methods, and improving the stability and accuracy of pose estimation.
Smart Images

Figure CN119941856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of tracking and positioning technology, and in particular to a posture optimization method, device, equipment and readable storage medium. Background Art
[0002] The SLAM system based on Gaussian point cloud can be widely used in future AR and VR scenarios. The SLAM system is a system that simultaneously locates and models objects; while Gaussian point cloud is a new technology that is different from traditional point cloud and can render realistic scene photos. Using the SLAM system based on Gaussian point cloud, we can quickly reconstruct indoor scenes, and the Gaussian point cloud data obtained after reconstruction can be directly used by AR and VR systems.
[0003] The existing SLAM system based on Gaussian point cloud has some problems in the tracking and positioning process. Generally, the gradient descent algorithm is used to calculate the camera pose we need, but the loss used generally only uses the pixel-level L1 loss function. This simple loss function is often difficult to provide sufficiently stable gradient information when dealing with dynamic scenes or fast camera movement, because even a small offset at the pixel level will cause a large fluctuation in the loss function, resulting in instability and reduced accuracy in the pose estimation process. Summary of the invention
[0004] The object of the present invention is to provide a posture optimization method, device, equipment and readable storage medium to improve the problems of instability and decreased accuracy in the above-mentioned posture estimation process.
[0005] In order to achieve the above objectives, the present application provides the following technical solutions:
[0006] On the one hand, an embodiment of the present application provides a posture optimization method, the method comprising:
[0007] During the tracking process of the SplaTAM system, the round L of the back-propagation gradient descent algorithm is set, and the back-propagation gradient descent algorithm is used to optimize the camera's pose at the current moment to obtain the camera pose optimization result; among them, multi-scale loss is used for optimization in the first K rounds, and depth and color loss is used for optimization in the last LK rounds. The multi-scale loss is key point loss, object-level loss or scene-level loss.
[0008] In a second aspect, an embodiment of the present application provides a posture optimization device, the device comprising:
[0009] The optimization module is used to set the round L of the back-propagation gradient descent algorithm during the tracking process of the SplaTAM system, and use the back-propagation gradient descent algorithm to optimize the camera's pose at the current moment to obtain the camera pose optimization result; wherein, multi-scale loss is used for optimization in the first K rounds, and depth and color loss is used for optimization in the last LK rounds. The multi-scale loss is key point loss, object-level loss or scene-level loss.
[0010] In a third aspect, an embodiment of the present application provides a posture optimization device, the device comprising a memory and a processor. The memory is used to store a computer program; the processor is used to implement the steps of the posture optimization method when executing the computer program.
[0011] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned posture optimization method are implemented.
[0012] The beneficial effects of the present invention are:
[0013] 1. The present invention uses a multi-scale loss function to optimize the SLAM system based on Gaussian point cloud to improve the accuracy and robustness of camera pose tracking. The system optimizes the estimation of camera pose by integrating information of different scales, and is particularly suitable for environments with violent camera motion or complex scene changes. The invention can be widely used in fields such as augmented reality, virtual reality, robot navigation, and three-dimensional map construction, providing more accurate visual positioning and environmental perception capabilities. At the same time, the present invention can greatly improve the tracking performance of the existing SLAM system based on Gaussian point cloud, and greatly reduce its pose deviation.
[0014] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or be understood by implementing the embodiments of the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 is a flow chart of the posture optimization method described in an embodiment of the present invention;
[0017] Figure 2 is a schematic structural diagram of a posture optimization device described in an embodiment of the present invention;
[0018] Figure 3 It is a schematic diagram of the structure of the posture optimization device described in an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0020] It should be noted that similar reference numerals or letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0021] Example 1
[0022] like Figure 1 As shown, this embodiment provides a posture optimization method, which includes step S1.
[0023] Step S1, during the tracking process of the SplaTAM system, set the round L of the back-propagation gradient descent algorithm, use the back-propagation gradient descent algorithm to optimize the camera's posture at the current moment, and obtain the camera posture optimization result; wherein, multi-scale loss is used for optimization in the first K rounds, and depth and color loss is used for optimization in the last LK rounds, and the multi-scale loss is key point loss, object level loss or scene level loss.
[0024] The SplaTAM system in the present invention is the SplaTAM system (SLAM system based on Gaussian point cloud) in the paper SplaTAM: Splat, Track & Map 3D Gaussians for DenseRGB-D SLAM; K can be 20% L, and the calculation of depth and color loss is to use the default depth loss and color loss function in the SplaTAM system, and the weighted sum of the two is taken as the final loss, and the weights of the two are 0.5 and 1 respectively; the present invention aims to use multi-scale loss as the loss function for optimization in the tracking process of the SplaTAM system; specifically, in the present invention, the specific implementation steps of using multi-scale loss for optimization during the first N optimizations include step S11;
[0025] Step S11, each time the multi-scale loss is used for optimization, it is determined whether the key point loss can be used as the loss for optimization. If so, the key point loss is used as the loss for optimization. If not, it is determined whether the object-level loss can be used as the loss for optimization. If so, the object-level loss is used as the loss for optimization. If not, the scene-level loss is used as the loss for optimization.
[0026] In step S11, it is determined whether the key point loss can be used as the loss for optimization. If so, the specific implementation steps of using the key point loss as the loss for optimization include step S111 and step S112;
[0027] Step S111, using the Gaussian point cloud system included in the SplaTAM system to render the original image acquired by the camera at the current moment to obtain a rendered image; inputting the original image and the rendered image into the key point matching model respectively to obtain the matching point pairs between the rendered image and the original image, and counting the number of matching point pairs. When the number is less than a preset first threshold, it is determined that the key point loss cannot be used as the loss for optimization; otherwise, the key point loss can be used as the loss for optimization;
[0028] In this step, the key point matching model is Super Point or SIFT; the matching point pair is two matching points in the original image and the rendered image; the first threshold can be 50;
[0029] Step S112: When the key point loss is used as the loss for optimization, the matching point pairs are first filtered by distance and angle to obtain the filtered matching point pairs; then the loss is calculated according to formula (1) and the filtered matching point pairs. Formula (1) is:
[0030]
[0031] In formula (1), N is the number of matching point pairs after filtering, is the coordinate of the i-th pair of matching points in the rendered image, are the coordinates of the i-th pair of matching points in the original image.
[0032] In this step, is the coordinates of the i-th pair of matching points in the rendered image, which can be understood as the coordinates of the matching points in a pair of matching points that exist in the rendered image; is the coordinates of the i-th pair of matching points in the original image, which can be understood as the coordinates of the matching points in a pair of matching points that exist in the original image;
[0033] Meanwhile, in this step, the matching point pairs are filtered by distance and angle, and the specific implementation steps of obtaining the filtered matching point pairs include step S1121 and step S1122;
[0034] Step S1121, when performing distance filtering, calculate the Euclidean distance between each pair of matching points, divide all the Euclidean distances into boxes with equal width to obtain multiple boxes, put each pair of matching points into a corresponding box according to the box in which the Euclidean distance between each pair of matching points falls, calculate the total number of matching point pairs in three adjacent boxes, and use the matching point pairs in the three adjacent boxes with the largest total number as the matching point pairs input for angle filtering;
[0035] Step S1122, when performing angle filtering, calculate the vector direction between each pair of matching points, divide the vector directions into circular equidistant boxes, put each pair of matching points into the corresponding box according to the box where the vector direction between each pair of matching points falls, calculate the total number of matching point pairs in three adjacent boxes, and use the matching point pairs in the three adjacent boxes with the largest total number as the filtered matching point pairs.
[0036] In step S11, it is determined whether the object-level loss can be used as the loss for optimization. If so, the specific implementation steps of using the object-level loss as the loss for optimization include step S113, step S114 and step S115;
[0037] Step S113: input the rendered image and the original image into the Yolo-World model respectively to obtain the target boxes in the original image and the rendered image;
[0038] Step S114: for each category of target frames, first set T pairs of target frames to match, where the value range of T is T<=M and T<=S, and the target frames in the original image and the rendered image are combined in pairs to form T pairs of target frames, and the target frame pairs formed by all the combinations are used as all target frame pairs corresponding to each T value, where S is the number of target frames contained in each category in the original image, M is the number of target frames contained in each category in the rendered image, and T is an integer; after all T values in the value range are taken, determine whether all target frame pairs corresponding to each T value are retained, and calculate the ranking scores of all target frames corresponding to each T value;
[0039] This step can be understood as follows: for each category, assuming that M is 2 and S is 3, then T can be 1 or 2, that is, matching one pair of target frames or two pairs of target frames; when T is 1, the target frames in the original image and the rendered image are arranged and combined in pairs, that is, The target frame pairs formed by all the combinations are taken as all the target frame pairs corresponding to T = 1; then all the target frame pairs corresponding to T = 2 are calculated according to the same logic; after T is taken, it is determined whether all the target frame pairs corresponding to the T value (1, 2) are retained, and the ranking scores of all the target frames corresponding to the T value (1, 2) are calculated;
[0040] At the same time, in this step, it is determined whether all target frame pairs corresponding to each T value are retained, and the specific implementation steps of calculating the ranking scores of all target frames corresponding to each T value include step S1141;
[0041] Step S1141, for each T value, analyze all the corresponding target frame pairs; wherein, for each target frame pair, record the target frame existing in the original image as the first target frame, and record the target frame existing in the rendered image as the second target frame; first calculate the Euclidean distance between the center point coordinate position of the first target frame and the center point coordinate position of the second target frame in each target frame pair; record the difference between the maximum value and the minimum value of all the distances corresponding to all the target frame pairs as the first value, multiply the length of the smallest side of the original image by a preset ratio, record it as the second value, and record the ratio between the second value and the first value as the first fraction; then establish the horizontal and vertical coordinate axes with the center point of the first target frame in each target frame pair as the origin, and record the second target frame as the origin. The center point of the target frame is connected to the origin to form a line segment, the angle between the line segment and the horizontal axis is calculated, the difference between the maximum and minimum values of all angles corresponding to all target frame pairs is recorded as the third value, and the ratio between the preset degree and the third value is recorded as the second score; finally, the similarity of the color histogram vectors of the first target frame and the second target frame in each target frame pair is calculated to obtain multiple similarity calculation results and take the average of all similarity calculation results as the third score; when the first score is greater than the preset second threshold and the second score is greater than the preset third threshold and each similarity calculation result is greater than the preset fourth threshold, all target frame pairs corresponding to this T value are retained, and the first score, the second score and the third score are weighted averaged to obtain the ranking score.
[0042] In this step, the preset ratio may be one sixth, the preset degree may be 30 degrees, the second threshold and the third threshold may both be 1, the fourth threshold may be 0.7, and the weights of the first score, the second score, and the third score may be 0.25, 0.25, and 0.5, respectively;
[0043] Step S115, count the number of all target frame pairs that are finally retained. When the number is less than the preset fifth threshold, it is determined that object-level loss cannot be used as the loss for optimization; otherwise, object-level loss can be used as the loss for optimization. When key point loss is used as the loss for optimization, all target frame pairs corresponding to the maximum ranking score are screened out, and the Euclidean distance between the coordinate positions of the two target frame center points of each target frame pair is calculated, and the average of all Euclidean distances is used as the loss.
[0044] In the step, the fifth threshold may be 2;
[0045] In step S11, the specific implementation steps of using the scene level loss as the loss for optimization include step S116;
[0046] Step S116: input the rendered image and the original image into the convolutional neural network respectively, use the output of the convolutional neural network as the feature vectors of the rendered image and the original image respectively, calculate the cosine similarity between the two feature vectors, and use it as the loss.
[0047] The purpose of the present invention is to solve the shortcomings of the existing SLAM system based on Gaussian point cloud, especially the problem of unstable pose estimation accuracy when the camera moves quickly or the scene changes are complex. Specifically, this embodiment introduces a multi-scale loss function to optimize the gradient descent process, making the gradient update smoother and hierarchical, avoiding the unstable fluctuation of gradient update in traditional methods, thereby improving the stability and accuracy of pose estimation.
[0048] Example 2
[0049] like Figure 2 As shown, this embodiment provides a posture optimization device, which includes an optimization module 1.
[0050] Optimization module 1 is used to set the round L of the back-propagation gradient descent algorithm during the tracking process of the SplaTAM system, and use the back-propagation gradient descent algorithm to optimize the camera posture at the current moment to obtain the camera posture optimization result; wherein, multi-scale loss is used for optimization in the first K rounds, and depth and color loss is used for optimization in the last LK rounds, and the multi-scale loss is key point loss, object level loss or scene level loss.
[0051] It should be noted that, regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here.
[0052] Example 3
[0053] Corresponding to the above method embodiments, the embodiments of the present disclosure further provide a posture optimization device, and the posture optimization device described below and the posture optimization method described above can be referred to each other.
[0054] Figure 3 is a block diagram of a posture optimization device 300 according to an exemplary embodiment. Figure 3 As shown, the posture optimization device 300 may include: a processor 301 and a memory 302. The posture optimization device 300 may also include one or more of a multimedia component 303, an I / O interface 304, and a communication component 305.
[0055] The processor 301 is used to control the overall operation of the posture optimization device 300 to complete all or part of the steps in the above-mentioned posture optimization method. The memory 302 is used to store various types of data to support the operation of the posture optimization device 300. These data may include, for example, instructions for any application or method used to operate on the posture optimization device 300, and application-related data, such as contact data, messages sent and received, pictures, audio, video, etc. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, referred to as EPROM), programmable read-only memory (Programmable Read-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, disk or optical disk. The multimedia component 303 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 302 or sent through the communication component 305. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 304 provides an interface between the processor 301 and other interface modules, and the above-mentioned other interface modules may be keyboards, mice, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 305 is used for wired or wireless communication between the posture optimization device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 305 may include: Wi-Fi module, Bluetooth module, NFC module.
[0056] In an exemplary embodiment, the posture optimization device 300 can be implemented by one or more application specific integrated circuits (Application Specific Integrated Circuit, referred to as ASIC), digital signal processors (Digital Signal Processor, referred to as DSP), digital signal processing devices (Digital Signal Processing Device, referred to as DSPD), programmable logic devices (Programmable Logic Device, referred to as PLD), field programmable gate arrays (Field Programmable Gate Array, referred to as FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned posture optimization method.
[0057] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by a processor, the steps of the above-mentioned posture optimization method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 302 including program instructions, and the above-mentioned program instructions can be executed by the processor 301 of the posture optimization device 300 to complete the above-mentioned posture optimization method.
[0058] Example 4
[0059] Corresponding to the above method embodiment, the embodiment of the present disclosure further provides a readable storage medium. The readable storage medium described below and the posture optimization method described above can refer to each other.
[0060] A readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the posture optimization method of the above method embodiment.
[0061] The readable storage medium may specifically be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or other readable storage medium that can store program codes.
[0062] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A posture optimization method, characterized in that: include: During the tracking process of the SplaTAM system, the round L of the back-propagation gradient descent algorithm is set, and the back-propagation gradient descent algorithm is used to optimize the camera's pose at the current moment to obtain the camera pose optimization result; among them, multi-scale loss is used for optimization in the first K rounds, and depth and color loss is used for optimization in the last LK rounds. The multi-scale loss is key point loss, object-level loss or scene-level loss.
2. The posture optimization method according to claim 1, characterized in that: Multi-scale loss is used for optimization in the first N optimizations, including: Each time multi-scale loss is used for optimization, it is determined whether key point loss can be used as the loss for optimization. If so, key point loss is used as the loss for optimization. If not, it is determined whether object-level loss can be used as the loss for optimization. If so, object-level loss is used as the loss for optimization. If not, scene-level loss is used as the loss for optimization.
3. The posture optimization method according to claim 2, characterized in that: Determine whether the key point loss can be used as the loss for optimization. If so, use the key point loss as the loss for optimization, including: The Gaussian point cloud system included in the SplaTAM system is used to render the original image acquired by the camera at the current moment to obtain a rendered image; the original image and the rendered image are respectively input into the key point matching model to obtain the matching point pairs between the rendered image and the original image, and the number of matching point pairs is counted. When the number is less than the preset first threshold, it is determined that the key point loss cannot be used as the loss for optimization; otherwise, the key point loss can be used as the loss for optimization; When the key point loss is used as the loss for optimization, the matching point pairs are first filtered by distance and angle to obtain the filtered matching point pairs; then the loss is calculated according to formula (1) and the filtered matching point pairs. Formula (1) is: In formula (1), N is the number of matching point pairs after filtering, is the coordinate of the i-th pair of matching points in the rendered image, are the coordinates of the i-th pair of matching points in the original image.
4. The posture optimization method according to claim 3, characterized in that: Filter the matching point pairs by distance and angle to obtain the filtered matching point pairs, including: When performing distance filtering, calculate the Euclidean distance between each pair of matching points, divide all the Euclidean distances into equal-width boxes to obtain multiple boxes, and put each pair of matching points into the corresponding box according to the box where the Euclidean distance between each pair of matching points falls. Calculate the total number of matching point pairs in three adjacent boxes, and use the matching point pairs in the three adjacent boxes with the largest total number as the matching point pairs input for angle filtering; When performing angle filtering, calculate the vector direction between each pair of matching points, divide the vector directions into a circular equidistant box, and put each pair of matching points into the corresponding box according to the box where the vector direction between each pair of matching points falls. Calculate the total number of matching point pairs in three adjacent boxes, and use the matching point pairs in the three adjacent boxes with the largest total number as the filtered matching point pairs.
5. The posture optimization method according to claim 3, characterized in that: Determine whether object-level loss can be used as loss for optimization. If so, use object-level loss as loss for optimization, including: Input the rendered image and the original image into the Yolo-World model respectively to obtain the target boxes in the original image and the rendered image; For each category of target frames, first set T pairs of target frames to match, where the value range of T is T<=M and T<=S. The target frames in the original image and the rendered image are combined in pairs to form T pairs of target frames. The target frame pairs formed by all combinations are used as all target frame pairs corresponding to each T value. S is the number of target frames in each category in the original image, M is the number of target frames in each category in the rendered image, and T is an integer. After all T values in the value range are taken, determine whether all target frame pairs corresponding to each T value are retained, and calculate the ranking scores of all target frames corresponding to each T value. The number of all target frame pairs finally retained is counted. When the number is less than the preset fifth threshold, it is determined that object-level loss cannot be used as the loss for optimization; otherwise, object-level loss can be used as the loss for optimization. When key point loss is used as the loss for optimization, all target frame pairs corresponding to the maximum ranking score are screened out, and the Euclidean distance between the coordinate positions of the two target frame center points of each target frame pair is calculated, and the average of all Euclidean distances is used as the loss.
6. The posture optimization method according to claim 5, characterized in that: Then determine whether all target frame pairs corresponding to each T value are retained, and calculate the ranking scores of all target frames corresponding to each T value, including: For each T value, all the corresponding target frame pairs are analyzed; wherein, for each target frame pair, the target frame existing in the original image is recorded as the first target frame, and the target frame existing in the rendered image is recorded as the second target frame; firstly, the Euclidean distance between the coordinate position of the center point of the first target frame and the coordinate position of the center point of the second target frame in each target frame pair is calculated; the difference between the maximum value and the minimum value of all the distances corresponding to all the target frame pairs is recorded as the first value, the length of the smallest side of the original image is multiplied by the preset ratio, recorded as the second value, and the ratio between the second value and the first value is recorded as the first fraction; then, the horizontal and vertical coordinate axes are established with the center point of the first target frame in each target frame pair as the origin, and the center point of the second target frame is recorded as the first fraction. The center point is connected to the origin to form a line segment, the angle between the line segment and the horizontal axis is calculated, the difference between the maximum and minimum values of all angles corresponding to all target frame pairs is recorded as the third value, and the ratio between the preset degree and the third value is recorded as the second score; finally, the similarity of the color histogram vectors of the first target frame and the second target frame in each target frame pair is calculated to obtain multiple similarity calculation results and take the average of all similarity calculation results as the third score; when the first score is greater than the preset second threshold and the second score is greater than the preset third threshold and each similarity calculation result is greater than the preset fourth threshold, all target frame pairs corresponding to this T value are retained, and the first score, the second score and the third score are weighted averaged to obtain the ranking score.
7. The posture optimization method according to claim 1, characterized in that: The scene-level loss is used as the loss for optimization, including: The rendered image and the original image are input into the convolutional neural network respectively, and the output of the convolutional neural network is used as the feature vectors of the rendered image and the original image respectively. The cosine similarity between the two feature vectors is calculated and used as the loss.
8. A posture optimization device, characterized in that: include: The optimization module is used to set the round L of the back-propagation gradient descent algorithm during the tracking process of the SplaTAM system, and use the back-propagation gradient descent algorithm to optimize the camera's pose at the current moment to obtain the camera pose optimization result; wherein, multi-scale loss is used for optimization in the first K rounds, and depth and color loss is used for optimization in the last LK rounds. The multi-scale loss is key point loss, object-level loss or scene-level loss.
9. A posture optimization device, characterized in that: include: Memory for storing computer programs; A processor, used to implement the steps of the posture optimization method as described in any one of claims 1 to 7 when executing the computer program.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the posture optimization method according to any one of claims 1 to 7.