Three-dimensional space material anchoring method based on calibrated arc projection
Patent Information
- Application Number
- CN202610926093.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-18
AI Technical Summary
这种方式虽然不需要环境感知计算,但数字内容的位置与终端屏幕坐标系绑定而非与真实空间绑定,当用户移动终端时,数字内容随屏幕同步移动,无法保持在真实空间中的固定位置,不具备空间锚定能力
1、锚定点的确定过程为一次性计算:根据用户操作确定标定弧线的落点位置后,结合终端的定位数据和姿态数据计算出锚定点的世界坐标,此后不再对周围环境进行持续的特征提取、匹配或建模。后续每帧仅需根据终端的姿态变化对锚定点执行屏幕投影坐标的更新运算,该运算为单点坐标变换,计算量与持续运行的环境识别和空间建模相比降低了多个数量级。由此,终端的处理器负荷维持在较低水平,设备发热和性能衰减得到缓解,能够支持数小时的连续拍摄或直播使用。同时,由于不需要将环境特征数据上传至服务器端进行运算或存储,服务器的计算负担和存储负担不随用户数量线性增长,使得面向消费者的规模化部署成为可能。
Smart Images

Figure CN122780480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of augmented reality, and in particular to a method for anchoring three-dimensional spatial materials based on calibration arc projection. Background Technology
[0002] Augmented reality (AR) technology requires locating digital content at a target position in real space, so that when a user observes it on a terminal screen, the digital content can visually blend with the real scene. To achieve this, existing AR solutions generally rely on the perception and modeling of the real environment: first, the visual features and spatial structure information of the environment are collected through the terminal's cameras and sensors; then, a 3D model or feature map of the environment is built in the background; and finally, the position of the digital content is matched and bound to a specific area in the model.
[0003] This location-based approach, which relies on environmental modeling, suffers from a core contradiction that persists throughout the entire deployment, usage, and expansion process: the coupling between location accuracy and computational overhead. To accurately anchor digital content to its target location, the system needs to continuously extract, match, and update environmental features and coordinates. These calculations occur throughout the entire content display process. This continuous computation causes the terminal's processor to operate under high load for extended periods, leading to device overheating and performance degradation; on mobile devices, it typically only maintains stable operation for a few minutes. Simultaneously, some environmental feature data and model data need to be uploaded to the server for processing or storage. When multiple users access the system simultaneously, the server's computational and storage burden increases linearly with the number of users, making it difficult to support large-scale deployments for ordinary consumers.
[0004] The aforementioned contradictions extend further to the deployment stage. Before being used in each new scenario, the operator needs to collect environmental data on-site and complete the modeling in the background, and then return to the site for positioning, debugging, and matching verification. This process usually requires multiple trips, and the deployment cost increases exponentially with the number of scenarios.
[0005] Another approach bypasses environment modeling by directly overlaying digital content onto a fixed pixel area of the terminal screen. While this method eliminates the need for environment-aware computation, the position of the digital content is bound to the terminal screen's coordinate system rather than to real space. As the user moves the terminal, the digital content moves synchronously with the screen, failing to maintain a fixed position in real space and lacking spatial anchoring capabilities.
[0006] Furthermore, regardless of the approach described above, both solutions face challenges in depth perception when users target digital content to distant objects. The terminal screen presents a two-dimensional projection of three-dimensional space; points at different depths along the same line of sight are projected to nearly the same position on the screen. Users find it difficult to intuitively discern from the screen whether the actual depth of the digital content matches the distance to the target object. In summary, existing solutions lack a mechanism to assist users in accurately anchoring digital content without relying on environmental modeling. Summary of the Invention
[0007] To assist users in accurately anchoring digital content, this application provides a three-dimensional spatial material anchoring method based on calibration arc projection.
[0008] This application provides a method for anchoring three-dimensional spatial materials based on calibrated arc projection, which adopts the following technical solution: A method for anchoring three-dimensional spatial materials based on calibration arc projection, characterized by comprising: S1. In response to a first user operation, a calibration arc based on arc shape parameters is generated on the viewfinder of the terminal, and the landing position of the calibration arc is determined according to the first user operation; wherein, the distance between the starting point and the landing position of the calibration arc is the depth distance. S2. In response to the second user operation, lock the landing point of the calibration arc as the anchor point, and establish a fixed relative positional relationship between the anchor point and the position of the terminal; S3. Obtain the positioning data and attitude data of the terminal, and calculate the three-dimensional coordinates of the anchor point in the world coordinate system based on the positioning data, the attitude data and the fixed relative position relationship; S4. Retrieve the target material from the material library and render and display the target material at the screen projection position corresponding to the anchor point in the shooting view.
[0009] Optionally, the calibration arc is a parabolic trajectory.
[0010] Optionally, in S1, the first user operation includes determining the projection direction and projection force, and determining the landing position of the calibration arc based on the projection direction and the projection force.
[0011] Optionally, the projection force is determined by touch input performed by the user on the screen of the terminal, wherein the sliding distance or sliding speed of the touch input is positively correlated with the projection force.
[0012] Optionally, in S1, the first user operation includes specifying a landing point position in the shooting view, and the calibration arc is adaptively generated according to the specified landing point position.
[0013] Optionally, in step S1, the shooting viewfinder also includes a projection trigger option; in response to the user triggering the projection trigger option, the projection process animation of the calibration arc from the starting point to the landing point is dynamically rendered in the shooting viewfinder.
[0014] Optionally, in step S4, the rendering size of the target material is determined by perspective projection calculation based on the real-time distance between the anchor point and the terminal, and the rendering size is negatively correlated with the real-time distance.
[0015] Optionally, before performing S1, the method further includes: The current location data of the terminal is obtained, and the current location data is compared with electronic fence data, wherein the electronic fence data defines at least one prohibited projection area; in response to the current location data being located within the prohibited projection area, the execution of steps S1 to S4 is prohibited; the electronic fence data is pre-installed in the terminal and updated via network connection.
[0016] Optionally, after S4, the following may also be included: In response to the user's material switching command, a replacement material is retrieved from the material library. While keeping the three-dimensional coordinates of the anchor point and the fixed relative positional relationship unchanged, the content rendered and displayed at the screen projection position corresponding to the anchor point in the captured view is replaced from the target material with the replacement material.
[0017] Optionally, the arc shape parameter is an adjustable parameter, and the value of the arc shape parameter is negatively correlated with the maximum range of the calibration arc under the same operating amplitude.
[0018] Optionally, the three-dimensional coordinates of the anchor point and the fixed relative position relationship are cleared when the terminal exits the shooting mode.
[0019] Optional, also includes: The environment type of the terminal is determined based on the signal characteristics of the positioning data; in response to the environment type being an indoor environment, the arc shape parameter is set to a first value, so that the maximum range of the calibration arc under the same operating amplitude is the first range; in response to the environment type being an outdoor environment, the arc shape parameter is set to a second value, so that the maximum range of the calibration arc under the same operating amplitude is the second range; the first range is less than the second range.
[0020] Optionally, the following steps may also be included: S5. In response to the lateral displacement of the terminal relative to the line of sight, obtain the first displacement amount of the anchor point in the shooting view and the second displacement amount of the scene image of the area where the anchor point is located in the shooting view, determine the depth matching state according to the comparison result of the first displacement amount and the second displacement amount, and generate depth matching feedback information according to the depth matching state. S6. In response to the depth matching state being either too close or too far in depth, receive the user's depth adjustment input, update the depth distance of the anchor point according to the depth adjustment input, and recalculate the three-dimensional coordinates of the anchor point in the world coordinate system according to the updated depth distance; S7. In response to the change in the motion state of the terminal, recalculate the screen projection position corresponding to the anchor point based on the fixed relative position relationship.
[0021] Optionally, step S5 includes the following sub-steps: S51. Obtain the first displacement of the anchor point in the captured viewfinder; S52. Perform inter-frame displacement analysis on the scene image of the area where the anchor point is located in the captured viewfinder, and obtain the second displacement amount of the scene image; S53. Calculate the difference between the first displacement and the second displacement; in response to the first displacement being greater than the second displacement and the absolute value of the difference being greater than a preset threshold, determine the depth matching state as close in depth; in response to the first displacement being less than the second displacement and the absolute value of the difference being greater than the preset threshold, determine the depth matching state as far in depth; in response to the absolute value of the difference not being greater than the preset threshold, determine the depth matching state as matching in depth. S54. Generate the depth matching feedback information based on the depth matching status.
[0022] Optionally, the depth matching feedback information is presented in the shooting view through visual indicator elements around the anchor point, and the display state of the visual indicator elements changes according to the depth matching state.
[0023] Optionally, in S7, the recalculation of the screen projection position corresponding to the anchor point is performed when the terminal is in a state of rotation in place or when the cumulative displacement of the terminal is within a preset displacement threshold; in response to the cumulative displacement of the terminal exceeding the preset displacement threshold, a reprojection prompt message is generated.
[0024] Optionally, during the execution of S6, a reference grid is rendered at the screen projection position corresponding to the anchor point, and the display size of the reference grid is determined based on the real-time distance between the anchor point and the terminal and preset physical size parameters.
[0025] Optionally, in step S3, obtaining the location data of the terminal includes the following sub-steps: Detect the signal quality parameters of the currently available positioning signal source for the terminal; In response to the signal quality parameters meeting the preset quality conditions, satellite positioning data is used as the positioning data, and the three-dimensional coordinates of the anchor point in the world coordinate system are calculated based on the satellite positioning data. In response to the signal quality parameters not meeting the preset quality conditions, the positioning mode of the terminal is switched to a local positioning mode. In the local positioning mode, a local coordinate system is established with the position of the terminal at the switching time as the origin and the orientation of the terminal at the switching time as the reference direction. The local coordinate system replaces the world coordinate system in executing S3 and S4. In the local positioning mode, the positioning data is provided by the cumulative displacement increment output by the inertial measurement unit of the terminal.
[0026] Optionally, in the local positioning mode, a position correction step is also included: The terminal acquires signal characteristic data of wireless signal sources that can be received by the terminal, wherein the wireless signal sources include Wi-Fi access points and / or Bluetooth beacons, and the signal characteristic data includes round-trip time measurements of the Wi-Fi access points and / or received signal strength values of the Bluetooth beacons. The wireless positioning estimate of the terminal is calculated based on the signal feature data through polygon ranging or fingerprint matching. The wireless positioning estimate is weighted and fused with the cumulative displacement increment output by the inertial measurement unit to obtain the fused terminal position. The fused terminal position is then used to replace the terminal position provided by the cumulative displacement increment alone in the calculation of S3.
[0027] Optionally, in the local positioning mode, a drift compensation step is also included: After the anchor point is established, the cumulative drift estimate of the inertial measurement unit is obtained at a preset time interval. The cumulative drift estimate is determined based on the acceleration integral residual of the inertial measurement unit in adjacent correction cycles. In response to the cumulative drift estimate exceeding a preset drift threshold, a wireless positioning estimate based on the wireless signal source is triggered, and the current position of the terminal in the local coordinate system is reset according to the wireless positioning estimate. The coordinates of the anchor point in the local coordinate system are recalculated based on the reset terminal position and the fixed relative position relationship.
[0028] Optionally, a positioning mode reversal step can also be included: During the operation of the local positioning mode, the signal quality parameters are continuously monitored; In response to the signal quality parameters recovering to meet the preset quality conditions and lasting for a duration exceeding the preset stable duration, the positioning mode is switched from the local positioning mode back to the global positioning mode. During the switching process, satellite positioning data at the switching time is acquired. Based on the satellite positioning data at the switching time and the current position of the terminal in the local coordinate system, the mapped coordinates of the origin of the local coordinate system in the world coordinate system are calculated. Based on the mapped coordinates, the coordinates of the anchor point in the local coordinate system are converted into three-dimensional coordinates in the world coordinate system.
[0029] Optionally, in the local positioning mode, in response to the fact that the number of available wireless signal sources is less than a preset minimum number of signal sources, a sparse visual initialization step is performed: At the moment the local coordinate system is established, a preset number of sparse feature points are extracted from the current frame of the captured view, and the pixel coordinates and feature descriptors of the sparse feature points are recorded as initialization reference frame data. In the subsequent drift compensation step, in response to the cumulative drift estimate exceeding the preset drift threshold, sparse feature points are extracted again from the current frame of the captured view, and the sparse feature points of the current frame are matched with the initial reference frame data in one step. The position offset correction of the terminal relative to the coordinate origin is estimated based on the pixel coordinate offset of the matched point pair and the attitude change of the terminal. The position offset correction is used to replace the wireless positioning estimate to perform the position reset in the drift compensation step.
[0030] Optionally, an anchoring persistence step may be included after S2: In response to the user's save command, an anchoring record is generated and stored in the terminal's local storage or a remote server; the anchoring record includes the three-dimensional coordinates of the anchoring point, the fixed relative position relationship, the terminal positioning data and attitude data at the time the anchoring point is created, and environmental keyframe data; the environmental keyframe data includes the pixel coordinates and feature descriptors of effective feature points extracted from the image area within a preset pixel range centered on the screen projection position of the anchoring point in the captured viewfinder at the time the anchoring point is created.
[0031] Optionally, the generation of the environmental keyframe data includes the following sub-steps: At the moment the anchor point is created, feature point extraction is performed on the image region within a preset pixel range centered on the screen projection position of the anchor point in the captured view, and a candidate feature point set is obtained; Calculate the image gradient magnitude of the local neighborhood where each candidate feature point is located, remove candidate feature points whose image gradient magnitude is lower than a preset gradient threshold, and take the remaining candidate feature points as valid feature points. In response to the number of effective feature points being lower than the preset minimum number of feature points, a keyframe quality insufficient prompt message is displayed in the captured view, guiding the user to adjust the orientation of the terminal so that the area around the anchor point contains texture features before re-triggering the save command.
[0032] Optionally, an anchoring recovery step may also be included: In response to the user's recovery command, the following two phases are executed: In the first stage, the current positioning data of the terminal is obtained, and candidate anchor records whose distance between the three-dimensional coordinates and the current positioning data is within a preset search radius are retrieved from the stored anchor records based on the current positioning data. In the second stage, the same feature point extraction and filtering process as the environmental key frame data is performed on the current shooting view of the terminal to obtain the effective feature points of the current frame. The feature descriptor of the effective feature points of the current frame is compared with the feature descriptor of the effective feature points in the candidate anchoring record to calculate the descriptor distance. Feature point pairs with a descriptor distance less than a preset matching threshold are selected as matching point pairs. In response to the number of matching point pairs being greater than a preset minimum number of matches, the pose deviation of the terminal relative to the moment the anchor point was created is calculated based on the pixel coordinate correspondence of the matching point pairs. The current pose of the terminal is corrected based on the pose deviation, and the three-dimensional coordinates of the anchor point in the world coordinate system are recovered based on the corrected pose and the fixed relative position relationship in the candidate anchor record.
[0033] Optionally, in the anchoring recovery step, in response to the number of matching point pairs not being greater than the preset minimum matching number, a matching failure prompt message is generated and a thumbnail of the environmental keyframe corresponding to the candidate anchoring record is displayed in the shooting view, guiding the user to adjust the terminal to a pose close to the shooting angle and scene range in the thumbnail before re-triggering the second stage.
[0034] Optionally, the anchoring recovery step further includes a recovery confidence assessment: In response to the number of matching point pairs being greater than the preset minimum number of matches, the average descriptor distance of the matching point pairs is calculated; In response to the average descriptor distance being greater than a preset confidence threshold, the recovery confidence is determined to be low. After the anchor point is recovered, steps S5 to S6 are automatically triggered to verify and adjust the depth distance of the recovered anchor point. In response to the average descriptor distance not being greater than the preset confidence threshold, the recovery confidence is determined to be high confidence, and the anchor point is directly recovered.
[0035] Optionally, after the anchoring recovery step is completed, a keyframe update step is also included: In response to the low confidence level of the recovery and the user's confirmation that the position of the recovered anchor point is correct, environmental keyframe data is regenerated using an image region within a preset pixel range centered on the screen projection position of the anchor point in the current viewfinder of the terminal, replacing the original environmental keyframe data in the candidate anchor record.
[0036] Optionally, in the anchoring persistence step, in response to the existence of multiple anchoring points in the same shooting session, the anchoring record also includes the relative position vector between each anchoring point; In the anchoring recovery step, in response to the candidate anchoring record containing multiple anchoring points, the second stage of pose correction is performed on the anchoring point with the highest feature matching confidence. The pose correction result is used as a reference to recover the three-dimensional coordinates of the remaining anchoring points in batches according to the relative position vector, so that the spatial relative relationship between each anchoring point remains consistent with that when it was saved.
[0037] In summary, this application includes at least one of the following beneficial technical effects: 1. The anchor point determination process is a one-time calculation: After determining the landing position of the calibration arc based on user operations, the world coordinates of the anchor point are calculated by combining the terminal's positioning and attitude data. Afterward, no further feature extraction, matching, or modeling of the surrounding environment is performed. Each subsequent frame only requires updating the screen projection coordinates of the anchor point based on the terminal's attitude changes. This is a single-point coordinate transformation, reducing the computational load by several orders of magnitude compared to continuous environmental recognition and spatial modeling. As a result, the terminal's processor load remains low, mitigating device overheating and performance degradation, enabling several hours of continuous shooting or live streaming. Furthermore, since environmental feature data does not need to be uploaded to the server for processing or storage, the server's computational and storage burden does not increase linearly with the number of users, making large-scale deployment for consumers possible.
[0038] 2. The anchor point is bound to the world coordinate system rather than the terminal screen coordinate system through a fixed relative position relationship. When the terminal shakes or rotates, the position of the anchor point in real space does not change with the movement of the terminal. Only its screen projection position is updated with the change of the terminal posture, so that the target material displayed in the rendering is visually kept in a fixed position in real space.
[0039] 3. The arc trajectory of the calibrated arc simultaneously encodes information in two dimensions on the shooting view: projection direction and depth distance. The extension direction of the arc corresponds to the projection direction, and the length and height of the arc from the starting point to the landing point correspond to the depth distance. This allows users to intuitively perceive the position of the landing point in three-dimensional space from the shape of the arc, reducing the difficulty of depth perception when performing three-dimensional spatial positioning on a two-dimensional screen.
[0040] 4. By comparing the displacement of the anchor point with the displacement of the scene image when the terminal undergoes lateral displacement, the principle of parallax is used to detect whether the distance between the anchor point and the target object matches without relying on a depth sensor. Combined with depth adjustment input and reference grid, a three-stage depth calibration process of coarse adjustment-verification-fine adjustment is formed. Attached Figure Description
[0041] Figure 1 This is a flowchart illustrating the overall process of a three-dimensional spatial material anchoring method based on calibration arc projection in one embodiment of the present invention.
[0042] Figure 2 This is a schematic diagram of the interface for marking arcs on the viewfinder screen captured by the terminal in one embodiment of the present invention.
[0043] Figure 3 This is a schematic diagram of the spatial relationship between the terminal, the anchor point, and the world coordinate system in one embodiment of the present invention.
[0044] Figure 4 This is a flowchart of the depth matching verification step in one embodiment of the present invention.
[0045] Figure 5 a is a parallax comparison diagram when the depth matching state is close in depth in one embodiment of the present invention.
[0046] Figure 5 b is a parallax comparison diagram when the depth matching state is longitudinal matching in one embodiment of the present invention.
[0047] Figure 5 c is a parallax comparison diagram when the depth matching state is remote in one embodiment of the present invention.
[0048] Figure 6 a is a schematic diagram of the interface between the reference grid and the visual indicator element in a state of near depth in one embodiment of the present invention.
[0049] Figure 6 b is a schematic diagram of the interface between the reference grid and the visual indicator element in a depth matching state according to an embodiment of the present invention. Detailed Implementation
[0050] The present application will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of the application and are not intended to limit the scope of the application.
[0051] Reference Figure 1 This application discloses a method for anchoring 3D spatial materials based on calibration arc projection. This method can be applied in application environments where the terminal and server communicate via a network. The terminal, also known as a client, refers to a program carrier that provides local services to the user. It can be an application installed on a smartphone, tablet, laptop, or portable wearable device, or a third-party mini-program embedded with other applications. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. It is used to store material resources, electronic fence data, and user account information, and provides video streaming services to the terminal in cloud mode. In offline mode, the terminal pre-downloads the materials to local storage, and the projection and playback process does not depend on the network connection with the server.
[0052] Before detailing the embodiments of this application, some terms will be explained first.
[0053] The calibration arc is an arc-shaped trajectory generated on the viewfinder of the device, used to calibrate the target position in three-dimensional space. The calibration arc has a starting point and an ending point. The starting point corresponds to the device's position in the image, and the ending point corresponds to the spatial position of the digital material the user wants to anchor. The distance between the starting and ending points is the depth distance. The shape of the calibration arc is controlled by arc shape parameters, which determine the degree of curvature and range of the arc.
[0054] An anchor point is a virtual spatial point with fixed three-dimensional coordinates in the world coordinate system, used to support and position the projected digital material. The anchor point maintains a fixed relative positional relationship with the terminal; when the terminal moves, the anchor point's position in real space remains unchanged.
[0055] The media library is a media management module within the terminal that stores pre-downloaded local video files or cloud video links for retrieval during projection.
[0056] Reference Figure 2 In step S1, in response to the first user operation, a calibration arc based on the arc shape parameters is generated on the terminal's viewfinder, and the landing point of the calibration arc is determined according to the first user operation. After the terminal starts shooting mode, the camera captures real-time images of the real scene and displays them on the screen, forming the viewfinder. When the user performs the first user operation on the viewfinder, the system renders a calibration arc in real time on the viewfinder according to the parameters of the user operation and the current arc shape parameters.
[0057] The calibrated arc trajectory simultaneously encodes information in two dimensions on the captured viewfinder: projection direction and depth. The direction of the arc's extension corresponds to the projection direction, while the arc length and height from the starting point to the landing point correspond to the depth distance. Compared to straight rays that only provide directional information or bullseye markers that only indicate planar positions, the calibrated arc allows users to intuitively perceive the landing point's location in three-dimensional space from its shape. The first user operation can include, but is not limited to, the following: the user slides their finger along a specific direction on the screen to determine the projection direction and force, or adjusts the projection angle and distance using a virtual joystick on the screen, or sets depth parameters by dragging a slider on the screen, or directly clicks on the target location in the captured viewfinder to specify the landing point.
[0058] For example, a user is standing on the street opposite a building and wants to anchor video footage onto the building's wall. After the user performs the first action, an arc appears in the viewfinder, curving from a starting point slightly below the center of the screen to a point on the building's wall. The longer and higher the arc, the farther the landing point will be. The user can adjust the parameters to change the shape of the arc until the landing point is at the desired target location.
[0059] The trajectory shape of the calibration arc is not limited to a single type. In different embodiments, the calibration arc can be a parabolic trajectory, a Bézier curve trajectory, or a spline curve trajectory. A Bézier curve trajectory defines the arc shape through several control points, and the user can drag the control points to adjust the curvature and landing position of the arc. A spline curve trajectory generates a smooth arc by fitting multiple interpolation nodes, which is suitable for scenarios requiring more complex trajectory shapes.
[0060] In some embodiments, the calibration arc is a parabolic trajectory. The parabolic trajectory is controlled by a term in the arc's shape parameters that is similar to the square of gravitational acceleration, giving the arc natural rising and falling segments. Compared to Bézier curves and spline curves, the parabolic trajectory has the following characteristics that make it a preferred embodiment for calibration arcs.
[0061] First, the landing point of a parabola is uniquely determined by the initial direction and force parameters. Users only need to control two intuitive input quantities to complete the positioning, without needing to understand or manipulate additional control points or interpolation nodes. The shape of a Bézier curve is affected by the position of the control points, and users need to adjust the control points to accurately change the landing point, making the operation more complex than that of a parabola.
[0062] Second, parabolas, constrained by their arc shape parameters, have a definite maximum range, naturally limiting the projection range to a reasonable spatial distance and preventing users from setting anchor points at unavailable distances. The extension range of Bézier curves and spline curves depends on the configuration of control points or nodes, lacking a built-in range constraint mechanism.
[0063] Third, parabolic trajectories exhibit high density of landing points at close range and low density at long range. Within the close range, small adjustments to operating parameters result in minimal changes to the landing point, leading to high positioning accuracy. However, at long range, larger adjustments are required to alter the landing point's position. This non-linear operational characteristic naturally aligns with the rendering requirements of perspective projection, which demands precise positioning at close range and allows for larger errors at greater distances. In contrast, the operational response characteristics of Bézier curves and spline curves depend on the distribution of control points and do not inherently possess this adaptive accuracy distribution.
[0064] Under the same operating range, the larger the arc shape parameter, the steeper the arc of the parabola and the closer the landing point, and the higher the fine adjustment accuracy within the close range; the smaller the arc shape parameter, the gentler the arc of the parabola and the farther the landing point, and the more distant the target can be reached.
[0065] In the embodiment where the calibration arc is a parabolic trajectory, step S1 involves the first user operation determining the projection direction and projection force, and then determining the landing position of the calibration arc based on the projection direction and projection force. The projection force is determined by the user's touch input on the terminal's screen, and the sliding distance or speed of the touch input is positively correlated with the projection force. For example, if the user quickly slides their finger along a certain direction on the screen, the sliding direction determines the projection direction, and the sliding speed determines the projection force. The faster the sliding speed, the greater the projection force, the farther the range of the calibration arc, and the farther the landing position is from the starting point. The user can also control the projection force using a virtual joystick or slider on the screen.
[0066] In an embodiment where the calibration arc is a Bézier curve trajectory, the first user operation includes dragging control points of the arc in the captured viewfinder. The system recalculates the shape and landing point of the Bézier curve in real time based on the position of the control points. In an embodiment where the calibration arc is a spline curve trajectory, the first user operation includes sequentially adding or dragging interpolation nodes in the captured viewfinder. The system fits the node sequence to generate a smooth spline arc and determines the landing point position.
[0067] In other embodiments, in step S1, the first user operation includes directly specifying the landing point position in the captured viewfinder, and the calibration arc is adaptively generated based on the specified landing point position. When the user directly specifies the landing point, the calibration arc generated by the system can help the user to verify whether the depth of the specified landing point is reasonable—if the generated arc shape does not match the user's intuitive perception of spatial distance (e.g., the arc is abnormally flat or abnormally steep), the user can determine that the specified landing point position needs to be corrected. For example, if the user clicks on a location on a building wall in the captured viewfinder, the system takes that clicked location as the landing point, calculates in reverse, and automatically generates a calibration arc extending from the starting point to the landing point.
[0068] Parabolic trajectories offer the following advantages in assisting users in calibrating anchor points. In existing augmented reality positioning solutions, users typically specify an anchor position by clicking a point or dragging a marker on the screen. These operations are performed only on the two-dimensional plane of the screen and fail to convey depth information about the target location. Users must guess or adjust the depth distance using other methods, resulting in a lack of intuitive spatial feedback.
[0069] The parabolic trajectory encodes depth information within the geometric shape of the arc. The arc height and length monotonically change with depth, allowing users to continuously receive visual feedback on the depth dimension from the changes in the arc's shape when adjusting operational parameters. When the landing point is closer, the arc is steeper and shorter; when the landing point is farther away, the arc is gentler and longer. This continuous geometric mapping allows users to intuitively perceive the depth position of the landing point from the screen without relying on numerical displays.
[0070] Furthermore, the arc of a parabola covers a continuous spatial area in the image from the starting point to the landing point. The area the arc passes through visually overlaps with reference objects in the real scene, such as buildings and the ground. The user's visual system can use the arc as a depth scale from near to far, utilizing the occlusion and overlap between the arc and scene references to help confirm the spatial location of the landing point. A straight ray, on the other hand, only appears as a directional indicator on the image and lacks this spatial reference capability that unfolds along the depth direction.
[0071] In some embodiments, step S1 further includes a projection trigger option on the viewfinder. In response to the user triggering the projection trigger option, an animation of the projection process of the calibration arc from the starting point to the landing point is dynamically rendered in the viewfinder.
[0072] Compared to a statically displayed complete arc, a dynamically rendered projection animation provides users with an additional dimension for depth verification. In static display mode, the complete calibration arc is presented on the screen at the same time, and users can only infer the landing point from the final shape of the arc, but cannot intuitively perceive the "landing" process of the arc when it reaches the landing point—whether the landing point just reaches the target surface, falls before the target surface, or crosses the target surface and continues to extend further away.
[0073] In dynamic rendering mode, the arc is drawn segment by segment from its starting point, progressing along its trajectory towards the landing point. As the leading edge of the arc approaches the target object in the frame, the user can observe the spatial relationship between the arc's leading edge and the target object: if the leading edge of the arc begins to fall and disappears before reaching the target object's position in the frame, it indicates the landing point is too close, and the arc has not reached the target surface; if the leading edge of the arc continues to extend after passing the target object's position in the frame, it indicates the landing point is too far, and the arc has crossed the target surface; if the leading edge of the arc converges and stops precisely at the target object's position in the frame, it indicates the landing point matches the target surface. This sequential visual process utilizes the human eye's ability to track motion trajectories, allowing the user to intuitively predict the landing point position during the arc's flight, thus deciding whether to adjust operation parameters before locking the anchor point.
[0074] Step S2: In response to the second user's operation, the landing point of the calibration arc is locked as the anchor point, establishing a fixed relative positional relationship between the anchor point and the terminal's position. (Refer to...) Figure 3 Once the user confirms the landing location, a second user action is performed, and the system fixes the current landing location as the anchor point. The second user action can include, but is not limited to, the following: the user clicks the confirmation button in the viewfinder, performs a double-tap gesture at the landing location, performs a long-press gesture at the landing location, or issues a lock confirmation via voice command. The relative positional relationship between the anchor point and the terminal is recorded at this moment as a fixed offset vector. This offset vector includes three components: the azimuth angle, pitch angle, and depth distance of the anchor point relative to the terminal, or equivalently, a three-dimensional offset in the terminal's local coordinate system.
[0075] The fixed relative positional relationship records the spatial offset of the anchor point relative to the terminal at the moment of locking. This offset information is entirely given by the landing direction and depth distance of the calibration arc in step S1, without relying on the identification or modeling of the surrounding environment. Once established, the fixed relative positional relationship remains unchanged. When the position or attitude of the terminal changes in subsequent steps, the system can continuously calculate the spatial position of the anchor point based on this fixed offset vector. This means that the positioning result completed by the user in step S1 through the calibration arc will not be lost due to the user's subsequent movement or rotation of the terminal—the position information of the anchor point has been independently saved in the form of an offset vector, decoupled from the current motion state of the terminal.
[0076] In some embodiments, the 3D coordinates and fixed relative positions of the anchor points are cleared when the terminal exits the shooting mode and are not persistently stored. Each use is a temporary projection, and the data disappears after exiting. The terminal does not need to maintain a historical anchor point database, and the server does not need to save spatial coordinate information for each user. The temporary lifecycle design means that the system does not need to maintain a persistent spatial anchor database on the terminal or the server. The storage overhead of a single shooting session is only one offset vector (three floating-point numbers), which is consistent with the deployment requirements for large-scale consumer users—each user's anchoring operation does not generate data that needs to be stored for a long time, and the server's storage pressure does not increase with the number of users.
[0077] Step S3: Acquire the terminal's positioning data and attitude data. Based on the positioning data, attitude data, and fixed relative position relationship, calculate the three-dimensional coordinates of the anchor point in the world coordinate system. The positioning data is acquired by the terminal's built-in GPS module, reflecting the terminal's current position in the world coordinate system. The attitude data is acquired by the terminal's built-in inertial measurement unit (IMU), which includes a gyroscope and an accelerometer, reflecting the terminal's current orientation and tilt angle. In each frame, the system jointly calculates the positioning data and attitude data with the fixed relative position relationship established in step S2: first, it constructs the terminal's attitude matrix in the world coordinate system based on the attitude data; then, it transforms the offset vector in the fixed relative position relationship from the terminal's local coordinate system to the world coordinate system through this attitude matrix; finally, it superimposes the transformed offset vector onto the terminal's world coordinates determined by the positioning data to obtain the three-dimensional coordinates of the anchor point in the world coordinate system. In the embodiments of this application, the world coordinate system adopted is the WGS-84 geographic coordinate system.
[0078] Through the calculation in step S3, the fixed relative positional relationship established in step S2 is transformed into the absolute coordinates of the anchor point in the world coordinate system. This makes the anchor point no longer just an offset relative to the terminal, but a fixed spatial position bound to the world coordinate system. When the terminal shakes or rotates, the position of the anchor point in real space does not change with the movement of the terminal; only its projected position on the screen updates with the change of the terminal's posture. For example, after the user sets the anchor point on the opposite building wall, even if the user holds the terminal and rotates it left or right, the anchor point always remains in the same position on the building wall, remaining relatively stationary with reference points such as windows and corners on the real building, as if there were a real physical object at that position.
[0079] This calculation process is a one-time establishment and continuous reuse model. The fixed relative positional relationship, determined once in step S2, remains unchanged. Each subsequent frame only needs to update the world coordinates of the anchor point based on the terminal's latest positioning and attitude data. This update operation is a single-point coordinate transformation and does not involve environmental feature extraction, matching, or modeling. For example, in conventional augmented reality solutions, the system needs to continuously perform feature point detection, descriptor calculation, and spatial map updates on the camera footage. These operations consume a significant amount of processor resources per frame. In this solution, the computational load per frame is only one matrix multiplication and one vector addition, reducing the computational load by several orders of magnitude compared to continuously running environment recognition. Therefore, the terminal's processor load remains at a low level, enabling several hours of continuous shooting or live streaming.
[0080] When the GPS positioning of the terminal drifts, the world coordinates of the anchor point are calculated based on the terminal's positioning data superimposed with a fixed offset vector. The calculated position of the anchor point drifts by the same amount as the terminal's position, and the relative relationship between the two remains unchanged. On the screen, the anchor point's projected position relative to the terminal does not shift due to GPS drift, and the visually anchored content remains stable.
[0081] Existing digital content overlay solutions mainly employ two approaches, each with its own limitations. The first approach fixes the image to screen pixel coordinates, such as overlaying a decorative pattern at a fixed position on the captured image. In this method, the image always remains within the same pixel area of the screen. When the user moves or rotates the device, the image moves synchronously with the screen, failing to maintain a fixed position in real space. Regardless of where the user points the device, the image always occupies the same area of the screen, lacking spatial anchoring capabilities.
[0082] The second approach involves anchoring textures to visual features in the image, such as anchoring virtual glasses to the eye area of a face in live video streaming. The system continuously detects feature points in the image to update the texture's position, allowing the texture to follow the feature points. This method maintains good tracking performance when feature points are consistently visible, but it suffers from the following problems: when the tracked features are occluded, deformed, or move away from the image, the texture may jump, jitter, or detach. For example, when virtual glasses are anchored to a face, a user quickly turning their head can cause facial feature point detection to fail, resulting in the virtual glasses drifting or briefly disappearing from the image; when the user blinks, the features in the eye area change, causing the virtual glasses to jitter. Furthermore, this method relies on continuous recognition of image content, requiring feature detection and matching operations to be performed every frame, increasing the computational load with the complexity of the image and the number of features.
[0083] The anchoring mechanism in this solution differs from the two methods mentioned above. The anchor point is bound to its absolute position in the world coordinate system through a fixed relative positional relationship, independent of screen pixel coordinates and visual features within the image. The screen projection position of the anchor point is calculated solely based on the terminal's own positioning and posture data, unaffected by changes in image content, feature occlusion, or lighting conditions. When the terminal shakes, the anchor point remains stationary in real space; changes in the scene captured by the camera do not interfere with the anchor point's position calculation.
[0084] Step S4: Retrieve the target material from the material library and render it at the screen projection position corresponding to the anchor point in the captured view. The system transforms the world coordinates of the anchor point calculated in step S3 to the screen pixel coordinate system through the terminal's current pose matrix and perspective projection matrix to obtain the screen projection position of the anchor point in the captured view. Then, it retrieves the target material selected by the user from the material library and overlays the rendered image of the target material onto the screen projection position.
[0085] In some embodiments, in step S4, the rendering size of the target material is determined by perspective projection calculation based on the real-time distance between the anchor point and the terminal, and the rendering size is negatively correlated with the real-time distance. For example, when the anchor point is 5 meters away from the terminal, the display width of a video clip on the screen is 200 pixels; when the anchor point is 10 meters away from the terminal, the display width of the same video clip shrinks to approximately 100 pixels. This perspective scaling, where objects appear larger when closer and smaller when farther away, makes the display effect of the material conform to the human eye's perception of the size of objects in real space.
[0086] In some embodiments, the media library includes two types of media: local media and cloud media. Local media consists of video files that users pre-download from app stores or content platforms to their device's local storage; the retrieval and playback process does not generate network traffic. Cloud media consists of video files stored on a remote server, which are transmitted to the device in real-time via a network connection in a streaming manner for playback. Since the calculation and rendering of anchor points only involve single-point coordinate transformations and do not consume additional network bandwidth, both local and cloud media playback can run smoothly under low computational load.
[0087] In some embodiments, a material switching step is included after step S4. In response to the user's material switching command, a replacement material is retrieved from the material library. While maintaining the three-dimensional coordinates and fixed relative position of the anchor point, the content rendered and displayed at the screen projection position corresponding to the anchor point in the captured view is replaced from the target material with the replacement material. For example, the user first plays an advertisement video at the anchor point position, and then selects an animation material from the material library to replace it. After replacement, the animation material appears in the exact same spatial position as the advertisement video, and the coordinates and offset relationship of the anchor point remain unchanged.
[0088] In some embodiments, before executing step S1, an electronic fence determination step is included: acquiring the terminal's current location data, comparing the current location data with electronic fence data, whereby the electronic fence data defines at least one prohibited projection area. In response to the current location data being located within the prohibited projection area, steps S1 to S4 are prohibited from execution. The electronic fence data is pre-installed in the terminal and updated via a network connection. For example, the operator of a landmark building reaches a cooperation agreement with a platform provider, designating a 200-meter radius around the building as a dedicated operating area. The platform provider marks this area as a prohibited projection area in the electronic fence data and pushes it to all terminals via the network. When an ordinary user enters this area and opens the application, the system detects that the current location is within the prohibited projection area and does not activate the calibration arc generation and anchor point projection functions.
[0089] It is important to note that the electronic fence determination is performed before step S1, constituting a pre-interception mechanism. Since the electronic fence data is pre-stored locally on the terminal, the determination process only requires comparing the current GPS coordinates with the local data; no query requests to the server are needed, resulting in low latency and no network bandwidth consumption. The electronic fence data is updated in a non-real-time, periodic manner, decoupled from the real-time projection process.
[0090] In some embodiments, the arc shape parameter is an adjustable parameter, and its value is negatively correlated with the maximum range of the calibration arc under the same operating amplitude. By adjusting the arc shape parameter, the user or system can switch between prioritizing accuracy and prioritizing range. For example, in indoor scenarios, where the target is typically within a range of 3 to 15 meters, setting the arc shape parameter to a larger value results in a shorter range for the calibration arc under the same operating amplitude, minimizing the change in the landing point during user fine-tuning and improving positioning accuracy. In outdoor scenarios, where the target may be tens or even hundreds of meters away, setting the arc shape parameter to a smaller value increases the range, allowing the user to reach distant building facades.
[0091] In some embodiments, the system can automatically switch the arc shape parameter. The system determines the environment type of the terminal based on the signal characteristics of the positioning data. In response to an indoor environment, the arc shape parameter is set to a first value, making the maximum range of the calibrated arc under the same operating amplitude the first range. In response to an outdoor environment, the arc shape parameter is set to a second value, making the maximum range of the calibrated arc under the same operating amplitude the second range. The first range is shorter than the second range. For example, when the GPS signal strength is lower than a preset signal threshold and the positioning accuracy factor exceeds a preset accuracy threshold, the system determines that the terminal is in an indoor environment and automatically switches the arc shape parameter to an indoor value, compressing the range to within 15 meters. When the GPS signal strength is higher than a preset signal threshold, the system determines that the terminal is in an outdoor environment and automatically switches the arc shape parameter to an outdoor value, extending the range to over 100 meters.
[0092] In the above embodiments, the system distinguishes between indoor and outdoor environments and adjusts the arc shape parameters based on the signal characteristics of GPS signals, but the GPS module still participates in the calculation of step S3 as the source of positioning data. In some scenarios, such as underground spaces, tunnels, or signal-shielded indoor areas, GPS signals are not only weakened but completely unusable, and the terminal cannot obtain satellite positioning data. To enable the anchoring method to still operate in these scenarios, in some embodiments, when obtaining the terminal's positioning data in step S3, the system switches between global positioning mode and local positioning mode based on the signal quality parameters of the currently available positioning signal source for the terminal.
[0093] Signal quality parameters include the signal-to-noise ratio (SNR) and positioning accuracy factor of the satellite positioning signal. The system sets preset quality conditions: the SNR is not lower than a preset SNR threshold, and the positioning accuracy factor does not exceed a preset accuracy factor threshold. When the signal quality parameters meet the preset quality conditions, the system uses satellite positioning data as the positioning data, and the coordinates of the anchor point are calculated in the WGS-84 world coordinate system, in the same way as in step S3 of the aforementioned embodiment. When the signal quality parameters do not meet the preset quality conditions, the system switches the positioning mode to local positioning mode.
[0094] For example, the preset signal-to-noise ratio threshold is 20dB-Hz, and the preset accuracy factor threshold is 6. After a user walks from outdoors into an underground parking lot, the GPS signal's signal-to-noise ratio drops from 35dB-Hz to 12dB-Hz, which is below the 20dB-Hz threshold. The system determines that the signal quality parameters do not meet the preset quality conditions and triggers a switch from global positioning mode to local positioning mode.
[0095] In local positioning mode, the system establishes a local coordinate system with the terminal's position at the time of switching as the origin and the terminal's orientation at the time of switching as the reference direction. The three axes of the local coordinate system correspond to the terminal's forward, right, and vertically upward directions, respectively, and are determined by the attitude data output by the inertial measurement unit at the time of switching. After the local coordinate system is established, it replaces the world coordinate system as the calculation reference for steps S3 and S4. The terminal's positioning data is no longer provided by the GPS module, but by the cumulative displacement increment output by the inertial measurement unit. The inertial measurement unit obtains the terminal's displacement relative to the origin by performing a double integration on the accelerometer output, and obtains the terminal's rotation relative to the reference direction through the gyroscope output.
[0096] The core logic of the anchoring mechanism is consistent between the local and global positioning modes. The anchor point is still determined by the calibration arc in step S1, a fixed relative position relationship is established in step S2, and the anchor point coordinates are calculated based on the positioning and attitude data in step S3. The calculation of the offset vector, the construction of the attitude matrix, and the updating of the screen projection position all follow the calculation process in the previous embodiments. The only difference is that the reference of the coordinate system changes from the WGS-84 world coordinate system to a local coordinate system with the terminal as the origin, and the source of the positioning data changes from the GPS module to the cumulative displacement increment of the inertial measurement unit.
[0097] For example, when a user opens the application in an underground exhibition hall, GPS signal is unavailable, and the system automatically enters local positioning mode. The terminal's current location is set as the origin (0, 0, 0), and the direction the terminal is facing is set as the front of the local coordinate system. The user performs a calibration arc projection operation towards an exhibit 8 meters in front, and the system establishes a fixed relative positional relationship offset by 8 meters along the front. Subsequently, the user rotates the terminal 90 degrees to the left. The inertial measurement unit detects the 90-degree rotation, and the system transforms the coordinates of the anchor point to the screen projection coordinates using the rotated attitude matrix. The anchor point moves from the center of the screen to the right edge of the screen, and the visual effect is consistent with the global positioning mode.
[0098] When local positioning mode relies solely on positioning data provided by the inertial measurement unit (IMU), a cumulative drift problem exists. The IMU's accelerometer and gyroscope introduce minute measurement noise in each sample. After integration, this noise is amplified, causing the deviation between the calculated and actual positions to increase over time. In short periods, such as tens of seconds to minutes, the cumulative drift is typically in the centimeter to decimeter range, with an acceptable impact on anchoring accuracy. However, as usage time increases to over ten minutes, the cumulative drift can reach the meter level, resulting in a noticeable shift in the anchor point's position on the screen.
[0099] To mitigate cumulative drift, in some embodiments, wireless signal source-assisted correction is introduced in the local positioning mode. The terminal detects receivable wireless signal sources in the current environment, including Wi-Fi access points and Bluetooth beacons. The round-trip time measurement of the Wi-Fi access point reflects the physical distance between the terminal and the access point, while the received signal strength of the Bluetooth beacon is negatively correlated with the distance between the terminal and the beacon. When there are three or more wireless signal sources with known locations in the environment, the system calculates the terminal's wireless positioning estimate using a polygon ranging algorithm. When the location of the wireless signal sources in the environment is unknown but a signal fingerprint database is available, the system compares the currently received signal features with pre-collected fingerprints in the database using a fingerprint matching algorithm to obtain the terminal's wireless positioning estimate.
[0100] The accuracy of wireless positioning estimates is typically between 1 and 5 meters, lower than that of GPS in open outdoor environments, but better than the drift accumulated by an inertial measurement unit (IMU) alone after long-term operation. The system weightedly fuses the wireless positioning estimate with the cumulative displacement increment output by the IMU. The IMU has high short-term accuracy but large long-term drift, while wireless positioning has high short-term noise but no cumulative drift; their error characteristics are complementary. The fusion weights are dynamically adjusted based on the IMU's operating time since the last calibration: the shorter the operating time, the higher the weight of the IMU; the longer the operating time, the higher the weight of the wireless positioning. The fused terminal position replaces the terminal position provided by the IMU alone in the calculation of step S3.
[0101] In some embodiments, the system does not perform wireless positioning estimation in every frame, but instead triggers drift compensation periodically at preset time intervals. The preset time interval can be 30 seconds or 60 seconds, set according to the nominal drift rate of the inertial measurement unit (IMU). At the end of each correction cycle, the system acquires the cumulative drift estimate of the IMU. The cumulative drift estimate is determined based on the acceleration integral residual of the IMU in adjacent correction cycles: theoretically, the IMU should output zero or constant acceleration when the terminal is stationary or moving at a constant speed; the cumulative amount of the deviation between the actual output and the theoretical value, after integration, is the cumulative drift estimate.
[0102] When the cumulative drift estimate exceeds a preset drift threshold, the system triggers a wireless positioning estimate based on the wireless signal source and resets the terminal's current position in the local coordinate system according to the wireless positioning estimate. After the reset, the coordinates of the anchor point in the local coordinate system are recalculated based on the reset terminal position and the fixed relative position relationship. After the reset is completed, the cumulative drift estimate returns to zero, and the measurement restarts in the next correction cycle.
[0103] For example, the preset drift threshold is 0.5 meters, and the preset time interval is 30 seconds. After establishing an anchor point in the indoor exhibition hall, the user used the system for 3 minutes. At the 90th second, the cumulative drift estimate of the inertial measurement unit (IMU) reached 0.6 meters, exceeding the 0.5-meter threshold. The system triggered a Wi-Fi round-trip time ranging (RTT) measurement, acquiring distance data from four surrounding Wi-Fi access points. Using a polygonal ranging algorithm, the system calculated the terminal's wireless positioning estimate as 3.2 meters to the right and 5.1 meters forward from the origin. The IMU then cumulatively calculated the terminal's position as 3.0 meters to the right and 5.5 meters forward. The system reset the IMU's position to the fused value and updated the anchor point coordinates accordingly.
[0104] In some scenarios, the number of available wireless signal sources in the terminal's environment is insufficient to perform multilateral ranging or fingerprint matching. For example, in enclosed warehouses without Wi-Fi coverage, construction sites, or large cave spaces, the number of available wireless signal sources may be less than three. In these scenarios, the system employs sparse visual initialization as an alternative correction method for drift compensation.
[0105] Sparse visual initialization performs feature acquisition once when the local coordinate system is established: the system extracts a preset number of sparse feature points from the current frame of the captured image, such as 100 corner features, and records the pixel coordinates and feature descriptors of these feature points as initialization reference frame data. The initialization reference frame data is only acquired once when the coordinate system is established, and no continuous feature tracking or environment modeling is performed thereafter.
[0106] At the subsequent drift compensation trigger moment, i.e., when the cumulative drift estimate exceeds the preset drift threshold, the system extracts sparse feature points again from the currently captured view and performs a one-time match between the feature descriptor of the current frame and the feature descriptor in the initialized reference frame data. The matching result gives the pixel position offset of the feature points between the current frame and the reference frame. Combining the terminal's attitude change between the two frames, the system estimates the translation direction and relative scale of the terminal relative to the origin through essential matrix factorization. Since monocular vision lacks an absolute scale reference, the absolute value of the translation is scaled by the cumulative displacement length of the inertial measurement unit within the correction period. The aligned displacement estimate replaces the wireless positioning estimate to perform position reset in drift compensation.
[0107] The computational complexity of sparse visual correction is comparable to that of the inter-frame displacement analysis performed on the region where the anchor point is located in step S52. Step S52 performs sparse feature point tracking on a local image patch of 100×100 pixels around the anchor point, while sparse visual correction extracts about 100 corner points from the entire image and performs descriptor matching. Both are local operations on a limited number of feature points and do not involve dense feature calculations or 3D map construction of the entire image.
[0108] It should be understood that the effectiveness of sparse visual correction depends on the fact that feature points in the reference frame are still visible in the current frame. When a user turns around and moves to a completely different scene area after establishing an anchor point, there may be no shared feature points between the current frame and the reference frame. In this case, the system skips the current visual correction, relies on the inertial measurement unit to maintain positioning, and waits for the user to return to a position with a shared viewing area with the reference frame before retrying in the next correction cycle. In addition, similar to the inter-frame displacement analysis in step S52, sparse visual correction cannot extract enough effective feature points in environments lacking texture, such as solid-color walls or uniform ground. When the number of extracted feature points is lower than the preset minimum number of feature points, the system also skips the current correction and relies on the inertial measurement unit to maintain positioning.
[0109] In some embodiments, the system continuously monitors the signal quality parameters of the satellite positioning signal during local positioning mode operation. When the signal quality parameters recover to meet preset quality conditions and the duration exceeds a preset stabilization period, the system switches the positioning mode from local positioning mode back to global positioning mode. The preset stabilization period is used to avoid frequent switching caused by instantaneous fluctuations in signal quality, such as in scenarios where the GPS signal is intermittent at a building entrance. The preset stabilization period can be 5 seconds or 10 seconds.
[0110] When switching back to global positioning mode, the system needs to convert the anchor point's coordinates from the local coordinate system to three-dimensional coordinates in the WGS-84 world coordinate system. The conversion process is as follows: First, acquire the satellite positioning data at the time of switching, which gives the terminal's current position in the world coordinate system. Simultaneously, acquire the terminal's current position in the local coordinate system, which is given by the cumulative displacement increment of the inertial measurement unit after wireless or visual correction. Based on the correspondence between these two positions, calculate the mapped coordinates of the origin of the local coordinate system in the world coordinate system. The mapped coordinates are equal to the satellite positioning coordinates at the time of switching minus the world coordinate system component of the terminal's displacement vector in the local coordinate system after attitude matrix transformation. After obtaining the mapped coordinates, superimpose the anchor point's coordinates in the local coordinate system onto the mapped coordinates to obtain the anchor point's three-dimensional coordinates in the world coordinate system.
[0111] For example, after a user returns to the ground floor from the underground exhibition hall via escalator, if the GPS signal signal-to-noise ratio remains above 30dB-Hz for 8 seconds, exceeding the preset stabilization time of 5 seconds, the system switches back to global positioning mode. At the time of the switch, the GPS positioning data shows the terminal is located at 31.2345°N, 121.4567°E. The terminal's position in the local coordinate system is 15 meters forward and 3 meters to the right, a position already calibrated twice with Wi-Fi assistance. Based on this, the system calculates the mapped coordinates of the local coordinate system origin in the world coordinate system and converts the anchor point coordinates from the local coordinate system to the world coordinate system. After the conversion, subsequent positioning data recovery is provided by the GPS module, and the anchor point coordinate updates are re-executed based on the WGS-84 world coordinate system. Before and after the switch, the anchor point's projection position on the screen in the captured viewfinder does not change, and the user is visually unaware of the coordinate system switching process.
[0112] The automatic switching of the arc shape parameters also applies in the aforementioned local positioning mode. When the system enters local positioning mode, it can simultaneously switch the arc shape parameters to indoor values, compressing the range to a suitable indoor distance. When switching back to global positioning mode, the arc shape parameters revert to outdoor values. The switching of arc shape parameters and the switching of positioning modes are triggered independently, but they usually occur simultaneously when entering and exiting indoor scenes. (Refer to...) Figures 4 to 6 In some embodiments, the method further includes depth matching verification steps S5 to S7. Since the terminal screen presents a two-dimensional projection of three-dimensional space, spatial points at different depth positions along the same viewing direction are projected onto approximately the same pixel position on the screen. Users can hardly distinguish whether the depth of the anchor point matches the actual distance to the target object based solely on a static image. Steps S5 to S7 utilize the parallax principle to provide users with depth verification and adjustment capabilities without relying on depth sensors and environmental modeling.
[0113] Step S5: In response to the lateral displacement of the terminal relative to the line of sight, the first displacement of the anchor point in the captured view and the second displacement of the scene image of the area where the anchor point is located in the captured view are obtained. The depth matching state is determined based on the comparison result of the first displacement and the second displacement, and depth matching feedback information is generated based on the depth matching state.
[0114] Reference Figure 5The principle of step S5 is as follows. When the terminal undergoes lateral displacement, objects at different depths in the image experience different displacements—objects closer to the terminal experience larger displacements, while objects farther from the terminal experience smaller displacements. The anchor point, as a virtual object, has its displacement in the image calculated by the system based on the anchor point's depth and the terminal's displacement; this is the first displacement. The scene image of the area where the anchor point is located is a real scene captured by the camera; its displacement is obtained through inter-frame displacement analysis; this is the second displacement. If the depth of the anchor point matches the actual distance to the target object in that area of the scene, such as... Figure 5 As shown in b, the first displacement and the second displacement should be approximately equal. If the anchor points are too close, such as... Figure 5 As shown in diagram a, its parallax is greater than that of the target object, and the first displacement is greater than the second displacement. If the anchor point is remote, such as... Figure 5 As shown in c, its parallax is smaller than that of the target object, and the first displacement is smaller than the second displacement.
[0115] In some embodiments, the specific calculation steps of step S5 are as follows.
[0116] Step S51: Obtain the first displacement of the anchor point in the captured viewfinder.
[0117] Step S52: Perform inter-frame displacement analysis on the scene image of the area where the anchor point is located in the captured view to obtain the second displacement amount of the scene image.
[0118] Step S53: Calculate the difference between the first displacement and the second displacement. If the first displacement is greater than the second displacement and the absolute value of the difference is greater than a preset threshold, determine the depth matching state as "depth too close". If the first displacement is less than the second displacement and the absolute value of the difference is greater than the preset threshold, determine the depth matching state as "depth too far". If the absolute value of the difference is not greater than the preset threshold, determine the depth matching state as "depth matched".
[0119] The principle behind the parallax comparison for depth detection is as follows: When the terminal moves laterally, objects at different depths in the image experience varying pixel displacements due to distance differences. Objects closer to the terminal experience larger displacements, while those farther away experience smaller displacements—this is the fundamental physical law of motion parallax. The first displacement is a deterministic value directly calculated by the system based on the set depth of the anchor point and the actual displacement of the terminal, eliminating measurement errors. The second displacement is the actual pixel displacement of the real scene in the area where the anchor point is located in the camera's view, caused by the terminal's movement, reflecting the true depth of the target object. When the set depth of the anchor point matches the true depth of the target object, the pixel displacements caused by the same segment of terminal displacement should be approximately equal. When they do not match, the direction of the difference directly indicates the direction of the deviation, and the magnitude of the difference reflects the degree of deviation. Users do not need to understand the parallax principle itself; the system converts the comparison results into visual indicators on the screen, and users only need to adjust the depth based on these visual indicators.
[0120] The inter-frame displacement analysis in step S52 can employ a local optical flow estimation algorithm. Specifically, a local image patch, such as a 100×100 pixel region, is extracted around the screen projection position corresponding to the anchor point. Sparse feature point tracking or template matching operations are performed on this region in both preceding and following frames to obtain the average pixel displacement of the scene image within this region as the second displacement. This operation is performed only on the local region near the anchor point, without performing dense optical flow calculations on the entire image, maintaining the same computational complexity as the lightweight computational design of the independent weighted scheme. It should be understood that local optical flow estimation relies on the existence of sufficient texture features in the image region for tracking. When the area aligned with the anchor point is a solid-color wall, sky, or uniform glass curtain wall, this region lacks traceable texture gradients, and the optical flow estimation may not yield a reliable displacement.
[0121] In some embodiments, the system calculates the image gradient magnitude of the local region before performing inter-frame displacement analysis. If the image gradient magnitude is lower than a preset gradient threshold, the system skips the disparity detection step and prompts the user to use a reference grid for depth verification. Furthermore, the preset threshold plays a role in absorbing measurement errors in the decision logic of step S53. During lateral movement, the handheld terminal inevitably generates image noise and motion blur, which can cause slight fluctuations in the output value of the optical flow estimation around the actual displacement. The preset threshold is set to a value higher than this fluctuation range, ensuring that the system only outputs a determination of depth proximity or distance when the difference between the first and second displacements clearly exceeds the measurement noise range, thus avoiding misjudging noise as depth mismatch.
[0122] For example, the preset threshold is 3 pixels. After the terminal moves 10 centimeters to the side, the anchor point moves 15 pixels in the image, and the building wall in the area where the anchor point is located moves 9 pixels in the image. The difference is 6 pixels, which is greater than the preset threshold of 3 pixels. Moreover, the first displacement is greater than the second displacement. The system determines that the depth matching status is too close in depth.
[0123] Step S54: Generate depth matching feedback information based on the depth matching state. In some embodiments, the depth matching feedback information is presented in the shooting view through visual indicator elements around the anchor point, and the display state of the visual indicator elements changes according to the depth matching state. For example, when the depth matching state is depth matching, such as... Figure 6 As shown in b, a solid circle is displayed around the anchor point; when the depth matching status is too close, a dashed circle is displayed around the anchor point, with an outward-pointing arrow indicator outside the circle, prompting the user to adjust the anchor point further away; when the depth matching status is too far away, as shown in b. Figure 6 As shown in Figure a, a dashed circle is displayed around the anchor point, with an inward-pointing arrow indicator on the outside of the circle, prompting the user to adjust the anchor point closer.
[0124] In close-range scenarios, such as when the anchor point is 5 to 15 meters from the terminal, the displayed size of the anchored content on the screen is relatively large, and visual flaws caused by depth deviations are easily noticed by users. For example, if a user anchors video footage onto a door 3 meters away, and the depth deviation of the anchor point is 0.5 meters, the alignment between the edge of the video and the door frame will be noticeably misaligned when the user shakes the phone, making it clear that the footage is not attached to the door. In such scenarios, high depth accuracy is required, and even a small lateral movement of the terminal can produce sufficient parallax difference. For example, if the anchor point is set at 5 meters and the target object is at 6 meters, a 10-centimeter lateral movement of the terminal results in an approximate displacement of 32 × arctan(0.1 / 5) ≈ 6.4 pixels for the anchor point and approximately 32 × arctan(0.1 / 6) ≈ 5.3 pixels for the target object, a difference of about 1.1 pixels. When the depth deviation increases to more than 1 meter, the difference rapidly increases to several pixels or more, exceeding the preset threshold, and the system can reliably detect the depth mismatch. In close-range scenarios requiring high precision, the parallax signal is strong enough, and the effectiveness of optical flow detection matches the perception requirements.
[0125] In long-distance scenarios, such as when the anchor point is more than 50 meters away from the terminal, the displayed size of the anchored content on the screen is small, possibly only a few dozen pixels wide. A 5-meter depth deviation of the anchor point has a very weak visual impact on the image—at 5 meters forward and 5 meters back, the difference in the projection position and display size of the material on the screen is not noticeable to the user. When the user shakes the phone, the relative movement between the distant object and the anchor point is very small, so even if the depth is not completely accurate, it will not visually reveal any obvious flaws. At this distance, the parallax difference caused by the same amplitude of lateral displacement of the terminal drops to the sub-pixel level, and the output accuracy of optical flow detection is insufficient to reliably distinguish between depth matching and depth deviation. However, the decrease in perception requirements and the decrease in detection accuracy occur simultaneously: at distances where high precision is not required, insufficient detection accuracy does not affect the user experience.
[0126] In distant scenes, there remains a boundary situation requiring precise alignment: users want to precisely embed footage into a specific feature of a distant building, such as aligning a video clip exactly within a window frame. In this case, the accuracy of parallax detection is insufficient to meet the alignment requirements, but a reference mesh can compensate. During depth adjustment, the user observes the proportional relationship between the displayed size of the reference mesh and the actual size of the window frame, visually confirming the accuracy of the depth. The verification of the reference mesh relies on the user's visual experience rather than pixel-level image calculations, and is unaffected by parallax signal attenuation.
[0127] Therefore, parallax detection covers the accuracy verification requirements in near and medium-range scenarios, while the reference grid covers the alignment verification requirements in long-range scenarios. These two verification methods complement each other to form depth calibration coverage across the entire range. It should be understood that the accuracy of parallax detection is related to the amount of lateral displacement of the terminal and the depth distance of the anchor point. In near and medium-range scenarios, a small lateral displacement of the terminal is sufficient to generate enough parallax difference for the system to detect. In long-range scenarios, the parallax difference generated by the same amount of lateral displacement decreases, and the detection reliability correspondingly decreases. In long-range scenarios, users can use visual comparison with the reference grid as an alternative verification method.
[0128] Step S6: In response to the depth matching status being either too close or too far in depth, receive the user's depth adjustment input, update the depth distance of the anchor point according to the depth adjustment input, and recalculate the three-dimensional coordinates of the anchor point in the world coordinate system based on the updated depth distance. (Refer to...) Figure 6 The user performs depth adjustment operations based on depth matching feedback information, inputting the depth adjustment amount via a slider on the screen or touch gestures. The system increases or decreases the depth distance of the anchor point along the projection direction vector and updates the world coordinates of the anchor point accordingly.
[0129] In some embodiments, during step S6, a reference grid is rendered at the screen projection position corresponding to the anchor point. The display size of the reference grid is determined based on the real-time distance between the anchor point and the terminal and preset physical size parameters. The reference grid represents a square grid of preset physical size, such as 1 meter × 1 meter. When the depth of the anchor point is exactly the same as the distance to the target object, the display size of the reference grid on the screen should appear the same size as an area of the same physical size on the target object. For example, the user projects the anchor point onto the wall of an opposite building, where each window is approximately 1.5 meters wide. The preset physical size of the reference grid is 1 meter × 1 meter. If the display width of the reference grid is approximately two-thirds of the window width, the depth is basically correct. If the reference grid is significantly wider than the window, the anchor point is too close, and the reference grid is magnified due to the principle of near objects appearing larger than distant ones. If the reference grid is significantly narrower than the window, the anchor point is too far away. The user adjusts the depth slider accordingly until the ratio between the reference grid and the window is reasonable.
[0130] Step S7: In response to changes in the terminal's motion state, the screen projection position corresponding to the anchor point is recalculated based on a fixed relative position relationship. When the terminal rotates or undergoes a small displacement, the system obtains the latest attitude data of the terminal from the inertial measurement unit every frame, transforms the world coordinates of the anchor point to the screen pixel coordinates of the current frame through perspective projection, and updates the display position of the anchor point and the materials on it in the image. This operation is a standard 3D rendering pipeline operation, and the computational load is comparable to the projection transformation of a single 3D point.
[0131] In some embodiments, in step S7, the recalculation of the screen projection position corresponding to the anchor point is performed when the terminal is rotating in place or when the cumulative displacement of the terminal is within a preset displacement threshold. In response to the cumulative displacement of the terminal exceeding the preset displacement threshold, a reprojection prompt is generated. For example, the preset displacement threshold is 2 meters. When the user rotates the phone in place or moves slightly within a 2-meter range, the anchor point remains stable. When the user walks more than 2 meters, the system displays a prompt on the viewfinder, guiding the user to re-perform the calibration arc projection at a new location to establish a new anchor point. This design stems from technical constraints: this solution maintains anchoring based on the relative positional relationship between GPS and the inertial measurement unit (IMU). The cumulative drift error of the IMU increases with the displacement distance, and the anchoring accuracy becomes unacceptable beyond a certain range.
[0132] In the aforementioned embodiments, the three-dimensional coordinates and fixed relative positions of the anchor points are cleared when the terminal exits the shooting mode, and each use is a temporary projection. In other embodiments geared towards commercial operation scenarios, operators need to repeatedly use the same anchor points at the same physical location, such as repeatedly playing the same promotional video daily at a fixed booth, or continuously projecting guided tour content at multiple viewpoints in a scenic area. For this purpose, the system provides an optional anchor persistence function.
[0133] In response to the user's save command, the system generates an anchoring record and stores it in the terminal's local storage or a remote server. The anchoring record includes the following data: the three-dimensional coordinates of the anchor point, the fixed relative positional relationship, the terminal's positioning and attitude data at the time of anchor point creation, and environmental keyframe data. The three-dimensional coordinates of the anchor point, the fixed relative positional relationship, and the terminal's positioning and attitude data are all data already calculated in steps S2 and S3, and can be directly read during saving without incurring additional computational overhead. The environmental keyframe data is new data introduced by the anchoring persistence function, used to assist in restoring the precise location of the anchor point during user revisits.
[0134] The generation process of environmental keyframe data is as follows. At the anchor point creation time, the system performs feature point extraction on the image region within a preset pixel range centered on the screen projection position of the anchor point in the captured viewfinder. The preset pixel range can be 300×300 pixels or 500×500 pixels, covering the local scene around the anchor point. Feature point extraction uses corner detection algorithms, such as ORB (Oriented Fast and Rotated BRIEF) or SIFT (Scale-Invariant Feature Transform), to obtain a set of candidate feature points. Each candidate feature point contains pixel coordinates and a feature descriptor, which is a compact mathematical representation of the local image texture surrounding the feature point.
[0135] The system performs quality screening on candidate feature points. It calculates the image gradient magnitude of the local neighborhood of each candidate feature point and removes those with gradient magnitudes below a preset threshold. A low gradient magnitude indicates weak texture variation in the region where the feature point is located, such as a solid-color wall or a uniform sky; these feature points are prone to mismatches in subsequent matching. The remaining candidate feature points after screening are considered valid feature points, and their pixel coordinates and feature descriptors constitute the feature description information of the environmental keyframe data.
[0136] The image gradient magnitude filtering logic described above is consistent with the texture check before inter-frame displacement analysis in step S52. Step S52 calculates the image gradient magnitude of the region where the anchor point is located before performing optical flow estimation; if the gradient magnitude is lower than a preset gradient threshold, disparity detection is skipped. The generation of environmental keyframe data uses the same gradient evaluation mechanism, performing the same culling operation on feature points with insufficient texture.
[0137] When the number of valid feature points is lower than the preset minimum, for example, less than 15, it indicates that the scene texture around the anchor point is insufficient to support subsequent visual matching. The system displays a keyframe quality insufficiency warning in the captured view, guiding the user to adjust the orientation of the terminal to include richer texture features in the area around the anchor point before re-triggering the save command. For example, if the user projects the anchor point onto a white wall with almost no texture, the system will extract only 6 valid feature points, less than the preset minimum of 15. The system will prompt the user to slightly adjust the terminal angle to include the door frame or corner lines of the wall edge in the image before saving. After the user adjusts, the image includes the edge texture of the door frame, increasing the number of valid feature points to 42, meeting the save condition. The system then completes the keyframe data generation and writes it to the anchor record.
[0138] The storage location of anchor records depends on the usage scenario. In personal user scenarios, anchor records are stored in the terminal's local storage, without consuming server resources. In commercial operation scenarios, anchor records are stored on a remote server. The operator can centrally manage anchor records from multiple locations in a backend management system and distribute them to multiple terminals for use by different users. The data size of a single anchor record is mainly determined by the number of feature points in the environmental keyframe data. Taking 50 valid feature points as an example, with each feature point's pixel coordinates occupying 8 bytes and the ORB feature descriptor occupying 32 bytes, the environmental keyframe data of a single anchor record is approximately 2KB. Including the 3D coordinates, offset vector, and terminal pose data, the total does not exceed 5KB. This data size is far smaller than the storage space of a complete image frame and does not constitute a storage burden on either the terminal's local storage or the server side.
[0139] When a user returns to the area where the anchor point was previously saved and wishes to restore the anchor, the system performs an anchor restoration step. Anchor restoration consists of two stages: the first stage is coarse localization retrieval, and the second stage is a one-time visual matching. The combination of these two stages forms a two-stage funnel, from narrowing down the spatial range to precise pose correction.
[0140] In the first phase, the system acquires the terminal's current location data. This data can come from a GPS module (outdoor scenarios) or a wireless location estimate from the aforementioned local positioning mode (indoor scenarios). Based on the current location data, the system retrieves candidate anchor records from the stored anchor records whose 3D coordinates are within a preset search radius of the current location data. The preset search radius can be 50 meters or 100 meters, depending on the accuracy of the location data source. GPS positioning accuracy is typically in the range of 5 to 10 meters, and Wi-Fi positioning accuracy is typically in the range of 1 to 5 meters. The preset search radius must be greater than the positioning error to avoid omissions. When multiple candidate anchor records are retrieved, the system arranges them from closest to furthest, prioritizing the closest one.
[0141] The first stage aims to narrow the search scope from all anchor records stored on the terminal to a few near the current location, avoiding the computational overhead of performing visual matching on every single record. For example, the terminal locally stores 200 anchor records distributed across different cities and locations. The user is currently located in a commercial plaza, with GPS indicating a location of 31.2345°N, 121.4567°E. The system uses a 100-meter search radius to select three candidate anchor records within 100 meters of the 200 records.
[0142] In the second stage, the system performs the same feature point extraction and filtering process on the current captured view from the terminal as it does when generating environmental keyframe data, obtaining valid feature points for the current frame. The system then calculates the descriptor distance between the feature descriptors of the valid feature points in the current frame and the feature descriptors of the valid feature points in the candidate anchor records. Descriptor distance measures the degree of local texture similarity between two feature points; a smaller distance indicates that the two feature points are more likely to correspond to the same visual feature at the same physical location. The system selects feature point pairs with a descriptor distance less than a preset matching threshold as matching point pairs.
[0143] When the number of matching point pairs exceeds the preset minimum number of matches, for example, more than 8 pairs, the system calculates the pose deviation of the terminal relative to the moment the anchor point was created, based on the pixel coordinate correspondence of the matching point pairs. The pose deviation includes the translation component between the terminal's current position and its position at the moment of creation, and the rotation component between the terminal's current orientation and its orientation at the moment of creation. The calculation method is as follows: the essential matrix is estimated from the pixel coordinates of the matching point pairs; the essential matrix is decomposed to obtain the rotation matrix and translation direction; the absolute scale of the translation is provided by the distance between the current positioning data and the positioning data at the moment of creation in the anchor record. The system corrects the terminal's current pose based on the pose deviation and recovers the three-dimensional coordinates of the anchor point in the world coordinate system based on the corrected pose and the fixed relative positional relationship in the candidate anchor records.
[0144] The second stage of visual matching is a one-time operation. The system performs feature extraction and descriptor matching only once when restoring the anchor point. After restoration, no further feature tracking or matching updates are performed on the screen. The restored anchor point is exactly the same as the normally created anchor point. Subsequent screen projection updates still rely solely on the terminal's positioning and attitude data, and the calculation mode reverts to the single-point coordinate transformation in steps S3 and S4.
[0145] For example, last week the operator created an anchor point on an exterior wall of a commercial plaza and saved the anchoring record. This week, the operator returned to the plaza, opened the application, and triggered a recovery command. In the first stage, the GPS showed that the current location was 35 meters away from the saved anchoring record, within a 100-meter search radius, and the record was listed as a candidate. In the second stage, the operator pointed the terminal towards the exterior wall, and the system extracted 58 valid feature points from the current image and matched them with the 46 valid feature points saved in the anchoring record, obtaining 22 matching point pairs, which is greater than the preset minimum matching number of 8 pairs. Based on this, the system calculated the deviation between the terminal's current pose and the pose at the time of creation, corrected it, and restored the world coordinates of the anchor point. After the recovery was completed, the promotional video on the exterior wall appeared in the same position in the image as last week.
[0146] Visual matching fails when the number of matching point pairs is less than the preset minimum number of matches. Matching failure may be caused by the following reasons: the user's current shooting direction is too different from the creation time, resulting in no shared viewing area between the current image and the scene in the keyframe; the scene has changed significantly, such as the billboard on the exterior wall being replaced or temporary structures being demolished, causing the feature points in the keyframe to no longer exist in the current scene; the lighting conditions are too different, such as keyframes saved during the day being restored at night, resulting in a decrease in the similarity of feature descriptors.
[0147] In some embodiments, when a match fails, the system displays a thumbnail of the environmental keyframe corresponding to the candidate anchor record in the viewfinder. The thumbnail is a scaled-down version of the image area surrounding the anchor point captured at the creation time, overlaid in a corner of the screen. By observing the scene content and shooting angle in the thumbnail, the user adjusts the terminal to a pose close to the shooting angle and scene range in the thumbnail, and then re-triggers the second stage of visual matching. This method of guiding users to self-correct via thumbnails does not introduce additional algorithmic complexity; the correction accuracy relies on the user's visual judgment and is suitable for situations where matching fails but the scene itself has not fundamentally changed.
[0148] In some embodiments, the anchoring recovery step further includes a recovery confidence assessment. When the number of matched point pairs exceeds a preset minimum number of matches, the matching quantity meets the recovery condition, but the matching quality still varies. The system calculates the average descriptor distance of the matched point pairs. The average descriptor distance reflects the overall texture similarity of the matched point pairs: the smaller the average descriptor distance, the higher the matching quality, and the more reliable the calculation result of the pose deviation.
[0149] When the average descriptor distance exceeds a preset confidence threshold, the system determines the recovered confidence level to be low. Low confidence means that although a sufficient number of matching point pairs have been found, the quality of these matches is not high, and there may be some mismatches. The calculated pose deviation may contain certain errors, and the depth of the recovered anchor point may deviate from the actual target position. In this case, the system automatically triggers the depth matching verification process in steps S5 to S6 after restoring the anchor point. The user confirms whether the depth of the anchor point matches the target object through disparity detection in step S5 and performs fine-tuning through depth adjustment in step S6. The execution method of the depth matching verification process is exactly the same as the verification after normal anchor point creation, requiring no additional algorithms or interaction design.
[0150] When the average descriptor distance is not greater than a preset confidence threshold, the system determines the recovery confidence level to be high and directly recovers the anchor point without triggering depth matching verification. High confidence level indicates that the texture similarity of the matched point pair is high, the calculation result of the pose deviation is reliable, and the distance between the recovered anchor point depth and the target object is accurate enough.
[0151] For example, the preset confidence threshold is a descriptor distance of 40 (ORB descriptor distance ranges from 0 to 256). In one anchor recovery, the system obtains 18 matching point pairs with an average descriptor distance of 28, which is less than the preset confidence threshold of 40. The system determines this as high confidence and directly recovers the anchor points. In another recovery, some features in the scene have changed appearance due to seasonal changes (leaves change from green to yellow). The system obtains 12 matching point pairs with an average descriptor distance of 52, which is greater than the preset confidence threshold of 40. The system determines this as low confidence, and after recovering the anchor points, it automatically enters step S5 for depth matching verification. Visual indicator elements appear on the screen to prompt the user to move the terminal to trigger parallax detection.
[0152] In some embodiments, after low-confidence recovery is completed and the user confirms the recovered anchor point position is correct, the system performs a keyframe update. The system regenerates environmental keyframe data using an image region within a preset pixel range centered on the screen projection position of the anchor point in the currently captured viewfinder, replacing the original environmental keyframe data in the candidate anchor records. The updated keyframes reflect the scene appearance at the current moment, providing higher-quality matching results in the next recovery.
[0153] Keyframe updates are triggered only after low-confidence recovery, not after high-confidence recovery. This is because high confidence indicates a good match between the original keyframe and the current scene, requiring no update; low confidence indicates an appearance difference between the original keyframe and the current scene. Without updates, subsequent recovery attempts will face the same low-confidence problem, potentially leading to matching failure as the scene continues to change. By continuously updating keyframes after low-confidence recovery, the system's environmental keyframe data always remains consistent with the scene appearance at the time of the most recent successful recovery, forming an adaptive update mechanism that gradually follows scene changes.
[0154] In some embodiments, when multiple anchor points exist in the same shooting session, the user can choose to save the multiple anchor points as a group in batches. During batch saving, in addition to containing the three-dimensional coordinates, fixed relative position relationships, and environmental keyframe data for each anchor point, the anchor record also records the relative position vectors between each anchor point. The relative position vector refers to the difference in three-dimensional coordinates between any two anchor points in the world coordinate system, which is obtained directly by subtracting the known coordinates of each anchor point during saving.
[0155] During batch recovery, the system does not perform the complete two-stage recovery process for each anchor point individually. Instead, it adopts a strategy of baseline anchor recovery plus relative position calculation. First, the system matches the environmental keyframe data of all anchor points with the current image, selecting the anchor point with the highest matching confidence as the baseline anchor point. A complete second-stage pose correction is then performed on the baseline anchor point. After correction, the system uses the relative position vectors saved in the anchor records, starting from the recovered 3D coordinates of the baseline anchor point, to batch calculate the 3D coordinates of the remaining anchor points using vector addition.
[0156] This strategy ensures that the spatial relative relationships between each anchor point are completely consistent with those at the time of saving. If visual matching and pose correction are performed independently for each anchor point, the errors of each matching are independent, which may lead to deviations in the relative positions between the restored anchor points. For example, if the distance between two anchor points is 5 meters at the time of saving, the distance may become 4.6 meters or 5.3 meters after independent restoration. By using a single reference point and calculating the remaining points using the relative vector at the time of saving, all anchor points share the same pose correction result, and their relative positions are guaranteed by the accurate data at the time of saving, unaffected by the independent accumulation of matching errors.
[0157] For example, the operator saves three anchor points A, B, and C on a walking path in a scenic area, located at the entrance, a pavilion in the middle section, and a viewing platform at the end of the path, respectively. At the time of saving, the relative position vector from A to B is offset 80 meters east and 15 meters north, and the relative position vector from B to C is offset 120 meters east and 5 meters south. During the restoration process, the operator stands near the pavilion, with the terminal screen facing the pavilion. The system matches the keyframes of the three anchor points with the current screen. B has the most matching point pairs and the smallest average descriptor distance, and is selected as the reference anchor point. The system performs a complete pose correction on B, restoring B's world coordinates. Subsequently, based on the saved relative position vectors, the world coordinates of A are equal to the world coordinates of B minus the offset of 80 meters east and 15 meters north, and the world coordinates of C are equal to the world coordinates of B plus the offset of 120 meters east and 5 meters south. All three anchor points are restored in a single pose correction operation, and their spacing is consistent with that at the time of saving.
[0158] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0159] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0160] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for anchoring three-dimensional spatial materials based on calibration arc projection, characterized in that, include: S1. In response to a first user operation, a calibration arc based on arc shape parameters is generated on the viewfinder of the terminal, and the landing position of the calibration arc is determined according to the first user operation; wherein, the distance between the starting point and the landing position of the calibration arc is the depth distance. S2. In response to the second user operation, lock the landing point of the calibration arc as the anchor point, and establish a fixed relative positional relationship between the anchor point and the position of the terminal; S3. Obtain the positioning data and attitude data of the terminal, and calculate the three-dimensional coordinates of the anchor point in the world coordinate system based on the positioning data, the attitude data and the fixed relative position relationship; S4. Retrieve the target material from the material library and render and display the target material at the screen projection position corresponding to the anchor point in the shooting view.
2. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 1, characterized in that, The calibration arc is a parabolic trajectory; in S1, the first user operation includes determining the projection direction and projection force, and determining the landing position of the calibration arc based on the projection direction and the projection force. The projection force is determined by the touch input performed by the user on the screen of the terminal, and the sliding distance or sliding speed of the touch input is positively correlated with the projection force.
3. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 1, characterized in that, In S1, the first user operation includes specifying the landing point position in the shooting view, and the calibration arc is adaptively generated according to the specified landing point position. And / or, in S1, the shooting viewfinder also includes a projection trigger option; in response to the user triggering the projection trigger option, the projection process animation of the calibration arc from the starting point to the landing point is dynamically rendered in the shooting viewfinder; And / or, in S4, the rendering size of the target material is determined by perspective projection calculation based on the real-time distance between the anchor point and the terminal, and the rendering size is negatively correlated with the real-time distance.
4. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 1, characterized in that, Before executing S1, the following is also included: The current location data of the terminal is obtained, and the current location data is compared with electronic fence data, wherein the electronic fence data defines at least one prohibited projection area; in response to the current location data being located within the prohibited projection area, the execution of steps S1 to S4 is prohibited; the electronic fence data is pre-installed in the terminal and updated via network connection. And / or, following S4, it further includes: In response to the user's material switching command, a replacement material is retrieved from the material library. While keeping the three-dimensional coordinates of the anchor point and the fixed relative positional relationship unchanged, the content rendered and displayed at the screen projection position corresponding to the anchor point in the captured view is replaced from the target material with the replacement material.
5. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 1, characterized in that, The arc shape parameter is an adjustable parameter, and the value of the arc shape parameter is negatively correlated with the maximum range of the calibration arc under the same operating amplitude. And / or, the three-dimensional coordinates of the anchor point and the fixed relative position relationship are cleared when the terminal exits the shooting mode.
6. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 5, characterized in that, Also includes: The environment type of the terminal is determined based on the signal characteristics of the positioning data; In response to the environment type being an indoor environment, the arc shape parameter is set to a first value, such that the maximum range of the calibration arc under the same operating amplitude is the first range; in response to the environment type being an outdoor environment, the arc shape parameter is set to a second value, such that the maximum range of the calibration arc under the same operating amplitude is the second range; the first range is less than the second range.
7. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 1, characterized in that, It also includes the following steps: S5. In response to the lateral displacement of the terminal relative to the line of sight, obtain the first displacement amount of the anchor point in the shooting view and the second displacement amount of the scene image of the area where the anchor point is located in the shooting view, determine the depth matching state according to the comparison result of the first displacement amount and the second displacement amount, and generate depth matching feedback information according to the depth matching state. S6. In response to the depth matching state being either too close or too far in depth, receive the user's depth adjustment input, update the depth distance of the anchor point according to the depth adjustment input, and recalculate the three-dimensional coordinates of the anchor point in the world coordinate system according to the updated depth distance; S7. In response to the change in the motion state of the terminal, recalculate the screen projection position corresponding to the anchor point based on the fixed relative position relationship.
8. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 7, characterized in that, S5 includes the following sub-steps: S51. Obtain the first displacement of the anchor point in the captured viewfinder; S52. Perform inter-frame displacement analysis on the scene image of the area where the anchor point is located in the captured viewfinder, and obtain the second displacement amount of the scene image; S53. Calculate the difference between the first displacement and the second displacement; in response to the first displacement being greater than the second displacement and the absolute value of the difference being greater than a preset threshold, determine the depth matching state as close in depth; in response to the first displacement being less than the second displacement and the absolute value of the difference being greater than the preset threshold, determine the depth matching state as far in depth; in response to the absolute value of the difference not being greater than the preset threshold, determine the depth matching state as matching in depth. S54. Generate the depth matching feedback information based on the depth matching status.
9. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 7, characterized in that, The depth matching feedback information is presented in the shooting view through visual indicator elements around the anchor point, and the display state of the visual indicator elements changes according to the depth matching state.
10. The method for anchoring three-dimensional spatial materials based on calibration arc projection according to claim 7, characterized in that, In step S7, the recalculation of the screen projection position corresponding to the anchor point is performed when the terminal is in a stationary rotation state or the cumulative displacement of the terminal is within a preset displacement threshold; in response to the cumulative displacement of the terminal exceeding the preset displacement threshold, a reprojection prompt message is generated. During the execution of S6, a reference grid is rendered at the screen projection position corresponding to the anchor point. The display size of the reference grid is determined based on the real-time distance between the anchor point and the terminal and preset physical size parameters.