A mobile phone real scene photo ranging and deviation correcting cultural wall modeling design method and system

CN122695103APending Publication Date: 2026-09-04ANHUI TSUEN SHUITING CULTURAL CREATIVITY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610928465.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

[0004]针对现有技术存在的缺乏统一表征和动态追踪多因素非线性耦合关系的度量框架,导致测距纠偏无法实现跨设备、跨时间、跨工况一致性校准的核心问题,本申请通过一种手机实景照片测距纠偏的文化墙建模设计方法、系统,实现基于统一度量参考系和纠偏映射模型对多因素耦合关系进行建模,并以前置去抖纠偏层与度量参考系联动进行运动模糊校正,从而提升文化墙三维建模的几何精度、模型一致性和现场采集效率

Benefits of technology

[0016] This application provides a method and system for modeling and designing cultural walls using mobile phone real-scene photo ranging and correction. By constructing a unified metric reference system, it provides an absolute metric benchmark across devices and time, eliminating the problem of inconsistent geometric scales caused by differences in intrinsic parameters of different devices. Simultaneously, it establishes a set of multi-factor influence relationships, explicitly modeling the nonlinear coupling relationships between factors such as light, angle, automatic exposure, and automatic focus, rather than treating them as independent variables, thus achieving accurate modeling of coupling effects. By employing an elastic weight consolidation strategy, the correction mapping model can be incrementally updated, enabling the system to track time-varying characteristics such as thermal drift of mobile phone intrinsic parameters, exhibiting adaptability. Therefore, by linking a pre-positioned anti-shake correction layer with the unified metric reference system, and using spatial nodes as reference anchor points for shake detection and motion blur correction, reliable anti-shake performance without external sensors is achieved, forming a closed-loop correction. Finally, through an end-to-end neural network, camera parameters are estimated in real time at native resolution, achieving instant on-site quality feedback, improving acquisition efficiency, and using the anti-shake intensity vector as auxiliary information input to the intrinsic parameter estimation network to improve the accuracy of parameter estimation. This application enables a comprehensive improvement in the geometric accuracy, model consistency, and on-site data acquisition efficiency of 3D modeling of cultural walls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122695103A_ABST
    Figure CN122695103A_ABST
Patent Text Reader

Abstract

The application relates to the field of cultural heritage digital protection and three-dimensional modeling technology, and provides a cultural wall modeling design method and system for mobile phone real scene photo ranging and deviation correction. The method comprises the following steps: acquiring a current frame image and a shooting state parameter of a target scene; performing ranging and deviation correction on the image based on a unified measurement reference system containing a space scale reference and a multi-factor coupling relationship and a deviation correction mapping model; detecting a motion blur state and performing deblurring correction with a space node as a reference anchor point to obtain a clear image; feeding back and updating a deblurring correction amount to a time-varying state parameter vector; and outputting a ranging and deviation correction result based on the clear image and the updated parameter vector. The application realizes unified calibration of multi-factor coupling deviation through the unified measurement reference system and the deviation correction mapping model, suppresses motion blur through the pre-deblurring deviation correction layer and the measurement reference system linkage, and improves the geometric precision, model consistency and on-site collection efficiency of cultural wall three-dimensional modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital preservation and 3D modeling of cultural heritage, and in particular to a method and system for designing and modeling cultural walls by measuring and correcting distances from real-world photos taken with a mobile phone. Background Technology

[0002] In the field of digital preservation of cultural heritage, using real-view photos from mobile phones to create 3D models and measure distances of cultural walls is an important technical means. Existing technologies typically employ traditional camera calibration methods, motion reconstruction-based modeling methods, or mobile phone-based distance measurement modeling applications. These methods rely on static intrinsic parameters calibrated at the device's factory or dedicated depth hardware, treating imaging parameters as independent static constants.

[0003] However, in actual digital acquisition and modeling of cultural walls, the imaging process of mobile phone photos is a highly coupled, dynamically changing nonlinear system involving multiple factors. For example, different brands of mobile phones may produce inconsistent ranging results due to hardware differences; continuous shooting on the same phone can cause time-varying drift in intrinsic parameters due to heat generation; motion blur caused by handheld shooting can lead to the loss of feature points; changes in ambient light and automatic exposure and autofocus parameters can couple and interfere with image quality; wide-angle lenses and filters can introduce additional distortion; and operators cannot judge the acquisition quality in real time on-site. Existing technologies treat all imaging parameters as independent static constants. This fundamental misalignment between the "static constant" assumption and the reality of "dynamic coupling" results in a lack of a unified metric framework that can characterize and dynamically track the nonlinear coupling relationships of multiple factors. This makes it impossible to achieve consistent calibration of ranging and correction across devices, time periods, and working conditions, ultimately leading to insufficient geometric accuracy, poor model consistency, and low on-site acquisition efficiency in the 3D modeling of cultural walls. Summary of the Invention

[0004] To address the core problem of existing technologies lacking a unified metric framework for representing and dynamically tracking the nonlinear coupling relationships of multiple factors, which prevents consistent calibration across devices, time periods, and operating conditions in distance measurement and correction, this application proposes a method and system for designing and modeling cultural walls using distance measurement and correction based on real-world mobile phone photos. This method and system model the coupling relationships of multiple factors based on a unified metric reference system and a correction mapping model, and performs motion blur correction in conjunction with a pre-positioned anti-shake correction layer and the metric reference system. This improves the geometric accuracy, model consistency, and on-site data acquisition efficiency of 3D modeling of cultural walls.

[0005] To achieve the above objectives, this application adopts the following technical solution: Firstly, a method for modeling and designing a cultural wall using mobile phone real-scene photos for ranging and correction includes: acquiring the current frame image of the target scene and the corresponding shooting state parameters; performing ranging and correction on the current frame image based on a unified metric reference system and a correction mapping model to obtain corrected coordinates and real spatial distance. The unified metric reference system includes a spatial scale benchmark with spatial nodes, a time-varying state parameter vector, and multi-factor coupling relationships. The correction mapping model is used to characterize the mapping relationship between image pixel coordinates and shooting state parameters and the corrected coordinates and real spatial distance; using spatial nodes as reference anchor points, detecting the motion blur state of the current frame image, and performing anti-shake correction on the current frame image based on the motion blur state to obtain a clear image; feeding back the correction amount generated by the anti-shake correction to update the time-varying state parameter vector of the unified metric reference system; and outputting the ranging and correction results and modeling and design method of the target scene based on the clear image and the updated time-varying state parameter vector.

[0006] The above scheme constructs a unified measurement reference system as an absolute measurement benchmark and establishes a correction mapping model to represent the coupling relationship of multiple factors. At the same time, it links the front-end anti-shake correction layer with the measurement reference system to perform motion blur detection and correction, thereby achieving unified calibration of ranging deviation caused by multi-factor coupling and effective suppression of motion blur.

[0007] As one implementation method, the construction process of the unified metric reference system includes: obtaining the precise three-dimensional spatial coordinates of physical reference points deployed in the target scene, the three-dimensional spatial coordinates being calibrated based on the world coordinate system; calculating the homography matrix from the world coordinate system to the image coordinate system based on the three-dimensional spatial coordinates of the physical reference points and their corresponding image pixel coordinates in the current frame image; generating a virtual perspective transformation grid on the surface of the target scene according to the homography matrix, the perspective transformation grid containing equally spaced grid nodes, defining the grid nodes as spatial nodes, and all spatial nodes constituting the spatial scale reference of the unified metric reference system; establishing a time-varying state parameter vector based on the spatial scale reference and the shooting time of the current frame image, the time-varying state parameter vector containing the camera intrinsic parameter matrix, distortion coefficients, camera extrinsic parameters, ambient lighting characteristic parameters, automatic exposure parameters, autofocus status, and lens mode identifier at the current time; spatially aligning and binding the time-varying state parameter vector with the spatial nodes, and constructing a multi-factor influence relationship matrix to generate multi-factor coupling relationships; obtaining the unified metric reference system based on the spatial scale reference, the time-varying state parameter vector, and the multi-factor influence relationship matrix. This implementation provides the system with an absolute three-dimensional measurement coordinate system of space, time, and factors that is independent of specific equipment by constructing a spatial scale benchmark based on high-precision physical reference points and establishing a vector containing multi-dimensional time-varying state parameters, combined with a multi-factor influence relationship matrix.

[0008] As one implementation method, the process of constructing a multi-factor influence relationship matrix and generating multi-factor coupling relationships includes: constructing an initial influence relationship matrix using the time-varying state parameter vector as the row and column dimensions of the matrix, where the element value in the i-th row and j-th column of the initial influence relationship matrix represents the influence intensity coefficient of the i-th factor on the j-th factor; estimating the parameters of each influence intensity coefficient in the initial influence relationship matrix based on a historical training dataset to obtain an initialized influence relationship matrix, where the historical training dataset contains the observed values ​​of each influencing factor in history, the corresponding image pixel coordinates, and the corresponding ground truth values ​​of target scene spatial ranging; obtaining the real-time observed values ​​of each influencing factor at the current moment, and performing online incremental updates to the initialized influence relationship matrix based on the real-time observed values ​​to obtain a multi-factor influence relationship matrix, which is used to characterize the multi-factor coupling relationships between multiple influencing factors. This implementation method, by constructing and dynamically updating the multi-factor influence relationship matrix, explicitly models the nonlinear coupling relationships between factors such as light, angle, automatic exposure, and automatic focus, rather than treating them as independent variables, thereby achieving accurate modeling of coupling effects.

[0009] As one implementation method, the construction process of the bias correction mapping model includes: acquiring a historical training dataset, which contains observed values ​​of each influencing factor, corresponding image pixel coordinates, and corresponding ground truth values ​​of target scene spatial ranging; analyzing the combinations of influencing factors in the historical training dataset to obtain a set of multi-factor influence relationships, where each influence relationship contains a coupling strength coefficient and a coupling function form between two influencing factors; constructing an initial bias correction mapping model, which is a multilayer perceptron (MLP), where the input layer of the MLP receives image pixel coordinates and a time-varying state parameter vector, and the output layer outputs corrected coordinates and the true spatial distance; initializing the parameters of the hidden layers of the MLP based on the set of multi-factor influence relationships, and training the parameter-initialized MLP using the historical training dataset to obtain the bias correction mapping model. This implementation method, by injecting the set of multi-factor influence relationships as prior knowledge into the parameter initialization process of the hidden layers of the MLP, enables the bias correction mapping model to have better interpretability and generalization ability.

[0010] As one implementation method, the method further includes: responding to the acquisition of the target scene, inputting the current frame image into the bias correction mapping model, and outputting the corrected coordinates and predicted spatial distance of each pixel in the current frame image; calculating the reprojection error based on the known three-dimensional spatial coordinates of the physical reference point in the target scene and the corresponding image in the current frame image; if the reprojection error is less than a preset threshold, then marking the current frame image as a valid new sample; obtaining the pixel coordinates of each image in the valid new sample, the corresponding time-varying state parameter vector, and the corresponding spatial ranging ground truth; calculating the Fisher information matrix of each model parameter in the bias correction mapping model based on the historical training dataset; constructing an elastic weight consolidation loss function based on the valid new samples and the Fisher information matrix; the elastic weight consolidation loss function includes a new sample loss term and a regularization term; the regularization term applies an importance-weighted penalty to each model parameter of the bias correction mapping model based on the Fisher information matrix; updating the model parameters of the bias correction mapping model using gradient descent based on the elastic weight consolidation loss function, and using the updated bias correction mapping model for ranging correction of the next frame image. This implementation adopts an elastic weight consolidation (EWC) strategy for incremental updates, enabling the bias correction mapping model to adapt to new data (such as thermal drift of tracking mobile phone intrinsic parameters) while preventing catastrophic forgetting, thus achieving system adaptability.

[0011] As one implementation method, the process of detecting motion blur in the current frame image using spatial nodes as reference anchor points includes: acquiring the current frame image and the previous frame image, and detecting the pixel positions of the spatial nodes; calculating the displacement vectors of the spatial nodes between adjacent frames, where the displacement vector is the difference between the pixel positions of the spatial nodes in the current frame image and the pixel positions in the previous frame image; calculating the statistics of the inter-frame displacement vectors based on the displacement vectors of the spatial nodes, where the statistics include at least one of the displacement mean, displacement standard deviation, and maximum displacement value; comparing the statistics with a preset threshold, and determining the jitter type and jitter intensity vector of the current frame image based on the comparison result, denoted as the motion blur state. This implementation method utilizes spatial nodes in a unified metric reference frame as natural reference anchor points, and detects jitter by analyzing the statistics of their inter-frame displacements, achieving jitter detection without the need for external sensors (such as IMUs), and has higher reliability and universality.

[0012] As one implementation method, the process of performing jitter correction on the current frame image based on motion blur to obtain a clear image includes: parsing the jitter type and jitter intensity vector in the motion blur state, and initializing the search space of the blur kernel; estimating the target blur kernel of the current frame image based on the search space, using the current frame image and the jitter intensity vector as input; obtaining the theoretical pixel positions of spatial nodes in a unified metric reference frame, and performing non-blind deconvolution on the current frame image using the target blur kernel; responding to the deconvolution activation signal, constructing a deconvolution objective function containing data fidelity terms and metric constraint terms, using the theoretical pixel positions of spatial nodes as constraints; iteratively optimizing the deconvolution objective function, and outputting a clear image when the convergence condition is met. This implementation method introduces metric grid constraints in the non-blind deconvolution process, that is, the position of spatial nodes in the corrected image should be consistent with the theoretical position, thereby improving the accuracy of jitter correction and forming a bidirectional closed loop between the jitter layer and the metric grid.

[0013] One implementation method further includes: inputting a clear image into an intrinsic parameter estimation network, where the clear image maintains the native resolution of the mobile phone camera; extracting multi-scale feature maps of the clear image through the convolutional backbone network of the intrinsic parameter estimation network; regressing the camera intrinsic parameter matrix and distortion coefficients based on the multi-scale feature maps using global average pooling and fully connected layers; estimating camera extrinsic parameters based on the multi-scale feature maps and the camera intrinsic parameter matrix using the cross-attention mechanism of the intrinsic parameter estimation network; and outputting the camera intrinsic parameter matrix, distortion coefficients, and camera extrinsic parameters, which are used to update the time-varying state parameter vector. This implementation method uses a lightweight end-to-end neural network to estimate camera intrinsic parameters, distortion coefficients, and extrinsic parameters in real time at the native resolution of the mobile phone, achieving immediate on-site quality feedback and shortening the acquisition-to-feedback cycle.

[0014] One implementation method further includes: acquiring a clear image and a jitter intensity vector from the shaking correction output; using the clear image as the main input to the intrinsic parameter estimation network and the jitter intensity vector as auxiliary information input to the conditional coding layer of the intrinsic parameter estimation network; mapping the jitter intensity vector to conditional coding features through the conditional coding layer, and fusing the conditional coding features with the multi-scale feature map output by the convolutional backbone network; regressing the camera intrinsic parameter matrix and distortion coefficients through global average pooling and fully connected layers based on the fused features; estimating the camera extrinsic parameters through a cross-attention mechanism based on the multi-scale feature map, the camera intrinsic parameter matrix, and the conditional coding features; outputting the camera intrinsic parameter matrix, distortion coefficients, and camera extrinsic parameters, and using the camera intrinsic parameter matrix, distortion coefficients, and camera extrinsic parameters to update the time-varying state parameter vector. This implementation method uses the jitter intensity vector output by the shaking correction layer as auxiliary information input to the intrinsic parameter estimation network, helping the network distinguish between "geometric changes caused by camera motion" and "geometric changes caused by scene structure," thus improving the accuracy of intrinsic parameter and pose estimation.

[0015] Furthermore, this application also provides a cultural wall modeling and design system for ranging and correcting real-scene photos taken with a mobile phone, applicable to any of the methods described above. This system specifically includes: a communication unit for acquiring the current frame image of the target scene and corresponding shooting state parameters; a processing unit for performing ranging and correcting on the current frame image based on a unified metric reference system and a correction mapping model to obtain corrected coordinates and the true spatial distance; detecting the motion blur state of the current frame image using spatial nodes in the unified metric reference system as reference anchor points, and performing anti-shake correction on the current frame image based on the motion blur state to obtain a clear image; feeding back the correction amount generated by the anti-shake correction to update the time-varying state parameter vector of the unified metric reference system; and outputting the ranging and correcting results of the target scene and the modeling and design method based on the clear image and the updated time-varying state parameter vector. The above system solution, through the collaborative work of the communication unit and the processing unit, achieves the technical functions corresponding to the aforementioned methods, thereby enabling deployment on mobile devices or servers to perform cultural wall modeling and design tasks.

[0016] This application provides a method and system for modeling and designing cultural walls using mobile phone real-scene photo ranging and correction. By constructing a unified metric reference system, it provides an absolute metric benchmark across devices and time, eliminating the problem of inconsistent geometric scales caused by differences in intrinsic parameters of different devices. Simultaneously, it establishes a set of multi-factor influence relationships, explicitly modeling the nonlinear coupling relationships between factors such as light, angle, automatic exposure, and automatic focus, rather than treating them as independent variables, thus achieving accurate modeling of coupling effects. By employing an elastic weight consolidation strategy, the correction mapping model can be incrementally updated, enabling the system to track time-varying characteristics such as thermal drift of mobile phone intrinsic parameters, exhibiting adaptability. Therefore, by linking a pre-positioned anti-shake correction layer with the unified metric reference system, and using spatial nodes as reference anchor points for shake detection and motion blur correction, reliable anti-shake performance without external sensors is achieved, forming a closed-loop correction. Finally, through an end-to-end neural network, camera parameters are estimated in real time at native resolution, achieving instant on-site quality feedback, improving acquisition efficiency, and using the anti-shake intensity vector as auxiliary information input to the intrinsic parameter estimation network to improve the accuracy of parameter estimation. This application enables a comprehensive improvement in the geometric accuracy, model consistency, and on-site data acquisition efficiency of 3D modeling of cultural walls.

[0017] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0018] Figure 1 A flowchart illustrating a method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction, as provided in an embodiment of this application. Figure 2 A schematic diagram illustrating the process of constructing a unified measurement reference system in a mobile phone real-scene photo ranging and correction method for cultural wall modeling and design provided in this application embodiment; Figure 3 A schematic diagram illustrating the multi-factor influence relationship matrix generation process in a mobile phone real-scene photo ranging and correction cultural wall modeling and design method provided in this application embodiment; Figure 4 A schematic diagram of the process for constructing and training the correction mapping model in a method for measuring and correcting the distance of a mobile phone real-scene photo for cultural wall modeling and design provided in this application embodiment; Figure 5 A schematic diagram of the incremental update mechanism of the correction mapping model in a method for modeling and designing a cultural wall using mobile phone real-scene photos for ranging and correction, provided in an embodiment of this application. Figure 6 In the cultural wall modeling and design method for ranging and correcting mobile phone real-scene photos provided in this application embodiment, the process of detecting motion blur state is based on spatial nodes as anchor points; Figure 7 In the cultural wall modeling and design method for distance measurement and correction of mobile phone real scene photos provided in this application embodiment, the process of performing shake correction based on the detection results and outputting a clear image is described. Figure 8 A schematic diagram of the intrinsic parameter estimation network architecture and processing flow in a mobile phone real-scene photo ranging and correction cultural wall modeling and design method provided in this application embodiment; Figure 9 This is a structural schematic diagram of a cultural wall modeling and design system for measuring and correcting distances from mobile phone real-scene photos, provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit the scope of this application.

[0021] Example 1: like Figure 1 As shown, this embodiment provides a method for modeling and designing cultural walls using mobile phone real-scene photos for distance measurement and correction. This method aims to solve the core problem in existing technologies where distance measurement deviations cannot be uniformly calibrated due to the coupling of multiple factors. The overall process of this method includes image acquisition, distance measurement and correction based on a unified metric reference system and a correction mapping model, motion blur detection and shaking correction based on spatial nodes, feedback and updating of time-varying state parameter vectors, and output of results.

[0022] Step S100: Obtain the current frame image of the target scene and the corresponding shooting status parameters.

[0023] The target scene refers to the surface of a cultural wall or other object with complex textures and geometric structures that requires 3D modeling and distance measurement.

[0024] The current frame image refers to a single photo taken by the mobile phone camera at a certain moment, which can be acquired in real time through the image sensor of the mobile phone camera.

[0025] Shooting status parameters refer to the comprehensive status information of the mobile phone camera and its surrounding environment when the photo is taken, such as, but not limited to, the camera's intrinsic parameter matrix, distortion coefficient, extrinsic parameters, ambient lighting characteristics, automatic exposure parameters, autofocus status, and lens mode identifier, which can be read from the mobile phone's system API or image EXIF ​​metadata.

[0026] Step S200: Based on the unified metric reference system and the correction mapping model, perform ranging and correction on the current frame image to obtain the corrected coordinates and the true spatial distance.

[0027] The unified measurement reference system is an absolute three-dimensional measurement coordinate system that is independent of any specific mobile device. It includes three-dimensional measurements of space, time, and factors, and has a spatial scale reference for spatial nodes, a time-varying state parameter vector, and multi-factor coupling relationships. Its spatial scale reference provides an absolute geometric reference for ranging; the time-varying state parameter vector records the dynamic state at the time of shooting; and the multi-factor coupling relationships characterize the mutual influence between the various state parameters.

[0028] The correction mapping model is used to characterize the mapping relationship between image pixel coordinates and shooting state parameters, and the correction coordinates and real spatial distance.

[0029] In some implementations, the pixel coordinates of the current frame image and the corresponding shooting state parameters are taken as input and fed into the correction mapping model. At this time, the model can output the corrected coordinates of the pixel in the unified metric reference system (i.e., the coordinates after eliminating distortion and perspective effects) and the corresponding real spatial distance (usually in millimeters) based on the mapping relationship learned internally.

[0030] It should be noted that by using the absolute benchmark and correction mapping model provided by the unified measurement reference system to model the coupling relationship of multiple factors, the problem of inconsistent ranging results under different equipment, different times, and different working conditions can be eliminated.

[0031] Step S300: Using spatial nodes as reference anchor points, detect the motion blur state of the current frame image, and perform anti-shake correction on the current frame image based on the motion blur state to obtain a clear image.

[0032] Spatial nodes are components of the spatial scale reference in a unified metric reference system. They are equidistant points distributed within a virtual perspective transformation grid generated on the surface of the target scene. The theoretical locations of these nodes in three-dimensional space are known, thus serving as natural reference anchor points for detecting image jitter.

[0033] Motion blur refers to the phenomenon where the edges of objects in an image appear blurred or trailed due to phone shaking during shooting. Its state information includes the type of shaking (such as global translation, rotation, scaling, etc.) and the shaking intensity vector.

[0034] In some implementations, the process of detecting motion blur may include: acquiring the current frame image and the previous frame image, detecting the pixel positions of spatial nodes in the two frames, calculating their displacement vectors, and determining the jitter type and intensity based on the statistics of the displacement vectors (such as mean and standard deviation).

[0035] Shake removal correction refers to the process of using estimated motion blur information to perform deconvolution and other processing on a blurred image to restore a clear image. Specifically, the execution logic of this step is as follows: First, using spatial nodes as references, the motion blur state of the current frame is accurately detected by analyzing their inter-frame displacements. Based on the detected blur state, the search space of the blur kernel is initialized, and the target blur kernel is estimated. Then, using the theoretical position of the spatial nodes as constraints, non-blind deconvolution is performed on the current frame image to obtain a clear image free from the effects of shake.

[0036] Step S400: Feedback the correction amount generated by the jitter correction to update the time-varying state parameter vector of the unified metric reference system.

[0037] Among them, the correction amount refers to the parameters calculated during the shake correction process to compensate for image blur, such as the estimated blur kernel parameters or the correction value of the camera motion trajectory.

[0038] The time-varying state parameter vector is a vector in a unified metric reference frame used to record all dynamic state parameters at the current moment.

[0039] In some implementations, the camera motion parameters (such as rotation angle and translation amount) estimated in the shaking correction step are fused or corrected with the camera extrinsic parameters in the current time-varying state parameter vector, thereby updating the motion-related states in the vector. Specifically, the correction information generated during the shaking correction process, which reflects the true motion state of the camera, is used as new observations to update the time-varying state parameter vector recorded in the unified metric reference frame.

[0040] Step S500: Based on the clear image and the updated time-varying state parameter vector, output the ranging and correction results of the target scene and the modeling design method.

[0041] The clear image is the image obtained after the shake correction in step S300, and its motion blur has been effectively suppressed.

[0042] The updated time-varying state parameter vector is the vector updated by feedback in step S400, and contains the latest dynamic state information.

[0043] Ranging correction results refer to the data obtained after accurately measuring the distance to each point in the target scene, which usually includes the corrected three-dimensional coordinates and the actual spatial distance.

[0044] Modeling and design methods refer to the specific technical approaches used to construct 3D models of cultural walls or other target scenes based on these precise distance measurement data.

[0045] In some implementations, the output process may include: inputting the sharpened image and the updated time-varying state parameter vector back into the correction mapping model for final ranging calculation; or, inputting the sharpened image into an end-to-end intrinsic parameter estimation network to estimate camera parameters and update vectors in real time before ranging. Specifically, the execution logic of this step is: using the corrected and updated image and state information, performing the final ranging calculation, and outputting the result in the form of data or a visualization model for subsequent 3D modeling, design, or analysis.

[0046] It should be noted that outputting based on clear images and updated state vectors ensures the geometric accuracy and consistency of the final results, enabling real-time quality feedback on-site, thereby improving the efficiency and quality of digital acquisition and modeling of cultural walls.

[0047] Based on the above technical solution, this embodiment constructs a complete methodological framework with a unified metric reference system as the core, a correction mapping model as the mapping tool, and a pre-positioned anti-shake correction layer as the quality assurance. Ke acquires raw data in step S100, performs ranging correction based on an absolute benchmark in step S200, performs anti-shake correction based on reference anchor points in step S300, implements dynamic feedback updates of the state in step S400, and finally outputs accurate results in step S500, forming a closed-loop process from data acquisition, processing, correction to output. This utilizes the unified metric reference system's three dimensions of metric standards—space, time, and multi-factor coupling—and enables the system to be adaptive through a feedback mechanism, thereby systematically solving the ranging inconsistency problem caused by multi-factor coupling and dynamic changes in existing technologies.

[0048] Example 2: like Figure 2-3 As shown in some embodiments, this embodiment describes in detail the construction process of the unified metric reference system. This process provides a specific and implementable construction scheme for the unified metric reference system on which step S200 in Embodiment 1 is based. It aims to establish an absolute three-dimensional metric coordinate system of space, time and factors that is independent of specific devices, thereby laying a technical foundation for solving the ranging deviation problem caused by multi-factor coupling.

[0049] Step S210: Obtain the precise three-dimensional spatial coordinates of the physical reference points deployed in the target scene. The three-dimensional spatial coordinates are calibrated based on the world coordinate system.

[0050] Physical reference points refer to markers with known and precise three-dimensional spatial coordinates that are pre-placed on or around the surface of the target scene (such as a cultural wall). These markers form a sparse three-dimensional control network, which is the physical framework of the unified metric reference system's spatial scale benchmark.

[0051] In some implementations, physical reference points can be represented by reflective markers. These markers can be reliably identified by the phone camera under various lighting conditions (such as strong light, shadow, and backlight), thus ensuring the robustness of reference point detection.

[0052] The three-dimensional spatial coordinates of the physical reference point can be measured and calibrated using a high-precision total station with a measurement accuracy of up to 0.5 millimeters, thus providing a high-precision absolute spatial reference for the entire system.

[0053] Specifically, when setting up physical reference points, they should be evenly distributed on the surface of the target scene. For example, they can be set up in a grid pattern or according to the positions of key feature points to ensure that the perspective transformation grid generated later can cover the entire target scene.

[0054] Step S220: Based on the three-dimensional spatial coordinates of the physical reference point and its corresponding image pixel coordinates in the current frame image, calculate the homography matrix from the world coordinate system to the image coordinate system.

[0055] The homography matrix is ​​a 3x3 matrix that describes the projection transformation relationship between two planes. A mathematical mapping relationship was established between the world coordinate system plane where the physical reference point is located (i.e., the surface of the cultural wall) and the image plane captured by the mobile phone camera.

[0056] In some implementations, the pixel coordinates of physical reference points in the current frame image are automatically identified through image processing algorithms (such as feature point detection and matching). These pixel coordinates and the known three-dimensional spatial coordinates of the physical reference points (usually their two-dimensional coordinates on the world coordinate system plane) can then be used as inputs to solve the homography matrix using the Direct Linear Transform (DLT) algorithm or the Random Sample Consensus (RANSAC) algorithm.

[0057] Specifically, homography matrix The following relationship must be satisfied: for any point on the world coordinate system plane Its projection point on the image plane It can be done through formula Perform calculations, where and All are represented using homogeneous coordinates.

[0058] Step S230: Based on the homography matrix, generate a virtual perspective transformation grid on the surface of the target scene. The perspective transformation grid contains grid nodes with equal spacing. Define the grid nodes as spatial nodes. All spatial nodes constitute the spatial scale reference of a unified metric reference system.

[0059] The perspective transformation grid is a virtual mesh with a regular geometric structure generated on the surface of the target scene. In the three-dimensional world coordinate system, this grid is evenly spaced (for example, the node spacing can be set to 5 centimeters), but on the image plane, due to perspective projection effects, its node distribution exhibits a pattern of larger nodes in the foreground and smaller nodes in the background.

[0060] A grid node is an intersection point on the virtual grid; these nodes are defined as spatial nodes.

[0061] In some implementations, a regular grid of points (e.g., a rectangular grid with a spacing of 5 cm) is defined on the world coordinate system plane. Then, the homography matrix H calculated in step S220 can be used to project each point in the grid onto the image plane to obtain the corresponding grid node position on the image plane.

[0062] Specifically, for a grid point on the world coordinate system plane Its corresponding point on the image plane It can be done through formula Calculated. All calculated in this way. The points form a perspective transformation grid on the image plane.

[0063] It should be noted that the theoretical positions of these spatial nodes in three-dimensional space are known (i.e., grid points on the world coordinate system plane), and their actual positions in the image are detectable. This correspondence between theory and reality makes spatial nodes natural reference anchors for detecting image jitter and performing metric correction. All spatial nodes together constitute the spatial scale benchmark of a unified metric reference system, providing an absolute geometric reference for subsequent ranging correction and jitter reduction.

[0064] Step S240: Based on the spatial scale reference and the shooting time of the current frame image, establish a time-varying state parameter vector. The time-varying state parameter vector includes the camera intrinsic parameter matrix, distortion coefficient, camera extrinsic parameters, ambient light characteristic parameters, automatic exposure parameters, autofocus status, and lens mode identifier at the current time.

[0065] The time-varying state parameter vector is a vector used to record all dynamic state parameters at the moment of shooting. It links the spatial scale benchmark with the time dimension and the influencing factor dimension, and is the core of the unified metric reference system's ability to track dynamic changes.

[0066] In some implementations, the time-varying state parameter vector It can be represented as a vector containing seven components: ,in The camera intrinsic parameter matrix at time t, including focal length , and principal point coordinates , ; The distortion coefficients at time t include radial distortion coefficients k1, k2, k3 and tangential distortion coefficients p1, p2; The camera extrinsic parameters at time t include the rotation matrix R and the translation vector T; The ambient lighting characteristics at time t can include color temperature, illuminance, and main lighting direction, etc. Automatic exposure parameters representing time t can include ISO sensitivity, shutter speed, aperture value, etc. The autofocus status at time t can include focus distance, depth of field range, etc. The lens mode identifier at time t indicates the currently used lens mode (such as wide-angle mode, standard mode, telephoto mode) and whether a filter is enabled.

[0067] It should be noted that the time-varying state parameter vector records and tracks multiple key factors affecting image quality (device intrinsic parameters, ambient light, shooting parameters, etc.) in a single vector, enabling the system to comprehensively describe the overall state at the moment of shooting, and providing complete state information for subsequent establishment of multi-factor coupling relationships and range measurement correction.

[0068] Step S250: Spatial alignment and binding of time-varying state parameter vectors and spatial nodes, and construction of multi-factor influence relationship matrix to generate multi-factor coupling relationship.

[0069] Spatial alignment binding refers to associating each component in a time-varying state parameter vector with a specific spatial node or spatial region. For example, it can be assumed that the state parameters of an image taken near a certain spatial node are mainly affected by factors such as illumination and angle in the region where that node is located.

[0070] In some implementations, this binding can be achieved by maintaining an independent copy of the time-varying state parameter vector for each spatial node or region, or through spatial interpolation methods.

[0071] The multi-factor influence matrix is ​​a matrix used to quantitatively describe the mutual influence relationships between the components in a time-varying state parameter vector.

[0072] It should be noted that binding state parameters to spatial nodes gives the unified metric reference frame not only a time dimension but also a spatial dimension, enabling a more precise characterization of state differences at different spatial locations. Furthermore, the introduction of the multi-factor influence matrix explicitly models the nonlinear coupling relationships between the various state parameters.

[0073] Step S260: Based on the spatial scale benchmark, time-varying state parameter vector, and multi-factor influence relationship matrix, a unified measurement reference system is obtained.

[0074] The unified metric reference system is a three-dimensional metric coordinate system composed of a spatial scale benchmark (composed of spatial nodes), a time-varying state parameter vector, and a multi-factor influence relationship matrix.

[0075] In some implementations, the reference frame can be understood as a dynamic knowledge base that simultaneously contains absolute spatial geometric information, time-varying state information, and coupling relationship information between various state factors.

[0076] Specifically, the spatial scale benchmark provides an absolute geometric reference, the time-varying state parameter vector records the dynamic state, and the multi-factor influence matrix describes the mutual influence between these states.

[0077] Step S270: Construct a multi-factor influence relationship matrix and generate multi-factor coupling relationships.

[0078] Among them, the multi-factor influence relationship matrix is ​​a square matrix, whose row dimension and column dimension both correspond to the various factors in the time-varying state parameter vector.

[0079] In some implementations, this matrix can be a 7x7 matrix, corresponding to the seven components of the time-varying state parameter vector. The value of the element in the i-th row and j-th column of the matrix is... This represents the influence strength coefficient of the i-th factor on the j-th factor. Specifically, the process of constructing this matrix includes the following sub-steps: Step S271: Construct an initial influence relationship matrix using the time-varying state parameter vector as the row and column dimensions of the matrix. The element value in the i-th row and j-th column of the initial influence relationship matrix represents the influence intensity coefficient of the i-th factor on the j-th factor.

[0080] The initial influence matrix is ​​the initial version of the multi-factor influence matrix.

[0081] In some implementations, the initialization of this matrix can be based on prior knowledge. For example, it can be assumed that ambient light L has a strong influence on automatic exposure A, while the camera intrinsic parameter K has no direct influence on ambient light L.

[0082] Specifically, during initialization, the diagonal elements (i.e., the influence of factors on themselves) can be set to 1, and the other elements can be set to an initial value between 0 and 1 based on prior knowledge.

[0083] Step S272: Based on the historical training dataset, perform parameter estimation on each influence intensity coefficient in the initial influence relationship matrix to obtain the initialized influence relationship matrix. The historical training dataset contains the observed values ​​of each historical influencing factor, the corresponding image pixel coordinates, and the corresponding ground truth values ​​of spatial ranging of the target scene.

[0084] The historical training dataset is a pre-collected dataset containing a large amount of historical data. Each record in this dataset contains observed values ​​of a set of influencing factors, corresponding image pixel coordinates, and ground truth values ​​of spatial ranging of the target scene obtained through high-precision measurement.

[0085] In some implementations, the parameter estimation process can employ statistical learning methods. For example, historical training datasets can be used to optimize the influence intensity coefficients in the initial influence relationship matrix through methods such as maximum likelihood estimation or Bayesian estimation, so that the matrix can best fit the relationships between factors observed in historical data.

[0086] Specifically, the optimization objective can be to minimize the prediction error, that is, to predict the ranging result using the current influence relationship matrix and factor observations, and then compare it with the true ranging value, and reduce the error by iteratively adjusting the matrix elements.

[0087] It should be noted that parameter estimation based on historical training datasets has transformed the multi-factor influence matrix from an initial matrix based on prior knowledge into a data-validated matrix that reflects the coupling relationships between real-world factors, thereby improving the accuracy and reliability of the matrix.

[0088] Step S273: Obtain the real-time observation values ​​of each influencing factor at the current moment, and perform online incremental updates on the initialized influence relationship matrix based on the real-time observation values ​​to obtain the multi-factor influence relationship matrix. The multi-factor influence relationship matrix is ​​used to characterize the multi-factor coupling relationship between multiple influencing factors.

[0089] Online incremental update refers to the process of continuously optimizing the influence relationship matrix using new data collected in real time during system operation.

[0090] In some implementations, the real-time observations of each influencing factor can be estimated in real time by reading the mobile phone system API, image EXIF ​​metadata, or through an intrinsic parameter estimation network. The real-time observations at the current moment can then be used together with historical data to reassess the accuracy of the influence relationship matrix.

[0091] If the prediction error of the matrix exceeds a preset threshold, an online learning algorithm (such as stochastic gradient descent) is used to fine-tune the matrix elements to adapt to the new data distribution. For example, when the intrinsic parameter K drifts due to heat generated by continuous shooting, the real-time observed K value differs from the K value in historical data. This difference will propagate to other factors through the influence relationship matrix. The online update mechanism can adjust the matrix elements in a timely manner to reflect this new coupling relationship.

[0092] Based on the above technical solution, this embodiment establishes an absolute spatial scale benchmark by acquiring high-precision physical reference points and calculating the homography matrix. This allows for the definition of a time-varying state parameter vector containing seven components, enabling comprehensive tracking of dynamic states. A multi-factor influence relationship matrix is ​​constructed and dynamically updated, explicitly modeling the nonlinear coupling relationships between various factors. This allows for the joint construction of a three-dimensional measurement coordinate system encompassing space, time, and factors, providing a unified, accurate, and adaptive measurement basis for subsequent ranging correction and jitter reduction, thus systematically solving the ranging deviation problem caused by multi-factor coupling.

[0093] Example 3: like Figure 4 As shown in some embodiments, this embodiment describes in detail the construction process of the bias correction mapping model. This process provides a specific and implementable construction scheme for the bias correction mapping model on which step S200 in Embodiment 1 is based. It aims to improve the interpretability and generalization ability of the model by injecting the multi-factor influence relationship as prior knowledge into the model, thereby distinguishing it from the deep learning calibration method in the prior art that only learns implicitly from data.

[0094] Step S310: Obtain the historical training dataset, which includes the observed values ​​of each influencing factor in history, the corresponding image pixel coordinates, and the corresponding ground truth values ​​of spatial ranging of the target scene.

[0095] The historical training dataset is a large dataset pre-collected for training the bias correction mapping model. Each record in this dataset contains observations of a set of influencing factors, corresponding image pixel coordinates, and ground truth values ​​of spatial ranging of the target scene obtained through high-precision measurement.

[0096] In some implementations, the data acquisition process may include: First, setting up high-precision physical reference points on the surface of the target scene (such as a cultural wall) and using equipment such as a total station to measure their precise three-dimensional spatial coordinates as the true values; then, using mobile phones of different brands and models, taking pictures of the target scene under the above-mentioned various working conditions, while recording the shooting state parameters (i.e., the observed values ​​of influencing factors) corresponding to each picture; finally, automatically identifying the pixel coordinates of the physical reference points in the pictures through image processing algorithms, and associating them with the known three-dimensional spatial coordinates, thereby constructing a complete data record containing the observed values ​​of influencing factors, image pixel coordinates, and the true values ​​of spatial ranging.

[0097] It should be noted that the collection of historical training datasets needs to cover a variety of working conditions to ensure that the model can learn the mapping relationship under various conditions. For example, working conditions may include different lighting conditions (such as early morning, noon, evening, and cloudy days), different shooting angles (such as 0°, 15°, 30°, and 45°), different autofocus states, different auto exposure parameters, whether filters are enabled, and whether a wide-angle lens is used, etc.

[0098] Step S320: Analyze the combination of each influencing factor in the historical training dataset to obtain a set of multi-factor influence relationships. Each influence relationship in the set of multi-factor influence relationships contains the coupling strength coefficient and coupling function form between two influencing factors.

[0099] Among them, the multi-factor influence relationship set records the pairwise coupling relationships between each factor in the time-varying state parameter vector.

[0100] In some implementations, the observed sequences of all influencing factors are extracted from the historical training dataset, and these factors are systematically analyzed in pairs, such as analyzing the relationship between ambient light and automatic exposure, or the relationship between automatic exposure and autofocus. For each pair of factors (e.g., factor i and factor j), the analysis process may include: calculating their correlation in historical data (e.g., Pearson correlation coefficient) to preliminarily determine the coupling strength; determining the form of the coupling function between them through regression analysis or nonlinear fitting, such as linear, polynomial, exponential, or piecewise functional relationships; and determining the confidence level of the coupling relationship based on the goodness of fit (e.g., R² value) and the sample size.

[0101] For example, analysis might reveal a strong coupling relationship between ambient lighting and automatic exposure (coupling strength coefficient close to 1), and that the coupling function approximates a logarithmic function, meaning that for every increase in light intensity by one level, the automatic exposure parameters (such as ISO) adjust logarithmically. All such pairwise relationships constitute the set R of multi-factor influence relationships.

[0102] Step S330: Construct an initial correction mapping model. The initial correction mapping model is a multilayer perceptron. The input layer of the multilayer perceptron receives the image pixel coordinates and time-varying state parameter vectors, and the output layer outputs the corrected coordinates and the true spatial distance.

[0103] The initial bias correction mapping model is an untrained multilayer perceptron neural network structure.

[0104] In some implementations, the multilayer perceptron can consist of an input layer, several hidden layers, and an output layer. The number of neurons in the input layer is determined by the dimension of the input data, which includes image pixel coordinates (usually two-dimensional, i.e., u and v) and a time-varying state parameter vector (containing seven components: K, D, E, L, A, F, W). Therefore, the number of neurons in the input layer can be 2 + 7 = 9.

[0105] The number of neurons in the output layer is determined by the dimension of the output data, which includes the corrected coordinates (usually two-dimensional, i.e., u' and v') and the true spatial distance (one-dimensional, i.e., d). Therefore, the output layer can have 3 neurons. The number of hidden layers and the number of neurons in each layer can be designed according to the complexity of the model and computational resources. For example, 2-3 hidden layers can be set, each containing 64 or 128 neurons.

[0106] Step S340: Initialize the parameters of the hidden layer of the multilayer perceptron based on the set of multi-factor influence relationships, and train the multilayer perceptron after parameter initialization using the historical training dataset to obtain the correction mapping model.

[0107] The parameter initialization of the hidden layer based on the set of multi-factor influence relationships refers to using the structured knowledge obtained in step S320 to set the initial value of the weight matrix of the hidden layer of the multilayer perceptron, rather than using random initialization.

[0108] In some implementations, the coupling strength coefficients and coupling function forms in the set of multi-factor influence relationships R are transformed into constraints on the network weights. For example, if it is known that factor i has a strong coupling relationship with factor j, then in the weight matrix from the input layer to the first hidden layer, the neuron connection weights corresponding to factors i and j can be initialized to a large value, or specifically initialized according to the coupling function form.

[0109] One possible implementation is to encode the set of multi-factor influence relationships R into a prior matrix P, and then use P to initialize the weight matrix W1 of the first hidden layer, for example, let ,in It is a coefficient between 0 and 1, used to control the balance between prior knowledge and random exploration. These are randomly initialized weights.

[0110] Using the historical training dataset obtained in step S310, the multilayer perceptron with initialized parameters is trained. The training process can employ the backpropagation algorithm and gradient descent optimizer (such as Adam), iteratively updating all parameters (including weights and biases) of the network by minimizing the loss function (such as mean squared error loss) between the predicted output (corrected coordinates and the distance to the real space) and the ground truth.

[0111] Based on the above technical solution, this embodiment obtains a historical training dataset covering various working conditions, providing a rich data foundation for model learning. By analyzing the set of multi-factor influence relationships and making the domain prior knowledge explicit and structured, a multilayer perceptron can be constructed, and parameters can be initialized based on the set of influence relationships. This realizes the injection of prior knowledge into the model, and finally, the final deviation correction mapping model is obtained through training. This results in a ranging deviation correction model with interpretability and strong generalization ability, enabling it to inject multi-factor coupling relationships as prior knowledge into the model training process, thereby systematically improving the model's ranging deviation correction performance in complex, dynamic, and multi-factor coupled environments.

[0112] Example 4: like Figure 5 As shown in some embodiments, this embodiment describes in detail the incremental update mechanism of the bias correction mapping model. This mechanism provides a specific and implementable solution for the "feedback update" involved in step S400 of Embodiment 1 and the core of the system's adaptability. This mechanism aims to enable the bias correction mapping model to continuously learn newly acquired data to track time-varying characteristics such as thermal drift caused by continuous shooting heat in the phone's intrinsic parameters. At the same time, it prevents the model from forgetting the old knowledge it has already learned when learning new knowledge through specific strategies, that is, to prevent catastrophic forgetting.

[0113] Step S410: In response to the acquisition of the target scene, input the current frame image into the correction mapping model and output the corrected coordinates and predicted spatial distance of each pixel in the current frame image.

[0114] Among them, "responding to target scene acquisition" refers to the system automatically triggering subsequent processing procedures when it receives a new frame image captured by the mobile phone.

[0115] The current frame image is the newly acquired, unprocessed original image. This image is input into the correction mapping model constructed and trained in Example 3. Based on its current internal parameters, the model calculates for each pixel in the image and outputs two key results: first, the corrected coordinates, which are the coordinates of the pixel after distortion and perspective correction in a unified metric reference system; and second, the predicted spatial distance, which is the actual physical spatial distance corresponding to the pixel.

[0116] Step S420: Based on the known three-dimensional spatial coordinates of the physical reference point in the target scene and the corresponding image in the current frame image, calculate the reprojection error. If the reprojection error is less than a preset threshold, then mark the current frame image as a valid new sample.

[0117] Among them, physical reference points refer to marker points with high-precision three-dimensional spatial coordinates that are pre-deployed in the target scene. Their coordinates are obtained by measuring equipment such as total stations and can be used as true values.

[0118] Reprojection error is an indicator of model prediction accuracy. Its calculation process involves first using the pixel coordinates of the physical reference point in the current frame image and the correction coordinates output by the correction mapping model, combined with the known three-dimensional spatial coordinates of the physical reference point, to calculate the theoretical projection position of the point on the image plane through the camera projection model. Then, the Euclidean distance between the theoretical projection position and the actual detected pixel position of the physical reference point in the image is calculated. This distance is the reprojection error.

[0119] The preset threshold is a value set according to the system's accuracy requirements; for example, it can be set to 2 pixels.

[0120] In some implementations, if the calculated reprojection error is less than a preset threshold, it indicates that the current frame image has high capture quality, the model's prediction on that image is close enough to the physical truth, and the information contained in the image is reliable and valuable; therefore, it is marked as a valid new sample. Conversely, if the error is too large, it may mean that the image has serious blurring, occlusion, or abnormal lighting, and its data quality is unreliable and unsuitable as a basis for model updates.

[0121] Step S430: Obtain the pixel coordinates of each image in the valid new sample, the corresponding time-varying state parameter vector, and the corresponding spatial ranging ground value.

[0122] Among them, a valid new sample refers to a current frame image that has been determined and labeled in step S420 and has reliable quality.

[0123] Acquiring data from this sample refers to extracting training data pairs for model updates from the image and its associated metadata. This requires extracting three parts of information: first, the coordinates of each pixel in the image; second, the time-varying state parameter vector corresponding to the time the image was captured, which contains seven components: camera intrinsic parameters, distortion coefficients, extrinsic parameters, ambient lighting, auto exposure, autofocus, and lens mode at the time of capture. This information can be read from the phone's system API or the image's EXIF ​​metadata; and third, the corresponding ground truth value for spatial ranging. For valid new samples, this ground truth value can be accurately calculated using the known three-dimensional spatial coordinates of the physical reference point and the camera projection model.

[0124] In some implementations, the above data will be organized into triples of (pixel coordinates, time-varying state parameter vectors, and spatial distance ground truth) to form a new sample dataset for model fine-tuning.

[0125] Step S440: Based on the historical training dataset, calculate the Fisher information matrix of each model parameter in the correction mapping model.

[0126] The historical training dataset refers to the dataset containing a large amount of historical data used in Example 3 for the initial training of the correction mapping model.

[0127] The Fisher information matrix is ​​used to quantify the importance of each parameter in the bias correction mapping model for fitting historical training data.

[0128] In some implementations, the model's output on the historical data is first calculated using a historical training dataset through forward propagation; then, the gradient of the loss function (e.g., mean squared error loss) with respect to each parameter of the model is calculated; finally, the diagonal elements of the Fisher information matrix are approximately equal to the expected value of the squared gradient of that parameter.

[0129] Specifically, for the i-th parameter in the model Fisher information It can be approximated as the average of the squared gradient of this parameter over historical data, i.e. Where N is the number of historical data samples, It is the loss of the nth sample.

[0130] Step S450: Based on the effective new samples and the Fisher information matrix, construct the elastic weight consolidation loss function. The elastic weight consolidation loss function includes a new sample loss term and a regularization term. The regularization term applies an importance-weighted penalty to each model parameter of the bias correction mapping model based on the Fisher information matrix.

[0131] The Elastic Weight Consolidation Loss function is a special type of loss function used to balance the fit to new data and the retention of old knowledge when the model learns new data. This loss function consists of two parts: a new sample loss term and a regularization term.

[0132] The new sample loss term measures the model's prediction error on valid new samples; for example, mean squared error loss can be used. ,Right now Where M is the number of pixels in the new sample. It is the true spatial distance of the m-th pixel. It is the spatial distance predicted by the model.

[0133] The regularization term is crucial for preventing catastrophic forgetting; it is constructed based on the Fisher information matrix calculated in step S440. Specifically, the regularization term... The form is ,in It is a hyperparameter used to control the strength of the regularization term; It is a parameter The corresponding Fisher information value; It is a parameter The old values ​​(i.e., the current model parameter values) after training on historical data.

[0134] In some implementations, the entire elastic weight consolidation loss function It can be represented as .

[0135] Step S460: Update the model parameters of the correction mapping model using gradient descent based on the elastic weight consolidation loss function, and use the updated correction mapping model for ranging correction of the next frame image.

[0136] Gradient descent is a commonly used algorithm for optimizing model parameters. The elastic weight consolidation loss function is constructed based on step S450. Calculate the loss function with respect to all parameters of the model. gradient Then, the model parameters are updated using gradient descent methods (such as stochastic gradient descent (SGD) or the Adam optimizer), with the update rule being... ,in It is the learning rate.

[0137] In some implementations, the update process can iterate for multiple cycles until the loss function converges or a preset number of iterations is reached. After the update is complete, the new model parameters... Will replace the old parameters This yields an updated correction mapping model. This updated model will be saved and used to process the next frame of the acquired image for ranging and correction.

[0138] Based on the above technical solution, this embodiment determines the quality of updated data by using reprojection error and physical reference points to identify valid new samples, and calculates the Fisher information matrix to quantify the historical importance of model parameters. This allows for the construction of an elastic weight consolidation loss function that includes a new sample loss term and an importance-weighted regularization term. This achieves the protection of old knowledge while learning new knowledge. Finally, the model parameters are updated through gradient descent, enabling the model to continuously track the time-varying characteristics of the device. This allows the elastic weight consolidation strategy to be applied to the field of mobile real-scene ranging and error correction, solving the problems of traditional static models being unable to adapt to dynamically changing environments and the catastrophic forgetting that online learning can easily lead to. This systematically improves the reliability and accuracy of the system under actual long-term, multi-condition operation.

[0139] Example 5: like Figure 5-6 As shown in some embodiments, this embodiment describes in detail the motion blur detection and shading correction process based on spatial nodes. This process provides a specific and implementable solution for "detecting motion blur state using spatial nodes as reference anchor points" and "performing shading correction based on motion blur state" involved in step S300 of Embodiment 1. This process aims to establish a bidirectional connection between the shading layer and a unified metric reference system, achieving reliable shading without external sensors (such as IMUs), thereby effectively suppressing motion blur caused by handheld shooting, restoring image sharpness, and providing high-quality image data for subsequent ranging correction and 3D modeling.

[0140] Step S510: Obtain the current frame image and the previous frame image, and detect the pixel position of the spatial node.

[0141] In some implementations, detecting the pixel position of spatial nodes refers to automatically identifying the pixel coordinates corresponding to these spatial nodes in the current frame image and the previous frame image using image processing algorithms (such as corner detection and feature point matching).

[0142] Specifically, since spatial nodes in an image are represented as regularly distributed corner points or feature points, they can be accurately located using sub-pixel level corner detection algorithms (such as Shi-Tomasi corner detection) or template matching-based methods.

[0143] For example, template image blocks of spatial nodes under ideal perspective transformation can be pre-stored, and then the node positions can be located in the current frame and the previous frame image by sliding window matching.

[0144] Step S520: Calculate the displacement vector of the spatial node between adjacent frames. The displacement vector is the difference between the pixel position of the spatial node in the current frame image and the pixel position in the previous frame image.

[0145] The displacement vector represents the difference in pixel coordinates of the same spatial node in two adjacent frames of the image.

[0146] In some implementations, for each spatial node i, its displacement vector It can be calculated as follows: ,in It is the pixel coordinate of node i in the current frame t. It is the pixel coordinate of node i in the previous frame t-1.

[0147] This allows for further analysis of the distribution characteristics of these displacement vectors to determine the overall motion pattern of the camera. For example, if the magnitude and direction of the displacement vectors of all nodes are basically the same, it indicates that the camera has undergone global translational motion; if the displacement vectors exhibit a rotational pattern centered on the image center, it indicates that the camera has undergone rotational motion.

[0148] Step S530: Based on the displacement vector of the spatial node, calculate the statistics of the inter-frame displacement vector. The statistics include at least one of the displacement mean, displacement standard deviation and maximum displacement value.

[0149] In some implementations, the displacement vectors of all N spatial nodes are collected. The mean displacement can then be calculated separately. standard deviation of displacement and maximum displacement value .

[0150] Among them, the mean displacement It reflects the overall translation trend of the camera, and the calculation formula is the arithmetic mean of the displacement vectors of all nodes; displacement standard deviation This reflects the dispersion of the displacement vector; a larger standard deviation indicates more uneven motion (possibly due to rotation or scaling); the maximum displacement value... It reflects the extreme situations of the sport.

[0151] Step S540: Compare the statistics with a preset threshold, and determine the jitter type and jitter intensity vector of the current frame image based on the comparison result, which is denoted as motion blur state.

[0152] The preset threshold is a value pre-set based on the system's accuracy requirements and typical jitter patterns, such as the average displacement threshold. It can be set to 2 pixels, displacement standard deviation threshold. It can be set to 1 pixel.

[0153] Shake type refers to the classification of camera motion patterns, such as no shake, global translation shake, rotation shake, zoom shake, or mixed shake. The shake intensity vector is a vector that quantifies the degree of shake, and its components can include translation intensity, rotation intensity, zoom intensity, etc. In some implementations, the determination logic can be as follows: If the mean displacement The modulus is less than and displacement standard deviation Less than If so, it is determined to be without jitter; If the mean displacement The model is longer than and displacement standard deviation Less than If so, it is determined to be global translation jitter. The jitter intensity vector can be set as If the displacement standard deviation Greater than If so, it may be determined as rotation or scaling jitter, requiring further analysis of the distribution pattern of the displacement vector to distinguish it.

[0154] For example, if the displacement vectors exhibit a tangential distribution centered on the image center, it is determined to be rotational jitter; if the displacement vectors exhibit a radial distribution from the center outwards, it is determined to be scaling jitter.

[0155] Step S550: Analyze the jitter type and jitter intensity vector in the motion blur state, and initialize the search space of the blur kernel.

[0156] The blur kernel (also known as the point spread function, PSF) is a mathematical model that describes the motion blur process, representing the shape of an ideal point light source diffused across an image during the exposure time.

[0157] The search space refers to the range of possible values ​​for the fuzzy kernel parameter.

[0158] In some implementations, after parsing the jitter type and intensity vector, different blur kernel models can be initialized according to different jitter modes. For example, for global translation jitter, the blur kernel can be initialized as a linear motion blur kernel, whose direction is consistent with the displacement mean vector and whose length is proportional to the translation intensity. For rotational jitter, the blur kernel can be initialized as an arc-shaped motion blur kernel, the curvature of which is related to the rotation intensity; for scaling jitter, the blur kernel can be initialized as a radial blur kernel.

[0159] Specifically, the process of initializing the search space can first select the parameterized model of the fuzzy kernel (such as linear model, arc model, radial model) according to the jitter type, and then determine the initial values ​​of the model parameters and the search range according to the jitter intensity vector.

[0160] Step S560: Based on the search space, using the current frame image and the jitter intensity vector as input, estimate the target blur kernel of the current frame image.

[0161] The target blur kernel refers to the blur kernel found in the search space that best describes the motion blur process of the current frame image.

[0162] In some implementations, the estimation process can employ gradient-based optimization methods. Specifically, a loss function can be defined to measure the degree of matching between the blur kernel and the image, such as minimizing the deconvolution residual: ,in It is the blurred image of the current frame. It is the fuzzy kernel to be estimated. It is an estimated clear image. 1 is the regularization coefficient. It is an L1 norm regularization term used to promote the sparsity of the fuzzy kernel.

[0163] The optimization process can begin with the search space initialized in step S550, iteratively updating the fuzzy kernel h using algorithms such as gradient descent, conjugate gradient, or expectation maximization (EM) until the loss function converges or the maximum number of iterations is reached. For example, iterative estimation can be performed using the Richardson-Lucy algorithm or a variant of the Wiener filter.

[0164] Step S570: Obtain the theoretical pixel positions of spatial nodes in the unified metric reference frame, and perform non-blind deconvolution on the current frame image using the target blur kernel.

[0165] The theoretical pixel position of a spatial node refers to the ideal position on the image plane that the spatial node should appear, calculated based on the state of the unified metric reference system (including the spatial scale reference and the time-varying state parameter vector).

[0166] Non-blind deconvolution refers to the process of deconvolving a blurred image with a known blur kernel to recover a clear image.

[0167] In some implementations, the three-dimensional coordinates of the spatial nodes in the world coordinate system can be obtained from a unified metric reference frame, which allows the use of the camera intrinsic parameter matrix in the current time-varying state parameter vector. and external references The theoretical pixel positions are obtained by projecting the three-dimensional coordinates onto the image plane using a camera projection model. The projection formula can be expressed as: ,in These are the homogeneous three-dimensional coordinates of the spatial nodes. It is an extrinsic parameter matrix.

[0168] Step S580: In response to the deconvolution activation signal, construct a deconvolution objective function containing data fidelity terms and metric constraint terms, using the theoretical pixel positions of spatial nodes as constraints.

[0169] The deconvolution activation signal refers to the conditional signal that triggers the execution of non-blind deconvolution, such as when step S540 determines that jitter exists, it is automatically activated.

[0170] The data fidelity term is a term in the deconvolution objective function that measures the consistency between the restored image and the original blurred image.

[0171] The metric constraint is a special term introduced in this application, which is used to force that the pixel position of the spatial node in the restored clear image should be consistent with the theoretical pixel position.

[0172] In some implementations, the deconvolution objective function can be constructed as: The first item is the data fidelity item, which ensures that a clear image can be approximately restored to the original blurred image after passing through a blur kernel convolution. Specifically, it includes: The blurred image of the current frame. The desired clear image, It is a known target fuzzy kernel. It's a convolution operation. It is the mean square error (L2 norm squared).

[0173] The second term is the metric constraint term, which injects the geometric constraints of the unified metric reference system into the deconvolution process by penalizing the deviation between the node positions in the sharp image and their theoretical positions. Specifically, it includes: These are the weighting coefficients for the metric constraint terms; It is the summation symbol. The actual pixel position of the i-th spatial node in the reconstructed sharp image. It is the theoretical pixel position of the i-th spatial node in a unified metric reference frame. It is the sum of squares of geometric deviations.

[0174] Step S590: Iteratively optimize the deconvolution objective function, and output a clear image when the convergence condition is met.

[0175] Iterative optimization refers to gradually adjusting a clear image using numerical algorithms. The pixel values ​​make the deconvolution objective function The process of its value continuously decreasing.

[0176] Convergence criteria refer to the criteria for stopping optimization, such as the change in the objective function value being less than a certain threshold, or reaching the maximum number of iterations.

[0177] In some implementations, the optimization process can employ gradient descent, conjugate gradient, or alternating direction multiplier (ADMM) methods. Taking gradient descent as an example, each iteration calculates the objective function with respect to... gradient Then follow the update rules Update, among which It is the learning rate.

[0178] The iterative process continues until... ( (This is either a preset small positive number) or the iteration count has reached its limit. The final output is... This refers to a clear image that has had motion blur removed.

[0179] Based on the above technical solution, this embodiment establishes the foundation for inter-frame motion analysis by acquiring adjacent frame images and detecting spatial node positions. It calculates displacement vectors and statistics to quantitatively detect jitter type and intensity, thereby analyzing the motion blur state and initializing the blur kernel search space, improving the efficiency of blur kernel estimation. The target blur kernel is then estimated to obtain key parameters describing motion blur, and the theoretical positions of spatial nodes are acquired. A deconvolution objective function containing metric constraints is constructed, injecting the geometric constraints of a unified metric reference system into the shaking process. Finally, a clear image is output through iterative optimization. This results in a shaking correction framework with a pre-positioned shaking correction layer as its core and bidirectional linkage with a unified metric reference system. This allows for jitter detection using spatial nodes as reference anchors and non-blind deconvolution with the theoretical positions of spatial nodes as constraints, achieving reliable shaking without external sensors, improving the accuracy of shaking correction, and enabling the system to effectively suppress motion blur caused by handheld shooting, recover high-frequency texture information of the image, and provide high-quality image data for subsequent feature point detection, ranging correction, and 3D modeling, solving the problems of feature loss and reconstruction holes caused by motion blur.

[0180] Example 6: like Figure 8 As shown in some embodiments, this embodiment details the architecture and training of the deep learning end-to-end intrinsic parameter estimation network, as well as its collaborative process with the shaking layer. This process provides a specific and implementable solution for the "output ranging and correction results" involved in step S500 of Embodiment 1 and the real-time performance of the system. This process aims to estimate camera intrinsic parameters, distortion coefficients, and extrinsic parameters in real time at the native resolution of the mobile phone using a lightweight neural network, achieving on-site instant quality feedback, and improving the accuracy of parameter estimation by utilizing the auxiliary information output by the shaking layer, thereby shortening the acquisition feedback cycle from the traditional several days to minutes.

[0181] Step S610: Input the clear image into the intrinsic parameter estimation network, and keep the clear image at the original resolution of the mobile phone camera.

[0182] The clear image is the image output after shaking correction in Example 5, in which motion blur has been effectively suppressed.

[0183] Intrinsic parameter estimation networks are deep learning neural networks used to directly regress the camera's intrinsic parameter matrix, distortion coefficients, and extrinsic parameters from a single image.

[0184] Maintaining the phone camera's native resolution means not downsampling or scaling the image, but directly inputting it into the network at its original pixel size (e.g., 4000x3000 pixels). This preserves all high-frequency texture details and geometric information in the image.

[0185] In some implementations, the shape of the sharp image tensor received by the input layer can be represented as [B,C,H,W], where B is the batch size, C is the number of channels (usually 3, corresponding to RGB), and H and W are the height and width of the image, i.e., the native resolution.

[0186] Step S620: Extract multi-scale feature maps of the clear image through the convolutional backbone network of the intrinsic parameter estimation network.

[0187] Among them, the convolutional backbone network is the core module responsible for feature extraction in the intrinsic parameter estimation network.

[0188] Multi-scale feature maps refer to the set of feature maps output by the backbone network at different levels (or different stages) with different spatial resolutions and semantic abstraction levels.

[0189] In some implementations, the convolutional backbone network can adopt a lightweight architecture design to adapt to the computing resources of mobile devices. For example, EfficientNet-B0 can be used as the backbone network, and a composite scaling method can be used to balance the depth, width, and resolution of the network, maintaining high accuracy while having low computational complexity.

[0190] Specifically, EfficientNet-B0 consists of multiple stages, each containing several convolutional layers, batch normalization layers, and activation functions. After a sharp image is input, it passes through these stages sequentially, with each stage outputting a feature map of a specific scale. For example, the feature map output from the first stage has a spatial resolution of 1 / 2 that of the original image, the second stage 1 / 4, and so on. Ultimately, multiple feature maps of different scales can be extracted, such as feature maps with resolutions of H / 8 x W / 8, H / 16 x W / 16, and H / 32 x W / 32. These multi-scale feature maps collectively constitute a hierarchical representation of the sharp image, from local details to global structure.

[0191] Step S630: Based on the multi-scale feature map, regress the camera intrinsic parameter matrix and distortion coefficients through global average pooling and fully connected layers.

[0192] Global average pooling is an operation that averages the feature map across spatial dimensions (height and width) to obtain a fixed-length feature vector.

[0193] A fully connected layer is a neural network layer in which each neuron is connected to all neurons in the previous layer, and is used to map feature vectors to the output space.

[0194] In some implementations, one or more feature maps with strong semantic information are selected from the multi-scale feature maps (for example, the highest-level feature map, which has the lowest spatial resolution but the most abstract semantics). Global average pooling is then performed on these feature maps to obtain multiple feature vectors. These feature vectors are then concatenated to form a comprehensive feature vector. Finally, this comprehensive feature vector is input into one or more fully connected layers.

[0195] The first fully connected layer maps the feature vector to an intermediate dimension, and the second fully connected layer outputs the four parameters of the intrinsic parameter matrix respectively. , , ,, The output intrinsic parameter matrix and distortion coefficients are used as inputs for subsequent pose estimation and state updates.

[0196] Step S640: Based on the multi-scale feature map and the camera intrinsic parameter matrix, estimate the camera extrinsic parameters through the cross-attention mechanism of the intrinsic parameter estimation network.

[0197] Cross-attention is a core component of the Transformer architecture, which allows one sequence (query sequence) to focus on relevant information in another sequence (key value sequence).

[0198] The camera extrinsic parameters include the rotation matrix R and the translation vector T, which are used to describe the camera's pose in the world coordinate system.

[0199] In some implementations, the camera intrinsic parameter matrix obtained from the regression in step S630 is first encoded to form an intrinsic parameter feature vector. This intrinsic parameter feature vector is used as a query, and the multi-scale feature map (after appropriate processing, such as flattening or pooling) is used as the key and value. Then, through the cross-attention mechanism, the similarity between the query and the key is calculated, and the values ​​are weighted and summed to obtain an extrinsic parameter feature vector that integrates intrinsic parameter information and image features. This extrinsic parameter feature vector is then input into a regression head (e.g., composed of fully connected layers), which outputs a rotation matrix R (usually represented by axis angles or quaternions, and then converted to a rotation matrix) and a translation vector T.

[0200] The rotation matrix R is a 3x3 orthogonal matrix that describes the camera's orientation; the translation vector T is a 3-dimensional vector that describes the camera's position in the world coordinate system.

[0201] Step S650: Obtain the clear image and jitter intensity vector of the jitter correction output.

[0202] The jitter intensity vector is the output of step S540 in Example 5, which is used to quantify the jitter level of the current frame image. Its components may include translation intensity, rotation intensity, scaling intensity, etc.

[0203] Step S660: Use the clear image as the main input to the intrinsic parameter estimation network, and use the jitter intensity vector as auxiliary information to input the conditional coding layer of the intrinsic parameter estimation network.

[0204] Among them, the conditional coding layer is a special module in the intrinsic parameter estimation network, which is used to encode auxiliary information (such as jitter intensity vector) into conditional features that can be fused with the main features.

[0205] In some implementations, the conditional coding layer can be a small multilayer perceptron or a series of fully connected layers. When the jitter intensity vector is input into the conditional coding layer, it maps it to a fixed-length conditionally encoded feature vector. This feature vector encapsulates information about the camera's motion state.

[0206] Step S670: Map the jitter intensity vector to conditional coding features through the conditional coding layer, and fuse the conditional coding features with the multi-scale feature map output by the convolutional backbone network.

[0207] Feature fusion refers to the process of combining feature representations from two or more different sources to form a richer and more discriminative feature representation.

[0208] In some implementations, there can be multiple fusion methods. For example, additive fusion can be used, which involves broadcasting the conditionally encoded feature vector to the same spatial dimension as the multi-scale feature map and then adding it directly to the multi-scale feature map; concatenation fusion can also be used, which involves broadcasting the conditionally encoded feature vector and then concatenating it with the multi-scale feature map in the channel dimension; more complex gating fusion or attention fusion mechanisms can also be used.

[0209] Specifically, taking additive fusion as an example, assuming that the shape of a certain layer in the multi-scale feature map is [B,C,H,W] and the shape of the conditional coding feature vector is [B,D], the conditional coding feature vector is first mapped to the number of channels C through a fully connected layer to obtain a vector with shape [B,C]. Then, it is reshaped into [B,C,1,1] and broadcast to [B,C,H,W]. Finally, it is added to the original multi-scale feature map.

[0210] Step S680: Based on the fused features, regress the camera intrinsic parameter matrix and distortion coefficients through global average pooling and fully connected layers.

[0211] In some implementations, the regression process is consistent with step S630, that is, global average pooling is performed on the fused feature map to obtain the feature vector, and then the intrinsic parameter matrix and distortion coefficients are output through a fully connected layer.

[0212] It should be noted that since the input features have already incorporated camera motion information, the intrinsic parameter matrix and distortion coefficients obtained from the regression may be more accurate, especially in situations with camera shake. This is because the network can use motion information to correct for estimation biases introduced by shake.

[0213] Step S690: Based on multi-scale feature maps, camera intrinsic matrix and conditional coding features, estimate camera extrinsic parameters through cross-attention mechanism.

[0214] In some implementations, in addition to using intrinsic parameter matrix encoding as the query, conditional encoding features can also be incorporated into the query, or conditional encoding features can be used as additional key-value pair inputs for cross-attention mechanisms.

[0215] One possible implementation is to concatenate or add the intrinsic parameter matrix encoding and the conditional encoding features to form a comprehensive query vector. This query vector is then cross-attentioned with the multi-scale feature map (as keys and values) to obtain the extrinsic parameter feature vector. The rotation matrix R and translation vector T can then be output through the regression head.

[0216] Step S6100: Output the camera intrinsic parameter matrix, distortion coefficients, and camera extrinsic parameters, and use the camera intrinsic parameter matrix, distortion coefficients, and camera extrinsic parameters to update the time-varying state parameter vector.

[0217] The output intrinsic parameter matrix, distortion coefficients, and extrinsic parameters are the final estimation results of the intrinsic parameter estimation network for the current frame of the sharp image.

[0218] In some implementations, the above output will be used to update the time-varying state parameter vector in the unified metric reference frame. Specifically, the update process may involve: converting the output intrinsic parameter matrix... Distortion coefficient and external references It is fused or replaced with the corresponding components K(t), D(t), and E(t) in the current time-varying state parameter vector S(t).

[0219] For example, a weighted average can be used to combine the new estimate with the old one, and the weights can be set according to the confidence level of the estimate.

[0220] It should be noted that by using an end-to-end neural network to estimate and update the time-varying state parameter vector in real time, on-site real-time quality feedback is achieved. After each photo is taken, the system can immediately output the estimated camera parameters and pose, allowing the operator to judge the acquisition quality on-site. If abnormal parameters (such as excessive intrinsic parameter drift) or unreasonable pose are found, reshoots can be performed immediately.

[0221] To more clearly illustrate the training process of the intrinsic parameter estimation network, the following supplementary explanation is provided.

[0222] In some implementations, the intrinsic parameter estimation network can be trained using a curriculum learning strategy. It is first pre-trained on a large-scale synthetic dataset (such as TartanAir), which provides accurate ground truth values ​​for camera parameters, helping the network learn basic geometric mappings. Then, it is fine-tuned on real-world scene datasets (such as ScanNet and a self-built cultural wall image dataset) to adapt the network to real-world image noise, lighting variations, and complex scenes. The loss function used during training can be a multi-task loss, for example:

[0223] in, It is the mean squared error loss of the intrinsic parameter estimation. It is the mean squared error loss of the distortion coefficient estimation. It is the geometric loss of pose estimation (e.g., the logarithmic mapping error of the rotation matrix plus the Euclidean distance error of the translation vector). It is the reprojection error loss, which projects three-dimensional spatial points onto the image plane and calculates the pixel distance between the projected points and the detected feature points. These are the weighting coefficients for each loss, used to balance the importance of different tasks.

[0224] Based on the above technical solution, this embodiment preserves all the detailed information of the image by maintaining the original resolution input. Combined with multi-scale features extracted by the convolutional backbone network, a hierarchical image representation is established. This allows for direct mapping from features to parameters through global average pooling and fully connected layers to regress intrinsic parameters and distortion coefficients. Simultaneously, extrinsic parameters are estimated through a cross-attention mechanism, fusing intrinsic parameter information. The shake intensity vector output from the shaking layer is then used as auxiliary information, input into the conditional coding layer, and fused with features, achieving the distinction between camera motion and scene structure changes. Finally, real-time output and updating of the time-varying state parameter vector enables immediate on-site quality feedback. This collectively constructs a lightweight, high-precision, motion-aware real-time intrinsic parameter estimation framework. Injecting motion information from the shaking layer as conditional encoding into the network allows parameter estimation to adapt to handheld shooting conditions, thereby improving the system's real-time performance, accuracy, and robustness in practical applications, providing high-quality real-time parameter support for cultural wall modeling and design.

[0225] Example 7: like Figure 9 As shown, this embodiment provides a cultural wall modeling and design system for measuring and correcting distances from real-scene photos taken with a mobile phone. This system is applied to the methods in the aforementioned embodiments and aims to achieve a complete process from image acquisition to accurate modeling output through the collaboration of the communication unit and the processing unit.

[0226] Its communication unit is an interface module for the system to interact with external data sources, used to obtain the current frame image of the target scene and the corresponding shooting status parameters.

[0227] In some implementations, the communication unit can be one or more hardware interfaces and corresponding drivers. For example, it can include an image sensor interface for connecting to the mobile phone camera module, a data bus interface for reading mobile phone system API or image EXIF ​​metadata, and a wireless or wired network interface for communicating with external storage devices or networks.

[0228] Specifically, when the system is deployed on a mobile device (such as a smartphone), the communication unit can directly call the phone's camera hardware and system services to acquire images and parameters; when the system is deployed on a server, the communication unit can receive image data packets and parameter information uploaded from the mobile device via the network.

[0229] The processing unit is the core computing module of the system, responsible for executing all data processing, analysis, and decision-making logic. It performs ranging and correction on the current frame image based on a unified metric reference system and a correction mapping model to obtain corrected coordinates and the true spatial distance. Using spatial nodes in the unified metric reference system as reference anchors, it detects the motion blur state of the current frame image and performs shaking correction based on the motion blur state to obtain a clear image. The correction amount generated by the shaking correction is fed back to update the time-varying state parameter vector of the unified metric reference system. Based on the clear image and the updated time-varying state parameter vector, it outputs the ranging and correction results of the target scene and the modeling design method.

[0230] In some implementations, the processing unit may consist of one or more processors (such as CPU, GPU, NPU) and software programs running on them. Specifically, the processing unit may be further divided into multiple functional sub-modules, which correspond to the steps in the aforementioned method embodiments, forming a pipelined processing architecture.

[0231] For example, the processing unit may include a ranging and correction module, corresponding to step S200 in the method, which is used to load and execute the correction mapping model, mapping the image pixel coordinates and shooting state parameters obtained by the communication unit into corrected coordinates and true spatial distance. The implementation of this module relies on a unified metric reference system, which can be stored in a dedicated reference system storage module. This storage module maintains the spatial scale benchmark, the time-varying state parameter vector, and the multi-factor influence relationship matrix.

[0232] The processing unit may further include a jitter correction module, corresponding to steps S300 and S400 in the method. This module uses spatial nodes in the reference frame storage module as reference anchor points, detects motion blur by analyzing the displacement of spatial nodes in adjacent frame images, and performs jitter correction. After correction, the module feeds back the generated correction amount to the reference frame storage module to update the time-varying state parameter vector, forming a closed loop.

[0233] The processing unit may also include an intrinsic parameter estimation module, corresponding to steps S610 to S6100 in the method. This module receives the sharp image and jitter intensity vector output by the shaking correction module, and estimates the camera's intrinsic parameter matrix, distortion coefficients, and extrinsic parameters in real time through its internal convolutional backbone network, conditional coding layer, global average pooling layer, fully connected layer, and cross-attention mechanism. These estimation results are then output to update the time-varying state parameter vector in the reference frame storage module, or directly used for the final modeling output.

[0234] Finally, the processing unit also includes an output module, corresponding to step S500 in the method. This module integrates the results from the ranging and correction module and the intrinsic parameter estimation module to generate the final ranging and correction results (e.g., corrected 3D point cloud or distance map) and modeling design methods (e.g., parameter suggestions for guiding 3D reconstruction), and outputs them to the user via the communication unit or directly through the user interface.

[0235] It should be understood that, although Figure 1 The diagram illustrates the linear connections between modules within the processing unit. However, in other embodiments, these modules can be connected using parallel processing or more complex network topologies, as long as efficient data flow and functional synergy between modules are achieved. For example, the jitter correction module and the intrinsic parameter estimation module can process the same frame of image in parallel to further improve processing speed.

[0236] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for modeling and designing cultural walls using mobile phone real-scene photos for distance measurement and correction, characterized in that, include: Obtain the current frame image of the target scene and the corresponding shooting status parameters; Based on a unified metric reference system and a correction mapping model, the current frame image is subjected to ranging and correction to obtain the corrected coordinates and the true spatial distance. The unified metric reference system includes a spatial scale reference with spatial nodes, a time-varying state parameter vector, and a multi-factor coupling relationship. The correction mapping model is used to characterize the mapping relationship between the image pixel coordinates and the shooting state parameters and the corrected coordinates and the true spatial distance. Using the spatial node as a reference anchor point, the motion blur state of the current frame image is detected, and the current frame image is rectified based on the motion blur state to obtain a clear image; The correction amount generated by the jitter correction is fed back to update the time-varying state parameter vector of the unified metric reference system; Based on the clear image and the updated time-varying state parameter vector, the ranging and correction results and modeling design method of the target scene are output.

2. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 1, characterized in that, The process of constructing the unified metric reference system includes: Obtain the precise three-dimensional spatial coordinates of physical reference points deployed in the target scene, wherein the three-dimensional spatial coordinates are calibrated based on the world coordinate system; Based on the three-dimensional spatial coordinates of the physical reference point and the corresponding image pixel coordinates in the current frame image, calculate the homography matrix from the world coordinate system to the image coordinate system; Based on the homography matrix, a virtual perspective transformation grid is generated on the surface of the target scene. The perspective transformation grid contains grid nodes that are evenly distributed. The grid nodes are defined as spatial nodes, and all spatial nodes constitute the spatial scale reference of the unified metric reference system. Based on the spatial scale reference and the shooting time of the current frame image, a time-varying state parameter vector is established. The time-varying state parameter vector includes the camera intrinsic parameter matrix, distortion coefficient, camera extrinsic parameters, ambient light characteristic parameters, automatic exposure parameters, automatic focus status, and lens mode identifier at the current time. The time-varying state parameter vector is spatially aligned and bound to the spatial nodes, and a multi-factor influence relationship matrix is ​​constructed to generate multi-factor coupling relationships; Based on the spatial scale benchmark, the time-varying state parameter vector, and the multi-factor influence relationship matrix, the unified metric reference system is obtained.

3. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 2, characterized in that, The process of constructing a multi-factor influence relationship matrix and generating multi-factor coupling relationships includes: Using the time-varying state parameter vector as the row and column dimensions of the matrix, an initial influence relationship matrix is ​​constructed. The element value in the i-th row and j-th column of the initial influence relationship matrix represents the influence intensity coefficient of the i-th factor on the j-th factor. Based on the historical training dataset, parameter estimation is performed on each influence intensity coefficient in the initial influence relationship matrix to obtain the initialized influence relationship matrix. The historical training dataset contains the observed values ​​of each historical influencing factor, the corresponding image pixel coordinates, and the corresponding target scene spatial ranging ground values. The real-time observation values ​​of each influencing factor at the current moment are obtained, and the initialized influence relationship matrix is ​​updated online based on the real-time observation values ​​to obtain a multi-factor influence relationship matrix. The multi-factor influence relationship matrix is ​​used to characterize the multi-factor coupling relationship between the multiple influencing factors.

4. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 1, characterized in that, The construction process of the correction mapping model includes: Obtain a historical training dataset, which includes historical observations of various influencing factors, corresponding image pixel coordinates, and corresponding ground truth values ​​for spatial ranging of the target scene. The combination of influencing factors in the historical training dataset is analyzed to obtain a set of multi-factor influence relationships. Each influence relationship in the set of multi-factor influence relationships contains the coupling strength coefficient and coupling function form between two influencing factors. An initial correction mapping model is constructed, which is a multilayer perceptron. The input layer of the multilayer perceptron receives the image pixel coordinates and the time-varying state parameter vector, and the output layer outputs the corrected coordinates and the true spatial distance. The parameters of the hidden layer of the multilayer perceptron are initialized based on the set of multifactor influence relationships, and the multilayer perceptron after parameter initialization is trained using the historical training dataset to obtain the bias correction mapping model.

5. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 4, characterized in that, Also includes: In response to the acquisition of the target scene, the current frame image is input into the correction mapping model, and the corrected coordinates and predicted spatial distance of each pixel in the current frame image are output. Based on the known three-dimensional spatial coordinates of the physical reference point in the target scene and the corresponding image in the current frame image, the reprojection error is calculated. If the reprojection error is less than a preset threshold, the current frame image is marked as a valid new sample. Obtain the pixel coordinates of each image in the effective new sample, the corresponding time-varying state parameter vector, and the corresponding spatial ranging ground value; Based on the historical training dataset, calculate the Fisher information matrix of each model parameter in the correction mapping model; Based on the effective new samples and the Fisher information matrix, an elastic weight consolidation loss function is constructed. The elastic weight consolidation loss function includes a new sample loss term and a regularization term. The regularization term applies an importance-weighted penalty to each model parameter of the bias correction mapping model based on the Fisher information matrix. The model parameters of the correction mapping model are updated by gradient descent based on the elastic weight consolidation loss function, and the updated correction mapping model is used for ranging correction of the next frame image.

6. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 1, characterized in that, The process of detecting the motion blur state of the current frame image using the spatial node as a reference anchor point includes: Acquire the current frame image and the previous frame image, and detect the pixel position of the spatial node; Calculate the displacement vector of the spatial node between adjacent frames, where the displacement vector is the difference between the pixel position of the spatial node in the current frame image and the pixel position in the previous frame image; Based on the displacement vector of the spatial node, calculate the statistics of the inter-frame displacement vector, which include at least one of the displacement mean, displacement standard deviation and maximum displacement value. The statistical quantity is compared with a preset threshold, and the jitter type and jitter intensity vector of the current frame image are determined based on the comparison result, which is denoted as the motion blur state.

7. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 1, characterized in that, The process of performing jitter correction on the current frame image based on the motion blur state to obtain a clear image includes: The jitter type and jitter intensity vector in the motion blur state are analyzed, and the search space of the blur kernel is initialized. Based on the search space, and using the current frame image and the jitter intensity vector as input, the target blur kernel of the current frame image is estimated. Obtain the theoretical pixel positions of spatial nodes in the unified metric reference system, and perform non-blind deconvolution on the current frame image using the target blur kernel; In response to the deconvolution activation signal, a deconvolution objective function containing a data fidelity term and a metric constraint term is constructed, with the theoretical pixel position of the spatial node as a constraint. The deconvolution objective function is iteratively optimized, and a clear image is output when the convergence condition is met.

8. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 1, characterized in that, Also includes: The sharp image is input into the intrinsic parameter estimation network, and the sharp image maintains the original resolution of the mobile phone camera; The convolutional backbone network of the intrinsic parameter estimation network extracts multi-scale feature maps of the sharp image; Based on the multi-scale feature map, the camera intrinsic parameter matrix and distortion coefficients are regressed through global average pooling and fully connected layers. Based on the multi-scale feature map and the camera intrinsic parameter matrix, the camera extrinsic parameters are estimated through the cross-attention mechanism of the intrinsic parameter estimation network. The camera intrinsic parameter matrix, the distortion coefficients, and the camera extrinsic parameters are output. The camera intrinsic parameter matrix, the distortion coefficients, and the camera extrinsic parameters are used to update the time-varying state parameter vector.

9. The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction according to claim 8, characterized in that, Also includes: Obtain the clear image and jitter intensity vector output by the jitter correction; The clear image is used as the main input to the intrinsic parameter estimation network, and the jitter intensity vector is used as auxiliary information input to the conditional coding layer of the intrinsic parameter estimation network. The jitter intensity vector is mapped to conditional coding features through the conditional coding layer, and the conditional coding features are fused with the multi-scale feature map output by the convolutional backbone network. Based on the fused features, the camera intrinsic parameter matrix and the distortion coefficients are regressed through the global average pooling and the fully connected layer. Based on the multi-scale feature map, the camera intrinsic parameter matrix, and the conditional coding features, the camera extrinsic parameters are estimated through the cross-attention mechanism. Output the camera intrinsic parameter matrix, the distortion coefficients, and the camera extrinsic parameters, and use the camera intrinsic parameter matrix, the distortion coefficients, and the camera extrinsic parameters to update the time-varying state parameter vector.

10. A cultural wall modeling and design system for distance measurement and correction using mobile phone real-scene photos, characterized in that, The method for modeling and designing a cultural wall using mobile phone real-scene photos for distance measurement and correction, as described in any one of claims 1-9, specifically includes: The communication unit is used to acquire the current frame image of the target scene and the corresponding shooting status parameters; The processing unit is used to perform ranging and correction on the current frame image based on a unified metric reference system and a correction mapping model to obtain corrected coordinates and true spatial distance; and to detect the motion blur state of the current frame image using spatial nodes in the unified metric reference system as reference anchor points, and to perform anti-shake correction on the current frame image based on the motion blur state to obtain a clear image. The correction amount generated by the jitter correction is fed back to update the time-varying state parameter vector of the unified metric reference system; Based on the clear image and the updated time-varying state parameter vector, the ranging and correction results and modeling design method of the target scene are output.