Marine park virtual scene construction method and system based on three-dimensional model
By unifying coordinate mapping and observation confidence decay, and combining cross-view structural stability and brightness fluctuation characterization, pseudo-geometric risk results are generated, and a scene state co-occurrence relationship graph is constructed. This solves the problem of pseudo-structure in the Ocean Park virtual scene, and achieves more stable global mesh reconstruction, which is suitable for Ocean Park virtual roaming and digital asset construction.
Patent Information
- Application Number
- CN202610643893.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-28
AI Technical Summary
Existing technologies struggle to effectively handle pseudo-structures caused by factors such as water disturbance, mirror reflection from acrylic observation windows, highlights from wet rock walls, and spray scattering when constructing virtual scenes of marine parks. This results in localized bulges, floating interlacing, mis-filled holes, and boundary distortion in the 3D reconstruction results, affecting the consistency of virtual roaming and interactive displays.
By employing unified coordinate mapping, observation confidence decay, and local structural fragment construction, combined with cross-view structural stability, interface perturbation characterization, and brightness fluctuation characterization, a set of pseudo-geometric candidate fragments and risk results are generated. Through scene state co-occurrence relationship graphs and mutual exclusion constraints, joint fragment screening and global mesh reconstruction are performed to form a stable set of scene fragments.
It effectively reduces the impact of water interfaces, acrylic viewing windows, and wet surface highlights on the original input, improves the accuracy of pseudo-geometry determination, accurately distinguishes between fixed scene segments and movable state segments, and generates global mesh reconstruction results with more stable structural continuity and state closure, making it suitable for virtual tours of marine parks and digital asset construction.
Smart Images

Figure CN122473354A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of 3D scene reconstruction technology, and in particular to a method and system for constructing virtual scenes of marine parks based on 3D models. Background Technology
[0002] Virtual scene construction in ocean parks typically serves applications such as immersive guided tours, remote previews, interactive educational demonstrations, and digital asset management for the park. The objects involved are not single static buildings, but rather complex spaces encompassing viewing corridors, curved acrylic observation windows, underwater viewing cabins, performance pools, wet rock formations, artificial reefs, misting systems, guide railings, and dynamic lighting installations. To enhance the realism and roamability of virtual scenes, existing technologies typically employ multi-view image acquisition, laser point cloud acquisition, or image-point cloud fusion modeling to perform 3D reconstruction of the park's structure, landscape facilities, and display interfaces. This is combined with processing steps such as dense matching, point cloud stitching, voxel filtering, random sampling consistency registration, and mesh reconstruction to form the scene model. This type of method is relatively mature in stable environments such as museums, factories, exhibition halls and conventional outdoor buildings. For example, it can achieve good reconstruction results for objects such as walls, columns, ground and fixed exhibition stands. However, in marine park scenes, there are complex factors such as blue light illumination, water surface reflection, ripple caustics, glass reflection, fog scattering, wet surface highlights and time-segmented acquisition conditions, which make the original observation data exhibit stronger uncertainty in both spatial structure and temporal state.
[0003] When performing 3D reconstruction for the aforementioned scenarios, existing technologies typically assume that cross-view observation differences mainly originate from changes in viewpoint and treat the reconstructed object as remaining relatively stationary during the acquisition period. This allows them to focus on routine issues such as outlier removal, local registration correction, and model surface smoothing. However, in practical applications at marine parks, on the one hand, water disturbances, specular reflections from acrylic observation windows, highlights from wet rock walls, and spray scattering can create pseudo-structures with continuous edges or surface textures in local areas. This can cause algorithms to misinterpret these as real geometry and write them into point clouds or meshes. On the other hand, the same exhibit area often requires data acquisition at different times and along multiple paths. Changes in feeding platform relocation, temporary installation of maintenance fencing, raising and lowering of performance devices, and changes in the occlusion status of swimming organisms can be incorrectly stitched into the same scene version, resulting in hybrid models with geometrically closed surfaces but incompatible temporal states. Existing technologies mostly focus on point-level noise removal and frame-level registration, lacking a joint judgment mechanism for the long-term stability of the structure and the rationality of state co-occurrence. This leads to the construction results being prone to local bulges, floating and interpenetrating, mis-filling holes and boundary distortion, which in turn affects the consistency of virtual roaming route generation, collision area delineation, interactive object binding and subsequent rendering and display. Summary of the Invention
[0004] This application proposes a method and system for constructing virtual scenes of marine parks based on 3D models, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, this application adopts the following technical solution: a method for constructing a virtual scene of a marine park based on a 3D model, comprising: S1. Acquire multi-source observation data of the marine park scene, perform unified coordinate mapping, observation confidence decay and local structure fragment construction, and form observation support information and standardized observation description based on the depth, curvature, edge gradient and brightness information corresponding to each local structure fragment. S2. Based on standardized observation description, calculate the cross-view structural stability results of each local structural segment between different effective observations, and combine interface perturbation characterization and brightness fluctuation characterization to generate a set of pseudo-geometric candidate segments and pseudo-geometric risk results. S3. Based on the structural stability results and the set of pseudo-geometric candidate segments, construct a scene state co-occurrence relationship graph, extract the spatial relationship consistency, temporal overlap, common observation support and spatial occupancy conflict between local structural segments, and generate state co-occurrence weights and mutual exclusion constraints. S4. Based on the structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutual exclusion constraints, perform joint fragment screening to obtain a set of stable scene fragments, and perform global mesh reconstruction based on the set of stable scene fragments to output the global mesh reconstruction results.
[0006] Furthermore, the multi-source observation data includes multi-view images, laser point clouds, acquisition poses, and timestamps of the marine park scene at multiple acquisition times. Perform unified coordinate mapping, including: establishing the coordinate correspondence between each multi-view image and each laser point cloud relative to the unified scene coordinate system based on the acquired pose, and mapping the observation area in the multi-view image and the spatial points in the laser point cloud to the unified scene coordinate system.
[0007] Furthermore, the observation confidence reduction and local structural fragment construction are performed, and observation support information and standardized observation descriptions are generated for each local structural fragment. This includes: based on the angle between the observation ray and the local surface normal, the confidence reduction or invalid observation removal is performed on the observation results passing through the water interface or acrylic observation window interface; based on the spatial adjacency relationship and normal continuity relationship of the mapped point cloud, local aggregation is performed on the scene point set to form local structural fragments; for each local structural fragment, the corresponding effective observation range and time support range are determined to generate observation support information; for the depth, curvature, edge gradient and brightness of each local structural fragment under the corresponding effective observation, benchmark conversion of the same quantity is performed to generate standardized observation descriptions.
[0008] Furthermore, based on standardized observation descriptions, the cross-view structural stability results of each local structural segment across different valid observations are calculated. This includes: for each local structural segment, extracting the standardized depth result, standardized curvature result, standardized edge gradient result, and normal consistency description corresponding to different valid observations; forming a cross-view structural difference description of the local structural segment based on the degree of depth deviation, curvature deviation, edge gradient deviation, and normal deviation between different valid observations; and performing suppression processing on the cross-view structural difference description by combining the observation confidence decay results corresponding to each valid observation, thereby generating cross-view structural stability results.
[0009] Furthermore, by combining interface perturbation characterization and brightness fluctuation characterization, a set of pseudo-geometric candidate segments and pseudo-geometric risk results are generated, including: for each local structural segment in the corresponding region of the image projection domain, water coverage area, specular reflection coverage area and fog scattering coverage area are identified, and interface perturbation characterization is formed based on the proportion of each coverage area in the corresponding projection range; for each local structural segment under different effective observations, brightness fluctuation characterization is formed based on its offset magnitude relative to the reference brightness. Based on cross-view structural stability results, interface perturbation characterization, and brightness fluctuation characterization, pseudo-geometric risk results are generated; local structural segments whose pseudo-geometric risk results meet preset conditions are written into the pseudo-geometric candidate segment set.
[0010] Furthermore, based on the structural stability results and the pseudo-geometric candidate fragment set, a scene state co-occurrence relationship graph is constructed. The spatial relationship consistency, temporal overlap, common observation support, and spatial occupancy conflict between local structural fragments are extracted to generate state co-occurrence weights, including: using each local structural fragment as a graph node; for spatially adjacent local structural fragment pairs, spatial relationship consistency is extracted based on the stability of the relative positional relationship under multiple common observations, temporal overlap is extracted based on the overlap of the corresponding temporal support range, common observation support is extracted based on the proportion of common effective observations, and spatial occupancy conflict is extracted based on the spatial encirclement overlap in the unified scene coordinate system. By combining spatial relationship consistency, temporal overlap, common observation support, and spatial occupancy conflict, state co-occurrence weights are generated between local structural segments, and a scene state co-occurrence relationship graph is established accordingly.
[0011] Furthermore, generating mutual exclusion constraint relationships includes: for any two local structural segments, if the degree of space occupation conflict between the two segments meets the preset conflict condition and the degree of time overlap meets the preset time incompatibility condition, then the two segments are determined as a pair of mutually exclusive segments and written into the mutual exclusion constraint relationship. For local structural fragments belonging to the pseudo-geometric candidate fragment set, when generating the co-occurrence weights of their participation in the state, the corresponding association strength is attenuated to reduce the impact of pseudo-geometric candidate fragments on the connectivity results of the scene state co-occurrence relationship graph.
[0012] Furthermore, joint fragment selection is performed based on structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutually exclusive constraint relationships, including: for each local structural fragment, forming fragment retention prior results based on the corresponding structural stability results and pseudo-geometric risk results; Based on the co-occurrence weights of states among local structural segments, a common retention constraint is applied to local structural segments with high co-occurrence consistency; based on mutual exclusion constraints, a simultaneous retention constraint is applied to local structural segments that have spatial conflicts and temporal incompatibility; under the combined effect of segment retention prior results, common retention constraints, and simultaneous retention constraints, the retention state of each local structural segment is determined, and a stable set of scene segments is obtained.
[0013] Furthermore, a global mesh reconstruction is performed based on a set of stable scene fragments, and the global mesh reconstruction result is output, including: extracting spatial points and unit normals from the set of stable scene fragments as reconstruction inputs; forming reconstruction weight results based on the prior results of each stable scene fragment; and using a weighted mesh reconstruction method to fuse and reconstruct the spatial points, unit normals and reconstruction weight results to obtain the global mesh reconstruction result.
[0014] A virtual scene construction system for marine parks based on 3D models includes: The multi-source observation processing module acquires multi-source observation data of the marine park scene, performs unified coordinate mapping, observation confidence decay and local structure fragment construction, and forms observation support information and standardized observation description based on the depth, curvature, edge gradient and brightness information corresponding to each local structure fragment. The pseudo-geometric candidate generation module calculates the cross-view structural stability results of each local structural segment between different effective observations based on standardized observation descriptions, and generates a set of pseudo-geometric candidate segments and pseudo-geometric risk results by combining interface perturbation characterization and brightness fluctuation characterization. The state closure construction module constructs a scene state co-occurrence relationship graph based on structural stability results and pseudo-geometric candidate fragment set, extracts the spatial relationship consistency, temporal overlap, common observation support and spatial occupancy conflict between local structural fragments, and generates state co-occurrence weights and mutual exclusion constraints. The joint reconstruction output module performs joint fragment filtering based on structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutual exclusion constraints to obtain a set of stable scene fragments. It then performs global mesh reconstruction based on the set of stable scene fragments and outputs the global mesh reconstruction results.
[0015] The beneficial effects of this invention are as follows: 1. This invention employs a processing approach that combines unified coordinate mapping, observation reliability attenuation, and local structural fragment construction. It first performs fragment-level processing on multi-source observation data in marine park scenes, then generates observation support information and standardized observation descriptions. This effectively reduces the one-time amplification effect of water interfaces, acrylic observation windows, wet surface highlights, and time-division acquisition differences on the original input. This ensures that subsequent processing is based on unified, comparable, and traceable data. Compared to directly reconstructing the original image and point cloud as a whole, this invention reduces perspective distortion, temporal mismatch, and local anomaly propagation during the input stage, improving the stability of subsequent scene analysis.
[0016] 2. This invention employs a joint judgment method combining cross-view structural stability results with interface perturbation characterization and brightness fluctuation characterization to perform pseudo-geometric recognition on local structural fragments. This method can more effectively suppress pseudo-surfaces, pseudo-boundaries, and pseudo-textures induced by water surface reflection, specular reflection, fog scattering, and wet surface highlights, which are common in marine park scenes. Compared to methods that rely solely on point-level denoising, frame-level registration, or single brightness anomaly detection, this invention can simultaneously utilize geometric consistency and optical perturbation features for comprehensive judgment, thereby improving the accuracy of the pseudo-geometric candidate fragment set and pseudo-geometric risk results, and reducing the situation where real structures are mistakenly deleted and pseudo structures are missed.
[0017] 3. This invention employs a version closure mechanism that combines scene state co-occurrence relationship graphs, state co-occurrence weights, and mutual exclusion constraints. It jointly models the spatial relationship consistency, temporal overlap, common observation support, and spatial occupancy conflict between different local structural fragments. This effectively identifies pseudo-stable states caused by time-sharing state differences such as feeding platform displacement, maintenance fence setup, performance prop elevation, and occlusion by swarming creatures. Compared to methods that splice based solely on spatial adjacency or filter based solely on temporal order, this invention can more accurately distinguish between fixed scene fragments and movable state fragments, preventing multiple incompatible states from being incorrectly spliced into the same scene version.
[0018] 4. This invention employs a joint fragment selection mechanism based on structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutually exclusive constraint relationships. Based on this, a stable set of scene fragments and a global mesh reconstruction result are formed. This mechanism simultaneously incorporates local geometric realism constraints and overall version consistency constraints into the final scene generation process. Compared to the linear processing method that sequentially performs denoising, splicing, and reconstruction, this invention reduces the impact of local bulges, floating interlacing, hole mis-patch, and version aliasing on the final model. The resulting global mesh reconstruction result is more stable in terms of structural continuity, state closure, and subsequent reusability, making it more suitable for applications such as virtual roaming displays in marine parks, digital asset construction, and scene interaction deployment. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort: Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system framework diagram of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1 like Figure 1 As shown, this invention discloses a method for constructing a virtual scene of a marine park based on a 3D model, including the following specific steps: In this embodiment, S1 is used to acquire multi-source observation data of the marine park scene, perform unified coordinate mapping, observation confidence decay and local structural fragment construction, and form observation support information and standardized observation description based on the depth, curvature, edge gradient and brightness information corresponding to each local structural fragment. S1 first organizes the original multi-source observation results in the marine park scene into fragment-level inputs that can be directly accepted later, and then provides a unified, continuous and comparable data foundation for the cross-view structural stability result calculation in S2, the scene state co-occurrence relationship diagram construction in S3 and the joint fragment screening in S4. This step does not focus on solving the input instability problem in the marine park scene caused by the water interface, acrylic observation window, wet surface highlights, blue light illumination and time-division acquisition differences.
[0022] First, multi-source observation data of the marine park scene at multiple acquisition times were acquired. This multi-source data includes multi-view images, laser point clouds, acquisition poses, and timestamps. Multi-view images are used to characterize edge textures, brightness distribution, and local highlight areas in the scene. Laser point clouds are used to characterize spatial location, local geometric continuity, and surface normal variations in the scene. Acquisition poses are used to characterize the position and orientation of each observation device relative to a unified scene coordinate system. Timestamps are used to characterize the order in which the observation results were acquired. The marine park scene is preferably selected to include an underwater tunnel, an underwater viewing cabin, an arc-shaped acrylic observation window, the boundary of the performance pool, and wet rocks. For areas near the main body, artificial reef, and misting devices, considering that water ripples, swimming organisms obstructing the view, and performance lighting changes can all alter the observed appearance in a short period, the time alignment error between the image and the point cloud is preferably controlled within 0.05 to 0.20 seconds. Under the condition of fixed track acquisition inside the venue, the time alignment error is preferably no greater than 0.10 seconds. At the edge of the performance pool and in the open viewing area, the time alignment error is preferably no greater than 0.20 seconds. The basis for this setting is that when the time misalignment exceeds 0.20 seconds, the high-frequency reflection on the water surface, the changes in the misting boundary, and the position of swimming organisms will increase significantly, and the stability of subsequent segment-level correspondences will decrease.
[0023] After acquiring multi-source observation data, a unified coordinate mapping is first performed. This can be achieved using extrinsic parameter calibration, pose calculation, and coordinate transformation methods. Specifically, based on the acquired pose, the coordinate correspondence between each multi-view image and each laser point cloud relative to a unified scene coordinate system is established. Then, the observation areas in the multi-view images and the spatial points in the laser point clouds are jointly mapped to the unified scene coordinate system. The unified scene coordinate system preferably uses the fixed ground reference plane of the marine park as the horizontal reference, the gravity direction as the vertical reference, and the direction of the main performance channel or the long side of the main pool as the horizontal reference. After adopting this method, the spatial position, surface orientation, and local adjacency relationship in each observation result are expressed under the same coordinate standard, which facilitates the subsequent construction of local structural segments and the comparison of spatial relationships between segments. To avoid residual misalignment after unified coordinate mapping, the observation window mounting frame point, column corner point, and stable ground reflection point can be selected as verification points. When the mapping deviation of the verification point is greater than twice the average point spacing of the scene, secondary registration or rejection processing is performed on the corresponding observation.
[0024] After completing the unified coordinate mapping, the observation confidence attenuation is performed. This process is one of the core methods in S1. The purpose is to suppress unreliable observations caused by interface refraction and specular reflection at the input end. Specifically, the water distribution area and acrylic observation window area under the unified scene coordinate system are combined to determine whether the image line of sight or point cloud sampling ray passes through the water interface or acrylic observation window interface. Then, based on the angle between the observation ray and the local surface normal, the confidence reduction or invalid observation removal is performed on the corresponding observation results. The local surface normal can be obtained by the local neighborhood plane fitting method.
[0025] The preferred method for grading the included angle is as follows: when the absolute value of the included angle is no greater than 30 degrees, the corresponding observation is recorded as a high-confidence observation; when the absolute value of the included angle is greater than 30 degrees but less than 45 degrees, the corresponding observation is recorded as a downweighted observation; when the absolute value of the included angle is no less than 45 degrees, the corresponding observation is recorded as an invalid observation. This threshold setting has a clear physical basis. Under small incident angle conditions, the refraction shift and specular highlights at the water interface and acrylic observation window interface are still within a controllable range, and a high correspondence is maintained between spatial points and image boundaries. When the included angle exceeds 30 degrees, the refraction path shift and specular reflection enhancement begin to become significant. This affects the stability of observations. When the included angle reaches 45 degrees or more, the combined effect of interface reflection and refraction is significantly enhanced. Retaining this observation will significantly increase the probability of pseudo-geometry entering the subsequent processing chain. For acrylic observation window areas with large curvature, the first threshold is preferably 25 to 30 degrees, and the second threshold is preferably 40 to 45 degrees. For flat observation windows and the outer area of open pool bodies, the first threshold is preferably 30 to 35 degrees, and the second threshold is preferably 45 to 50 degrees. This range comes from the differences in curvature of the observation interface, the range of changes in viewing angle, and the differences in indoor lighting conditions, which are consistent with the actual acquisition environment.
[0026] After the observation reliability decays, local structural fragment construction is performed. This process is also a core method in S1 because the subsequent steps do not directly use a single point or a single pixel as the processing object, but use local structural fragments as the basic carrier. If the point-level results are directly used in the subsequent stability analysis, water surface reflection, spray boundary and local highlights will form a large number of discrete noise points, making it difficult to express the regional continuity of the real structure. If the entire region is processed directly, it will cover up geometric details such as the observation window frame, pool wall corner and rock protrusion. Therefore, this implementation method performs local aggregation on the scene point set based on the spatial adjacency relationship and normal continuity relationship of the mapped point cloud to form local structural fragments. This processing can be implemented by supervoxel segmentation or local aggregation based on neighborhood growth.
[0027] In practice, a basic neighborhood is first established based on the average point spacing of the scene. Then, spatial distance and normal deviation are used as the basis for local aggregation. The average point spacing of point clouds commonly found in marine park scenes is preferably 0.01 meters to 0.06 meters, and the corresponding local aggregation radius is preferably 0.05 meters to 0.30 meters. The normal continuity threshold is preferably 10 degrees to 20 degrees. When the aggregation radius is less than 0.05 meters, a stable facility surface will be divided into too many fragmented segments, and the subsequent observation support information will be too discrete. When the aggregation radius is greater than 0.30 meters, the pool wall edge, the observation window frame, and rock protrusions will be excessively merged, and the ability to distinguish the subsequent cross-view structural stability results will decrease. When the normal continuity threshold is less than 10 degrees, wet surfaces and curved facilities are easily segmented too finely. When the normal continuity threshold is greater than 20 degrees, different surfaces with obvious geometric transitions are easily mistakenly merged into the same local structural segment. After aggregation, a set of local structural segments is formed, and a unique segment identifier is assigned to each local structural segment so as to establish a stable correspondence between different observations.
[0028] After the local structural fragments are constructed, observation support information is generated for each local structural fragment. The observation support information includes at least the effective observation range and the time support range. The effective observation range refers to the observation range in which the projection coverage of the local structural fragment meets the effective coverage condition, the occlusion meets the allowable occlusion condition, and it is not judged as an invalid observation under multi-view image or laser point cloud observation.
[0029] Preferably, the effective projected area of the local structural segment in the image domain accounts for no less than 60% of the theoretical projected area, the occlusion ratio is no more than 40%, and the corresponding observation is not determined as an invalid observation by the aforementioned observation confidence decay rule. The time support range refers to the set of time intervals that can continuously provide effective observations of the local structural segment on the time axis. The reason for writing the effective observation range and the time support range into the observation support information is that the subsequent S2 needs to compare the cross-view structural stability of the same local structural segment based on multiple effective observations, and the subsequent S3 also needs to compare the degree of temporal overlap and the degree of common observation support between different local structural segments. Therefore, simply retaining the frame number or visibility marker is not enough to support the subsequent processing chain, and the segment-level observation support foundation must be established synchronously in the S1 stage.
[0030] After organizing the observation support information, standardized observation descriptions are formed for each local structural segment. First, depth, curvature, edge gradient, and brightness information are extracted from the effective observations corresponding to each local structural segment. Depth information is preferably represented by the median or average distance from each spatial point in the local structural segment to the observation center, in meters. Curvature information is preferably represented by the absolute value of curvature obtained by the local quadratic surface fitting method or the covariance eigenvalue analysis method, in meters. Edge gradient information is preferably represented by the average gradient magnitude obtained by back-projecting the local structural segment to the image domain and using the edge detection method. The initial unit can be regarded as a combination of grayscale change and pixel spacing. Brightness information is preferably represented by the average brightness after camera response correction and exposure equalization. After correction, the brightness is preferably first mapped to the normalized brightness range of 0 to 1. Since the above four types of quantities correspond to meters, meters, edge response magnitude, and normalized brightness, respectively, they cannot directly participate in the same comparison process. Therefore, the benchmark conversion of the same type of quantity must be completed in stage S1 before proceeding to the subsequent steps.
[0031] Specifically, the standardization of depth information can be achieved by using the median depth of the same local structural segment under all valid observations as the depth benchmark, and then the depth values under each valid observation are compared with this depth benchmark to obtain the dimensionless standardized depth result. The reason for using the median depth as the depth benchmark is that the median depth is less sensitive to local occlusion and a few abnormal depth points than the average value. The standardization of curvature information should not directly use the median curvature as the denominator for ratio processing, because the flat pool wall and the surface of the observation window may be close to zero curvature in a local range, and direct ratio processing will amplify small noise.
[0032] Therefore, curvature information is preferably calculated using the characteristic length of the local structural segment. The absolute value of curvature is multiplied by the equivalent scale of the local structural segment to obtain a dimensionless standardized curvature result. The equivalent scale of the local structural segment is preferably the diagonal length or equivalent radius of the segment's bounding box, preferably ranging from 0.05 meters to 0.50 meters. The standardization of edge gradient information preferably uses the gradient statistics of the same local structural segment under effective observation as a reference benchmark, and converts it into a dimensionless standardized edge gradient result through quantile normalization or benchmark amplitude normalization. The reference benchmark preferably uses the 95th percentile gradient amplitude to reduce the influence of individual highlight peaks on the overall gradient benchmark. The standardization of brightness information preferably uses the normalized brightness median value of the same local structural segment under effective observation as the brightness benchmark, and then the normalized brightness under each effective observation is offset or scaled relative to this brightness benchmark to obtain a dimensionless standardized brightness result.
[0033] To avoid abnormal amplification of sensor noise in extremely dark areas, the lower limit of the brightness reference is preferably set to 0.03 to 0.08; to avoid abnormal amplification of random disturbances in extremely weak edge areas, the lower limit of the edge gradient reference is preferably set to 0.02 to 0.05 of the corresponding normalized gradient scale. After the above conversion is completed, standardized depth results, standardized curvature results, standardized edge gradient results and standardized brightness results are formed, which together constitute a standardized observation description.
[0034] Through the above processing, S1 ultimately outputs a set of local structural fragments, observation support information corresponding to each local structural fragment, and standardized observation descriptions. Among them, the set of local structural fragments provides a unified fragment-level processing object for the calculation of cross-view structural stability results in S2; the observation support information provides a direct basis for the extraction of temporal overlap and common observation support in S3; and the standardized observation descriptions provide a unified input for the analysis of depth deviation, curvature deviation, edge gradient deviation, and normal deviation in S2. The technical effect of this step is that, from the input end, the unstable observations in the marine park scene caused by the water interface, acrylic observation window, wet surface, and time-division acquisition are subject to credibility control, spatial organization, and scale unification. This ensures that the subsequent identification of pseudo-geometric candidate fragments and the construction of scene state co-occurrence relationship graphs are based on data that has completed coordinate unification, credibility filtering, fragment carrier construction, and standardized conversion of similar quantities. This can significantly reduce the risk of floating structures, misaligned splicing, and pseudo-stable states being mixed in in the subsequent global mesh reconstruction results.
[0035] In this embodiment, S2 calculates the cross-view structural stability results of each local structural segment between different effective observations based on the standardized observation description, and generates a set of pseudo-geometric candidate segments and pseudo-geometric risk results by combining interface perturbation characterization and brightness fluctuation characterization. S2 takes the set of local structural segments, observation support information and standardized observation description output by S1. The task is to first identify whether the geometry is stable based on the segment-level standardized observation results, and then identify whether pseudo-geometry is formed due to water light, reflection or scattering perturbation, so as to provide clean and comparable input for the construction of the scene state co-occurrence relationship graph in S3.
[0036] First, for each local structural segment, a set of valid observation pairs is constructed within its valid observation set. Here, the valid observations have already undergone unified coordinate mapping, observation confidence decay, and invalid observation removal in S1. Therefore, each observation entering S2 has a comparable spatial basis and confidence basis. If a local structural segment retains only one valid observation in the output of S1, then the local structural segment does not meet the conditions for cross-view comparison. In this implementation, it is directly marked as a low-stability candidate segment and written into the segment set to be jointly screened and verified in the subsequent process, instead of entering the main calculation process of the cross-view structural stability result. The reason for this is that the cross-view structural stability result essentially depends on the stable performance of the same local structural segment under multiple views, and single-view observations cannot support the reliable generation of this result.
[0037] Under the condition of having at least two valid observations, for the first A local structural fragment Construct a set of effective observation pairs within its effective observation set. Then, for each valid observation pair, the consistency of the normalized depth, normalized curvature, normalized edge gradient, and unit normal vector of the local structural segment under the two valid observations is compared, and combined with the corresponding observation confidence decay result to form a cross-view structural difference quantity. The calculation relationship used is as follows: , ; in, Indicates the first Cross-view structural differences of a local structural segment Indicates the first Cross-view structural stability results for individual local structural segments Indicates the first A set of effective observation pairs for a local structural segment They represent the first The local structural fragment in the first The first valid observation and the first The observation reliability decay coefficient under a valid observation, They represent the first The normalized depth results of a local structural segment under two valid observations. These represent the results of the standardized curvature. These represent the results of the standardized marginal gradients. They represent the unit normal vectors, Let represent the non-negative weights of the depth difference term, curvature difference term, edge gradient difference term, and normal difference term, respectively, and satisfy the condition that the sum of the weights is 1. This represents a stable term to prevent the denominator from being zero. All values are dimensionless results obtained after standardization of S1. The unit normal vector itself is also dimensionless. Therefore, all differences in the terms entering this equation are dimensionless. The normal difference term adopts... The form is because the dot product of two unit normal vectors ranges from -1 to 1. After this conversion, it can be stably mapped to the interval of 0 to 1, which makes it easier to integrate with other dimensionless difference terms under the same weighting framework.
[0038] The above formula is formed logically as follows: If the same local structural segment corresponds to a real stable structure, then in a unified scene coordinate system, the standardized depth result, standardized curvature result, standardized edge gradient result, and unit normal vector under different effective observations should maintain high consistency. Therefore, each difference term is small, resulting in a small cross-view structural difference. If the same local structural segment is affected by reflection, refraction, scattering, or local occlusion, then significant depth drift, curvature distortion, edge response drift, or normal fluctuation will occur under different effective observations, increasing each difference term and resulting in a larger cross-view structural difference. Since large incident angle observations have already undergone graded control through observation reliability decay in S1, this formula further utilizes... Weighted suppression of the contribution of observations can further reduce the bias of low-confidence observations on the results. Finally, the exponential mapping is used to convert the difference into a structural stability result. This is to monotonically map the cross-view structural difference to the interval between 0 and 1, so that it can be entered into S3 and S4 together with the pseudo-geometric risk result. The mapping result here satisfies the following: the smaller the cross-view structural difference, the closer the structural stability result is to 1; the larger the cross-view structural difference, the closer the structural stability result is to 0.
[0039] Regarding the weighting in this formula, this implementation prefers to assign higher weights to the depth difference term and the normal difference term, and to assign auxiliary weights to the curvature difference term and the edge gradient difference term. This is because the pseudo-geometry in the ocean park scene is primarily manifested as instability in spatial position and surface orientation, and only secondarily as changes in surface curvature and image edge response. Based on this physical law... The preferred value is 0.35 to 0.45. The preferred value is 0.20 to 0.30. The preferred value is 0.15 to 0.25. The preferred value is 0.10 to 0.20, and the preferred combination is [missing value]. The rationale for this combination is as follows: depth stability provides the most direct assessment of the true structure, hence its highest value; normal stability reflects the consistency of local surface orientation, hence its second highest value; curvature is used to compensate for changes in surface structure; edge gradient is easily affected by local brightness conditions, and although it has discernible value, its weight is lower than the depth and normal terms to avoid amplifying image noise in blue light and highlight areas, and the stability term... Preferred selection To adapt to different orders of magnitude of effective observation pairs without changing the dominant trend of structural difference, when the total weight of effective observation pairs is close to 0, it indicates that although the local structural segment has passed through S1, it still lacks sufficient high-confidence observation support. At this time, the cross-view structural difference is no longer used as the main basis for stability determination, but the local structural segment is marked as a low-stability candidate segment and handed over to the subsequent joint segment screening and comprehensive processing.
[0040] After obtaining the cross-view structural stability results, secondary core processing is performed around the pseudo-geometric judgment logic. Specifically, the water coverage area, specular reflection coverage area, and fog scattering coverage area corresponding to each local structural segment are first identified within the image projection domain. This identification can be achieved using a joint segmentation method based on brightness gradient and saturation, or by using reflection area detection and fog area detection methods. This part belongs to the supporting existing technology processing, and the algorithm details are not elaborated here. After the identification is completed, the area ratio of the above three types of coverage areas within the projection range of each local structural segment is calculated to form an interface perturbation characterization. Since the area ratio itself is a dimensionless ratio between 0 and 1, the interface perturbation characterization is a dimensionless quantity and can be directly used in subsequent risk judgment. Furthermore, for the same local The standardized brightness results of the structural segment under different effective observations are used to obtain the reference brightness of the local structural segment by weighted averaging. The weight is preferably the observation confidence attenuation coefficient of the corresponding observation. The standardized brightness result is used instead of the original brightness value because the original brightness is still affected by the exposure difference of the acquisition equipment and the absolute intensity of scene lighting. The standardized brightness result has been benchmarked by S1 and is a dimensionless quantity that can be compared across observations within the segment. With the reference brightness as a reference, the deviation of the standardized brightness result relative to the reference brightness under each effective observation is statistically analyzed to form a brightness fluctuation characterization. The brightness fluctuation characterization is also a dimensionless quantity. The larger the value, the more obvious the time domain or view domain optical drift of the local structural segment under different effective observations.
[0041] After the interface perturbation characterization and brightness fluctuation characterization are formed, a pseudo-geometric risk result is generated based on the cross-view structural stability result, interface perturbation characterization, and brightness fluctuation characterization. This can be achieved through a hierarchical fusion approach: first, the structural stability result is converted into a structural instability level; then, it is fused with the interface perturbation characterization and brightness fluctuation characterization according to preset weights; finally, the fused result is restricted to the range of 0 to 1 to obtain the pseudo-geometric risk result. The fusion logic here follows this rule: when the structural stability result is low, the interface perturbation characterization is high, and the brightness fluctuation characterization is high, the pseudo-geometric risk result increases; when the structural stability result is high and both the interface perturbation characterization and brightness fluctuation characterization are low, the pseudo-geometric risk result decreases. The preset weights prioritize the structural instability level, making the interface perturbation characterization the dominant factor. The fusion weights for structural instability and brightness fluctuation are assigned as auxiliary weights. Preferably, the fusion weights for structural instability can be 0.40 to 0.55, interface disturbance can be 0.20 to 0.35, and brightness fluctuation can be 0.20 to 0.30. A preferred combination is 0.50, 0.25, and 0.25. The basis for this setting is that the essence of pseudo-geometry is still geometric instability. However, in the marine park scene, interface disturbance and brightness fluctuation become typical causes. Therefore, all three need to be evaluated together, but geometric instability should be dominant. If the weights for interface disturbance and brightness fluctuation are too high, the fixed structure may be mistakenly elevated in the performance lighting switching area. If the weights for the two are too low, the pseudo-structure caused by water surface reflection and specular reflection may escape the risk identification range.
[0042] After the pseudo-geometric risk results are generated, local structural fragments that meet the preset conditions are written into the pseudo-geometric candidate fragment set according to a preset risk threshold. The preset risk threshold is preferably between 0.55 and 0.70. For areas with dense blue light, spray, and curved acrylic observation windows, the disturbances caused by reflection and scattering are relatively strong. To avoid a large number of boundary fragments being incorrectly written into the pseudo-geometric candidate fragment set, the preset risk threshold is preferably between 0.60 and 0.70. For areas with straight pool walls and dense fixed facilities, the preset risk threshold is preferably between 0.55 and 0.65 to improve the sensitivity to local pseudo-geometry. This range is derived from the balance between the false positive rate and the false negative rate in the labeled area samples: when the risk threshold is below 0.55, real fixed structures are mistakenly written into the pseudo-geometric candidate fragment set. The proportion of matching increases significantly; when the risk threshold is higher than 0.70, pseudo-structures induced by high reflectivity of the water surface and high reflectivity of wet surfaces are easily missed. It should be emphasized that this implementation does not directly remove pseudo-geometric candidate segments in S2, but only writes them into the pseudo-geometric candidate segment set, and outputs the pseudo-geometric risk results to the subsequent S3 and S4. The reason for this is that there is a type of edge segment in the marine park scene. Its local optical perturbation is strong, but it is still adjacent to the real fixed structure. If it is directly removed in the S2 stage, it is easy to cause gaps in the real structure when closing the state version and reconstructing the global mesh. Retaining it as a candidate and processing it in a unified manner in combination with the scene state co-occurrence relationship graph and joint segment screening can reduce the risk of accidentally deleting real structures while maintaining the pseudo-geometric recognition capability.
[0043] Through the above processing, S2 finally outputs cross-view structural stability results, a set of pseudo-geometric candidate segments, and pseudo-geometric risk results. The cross-view structural stability results characterize whether each local structural segment maintains consistency in geometric and boundary responses under different effective observations. The pseudo-geometric risk results characterize whether each local structural segment is significantly affected by water reflection, specular reflection, wet surface highlights, or spray scattering. The set of pseudo-geometric candidate segments serves as a key constraint for subsequent scene state co-occurrence relationship graph construction and joint segment selection. Compared with the common processing methods of direct denoising or direct removal of outliers in existing technologies, this implementation first completes the calculation of cross-view structural stability results at the segment level, then introduces interface perturbation characterization and brightness fluctuation characterization to generate pseudo-geometric risk results, and retains them as input for subsequent state version closure. This connects the pseudo-geometric identification in the first-level problem with the pseudo-stable state determination in the second-level problem, which can improve the accuracy of subsequent joint segment selection and the authenticity of global mesh reconstruction results without prematurely destroying the continuity of the real structure.
[0044] In this embodiment, S3 constructs a scene state co-occurrence relationship graph based on the structural stability results and the pseudo-geometric candidate fragment set. It extracts the spatial relationship consistency, temporal overlap, common observation support, and spatial occupancy conflict between local structural fragments, and generates state co-occurrence weights and mutual exclusion constraints. S3 takes the structural stability results, pseudo-geometric candidate fragment set, and pseudo-geometric risk results output by S2. The processing goal is to determine whether different local structural fragments can coexist in the same scene version, thereby solving the problem of pseudo-stable state mixing caused by feeding platform displacement, maintenance fence deployment, performance prop lifting and lowering, and occlusion by swarming creatures under time-sharing acquisition conditions. The core processing of this step is to construct a scene state co-occurrence relationship graph based on the spatial compatibility, temporal compatibility, and common observation support between local structural fragments, and then generate mutual exclusion constraints using spatial occupancy conflict and temporal incompatibility, providing version closure basis for subsequent joint fragment selection. Centroid extraction, 3D bounding volume establishment, and temporal interval overlap statistics can be implemented using point cloud statistical processing and temporal interval operation methods.
[0045] In practice, firstly, a basic node set is established using each local structural fragment as a graph node. Then, in a unified scene coordinate system, a candidate edge set is established for spatially adjacent local structural fragments. Spatially adjacent is preferably defined as the minimum bounding volume distance between two local structural fragments being no greater than 0.30 meters to 0.80 meters, or the two fragments having a direct adjacency relationship in at least one common valid observation of their projected boundaries. The reason for adopting this adjacency condition is that the version closure problem mainly occurs in the combination of fragments in the same region or neighboring regions. If an edge candidate relationship is established for every pair of fragments globally, This will introduce a large number of invalid associations and increase the probability of erroneous connectivity. For small-scale local structural segments such as the observation window frame, nozzle outer edge, and rock protrusion, the adjacency distance is preferably 0.30 meters to 0.50 meters. For larger-scale local structural segments such as pool walls, walkway boundaries, and large artificial reefs, the adjacency distance is preferably 0.50 meters to 0.80 meters. This range is derived from the back-calculation results of the typical size of marine park fixed facilities, point cloud sampling density, and local structural segment segmentation granularity. It can cover real adjacent segments while suppressing distant irrelevant segments from entering the same version of the decision chain.
[0046] After establishing the candidate edge set, for any two spatially adjacent local structural segments, first extract their common observation support information and statistically analyze their relative position changes under common effective observations. Specifically, the effective observation set where the two local structural segments coexist can be determined based on the observation support information output by S1. Then, the distance between their centroids is calculated within the common effective observation set. The centroid position can be obtained using the arithmetic mean coordinates of the spatial points contained in the local structural segment. If there are still a few discrete outliers in the local structural segment, it is preferable to preprocess the spatial points using a statistical outlier removal method before calculating the centroid position. The dimension of the distance between centroids is meters. Subsequently, the distance sequence between centroids of the same segment obtained under all common effective observations is statistically analyzed, and the median distance is used as the reference distance. The spatial relationship is then formed by the ratio of the distance fluctuation to the reference distance. Consistency, where both the distance fluctuation and the reference distance are in meters, is converted into dimensionless quantities after ratioization. Spatial relationship consistency can therefore participate in the unified expression of subsequent state co-occurrence weights. The higher the spatial relationship consistency, the more stable the relative spatial position of the two local structural segments is under multiple observations; the lower the spatial relationship consistency, the greater the relative position fluctuation of the two under multiple observations, and the more likely they correspond to different segment combinations in the process of movable state change. To avoid the normalization result being abnormally amplified under the condition that the reference distance is too small, when the reference distance is less than twice the average point spacing of the local structural segments, it is preferable to limit the minimum reference scale to twice the average point spacing of the local structural segments. After adopting this limitation, the locally adjacent segments of the observation window frame, the pool wall splicing segments, and the rock mass edge segments can still be stably compared under short distance conditions, and the position fluctuation will not be exaggerated due to the small denominator.
[0047] After obtaining the spatial relationship consistency, the temporal overlap degree is extracted around the temporal support ranges corresponding to the two local structural segments. Specifically, the temporal support ranges output by S1 are directly read, and the overlap duration and union duration of the two temporal support ranges on the time axis are calculated. The overlap duration is then divided by the union duration to obtain the temporal overlap degree. The initial units of the overlap duration and the union duration are seconds, which are converted into dimensionless quantities after ratioization. The higher the temporal overlap degree, the more sufficient the evidence of temporal coexistence between the two local structural segments during the actual acquisition process; the lower the temporal overlap degree, the more likely they are to appear in different states at different acquisition stages. For fixed facility segments such as observation window frames, pool wall finishes, and walkway railings, adjacent segments typically have a high degree of temporal overlap. For segments corresponding to feeding platforms, maintenance enclosures, and performance props, the degree of temporal overlap is typically low. To improve the robustness of temporal relationship determination, when the total duration of the union of two local structural segments is less than 0.50 seconds, it is preferable to mark the segment pair as a low-temporal-sequence support segment pair and retain it only as a weak edge candidate, rather than directly using it for mutual exclusion constraint relationship determination. The basis for this setting is that when the time window is too short, occasional occlusion and short-term lighting changes will significantly amplify the temporal differences between segment pairs, which is insufficient to support reliable version compatibility determination.
[0048] Subsequently, the degree of common observation support is extracted based on the number of common effective observations and the total number of effective observations of the two local structural segments. In practice, the ratio of the number of common effective observations to the union of the total number of effective observations of the two segments can be converted to obtain the dimensionless degree of common observation support in the interval of 0 to 1. The higher the degree of common observation support, the more sufficient the common imaging basis of the two local structural segments. The lower the degree of common observation support, the more the co-occurrence relationship is supported by a small number of observations, and it is not advisable to assign a high connectivity confidence score. Here, the ratio of the number of observations is used instead of the absolute value of the number of observations because the absolute number of observations is not directly comparable due to differences in the degree of occlusion, the degree of openness of the viewing angle, and the acquisition path of different local structural segments. After the ratio processing, the degree of common observation support can be used in conjunction with the consistency of spatial relationship and the degree of temporal overlap under the same dimensionless framework.
[0049] After extracting spatial relationship consistency, temporal overlap, and common observation support, the spatial occupancy conflict degree is extracted. Specifically, for two local structural segments, their respective three-dimensional bounding volumes are first established based on the spatial point distribution in a unified scene coordinate system. Then, the overlap volume of the two three-dimensional bounding volumes is calculated. The initial dimension of the overlap volume and the volume of each bounding volume is cubic meters. To ensure that the spatial occupancy conflict degree can participate in the generation of state co-occurrence weights along with the aforementioned three types of dimensionless quantities, this implementation uses the ratio of the overlap volume to the smaller bounding volume of the two segments as the spatial occupancy conflict degree. After ratioization, the spatial occupancy conflict degree is a dimensionless quantity used to characterize the two segments. Whether a local structural fragment experiences a state conflict that cannot be simultaneously established at the same spatial location in a unified scene coordinate system, the higher the degree of spatial occupancy conflict, the more likely the two correspond to different versions of the state; the lower the degree of spatial occupancy conflict, the more compatible the spatial occupancy relationship between the two. The reason for using a smaller bounding volume as a normalization reference is that if the union volume is used as the normalization benchmark, the intrusion conflict of small local fragments on large fixed structures will be significantly diluted; after using a smaller bounding volume as a reference, the occupancy conflict of local fragments can be more sensitively displayed, which is suitable for the mutual exclusion determination of local movable landscape fragments such as maintenance fences, performance props, and movable feeding platforms.
[0050] Based on the above, the calculation relationship for the co-occurrence weight of the construction state for each local structural segment is as follows: ; in, Representing local structural fragment pairs State co-occurrence weights This indicates the consistency of the spatial relationships between the local structural fragments. This indicates the degree of temporal overlap between the local structural segments. This indicates the degree of common observational support for the pair of local structural segments. This indicates the degree of spatial conflict between the local structural segments. These represent the non-negative weights of the spatial relationship consistency term, the temporal overlap term, the common observation support term, and the spatial occupancy conflict term, respectively, and the sum of the four is 1. Since spatial relationship consistency is derived from the distance ratio, temporal overlap is derived from the time length ratio, common observation support is derived from the observation quantity ratio, and spatial occupancy conflict is derived from the volume ratio, the four types of quantities entering the above expression are all dimensionless quantities and can be directly entered into the same weight expression.
[0051] The physical meaning of state co-occurrence weight is as follows: when two local structural segments maintain a stable relative positional relationship under multiple common effective observations, have high evidence of temporal coexistence and a common imaging basis, and there is no obvious spatial occupancy conflict in a unified scene coordinate system, their state co-occurrence weight increases; when two local structural segments have obvious spatial occupancy conflicts, or only alternate within a very short time window, their state co-occurrence weight decreases. In combination with the actual application of marine park scenes, fixed landscape structural segments usually have high spatial relationship consistency, high temporal overlap, and high common observation support, while movable state segments caused by feeding platform relocation, maintenance fence deployment, performance prop lifting and lowering, and occupancy obstruction by swimming organisms usually exhibit high spatial occupancy conflict and low temporal overlap. Based on this rule, this implementation method preferably assigns higher weights to the spatial relationship consistency item and the spatial occupancy conflict item, medium weights to the temporal overlap item, and auxiliary weights to the common observation support item.
[0052] Specifically, The preferred value is 0.25 to 0.35. The preferred value is 0.20 to 0.30. The preferred value is 0.15 to 0.25. The preferred value is 0.25 to 0.35, and the preferred combination is... With this setting, spatial relationship consistency directly supports the stable connection of fixed structures, spatial occupancy conflict degree effectively suppresses movable state fragments from entering the same version at the same time, temporal overlap degree provides a basis for temporal compatibility, and common observation support degree provides sample support supplement. If the weight of common observation support degree is too high, occasional common observations in short-term occlusion areas will be over-amplified; if the weight of spatial occupancy conflict degree is too low, different state fragments may be incorrectly retained in the same scene version.
[0053] After generating the state co-occurrence weights, mutual exclusion constraints are generated. Specifically, for any two local structural segments, if the spatial occupancy conflict degree of the two segments meets the preset conflict condition and the temporal overlap degree meets the preset temporal incompatibility condition, then the two segments are identified as a pair of mutually exclusive segments and written into the mutual exclusion constraint relationship. The preset conflict condition is preferably a spatial occupancy conflict degree of not less than 0.60 to 0.80, and the preset temporal incompatibility condition is preferably a temporal overlap degree of not more than 0.20 to 0.35. For the areas where performance props, maintenance fences, and feeding platforms are located, due to the large range of movable state changes, the preset conflict condition is preferably 0.60 to 0.70, and the preset temporal incompatibility condition is preferably 0.25 to 0.35. For the pool wall boundary, the observation window frame, and the adjacent area of the artificial reef landscape... To prevent real fixed structures from being misjudged as mutually exclusive fragment pairs, the preset conflict condition is preferably set to 0.70 to 0.80, and the preset time incompatibility condition is preferably set to 0.20 to 0.30. This range is derived from the statistical results of the occupancy and temporal relationships of real fixed fragment pairs and movable state fragment pairs in the museum area samples. When the space occupancy conflict degree is lower than 0.60, many real adjacent fragments that only partially overlap may also be mistakenly written into the mutually exclusive constraint relationship; when the space occupancy conflict degree is higher than 0.80, some real movable state fragment pairs will be missed due to insufficient occupancy overlap. If the upper limit of the time overlap degree is higher than 0.35, real fragment pairs with strong synchronous support will be incorrectly rejected; if it is lower than 0.20, some mutually exclusive fragment pairs formed by time-sharing acquisition will be missed into the same scene version.
[0054] Meanwhile, to reduce the contamination effect of the pseudo-geometric candidate fragment set on the scene state co-occurrence relationship graph, for local structural fragments belonging to the pseudo-geometric candidate fragment set, the corresponding association strength is attenuated when they participate in the generation of state co-occurrence weights. Specifically, while maintaining the original physical meaning of spatial relationship consistency, temporal overlap, common observation support, and spatial occupancy conflict, the state co-occurrence weights of the local structural fragments are reduced by a preset attenuation ratio, preferably 20% to 50%. In the high-contrast blue light venue and spray performance area, since the pseudo-geometric candidate fragments have a higher risk of pulling erroneous state chains, the attenuation ratio is preferably 35% to 50%. In areas with dense fixed facilities and low reflection areas, the attenuation ratio is preferably 20% to 35%. The reason for adopting attenuation processing is that pseudo-geometric candidate fragments often have local adjacency relationships with the boundaries of real fixed structures. Directly breaking all their associations will destroy the continuity of local structures, while reducing the association strength proportionally can reduce the formation of erroneous state chains while preserving the topological continuity of the structure.
[0055] Through the above processing, S3 finally outputs the state co-occurrence weight and mutual exclusion constraint relationship. The state co-occurrence weight is used to characterize the version compatibility between different local structural fragments based on space, time and common observation. The mutual exclusion constraint relationship is used to prevent local structural fragments with obvious spatial occupancy conflicts and temporal incompatibility from entering the same scene version. This output result will be directly used as the version closure basis for joint fragment screening in S4. Compared with the processing method of splicing based only on spatial adjacency relationship or screening based only on temporal order relationship, this implementation method introduces spatial relationship consistency, temporal overlap degree, common observation support degree and spatial occupancy conflict degree at the same time. Furthermore, it combines the pseudo-geometric candidate fragment set to perform attenuation processing on the state co-occurrence weight, which can continue to pass the pseudo-geometric recognition result in S2 to the version closure link. This allows the scene state co-occurrence relationship graph to retain the stable connection of real fixed structure and prevent movable state fragments and pseudo-geometric fragments from forming erroneous state chains, providing a more reliable version closure basis for subsequent joint fragment screening.
[0056] In this implementation, S4 performs joint fragment screening based on structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutual exclusion constraints to obtain a set of stable scene fragments. Based on the set of stable scene fragments, global mesh reconstruction is performed, and the global mesh reconstruction result is output. S4 takes the structural stability results, pseudo-geometric risk results, and pseudo-geometric candidate fragment set output by S2, as well as the state co-occurrence weights and mutual exclusion constraints output by S3. The processing goal is to simultaneously write structural authenticity constraints and version consistency constraints within the same screening framework, determine the local structural fragments that can jointly enter the final scene version, and form a set of stable scene fragments and a global mesh reconstruction result based on this.
[0057] In specific implementation, a priori results for segment retention are first formed for each local structural segment. These priori results are derived from the joint determination of structural stability results and pseudo-geometric risk results. Both structural stability results and pseudo-geometric risk results are formed in S2, and both are dimensionless quantities located in the 0 to 1 interval. Therefore, they can be directly combined and used in the same priori formation relation. The formation logic of the segment retention priori results is as follows: the higher the structural stability result, the higher the segment retention priori result; the higher the pseudo-geometric risk result, the lower the segment retention priori result. Specifically, the structural stability result can be used as a positive contribution term first, and the pseudo-geometric risk result multiplied by the pseudo-geometric risk penalty coefficient can be used as a negative suppression term. The difference between the two is then processed. The results of the subsequent difference processing are truncated to non-negative values, and upper bound constraints are applied to results exceeding 1. This results in segments within the 0 to 1 range that retain prior results. The reason for this processing logic is that even if a local structural segment has a high structural stability result, if it also has a high pseudo-geometric risk result, its corresponding region may still be dominated by water reflection, specular reflection, wet surface highlights, or fog scattering. Directly entering the final scene version would significantly increase the probability of pseudo-structure residue. Even if a local structural segment is written into the pseudo-geometric candidate segment set, as long as the structural stability result is at a high level and the pseudo-geometric risk result is at a medium to low level, it still retains the opportunity to enter the subsequent joint segment selection, thereby reducing the premature rejection of real boundary segments.
[0058] The pseudo-geometric risk penalty coefficient needs to be set by considering both the risk of accidental deletion of real boundary segments and the risk of residual pseudo-geometric segments in the marine park scene. The optimal pseudo-geometric risk penalty coefficient is between 0.45 and 0.65. This range is derived from the inverse balance result of the accidental deletion rate of real fixed structures and the residual rate of pseudo-geometric segments in the labeled aquarium area samples. When the pseudo-geometric risk penalty coefficient is below 0.45, high-risk local structural segments corresponding to reflective and scattering boundaries may still be retained in subsequent screening due to their higher structural stability, leading to a significant increase in the proportion of residual pseudo-structures. When the value is higher than 0.65, the true local structural fragments near the observation window frame, pool boundary, and rock corner will be significantly suppressed due to local optical disturbances, resulting in gaps in the true boundary in subsequent reconstruction. For dense areas of curved acrylic observation windows and spray performance areas, the pseudo-geometric risk penalty coefficient is preferably taken as 0.55 to 0.65; for areas with straight pool walls and dense fixed facilities, the pseudo-geometric risk penalty coefficient is preferably taken as 0.45 to 0.55. After adopting this zoning value method, a relatively consistent true structure retention rate can be maintained in different venue areas.
[0059] After forming the prior results of fragment retention, common retention constraints are executed based on state co-occurrence weights, and simultaneous retention restrictions are executed based on mutual exclusion constraints. The meaning of common retention constraints is: for local structural fragment pairs with higher state co-occurrence weights, if one local structural fragment is retained, the other local structural fragment should also be retained first, thereby maintaining the closure of local structures within the same scene version. The meaning of simultaneous retention restrictions is: for local structural fragment pairs with mutual exclusion constraints, they cannot be retained simultaneously in the same joint fragment selection, thereby preventing fragments from different scene versions or fragments with conflicting space occupancy from entering the final scene together. Here, the state co-occurrence weights come from S3 and are dimensionless; the mutual exclusion constraints come from the set of mutually exclusive fragment pairs generated by S3 and are discrete constraints. In S4, the two play opposite but complementary roles: encouraging coexistence of the same version and prohibiting conflicting coexistence.
[0060] Under the combined effect of common retention constraints and simultaneous retention restrictions, joint fragment filtering is performed on all local structural fragments. The target relationship for joint fragment filtering is as follows: ; Satisfy constraints: .
[0061] in, Indicates the first Preservation markers for local structural segments, This indicates that the corresponding local structural fragment is preserved. This indicates that the corresponding local structural fragment is not preserved; Indicates the first Each local structural segment retains the prior value corresponding to the prior result; Indicates the state co-occurrence gain coefficient; Indicates the first The local structural fragment and the first State co-occurrence weights among local structural segments; This represents the set of mutually exclusive fragment pairs corresponding to a mutual exclusion constraint relationship. This represents the total number of local structural fragments participating in the joint fragment selection. In this objective relation, the first part is used to prioritize the retention of local structural fragments with stable structures and low pseudo-geometric risks, and the second part is used to encourage the simultaneous retention of local structural fragments that can be established in the same time version. The mutual exclusion constraint relation is used to prevent local structural fragments with obvious spatial occupancy conflicts and insufficient temporal compatibility from entering the final scene at the same time. Since the prior values and state co-occurrence weights corresponding to the fragment retention prior results are both dimensionless, and the state co-occurrence gain coefficient is also set to a dimensionless coefficient, there is no dimension inconsistency problem in the entire objective relation.
[0062] The state co-occurrence gain coefficient needs to be set based on the average connectivity between local structural fragments and the dispersion of the prior results of fragment preservation. The optimal range for the state co-occurrence gain coefficient is 0.15 to 0.35. This range is derived from the inverse balance between the authenticity of a single fragment and the closure of multiple fragment versions in the labeled library samples. When the state co-occurrence gain coefficient is below 0.15, the effect of the co-preservation constraint is too weak, and the local structural fragments are mainly determined by the prior results of fragment preservation alone, leading to the easy occurrence of fragmented preservation of adjacent fragments in the same scene version. When the state co-occurrence gain coefficient is above 0.35, ... The connectivity effect of state co-occurrence weights will outweigh the constraint effect of preserving prior results in fragments. Some pseudo-geometric candidate fragments may be pulled back into the final version by means of local connectivity. For areas with dense fixed facilities and flat pool walls, the state co-occurrence gain coefficient is preferably 0.20 to 0.30; for areas with dense spray performances and observation windows, the state co-occurrence gain coefficient is preferably 0.15 to 0.25. With this partitioning setting, the joint fragment screening can maintain the overall coherence of the real scene version and will not form a large-scale erroneous state chain due to local erroneous connections.
[0063] After solving the joint fragment selection objective, the local structural fragments marked as 1 are retained and written into the stable scene fragment set. The solution process here can be implemented using integer programming methods from 0 to 1, such as branch and bound, graph cut, or heuristic iterative methods. Since the joint fragment selection objective has already written the fragment retention prior results, state co-occurrence weights, and mutual exclusion constraints into the same solution framework, no matter which existing solution method is used, as long as the above objective and constraint relationships are satisfied, feasible retention results can be obtained. This process is significantly different from the common sequential method of denoising, splicing, and reconstruction. The final scene version is directly determined through a single joint fragment selection. The structural authenticity constraint and version consistency constraint work together within the same objective framework. Therefore, the high-confidence local structural fragments obtained in the previous layer will enhance the accuracy of the version closure in the next layer, and the version closure results obtained in the next layer will in turn suppress any pseudo-geometric fragments or pseudo-stable state fragments that may remain in the previous layer from entering the final model.
[0064] After the set of stable scene fragments is formed, the reconstruction weight results are generated based on the fragment retention prior results corresponding to each stable scene fragment. The reconstruction weight results are used to characterize the contribution of each stable scene fragment in the final mesh reconstruction. Since the fragment retention prior results are in the range of 0 to 1 after the aforementioned non-negative truncation and upper limit constraint, non-negative truncation processing can be performed on the fragment retention prior results corresponding to all selected local structure fragments first, and then normalization processing can be performed on the truncated results to obtain the reconstruction weight results. After normalization, the reconstruction weight results corresponding to each stable scene fragment are dimensionless, and the sum of all reconstruction weight results is 1. The fragment retention prior results are used as the basis for the reconstruction weight results because the fragment retention prior results have comprehensively reflected the structural stability results and pseudo-geometric risk results, and have a good ability to distinguish between real fixed structures and high-risk pseudo structures. If the sum of the fragment retention prior results corresponding to all selected local structure fragments after non-negative truncation is lower than the preset lower limit, it is preferable to regenerate the reconstruction weight results according to the proportion of spatial points of each selected local structure fragment. The preset lower limit is preferably taken as... The reason for adopting this fallback strategy is that in extremely low confidence scenarios or in areas with strong local disturbances, the prior results of the fragment retention of all fragments may be close to 0 at the same time. If direct normalization is continued in this case, it will cause numerical instability. After forming the reconstruction weight result according to the proportion of spatial points, it can still be guaranteed that the stable scene fragment set can obtain implementable input weights in the subsequent global grid reconstruction.
[0065] Subsequently, spatial points and unit normals from the set of stable scene fragments are extracted as reconstruction inputs, and global mesh reconstruction is performed using a weighted mesh reconstruction method, outputting the global mesh reconstruction result. The weighted mesh reconstruction method preferably adopts the weighted Poisson mesh reconstruction method or the weighted masked Poisson mesh reconstruction method. The reason for choosing the Poisson reconstruction method is that it can stably handle the oriented point set with unit normals, and the contribution of different stable scene fragments can be directly adjusted using the reconstruction weight results. For stable scene fragments with high prior results and high structural stability, their reconstruction weight results are larger, and they contribute more to the final surface fitting. For boundary fragments that are retained but have low prior results, their reconstruction weight results are smaller, and they only play an auxiliary splicing role in the final surface fitting. In this way, the global mesh reconstruction result can reduce the pulling effect of boundary pseudo-structures and version conflict fragments on the final scene model while maintaining local geometric continuity.
[0066] Through the above processing, S4 finally outputs a stable scene fragment set and a global mesh reconstruction result. The stable scene fragment set represents the set of local structural fragments that can enter the same scene version under the combined effect of structural realism constraints and version consistency constraints. The global mesh reconstruction result represents the stable 3D scene model reconstructed based on the stable scene fragment set. The technical effect of this step is that the final scene model is generated through a joint fragment selection framework. The structural stability result, state co-occurrence weight, pseudo-geometric risk result, and mutual exclusion constraint relationship work together in the same solution process, ensuring both local geometric realism and overall version closure. Compared with the processing method of performing denoising, splicing, and reconstruction in sequence, the global mesh reconstruction result output by this implementation method can better maintain the continuity of real fixed structures and effectively prevent movable state fragments and pseudo-geometric fragments from mixing into the final 3D model, thereby improving the realism, integrity, and reusability of the virtual scene construction result of the Ocean Park.
[0067] Example 2 like Figure 2 As shown, the present invention also discloses a virtual scene construction system for marine parks based on three-dimensional models, including: The multi-source observation processing module acquires multi-source observation data of the marine park scene, performs unified coordinate mapping, observation confidence decay and local structure fragment construction, and forms observation support information and standardized observation description based on the depth, curvature, edge gradient and brightness information corresponding to each local structure fragment. The pseudo-geometric candidate generation module calculates the cross-view structural stability results of each local structural segment between different effective observations based on standardized observation descriptions, and generates a set of pseudo-geometric candidate segments and pseudo-geometric risk results by combining interface perturbation characterization and brightness fluctuation characterization. The state closure construction module constructs a scene state co-occurrence relationship graph based on structural stability results and pseudo-geometric candidate fragment set, extracts the spatial relationship consistency, temporal overlap, common observation support and spatial occupancy conflict between local structural fragments, and generates state co-occurrence weights and mutual exclusion constraints. The joint reconstruction output module performs joint fragment filtering based on structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutual exclusion constraints to obtain a set of stable scene fragments. It then performs global mesh reconstruction based on the set of stable scene fragments and outputs the global mesh reconstruction results.
[0068] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a virtual scene of a marine park based on a three-dimensional model, characterized in that, include: S1. Acquire multi-source observation data of the marine park scene, perform unified coordinate mapping, observation confidence decay and local structure fragment construction, and form observation support information and standardized observation description based on the depth, curvature, edge gradient and brightness information corresponding to each local structure fragment. S2. Based on standardized observation description, calculate the cross-view structural stability results of each local structural segment between different effective observations, and combine interface perturbation characterization and brightness fluctuation characterization to generate a set of pseudo-geometric candidate segments and pseudo-geometric risk results. S3. Based on the structural stability results and the set of pseudo-geometric candidate segments, construct a scene state co-occurrence relationship graph, extract the spatial relationship consistency, temporal overlap, common observation support and spatial occupancy conflict between local structural segments, and generate state co-occurrence weights and mutual exclusion constraints. S4. Based on the structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutual exclusion constraints, perform joint fragment screening to obtain a set of stable scene fragments, and perform global mesh reconstruction based on the set of stable scene fragments to output the global mesh reconstruction results. 2.The method of claim 1, wherein, Multi-source observation data includes multi-view images, laser point clouds, acquisition poses, and timestamps of the marine park scene at multiple acquisition times; Perform unified coordinate mapping, including: establishing the coordinate correspondence between each multi-view image and each laser point cloud relative to the unified scene coordinate system based on the acquired pose, and mapping the observation area in the multi-view image and the spatial points in the laser point cloud to the unified scene coordinate system. 3.The method of claim 2, wherein, The process involves performing observation confidence reduction and constructing local structural fragments, generating observation support information and standardized observation descriptions for each fragment. This includes: reducing the confidence level or removing invalid observations for observations passing through water interfaces or acrylic observation windows based on the angle between the observation ray and the local surface normal; performing local aggregation on scene point sets based on the spatial adjacency and normal continuity of the mapped point cloud to form local structural fragments; determining the corresponding effective observation range and time support range for each fragment to generate observation support information; and performing benchmarking conversions of the same quantities for the depth, curvature, edge gradient, and brightness of each fragment under the corresponding effective observations to generate standardized observation descriptions. 4.The method of claim 1, wherein, The cross-view structural stability results of each local structural segment across different valid observations are calculated based on standardized observation descriptions. This includes: extracting the standardized depth, standardized curvature, standardized edge gradient, and normal consistency descriptions for each local structural segment under different valid observations; forming a cross-view structural difference description of the local structural segment based on the degree of depth deviation, curvature deviation, edge gradient deviation, and normal deviation between different valid observations; and performing suppression processing on the cross-view structural difference description by combining the observation confidence decay results corresponding to each valid observation to generate cross-view structural stability results.
5. The method of claim 4, wherein, Combining interface perturbation characterization and brightness fluctuation characterization, a set of pseudo-geometric candidate segments and pseudo-geometric risk results are generated, including: for each local structural segment in the corresponding region of the image projection domain, water coverage area, specular reflection coverage area and fog scattering coverage area are identified, and interface perturbation characterization is formed based on the proportion of each coverage area in the corresponding projection range; for each local structural segment under different effective observations, brightness fluctuation characterization is formed based on its offset magnitude relative to the reference brightness. Based on cross-view structural stability results, interface perturbation characterization, and brightness fluctuation characterization, pseudo-geometric risk results are generated; local structural segments whose pseudo-geometric risk results meet preset conditions are written into the pseudo-geometric candidate segment set. 6.The method of claim 1, wherein, Based on the structural stability results and the pseudo-geometric candidate fragment set, a scene state co-occurrence relationship graph is constructed. The spatial relationship consistency, temporal overlap, common observation support, and spatial occupancy conflict between local structural fragments are extracted to generate state co-occurrence weights. This includes: using each local structural fragment as a graph node; for spatially adjacent local structural fragment pairs, spatial relationship consistency is extracted based on the stability of the relative positional relationship under multiple common observations, temporal overlap is extracted based on the overlap of the corresponding temporal support range, common observation support is extracted based on the proportion of common effective observations, and spatial occupancy conflict is extracted based on the spatial enclosure overlap in the unified scene coordinate system. By combining spatial relationship consistency, temporal overlap, common observation support, and spatial occupancy conflict, state co-occurrence weights are generated between local structural segments, and a scene state co-occurrence relationship graph is established accordingly.
7. The method of claim 6, wherein, Generate mutual exclusion constraint relationships, including: for any two local structural segments, if the degree of space occupation conflict between the two segments meets the preset conflict condition and the degree of time overlap meets the preset time incompatibility condition, then the two segments are determined as a pair of mutually exclusive segments and written into the mutual exclusion constraint relationship. For local structural fragments belonging to the pseudo-geometric candidate fragment set, when generating the co-occurrence weights of their participation in the state, the corresponding association strength is attenuated to reduce the impact of pseudo-geometric candidate fragments on the connectivity results of the scene state co-occurrence relationship graph. 8.The method of claim 1, wherein, Joint fragment selection is performed based on structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutually exclusive constraint relationships. This includes: for each local structural fragment, forming fragment retention prior results based on the corresponding structural stability results and pseudo-geometric risk results. Based on the co-occurrence weights of states among local structural segments, a common retention constraint is applied to local structural segments with high co-occurrence consistency; based on mutual exclusion constraints, a simultaneous retention constraint is applied to local structural segments that have spatial conflicts and temporal incompatibility; under the combined effect of segment retention prior results, common retention constraints, and simultaneous retention constraints, the retention state of each local structural segment is determined, and a stable set of scene segments is obtained.
9. The method of claim 8, wherein, Global mesh reconstruction is performed based on a set of stable scene fragments, and the output of the global mesh reconstruction result includes: extracting spatial points and unit normals from the set of stable scene fragments as reconstruction inputs; forming reconstruction weight results based on the prior results of each stable scene fragment; and using a weighted mesh reconstruction method to fuse and reconstruct the spatial points, unit normals and reconstruction weight results to obtain the global mesh reconstruction result.
10. A system for constructing a virtual scene of a marine park based on a three-dimensional model, which applies the method for constructing a virtual scene of a marine park based on a three-dimensional model according to any one of claims 1 to 9, characterized in that, include: The multi-source observation processing module acquires multi-source observation data of the marine park scene, performs unified coordinate mapping, observation confidence decay and local structure fragment construction, and forms observation support information and standardized observation description based on the depth, curvature, edge gradient and brightness information corresponding to each local structure fragment. The pseudo-geometric candidate generation module calculates the cross-view structural stability results of each local structural segment between different effective observations based on standardized observation descriptions, and generates a set of pseudo-geometric candidate segments and pseudo-geometric risk results by combining interface perturbation characterization and brightness fluctuation characterization. The state closure construction module constructs a scene state co-occurrence relationship graph based on structural stability results and pseudo-geometric candidate fragment set, extracts the spatial relationship consistency, temporal overlap, common observation support and spatial occupancy conflict between local structural fragments, and generates state co-occurrence weights and mutual exclusion constraints. The joint reconstruction output module performs joint fragment filtering based on structural stability results, state co-occurrence weights, pseudo-geometric risk results, and mutual exclusion constraints to obtain a set of stable scene fragments. It then performs global mesh reconstruction based on the set of stable scene fragments and outputs the global mesh reconstruction results.