AI-driven aerial hovering imaging methods, devices, equipment, and media
By constructing a three-dimensional spatial model and using feature fusion technology, combined with user interaction and environmental factors, the aerial suspension imaging parameters are dynamically adjusted, solving the problem that the imaging effect is limited to a preset template in the existing technology, and improving the imaging quality and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUIZHI WORLD (HANGZHOU) TECHNOLOGY CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-26
AI Technical Summary
Existing aerial levitation imaging technology is difficult to dynamically adjust according to users' real-time needs and environmental changes. The imaging effect is limited to preset scene templates and cannot achieve personalized optimization.
By acquiring multi-view depth and color image data, a three-dimensional spatial model is constructed, high-level structural and texture features are extracted, feature information is fused, and model reconstruction is performed by combining user interaction information and environmental factors. The imaging size and viewing angle parameters are adjusted, and a suitable levitation imaging device is selected for display.
It enables dynamic adjustments based on user needs and environmental changes, improving the quality and visual experience of aerial levitation imaging and ensuring that the imaging effect meets user expectations.
Smart Images

Figure CN121582482B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure relate to the field of aerial levitation imaging, and more specifically, to an AI-driven aerial levitation imaging method, apparatus, device, and medium. Background Technology
[0002] With the rapid development of artificial intelligence and optical display technology, human-computer interaction is evolving from traditional touch and remote control to natural multimodal interaction. Aerial levitation imaging, as a new type of medium-free display, can project images into the air without a physical screen, bringing a stronger sense of immersion and technology.
[0003] In related technologies, aerial levitation imaging often relies on fixed parameter settings or simple user input commands for imaging control. Its imaging effect is often limited to preset scene templates, making it difficult to dynamically adjust and personalize according to the user's real-time needs, environmental changes, and interactive intentions, thus affecting the imaging quality. Summary of the Invention
[0004] The embodiments described herein provide an AI-driven aerial levitation imaging method, apparatus, device, and medium that overcomes the aforementioned problems.
[0005] Firstly, according to the content of this disclosure, an AI-driven aerial levitation imaging method is provided, comprising:
[0006] Acquire multi-view depth image data and multi-view color image data of the target imaging scene;
[0007] Based on the depth values of each pixel in the multi-view depth image data, the position information of the object surface points in the target imaging scene in the three-dimensional coordinate system is determined, and three-dimensional point cloud data under different viewpoints is constructed using the position information; and a three-dimensional spatial model corresponding to the target imaging scene is generated based on the three-dimensional point cloud data under different viewpoints.
[0008] The multi-view color image data is subjected to layer-by-layer feature learning through a preset feature extraction network to extract high-level structural features and high-level texture features; and the high-level structural features and high-level texture features are associated and integrated through a cross-view feature matching and aggregation strategy to obtain structural feature information and texture feature information of the same key object under different viewpoints.
[0009] The structural and texture features of the same key object from different perspectives are fused into the three-dimensional spatial model corresponding to the target imaging scene to obtain a fused scene model that includes spatial location and visual details.
[0010] The fused scene model is reconstructed once based on user-defined scene planning information; and the fused scene model is reconstructed a second time based on environmental influencing factors within a preset imaging area to obtain an enhanced scene model.
[0011] Based on the enhanced scene model, the aerial imaging size and imaging angle parameters corresponding to aerial levitation imaging are determined according to user interaction information; the aerial imaging size is adjusted according to the user's head posture data relative to the preset imaging area and the spatial distribution relationship of key objects in the enhanced scene model; and the imaging angle parameters are adjusted according to the relative azimuth angle between the user's spatial position and the preset imaging area.
[0012] The levitation imaging device is selected based on the straight-line distance between the user and the preset imaging area; and the levitation imaging device is controlled to perform aerial levitation imaging display of the enhanced scene model within the preset imaging area according to the aerial imaging size and the imaging angle parameters.
[0013] Secondly, according to the present disclosure, an AI-driven aerial levitation imaging device is provided, comprising:
[0014] The acquisition module is used to acquire multi-view depth image data and multi-view color image data of the target imaging scene;
[0015] The first determining module is used to determine the position information of the object surface points in the target imaging scene in the three-dimensional coordinate system based on the depth value of each pixel in the multi-view depth image data, construct three-dimensional point cloud data under different viewpoints based on the position information, and generate a three-dimensional spatial model corresponding to the target imaging scene based on the three-dimensional point cloud data under different viewpoints.
[0016] The second determining module is used to perform layer-by-layer feature learning on the multi-view color image data through a preset feature extraction network to extract high-level structural features and high-level texture features; and to associate and integrate the high-level structural features and high-level texture features through a cross-view feature matching and aggregation strategy to obtain structural feature information and texture feature information of the same key object under different views.
[0017] The fusion module is used to fuse the structural feature information and texture feature information of the same key object from different perspectives in the three-dimensional spatial model corresponding to the target imaging scene, so as to obtain a fused scene model containing spatial position and visual details.
[0018] The reconstruction module is used to perform a first reconstruction of the fused scene model based on user-defined scene planning information; and to perform a second reconstruction of the fused scene model based on environmental influence factors within a preset imaging area to obtain an enhanced scene model.
[0019] The third determining module is used to determine the aerial imaging size and imaging angle parameters corresponding to aerial levitation imaging based on the enhanced scene model and user interaction information; adjust the aerial imaging size based on the user's head posture data relative to the preset imaging area and the spatial distribution relationship of key objects in the enhanced scene model; and adjust the imaging angle parameters based on the relative azimuth angle between the user's spatial position and the preset imaging area.
[0020] The selection module is used to select a levitation imaging device based on the straight-line distance between the user and the preset imaging area; and to control the levitation imaging device to perform aerial levitation imaging display of the enhanced scene model within the preset imaging area according to the aerial imaging size and the imaging angle parameters.
[0021] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the AI-driven aerial levitation imaging method as described in any of the above embodiments.
[0022] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, and when executed by a processor, the computer program implements the steps of the AI-driven aerial levitation imaging method as described in any of the above embodiments.
[0023] The AI-driven aerial levitation imaging method provided in this application acquires multi-view depth image data and multi-view color image data of a target imaging scene; determines the position information of object surface points in the target imaging scene in a three-dimensional coordinate system based on the depth values of each pixel in the multi-view depth image data, and constructs three-dimensional point cloud data from different viewpoints using the position information; generates a three-dimensional spatial model corresponding to the target imaging scene based on the three-dimensional point cloud data from different viewpoints; performs layer-by-layer feature learning on the multi-view color image data through a preset feature extraction network to extract high-level structural features and high-level texture features; and integrates the high-level structural features and high-level texture features through cross-view feature matching and aggregation strategies to obtain structural feature information and texture feature information of the same key object from different viewpoints; and fuses the structural feature information and texture feature information of the same key object from different viewpoints into the three-dimensional spatial model corresponding to the target imaging scene. The process involves processing feature information to obtain a fused scene model containing spatial location and visual details. Based on user-defined scene planning information, the fused scene model undergoes a primary reconstruction. Then, based on environmental influencing factors within a preset imaging area, a secondary reconstruction is performed to obtain an enhanced scene model. Based on this enhanced scene model, the aerial imaging size and viewing angle parameters for aerial levitation imaging are determined according to user interaction information. The aerial imaging size is adjusted based on the user's head posture data relative to the preset imaging area and the spatial distribution of key objects in the enhanced scene model. The viewing angle parameters are adjusted based on the relative azimuth angle between the user's spatial location and the preset imaging area. A levitation imaging device is selected based on the straight-line distance between the user and the preset imaging area. Finally, the levitation imaging device is controlled to perform aerial levitation imaging of the enhanced scene model within the preset imaging area according to the aerial imaging size and viewing angle parameters. In this way, by pre-reconstructing a 3D model of the imaging scene based on user needs and the imaging environment, and then selecting an imaging device adapted to the user's spatial perspective for aerial levitation imaging, the imaging effect can meet user needs while providing a better visual experience, effectively improving the quality of aerial levitation imaging.
[0024] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein:
[0026] Figure 1This is a flowchart illustrating an AI-driven aerial levitation imaging method disclosed herein.
[0027] Figure 2 This is a schematic diagram of the structure of an AI-driven aerial levitation imaging device disclosed herein.
[0028] Figure 3 This is a schematic diagram of the structure of a computer device provided in this disclosure.
[0029] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.
[0031] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, the statement of “connecting” or “coupling” two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.
[0032] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0033] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or part of a component) from another component (or another part of a component).
[0034] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).
[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0036] Figure 1 This is a flowchart illustrating an AI-driven aerial levitation imaging method provided in this disclosure. Figure 1 As shown, the specific process of the AI-driven aerial levitation imaging method includes:
[0037] S110. Acquire multi-view depth image data and multi-view color image data of the target imaging scene; determine the position information of the object surface points in the target imaging scene in the three-dimensional coordinate system based on the depth value of each pixel in the multi-view depth image data, construct three-dimensional point cloud data under different viewpoints based on the position information, and generate a three-dimensional spatial model corresponding to the target imaging scene based on the three-dimensional point cloud data under different viewpoints.
[0038] The target imaging scene is a virtual / real scene to be imaged while suspended in the air. Multi-view depth image data is obtained by acquiring images of the target imaging scene from different angles using a depth sensor. Each angle corresponds to a set of depth images containing distance information to the surfaces of objects in the scene. Multi-view color image data is simultaneously acquired by a color camera to record the texture, color, and other appearance features of the target imaging scene from different perspectives.
[0039] When determining the position information of a point on the object's surface in a three-dimensional coordinate system, the optical center of the depth sensor can be used as the origin. Combining the pixel's coordinates in the image coordinate system with its corresponding depth value, the precise coordinates of the surface point in three-dimensional space can be calculated using a coordinate transformation formula. After constructing three-dimensional point cloud data from different viewpoints using the position information, point cloud registration algorithms can be used to fuse these point cloud data from different viewpoints. This eliminates local occlusion and data redundancy caused by viewpoint differences, thereby constructing a three-dimensional spatial model that can completely reflect the geometric structure of the target imaging scene. The three-dimensional spatial model includes the shape, size, and relative positional relationships of each object in the target imaging scene.
[0040] S120. Through a preset feature extraction network, multi-view color image data is subjected to layer-by-layer feature learning to extract high-level structural features and high-level texture features; and through cross-view feature matching and aggregation strategies, high-level structural features and high-level texture features are associated and integrated to obtain structural feature information and texture feature information of the same key object under different views.
[0041] The preset feature extraction network can adopt an improved residual network architecture. While preserving the low-level detailed features, it introduces an attention mechanism module to dynamically adjust the weights of different channel features, so that the network can focus more on the salient areas of key objects, such as edge contours and unique textures, thereby improving the discriminative power of high-level features.
[0042] Key objects are core target objects in the target imaging scene that have clear geometric contours, unique texture features, or specific functional attributes, such as furniture (such as sofas, tables and chairs) and electrical appliances (such as televisions and air conditioners) in indoor scenes; and buildings, trees, and traffic signs in outdoor scenes.
[0043] In some embodiments, a cross-view feature matching and aggregation strategy is used to associate and integrate high-level structural features and high-level texture features to obtain structural and texture feature information of the same key object from different viewpoints. This includes: constructing a multi-view feature association matrix based on high-level structural features and high-level texture features; and using a graph neural network to aggregate features from different viewpoints based on the multi-view feature association matrix to perform cross-view feature fusion of high-level structural features and high-level texture features to obtain structural and texture feature information of the same key object from different viewpoints.
[0044] In this matrix, the row vectors of the multi-view feature association matrix correspond to the high-level structural feature vectors from different viewpoints, and the column vectors correspond to the high-level texture feature vectors from different viewpoints. The matrix element values represent the similarity between the corresponding structural features and texture features. Similarity can be calculated using metrics such as cosine similarity or Euclidean distance, by quantifying the correlation strength between the high-level structural feature vectors and high-level texture feature vectors from different viewpoints.
[0045] When performing graph neural network aggregation based on multi-view feature association matrices, a graph structure representing the multi-view feature associations can be constructed by using the feature vectors of each viewpoint as graph nodes and the matrix element values as edge weights between nodes. Graph neural networks can leverage message passing mechanisms to enable each node to aggregate feature information from other nodes with high similarity associations, achieving deep interaction and complementarity between high-level structural features and high-level texture features from different viewpoints. For example, when the high-level structural features of a key object are relatively clear from a certain viewpoint but lack texture details, a graph neural network can aggregate rich texture information from nodes with highly similar high-level texture features from other viewpoints, and vice versa. This ensures that the generated structural and texture feature information of the same key object possesses both completeness and accuracy across multiple viewpoints.
[0046] S130. In the three-dimensional spatial model corresponding to the target imaging scene, the structural feature information and texture feature information of the same key object under different perspectives are fused to obtain a fused scene model containing spatial position and visual details; the fused scene model is reconstructed once according to the user-defined scene planning information; and the fused scene model is reconstructed a second time according to the environmental influence factors in the preset imaging area to obtain an enhanced scene model.
[0047] The user-defined scene planning information can include parameters such as imaging perspective, zoom level, key object highlighting mode, virtual objects, and environment configuration. For example, users can use interactive commands to set a key object as the imaging center and zoom in to a specific scale, or adjust the overall scene viewing angle to highlight the spatial layout of the target area. Preset environmental influencing factors within the imaging area can include physical parameters such as light intensity, air refractive index, and density of the suspended imaging medium. Sensors collect dynamic change data of these environmental influencing factors in real time, and combined with physical optics simulation algorithms, the lighting rendering effect, object edge sharpness, and suspension imaging stability of the fused scene model are corrected. For example, when a sudden increase in light intensity in the imaging area is detected, causing local overexposure, the pixel brightness and contrast parameters of the corresponding area will be automatically adjusted to ensure that the enhanced scene model can present an aerial suspension imaging effect that conforms to human visual habits in the actual physical environment.
[0048] In some embodiments, a first-time reconstruction of the fused scene model is performed based on user-defined scene planning information, including: obtaining scene layout parameters and environment layout parameters from the user-defined scene planning information; and updating the fused scene model based on the scene layout parameters and environment layout parameters to complete the first-time reconstruction of the fused scene model.
[0049] The scene layout parameters include: the type, quantity, three-dimensional coordinates, size ratio, and spatial arrangement of virtual objects; the environment layout parameters include: the location coordinates of ground textures and wall markers within the preset imaging area, as well as the distribution of light intensity.
[0050] Specifically, based on the type of virtual object, the corresponding 3D base model is retrieved from the model database. The number of model instantiations is then determined by combining this with quantity information. Each virtual object instance is precisely placed in the 3D spatial coordinate system of the fused scene model according to its 3D coordinates and size ratio. Simultaneously, the relative positions, orientations, and occlusion levels of the virtual objects are adjusted based on their spatial arrangement. Corresponding texture maps are loaded according to the type of ground texture within the preset imaging area (e.g., marble, wood, grass), and seamlessly stitched together based on their position coordinates in the 3D coordinate system, ensuring the ground texture matches the actual physical boundary of the imaging area. The model size is adjusted according to the actual size ratio of the landmark to create a natural spatial integration with the virtual wall. Light intensity distribution parameters are set as the basic lighting parameters for ambient light in the fused scene model, such as the light direction, color temperature, and intensity value for each area. This achieves precise updates to the spatial structure, object distribution, and basic environmental representation of the fused scene model.
[0051] In some embodiments, the fusion scene model is reconstructed a second time based on environmental impact factors within a preset imaging area to obtain an enhanced scene model. This includes: inputting environmental impact factors within the preset imaging area into a pre-trained scene enhancement network to obtain imaging parameter compensation values for virtual objects; and optimizing the fusion scene model after the first reconstruction frame by frame using the imaging parameter compensation values for virtual objects to obtain an enhanced scene model.
[0052] The scene enhancement network employs an architecture combining a deep residual network and an attention mechanism. The input layer receives temporal sequence data of environmental influencing factors and the feature map of the fused scene model after a single reconstruction. The intermediate layers extract the correlation features between environmental parameters and the imaging quality of virtual objects through convolutional blocks. The attention module focuses on key environmental factors that significantly affect image sharpness, color fidelity, and suspension stability. The output layer generates imaging parameter compensation values, such as spatial position offset compensation values for virtual objects, texture detail enhancement coefficients, illumination attenuation correction parameters, and dynamic blur suppression weights. Subsequently, based on the imaging parameter compensation values, the 3D coordinates of each virtual object in the fused scene model are fine-tuned at the sub-pixel level. The RGB channel values of the object surface texture are enhanced pixel-by-pixel, the ambient light attenuation function is corrected in real time, and the imaging motion blur caused by airflow disturbances is eliminated through the inverse operation of the motion blur convolution kernel. This allows the optimized enhanced scene model to dynamically adapt to complex environmental changes in the preset imaging area, ensuring the stability and realism of the aerial suspension imaging effect.
[0053] S140. Based on the enhanced scene model, determine the aerial imaging size and imaging angle parameters corresponding to aerial levitation imaging according to user interaction information; adjust the aerial imaging size according to the relationship between the user's head posture data relative to the preset imaging area and the spatial distribution of key objects in the enhanced scene model; and adjust the imaging angle parameters according to the relative azimuth angle between the user's spatial position and the preset imaging area.
[0054] The user interaction information may include imaging size scaling requests (such as "zoom in to 150% of the original size" or "zoom out to a size suitable for one-handed operation") and viewpoint switching commands (such as "switch to top view" or "display with object A as the center") input through voice commands, gestures, or eye tracking. Imaging viewpoint parameters include: horizontal viewpoint, vertical viewpoint, and viewpoint rotation angle.
[0055] In some embodiments, adjusting the aerial imaging size based on the user's head pose data relative to a preset imaging area and the spatial distribution relationship of key objects in the enhanced scene model includes: constructing a three-dimensional head pose matrix based on the user's head pose data relative to the preset imaging area; calculating the spatial distribution density of key objects in the enhanced scene model based on the spatial distribution relationship of key objects in the enhanced scene model; performing a correlation analysis on the user's three-dimensional head pose matrix and the spatial distribution density of key objects in the enhanced scene model; and optimizing the aerial imaging size based on the correlation analysis results.
[0056] The user's head 3D pose matrix includes attitude parameters such as pitch, yaw, and roll angles in the spatial coordinate system. The spatial distribution density of key objects in the enhanced scene model can be obtained by calculating the number of key objects per unit volume and the distance distribution between them. For example, when key objects are densely distributed in front of a preset imaging area, the spatial distribution density value is high.
[0057] The posture parameters in the user's 3D head posture matrix and the spatial distribution density values of key objects can be used as feature inputs. The correlation analysis results between the two can be obtained by training a neural network model or a multivariate regression algorithm. If the analysis results show that the user's head posture tends to be forward and the distribution density of key objects is high, it indicates that the user needs to observe details at close range. At this time, the aerial imaging size can be automatically increased to ensure that the imaging content is clear and distinguishable. If the user's head is tilted back and the distribution of key objects is sparse, the imaging size can be appropriately reduced to avoid the imaging area being too large and affecting the user's perception of the overall scene.
[0058] In some embodiments, the imaging viewpoint parameters are adjusted according to the relative azimuth angle between the user's spatial position and the preset imaging area, including: determining the deflection angle adjustment amount and the pitch angle adjustment amount of the imaging viewpoint according to the relative azimuth angle between the user's spatial position and the preset imaging area; and optimizing the imaging viewpoint parameters according to the deflection angle adjustment amount and the pitch angle adjustment amount.
[0059] The relative azimuth angle can be calculated using the coordinates of the user's spatial position and the center coordinates of the preset imaging area. For example, a spatial rectangular coordinate system can be established with the center of the preset imaging area as the origin. The user's spatial position can be projected onto the XOY and XOZ planes of this coordinate system, and the horizontal deflection angle and vertical pitch angle can be calculated respectively. The deflection angle adjustment is positively correlated with the horizontal relative azimuth angle (i.e., the horizontal deflection angle). When the user is to the left of the center of the preset imaging area, the deflection angle adjustment is positive, controlling the imaging viewpoint to deflect to the left by the corresponding angle; when the user is to the right, the adjustment is negative, and the imaging viewpoint deflects to the right. The pitch angle adjustment corresponds to the vertical relative azimuth angle (i.e., the vertical pitch angle). When the user's position is higher than the center of the imaging area, the pitch angle adjustment is positive, and the imaging viewpoint tilts upward; when the user's position is lower than the center of the area, the pitch angle adjustment is negative, and the imaging viewpoint tilts downward.
[0060] Therefore, by collecting the user's spatial position coordinates in real time and dynamically updating the relative azimuth angle and the corresponding deflection and pitch angle adjustments, the optimized imaging viewpoint is always oriented towards the user's current position, ensuring that the user can observe the aerial imaging content from a frontal and comfortable perspective, thus enhancing the naturalness and immersion of the viewing experience.
[0061] S150. Select a levitation imaging device based on the straight-line distance between the user and the preset imaging area; and control the levitation imaging device to perform aerial levitation imaging display of the enhanced scene model within the preset imaging area according to the aerial imaging size and imaging angle parameters.
[0062] The levitation imaging device can be composed of a light source assembly, a beam splitter assembly, and a reflective freeform mirror. Light emitted from the light source assembly is partially reflected and partially transmitted by the beam splitter. The reflected light is then reflected again by the freeform mirror, passes through the beam splitter assembly, and forms a levitation image in the air. Specifically, the beam splitter assembly can employ a multi-layer dielectric film structure with a reflectivity-to-transmittance ratio between 30% and 70% to achieve high brightness and high definition levitation images. The surface shape of the freeform mirror can be precisely calculated and optimized to facilitate accurate control of the incident light, effectively eliminating aberrations and distortions during the imaging process, ensuring the levitation image has good stereoscopic effect and spatial positioning accuracy. Simultaneously, the light source assembly can use a high-brightness LED array or laser light source, coupled with a dedicated driving circuit to achieve brightness adjustment and color temperature control, adapting to imaging needs under different ambient light conditions and ensuring clear visibility of the levitation image in both strong and low light environments.
[0063] In some embodiments, selecting a levitation imaging device based on the straight-line distance between the user and a preset imaging area includes: comparing the straight-line distance between the user and the preset imaging area with the effective imaging distance range of each imaging device in the database to obtain a set of candidate imaging devices; giving each imaging device a comprehensive score based on the resolution attenuation coefficient of each imaging device in the candidate imaging device set at the straight-line distance; and determining the imaging device with the highest comprehensive score as the levitation imaging device.
[0064] The database pre-stores the effective imaging distance range of each imaging device. This range is determined through multiple experiments measuring the imaging clarity, stability, and levitation effect of the devices at different distances. The resolution attenuation coefficient measures the degree of resolution degradation of an imaging device at a specific straight-line distance. A larger value indicates more severe resolution attenuation. The resolution attenuation coefficient is calculated by comparing the reference resolution of the imaging device at a standard distance (e.g., 1 meter) with its actual resolution at a straight-line distance. The formula is: Resolution attenuation coefficient = (Reference resolution - Actual resolution) / Reference resolution.
[0065] When conducting a comprehensive evaluation, the resolution attenuation coefficient can be combined with the imaging device's response speed, power consumption, and cost. For example, the weights of the resolution attenuation coefficient (0.5), response speed (0.2), power consumption (0.15), and cost (0.15) can be weighted and summed. This facilitates the selection of the imaging device with the highest user compatibility.
[0066] Furthermore, if there are at least two imaging devices with the highest overall scores, the system can interact with the user to allow them to select one of these devices as the levitation imaging device based on their preference. During this interaction, the system can display detailed parameter comparison information for each imaging device with the highest overall score, including but not limited to the specific value of the resolution attenuation coefficient, measured data on response speed (such as the time taken from startup to stable imaging), power consumption curves under different operating modes, and detailed equipment procurement and maintenance costs, presented in visual charts to help users intuitively understand the differences between the devices. Simultaneously, the system can proactively recommend device options that better suit the actual application conditions based on the user's historical usage preferences (such as prioritizing response speed or cost in similar scenarios) or the specific needs of the current imaging task (such as precision detection scenarios with extremely high requirements for imaging stability, or outdoor mobile scenarios with specific limitations on device portability), further improving the efficiency and accuracy of the user's decision-making.
[0067] In this embodiment, multi-view depth image data and multi-view color image data of the target imaging scene are acquired; the position information of the object surface points in the target imaging scene in the three-dimensional coordinate system is determined based on the depth value of each pixel in the multi-view depth image data, and three-dimensional point cloud data under different viewpoints is constructed using the position information; and a three-dimensional spatial model corresponding to the target imaging scene is generated based on the three-dimensional point cloud data under different viewpoints; a preset feature extraction network is used to perform layer-by-layer feature learning on the multi-view color image data to extract high-level structural features and high-level texture features; and a cross-view feature matching and aggregation strategy is used to associate and integrate the high-level structural features and high-level texture features to obtain the structural feature information and texture feature information of the same key object under different viewpoints; the structural feature information and texture feature information of the same key object under different viewpoints are fused in the three-dimensional spatial model corresponding to the target imaging scene to obtain a package The process involves: creating a fused scene model incorporating spatial location and visual details; reconstructing the fused scene model based on user-defined scene planning information; reconstructing the fused scene model a second time based on environmental influencing factors within a preset imaging area to obtain an enhanced scene model; determining the aerial imaging size and imaging perspective parameters for levitation imaging based on the enhanced scene model and user interaction information; adjusting the aerial imaging size based on the user's head posture data relative to the preset imaging area and the spatial distribution of key objects in the enhanced scene model; adjusting the imaging perspective parameters based on the relative azimuth angle between the user's spatial location and the preset imaging area; selecting a levitation imaging device based on the straight-line distance between the user and the preset imaging area; and controlling the levitation imaging device to display the enhanced scene model in levitation imaging within the preset imaging area according to the aerial imaging size and imaging perspective parameters. In this way, by pre-reconstructing the 3D model of the imaging scene based on user needs and the imaging environment, and then selecting an imaging device adapted to the user's spatial perspective for levitation imaging, the system can provide users with a better visual experience while meeting their imaging needs, effectively improving the quality of levitation imaging.
[0068] Figure 2 This is a schematic diagram of the structure of an AI-driven aerial levitation imaging device provided in this embodiment. The AI-driven aerial levitation imaging device may include:
[0069] The acquisition module 210 is used to acquire multi-view depth image data and multi-view color image data of the target imaging scene.
[0070] The first determining module 220 is used to determine the position information of the object surface points in the target imaging scene in the three-dimensional coordinate system based on the depth value of each pixel in the multi-view depth image data, construct three-dimensional point cloud data under different viewpoints through the position information, and generate a three-dimensional spatial model corresponding to the target imaging scene based on the three-dimensional point cloud data under different viewpoints.
[0071] The second determining module 230 is used to perform layer-by-layer feature learning on multi-view color image data through a preset feature extraction network to extract high-level structural features and high-level texture features; and to associate and integrate high-level structural features and high-level texture features through cross-view feature matching and aggregation strategies to obtain structural feature information and texture feature information of the same key object under different viewpoints.
[0072] The fusion module 240 is used to fuse the structural and texture features of the same key object from different perspectives in the three-dimensional spatial model corresponding to the target imaging scene, so as to obtain a fused scene model containing spatial position and visual details.
[0073] The reconstruction module 250 is used to perform a first reconstruction of the fused scene model based on user-defined scene planning information; and to perform a second reconstruction of the fused scene model based on environmental influence factors within a preset imaging area, to obtain an enhanced scene model.
[0074] The third determining module 260 is used to determine the aerial imaging size and imaging angle parameters corresponding to aerial levitation imaging based on the enhanced scene model and user interaction information; adjust the aerial imaging size based on the user's head posture data relative to the preset imaging area and the spatial distribution relationship of key objects in the enhanced scene model; and adjust the imaging angle parameters based on the relative azimuth angle between the user's spatial position and the preset imaging area.
[0075] The selection module 270 is used to select a levitation imaging device based on the straight-line distance between the user and the preset imaging area; and to control the levitation imaging device to perform aerial levitation imaging display of the enhanced scene model within the preset imaging area according to the aerial imaging size and imaging angle parameters.
[0076] In this embodiment, optionally, the reconstruction module 250 is specifically used for:
[0077] The scene layout parameters and environment layout parameters are obtained from the user-defined scene planning information. The scene layout parameters include: the type, quantity, three-dimensional coordinates, size ratio and spatial arrangement of virtual objects. The environment layout parameters include: the location coordinates of ground texture and wall markers in the preset imaging area and the distribution of light intensity. The fused scene model is updated according to the scene layout parameters and environment layout parameters to complete the first reconstruction of the fused scene model.
[0078] In this embodiment, optionally, the reconstruction module 250 is specifically used for:
[0079] Environmental impact factors within a preset imaging area are input into a pre-trained scene enhancement network to obtain imaging parameter compensation values for virtual objects. The fused scene model after one reconstruction is then optimized frame by frame using the imaging parameter compensation values of the virtual objects to obtain the enhanced scene model.
[0080] In this embodiment, optionally, the third determining module 260 is specifically used for:
[0081] A three-dimensional head pose matrix is constructed based on the user's head pose data relative to the preset imaging area; the spatial distribution density of key objects in the enhanced scene model is calculated based on the spatial distribution relationship of key objects in the enhanced scene model; a correlation analysis is performed on the user's three-dimensional head pose matrix and the spatial distribution density of key objects in the enhanced scene model, and the aerial imaging size is optimized based on the correlation analysis results.
[0082] In this embodiment, optionally, the third determining module 260 is specifically used for:
[0083] Based on the relative azimuth angle between the user's spatial location and the preset imaging area, determine the deflection angle adjustment and pitch angle adjustment of the imaging viewpoint; based on the deflection angle adjustment and pitch angle adjustment, optimize the imaging viewpoint parameters.
[0084] In this embodiment, optionally, the selection module 270 is specifically used for:
[0085] The straight-line distance between the user and the preset imaging area is compared with the effective imaging distance range of each imaging device in the database to obtain a set of candidate imaging devices. Based on the resolution attenuation coefficient of each imaging device in the candidate imaging device set at the straight-line distance, each imaging device is comprehensively scored. The imaging device with the highest comprehensive score is determined as the levitation imaging device.
[0086] In this embodiment, optionally, the second determining module 230 is specifically used for:
[0087] A multi-view feature association matrix is constructed based on high-level structural features and high-level texture features. The row vectors of the multi-view feature association matrix correspond to the high-level structural feature vectors under different views, and the column vectors correspond to the high-level texture feature vectors under different views. The matrix element values represent the similarity between the corresponding structural features and texture features. Based on the multi-view feature association matrix, a graph neural network is used to aggregate the features from different views to perform cross-view feature fusion of high-level structural features and high-level texture features, thereby obtaining the structural feature information and texture feature information of the same key object under different views.
[0088] The AI-driven aerial levitation imaging device provided in this disclosure can execute the above-described method embodiments. For its specific implementation principle and technical effects, please refer to the above-described method embodiments. This disclosure will not repeat them here.
[0089] This application also provides a computer device. Please refer to the following for details. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.
[0090] The computer device includes a memory 310 and a processor 320 that are interconnected via a system bus. It should be noted that only a computer device with memory 310 and processor 320 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0091] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0092] The memory 310 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, or flash card equipped on the computer device. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory 310 is typically used to store the operating system and various application software installed on the computer device, such as the program code of the method described above. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.
[0093] Processor 320 is typically used to perform overall operations of a computer device. In this embodiment, memory 310 is used to store program code or instructions, including computer operation instructions, and processor 320 is used to execute the program code or instructions stored in memory 310 or process data, such as program code that runs the methods described above.
[0094] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0095] Another embodiment of this application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads computer-readable program code stored in the computer-readable medium, enabling the processor to execute the functional actions specified in each step or combination of steps in the above method; and to generate means for implementing the functional actions specified in each block or combination of blocks in the block diagram.
[0096] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, wherein the memory is used to store program code or instructions, the program code including computer operation instructions, and the processor is used to execute the program code or instructions of the above-described methods stored in the memory.
[0097] The definitions of memory and processor can be found in the description of the foregoing computer device embodiments, and will not be repeated here.
[0098] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0099] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0100] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0101] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" as described in this application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several units of these means may be embodied by the same item of hardware. The use of "first," "second," and "third," etc., does not indicate any order and these words should be interpreted as names. Unless otherwise specified, the steps in the above embodiments should not be construed as limiting the order of execution.
[0102] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An AI-driven aerial levitation imaging method, characterized in that, include: Acquire multi-view depth image data and multi-view color image data of the target imaging scene; Based on the depth values of each pixel in the multi-view depth image data, the position information of the object surface points in the target imaging scene in the three-dimensional coordinate system is determined, and three-dimensional point cloud data under different viewpoints is constructed using the position information; and a three-dimensional spatial model corresponding to the target imaging scene is generated based on the three-dimensional point cloud data under different viewpoints. The multi-view color image data is subjected to layer-by-layer feature learning through a preset feature extraction network to extract high-level structural features and high-level texture features. Furthermore, the high-level structural features and high-level texture features are associated and integrated through cross-view feature matching and aggregation strategies to obtain structural feature information and texture feature information of the same key object from different viewpoints. The structural and texture features of the same key object from different perspectives are fused into the three-dimensional spatial model corresponding to the target imaging scene to obtain a fused scene model that includes spatial location and visual details. The fused scene model is reconstructed once based on user-defined scene planning information; and the fused scene model is reconstructed a second time based on environmental influencing factors within a preset imaging area to obtain an enhanced scene model. Based on the enhanced scene model, the aerial imaging size and imaging angle parameters corresponding to aerial hovering imaging are determined according to user interaction information. The aerial imaging size is adjusted based on the user's head pose data relative to the preset imaging area and the spatial distribution relationship of key objects in the enhanced scene model. The imaging angle parameters are adjusted according to the relative azimuth angle between the user's spatial location and the preset imaging area. The levitation imaging device is selected based on the straight-line distance between the user and the preset imaging area; The system controls the levitation imaging device to perform aerial levitation imaging display of the enhanced scene model within the preset imaging area, according to the aerial imaging size and the imaging angle parameters.
2. The method according to claim 1, characterized in that, Based on user-defined scene planning information, the fused scene model is reconstructed once, including: The scene layout parameters and environment layout parameters are obtained from the user-defined scene planning information. The scene layout parameters include: the type, quantity, three-dimensional coordinates, size ratio, and spatial arrangement of virtual objects. The environment layout parameters include: the location coordinates of the ground texture and wall markers within the preset imaging area, as well as the light intensity distribution. The fused scene model is updated based on the scene layout parameters and the environment layout parameters to complete the first reconstruction of the fused scene model.
3. The method according to claim 1, characterized in that, Based on environmental influence factors within a preset imaging area, the fused scene model undergoes secondary reconstruction to obtain an enhanced scene model, including: The environmental impact factors within the preset imaging area are input into the pre-trained scene enhancement network to obtain the imaging parameter compensation values of the virtual object. The fused scene model after one reconstruction is optimized frame by frame by using the imaging parameter compensation values of the virtual object to obtain the enhanced scene model.
4. The method according to claim 1, characterized in that, Based on the user's head pose data relative to the preset imaging area and the spatial distribution relationship of key objects in the enhanced scene model, the aerial imaging size is adjusted, including: A three-dimensional head pose matrix is constructed based on the user's head pose data relative to the preset imaging area; and the spatial distribution density of key objects in the enhanced scene model is calculated based on the spatial distribution relationship of key objects in the enhanced scene model. A correlation analysis is performed on the three-dimensional pose matrix of the user's head and the spatial distribution density of key objects in the enhanced scene model, and the aerial imaging size is optimized based on the correlation analysis results.
5. The method according to claim 1, characterized in that, Adjusting the imaging viewpoint parameters based on the relative azimuth angle between the user's spatial location and the preset imaging area includes: Based on the relative azimuth angle between the user's spatial location and the preset imaging area, determine the deflection angle adjustment amount and the pitch angle adjustment amount of the imaging viewpoint; The imaging viewpoint parameters are optimized based on the deflection angle adjustment and the pitch angle adjustment.
6. The method according to claim 1, characterized in that, Selecting a levitation imaging device based on the straight-line distance between the user and the preset imaging area includes: The straight-line distance between the user and the preset imaging area is compared with the effective imaging distance range of each imaging device in the database to obtain a set of candidate imaging devices. Based on the resolution attenuation coefficient of each imaging device in the candidate imaging device set at the straight-line distance, a comprehensive score is given to each imaging device; and the imaging device with the highest comprehensive score is determined as the levitation imaging device.
7. The method according to claim 1, characterized in that, The high-level structural features and high-level texture features are associated and integrated through a cross-view feature matching and aggregation strategy to obtain structural and texture feature information of the same key object from different viewpoints, including: A multi-view feature association matrix is constructed based on the high-level structural features and the high-level texture features; the row vectors of the multi-view feature association matrix correspond to the high-level structural feature vectors under different views, the column vectors of the multi-view feature association matrix correspond to the high-level texture feature vectors under different views, and the matrix element values represent the similarity between the corresponding structural features and texture features. Based on the multi-view feature association matrix, a graph neural network is used to aggregate features from different perspectives to perform cross-view feature fusion of the high-level structural features and the high-level texture features, thereby obtaining structural feature information and texture feature information of the same key object from different perspectives.
8. An AI-driven aerial levitation imaging device, characterized in that, include: The acquisition module is used to acquire multi-view depth image data and multi-view color image data of the target imaging scene; The first determining module is used to determine the position information of the object surface points in the target imaging scene in the three-dimensional coordinate system based on the depth value of each pixel in the multi-view depth image data, construct three-dimensional point cloud data under different viewpoints based on the position information, and generate a three-dimensional spatial model corresponding to the target imaging scene based on the three-dimensional point cloud data under different viewpoints. The second determining module is used to perform layer-by-layer feature learning on the multi-view color image data through a preset feature extraction network to extract high-level structural features and high-level texture features. Furthermore, the high-level structural features and high-level texture features are associated and integrated through cross-view feature matching and aggregation strategies to obtain structural feature information and texture feature information of the same key object from different viewpoints. The fusion module is used to fuse the structural feature information and texture feature information of the same key object from different perspectives in the three-dimensional spatial model corresponding to the target imaging scene, so as to obtain a fused scene model containing spatial position and visual details. The reconstruction module is used to perform a first reconstruction of the fused scene model based on user-defined scene planning information; and to perform a second reconstruction of the fused scene model based on environmental influence factors within a preset imaging area to obtain an enhanced scene model. The third determining module is used to determine the aerial imaging size and imaging angle parameters corresponding to aerial levitation imaging based on the enhanced scene model and user interaction information. The aerial imaging size is adjusted based on the user's head pose data relative to the preset imaging area and the spatial distribution relationship of key objects in the enhanced scene model. The imaging angle parameters are adjusted according to the relative azimuth angle between the user's spatial location and the preset imaging area. The selection module is used to select a levitation imaging device based on the straight-line distance between the user and the preset imaging area; The system controls the levitation imaging device to perform aerial levitation imaging display of the enhanced scene model within the preset imaging area, according to the aerial imaging size and the imaging angle parameters.
9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the AI-driven aerial levitation imaging method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the AI-driven aerial levitation imaging method as described in any one of claims 1 to 7.