AR glasses real-time object scanning and modeling system

By integrating high-precision cameras and depth sensors in AR glasses, combining parallel computing and advanced algorithms, real-time object scanning and modeling is realized, solving the problems of insufficient data processing efficiency and model accuracy in the existing technology, and achieving efficient and accurate three-dimensional modeling.

CN120088381APending Publication Date: 2025-06-03GUDONG TECH CO LTD

Patent Information

Application Number
CN202510253612.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing object scanning and modeling technologies are difficult to achieve ideal states in terms of data processing efficiency and model accuracy. Traditional algorithms have a long calculation time when processing large amounts of scanned data, and real-time modeling is not possible, and the generated models are insufficient in detail restoration.

Method used

The real-time object scanning and modeling system of AR glasses is adopted. By integrating high-precision cameras and depth sensors, combining parallel computing technology and advanced algorithms, such as feature extraction, point cloud generation and three-dimensional reconstruction algorithms based on convolutional neural networks, real-time scanning and modeling is realized, and a variety of interaction methods are provided.

Benefits of technology

Real-time modeling and rich details are realized. The hardware synchronous calibration module and parallel computing technology further improve system performance, simplify equipment use, and improve work efficiency and model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088381A_ABST
    Figure CN120088381A_ABST
Patent Text Reader

Abstract

The invention discloses an AR glasses real-time object scanning and modeling system, and the system comprises a hardware layer which comprises an AR glasses main body, and a high-precision camera, a depth sensor, a processor and a storage module which are integrated on the AR glasses main body; the data acquisition layer is used for carrying out multi-angle shooting on an object at a high frame rate based on the high-precision camera to obtain a two-dimensional image sequence, synchronously obtaining depth information of each point on the surface of the object by the depth sensor, and carrying out accurate time and space alignment on data of the depth information and the depth information; the algorithm layer comprises a feature extraction algorithm, a point cloud generation algorithm and a three-dimensional reconstruction algorithm; and the display layer is used for rendering the constructed three-dimensional model into a display interface of AR glasses in real time and displaying the model and a real scene in a fused manner. Real-time modeling is achieved, details are rich, a display layer renders in real time, various interactions are provided, and operation is convenient and visual. And a hardware synchronous calibration module and a parallel computing technology further improve the system performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of AR glasses, and in particular to a real-time object scanning and modeling system for AR glasses. Background Art

[0003] In today's digital age, object scanning and modeling technologies have extensive application requirements in many fields such as industrial design, cultural relic protection, and medical rehabilitation. Traditional object scanning and modeling methods often rely on large professional devices, such as laser scanners. Although these devices can obtain relatively accurate object data.

[0004] Chinese Patent Grant Publication No. CN210109870U, Grant Publication Date February 21, 2020, an AR-based model reconstruction system, including an AR scanner for scanning an actual object to obtain a point cloud of the geometric surface of the actual object, the point cloud including the three-dimensional coordinates and color information of the actual object; a controller, signal-connected to the AR scanner, for receiving the point cloud and performing calculations to generate a modeling signal representing the three-dimensional modeling of the actual object; a VR imaging glasses, signal-connected to the controller, for restoring the three-dimensional modeling according to the modeling signal.

[0005] There are many limitations in the prior art. From the perspective of data processing, the existing scanning and modeling technologies are difficult to achieve an ideal state in terms of data processing efficiency and model accuracy. When traditional algorithms process a large amount of scanning data, the calculation time is long and real-time modeling cannot be achieved. For example, when scanning a complex object, conventional point cloud generation algorithms and three-dimensional reconstruction algorithms may take hours or even days to complete the modeling, seriously affecting work efficiency. Moreover, due to the limitations of the algorithms, the generated models have deficiencies in detail restoration. For some objects with complex textures and shapes, the reconstructed models cannot accurately present their true features, which makes it difficult to be efficiently applied in complex and changing real-world scenarios. For this reason, we propose a real-time object scanning and modeling system for AR glasses. Summary of the Invention

[0007] An object of the present invention is to solve at least the above problems and provide at least the advantages described later.

[0008] Another object of the present invention is to provide a real-time object scanning and modeling system for AR glasses, which can achieve real-time modeling with rich details, the display layer performs real-time rendering and provides diverse interactions, and the operation is convenient and intuitive. The hardware synchronization calibration module and parallel computing technology further improve the system performance.

[0009] To achieve the above object and some other objects, the present invention adopts the following technical solutions: An AR glasses real-time object scanning and modeling system, comprising: a hardware layer, including an AR glasses main body, as well as a high-precision camera, a depth sensor, a processor, and a storage module integrated on the AR glasses main body; the high-precision camera is used to capture the texture information of the object surface, the depth sensor is used to obtain the depth information of each point on the object surface, the processor is used to process the collected data and run a 3D reconstruction algorithm, and the storage module is used to temporarily store the collected data and intermediate calculation results; A data acquisition layer, based on the high-precision camera, takes multi-angle pictures of the object at a high frame rate to obtain a two-dimensional image sequence. At the same time, the depth sensor synchronously obtains the depth information of each point on the object surface, and precisely aligns the two sets of data in terms of time and space; An algorithm layer, including a feature extraction algorithm, a point cloud generation algorithm, and a 3D reconstruction algorithm; the feature extraction algorithm uses a model based on a convolutional neural network to extract features from the collected image data, the point cloud generation algorithm combines image feature points and depth data to calculate the coordinates of each point on the object surface in 3D space and generates point cloud data, and the 3D reconstruction algorithm constructs and optimizes the surface mesh model of the object based on the point cloud data; A display layer, which renders the constructed 3D model in real time onto the display interface of the AR glasses, fuses and displays it with the real-world scene, and provides interaction methods such as gesture control and voice control.

[0010] Preferably, the feature extraction model based on the convolutional neural network adopts an improved ResNet architecture. The improved ResNet architecture reduces some redundant layers of the original ResNet and adds an attention mechanism module to automatically learn the importance of different channel features and recalibrate the feature map.

[0011] Preferably, the training process of the feature extraction model based on the convolutional neural network includes: collecting a large amount of object image data covering different shapes, materials, and color features, dividing it into a training set, a validation set, and a test set; using the stochastic gradient descent algorithm combined with a momentum optimizer to update the network parameters, adopting a dynamic adjustment strategy for the learning rate, and adding a Dropout layer to the network to prevent overfitting.

[0012] Preferably, the feature point matching in the point cloud generation algorithm uses a descriptor-based matching method. Feature points are extracted from images taken from different angles through a corner detection algorithm, descriptors are calculated for each feature point, and feature point matching is performed by calculating the similarity between descriptors.

[0013] Preferably, the three-dimensional coordinate calculation in the point cloud generation algorithm is based on the principle of triangulation, combined with the internal and external parameters of the camera, the pixel coordinates of the matched feature points in different images, and the depth information of the corresponding points obtained by the depth sensor, and the coordinates of the object surface points in the three-dimensional space are solved through a system of equations.

[0014] Preferably, the point cloud generation and preprocessing in the point cloud generation algorithm include: generating point cloud data after calculating the three-dimensional coordinates of the object surface points, removing noise points by using a statistical filtering method, and using a voxel grid filtering method to reduce the density of the point cloud data and retain the main features.

[0015] Preferably, in the three-dimensional reconstruction algorithm, a surface mesh model is constructed based on the Poisson reconstruction algorithm. The point cloud data is regarded as the sampling points of a scalar field. By constructing a three-dimensional voxel grid containing the point cloud data, calculating the voxel scalar values, using a discrete Poisson equation solver to solve the Poisson equation, extracting the zero isosurface through the marching cubes algorithm to generate the surface mesh model, and performing smoothing processing and hole filling on the model.

[0016] Preferably, the model optimization of the three-dimensional reconstruction algorithm based on the moving least squares reconstruction algorithm is to construct a local polynomial approximation function in the neighborhood of each point cloud point, determine the polynomial coefficients by minimizing the sum of the squares of the distances from the points in the neighborhood to the approximation function, iteratively update the fitted surface, and perform model refinement and smoothing processing by increasing the number of neighborhood points, adjusting the polynomial order, etc.

[0017] Preferably, it further includes a hardware synchronization and calibration module, which is used to achieve a time synchronization accuracy of microseconds when the high-precision camera and the depth sensor collect data, accurately calibrate the two, and establish a spatial mapping relationship.

[0018] Preferably, the algorithms in the algorithm layer are optimized and accelerated by using the multi-core performance of the processor through parallel computing technology to improve the real-time performance of the system.

[0019] The present invention has at least the following beneficial effects: 1. Compared with the existing AR-based model reconstruction system, this AR glasses real-time object scanning and modeling system innovatively integrates a high-precision camera and a depth sensor in the AR glasses. The hardware works collaboratively and, combined with advanced algorithms, realizes real-time scanning and modeling. The prior art may have insufficient hardware integration and cannot achieve such a tightly integrated real-time function. This integration method reduces the complexity of the device, and users do not need to carry multiple independent devices for scanning, improving the convenience.

[0020] 2. Compared with the existing AR-based model reconstruction systems, this AR glasses real-time object scanning and modeling system adopts an improved ResNet architecture and adds an attention mechanism module. Compared with traditional feature extraction algorithms, it can extract object features more accurately. Traditional algorithms may not perform well in extracting features of complex objects or under different lighting conditions. However, by optimizing the network structure and introducing the attention mechanism, this system enhances the learning ability of key features. Even in the face of complex scenes and diverse objects, it can effectively extract features and improve the accuracy of subsequent modeling.

[0021] 3. Compared with the existing AR-based model reconstruction systems, this AR glasses real-time object scanning and modeling system is innovative in the way of generating point clouds by combining image feature points and depth data, as well as its unique denoising and filtering preprocessing methods. Compared with some existing point cloud generation technologies, it uses a descriptor-based matching method to improve the accuracy of feature point matching, thereby enhancing the precision of point cloud generation. In terms of denoising and filtering, the combination of statistical filtering and voxel grid filtering can more effectively remove noise and reasonably reduce the data density. While retaining the main features of the point cloud, it reduces the data volume and improves the efficiency of subsequent processing.

[0022] 4. Compared with the existing AR-based model reconstruction systems, this AR glasses real-time object scanning and modeling system is innovative in using the Poisson reconstruction algorithm to construct a surface mesh model and combining the moving least squares method for model optimization. Compared with single 3D reconstruction algorithms, the Poisson reconstruction algorithm is based on the idea of implicit surface reconstruction and can effectively handle the surface reconstruction of objects with complex shapes, but there may be some local inaccuracies. The moving least squares method optimizes the model after Poisson reconstruction through local approximation, can adaptively adjust local parameters, make the model more conform to the real object surface, improve the accuracy and detail expression of the model, and generate a more complete and accurate 3D model.

[0023] Other advantages, objectives, and features of the present invention will be partially reflected by the following description and partially understood by those skilled in the art through the research and practice of the present invention. Brief Description of the Drawings

[0025] Figure 1 It is a system schematic diagram of an AR glasses real-time object scanning and modeling system of the present invention. Detailed Description of the Invention

[0027] The following will describe the present invention in detail with reference to the drawings, so that those of ordinary skill in the art can implement it after referring to this specification.

[0028] As Figure 1As shown in the figure, a real-time object scanning and modeling system for AR glasses, comprising: Hardware layer AR glasses main body: As the carrier of the entire system, it needs to have a lightweight and comfortable wearing design, and at the same time be able to stably run the system software and provide a clear display interface; High-precision camera: Select a camera with high resolution and large aperture, which can quickly capture the texture information of the object surface. For example, use a high-pixel camera of the Sony IMX series, with a resolution of up to 48 million pixels or more, to ensure clear capture of image details; Depth sensor: Integrate advanced depth sensors, such as structured light depth sensors or lidar depth sensors, to obtain the depth information of each point on the object surface. Taking the structured light depth sensor as an example, it projects a specific structured light pattern onto the object surface, uses the camera to capture the reflected pattern, and calculates the depth of each point on the object surface through the principle of triangulation, with an accuracy of up to millimeter level; Processor: Equip with a high-performance processor, such as the Snapdragon XR series chips, which have powerful computing capabilities, can quickly process the data collected by the camera and depth sensor, and at the same time run complex operations such as 3D reconstruction algorithms; Storage module: Adopt a large-capacity and high-speed storage chip to temporarily store the collected data and intermediate calculation results to ensure the continuity of data processing.

[0029] Data acquisition layer Image acquisition: The camera takes multi-angle pictures of the object at a high frame rate to obtain a two-dimensional image sequence of the object surface. Through intelligent control of the shooting angle and frame rate, ensure that the collected images can completely cover the object surface without obvious overlap or omission; Depth data acquisition: The depth sensor works synchronously to obtain the depth information of each point on the object surface in real time. Perform precise alignment in time and space with the image acquisition data to provide an accurate data basis for subsequent 3D reconstruction.

[0030] Algorithm layer Feature extraction algorithm: A feature extraction model based on convolutional neural network (CNN): Adopt the ResNet architecture. ResNet solves the problems of gradient disappearance and gradient explosion in the training process of deep neural networks by introducing residual blocks, enabling the network to be constructed deeper, so as to learn more complex features. In this system, the original ResNet is optimized, some redundant layers are reduced, and at the same time, an attention mechanism module (Squeeze-and-Excitation module) is added to automatically learn the importance of different channel features, recalibrate the feature map, and enhance the expression ability of key features; Collect a large amount of image data from different object categories, covering various shape, material, and color characteristics. Divide these image data into a training set, a validation set, and a test set according to a certain ratio. During the training process, use the stochastic gradient descent algorithm combined with a momentum optimizer to update the network parameters. The learning rate adopts a dynamic adjustment strategy, initially set to a relatively large value (such as 0.01), and as the training progresses, it decays by a certain ratio (such as decaying to 0.1 of the original value every 10 epochs) according to the performance on the validation set. At the same time, to prevent overfitting, add a Dropout layer to the network to randomly discard some neurons; Input the collected two-dimensional image of the object surface. The image first undergoes downsampling through a series of convolutional layers and pooling layers to gradually extract low-level features of the image, such as edges, textures, etc. Then, through multiple residual blocks and attention mechanism modules, the features are deeply learned and fused to obtain representative high-level features. Finally, the extracted features are output for subsequent matching and three-dimensional coordinate calculation.

[0031] Point cloud generation algorithm: Use descriptor-based matching methods, such as SIFT (Scale-Invariant Feature Transform) descriptors or ORB (Oriented FAST and Rotated BRIEF) descriptors. Taking the ORB descriptor as an example, for images taken from different angles, first extract the corner points in the image as feature points through the FAST corner detection algorithm. Then, calculate the ORB descriptor for each feature point. This descriptor is a binary string that describes the image features in the neighborhood of the feature point. Calculate the similarity between the feature point descriptors in different images through the Hamming distance to perform feature point matching and find the positions of the corresponding points on the same object surface in different images; According to the principle of triangulation, given the internal parameters (focal length, optical center position, etc.) and external parameters (rotation and translation matrices) of the camera, as well as the pixel coordinates of the matched feature points in different images, combined with the depth information of the corresponding points obtained by the depth sensor, the coordinates of the point in three-dimensional space can be calculated. The specific calculation process is as follows: Let the optical centers of two cameras be O_1 and O_2, and the projection points of a certain feature point on the image planes of the two cameras be p_1 and p_2 respectively. The distance from this point to the camera is obtained by the depth sensor as d. According to the principle of similar triangles and the internal and external parameter relationships of the camera, a system of equations can be listed to solve the coordinates (X, Y, Z) of the point in three-dimensional space; After calculating the three-dimensional coordinates of a large number of points on the object surface through the above method, point cloud data is generated. To improve the quality of the point cloud data, denoising and filtering processes are performed on it. The statistical filtering method is adopted to calculate the distance between each point and its neighboring points, and according to the statistical distribution of the distances, the points with abnormal distances, that is, the noise points, are removed. Then, the voxel grid filtering method is used to divide the point cloud data into three-dimensional voxel grids, and only one representative point is retained within each voxel, thereby reducing the density of the point cloud data, reducing the amount of data, and at the same time retaining the main features of the point cloud.

[0032] Three-dimensional reconstruction algorithm: Construct a surface mesh model based on the Poisson reconstruction algorithm: Construct a signed distance field: For each voxel, calculate its distance to the nearest point cloud point, and determine the positive or negative sign of the distance according to the position relationship between the voxel and the point cloud, and construct a signed distance field; Use the signed distance field as a boundary condition, and by iteratively solving the Poisson equation, obtain the scalar function values on the voxel grid; Apply the marching cubes algorithm, and according to the scalar function values, extract the zero isosurface in the voxel grid to generate a triangular mesh to form the surface model of the object; Smooth the generated surface mesh model. The Laplace smoothing algorithm is adopted to make the mesh surface smoother by adjusting the positions of the vertices. At the same time, check whether there are holes in the mesh model. For the detected holes, use the hole filling algorithm based on Delaunay triangulation to fill them to obtain a complete and accurate three-dimensional model.

[0033] Model optimization based on the moving least squares reconstruction algorithm: Construction of local approximation function: For each point cloud point, determine its neighborhood range, usually using the neighborhood search method based on radius. Within the neighborhood, construct a local approximation function based on polynomials, such as linear polynomials or quadratic polynomials. By minimizing the sum of the squares of the distances from the points in the neighborhood to the approximation function, determine the coefficients of the polynomial; According to the local approximation function, calculate the projection points of each point cloud point on the fitted surface to obtain a preliminary fitted surface. Then, iteratively update the fitted surface. Each time it is updated, recalculate the local approximation function and adjust the positions of the projection points to make the fitted surface closer to the real object surface; After completing the basic surface fitting, refine the model. By increasing the number of neighborhood points, adjusting the order of the polynomial, etc., improve the accuracy and detail expressiveness of the model. At the same time, smooth the model to remove the discontinuities caused by local approximation and make the model surface smoother and more natural.

[0034] Display layer: AR display interface: Render the constructed 3D model in real time onto the display interface of the AR glasses and integrate it with the real-world scene for display. By optimizing the rendering algorithm, improve the rendering efficiency and ensure the smooth and realistic display effect of the 3D model. At the same time, provide users with various interaction methods, such as gesture control, voice control, etc., to facilitate users to view, rotate, zoom, etc. the 3D model.

[0035] Embodiment 1 (Industrial Design and Manufacturing): The designer transmits the constructed 3D model in real time through the network to the AR glasses of team members located in different regions. The team members view the 3D model integrated with the real-world scene through the display layer of the AR glasses. Use interaction methods such as gesture control and voice control provided by the AR glasses to view, rotate, zoom, etc. the model; Team members in different regions can discuss the design optimization plan of the model in real time. For example, an engineer in region A discovers that the design of a certain connection part in the model may affect the assembly efficiency and communicates with the designer in region B through voice. The designer uses the annotation function of the AR glasses to mark the position that needs to be modified on the model in real time and makes a preliminary adjustment to the model through gesture operations. The adjusted model is displayed in real time on the AR glasses of other members, and everyone can continue to discuss and put forward further optimization suggestions; The design team makes multiple modifications and optimizations to the model in the virtual environment. After each modification, the system uses parallel computing technology to optimize and accelerate the algorithm through the multi-core performance of the processor, quickly generating a new 3D model to ensure the high efficiency and smoothness of the entire design optimization process. Finally, the optimal design plan for the engine cylinder block is determined.

[0036] Embodiment 2 (Education Field - Physics): The physics teacher sets up a virtual mechanics experiment scene in the AR glasses system, including various mechanics experiment devices, such as inclined planes, sliders, spring dynamometers, etc. Students wear AR glasses and enter this scene as if they were in a real physics laboratory; Students use the gesture control function of the AR glasses to conduct mechanics experiment operations. For example, in the experiment of "exploring the influencing factors of sliding friction", students place sliders of different masses on inclined planes with different roughness levels through gestures, use a virtual spring dynamometer to pull the sliders, and observe the changes in the readings of the spring dynamometer; The system simulates the mechanical process in real time, and uses physical formulas and algorithms to calculate the corresponding mechanical data based on the students' operating parameters, such as the mass of the slider, the angle of the inclined plane, the roughness of the contact surface, etc., and displays it to the students through AR glasses. Students can use voice commands to "show the force analysis diagram", and the system will display the force analysis diagram of the slider in the virtual scene to help students better understand the principles of mechanics.

[0037] Students can change the experimental conditions many times, repeat the experiment, observe the changes in the experimental results, and summarize which factors are related to sliding friction. Teachers can guide students to think and organize group discussions to deepen their understanding and mastery of physical knowledge.

[0038] Example 3 (Entertainment and Games): When players or artists want to create something creative, they first use AR glasses to scan real objects. Taking a toy car as an example, the scanning process is the same as the scanning process in game scene construction. The texture and depth information of the object are obtained through a high-precision camera and depth sensor. After hardware synchronization and calibration, the feature extraction model based on the convolutional neural network, the point cloud generation algorithm and the 3D reconstruction algorithm are used to generate an accurate 3D model of the toy car. In the display interface of AR glasses, players or artists use gesture control and voice commands to create and modify the generated 3D model. The creation platform provides a rich collection of creation tools and material libraries. Players can select various decorative elements from the material library, such as stickers, patterns, etc., and paste them onto the model through gestures. For example, a player selects a star-patterned sticker, adjusts its size and position through gestures, and pastes it on the body of a toy car. During the creation process, the system saves the player's operation records in real time, allowing the player to undo or restore the operation at any time. When the player completes the creation, the modified 3D model can be saved to the game's creative library for other players to use in the game.

[0039] Example 4 (Medical and Health Field): In a virtual environment, doctors can use the display layer of AR glasses to observe the three-dimensional model of the patient's body parts in a stereoscopic and intuitive way. Using gesture control, voice control and other interactive methods provided by AR glasses, doctors can rotate, scale, and slice the model to observe the spatial relationship between the lesion and the surrounding tissues and organs from different angles. For example, through the voice command "zoom in on the tumor site", the model will automatically zoom in to display the tumor area, allowing doctors to more clearly observe the shape, size and boundaries of the tumor; The surgical simulation software combines medical knowledge bases and clinical experience to simulate various surgical procedures on a 3D model. Doctors can simulate actions such as cutting and suturing with a scalpel through gestures. The system real-time simulates the impact of surgical operations on surrounding tissues, such as tissue deformation and bleeding conditions, and visually presents them to doctors through color changes, numerical displays, etc. Doctors can adjust the surgical plan based on the simulation results, select the best surgical path and operation method, and improve the success rate and safety of the surgery. For example, when simulating a liver surgery, doctors determine the optimal resection range to avoid important blood vessels and bile ducts through simulation operations, reducing surgical risks; Team members of doctors can use the remote collaboration function to view and discuss the surgical simulation process simultaneously through AR glasses. Experts from different regions can see the surgical model and the simulation operations of the surgeon in chief in their own AR glasses in real time, and put forward suggestions and opinions through voice communication and annotation functions to jointly improve the surgical plan.

[0040] Example 5 (Retail and Shopping Field): Before putting a product on the shelf for display, a merchant uses AR glasses to scan the product. Taking a smartwatch as an example, a high-precision camera and a depth sensor work together to obtain the texture and depth information of various parts of the watch, such as the appearance, dial, and strap. After hardware synchronization and calibration, the data enters the algorithm layer; The feature extraction model based on a convolutional neural network extracts the key features of the watch. The point cloud generation algorithm combines image feature points and depth data to generate point cloud data, and performs denoising and filtering processing. A surface mesh model is constructed based on the Poisson reconstruction algorithm, and then the model is optimized using the reconstruction algorithm based on moving least squares to obtain an accurate 3D model of the smartwatch. The merchant can add information such as product function introductions and usage instructions to the model to enrich the product display content; When a customer enters the electronic product display area of a shopping mall and wears AR glasses, they can see the 3D models of various products presented in a three-dimensional and vivid form in front of their eyes. The customer selects the product they want to know about through gesture control or voice commands, such as "Show the smartwatch". The system magnifies the 3D model of the smartwatch and displays it in the center of the customer's field of vision; The product function demonstration program preset by the merchant is launched. For example, for the smartwatch, the system demonstrates functions such as time display, health monitoring, and information reminder of the watch through animation demonstrations and voice explanations. The customer can simulate the actual operation process through gesture operations, such as clicking on the screen of the virtual watch, to view the specific implementation effects of the functions. During the demonstration process, the system uses parallel computing technology to ensure the smoothness and real-time nature of the demonstration through the multi-core performance of the processor, providing the customer with a clear and vivid product display experience and attracting the customer's attention.

[0041] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the examples shown and described herein.

Claims

1. A real-time object scanning and modeling system for AR glasses, wherein: include: The hardware layer includes an AR glasses body, and a high-precision camera, a depth sensor, a processor, and a storage module integrated on the AR glasses body; The high-precision camera is used to capture texture information on the surface of an object, the depth sensor is used to obtain depth information of each point on the surface of the object, the processor is used to process the collected data and run a three-dimensional reconstruction algorithm, and the storage module is used to temporarily store the collected data and intermediate calculation results; The data acquisition layer uses the high-precision camera to shoot the object at multiple angles at a high frame rate to obtain a two-dimensional image sequence, while the depth sensor synchronously obtains the depth information of each point on the surface of the object, and accurately aligns the data of the two in time and space; The algorithm layer includes a feature extraction algorithm, a point cloud generation algorithm and a 3D reconstruction algorithm; the feature extraction algorithm uses a model based on a convolutional neural network to extract features from the collected image data; the point cloud generation algorithm combines image feature points and depth data to calculate the coordinates of each point on the object surface in 3D space and generate point cloud data; the 3D reconstruction algorithm constructs a surface mesh model of the object based on the point cloud data and performs optimization; The display layer renders the constructed 3D model in real time to the display interface of AR glasses, integrates it with the real scene, and provides interactive methods such as gesture control and voice control.

2. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The convolutional neural network-based feature extraction model adopts an improved ResNet architecture, which reduces some redundant layers of the original ResNet and adds an attention mechanism module to automatically learn the importance of different channel features and recalibrate the feature map.

3. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The training process of the feature extraction model based on convolutional neural network includes: collecting a large amount of object image data covering different shapes, materials, and color features, and dividing them into training sets, validation sets, and test sets; using a stochastic gradient descent algorithm combined with a momentum optimizer to update network parameters, using a dynamic adjustment strategy for the learning rate, and adding a Dropout layer to the network to prevent overfitting.

4. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The feature point matching in the point cloud generation algorithm utilizes a descriptor-based matching method to extract feature points from images taken at different angles through a corner detection algorithm, calculate a descriptor for each feature point, and perform feature point matching by calculating the similarity between descriptors.

5. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The three-dimensional coordinate calculation in the point cloud generation algorithm is based on the principle of triangulation. It combines the intrinsic and extrinsic parameters of the camera, the pixel coordinates of the matched feature points in different images, and the depth information of the corresponding points obtained by the depth sensor, and solves the coordinates of the surface points of the object in three-dimensional space through a set of equations.

6. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The point cloud generation and preprocessing in the point cloud generation algorithm include: generating point cloud data after calculating the three-dimensional coordinates of the surface points of the object, removing noise points by using a statistical filtering method, and reducing the density of the point cloud data by using a voxel grid filtering method while retaining the main features.

7. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The 3D reconstruction algorithm constructs a surface mesh model based on the Poisson reconstruction algorithm, regards the point cloud data as the sampling points of the scalar field, constructs a 3D voxel mesh containing the point cloud data, calculates the voxel scalar value, solves the Poisson equation using a discrete Poisson equation solver, extracts the zero isosurface using a marching cube algorithm to generate a surface mesh model, and smoothes the model and fills holes.

8. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The model optimization based on the moving least squares reconstruction algorithm in the three-dimensional reconstruction algorithm constructs a local polynomial approximation function in the neighborhood of each point cloud point, minimizes the sum of squares of distances from points in the neighborhood to the approximation function to determine the polynomial coefficients, iteratively updates the fitting surface, and refines and smoothes the model by increasing the number of neighborhood points, adjusting the order of the polynomial, etc.

9. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: It also includes a hardware synchronization and calibration module, which is used to achieve time synchronization accuracy of the high-precision camera and depth sensor during data acquisition to the microsecond level, and to accurately calibrate the two to establish a spatial mapping relationship.

10. The AR glasses real-time object scanning and modeling system according to claim 1, wherein: The algorithms in the algorithm layer are optimized and accelerated by using parallel computing technology and the multi-core performance of the processor to improve the real-time performance of the system.

Citation Information

Patent Citations

  • AR-based model reconstruction system

    CN210109870U

Cited By

  • Vacuum cup outer surface dynamic display system based on AR technology

    CN120876794A