Real-time pattern recognition and digital interaction system based on AR (Augmented Reality) glasses

By realizing a real-time graphics recognition and digital interaction system on AR glasses, the shortcomings of existing AR systems in image recognition accuracy, real-time and system independence are solved, and the response time and user experience effect are improved.

CN120066263AInactive Publication Date: 2025-05-30HUNAN NORMAL UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510128012.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing AR systems have shortcomings in the accuracy, real-timeness and system independence of image recognition, resulting in long response time and poor user experience.

Method used

A real-time graphics recognition and digital interaction system based on AR glasses is proposed, and efficient recognition and interaction of real-time graphics can be achieved through clarity evaluation, image processing and three-dimensional modeling.

Benefits of technology

It improves the response time and precision of user interaction of AR glasses, enhances the user experience effect, and makes full use of the application potential of AR glasses in occasions such as education and art display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066263A_ABST
    Figure CN120066263A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time graph recognition and digital interaction system based on AR glasses, and relates to an augmented reality (AR) technology and image processing. According to real-time graphs acquired by the AR glasses, each image is analyzed, and the clearness coefficient of the corresponding image is calculated; comparing the definition coefficient of each image with a definition coefficient threshold value of a preset image, and judging the image with the highest definition qualification degree; processing the image with the highest definition qualification degree to obtain a target image; segmenting the target image into different areas, allocating independent labels and depth information to each area, performing three-dimensional modeling on the target image according to the independent labels and depth information allocated to each area, and interacting with a user according to the three-dimensional modeling; in this way, it is ensured that the response time of the AR glasses is short, user interaction is closer, and the user experience effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of augmented reality technology and image processing technology, and particularly to a real-time graphic recognition and digital interaction system based on AR glasses. Background Art

[0002] AR technology has been widely applied in multiple fields such as education, art, and advertising. However, most current systems still face some technical limitations, especially in terms of the accuracy, real-time performance, and independence of image recognition. Taking traditional AR systems as an example, they mainly rely on external devices for image processing and data analysis, which not only prolongs the response time but also limits the mobility of the system and the fluency of the user experience. In recent years, continuously iterated AR glasses, such as Google Glass or Mijia Glass Camera, mainly focus on basic information overlay and simple user interaction in function research and development, without fully utilizing image recognition technology to achieve complex visual conversion, thus limiting their application potential in occasions that require high visual interaction such as education and art exhibitions.

[0003] Therefore, the main deficiencies of the existing technology lie in its relatively long response time and often simple interaction with users, resulting in a poor user experience. Summary of the Invention

[0004] The object of the present invention is to solve the above-mentioned problems and provide a real-time graphic recognition and digital interaction system based on AR glasses.

[0005] The present invention provides a real-time graphic recognition and digital interaction system based on AR glasses, which includes:

[0006] Clarity module: Analyze the real-time graphics obtained by the AR glasses and calculate the clarity coefficient of each corresponding image.

[0007] Judgment module: Compare the clarity coefficient of each image with the preset clarity coefficient threshold of the image to determine the image with the highest qualified degree of clarity.

[0008] Processing module: Process the image with the highest qualified degree of clarity to obtain a target image.

[0009] Modeling interaction module: Segment the target image into different regions, assign independent labels and depth information to each region, perform three-dimensional modeling of the target image based on the independent labels and depth information assigned to each region, and interact with the user according to the three-dimensional modeling.

[0010] Optionally, analyzing the real-time graphics obtained by the AR glasses and calculating the clarity coefficient of each corresponding image includes:

[0011] Convert each image into a grayscale image, perform a convolution operation on the grayscale image using the Laplacian operator to obtain a Laplacian image;

[0012] Calculate the variance σ of all pixel points in the Laplacian image 2 L, and the calculation formula is:

[0013]

[0014] In the formula, L(x, y) is the pixel value of the Laplacian image at the pixel position (x, y), M is the height of the Laplacian image, N is the width of the Laplacian image, and μL is the average value of the pixel values of the Laplacian image. The calculation formula is:

[0015] Take the variance σ of the image 2 L as the clarity coefficient of the corresponding image.

[0016] Optionally, compare the clarity coefficient of each image with the clarity coefficient threshold of the preset image to determine the image with the highest clarity qualification degree, including:

[0017] Compare the clarity coefficient of each image with the clarity coefficient threshold of the preset image. If the clarity coefficient is greater than the clarity coefficient threshold of the preset image, mark the corresponding image as the image with the qualified clarity degree, and mark the image with the largest clarity coefficient as the image with the highest clarity qualification degree.

[0018] Optionally, process the image with the highest clarity qualification degree to obtain the target image, including:

[0019] Denoise the image with the highest clarity qualification degree through Gaussian filtering and median filtering;

[0020] Use perspective correction technology to adjust the image to a front view;

[0021] Use techniques such as histogram equalization or gamma correction to enhance the contrast of the image;

[0022] Identify the main structures and contours in the image through edge detection;

[0023] Use algorithms such as sharpening filtering or Laplacian enhancement to enhance the details of the image.

[0024] Optionally, segment the target image into different regions and assign independent labels and depth information to each region, including:

[0025] Classify and segment different parts in the target image through a semantic segmentation algorithm, and each region will be marked with a different label;

[0026] For each region, use a preset depth estimation model based on a convolutional neural network to predict the depth value of each pixel, and obtain the depth information of each region.

[0027] Optionally, performing three-dimensional modeling of the target image according to the independently assigned labels and depth information for each region includes:

[0028] Convert the two-dimensional image corresponding to each region in the target image into a three-dimensional point cloud through stereovision or structured light technology;

[0029] Use a mesh reconstruction algorithm to convert the three-dimensional point cloud into a polygon mesh to obtain the three-dimensional point cloud of the two-dimensional image corresponding to each region finally;

[0030] Combine the three-dimensional point clouds of the two-dimensional images corresponding to each region finally to construct a complete three-dimensional model of the target image.

[0031] Optionally, it is characterized in that the interaction with the user according to the three-dimensional model includes:

[0032] Embed a gesture control sensor and a gaze tracking sensor into the terminal of the AR glasses;

[0033] Use the gesture sensor to capture the user's gestures and convert them into control signals.

[0034] Use the gaze tracking sensor to track the user's eye movement information in real time and obtain the region stared at by the user;

[0035] According to the gesture data, adjust the display parameters of the three-dimensional model in real time;

[0036] According to the gaze tracking data, dynamically adjust the focus position and detail enhancement of the three-dimensional model.

[0037] Advantages of the present invention:

[0038] The present invention proposes a real-time graphic recognition and digital interaction system based on AR glasses. According to the real-time graphics obtained by the AR glasses, analyze each image to calculate the clarity coefficient of the corresponding image; compare the clarity coefficient of each image with the clarity coefficient threshold of the preset image to determine the image with the highest qualified degree of clarity; process the image with the highest qualified degree of clarity to obtain the target image; segment the target image into different regions, and assign independent labels and depth information to each region, and perform three-dimensional modeling of the target image according to the independently assigned labels and depth information for each region, and interact with the user according to the three-dimensional model; in this way, ensure that the response time of the AR glasses is relatively short, the user interaction is more precise, and the user experience effect is increased. Description of the Drawings

[0039] The present invention will be further described below with reference to the accompanying drawings.

[0040] Figure 1 It is a framework diagram of a real-time graphic recognition and digital interaction system based on AR glasses. Specific implementation manners

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] The embodiments of the present invention provide a real-time graphic recognition and digital interaction system based on AR glasses. Refer to Figure 1 , Figure 1 which is a framework diagram of the real-time graphic recognition and digital interaction system based on AR glasses provided by the embodiments of the present invention. The system includes:

[0044] Clarity module: Analyze each image according to the real-time graphics obtained by the AR glasses and calculate the clarity coefficient of the corresponding image;

[0045] Judgment module: Compare the clarity coefficient of each image with the clarity coefficient threshold of the preset image, and judge the image with the highest qualified clarity degree;

[0046] Processing module: Process the image with the highest qualified clarity degree to obtain a target image;

[0047] Modeling and interaction module: Segment the target image into different regions, assign independent labels and depth information to each region, perform three-dimensional modeling of the target image according to the independent labels and depth information assigned to each region, and interact with the user according to the three-dimensional modeling.

[0048] Based on the real-time graphic recognition and digital interaction system based on AR glasses provided by the embodiments of the present invention, according to the real-time graphics obtained by the AR glasses, analyze each image to calculate the clarity coefficient of the corresponding image; compare the clarity coefficient of each image with the clarity coefficient threshold of the preset image to determine the image with the highest qualified degree of clarity; process the image with the highest qualified degree of clarity to obtain the target image; segment the target image into different regions, assign independent labels and depth information to each region, and perform three-dimensional modeling of the target image according to the independent labels and depth information assigned to each region, and interact with the user according to the three-dimensional modeling; in this way, ensure that the response time of the AR glasses is short, the user interaction is closer, and the user experience effect is increased.

[0049] In one embodiment, analyzing each image to calculate the clarity coefficient of the corresponding image according to the real-time graphics obtained by the AR glasses includes:

[0050] Convert each image into a grayscale image, perform convolution operation on the grayscale image using the Laplace operator to obtain a Laplace image;

[0051] Calculate the variance σ 2 L of all pixel points in the Laplace image, and the calculation formula is:

[0052]

[0053] In the formula, L(x,y) is the pixel value of the Laplace image at the pixel position (x, y), M is the height of the Laplace image, N is the width of the Laplace image, and μL is the average value of the pixel values of the Laplace image. The calculation formula is:

[0054] Take the variance σ 2 L of the image as the clarity coefficient of the corresponding image.

[0055] It should be noted that the Laplace operator is a second-order differential operator used to detect edges and details in images. Its basic idea is to perform a convolution operation on the image to emphasize the regions of rapid change in the image, that is, the high-frequency parts, which usually correspond to the edge and detail information of the image. In the field of image processing, the Laplace operator is widely used in tasks such as edge detection, texture analysis, and clarity evaluation; the variance calculated by the Laplace operator is used as the clarity coefficient of the image. Its advantage is that it can effectively capture the fine changes in the image, especially at the edges. Blurred or out-of-focus regions usually exhibit low variance, while clear images have a larger variance. This method can not only reflect the overall clarity of the image but also eliminate the interference of factors such as low contrast and noise, thereby improving the reliability of image quality evaluation. Therefore, using the Laplace operator to calculate the clarity coefficient can determine the clarity of the real-time graphics obtained by the AR glasses;

[0056] In one embodiment, comparing the clarity coefficient of each image with the clarity coefficient threshold of the preset image, the images with the highest qualified degree of clarity are determined as follows:

[0057] Compare the clarity coefficient of each image with the clarity coefficient threshold of the preset image. If the clarity coefficient is greater than the clarity coefficient threshold of the preset image, mark the corresponding image as an image with a qualified degree of clarity, and mark the image with the largest clarity coefficient as the image with the highest qualified degree of clarity.

[0058] It should be noted that the clarity coefficient threshold of the preset image is set by professionals according to the actual situation and will not be specifically limited and elaborated here.

[0059] It should be noted that by analyzing the clarity coefficient of the real-time images obtained by the AR glasses, the accuracy, real-time performance, and system independence of graphic recognition can be effectively improved. First, the calculation of the clarity coefficient can quantify the quality of each image, ensuring that only images meeting the clarity requirements are used for further processing. This screening mechanism reduces misrecognition caused by blurred or unclear images and improves the accuracy of image recognition. Second, by comparing the clarity coefficient of each image with the preset threshold, the system can quickly determine the most suitable image for modeling or recognition, which avoids the processing of meaningless images, thereby shortening the processing time and improving the response speed and real-time performance of the system. In addition, the clarity evaluation and processing are performed on the device side, which can reduce the dependence on external servers or complex cloud computing, thereby enhancing the system independence and enabling the AR glasses to operate more flexibly in various scenarios in real time. This process avoids the limitations of traditional AR systems that rely on external hardware for data processing, thereby reducing the overall response time and providing a smoother and more immediate user experience.

[0060] It should be noted that in order to accelerate the recognition speed of real-time graphics on the AR glasses, different intelligent devices can also be added to and embedded in the AR glasses. For example, the AR glasses can integrate advanced processors, dedicated image processing accelerators, or deep learning chips. These hardware components can be specifically used for rapid image processing and analysis, reducing the burden on computing resources, thereby improving the speed of graphics recognition. In addition, it can work in cooperation with external sensors (such as lidar, infrared sensors, etc.) to enhance the depth perception and environmental adaptability of image recognition, especially in complex real-time scenarios, such as dynamic environments or low-light conditions. Combining these devices, the AR glasses can not only improve the clarity and depth perception of images, but also further optimize the image recognition algorithm through data fusion of sensors, accurately capture the features of target objects, and thus improve the real-time performance and accuracy of the system. At the same time, the efficient cooperation of intelligent devices can allocate some computing tasks to dedicated hardware, reducing the pressure on the main processor and improving the response speed and stability of the overall system. This way of multi-device cooperation makes the AR glasses more efficient, flexible, and system-independent in real-time graphics recognition.

[0061] In one embodiment, processing the image with the highest qualified degree of clarity to obtain the target image includes:

[0062] Denoise the image with the highest qualified degree of clarity through Gaussian filtering and median filtering;

[0063] Use perspective correction technology to adjust the image to a front view;

[0064] Use techniques such as histogram equalization or gamma correction to enhance the contrast of the image;

[0065] Identify the main structures and contours in the image through edge detection;

[0066] Use algorithms such as sharpening filtering or Laplacian enhancement to enhance the details of the image.

[0067] It should be noted that the denoising process: Function: The purpose of denoising is to remove noise in the image caused by environmental interference, sensor problems, or poor lighting conditions. This can make subsequent image analysis more accurate and reduce misrecognition.

[0068] Steps: Adopt common denoising algorithms, such as Gaussian filtering, median filtering, etc.; For example: For the "Taiji Diagram" image, random noise in the background can be removed through denoising filtering to ensure that only the main part of the pattern is retained, avoiding affecting edge detection and subsequent feature extraction.

[0069] Image Correction: Function: The goal of image correction is to correct geometric distortions in the image to ensure spatial accuracy, especially when capturing planar graphs with complex geometries. For example, for images with regular patterns like the "Taiji Diagram", correction can ensure that circular patterns do not distort due to aberration.

[0070] Steps: First, eliminate camera distortion through camera calibration and use perspective correction technology to adjust the image to a front view, so that the central part of the "Taiji Diagram" in the image is not affected by angles or tilts, ensuring accurate capture of its geometric features.

[0071] Contrast Adjustment: Function: By enhancing contrast, key features in the image become more prominent, especially in low-light or uneven brightness situations, which helps improve image recognition.

[0072] Steps: Use techniques such as histogram equalization or gamma correction to enhance the contrast of the image. For the "Taiji Diagram", by increasing the contrast between black and white patterns, the shapes of the Yin and Yang fish in the Taiji can be made clearer, facilitating subsequent edge detection and feature extraction.

[0073] Edge Detection: Function: Edge detection is used to identify the main structures and contours in the image, which is crucial for complex boundaries and symmetries in the graph. By extracting edge information, it can assist in subsequent feature matching and 3D reconstruction.

[0074] Steps: Use edge detection algorithms such as Sobel and Canny to extract the edges of the boundary between the Yin and Yang fish, the outer circle, and the internal black and white blocks in the "Taiji Diagram". This step can effectively extract the structural information of the pattern, providing key data for subsequent feature matching and 3D modeling.

[0075] Image Enhancement: Function: Image enhancement makes object details more visible by improving details and contrast. This helps improve image quality, especially when clear object surface information is required for 3D reconstruction.

[0076] Steps: Use algorithms such as sharpening filters or Laplacian enhancement to enhance the details of the image. For example, in the "Taiji Diagram", by enhancing the details of the pattern, the shapes of the Yin and Yang fish and their internal symmetry structures can be clearly distinguished, providing more accurate information for subsequent modeling.

[0077] In one implementation, through these processing steps, the obtained image is not only visually clearer, but also can provide more reliable data in subsequent 3D modeling and recognition. For the "Taiji Diagram", these steps can ensure that every detail in the image is effectively captured, so that when it is transformed into a three-dimensional model that is three-dimensional, dynamic, and interactive, it can accurately reflect the details and spatial structure of the pattern, enabling users to interact with it in the AR environment and enhancing the sense of immersion and experience.

[0078] In one embodiment, the target image is segmented into different regions, and independent labels and depth information are assigned to each region, including:

[0079] The different parts in the target image are classified and segmented through a semantic segmentation algorithm, and each region will be marked with a different label;

[0080] For each region, a preset depth estimation model based on a convolutional neural network is used to predict the depth value of each pixel to obtain the depth information of each region.

[0081] It should be noted that the above steps are described. For example, for the "Taiji Diagram", the purpose of image segmentation is to decompose a complex image into multiple regions to help visually identify different parts in the pattern, so as to perform depth estimation and 3D modeling on each part separately. For the "Taiji Diagram", the main regions in the image include the black and white parts of the Yin and Yang fish, the circular border, and the central part.

[0082] Semantic segmentation or instance segmentation: Use a semantic segmentation algorithm (such as U-Net, DeepLabV3) or an instance segmentation algorithm (such as Mask R-CNN) to perform image segmentation on the "Taiji Diagram". Through these methods, different regions in the "Taiji Diagram" can be identified:

[0083] The black region (representing a part of the Yin and Yang fish); the white region (representing the other part of the Yin and Yang fish); the central part (symbolizing the intersection of Yin and Yang); the outer ring, etc.;

[0084] Each region will be assigned an independent label for subsequent processing.

[0085] Details of the segmentation algorithm: Use a preset depth estimation model based on a convolutional neural network (CNN) to train the image, learn to identify different regions in the "Taiji Diagram", and mark the boundaries of each region; segment through the boundaries of each region to ensure that each part can be processed separately and generate a different label for each region.

[0086] Depth Information Estimation: The goal of depth information estimation is to assign corresponding depth values to each segmented region so that the planar graph can be transformed into a 3D model. Specifically, for each region, a depth estimation model based on a convolutional neural network (CNN) (such as Monodepth or MiDaS) is used to predict the depth value of each pixel. For example, the black region of the Yin-Yang fish may be inferred as a relatively foreground object, while the depth value of the outer ring region may be estimated as a more distant object. The depth value can be estimated based on the grayscale, texture, and geometric structure of the image.

[0087] Through the trained depth network, the system can generate a depth map, and the depth information of each segmented region will be embedded into the corresponding region label to ensure the accurate matching of the depth information of each region.

[0088] After obtaining the depth information and labels of each region, this information can be used for 3D modeling, that is, to reconstruct the 3D form of the Planar Diagram of the Taiji Diagram according to these depths and region labels.

[0089] In one embodiment, the 3D modeling of the target image according to the independent labels and depth information assigned to each region includes:

[0090] Through stereovision or structured light technology, convert the 2D image corresponding to each region in the target image into a 3D point cloud;

[0091] Use a mesh reconstruction algorithm to convert the 3D point cloud into a polygon mesh to obtain the final 3D point cloud corresponding to the 2D image of each region;

[0092] Finally, combine the 3D point clouds corresponding to the 2D images of each region to construct the complete 3D modeling of the target image.

[0093] It should be noted that an example of the above steps is as follows: First, use stereovision or structured light technology to obtain the depth information of each region of the Taiji Diagram.

[0094] Stereovision: Stereovision technology is based on using two or more cameras to capture the same scene from different angles. By matching the feature points in two or more images, calculating the disparity between these points, and then inferring the depth information.

[0095] In the 3D modeling of the Taiji Diagram, use stereovision to capture each region in the image (such as the black and white parts of the Yin-Yang fish, the ring, etc.) from different perspectives. The depth information of each region can be obtained by matching the feature points (such as edges, corners, etc.) in the images. The calculated disparity map estimates the 3D coordinates of each pixel point, and finally generates a point cloud.

[0096] Structured light: Structured light technology projects a known light pattern (such as stripes or grids) onto the planar graph of the "Taiji Diagram". The camera captures the deformed light pattern and, combined with the known projected pattern, calculates the depth information of each area.

[0097] For the circular pattern of the "Taiji Diagram", a stripe pattern can be projected by structured light, and the reflection deformation of the pattern in areas such as the black and white parts of the Yin and Yang fish, and the circular ring, etc. can be captured. Structured light can provide more accurate surface depth information, especially for these fine pattern parts, and it can help determine the three-dimensional structure and relative positions of different areas.

[0098] Through stereovision or structured light technology, three-dimensional point cloud data has been obtained for each area of the "Taiji Diagram" (such as the black and white parts of the Yin and Yang fish, the circular ring, and the center). The depth information of each area is combined with the two-dimensional image, thus laying a foundation for subsequent three-dimensional modeling.

[0099] The mesh reconstruction algorithm converts the three-dimensional point cloud into a polygon mesh:

[0100] After obtaining the three-dimensional point cloud, the next task is to convert this point cloud data into a tangible three-dimensional model through a mesh reconstruction algorithm. In this process, the point cloud data is converted into a polygon mesh, and then a complete three-dimensional structure is formed;

[0101] Point cloud to mesh conversion: Poisson reconstruction algorithm: This is a common point cloud reconstruction algorithm, especially suitable for generating a smooth three-dimensional surface from a dense point cloud. Based on the normal information of the points in the point cloud, by calculating the curvature and surface normal of the point cloud, the point cloud data is "wrapped" into a smooth polygon mesh.

[0102] Delaunay triangulation: This algorithm reconstructs the surface by dividing the point cloud into a series of triangular meshes, and can effectively maintain the geometric shape of the point cloud. For the "Taiji Diagram", Delaunay triangulation can help construct a detailed three-dimensional mesh structure from different areas of the Yin and Yang fish and the circular ring part of the point cloud.

[0103] During the process of the "Taiji Diagram" model, the algorithm generates an independent polygon mesh for each segmented area (such as the black and white parts of the Yin and Yang fish, the circular ring, etc.). The mesh of each area will accurately represent its three-dimensional shape and structure, retaining the details in the image (such as the shading effects of the black and white parts).

[0104] Combine the three-dimensional point clouds of each area to construct a complete three-dimensional modeling:

[0105] Once the three-dimensional point clouds and meshes of all areas are constructed, the next step is to combine the three-dimensional data of these areas into a complete three-dimensional model of the "Taiji Diagram". This process includes the following steps:

[0106] Region combination: In this step, the independent 3D point clouds of each region are merged. Each region will have its own depth data (such as different depth information for black and white regions, rings, and the central part), and these data will be combined according to the labels of the regions in the original image. The depth information between regions will help construct a three-dimensional and hierarchical model.

[0107] For the "Taiji Diagram", this means that the yin and yang fish in the black and white parts will show different levels in three-dimensional space, and the depth value of the outer ring will be farther than that of the central part.

[0108] Multi-region synthesis: Once the 3D meshes of all regions are synthesized, they will be accurately docked through a point cloud alignment algorithm to ensure smooth transitions between each region and conform to the physical form. For example, the black and white parts of the yin and yang fish should have a smooth transition, and the construction of the outer ring should maintain a certain proportional relationship with the inner ring.

[0109] By optimizing the point cloud data of each region, a unified 3D mesh model is finally formed, making the overall shape of the "Taiji Diagram" look three-dimensional and realistic from all angles.

[0110] In one embodiment, the interaction with the 3D modeling and the user includes:

[0111] Install gesture control sensors and eye tracking sensors in the terminal of the AR glasses; ensure that the gesture actions and eye movement information of the user can be captured.

[0112] Use the gesture sensor to capture the user's gestures (such as rotation, scaling, panning) and convert them into control signals.

[0113] Use the eye tracking sensor to track the user's eye movement information in real time and obtain the area where the user is gazing.

[0114] According to the gesture data, adjust the display parameters such as rotation, scaling, and panning of the 3D model in real time.

[0115] According to the eye tracking data, dynamically adjust the focus position and detail enhancement of the 3D model to ensure that the area of interest to the user is clearly displayed.

[0116] The system immediately feeds back the user's interaction operations and shows the changes in the 3D model to ensure the real-time nature of the operations.

[0117] Provide personalized adjustment options to dynamically optimize the interaction experience according to the user's operation habits or preferences.

[0118] It should be noted that in this process, the GPU is utilized to accelerate the real-time rendering of 3D models to ensure the smoothness of the interaction process; the rendering details are automatically adjusted through optimized algorithms to guarantee high display performance; and the interaction mode is adjusted according to the user's historical operations and preferences to make the experience more personalized and meet the requirements.

[0119] In one implementation manner, the benefits brought by the present invention are as follows:

[0120] (1) Precise image recognition ability: The image recognition algorithm adopted by the present invention is optimized on the basis of existing deep learning technologies and is particularly suitable for the recognition of complex graphics and artworks. The algorithm can highly accurately capture the details and structural information in two-dimensional images to ensure the accuracy of the conversion process and the reliability of the effect.

[0121] (2) Two-dimensional to three-dimensional conversion mechanism: It enhances the visual effect and at the same time provides a more rich interactive experience, effectively generating three-dimensional dynamic views and presenting a three-dimensional and dynamic visual feeling to the user.

[0122] (3) Enhanced real-time performance and interactivity: This system implements instant image processing and presentation on AR glasses, reducing the latency of traditional processing technologies; users can adjust the display parameters of three-dimensional images in real time through intuitive interactive operations such as gesture control or gaze tracking, improving the interactivity and personalized experience.

[0123] (4) Application prospects in multiple fields: This technology is not only applicable to the enhanced display of artworks, but can also be widely applied in multiple fields such as education and design. By adding three-dimensional dynamic displays to two-dimensional images, the present invention can significantly improve the efficiency and attractiveness of information transmission, opening up new application scenarios and market opportunities.

[0124] (5) Deep customization of user experience: The present invention supports users to customize various parameters of three-dimensional displays according to their personal preferences and needs, such as the speed of dynamic effects, viewing angles, etc., so that each user can obtain the viewing experience that best meets their expectations.

[0125] The above has described a detailed description of an embodiment of the present invention, but the content described is only a preferred embodiment of the present invention and cannot be artificially used to limit the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the patent coverage scope of the present invention.

Claims

1. A real-time graphic recognition and digital interaction system based on AR glasses, characterized by: The system comprises: Clarity module: Analyzes and calculates the clarity coefficient of each image based on the real-time graphics obtained by AR glasses; Judgment module: compares the clarity coefficient of each image with the preset image clarity coefficient threshold, and determines the image with the highest clarity qualification; Processing module: Process the image with the highest definition to obtain the target image; Modeling and interaction module: Segment the target image into different regions, assign independent labels and depth information to each region, perform three-dimensional modeling of the target image based on the independent labels and depth information assigned to each region, and interact with the user based on the three-dimensional modeling.

2. The real-time graphic recognition and digital interaction system based on AR glasses according to claim 1 is characterized in that: According to the real-time graphics obtained by AR glasses, the clarity coefficients of the corresponding images are analyzed and calculated, including: Convert each image into a grayscale image, and perform convolution operation on the grayscale image using the Laplacian operator to obtain a Laplacian image; Calculate the variance σ of all pixels in the Laplacian image 2 L, the calculation formula is: Where L(x, y) is the pixel value at the pixel position (x, y) of the Laplace image, M is the height of the Laplace image, N is the width of the Laplace image, μL is the average value of the pixel value of the Laplace image, and the calculation formula is: The variance σ of the image 2 L is the clarity coefficient of the corresponding image.

3. The real-time graphic recognition and digital interaction system based on AR glasses according to claim 1 is characterized in that: The clarity factor of each image is compared with the preset image clarity factor threshold, and the images with the highest clarity qualification are judged to include: The clarity coefficient of each image is compared with the preset clarity coefficient threshold of the image. If the clarity coefficient is greater than the preset clarity coefficient threshold of the image, the corresponding image is recorded as an image with a qualified clarity degree, and the image with the largest clarity coefficient is recorded as the image with the highest qualified clarity degree.

4. The real-time graphic recognition and digital interaction system based on AR glasses according to claim 1, characterized in that: The target image obtained by processing the image with the highest definition includes: The image with the highest definition is denoised by Gaussian filtering and median filtering; Use perspective correction technology to adjust the image to a front view; Use techniques such as histogram equalization or gamma correction to enhance the contrast of the image; Identify key structures and contours in an image through edge detection; Enhance the details of the image using algorithms such as sharpening filtering or Laplacian enhancement.

5. The real-time graphic recognition and digital interaction system based on AR glasses according to claim 1, characterized in that: Segment the target image into different regions and assign independent labels and depth information to each region, including: The different parts of the target image are classified and segmented through the semantic segmentation algorithm, and each area will be marked with a different label; For each area, a preset convolutional neural network-based depth estimation model is used to predict the depth value of each pixel to obtain the depth information of each area.

6. The real-time graphic recognition and digital interaction system based on AR glasses according to claim 1, characterized in that: The three-dimensional modeling of the target image is performed by assigning independent labels and depth information to each region, including: Through stereo vision or structured light technology, the two-dimensional image corresponding to each area in the target image is converted into a three-dimensional point cloud; The 3D point cloud is converted into a polygonal mesh using a mesh reconstruction algorithm to obtain a 3D point cloud of the 2D image corresponding to each region; Finally, the three-dimensional point cloud of the two-dimensional image corresponding to each area is combined to construct a complete three-dimensional modeling of the target image.

7. The real-time graphic recognition and digital interaction system based on AR glasses according to claim 1, characterized in that: Interaction with users based on 3D modeling includes: Embed gesture control sensors and gaze tracking sensors into the terminals of AR glasses; Use gesture sensors to capture user gestures and convert them into control signals. Use the gaze tracking sensor to track the user's eye movement information in real time and obtain the area where the user is gazing; According to the gesture data, the display parameters of the 3D model are adjusted in real time; Dynamically adjust the focus position and detail enhancement of the 3D model based on the gaze tracking data.

Citation Information

Patent Citations

  • Exhibition system capable of adjusting three-dimensional models according to sight lines of visitors and method implemented by exhibition system

    CN104076915A

  • Image recognition system based on AR intelligent glasses

    CN112866678A

  • Image recognition system based on AR intelligent glasses

    CN113794872A

  • Man-machine interaction method based on museum guide AR glasses

    CN114647315A

  • Power equipment blurred image recognition method and system based on Laplacian operator

    CN115731185A