Four-camera AI intelligent glasses

By utilizing the built-in chip and deep learning network of the quad-camera AI smart glasses, the problem of image fluctuations caused by camera tilt is solved, enabling adaptive image aggregation and safety alerts, thus improving the stability and security of the smart glasses.

CN121596566APending Publication Date: 2026-03-03HANGZHOU RUISHI FUTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511775407.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing AI smart glasses suffer from camera misalignment due to the reliance on the temples for support, resulting in fluctuating image information aggregation and an inability to effectively track changes in camera angle.

Method used

The AI ​​smart glasses feature a quad-camera setup. They use a built-in chip and deep learning network to perform image time alignment and angle correction, aggregate features using an attention mechanism, configure an adaptive decoder, monitor the bending angle and fatigue limit of the temples, and achieve adaptive updates.

Benefits of technology

It enables adaptive shooting angle changes when worn by different people, and configures a decoder for different angle images for each user, reducing fluctuations in image information aggregation and providing security prompts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121596566A_ABST
    Figure CN121596566A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent equipment, and particularly relates to four-camera AI intelligent glasses which comprise a glasses frame, glasses legs, and a first camera and a second camera which are used for shooting original images, and further comprise built-in chips installed in the glasses legs, and the built-in chips comprise data input interfaces which are used for connecting the first camera and the second camera; the encoder extracts a first feature in the original image through a CNN (Convolutional Neural Network); the aggregation module is used for aggregating the first features through an attention mechanism and outputting global features; the first decoder is used for aggregating into a panoramic image according to the global features in combination with the spatial tensor; and the angle correction module is used for monitoring the bending angle of the glasses leg. According to the invention, after a plurality of angle images are acquired by different cameras, time alignment is carried out to synthesize the panoramic image, different decoders are learned to be configured for the different angle images to realize adaptive updating, and an alarm is given when deformation reaches a fatigue limit or a bending angle limit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart device technology, specifically relating to a quad-camera AI smart glasses. Background Technology

[0002] AI smart glasses are wearable devices that combine miniature cameras and infrared sensors to scan the environment in real time. The scanned physical data is fed back through AI algorithms, and digital information is superimposed on the real field of vision through a transparent screen on the lens using optical waveguide technology.

[0003] Currently available AI smart glasses, such as the Quark AI Glasses launched by Alibaba Group, use a dual-chip architecture of Qualcomm Snapdragon AR1 and Hengxuan BES2800, are equipped with a hot-swappable battery system to achieve 24-hour battery life, enable phoneless QR code payment through voiceprint multi-factor authentication, and have jointly developed a near-eye display navigation system adapted for cycling and walking scenarios with Gaode Maps; or the Xiaomi AI Glasses launched by Xiaomi Technology Co., Ltd., which are equipped with a 12-megapixel high-definition image-stabilized camera, can achieve first-person perspective shooting, and generate immersive panoramic content by combining multiple cameras.

[0004] When worn, AI smart glasses can acquire images from different angles through different cameras. However, some cameras rely on the support of the temples of the glasses. Since the user's head will push the two temples apart, the cameras on the temples will be tilted. Existing smart glasses have failed to track changes in camera angle, which ultimately results in fluctuations in the image information aggregation results. Summary of the Invention

[0005] The purpose of this invention is to provide a quad-camera AI smart glasses that acquires images from multiple angles using different cameras, then performs time alignment and aggregation to form a panoramic image. It learns to configure different decoders for images from different angles to achieve adaptive updates, and issues an alert when the deformation reaches the fatigue limit or bending angle limit.

[0006] The specific technical solution adopted by this invention is as follows: A quad-camera AI smart glasses, including a frame, temples, a first camera and a second camera for capturing raw images, and also including: An embedded chip installed inside the temple of the eyeglasses, the embedded chip comprising: A data input interface is used to connect the first camera and the second camera; The encoder extracts the first feature from the original image using a CNN convolutional neural network; The aggregation module aggregates the first feature through an attention mechanism and outputs the global feature; The first decoder, based on global features and combined with spatial tensors, aggregates them into a panoramic image; An angle correction module is used to monitor the bending angle of the temples of the glasses and automatically correct the panoramic image based on the bending angle. The fatigue monitoring module is used to determine the bending fatigue limit at the bending angle.

[0007] As an alternative, the data input interface captures the original images at the same time, corrects the original images, and stacks them as the input feature map of the CNN convolutional neural network.

[0008] As an alternative, the encoder extracts a first feature from the original image using ResNet, and the encoder includes the following parts: Convolutional layers perform discrete convolution operations; The residual block processes a single input original image and outputs a feature map.

[0009] As an alternative, the encoder also includes a global average pooling operation that transforms the feature map into a one-dimensional first feature vector.

[0010] As an optional approach, the aggregation module aggregates the first feature through an attention mechanism, including the following steps: Dynamically learn the importance weights of each perspective and calculate the perspective weights; For each viewpoint's feature vector, an attention score is obtained through a neural network; Perform a weighted summation and output the global features.

[0011] As an optional approach, the first decoder aggregates panoramic images by including the following steps: The global features and spatial tensor are input together into the first decoder; The global features are input into the fully connected network, which is the first decoder. Each layer generates one set of style parameters; Starting with the spatial tensor through the first decoder, in Guided by this, upsampling is performed block by block; High-resolution feature maps are mapped to RGB space using a single convolutional layer, and then... The function constrains the value to .

[0012] As an optional solution, the angle correction module includes: Strain gauges are used to monitor the bending angle of the temples of the eyeglasses. The unit establishes a second decoder to automatically correct the panoramic image based on the curvature angle.

[0013] As an optional solution, the strain gauge is connected to the built-in chip; When the temple of the eyeglasses deforms, the axial strain of the strain gauge produces a bending angle; The bending angle value is derived based on the proportional relationship between the resistance change rate of the strain gauge and the axial strain.

[0014] As an alternative, the second decoder injects the bending angle into the decoding process, uses a supernetwork or modulation network to dynamically generate the influence parameters and output parameters of the second decoder, combines the influence parameters and output parameters to form specific parameters for the viewing angle, and performs decoding based on the specific parameters.

[0015] As an optional solution, the fatigue monitoring module determines the bending fatigue limit of the bending angle by the following steps: The document records the experience of the temples of the glasses during use. Variations in bending angle Each angle occurs Second-rate; Establish the relationship between bending angle and stress; The final number of bends corresponding to each stress level is determined by the SN curve. Calculate the total damage according to Miner's rule. ; When total damage At that time, the temple of the glasses reaches its fatigue limit.

[0016] The technical effects achieved by this invention are as follows: This invention acquires images from multiple angles using different cameras, then aligns these images in time using an end-to-end stitching network based on deep learning. It utilizes an attention mechanism to automatically focus on the second features of overlapping areas and ignores interference from moving objects, as well as to pay attention to the angle changes in the corresponding images caused by the deformation of the temples of the glasses. This enables the end-to-end stitching network to learn to configure different decoders for images from different angles, achieve adaptive updates, and monitor this deformation. When the deformation reaches the fatigue limit or bending angle limit, an alarm is issued.

[0017] This invention can adapt to changes in shooting angle when worn by different people, and configure decoders for different angle images for each user to achieve configuration updates for smart glasses. Attached Figure Description

[0018] Figure 1 This is a perspective view of the smart glasses of the present invention; Figure 2 This is a top view of the smart glasses of the present invention; Figure 3 This is a logical schematic diagram of the control signal transmission state of the smart glasses of the present invention; Figure 4This is a system block diagram of the built-in chip of the present invention.

[0019] The attached diagram lists the components represented by each number as follows: 1. Frame; 2. Temples; 3. First camera; 4. Second camera; 5. Built-in chip; 6. Central processing module; 7. Angle sensor; 8. Communication module; 9. Data input interface; 10. Encoder; 11. Aggregation module; 12. First decoder; 13. Angle correction module; 14. Second decoder; 15. Fatigue monitoring module; 16. Bending limit warning module. Detailed Implementation

[0020] To make the objectives and advantages of this invention clearer, the invention will be specifically described below with reference to embodiments. It should be understood that the following text is merely used to describe one or more specific embodiments of the invention and does not strictly limit the scope of protection specifically claimed by the invention.

[0021] like Figures 1-4 As shown, a quad-camera AI smart glasses includes a frame 1 and temples 2 mounted together by screws. At least one first camera 3 is embedded in the front of the frame 1, and at least one second camera 4 is embedded in the outer side of the temples 2. The first camera 3 and the second camera 4 are electrically connected and controlled by a built-in chip 5 encapsulated inside the temples 2. The built-in chip 5 can be equipped with a communication module 7 to transmit signals to the cloud. When worn, the first image is captured by the first camera 3, and the second image is captured by the second camera 4, which captures the image from the side. The first image and the second image are cropped and stitched together by an end-to-end stitching network based on deep learning preset by the built-in chip 5 to form a panoramic image. The image is then projected onto the lenses inside the frame 1 using optical waveguide technology, which can eliminate lateral blind spots and is used for guidance for the blind, reducing the occurrence of safety accidents caused by lateral blind spots.

[0022] As an optional embodiment, see Figure 3 Since the temple 2 contains a power supply, charging interface, and built-in chip 5 and their supporting circuits, the frame 1 needs to be in a relatively stable environment with the temple 2. Therefore, it is necessary to fix the temple 2 to the back of the frame 1. For the safety of the supporting circuits, sealant is needed to seal the connection between the frame 1 and the temple 2 to prevent water leakage.

[0023] See attached document Figure 3 and Figure 4To deploy a deep learning-based end-to-end stitching network at the edge, this embodiment develops a data input interface 9, an encoder 10, an aggregation module 11, and a decoder 12 in the built-in chip 5. The first and second images are input as raw data to the encoder 10, processed by the deep learning-based end-to-end stitching network preset in the aggregation module 11, and decoded by the decoder 12 to form a panoramic image, which is then output to the lens inside the frame 1. The configurations of each part are as follows: Pre-processing: The built-in chip 5 simultaneously sends trigger signals to the first camera 3 and the second camera 4 to control both to expose at the same time, or to add a high-precision timestamp to each frame of the image, so as to facilitate time alignment in the later stage through an end-to-end stitching network based on deep learning. At the same time, the intrinsic parameters such as focal length, optical center, and distortion coefficient of each camera, as well as the extrinsic parameters such as relative position and orientation, are obtained. Based on the intrinsic and extrinsic parameters, a spatial coordinate system is constructed with the midpoint of frame 1 as the origin. The first and second images are corrected using the preset distortion coefficients to obtain the original image without distortion. In order to eliminate the distortion of the camera by projecting different original images onto this unified spatial coordinate system. Data input interface 9, with Figure 1 For example, two first cameras 3 and two second cameras 4 are connected to capture four original images of the front, left and right at the same time. After the four original images are corrected, they are stacked into a 12-channel tensor, that is, 4 images * 3 RGB channels, which is used as the input feature map of the end-to-end stitching network based on deep learning. Encoder 10 extracts the first feature from multi-view images using a CNN convolutional neural network, such as ResNet or MobileNetV3-Lightweight. This embodiment uses ResNet as an example to extract the first feature from multi-view images. The specific steps are as follows: Convolutional layer: Perform discrete convolution operation and calculate as follows (1): Formula (1) in, It is the input feature map. These are the weights of the convolution kernel. It is the spatial location on the output feature map. It is the traversal index on the convolution kernel; The core of ResNet is the residual block, which solves the gradient vanishing problem in deep networks. A basic residual block can be expressed as formula (2): Formula (2) in, For the input of the i-th residual block, For the output of the i-th residual block, Indicates the residual portion. This represents the set of weights for the residual. For example, activation functions ; Assume there is A residual block, CNN convolutional neural network The processing of a single input feature map can be expressed as formula (3): Formula (3) in, For the first The input original image from each perspective, For all learnable parameters of ResNet, For from the first Feature maps extracted from each viewpoint The number of channels in the original input image. The height of the original image is input. The width of the original input image; As an alternative embodiment, global average pooling is used at the end of the CNN convolutional neural network to convert the feature map into a one-dimensional feature vector, calculated using the following formulas (4) and (5): Formula (4) Formula (5) in, For the first The first feature vector from each perspective; Aggregation module 11 aggregates the first feature vector based on an attention mechanism, allowing the attention mechanism to automatically focus on the second features of overlapping regions and ignore interference from moving objects. Here, the attention mechanism learns how to align features and understands the relationship between overlapping and non-overlapping regions of the four original images. The specific steps are as follows: The attention mechanism can dynamically learn the importance weight of each perspective, first calculating the perspective weight; Feature vectors for each viewpoint Through a small neural network, such as a single fully connected layer and Receive attention score Calculate using the following formulas (6) and (7): Formula (6) Formula (7) in, and These are preset, learnable parameters; Then, perform a weighted summation and calculate the following formula (8): Formula (8) in, For the final global features, a more informative perspective will contribute more to the final global features.

[0024] The first decoder 12, based on the global features of the final output. The output is a fixed, learnable space tensor. The images are then aggregated into a single panoramic image. The specific steps are as follows: A 4×4×512 space tensor , and global features Input the first decoder 12 together; global features Input a small fully connected network , for the first decoder 12 Each layer generates a set of style parameters. The calculation is performed using the following formula (9); Formula (9) via the first decoder 12 Starting from, in Guided by this, upsampling is performed block by block: Finally, a single convolutional layer maps the high-resolution feature map to the RGB space, and uses... The function constrains the value to Calculate using the following formula (10); Formula (10) in, This is the final output layer.

[0025] Since the second camera 4 relies on the support of the temples 2, different users' heads will spread the two temples 2 apart, causing the second camera 4 on the temples 2 to tilt. In order to track the change in the angle of the second camera 4 and facilitate the end-to-end stitching network to adapt to this angle change, this embodiment has strain gauges bonded inside the two temples 2 to monitor this angle change and to offset the fluctuations caused by the angle change during the image information aggregation process.

[0026] See attached document Figure 4Angle correction module 13 is used to monitor the bending angle of the temple 2 when it deforms left and right in real time. Angle correction module 13 can be made of strain gauges made of metal wire. Since the resistance change rate of the metal wire is proportional to the axial strain, the following formula (11) is used to calculate it: Formula (11) in, The resistance change rate of the strain gauge. The sensitivity coefficient of the strain gauge. The strain is the axial strain of the strain gauge; When a strain gauge bends on a horizontal plane, the ratio of surface strain to radius of curvature is calculated using the following formula (12): Formula (12) in, for, The radius of curvature; Furthermore, radius of curvature It can be expressed as the following formula (13): Formula (13) Among them, see Figure 2 , For the bending angle, The length of the curved arc; Substituting formula (13) into formula (12), we get formula (14): Formula (14) Substituting the bending strain formula (12) into the resistance change rate formula (11), we obtain the formula (15) for deriving the bending angle through the resistance change: Formula (15) Further deformation is used to solve for the bending angle. Calculate using the following formula (16): Formula (16) Among them, bending angle The unit is degrees, centered on the normal of frame 1, and extending to one side of the normal. The value is negative, and the other side is... It is a positive value; At this point, we can demonstrate this by constructing a single mathematical expression. This directly affects the relationship between global feature generation and the final output layer. Assuming a layer based on... Its expressive form can encode multi-view images into a single continuous 3D scene representation; Unit, establish a second decoder 14 based on, and Injection decoding process, assuming input is a 3D coordinate and viewing direction Output color and density The output set is calculated using the following formula (17): Formula (17) in, To affect the parameters, but in formula (17) With only one input value, the parameters affected are... It has no effect in itself; For more powerful modeling Due to the influence of this, this embodiment uses a supernetwork or modulation network, which is viewed from the perspective of... As input, a second decoder is dynamically generated 14 The parameters that affect the modulation of the second decoder 14, or the parameters that generate intermediate features of the modulation. Configure one modulation network Its parameters are The receiving perspective is Output one set of parameters Calculate using the following formula (18); Formula (18) The original influence parameters through the second decoder 14 With output parameters Combined, specific parameters for the viewpoint are formed, and the following formula (19) is used for calculation: Formula (19) in, This is the scaling factor; Furthermore, decoding is performed using the specific parameters of the modulation, and the following formula (20) is calculated: Formula (20) This directly determines the specific parameters of the second decoder 14. The model can then learn to perform different tasks. The second decoder 14 has a different configuration.

[0027] See attached document Figure 4 The fatigue monitoring module 15 is used to record the fatigue experience of the temples 2 during a single use. Variations in bending angle Each angle occurred Second-rate; Establish the relationship between bending angle and stress, and calculate using the following formula (21): Formula (21) in, The quality function of the temple 2 itself is preset based on material mechanics and structure; Then, using the SN curve, the final number of bends corresponding to each stress level is determined. Calculate using the following formula (21): Formula (21) Finally, the total damage is calculated according to Miner's rule. Calculate using the following formula (22): Formula (22) When total damage At that time, the second temple of the glasses reached its fatigue limit, and the total number of bends was set to be... It can make Solve the following formula (23); Formula (23) in, For the bending angle of the temples 2 during each use The probability of occurrence, and satisfying and .

[0028] See attached document Figure 4 The bending upper limit warning module 16 is used to set the bending angle limit threshold and to set the bending angle for each bending angle. The bending angle is compared with the bending angle limit threshold. If the bending angle limit threshold is exceeded, the second camera 4 at the corresponding position is turned off, and an alarm model is projected onto the lens of the frame 1. Specifically, this includes: The bending angle limit determination unit is used to set the bending angle limit threshold. , each bending angle With bending angle limit threshold A comparison is performed; if the bending angle exceeds the limit threshold... If the second camera 4 at the corresponding location is turned off, or if the bending angle limit threshold is not exceeded, then the second camera 4 at the corresponding location will be turned off. Normally, the second camera 4 is activated; The warning unit stores warning models, such as a triangular digital model with an exclamation mark. When the bending angle exceeds the limit threshold, the warning model is projected onto the lens using light wave technology to alert the user.

[0029] In summary, this invention acquires images from multiple angles using different cameras, then uses a deep learning-based end-to-end stitching network to time-align these images. It automatically focuses on the second feature of overlapping areas using an attention mechanism, ignores interference from moving objects, and pays attention to the angle changes in the corresponding images caused by the deformation of the temple 2. This enables the end-to-end stitching network to learn to configure different decoders for images from different angles, achieve adaptive updates, and monitor this deformation. When the deformation reaches the fatigue limit or bending angle limit, an alarm is issued.

[0030] The above description is merely an optional embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention, unless otherwise specified or limited, shall be implemented according to conventional means in the art.

Claims

1. A quad-camera AI smart glasses, comprising a frame (1), temples (2), a first camera (3) for capturing raw images, and a second camera (4), characterized in that, Also includes: An embedded chip (5) installed inside the temple (2) of the eyeglasses, the embedded chip (5) comprising: Data input interface (9) is used to connect the first camera (3) and the second camera (4); The encoder (10) extracts the first feature from the original image through a CNN convolutional neural network; The aggregation module (11) aggregates the first feature through an attention mechanism and outputs the global feature; The first decoder (12) aggregates global features and spatial tensors to form a panoramic image; An angle correction module (13) is used to monitor the bending angle of the temple (2) of the glasses and automatically correct the panoramic image according to the bending angle; The fatigue monitoring module (15) is used to determine the bending fatigue limit of the bending angle.

2. The smart glasses according to claim 1, characterized in that: The data input interface (9) captures the original images at the same time, corrects the original images and stacks them as the input feature map of the CNN convolutional neural network.

3. The smart glasses according to claim 1, characterized in that: The encoder (10) extracts a first feature from the original image using ResNet, and the encoder (10) includes the following parts: Convolutional layers perform discrete convolution operations; The residual block processes a single input original image and outputs a feature map.

4. The smart glasses according to claim 3, characterized in that: The encoder (10) further includes a global average pooling operation that converts the feature map into a one-dimensional first feature vector.

5. The smart glasses according to claim 1, characterized in that: The aggregation module (11) aggregates the first feature through an attention mechanism, including the following steps: Dynamically learn the importance weights of each perspective and calculate the perspective weights; For each viewpoint's feature vector, an attention score is obtained through a neural network; Perform a weighted summation and output the global features.

6. The smart glasses according to claim 1, characterized in that: The first decoder (12) aggregates panoramic images by the following steps: The global features and spatial tensor are input together into the first decoder (12). The global features are input into the fully connected network to serve as the first decoder (12). Each layer generates one set of style parameters; Starting with the spatial tensor through the first decoder (12), in Guided by this, upsampling is performed block by block; The feature map is mapped to the RGB space through a convolutional layer, and then... The function constrains the value to .

7. The smart glasses according to claim 1, characterized in that, The angle correction module (13) includes: Strain gauges are used to monitor the bending angle of the temple (2) of the eyeglasses; Unit, establish a second decoder (14), automatically correct the panoramic image according to the curvature angle.

8. The smart glasses according to claim 7, characterized in that: The strain gauge is connected to the built-in chip (5). When the temple (2) of the eyeglasses deforms, the axial strain of the strain gauge produces a bending angle; The bending angle value is derived based on the proportional relationship between the resistance change rate of the strain gauge and the axial strain.

9. The smart glasses according to claim 7, characterized in that: The second decoder (14) injects the bending angle into the decoding process, uses a super network or modulation network to dynamically generate the influence parameters and output parameters of the second decoder (14), combines the influence parameters and output parameters to form specific parameters for the viewpoint, and performs decoding based on the specific parameters.

10. The smart glasses according to claim 1, characterized in that: The fatigue monitoring module (15) determines the bending fatigue limit of the bending angle by the following steps: The record shows that the temples (2) of the glasses experienced the following during use: Variations in bending angle Each angle occurs Second-rate; Establish the relationship between bending angle and stress; The final number of bends corresponding to each stress level is determined by the SN curve. Calculate the total damage according to Miner's rule. ; When total damage At that time, the temple (2) of the glasses reaches its fatigue limit.