Visual field optimization method, device and equipment and computer readable storage medium

By dynamically acquiring the wearer's posture data in AR glasses, activating the lower camera to collect and process keyboard images, the problem of AR glasses obstructing the field of vision is solved, achieving high-precision keyboard visual restoration and an immersive operating experience.

CN121613624APending Publication Date: 2026-03-06GOERTEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511922597.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Current mainstream AR glasses designs obstruct the lower half of the user's field of vision, making it impossible to directly see the hands and the keyboard on the desktop, affecting typing efficiency and user experience.

Method used

By dynamically acquiring the posture data of the smart glasses wearer, using an inertial measurement unit and an eye-tracking module to determine the probability of looking down, activating the lower camera to capture images, acquiring a keyboard model based on image processing technology, and performing image restoration and rendering, which is then mapped into the main field of view.

Benefits of technology

It achieves seamless and precise system triggering, improves image segmentation accuracy and robustness under occlusion and lighting changes, and provides an immersive visual experience that goes beyond simple video perspective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121613624A_ABST
    Figure CN121613624A_ABST
Patent Text Reader

Abstract

The invention discloses a visual field optimization method, device and equipment and a computer readable storage medium, and belongs to the technical field of computer vision. The visual field optimization method is applied to the intelligent glasses comprising a lower camera, and comprises the following steps: dynamically obtaining posture data of a wearer of the intelligent glasses; determining the head lowering probability of the wearer according to the posture data; under the condition that the head lowering probability is greater than a preset threshold value, activating a lower camera; acquiring an acquired image of the lower camera, and acquiring a current keyboard model according to the acquired image; performing image processing based on the acquired image and the current keyboard model to obtain a beautified image; and mapping the beautified image into a main view field of the intelligent glasses for a wearer to view. The visual field optimization method provided by the invention can solve the problem that the intelligent glasses can shield part of the visual field of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to methods, apparatus, devices and computer-readable storage media for field of view optimization. Background Technology

[0002] With the rapid development of Augmented Reality (AR) technology, AR glasses are playing a wide role in productivity scenarios such as virtual desktops, remote collaboration, programming, and document processing. Users expect to be able to efficiently operate their physical devices, such as keyboards and mice, while using AR virtual screens to complete tasks and work in their daily lives.

[0003] However, current mainstream AR glasses have a significant interaction flaw in their design: the device itself usually obstructs the lower half of the user's field of vision, especially the area directly in front of the nose and cheeks. This prevents the user from directly seeing their hands and the keyboard on the table. Summary of the Invention

[0004] The main objective of this application is to provide a method, apparatus, device, and computer-readable storage medium for optimizing field of vision, aiming to solve the technical problem that smart glasses may obstruct part of the user's field of vision.

[0005] To achieve the above objectives, this application provides a field-of-view optimization method, which is applied to smart glasses including a downward-facing camera, comprising: Dynamically acquire the posture data of the wearer of the smart glasses; The probability of the wearer looking down is determined based on the posture data; When the probability of looking down exceeds a preset threshold, the lower-mounted camera is activated; Acquire the image captured by the lower camera, and obtain the current keyboard model based on the captured image; A beautified image is obtained by performing image processing based on the acquired image and the current keyboard model; The beautified image is mapped onto the main field of view of the smart glasses for the wearer to view.

[0006] In one embodiment, the posture data includes head posture data and eye-tracking data; the step of dynamically acquiring the posture data of the wearer of the smart glasses includes: Dynamically acquire the head posture data collected by the preset inertial measurement unit in the smart glasses; The eye-tracking data collected by the preset eye-tracking module in the smart glasses is dynamically acquired.

[0007] In one embodiment, the step of determining the wearer's head-down probability based on the posture data includes: Determine prior probabilities based on attitude data; Obtain the likelihood probability determined based on preset training data; The prior probability and the likelihood probability are substituted into a preset Bayesian inference model for calculation to obtain the posterior probability, which is the probability of looking down.

[0008] In one embodiment, the step of obtaining the current keyboard model based on the acquired image includes: Identify the keyboard model in the captured image; Determine whether a keyboard model matching the specified model exists in the local storage space; If so, retrieve the current keyboard model from local storage. If not, download the current keyboard model from the cloud model library.

[0009] In one embodiment, the step of obtaining a beautified image by image processing based on the acquired image and the current keyboard model includes: The camera pose of the current keyboard model is calculated using a preset PnP algorithm; A precise segmentation mask is obtained by performing image segmentation on the hand and keyboard images in the acquired images using a preset energy minimization segmentation algorithm. Based on the overlapping area and depth information of the precisely segmented mask, the interaction state between the hand and the keyboard is determined. Image restoration is performed based on the interaction state, the camera pose, and the preset perspective projection formula to obtain an enhanced image.

[0010] In one embodiment, the step of image inpainting based on the parameters of the lower-mounted camera, the camera pose, and the current camera model includes: The occlusion area is determined based on the interaction state; The parameters of the lower camera, the camera pose, and the 3D coordinates of the keycap point in the current camera model are substituted into a preset perspective projection formula to calculate the content to be displayed. Based on the content to be displayed and the pre-stored keycap template image, the image is drawn onto the occluded area after homography transformation to complete the image restoration.

[0011] In one embodiment, the step of mapping the beautified image onto the main field of view of the smart glasses for the wearer to view includes: The beautified image is then fused and rendered with the main field of view of the smart glasses for the wearer to view.

[0012] Furthermore, to achieve the above objectives, this application also provides a field-of-view optimization device, which is applied to smart glasses including a downward-facing camera, comprising: An acquisition module is used to dynamically acquire the posture data of the wearer of the smart glasses; The processing module is used to determine the probability of the wearer looking down based on the posture data; The processing module is also used to activate the lower camera when the probability of looking down is greater than a preset threshold; The acquisition module is also used to acquire the captured image from the lower camera and to acquire the current keyboard model based on the captured image; The processing module is also used to perform image processing based on the acquired image and the current keyboard model to obtain a beautified image; The processing module is also used to map the beautified image onto the main field of view of the smart glasses for the wearer to view.

[0013] In addition, to achieve the above objectives, this application also provides smart glasses, the smart glasses comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the vision optimization method as described above.

[0014] In addition, to achieve the above objectives, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the field-of-view optimization method as described above.

[0015] This application proposes a field-of-view optimization method, apparatus, device, and computer-readable storage medium. In this method, by dynamically acquiring the wearer's posture data of the smart glasses and determining the probability of the wearer looking down based on the posture data, the user's interaction intent can be intelligently inferred. By activating the lower-facing camera when the probability of looking down exceeds a preset threshold, a seamless and precise system triggering can be achieved. By acquiring images from the lower-facing camera and obtaining the current keyboard model based on these images, image processing is performed on the acquired images and the current keyboard model to obtain an enhanced image. This fundamentally improves the accuracy and robustness of image segmentation in complex environments such as occlusion and lighting changes. High-fidelity virtual repair and enhanced rendering are performed on the keycap visual information missing due to finger occlusion. The enhanced image is mapped onto the main field of view of the smart glasses for the wearer to view, providing a visual experience that surpasses simple video perspective. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only a part of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating a field-of-view optimization method provided in an embodiment of this application; Figure 2 for Figure 1 A detailed flowchart of step S20; Figure 3 for Figure 1 A detailed flowchart of step S40; Figure 4 for Figure 1 A detailed flowchart of step S50; Figure 5 for Figure 4 A detailed flowchart of step S54; Figure 6 This is a schematic diagram of a field-of-view optimization device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a smart glasses provided in an embodiment of this application. Detailed Implementation

[0018] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that the embodiments of this application can also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the embodiments of this application with unnecessary detail.

[0019] With the rapid development of augmented reality technology, AR glasses and XR (Extended Reality) glasses are playing a wide role in productivity scenarios such as virtual desktops, remote collaboration, programming, and document processing. Users expect to be able to efficiently operate their physical devices, such as keyboards and mice, while using AR virtual screens to complete tasks and work in their daily lives.

[0020] However, current mainstream AR glasses have a significant design flaw: the device itself typically obstructs the lower half of the user's field of vision, especially the area directly in front of the nose and cheeks. This prevents users from directly seeing their hands and the keyboard on the table, leading to the following problems: 1) Typing difficulties: Users cannot accurately place their fingers on the correct keys, resulting in decreased efficiency. 2) Disjointed experience: Users need to frequently remove or move the glasses to confirm hand positions, severely impacting the AR glasses experience. 3) High learning curve: For non-professionals, frequently switching between different keyboards degrades the user experience.

[0021] To address the aforementioned issues, related technologies include the following solutions, which still have shortcomings: 1) Simple Video See-Through (VST): Some high-end AR devices are equipped with see-through cameras, but simply displaying the video stream from the lower camera in full-screen or floating mode excessively attracts user attention, disrupting the immersive experience of the main display area. Furthermore, the unprocessed raw video footage clashes with the style of the virtual environment, appearing jarring. 2) Virtual Keyboard Projection: This method uses a camera to recognize hand gestures for input on a virtual keyboard, but it suffers from slow input speed, low accuracy, and fatigue, failing to meet the demands of long-term, high-efficiency productivity. 3) Basic Image Recognition: Relying solely on simple color segmentation or template matching leads to a sharp decline in recognition accuracy and robustness under complex lighting, occlusion, or when using different keyboards, resulting in an unreliable user experience.

[0022] Based on this, embodiments of this application provide a field-of-view optimization method, apparatus, device, and computer-readable storage medium. By dynamically acquiring the posture data of the wearer of smart glasses and determining the probability of the wearer looking down based on the posture data, the user's interaction intention can be intelligently inferred. By activating the lower camera when the probability of looking down is greater than a preset threshold, a seamless and precise system trigger can be achieved. By acquiring the captured image from the lower camera and obtaining the current keyboard model based on the captured image, image processing is performed on the captured image and the current keyboard model to obtain an enhanced image, which can fundamentally improve the image segmentation accuracy and robustness in complex environments such as occlusion and changes in lighting. High-fidelity virtual repair and enhanced rendering are performed on the visual information of the keycaps that are missing due to finger occlusion. The enhanced image is mapped onto the main field of view of the smart glasses for the wearer to view, providing a visual experience that goes beyond simple video perspective.

[0023] The field of view optimization method, apparatus, device, and computer-readable storage medium provided in the embodiments of this application are specifically described through the following embodiments. First, the field of view optimization method in the embodiments of this application is described.

[0024] This application provides a field-of-view optimization method, referring to... Figure 1 , Figure 1This is a flowchart illustrating a field-of-view optimization method provided in an embodiment of this application. This field-of-view optimization method can be applied to smart glasses that include a lower-mounted camera, such as... Figure 1 As shown, the field of view optimization method provided in this embodiment includes steps S10 to S60.

[0025] Step S10: Dynamically acquire the posture data of the wearer of the smart glasses; Step S20: Determine the probability of the wearer looking down based on posture data; In this embodiment, the smart glasses may include a bottom-mounted wide-angle camera located at the bottom edge of the device, a preset inertial measurement unit (IMU), a preset eye-tracking module, and a processing and computing unit. The unique design and layout of the camera ensures that its field of view can completely cover the wearer's hands and keyboard operation area. The smart glasses can continuously monitor the wearer's posture data to see if there is an intention to look down at the keyboard through the inertial measurement unit and the eye-tracking module. For example, the head posture data collected by the preset inertial measurement unit in the smart glasses is dynamically acquired, and the eye-tracking data collected by the preset eye-tracking module in the smart glasses is dynamically acquired. Then, the probability of the wearer looking down is determined based on the head posture data and the eye-tracking data. If the probability of looking down is greater than a preset threshold, it is considered that the wearer has an intention to look down at the keyboard. If the probability of looking down is not greater than the preset threshold, it is considered that the wearer does not have an intention to look down at the keyboard.

[0026] Step S30: Activate the lower camera when the probability of looking down is greater than a preset threshold; Understandably, if the probability of looking down exceeds a preset threshold, and it is determined that the wearer intends to look down at the keyboard, the lower camera can be activated to capture images of the field of vision blocked by the smart glasses. These images can then serve as the basis for providing the wearer with a complete field of vision.

[0027] Reference Figure 2 In some feasible embodiments, step S20 above may include: Step S21: Determine the prior probability based on the attitude data; Step S22: Obtain the likelihood probability determined based on the preset training data; Step S23: Substitute the prior probability and likelihood probability into the preset Bayesian inference model for calculation to obtain the posterior probability as the probability of looking down.

[0028] In this embodiment, the smart glasses can calculate the posterior probability that the user has the intention to view the keyboard I using head posture data H provided by the inertial measurement unit, eye tracking data (gaze vector G) provided by the eye tracking module, and a Bayesian inference model: P(I|H,G)=[P(H,G|I)·P(I)] / P(H,G); where P(I) is the prior probability, which can be dynamically updated based on the user's recent behavior, i.e., head posture data H and gaze vector G; P(H,G|I) is the likelihood probability, which can be described by a probability model established by a large amount of training data; and when the posterior probability P(I|H,G) exceeds the preset threshold τ_activate, the lower camera can be automatically activated to realize the subsequent image processing process.

[0029] This embodiment can realize intelligent intent reasoning by using a multimodal sensor fusion algorithm based on Bayes' theorem to intelligently infer the wearer's interaction intent, thereby achieving seamless and accurate system triggering.

[0030] Step S40: Obtain the image captured by the lower camera, and obtain the current keyboard model based on the captured image; In this embodiment, after activating the lower camera, a frame of image can be captured first. The image is then preprocessed by adjusting brightness, contrast, and denoising. It is then detected and recognized by a lightweight convolutional neural network. After the keyboard model is successfully recognized, the keyboard model can be obtained.

[0031] Reference Figure 3 In some feasible embodiments, the step of obtaining the current keyboard model based on the acquired image in step S40 above may include: Step S41: Identify the keyboard model in the captured image; Step S42: Determine whether a keyboard model matching the model exists in the local storage space; Step S43: If yes, retrieve the current keyboard model from the local storage space; Step S44: If not, download the current keyboard model from the cloud model library.

[0032] In this embodiment, after the smart glasses successfully identify the keyboard model in the keyboard image captured by the lower camera, it will first search the local storage space for a keyboard model that matches the model. If it does, it will directly obtain the current keyboard model from the local storage space. If it does not, it will download the high-precision 3D model of the current keyboard from the cloud model library.

[0033] Step S50: Perform image processing based on the acquired image and the current keyboard model to obtain a beautified image; It should be noted that since the captured images may include not only the keyboard image but also the image of the wearer's hands operating the keyboard, the hand image may obscure the keyboard image. Therefore, in order to allow the wearer to clearly see the entire keyboard and facilitate subsequent operation, it is necessary to perform image processing on the captured images to obtain an enhanced image based on the current keyboard model, rather than directly providing the wearer with the real-time captured image.

[0034] Reference Figure 4 In some feasible embodiments, step S50 above may include: Step S51: Calculate the camera pose of the current keyboard model using the preset PnP algorithm; Step S52: Use a preset energy minimization segmentation algorithm to perform image segmentation on the hand image and keyboard image in the acquired image to obtain a precise segmentation mask; Step S53: Determine the interaction state between the hand and the keyboard based on the overlapping area and depth information of the precisely segmented mask. Step S54: Perform image restoration based on the interaction state, camera pose, and preset perspective projection formula to obtain an enhanced image.

[0035] In this embodiment, the precise pose of the model relative to the camera coordinate system (rotation matrix R and translation vector t) can be solved first using the PnP algorithm. The solution process of the PnP algorithm is as follows: Based on the high-precision 3D model M of the current keyboard downloaded from the cloud model library, the 3D model M contains the precise 3D coordinates (X, Y, Z) of each key point on the keyboard (keys in the four corners, the center of the function keys, etc.); through the keyboard detection neural network, the 2D pixel coordinates (u, v) corresponding to the aforementioned 3D key points are detected on the image containing the keyboard captured by the lower camera; PnP calculation is performed to obtain two sets of corresponding point sets: 3D point set: {P1, P2, P3, ..., Pn} (from the model); 2D point set: {p1, p2, p3, ..., pn} (from the image). Using these two sets of points, the camera pose [R|t] of the lower camera relative to the current keyboard model is calculated.

[0036] In this embodiment, the segmentation problem in the acquired image can be transformed into solving the label field L (L_P represents the label of pixel p, such as background, keyboard, hand) to minimize the following energy function E(L): E(L) = λ·E_data(L) + E_smooth(L) + E_model(L); where, the data term E_data(L): based on pixel color, depth and other features, measures the label assignment cost, E_data(L) = Σ_p[-logP(I_p|L_p)]; the smoothing term E_smooth(L): The Potts model is adopted to encourage spatial continuity: E_smooth(L)=Σ_{p,q∈N}[1-δ(L_p,L_q)]•exp(-β||I_p-I_q||²). The model prior term E_model(L) introduces a 3D keyboard model M as a strong constraint to penalize erroneous segmentations that deviate from the model's projection region: E_model(L)=γ•Σ_p|D(p,M)|•[1-δ(L_p,L_keyboard)], where D(p,M) is the geometric distance from pixel p to the projected contour of model M. This embodiment achieves accurate segmentation of the keyboard and hand by minimizing the energy function E(L), and the innovatively introduced model prior term E_model(L) utilizes the known keyboard geometry as a strong constraint, ensuring extremely high robustness of the segmentation.

[0037] This embodiment can achieve accurate and robust segmentation. It uses the precise 3D geometric model of the keyboard as prior knowledge and guides image segmentation through an energy function minimization framework, which fundamentally improves the segmentation accuracy and robustness in complex environments such as occlusion and lighting changes.

[0038] In this embodiment, the interaction state of the hand and keyboard ("pressed", "hovered" or "no interaction") can be determined based on the overlapping area and depth information of the accurately segmented mask, and it can be determined whether there is an occlusion relationship between the two. Then, the corresponding image repair and enhancement rendering methods are adopted to output the processed image. For example, the occluded area can be virtually repaired by using the perspective projection principle based on the pose parameters [R|t] calculated by the PnP algorithm in the aforementioned steps.

[0039] Reference Figure 5 In some feasible embodiments, the step of image repair based on the interaction state, camera pose, and preset perspective projection formula in step S54 above may include: Step S541: Determine the occlusion area based on the interaction state; Step S542: Substitute the parameters of the lower camera, the camera pose, and the 3D coordinates of the keycap point in the current camera model into the preset perspective projection formula to calculate the content to be displayed. Step S543: Based on the content to be displayed and the pre-stored keycap template image, the image is drawn onto the occluded area after homography transformation to complete the image restoration.

[0040] In this embodiment, for the keycap area determined to be occluded, the keyboard model M and the camera parameters of the lower-facing camera from the aforementioned steps can be used to calculate the content to be displayed using the perspective projection formula: p_screen=K·[R|t]·P_model; where K is the intrinsic parameter matrix of the lower-facing camera, and P_model is the 3D coordinates of a point on the keycap. Based on this projection relationship, the pre-stored keycap template image can be transformed using homography and then drawn onto the occluded area, thereby completing the visual restoration. As an example, aesthetic enhancements such as semi-transparency and luminous contours can also be applied to the keyboard and hand.

[0041] This embodiment enables realistic visual restoration. Based on the principles of perspective geometry, it performs high-fidelity virtual restoration and enhanced rendering of the visual information of the keycaps missing due to finger obstruction, providing a visual experience that surpasses simple video perspective.

[0042] Step S60: Map the beautified image onto the main field of view of the smart glasses for the wearer to view.

[0043] In this embodiment, the beautified image can be fused and rendered with the main field of view of the smart glasses. That is, the processed keyboard and hand images are fused into the user's AR main field of view in a low-latency and high-fidelity manner for the wearer to view, thereby achieving a seamless, immersive, blind-spot-free visual experience.

[0044] This embodiment can also achieve closed-loop optimization. The smart glasses have the ability to learn adaptively based on the wearer's habits, and can adaptively adjust parameters to continuously optimize the personalized experience.

[0045] This embodiment provides a field-of-view optimization method, which brings the following significant technical effects: 1) High-reliability interaction: Intent recognition based on a probabilistic model significantly reduces the false trigger rate, making the system response more in line with the user's natural expectations; 2) Exceptional accuracy and robustness: By combining the energy-minimizing segmentation framework with a geometric prior model, even in challenging environments such as strong light, shadows, and partial occlusion, segmentation accuracy far exceeds that of traditional algorithms, laying a solid foundation for reliable interaction; 3) Ultimate immersive visual experience: Based on physically based perspective projection and image restoration technology, a keyboard screen with extremely high visual consistency and no missing information is generated, completely eliminating the "disjointed feeling" of traditional video perspective and providing a truly immersive blind-spot-free office experience; 4) High-efficiency computing performance: The introduction of a prior model constrains the solution space, reducing the dependence on complex calculations. Combined with lightweight network design, it ensures that the entire system can run in real time on a mobile computing platform, meeting the power consumption and performance requirements of smart glasses such as AR glasses.

[0046] Furthermore, embodiments of this application also propose a field-of-view optimization device, referring to... Figure 6 , Figure 6 This is a schematic diagram of a field-of-view optimization device provided in an embodiment of this application, as shown below. Figure 6 As shown, in this embodiment, the field of view optimization device can be applied to smart glasses containing a lower-mounted camera, including: an acquisition module 10 and a processing module 20.

[0047] The acquisition module 10 is used to dynamically acquire the posture data of the wearer of the smart glasses; Processing module 20 is used to determine the probability of the wearer looking down based on posture data; The processing module 20 is also used to activate the lower camera when the probability of looking down is greater than a preset threshold; The acquisition module 10 is also used to acquire images captured by the lower camera and to obtain the current keyboard model based on the captured images; The processing module 20 is also used to perform image processing based on the acquired image and the current keyboard model to obtain an enhanced image; The processing module 20 is also used to map the beautified image onto the main field of view of the smart glasses for the wearer to view.

[0048] In some feasible embodiments, the pose data includes head pose data and eye-tracking data; the acquisition module 10 is also used for: Dynamically acquire head posture data collected by the preset inertial measurement unit in the smart glasses; Dynamically acquire eye-tracking data collected by the preset eye-tracking module in the smart glasses.

[0049] In some feasible embodiments, the processing module 20 is further configured to: Determine prior probabilities based on attitude data; Obtain the likelihood probability determined based on preset training data; Substitute the prior probability and likelihood probability into the preset Bayesian inference model to calculate the posterior probability, which serves as the probability of looking down.

[0050] In some feasible embodiments, the acquisition module 10 is further configured to: Identify the keyboard model in the captured image; Determine if a keyboard model matching the specified model exists in the local storage space; If so, retrieve the current keyboard model from local storage. If not, download the current keyboard model from the cloud model library.

[0051] In some feasible embodiments, the processing module 20 is further configured to: The camera pose of the current keyboard model is calculated using a preset PnP algorithm. A precise segmentation mask is obtained by performing image segmentation on the hand and keyboard images in the acquired images using a preset energy minimization segmentation algorithm. Based on the overlapping area and depth information of the precisely segmented mask, the interaction state between the hand and the keyboard is determined. Image restoration is performed based on the interaction state, the camera pose, and the preset perspective projection formula to obtain an enhanced image.

[0052] In some feasible embodiments, the processing module 20 is further configured to: Determine the occlusion area based on the interaction state; The parameters of the lower camera, the camera pose, and the 3D coordinates of the keycap points in the current camera model are substituted into the preset perspective projection formula to calculate the content to be displayed. The image restoration is completed by applying homography transformation to the occluded area based on the content to be displayed and the pre-stored keycap template image.

[0053] In some feasible embodiments, the processing module 20 is further configured to: The enhanced image is blended and rendered with the main field of view of the smart glasses for the wearer to view.

[0054] The field of view optimization device provided in this embodiment belongs to the same technical concept as the field of view optimization method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in any of the above embodiments. Furthermore, this embodiment has the same beneficial effects as the field of view optimization method.

[0055] Furthermore, this application also provides a smart glasses, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the field of view optimization method in any of the above embodiments.

[0056] The following is for reference. Figure 7 The diagram illustrates a structure suitable for implementing smart glasses according to embodiments of this application. The smart glasses in these embodiments may include, but are not limited to, AR glasses, XR glasses, etc. Figure 7 The smart glasses shown are merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0057] like Figure 7 As shown, the smart glasses may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the smart glasses. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 allows the smart glasses to communicate wirelessly or wiredly with other devices to exchange data. While the diagram shows smart glasses with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0058] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0059] The beneficial effects of the smart glasses provided in this application are the same as those of the field of vision optimization method provided in the above embodiments, and other technical features of the smart glasses are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.

[0060] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0061] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0062] Furthermore, embodiments of this application also provide a computer-readable storage medium, which may be a non-volatile computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the field-of-view optimization method provided in any of the above embodiments.

[0063] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0064] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0065] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, enable the electronic device to implement the aforementioned field-of-view optimization method.

[0066] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0067] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0068] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0069] The readable storage medium provided in this embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for performing the above-described field of view optimization method. The beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the field of view optimization method provided in the above-described embodiment, and will not be repeated here.

[0070] Furthermore, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the field-of-view optimization method provided in any of the above embodiments.

[0071] The computer program product provided in this embodiment belongs to the same technical concept as the field of view optimization method proposed in the above embodiment. Compared with related technologies, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the field of view optimization method provided in the above embodiment, and will not be repeated here.

[0072] It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown in the flowchart. The terms "first," "second," etc., in the specification, claims, and drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0073] It should also be understood that references to "one embodiment" or "some embodiments" in the specification of embodiments of this application mean that one or more embodiments of this application include the specific features, structures, or characteristics described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0074] The above describes some implementation methods of the embodiments of this application. However, the embodiments of this application are not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.

Claims

1. A method of field optimization, characterized by, The field of view optimization method is applied to smart glasses comprising a downward camera, and comprises the following steps: dynamically obtaining posture data of a wearer of the smart glasses; determining a probability of the wearer lowering his head according to the posture data; activating the downward camera if the probability of the wearer lowering his head is greater than a preset threshold; obtaining a captured image of the downward camera, and obtaining a current keyboard model according to the captured image; performing image processing based on the captured image and the current keyboard model to obtain a beautified image; mapping the beautified image into a main field of view of the smart glasses for the wearer to view.

2. The field of view optimization method of claim 1, wherein, The posture data comprises head posture data and eye tracking data; The step of dynamically obtaining the posture data of the wearer of the smart glasses comprises the following steps: dynamically obtaining the head posture data collected by a preset inertial measurement unit in the smart glasses; dynamically obtaining the eye tracking data collected by a preset eye tracking module in the smart glasses.

3. The field of view optimization method of claim 1, wherein, The step of determining the probability of the wearer lowering his head according to the posture data comprises the following steps: determining a prior probability based on the posture data; obtaining a likelihood probability determined based on preset training data; substituting the prior probability and the likelihood probability into a preset Bayesian inference model to obtain a posterior probability as the probability of the wearer lowering his head.

4. The field of view optimization method of claim 1, wherein, The step of obtaining the current keyboard model according to the captured image comprises the following steps: identifying a model of the keyboard in the captured image; determining whether a keyboard model matching the model exists in a local storage space; if yes, obtaining the current keyboard model from the local storage space; if no, downloading the current keyboard model from a cloud model library.

5. The field of view optimization method of claim 1, wherein, The step of performing image processing based on the captured image and the current keyboard model to obtain the beautified image comprises the following steps: calculating a camera pose of the current keyboard model by a preset PnP algorithm; performing image segmentation on hand images and keyboard images in the captured image by a preset energy minimization segmentation algorithm to obtain a precise segmentation mask; determining an interaction state of the hand and the keyboard according to overlapping regions and depth information of the precise segmentation mask; performing image restoration according to the interaction state, the camera pose, and a preset perspective projection formula to obtain the beautified image.

6. The field of view optimization method of claim 5, wherein, The step of performing image restoration according to the interaction state, the camera pose, and the preset perspective projection formula comprises the following steps: determining an occlusion region based on the interaction state; substituting parameters of the downward camera, the camera pose, and 3D coordinates of a point on a keycap in the current camera model into a preset perspective projection formula to obtain content that should be displayed; drawing the content that should be displayed and a pre-stored keycap template image to the occlusion region after homographic transformation to complete image restoration.

7. The field of view optimization method of claim 1, wherein, The step of mapping the beautified image into the main field of view of the smart glasses for the wearer to view comprises the following step: fusing and rendering the beautified image with the main field of view of the smart glasses for the wearer to view.

8. A field of view optimization device, characterized in that The field of view optimization device is applied to smart glasses comprising a downward camera, and comprises the following steps: an obtaining module, which is configured to dynamically obtain posture data of a wearer of the smart glasses; a processing module, configured to determine a probability of the wearer lowering his head according to the attitude data; the processing module is further configured to activate the downward camera when the probability of the wearer lowering his head is greater than a preset threshold; the acquisition module is further configured to acquire an image collected by the downward camera, and acquire a current keyboard model according to the image collected by the downward camera; the processing module is further configured to perform image processing based on the image collected by the downward camera and the current keyboard model to obtain a beautified image; the processing module is further configured to map the beautified image to a main visual field of the smart glasses for the wearer to view.

9. An intelligent eyewear, characterized in that, The smart glasses comprise a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the visual field optimization method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the visual field optimization method according to any one of claims 1 to 7.