Object selection method, selection frame display method, and device and medium
By presenting selection boxes with different transmittance on the lenses of smart head-mounted devices and combining them with an FOV mapping algorithm, the problem of cumbersome object selection interaction in smart head-mounted devices is solved, achieving an efficient and intuitive user interaction experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-06-30
- Publication Date
- 2026-04-23
AI Technical Summary
Existing smart head-mounted devices suffer from cumbersome and inefficient interaction processes during object selection and interaction, resulting in a poor user experience. In particular, when using voice interaction and eye-tracking technologies, they suffer from high costs, high power consumption, and difficulty in guaranteeing accuracy.
By displaying selection boxes with different transmittance on the lenses of smart head-mounted devices, the FOV mapping algorithm is used to accurately locate the object selected by the user. Combined with electrochromic technology or other optical display methods, the selection box can be accurately presented and adjusted.
It simplifies the user interaction process, improves the intuitiveness and accuracy of the interaction, reduces the system's computing power requirements, and enhances interaction efficiency and user experience.
Smart Images

Figure CN2025105930_23042026_PF_FP_ABST
Abstract
Description
Object selection method, selection box rendering method, device and medium
[0001] This application claims priority to Chinese Patent Application No. 202411457859.5, filed on October 17, 2024, entitled "Object Selection Method, Selection Box Presentation Method, Device and Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of software technology, and in particular to an object selection method, a selection box presentation method, a device, and a medium. Background Technology
[0003] Currently, the technology and products of smart head-mounted devices are gradually maturing. These devices are becoming increasingly intelligent, lightweight, and low-cost, and are gradually integrating into people's lives, becoming wearable smart hardware devices in people's daily lives, and even becoming smart tools to help people enhance certain abilities.
[0004] For example, in scenarios such as travel, shopping, and home use, when users wear smart glasses and use artificial intelligence (AI) assistants, the cameras on smart glasses are typically ultra-wide-angle cameras, capturing a large number of objects and a large amount of information. When users want to select a specific object within their field of view or further narrow down the selection, they usually use voice interaction. This makes the interaction process cumbersome, inefficient, and results in a poor user experience.
[0005] For example, when using electronic devices such as smartphones, tablets, and personal computers, users typically select content displayed on the screen (including text and images) using their fingers. Smart glasses can also serve as a content selection / interaction device for these devices, allowing users to select screen content without using their hands. However, when users do select screen content, they usually rely on voice interaction, which is cumbersome and inefficient, resulting in a poor user experience. Summary of the Invention
[0006] Some embodiments of this application provide an object selection method, a selection box presentation method, an apparatus, and a medium, which enable users to easily, directly, and accurately select the range / object they wish to select using a selection box. The following describes this application from multiple aspects, and the embodiments and beneficial effects described below can be referenced interchangeably.
[0007] In a first aspect, embodiments of this application provide an object selection method applied to a smart head-mounted device, the smart head-mounted device being equipped with lenses, the method comprising:
[0008] Upon receiving a first instruction, a first selection box is displayed on the lens of the smart head-mounted device; the lens within the first selection box displays a first transmittance, and the lens outside the first selection box displays a second transmittance, wherein the first transmittance and the second transmittance are different;
[0009] In response to the user's second instruction, determine the object selected by the first selection box.
[0010] In some implementations, the object selected by the first selection box is the object that the user sees through the first selection box when looking out through the lenses of the smart head-mounted device.
[0011] According to the embodiments of this application, when a user issues a first command (e.g., a selection box activation command), the smart head-mounted device (e.g., smart glasses) can display a selection box on the lenses. When the user needs to select an object within their field of vision, the selection box can serve as a selection method; when the selection box frames an object, it indicates that the object is selected. By displaying the selection box on the lenses of the smart head-mounted device, the user can accurately select the desired range / object, precisely expressing their intent. The processing device, such as the smart head-mounted device, can also accurately obtain the range / object seen by the user within the selection box, thereby accurately understanding the user's intent.
[0012] In some implementations, the first transmittance is greater than the second transmittance.
[0013] According to the embodiments of this application, since the scenery in areas with higher transmittance appears clearer and the scenery in areas with lower transmittance appears relatively less clear, by making the transmittance of the lens inside the selection frame higher than that of the lens outside the selection frame, the user can clearly see the range / object to be selected through the selection frame.
[0014] In some implementations, a first selection box is presented at the center of the lens, and the shape of the first selection box is rectangular, circular, elliptical, horizontal or vertical.
[0015] In some implementations, the method further includes:
[0016] In response to a third instruction from the user, at least one of the following is adjusted: the position of the first selection box on the lens, the shape of the first selection box, and the size of the first selection box.
[0017] According to the embodiments of this application, by adjusting parameters such as the position, shape, and size of the selection box, the selection box can adapt to different users' face shapes, interpupillary distances, and different user habits, thereby further improving the user experience.
[0018] In some implementations, a camera is provided on the smart head-mounted device;
[0019] Determine the objects selected by the first selection box, including:
[0020] Determine the first field of view range formed by the first selection box with the first preset position as the reference point, and the second field of view range of the camera;
[0021] Determine the mapping relationship from the first field of view to the second field of view;
[0022] Based on the mapping relationship, a second image corresponding to the object selected by the first selection box is determined from the first image captured by the camera.
[0023] In some implementations, the first preset position is used to identify the position of the user's eyes when the user is wearing the smart head-mounted device; the first field of view is the range seen by the user through the first selection box when looking outward through the lenses of the smart head-mounted device.
[0024] According to the embodiments of this application, when a user issues an object selection command, the smart head-mounted device can respond to the object selection command and accurately capture the object selected by the user through the selection box from the image captured by the camera through algorithms such as field of view (FOV) mapping.
[0025] In some implementations, determining the first field of view range formed by the first selection box with a first preset position as a reference point includes:
[0026] Get the first selection box parameters, which include the size and shape of the first selection box;
[0027] Obtain a first distance, which is used to characterize the distance between a first preset position and a first plane where the lenses of the smart head-mounted device are located;
[0028] The first field of view is determined based on the parameters of the first selection box and the first distance.
[0029] In some implementations, determining the mapping relationship from the first field of view to the second field of view includes:
[0030] By center-aligning the first field of view to the second field of view, a mapping relationship is obtained.
[0031] or,
[0032] Get the position of the first selection box and the position of the camera;
[0033] Based on the position of the first selection box and the position of the camera, determine the relative position information between the first selection box and the camera;
[0034] The mapping position of the first field of view range within the second field of view range is determined based on relative position information;
[0035] The mapping relationship is obtained by mapping the first field of view to the mapping position in the second field of view.
[0036] According to the embodiments of this application, FOV mapping using a center-aligned method can improve system processing efficiency. By utilizing the relative positional relationship between the selection box and the camera to determine the mapping position of the first field of view range within the second field of view range, the accuracy of the mapping relationship can be improved, and the error in determining the object selected by the selection box can be reduced.
[0037] In some implementations, the method further includes:
[0038] Processing the object selected by the first selection box includes:
[0039] The object selected by the first selection box is identified, its type is obtained and output; and / or,
[0040] Analyze the object selected by the first selection box, obtain the object's feature information and output it; and / or,
[0041] The second instruction is responded to based on the object selected in the first selection box.
[0042] According to the embodiments of this application, the intelligent head-mounted device identifies and analyzes the object selected by the selection box (e.g., AI analysis), and interacts with the user based on the results. This greatly simplifies the interaction process, making it more direct and intuitive, and significantly improving accuracy, thus enhancing interaction efficiency and user experience. Furthermore, by further narrowing the selection and analysis scope through the selection box, the computational power requirement can be reduced, ensuring the system's real-time performance requirements.
[0043] Secondly, embodiments of this application provide a selection box presentation method applied to a smart head-mounted device, the smart head-mounted device being equipped with lenses and a camera, the method comprising:
[0044] Acquire a third image captured by the camera;
[0045] Identify the third image to obtain the image region where the object of interest is located in the third image;
[0046] A second selection box is determined based on the image region where the object of interest is located; wherein, when the second selection box is presented on the lens of the smart head-mounted device, the field of view range formed with the first preset position as the reference point can cover at least a part of the field of view range corresponding to the image region;
[0047] A second selection box is displayed on the lens of the smart head-mounted device.
[0048] According to the embodiments of this application, a smart head-mounted device (such as smart glasses) can perform image analysis on images captured by a camera, automatically identify objects of interest to the user, determine a selection box on the lenses based on the image area where the object of interest is located, and finally present it. In this way, the smart head-mounted device can proactively label objects that the user may be interested in and recommend them to the user. The user can quickly select the object based on the recommendation, thus improving the user experience.
[0049] In some implementations, determining a second selection box based on the image region where the object of interest is located includes:
[0050] Based on the image region where the object of interest is located and the second field of view of the camera, determine the third field of view range corresponding to the image region;
[0051] The parameters of the second selection box are determined based on at least the third field of view.
[0052] According to the embodiments of this application, the smart head-mounted device determines the corresponding field of view range based on the image region where the object of interest is located, and calculates the selection box parameters presented on the lens of the selection box forming the field of view range, thereby accurately determining the selection box corresponding to the image region where the object of interest is located.
[0053] In some implementations, the second selection box parameters include the size of the second selection box and the position of the second selection box;
[0054] The second selection box parameters are determined based on at least the third field of view, including:
[0055] Obtain a first distance, which is used to characterize the distance between a first preset position and a first plane where the lenses of the smart head-mounted device are located;
[0056] The size of the second selection box is determined based on the third field of view range and the first distance;
[0057] The position of the second selection box is determined based on the position of the third field of view within the second field of view.
[0058] In some implementations, the second selection box parameter further includes the shape of the second selection box;
[0059] Determine the second selection box parameters of the second selection box based at least on the third field of view angle range, including:
[0060] Determine the shape feature information of the third field of view angle range;
[0061] Determine the second selection box shape based on the shape feature information of the third field of view angle range;
[0062] Alternatively, determine the shape feature information of the object of interest; determine the second selection box shape based on the shape feature information of the object of interest;
[0063] Alternatively, determine that the second selection box shape is a preset shape.
[0064] According to the embodiments of the present application, the intelligent head-mounted device can determine the shape of the selection box presented on the lens according to the shape features of the object of interest or the shape features of the field of view angle range formed by the selection box, so as to present the selection box in the most suitable shape, further improving the user experience.
[0065] In a third aspect, an embodiment of the present application provides a method for presenting a selection box, which is applied to an electronic device. The method includes:
[0066] Display a first interface;
[0067] Display a fourth selection box in the first interface;
[0068] Detect that the third selection box presented on the lens of the intelligent head-mounted device changes from a first position to a second position relative to the electronic device;
[0069] Adjust the display position of the fourth selection box based on the position change of the third selection box.
[0070] According to the embodiments of the present application, when the user issues a selection box activation instruction, the intelligent head-mounted device (such as smart glasses) can present a selection box on the lens for linkage and cooperation with other electronic devices to select the content on the display screen of the electronic device, liberating the user's hands. A user-visible selection box can also be presented on the display screen of the electronic device for the user to determine the range / object selected by the selection box. When the user adjusts the position of the selection box on the lens of the intelligent head-mounted device relative to the electronic device by turning the head, the electronic device can track the selection box on the intelligent head-mounted device in real time, and through coordinate mapping, map the position change of the selection box on the intelligent head-mounted device to the plane where the display screen of the electronic device is located, so as to adjust the selection box on the display screen of the electronic device and change the object in the display screen of the electronic device selected by the selection box. This interaction method is more intuitive and does not require the user to operate with both hands, and can be applied to barrier-free scenarios.
[0071] Furthermore, by directly tracking the selection boxes on the lenses of the smart head-mounted device, the tracking method offers higher accuracy and greater adaptability compared to direct eye tracking due to the obvious feature points of the selection boxes. It eliminates the need for eye calibration and adaptation for different user groups. In some implementations, selection boxes on both lenses can be tracked simultaneously, thereby improving tracking accuracy and stability, and further enhancing the accuracy and stability of selecting content on electronic device displays using these selection boxes.
[0072] In some implementations, a fourth selection box is displayed in the first interface, including:
[0073] A fourth selection box is displayed at a second preset position on the display screen of the electronic device.
[0074] In some implementations, the fourth selection box is displayed corresponding to the third selection box;
[0075] The fourth selection box is displayed in the first interface, including:
[0076] Acquire a fourth image, which includes a third selection box presented on the lenses of the smart head-mounted device;
[0077] The first position of the third selection box relative to the electronic device is determined based on the fourth image;
[0078] The third position of the fourth selection box on the display screen of the electronic device is determined based on the first position of the third selection box;
[0079] A fourth selection box is displayed in a third position on the screen of the electronic device.
[0080] According to the embodiments of this application, the initial position of the selection box on the display screen of the electronic device can be set at a preset position, or it can correspond to the position of the selection box on the smart head-mounted device.
[0081] In some implementations, the display position of the fourth selection box is adjusted based on the change in the position of the third selection box, including:
[0082] Determine the amount of position change of the third selection box from the first position to the second position;
[0083] The second position change of the fourth selection box on the display screen of the electronic device is determined based on the first position change.
[0084] The fourth position of the fourth selection box on the display screen of the electronic device is determined based on the change in the second position.
[0085] Move the fourth selection box to the fourth position.
[0086] According to an embodiment of this application, the electronic device can accurately obtain the position change of the selection box on the display screen by mapping the position change of the selection box on the smart head-mounted device onto the plane where the electronic device's display screen is located.
[0087] In some implementations, determining the second position change of the fourth selection box on the display screen of the electronic device based on the first position change includes:
[0088] The change in the first position is multiplied by a preset scaling factor to obtain the change in the second position.
[0089] In some implementations, the method further includes:
[0090] Define the object selected by the fourth selection box;
[0091] Perform operations on the objects selected by the fourth selection box.
[0092] According to the embodiments of this application, the electronic device can directly operate on the object selected by the selection box on the display screen of the electronic device, further improving the interaction efficiency between the smart head-mounted device and the electronic device, and freeing the user's hands.
[0093] Fourthly, embodiments of this application provide an electronic device, including: a memory for storing instructions executable by one or more processors of the electronic device; and a processor, which, when executing the instructions in the memory, causes the electronic device to perform the method provided by any embodiment of the first, second, or third aspect of this application. The beneficial effects achievable by the fourth aspect can be referred to the beneficial effects of any embodiment of the first, second, or third aspect of this application, and will not be repeated here.
[0094] Fifthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method provided by any embodiment of the first, second, or third aspect of this application. The beneficial effects achievable through the fifth aspect can be found in the beneficial effects of any embodiment of the first, second, or third aspect of this application, and will not be repeated here.
[0095] Sixthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the method provided in any of the first, second, or third embodiments of this application. The beneficial effects achievable through the sixth aspect can be found in the beneficial effects of any of the first, second, or third embodiments of this application, and will not be repeated here. Attached Figure Description
[0096] Figure 1a is a schematic diagram of an object selection process provided in some embodiments of the prior art;
[0097] Figure 1b is a schematic diagram of an object selection process provided in some other embodiments of the prior art;
[0098] Figure 1c is a schematic diagram of an object selection process provided in some other embodiments of the prior art;
[0099] Figure 1d is a schematic diagram of an object selection process provided in some other embodiments of the prior art;
[0100] Figure 2a is a schematic diagram of an application scenario provided by an implementation example of this application;
[0101] Figures 2b to 2e are schematic diagrams of the selection box provided in one implementation example of this application;
[0102] Figures 3a and 3b are schematic diagrams of an implementation example of the electrochromic technology provided in this application;
[0103] Figures 4a and 4b are schematic diagrams of application scenarios provided by another implementation example of this application;
[0104] Figure 5a is a software structure block diagram of an electronic device provided in an embodiment of this application;
[0105] Figure 5b is a schematic diagram of the system architecture of a smart head-mounted device provided in an embodiment of this application;
[0106] Figure 5c is a schematic diagram of the system architecture of a smart head-mounted device provided in another embodiment of this application;
[0107] Figure 5d is a schematic diagram of the system architecture of an electronic device provided in an embodiment of this application;
[0108] Figure 6 is a flowchart of an object selection method provided in an embodiment of this application;
[0109] Figure 7a is a vertical view of a user wearing smart glasses according to an embodiment of this application;
[0110] Figure 7b is a horizontal view of a user wearing smart glasses according to an embodiment of this application;
[0111] Figure 7c is a schematic diagram of the actual scene seen by the user before and after the selection box is opened according to an embodiment of this application;
[0112] Figure 8a is a flowchart of an object selection method provided in another embodiment of this application;
[0113] Figure 8b is a schematic diagram illustrating the determination of the FOV1 range according to an embodiment of this application;
[0114] Figure 8c is a schematic diagram showing the mapping of the selection box FOV1 range to the camera FOV2 range according to an embodiment of this application;
[0115] Figure 8d is a schematic diagram of a region image cropping provided in an embodiment of this application;
[0116] Figure 8e is a schematic diagram of region image cropping provided in another embodiment of this application;
[0117] Figure 9a is a flowchart of an object selection method provided in another embodiment of this application;
[0118] Figure 9b is a schematic diagram of an object selection process provided in an embodiment of this application;
[0119] Figure 9c is a schematic diagram of an object selection process provided in another embodiment of this application;
[0120] Figure 9d is a schematic diagram of an object selection process provided in another embodiment of this application;
[0121] Figure 9e is a schematic diagram of an object selection process provided in another embodiment of this application;
[0122] Figure 10a is a flowchart of a selection box presentation method provided in an embodiment of this application;
[0123] Figure 10b is a schematic diagram of an image captured by a camera according to an embodiment of this application;
[0124] Figure 10c is a schematic diagram of the coordinates of the selection box in the plane where the lens is located according to an embodiment of this application;
[0125] Figure 11a is a flowchart of a selection box presentation method provided in another embodiment of this application;
[0126] Figures 11b and 11c are schematic diagrams of the Z1 plane and Z2 plane provided in one embodiment of this application;
[0127] Figure 12a is a schematic diagram of the application page of a reading application provided in an embodiment of this application;
[0128] Figure 12b is a schematic diagram of a system desktop provided in one embodiment of this application;
[0129] Figure 12c is a schematic diagram of the application page of a video application provided in an embodiment of this application;
[0130] Figure 13 is a block diagram of an electronic device provided in one embodiment of this application;
[0131] Figure 14 is a block diagram of a System on Chip (SOC) provided in one embodiment of this application. Detailed Implementation
[0132] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0133] This application's embodiments can be applied to scenarios where a smart head-mounted device is used alone to select objects within the user's field of vision. Alternatively, it can be applied to scenarios where a smart head-mounted device works in conjunction with an electronic device to select content displayed on the screen of an electronic device. The following describes this application's embodiments in detail, using smart glasses as an example of a smart head-mounted device.
[0134] In this scenario, existing technologies typically rely on voice interaction for selection, which is cumbersome, inefficient, and results in a poor user experience. The following examples illustrate the inconvenience of voice interaction.
[0135] In existing technologies, when users use smart glasses for real-world recognition, they can interact via voice commands. For example, by saying "Please analyze the scene in front of you," the smart glasses can be controlled to identify specific objects within the user's field of vision. Alternatively, an eye-tracking module can be integrated into the smart glasses to track the user's gaze point and determine the direction of the user's gaze, thereby selecting specific objects within the user's field of vision. Simultaneously, a display module needs to be added to the smart glasses to display the selected object to the user, allowing the user to clearly see that the object they are looking at has been recognized by the system.
[0136] For example, referring to Figure 1a of the reference manual, in a museum guided tour scenario, after wearing the smart glasses, the user can say, "Please describe the third porcelain piece from the left in front of me." Upon receiving this voice, the smart glasses first need to analyze and segment the overall image to accurately locate the third porcelain piece from the left. If it can be accurately located, the smart glasses can identify the porcelain as blue and white porcelain and output, "This porcelain is Ming Dynasty blue and white porcelain." If it cannot be accurately located, the smart glasses need to interact with the user for further confirmation, for example, outputting, "Excuse me, I'd like to confirm again which porcelain piece you want to know about. Do you want to know about the shortest one?" The user needs to confirm via voice, and if the user feels the selected object is incorrect, they need to continue using voice description.
[0137] For example, referring to Figure 1b of the reference specification, in a shopping scenario, after the user puts on the smart glasses, they can issue a voice command saying, "Please add the tent in front of me to your shopping cart." Since there may be multiple tents in front of the user, the smart glasses may not be able to accurately locate the tent the user needs after receiving this voice command. The smart glasses need to interact with the user to further confirm, for example, by outputting a voice command saying, "Excuse me, I'd like to confirm again which tent you need." The user then needs to issue another voice command saying, "It's the rectangular tent in front of me," to provide further details.
[0138] However, this voice-based control and interaction method is inefficient and lacks intuitiveness and accuracy. Integrating eye-tracking modules into smart glasses is costly, consumes a lot of power, and is difficult to guarantee in terms of accuracy. Moreover, eye-tracking technology tracks points, not a range of objects.
[0139] In existing technologies, when users use smart glasses to work in conjunction with electronic devices, an eye-tracking module and a display module can be integrated into the smart glasses. The screen of the electronic device can be projected onto the smart glasses, allowing reverse control of the electronic device's display screen via eye movements, gestures, or a controller. Alternatively, the user's eyes can be tracked using the front-facing camera of the electronic device, and eye-tracking technology can be used to track the user's gaze point to determine the position of the user's gaze on the screen, thereby selecting the content displayed on the electronic device's screen (unrelated to the smart glasses themselves). Alternatively, the electronic device can also be selected or controlled via voice. For example, a user can say "select row xx, column xx," and the smart glasses' microphone can pick up the user's voice, recognize the voice, and determine the content displayed on the electronic device's screen that the user wants to select.
[0140] For example, referring to Figure 1c of the reference specification, when a user uses smart glasses in conjunction with an electronic device, eye-tracking technology can be used to track the user's gaze point, determining that the user is looking at the middle left side of the screen, thereby determining that the content the user wants to select is the content in rows 6 to 9 on the left. Referring to Figure 1d of the reference specification, when a user uses smart glasses in conjunction with an electronic device, the smart glasses' microphone can pick up the user's voice, "Select the first app icon in the second row," thereby determining that the content the user wants to select is the first app icon in the second row, i.e., the icon for the Sports & Health app. If it cannot be accurately found, the smart glasses need to interact with the user for further confirmation, for example, it can output the voice message, "Excuse me, I would like to confirm again the app you need to select. Do you want to select the Sports & Health app?" The user needs to confirm by voice, and if the user feels that the selected object is incorrect, they need to continue to describe it by voice.
[0141] However, integrating a display module into smart glasses is costly, consumes a lot of power, and increases the weight of the glasses. Furthermore, integrating an eye-tracking module into smart glasses is costly, consumes a lot of power, and its accuracy is difficult to guarantee. Tracking the user's eyes using the front-facing camera of an electronic device has poor adaptability. Additionally, eye-tracking technology tracks points, not a range of objects. Moreover, voice control is inefficient and lacks intuitiveness and accuracy.
[0142] Therefore, a more convenient and direct interaction / selection method is needed to accurately express user commands.
[0143] Therefore, this application provides an object selection method, a selection box presentation method, a device, and a medium that enable users to easily, directly, and accurately select the range / objects they want to select through the selection box.
[0144] In one embodiment, the method can be used in scenarios such as travel, shopping, and home use, to select specific objects within the user's field of vision, or to further narrow down the selected range, while the user is wearing smart glasses.
[0145] In this scenario, as shown in Figure 2a, the smart glasses can display a selection box 110 on the lens. The selection box 110 can be one or multiple. The selection box 110 can be displayed as a monocular selection box as shown in Figure 2a(1) or as a binocular selection box as shown in Figure 2a(2). That is, the selection box 110 can be displayed on only one lens or on both lenses. This application embodiment does not specifically limit the number or display method of the selection box 110.
[0146] To ensure the selection box 110 is clearly visible to the user, the lenses within and outside the selection box 110 can have different transmittance. The lenses outside the selection box 110 can be a portion of the lenses immediately adjacent to it. For example, as shown in Figures 2b(1) and (3), the lenses within the selection box 110 can have higher transmittance, while the lenses outside can have lower transmittance; that is, the transmittance of the lenses within the selection box 110 is greater than that of the lenses outside. Alternatively, as shown in Figure 2b(2), the lenses within the selection box 110 can have lower transmittance, while the lenses outside can have higher transmittance; that is, the transmittance of the lenses within the selection box 110 is less than that of the lenses outside. In this case, the area with higher transmittance appears clearer, while the area with lower transmittance appears relatively less clear.
[0147] The shape of the selection box 110 can be preset or adjusted by the user as needed, as long as it is clearly visible to the user. For example, as shown in FIG2b, the shape of the selection box 110 can be set to include, but is not limited to, rectangles (rectangles mentioned in this embodiment can include rectangles with rounded corners), circles, ellipses, horizontal bars, and vertical bars, etc., and this embodiment does not specifically limit this. When the smart glasses present a dual-lens selection box, the shapes of the selection boxes 110 on the two lenses of the smart glasses can be the same or different. For example, the shapes of the selection boxes 110 on both lenses of the smart glasses can both be rectangles as shown in FIG2b, or one can be a rectangle as shown in FIG2b and the other an ellipse as shown in FIG2b, and this embodiment does not specifically limit this.
[0148] In some embodiments, the smart glasses can also adaptively adjust the transmittance of the lenses based on the actual ambient brightness (e.g., indoor, outdoor, different lighting conditions, etc.) collected by the sensors to achieve the best selection box rendering effect. For example, as shown in FIG2c, when the ambient brightness is low, the overall transmittance of the lenses can be higher; when the ambient brightness is high, the overall transmittance of the lenses can be lower.
[0149] The position of the selection box 110 can be preset, for example, the initial position of the selection box 110 can be set at the center of the lens. Furthermore, since different users have different face shapes, interpupillary distances, and habits, the position of the selection box 110 can also be adjusted by the user as needed, so that the selection box 110 is presented in the most comfortable position for the user's viewing. For example, as shown in FIG2d, the position of the selection box 110 can be set at, but is not limited to, the center of the lens, the left side of the lens, the right side of the lens, the upper part of the lens, and the lower part of the lens, etc., and this application embodiment does not specifically limit this.
[0150] The size of the selection box 110 can be preset or adjusted by the user as needed to facilitate the selection of objects of different sizes in various scenarios. This embodiment does not impose specific limitations on this. For example, as shown in FIG2e, the size of the selection box 110 can be set to different sizes according to different user needs. In one embodiment, since the selection box 110 is very close to the user's eyes, it is usually set to a relatively small size so that the user can clearly observe it.
[0151] In this embodiment, an electrochromic selection box 110 can be displayed on the lens using electrochromic technology. Specifically, an electrochromic film can be attached to the lens of the smart glasses (either the inner or outer side of the lens). The transmittance of the electrochromic film is controlled in sections by a driving circuit (Integrated Circuit, IC) to display the selection box 110 on the side of the glasses. Electrochromic technology has the advantages of easy integration with lenses, low cost, and low power consumption. Furthermore, the user can clearly observe the selection box 110 and use it to select objects and express user intentions / instructions.
[0152] Referring to Figure 3a in the reference specification, electrochromic technology uses dye-based liquid crystals to form a host-guest system. Dye molecules align with the liquid crystal molecules, and the rotation of the liquid crystal molecules causes the dye molecules to rotate. Specifically, when the applied voltage changes, the liquid crystal molecules flip, simultaneously causing the dye molecules to rotate, thus exhibiting different transmittances. The dichroic ratio of the dye is N = D. / / / D ⊥ This determines the range of the brightest and darkest areas, where D / / =-logT / / D ⊥ =-logT ⊥ T / / T represents the platform's transmittance. ⊥ This represents the transmittance in the vertical state. The more orderly the dye liquid crystal arrangement and the greater the dichroism ratio, the wider the light modulation range can be obtained.
[0153] Referring to Figure 3b in the reference specification, by using electrodes for zoned power supply, independent control of different pixels can be achieved. By applying different voltages to the pixels within and outside the selection frame 110, the transmittance of the electrochromic film within and outside the selection frame 110 can be made different. The size and shape of the pixels can be preset according to actual needs, and the electrodes can be transparent electrodes, or single-layer, double-layer, or multi-layer electrodes.
[0154] In some possible embodiments, other optical display methods can also be used to present the selection box 110, such as including but not limited to Micro LED combined with waveguide optics, Micro OLED combined with BB optics, Laser Beam Scanning (LBS) technology combined with waveguide optics, etc. The embodiments of this application do not specifically limit the implementation method of the selection box 110.
[0155] It should be noted that users can adjust the position, size, shape, and other information of the selection box 110 through touch, buttons, and knobs. This application embodiment does not specifically limit the adjustment method. For example, the selection box 1 can be enlarged or reduced by clicking the touchpad on the smart glasses (e.g., on the temple), pressing the button on the smart glasses (e.g., on the temple), or rotating the knob on the smart glasses (e.g., on the temple).
[0156] In this scenario, as shown in Figure 2a, the smart glasses can also be equipped with one or more cameras 120. The smart glasses can acquire a field of view (FOV) through the cameras 120. This FOV includes the horizontal field of view (HFOV) and the vertical field of view (VFOV). FOV is defined as the angle formed by the two edges of the maximum range through which the image of the target object can pass through the lens, with the camera of the optical instrument as the vertex. The size of the field of view determines the field of view of the optical instrument; the larger the field of view, the larger the field of view, and the smaller the optical magnification. Simply put, if the target object is beyond this angle, it will not be captured by the lens.
[0157] In this embodiment, the selection box 110 can be used to select objects within the user's field of vision or to express the user's intentions or instructions. When the user wears the smart glasses and issues a selection box activation command, the smart glasses can respond to the command by displaying the selection box 110 on the lenses. When the user wants to perform a selection operation, they can issue an object selection command. The smart glasses can respond to the object selection command by segmenting the image corresponding to the object selected by the selection box 110 from the image acquired by the camera 120. The smart glasses can also further analyze the segmented image to determine the object selected by the selection box 110, and then respond to the user based on the determined object.
[0158] In another embodiment, the method can be used to select content such as text and icons displayed on the screen of an electronic device when the smart glasses work together with the electronic device, i.e., when a user is wearing smart glasses and looking at the screen of the electronic device.
[0159] In this scenario, as shown in Figures 4a and 4b, the smart glasses can display a selection box 210 on the lens, and the electronic device can also display the corresponding selection box 220 on the display screen in a mapped manner. The number of selection boxes 210 and 220 can be one or more. The selection box 210 can be presented as a monocular selection box as shown in Figure 4a, or as a binocular selection box as shown in Figure 4b. That is, the selection box 210 can be presented on only one lens or on both lenses. This application embodiment does not specifically limit the number and presentation method of the selection boxes 210, nor the number of selection boxes 220.
[0160] The size of the selection box 210 can be preset or adjusted by the user as needed; this embodiment does not specifically limit this. In one embodiment, when the smart glasses work in conjunction with an electronic device, the selection box 210 on the smart glasses can be set relatively large, for example, the internal area of the selection box 210 can occupy more than 70% of the lens area. Because the selection box 210 is very close to the user's eyes, the user can hardly observe the selection box 210 on the smart glasses, and it is also convenient for the camera of the electronic device to track the selection box 210. At this time, the selection box 220 displayed on the electronic device screen can be seen by the user, that is, the user can only observe the selection box 220 on the electronic device screen.
[0161] The shape, position, and size of the selection box 220 can be preset, adjusted by the user as needed, and adaptively adjusted according to the object selected by the selection box 220. This embodiment of the application does not limit these aspects. For example, the initial position of the selection box 220 can be set at the center of the electronic device's display screen, and the initial shape of the selection box 220 can be set to a rectangle, circle, or ellipse, etc.
[0162] It should be noted that other contents of selection box 210 can be referred to the relevant contents of selection box 110 in the embodiments shown in Figures 2a to 3b, and will not be repeated here in the embodiments of this application.
[0163] In this scenario, a camera 230 may be installed on the side where the display screen of the electronic device is located. The smart glasses may or may not have a camera installed, or may have one or more cameras installed. This application embodiment does not specifically limit this.
[0164] In this embodiment, selection boxes 210 and 220 can be used to select content displayed on the screen of an electronic device. When a user wears the smart glasses and issues a selection box activation command, the smart glasses can respond to the command and display selection box 210 on the lenses. The smart glasses can also send a selection box activation request to the electronic device, which can respond to the request and display the corresponding selection box 220 on the display screen. Subsequently, when the user adjusts the position of selection box 210 relative to the electronic device by rotating their head, the camera 230 of the electronic device can track the selection box 210 on the smart glasses to determine the position change of selection box 210 relative to the electronic device. Then, the position change of selection box 210 can be mapped to the plane where the display screen of the electronic device is located to obtain the position change of selection box 220 on the plane where the display screen of the electronic device is located. Finally, selection box 220 is updated according to the position change of selection box 220, and the object selected by the updated selection box 220 is determined from the content displayed on the display screen, and then the user is responded to based on the determined object.
[0165] As can be seen, by setting a selection box, users can intuitively and accurately select the range / object they want to select, accurately expressing their intentions. Smart glasses and other processing devices can also accurately obtain the range / object seen in the selection box, thus accurately understanding the user's intentions. This method of interacting with users based on selection boxes greatly simplifies the interaction process, making it more direct and intuitive, significantly improving accuracy, and enhancing interaction efficiency and experience.
[0166] According to one embodiment of this application, when a user issues a selection box activation command, a smart head-mounted device (e.g., smart glasses) can display a selection box on the lenses. When a user needs to select an object within their field of vision, the selection box can serve as a selection method; when the selection box frames an object, it indicates that the object is selected. Furthermore, when a user issues an object selection command, the smart head-mounted device can respond to the command and accurately capture the object selected by the user through the selection box from the image captured by the camera using algorithms such as FOV mapping. By setting a selection box, users can precisely select the desired range / object, accurately expressing their intent. The processing device, such as the smart head-mounted device, can also accurately obtain the range / object seen by the user within the selection box, thereby accurately understanding the user's intent.
[0167] According to another embodiment of the present application, a smart head-mounted device (such as smart glasses) can perform image analysis on the images captured by a camera, automatically identify the user's interested objects, determine the selection frame presented on the lens according to the image area where the interested object is located, and finally present it. In this way, the smart head-mounted device can actively mark the objects that the user may be interested in and recommend them to the user. The user can quickly select the object according to the recommendation, which can improve the user experience.
[0168] According to another embodiment of the present application, when the user issues a selection frame activation instruction, the smart head-mounted device (such as smart glasses) can present a selection frame on the lens for linkage and cooperation with other electronic devices to select the content on the display screen of the electronic device, liberating the user's hands. A selection frame visible to the user can also be presented on the display screen of the electronic device for the user to determine the range / object selected by the selection frame. When the user adjusts the position of the selection frame on the lens of the smart head-mounted device relative to the electronic device by turning the head, the electronic device can use the camera to track the selection frame on the smart head-mounted device in real time, and through coordinate mapping, map the position change of the selection frame on the smart head-mounted device to the plane where the display screen of the electronic device is located, so as to adjust the selection frame on the display screen of the electronic device and change the object selected by the selection frame in the display screen of the electronic device. This interaction method is more intuitive and does not require the user to operate with both hands, and can be applied to barrier-free scenarios.
[0169] Moreover, by directly tracking the selection frame on the lens of the smart head-mounted device, since the feature points of the selection frame are obvious, compared with directly performing eye movement tracking, this tracking method has higher accuracy and stronger adaptability, does not require calibration of the user's eyeballs, and does not require adaptation for different populations. In some embodiments, the selection frames on the two lenses can be tracked simultaneously, so as to improve the accuracy and stability of the tracking, and further improve the accuracy and stability of selecting the content on the display screen of the electronic device through the selection frame.
[0170] The embodiments of the present application do not limit the forms of the smart head-mounted device and the electronic device. The smart head-mounted device can be a smart glasses, a smart helmet, an augmented reality (AR) / virtual reality (VR) glasses and other devices. The electronic device can be a mobile phone, a tablet computer, a notebook computer, a cellular phone, a personal computer (PC), a personal digital assistant (PDA), a television, a smart screen, a vehicle-mounted device (such as a car machine, a vehicle-mounted navigator), etc. The system installed in the electronic device can be Android, IOS, HarmonyOS, etc., which is not limited here.
[0171] The software structure of the smart head-mounted device and the electronic device interacting with the smart head-mounted device according to embodiments of this application is described below with reference to FIG5a.
[0172] Figure 5a is a software structure block diagram of an electronic device provided in one embodiment of this application. As shown in Figure 5a, the layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer (APP), the application framework layer (APP Framework), the Android Runtime and system libraries, and the kernel layer.
[0173] The application layer can include a series of application packages.
[0174] As shown in Figure 5a, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS. The application layer may also include other applications (APPs) besides those shown in Figure 5a, such as games, settings, and social networking, which will not be described in detail in this embodiment.
[0175] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0176] As shown in Figure 5a, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0177] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0178] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0179] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0180] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).
[0181] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0182] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0183] The Android Runtime consists of core libraries and a virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0184] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0185] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0186] System libraries can include multiple functional modules. For example: Surface Manager, Media Libraries, 3D graphics processing libraries (e.g., OpenGLES), 2D graphics engines (e.g., SGL), and selection box processing algorithm engines, etc.
[0187] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0188] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0189] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0190] A 2D graphics engine is a graphics engine for 2D drawing.
[0191] The selection box processing algorithm engine can be used to execute operations based on images of selection boxes displayed on the lenses of smart head-mounted devices. It can track the position of the selection boxes on the lenses of the smart head-mounted devices in real time and, through coordinate mapping, map the position (or position changes) of the selection boxes on the smart head-mounted devices to the plane where the electronic device's display screen is located, thereby determining the display position of the selection boxes on the electronic device's display screen. The selection box processing algorithm engine can also be used to send the display position of the selection boxes on the electronic device's display screen to the window manager, thereby enabling the selection boxes to be displayed on the electronic device's display screen.
[0192] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0193] It should be noted that, in the embodiments of this application, the smart head-mounted device may include some or all of the structures in the software structure shown in Figure 5a, and the embodiments of this application do not specifically limit this.
[0194] The system architecture of the smart head-mounted device and the electronic device interacting with the smart head-mounted device according to embodiments of this application is described below with reference to Figures 5b to 5d.
[0195] Referring to Figure 5b of the specification, Figure 5b illustrates a schematic diagram of the system architecture of a smart head-mounted device according to an embodiment of this application. As shown in Figure 5b, the system architecture of the smart head-mounted device may include a camera 501, a processor 503, a control module 505, a communication module 507, a driver IC 509, and an electrochromic film 511. The processor 503 is communicatively connected to the camera 501, the control module 505, the communication module 507, and the driver IC 509, respectively. The driver IC 509 is electrically connected to the electrochromic film 511.
[0196] Camera 501 can capture images within the user's field of view and send the captured images to processor 503. Processor 503 can receive instruction signals sent by control module 505, such as instructions to select objects, open / close selection boxes, or adjust the size of selection boxes. Processor 503 can also process the images captured by camera 501 in response to the received instruction signals and output control signals to driver IC 509 based on the processing results, or it can directly output control signals to driver IC 509 based on the received instruction signals. Control module 505 can receive user control commands, such as instructions to select objects, open / close selection boxes, or adjust the size of selection boxes, and send the corresponding instruction signals to processor 503. Communication module 507 can communicate with other electronic devices such as mobile phones. Driver IC 509 can output different voltages and currents in response to control signals sent by processor 503 to control the transmittance of the electrochromic film 511 on the lenses of the smart head-mounted device, and can perform regional control.
[0197] In some embodiments, as shown in FIG5c, the system architecture of the smart head-mounted device may not include the camera 501, but only include the processor 503, control module 505, communication module 507, driver IC 509, and electrochromic film 511. The specific details of the processor 503, control module 505, communication module 507, driver IC 509, and electrochromic film 511 can be found in the embodiment shown in FIG5b, and will not be repeated here.
[0198] Referring to Figure 5d in the specification, Figure 5d illustrates a schematic diagram of the system architecture of an electronic device according to an embodiment of this application. As shown in Figure 5d, the system architecture of the electronic device may include a camera 513, a processor 515, a communication module 517, a driver IC 519, and a display screen 521. The processor 515 is communicatively connected to the camera 513, the communication module 517, and the driver IC 519, respectively, and the driver IC 519 is electrically connected to the display screen 521.
[0199] Camera 513 can be used to capture images and send the captured images to processor 515. Processor 515 can process the images captured by camera 513 and output control signals to driver IC 519 based on the processing results. Communication module 517 can be used to communicate with smart head-mounted devices. Driver IC 519 can respond to the control signals sent by processor 515 to control the display screen 521 of the electronic device to drive the display of selection boxes on the display screen 521 for easy observation by the user.
[0200] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device or smart head-mounted device. In other embodiments of this application, the electronic device or smart head-mounted device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0201] The following describes in detail the specific process of an object selection method and a selection box presentation method provided in one embodiment of this application, based on a scenario of using smart glasses for real-world recognition. The real-world mentioned in this embodiment refers to objects within the user's field of vision when looking out through the lenses of the smart glasses, not objects projected by the smart glasses.
[0202] The following describes in detail an object selection method provided in one embodiment of this application, using smart glasses as an example. The smart glasses are equipped with lenses and a camera, which can be of various types. This embodiment does not specifically limit the camera installed on the smart glasses. Referring to Figure 6, which shows a flowchart of an object selection method provided in one embodiment of this application, the method may include steps S601-S602.
[0203] S601: The smart glasses receive a selection box activation command (as an example of a first command) and display selection box 1 (as an example of a first selection box) on the lens of the smart glasses.
[0204] In this embodiment, after the user puts on the smart glasses, they can issue a selection box activation command. The smart glasses can receive the selection box activation command issued by the user, which instructs the smart glasses to display a selection box 1 on the lens. The selection box activation command can be a voice command, or it can be triggered by clicking the touchpad on the smart glasses, pressing a button on the smart glasses, or rotating a knob on the smart glasses. This embodiment does not specifically limit the method by which the user issues the selection box activation command.
[0205] In this embodiment, after the smart glasses receive a selection box activation command, they can display selection box 1 on the lens in response to the command. The transmittance of the lens within selection box 1 differs from the transmittance of the lens outside selection box 1. In some embodiments, the transmittance of the lens within selection box 1 is greater than the transmittance of the lens outside selection box 1.
[0206] In this embodiment, electrochromic technology can be used to display a selection box 1 on the lens of smart glasses. The specific details regarding the position, size, shape, and presentation method of the selection box 1 can be found in the relevant content of the selection box 110 in the embodiments shown in Figures 2a to 3b, and will not be repeated here.
[0207] The principle of selecting objects using a selection box is explained below with reference to the accompanying drawings. For example, referring to Figures 7a and 7b of the specification, Figures 7a and 7b respectively show a vertical and horizontal view of a user wearing smart glasses according to an embodiment of this application. As shown in Figures 7a and 7b, a camera is mounted on the frame of the smart glasses, and a selection box 1 with high transmittance in the central area and low transmittance in the peripheral area can be formed on the lenses using an electrochromic film. The user can see at least a portion of the actual scene within the camera's field of view through the selection box 1, thereby accurately selecting the desired range / object.
[0208] In this embodiment, the user can quickly and conveniently adjust the selection box to dynamically select different objects / content. When the user needs to adjust at least one of the position, shape, and size of the selection box 1 on the lens, they can issue a selection box adjustment command (as an example of a third command). The smart glasses can obtain the selection box adjustment command issued by the user and adjust at least one of the position, shape, and size of the selection box 1 on the lens in response to the command. The selection box adjustment command can be an adjustment command issued via voice, or it can be an adjustment command triggered by clicking the touchpad on the smart glasses, pressing a button on the smart glasses, or rotating a knob on the smart glasses, etc. This embodiment does not specifically limit the way the user issues the selection box adjustment command.
[0209] For example, users can adjust the size, position, and shape of the selection box 1 using touch, buttons, and knobs. For instance, the selection box 1 can be enlarged or reduced by tapping a touchpad on the smart glasses (e.g., on the temples), pressing a button on the smart glasses (e.g., on the temples), or rotating a knob on the smart glasses (e.g., on the temples). Users can also move the selection box 1 by changing their head position, thus enabling the selection of objects in different areas.
[0210] In this embodiment, when the selection box 1 is displayed on the lens of the smart glasses, it can form a field of view (FOV1) with the human eye (as an example of the first preset position) as the reference point. This represents the field of view seen by the human eye through the selection box 1 when the user looks outward through the lens of the smart glasses. The camera mounted on the smart glasses has a field of view (FOV2). It is evident that the FOV2 of the camera is larger, while the FOV1 seen by the human eye through the selection box 1 is smaller. The FOV1 range can include VFOV1 and HFOV1, and the FOV2 range can include VFOV2 and HFOV2. As shown in Figure 7a, VFOV1 represents the vertical field of view seen by the human eye through the selection box 1, and VFOV2 represents the vertical field of view corresponding to the camera. As shown in Figure 7b, HFOV1 represents the horizontal field of view seen by the human eye through the selection box 1, and HFOV2 represents the horizontal field of view corresponding to the camera.
[0211] Referring to Figure 7c of the specification, Figure 7c shows a real-world view of the user before and after the selection box is opened according to an embodiment of this application. As shown in Figure 7c (1), before the selection box is opened, the user can clearly see all the food on the table through the smart glasses. After the selection box is opened, the smart glasses can display the selection box 1 on the lens through an electrochromic film. The selection box 1 can be displayed in front of the eyes and overlapped with all the food on the table. If the selection box 1 is a rectangular frame 710, and the transmittance of the middle area is high and the transmittance of the outer area is low, then as shown in Figure 7c (2), the user can clearly see some of the food on the table through the selection box 1, while the food seen through other areas is unclear.
[0212] Next, returning to Figure 6, we will explain the object selection method of this application embodiment.
[0213] S602: The smart glasses respond to the user's object selection instruction (as an example of a second instruction) and determine the object selected by selection box 1.
[0214] In this embodiment, the selection box 1 can be a selection method. When the selection box 1 frames an object, it means that the object is selected. The object selected by the selection box 1 can be the object that the user sees through the selection box 1 when looking out through the lenses of the smart glasses.
[0215] In this embodiment, after the user puts on the smart glasses and opens the selection box, if they need to select a real-world object, they can issue an object selection command. The smart glasses can acquire the object selection command issued by the user, which is used to instruct the selection of a real-world object located within the selection box 1. The object selection command can be a voice command, or it can be triggered by clicking a touchpad, pressing a button, or rotating a knob on the smart glasses, etc. This embodiment does not specifically limit the way the user issues the object selection command. For example, when performing AI object recognition on images captured by the smart glasses' camera, if there are many objects in the image and the user needs to select one, the user can say, "Please select an item from the selection box." The smart glasses can acquire this voice and analyze it to determine that the user has issued an object selection command.
[0216] In this embodiment, after the smart glasses receive an object selection instruction, they can respond to the instruction by mapping the information of the selection box onto the image captured by the camera through an algorithm, thereby determining the object selected by selection box 1 from the image captured by the camera. The specific determination method will be described in detail later.
[0217] In summary, this application embodiment presents a selection box on the lenses of smart glasses, enabling users to accurately select the desired range / object and precisely express their intent. The processing device, such as the smart glasses, can also accurately obtain the range / object seen by the user within the selection box, thereby accurately understanding the user's intent.
[0218] The object selection method provided in this application will be described in detail below with reference to a specific embodiment.
[0219] Referring to FIG8a, FIG8a shows a flowchart of an object selection method provided in another embodiment of the present application, which may include steps S801-S804.
[0220] S801: The smart glasses receive a selection box activation command (as an example of a first command) and display selection box 1 (as an example of a first selection box) on the lens of the smart glasses.
[0221] S802: In response to the user's object selection instruction (as an example of a second instruction), the smart glasses determine the FOV1 range (as an example of a first field of view range) formed by the selection box 1 with the human eye position (as an example of a first preset position) as the reference point, and the FOV2 range of the camera (as an example of a second field of view range).
[0222] The first preset position can be used to identify the location of the user's eyes when wearing the smart glasses. For example, the first preset position can be the location of the user's eyes (i.e., the position of the human eye) when wearing the smart glasses. The first preset position can be the same or different for different users. In some embodiments, the projection of the user's eyes onto the plane of the smart glasses lenses can be located within the selection box 1, but is not limited to this.
[0223] In practical applications, once a user wears smart glasses, the glasses can detect the user's eyes using sensors (such as eye-tracking sensors) to determine the position of the eyes. They can also determine the distance between the eye position and the plane containing the lenses of the smart glasses.
[0224] It should be noted that the aforementioned eye position can also be an empirical value, that is, the flash memory of the smart glasses can store a default position as the eye position (or the first preset position). This default position can be preset based on the eye position of most users after wearing the smart glasses. For example, it can be set to a position 6-10mm away from the plane where the lenses of the smart glasses are located. This application embodiment does not specifically limit this.
[0225] In this embodiment, after the smart glasses receive an object selection instruction, they can calculate the FOV1 range formed by the selection box 1 with a preset position as the reference point based on the parameter information of the selection box 1. The FOV2 range of the camera can also be determined based on the camera's own hardware parameters. The parameter information of the selection box 1 can be obtained in advance or after receiving the object selection instruction; this embodiment does not specifically limit this. The method for determining the FOV1 range will be described in detail later.
[0226] S803: The smart glasses determine the mapping relationship between the FOV1 range and the FOV2 range.
[0227] In this embodiment, after determining the FOV1 and FOV2 ranges, the FOV1 range can be mapped to the FOV2 range to determine the mapping relationship between them. This mapping relationship can be represented by the position and size of the FOV1 range within the FOV2 range. The method for determining the mapping relationship will be described in detail later.
[0228] S804: The smart glasses determine the region image (as an example of the second image) corresponding to the object selected by selection box 1 from the image captured by the camera (as an example of the first image) based on the mapping relationship.
[0229] In this embodiment, the camera mounted on the smart glasses can capture images within the user's field of view, which may include all objects within the FOV2 range. After determining the mapping relationship between the FOV1 range and the FOV2 range, the image captured by the camera can be proportionally cropped based on the HFOV and VFOV values corresponding to the FOV1 and FOV2 ranges, resulting in the image pixel region corresponding to the FOV1 range, which serves as the region image corresponding to the object selected by selection box 1.
[0230] For example, suppose FOV1 is a rectangular area with HFOV1 = 20° and VFOV1 = 20°, while FOV2 is a rectangular area with HFOV2 = 60° and VFOV2 = 60°, and the image captured by the camera is 1920*1920 pixels. Then, by performing proportional cropping—that is, first determining the ratio between HFOV1 and HFOV2 to be 1:3, and the ratio between VFOV1 and VFOV2 to also be 1:3—the cropped area image size can be determined to be 640*640 pixels. Then, based on the mapping relationship between FOV1 and FOV2, the cropping position is determined, and a 640*640 pixel area image is cropped at that position. For example, if the cropping position is the center, then the center of the image captured by the camera can be used as the center, and a 640*640 pixel area image can be directly cropped.
[0231] It should be noted that other contents in steps S801 to S804 in the embodiments of this application can be referred to the relevant contents in the embodiments shown in Figures 6 to 7c, and will not be repeated here.
[0232] In summary, in the embodiments of this application, the smart glasses can respond to object selection instructions and accurately capture the object selected by the user through the selection box from the image captured by the camera through algorithms such as field of view mapping.
[0233] The method for determining the FOV1 range will be described in detail below with reference to Figure 8b.
[0234] In some embodiments, the method for determining the FOV1 range formed by the selection box 1 with a preset position as a reference point may include: obtaining selection box parameters of the selection box 1, the selection box parameters including the selection box size and selection box shape of the selection box 1; obtaining a first distance, the first distance being used to characterize the distance between the human eye position and the plane where the lens of the smart glasses is located; and determining the FOV1 range based on the selection box parameters of the selection box 1 and the first distance.
[0235] In this embodiment, the VFOV1 and HFOV1 values of the FOV1 range can be calculated first using trigonometric functions based on the size of the selection box 1 (including the vertical height and horizontal width of the selection box) and the first distance between the human eye position and the plane where the lens is located. Then, the shape of the FOV1 range is determined based on the shape of the selection box 1. The shape of the FOV1 range can be the same as or only similar to the shape of the selection box 1. For example, if the selection box 1 is rectangular, the shape of the FOV1 range is also rectangular. This embodiment does not specifically limit this.
[0236] By way of example, referring to Figure 8b of the specification, Figure 8b shows a schematic diagram of determining the FOV1 range according to an embodiment of this application. As shown in Figure 8b, the shape of the selection box 1 can be obtained as a rectangle (or a rectangle with rounded corners), the vertical height of the selection box 1 is a, the horizontal width is b, and the first distance between the human eye position and the plane where the lens is located is d. Then, the VFOV1 value of the FOV1 range can be calculated using trigonometric functions based on d and a, and the HFOV1 value of the FOV1 range can be calculated using trigonometric functions based on d and b.
[0237] The method for determining the mapping relationship between the FOV1 range and the FOV2 range is described in detail below with reference to Figures 8c to 8e.
[0238] In some embodiments, determining the mapping relationship between the FOV1 range and the FOV2 range may include: mapping the FOV1 range to the FOV2 range in a center-aligned manner to obtain the mapping relationship.
[0239] In this embodiment, the FOV1 range can be directly mapped to the FOV2 range using center alignment. That is, as shown in Figure 8c, the FOV1 range can be directly mapped to the center position of the FOV2 range, and when cropping the image captured by the camera proportionally, the cropping can also be performed directly from the center position of the image.
[0240] For example, suppose FOV1 is a rectangular area with HFOV1 = 20° and VFOV1 = 20°, while FOV2 is a rectangular area with HFOV2 = 60° and VFOV2 = 60°, and the image captured by the camera is 1920*1920 pixels. Then, as shown in Figure 8d, a 640*640 pixel area can be cropped directly from the center of the image captured by the camera.
[0241] It is understood that the embodiments of this application can improve system processing efficiency by performing FOV mapping in a center-aligned manner.
[0242] In some embodiments, determining the mapping relationship between the FOV1 range and the FOV2 range may include: obtaining the selection box position of the selection box 1 and the camera position of the camera; determining the relative position information between the selection box 1 and the camera based on the selection box position of the selection box 1 and the camera position; determining the mapping position of the FOV1 range in the FOV2 range based on the relative position information; and mapping the FOV1 range to the mapping position in the FOV2 range to obtain the mapping relationship.
[0243] In this embodiment, since the FOV1 and FOV2 ranges may not be perfectly center-aligned, using center alignment for mapping would introduce errors, making the actual selected object inaccurate. Therefore, further error elimination is necessary. Specifically, the mapping position of the FOV1 range within the FOV2 range can be determined first based on the relative position information between the selection box 1 and the camera. Then, the FOV1 range can be mapped to its mapping position within the FOV2 range to obtain a more accurate mapping relationship.
[0244] In this embodiment, the position of the selection box 1 and the camera position both refer to the position of their respective center points on the plane where the lenses of the smart glasses are located, and can be directly obtained from the Flash memory of the smart glasses. The camera position and the initial position of the selection box 1 can be pre-tested and calibrated during the manufacturing process of the smart glasses and stored in the Flash memory. If the user adjusts the position of the selection box 1, the smart glasses can also store the adjusted position in the Flash memory.
[0245] In practical applications, the position (x0, y0) of the camera in the plane where the lens is located can be set as the origin (0, 0) of that plane, and the position of selection box 1 in that plane can be recorded as (x1, y1). Then, the relative position coordinates between selection box 1 and the camera can be determined as (x1, y1). Using the relative position coordinates (x1, y1), the position (x2, y2) of the center point of the FOV1 range relative to the center point of the FOV2 range can be determined, which is the mapped position of the FOV1 range within the FOV2 range. Specifically, the relative position coordinates (x1, y1) can be directly multiplied by a first scaling factor to obtain (x2, y2). This first scaling factor can be predetermined based on the camera's hardware parameters; this embodiment does not specifically limit its application. Then, the FOV1 range can be mapped to this mapped position within the FOV2 range. When cropping the image captured by the camera proportionally, the cropping can be performed from the mapped position in the image.
[0246] For example, suppose FOV1 is a rectangular area with HFOV1 = 20° and VFOV1 = 20°, while FOV2 is a rectangular area with HFOV2 = 60° and VFOV2 = 60°, and the image captured by the camera is 1920*1920 pixels. Then, as shown in Figure 8e, a region of 640*640 pixels can be cropped with the mapped position (x2, y2) in the image as the center.
[0247] In some embodiments, the FOV1 range and FOV2 range can be corrected to center alignment according to the relative position coordinates (x1, y1) and the first scaling factor. That is, the image captured by the camera and the image of the area to be cropped are corrected to center alignment, and then cropped from the center of the image to obtain the area image corresponding to the object selected by the selection box 1.
[0248] It is understood that the embodiments of this application utilize the relative positional relationship between the selection box and the camera to determine the mapping position of the FOV1 range within the FOV2 range, which can improve the accuracy of the mapping relationship and reduce the error in determining the object selected by the selection box.
[0249] The following describes in detail the object selection method provided in the embodiments of this application, taking the scenario of using an AI assistant as an example.
[0250] The following describes in detail an object selection method provided in one embodiment of this application, using smart glasses as an example. The smart glasses are equipped with lenses and a camera, which can be of various types. This embodiment does not specifically limit the camera installed on the smart glasses. Referring to FIG9a, FIG9a shows a flowchart of an object selection method provided in one embodiment of this application, which may include steps S901-S907.
[0251] S901: The smart glasses receive a selection box activation command (as an example of a first command) and display selection box 1 (as an example of a first selection box) on the lens of the smart glasses.
[0252] S902: In response to the user's object selection instruction (as an example of a second instruction), the smart glasses determine the FOV1 range (as an example of a first field of view range) formed by the selection box 1 with the human eye position (as an example of a first preset position) as the reference point, and the FOV2 range of the camera (as an example of a second field of view range).
[0253] S903: The smart glasses determine the mapping relationship between the FOV1 range and the FOV2 range.
[0254] S904: The smart glasses acquire images captured by the camera (as an example of a first image).
[0255] It should be noted that steps S904, S902, and S903 can be performed sequentially or simultaneously, and this application embodiment does not specifically limit this.
[0256] S905: The smart glasses determine the region image (as an example of the second image) corresponding to the object selected by selection box 1 from the image captured by the camera based on the mapping relationship.
[0257] S906: The smart glasses recognize the area image and obtain the object selected by selection box 1.
[0258] In this embodiment, image recognition processing can be performed on the captured region image to obtain the object selected by selection box 1. Specific recognition methods will not be detailed here.
[0259] S907: The smart glasses process the object selected by selection box 1.
[0260] In this embodiment, after determining the object selected by selection box 1, further processing can be performed on the object. For example, the object selected by selection box 1 can be identified, its type can be obtained and output. Alternatively, the object selected by selection box 1 can be analyzed to obtain its feature information and output. Alternatively, the user's object selection command can be responded to directly based on the object selected by selection box 1. Specific identification, analysis, and response methods will not be elaborated upon in this embodiment.
[0261] In some embodiments, different analyses can be performed on different types of objects. For example, when the object selected by selection box 1 is food, the object selected by selection box 1 can be processed, including but not limited to calorie analysis and freshness analysis. When the object selected by selection box 1 is an item, parameters such as the size, brand, and price of the object selected by selection box 1 can be analyzed.
[0262] In some embodiments, the system can also generate and output response information for the user's object selection instruction based on the processing result, thereby completing the interaction with the user. It can also perform operations on the object based on the processing result, such as search operations, parameter query operations, or control operations, and generate and output response information for the user's object selection instruction based on the operation result, thereby completing the interaction with the user.
[0263] For example, referring to Figure 9b of the reference manual, in a museum tour scenario, after the user puts on the smart glasses and opens the selection box, they can see the screen shown in Figure 9b. At this time, the user can say, "Please introduce the porcelain in the selection box in front of me." After receiving the voice, the smart glasses can determine that the user has issued an object selection command. At this time, the smart glasses can segment the area image 911 corresponding to the object selected by selection box 1 from the image 910 captured by the camera. The smart glasses can also further analyze the segmented area image 911 to obtain the analysis result, that is, to determine that the object selected by selection box 1 is "blue and white porcelain" and the corresponding dynasty is Ming Dynasty. Then, the smart glasses can output the voice, "This porcelain is blue and white porcelain from the Ming Dynasty." During the analysis process, dynamic effects of the selection box can also be added to indicate that the search is in progress, and the voice broadcast will be directly performed after the analysis is completed. Afterwards, the user can also move their head to move selection box 1 to another piece of porcelain and say, "Please introduce the porcelain in the selection box in front of me." After receiving the voice message, the smart glasses can perform the above analysis process again to obtain the analysis result, namely, confirming that the object selected in selection box 1 is "enamel porcelain" and the corresponding dynasty is Qing Dynasty. Then, the smart glasses can output the voice message "This porcelain is enamel porcelain from the Qing Dynasty".
[0264] For example, referring to Figure 9c of the specification, in a shopping scenario, after the user puts on the smart glasses and opens the selection box, they will see the screen shown in Figure 9c. At this time, the user can issue a voice command saying, "Please search for this tent in the XX shopping app and add it to my cart." Upon receiving this voice command, the smart glasses can determine that the user has issued an object selection instruction. The smart glasses can then segment the area image 921 corresponding to the object selected by selection box 1 from the image 920 captured by the camera. The smart glasses can further analyze the segmented area image 921 to obtain the analysis result, namely, determining that the object selected by selection box 1 is "tent." Then, the smart glasses can search for this tent in the XX shopping app, add it to the shopping cart, and then output the voice command "Okay, added."
[0265] For example, referring to Figure 9d in the reference specification, in an AI object recognition scenario, after the user puts on the smart glasses and opens the selection box, they can see the screen shown in Figure 9d. At this time, the user can ask, "What kind of flower is this?" Upon receiving this voice, the smart glasses can determine that the user has issued an object selection command. The smart glasses can then segment the area image 931 corresponding to the object selected by selection box 1 from the image 930 captured by the camera. The smart glasses can further analyze the segmented area image 931 to obtain the analysis result, namely, determining that the object selected by selection box 1 is "Gesang flower". Then, the smart glasses can output the voice message, "This is a Gesang flower, from the Qinghai-Tibet Plateau region of China. It loves sunshine and has extremely strong cold resistance."
[0266] For example, referring to Figure 9e of the reference specification, in a whole-house smart home scenario, after the user puts on the smart glasses and opens the selection box 1, and moves their head to select the smart devices (such as the living room air conditioner) within their field of vision, they will see the screen shown in Figure 9e. At this time, the user can directly issue a voice command, "Adjust to xx degrees." After receiving this voice command, the smart glasses can determine that the user has issued an object selection instruction. At this time, the smart glasses can segment the area image 941 corresponding to the object selected by selection box 1 from the image 940 captured by the camera. The smart glasses can also further analyze the segmented area image 941 to obtain the analysis results, such as identifying the object selected by selection box 1 as the living room air conditioner through an image recognition algorithm. Then, the smart glasses can adjust the temperature of the living room air conditioner to xx degrees and output the voice command, "Okay, adjusted to xx degrees."
[0267] In some embodiments, if the smart glasses fail to recognize the object selected by selection box 1, they can interact with the user to guide the user to adjust the position of selection box 1, or adjust the position and / or size of selection box 1 itself, and then reselect the object to finally determine the object the user needs to select. For example, the smart glasses can output the voice message "Selection failed, please adjust."
[0268] In this embodiment, the user can adjust the selection box to more precisely select objects. For example, the user can change the position of their head, causing the selection box 1 to move with their head, and observe the selection box to more precisely select objects, as well as select objects in different areas. Alternatively, the user can adjust the position and / or size of the selection box 1 using touch, buttons, and knobs to more precisely select objects. After the user has made the adjustments, they can issue another object selection command, and the smart glasses can again use the above object selection method to select objects.
[0269] It should be noted that other contents in steps S901 to S905 in the embodiments of this application can be referred to the relevant contents in the embodiments shown in Figures 6 to 8e, and will not be repeated here.
[0270] In summary, smart glasses use a selection box to identify and process the selected object, and interact with the user based on the recognition results. This greatly simplifies the interaction process, making it more direct and intuitive, and significantly improving accuracy, thus enhancing interaction efficiency and user experience. Furthermore, by further narrowing the selection and analysis scope using the selection box, the demand for AI computing power can be reduced, ensuring the system's real-time performance requirements are met.
[0271] The following uses smart glasses as an example to describe in detail the selection box presentation method provided in this application embodiment. This method can be executed when the object selected by selection box 1 is not accurately determined in the embodiment shown in FIG. 6, or it can be executed when the user triggers an AI analysis command. The smart glasses are equipped with lenses and a camera, which can be of various types; this application embodiment does not specifically limit the camera installed on the smart glasses. Referring to FIG. 10a, FIG. 10a shows a flowchart of a selection box presentation method provided in an embodiment of this application, which may include steps S1001-S1004.
[0272] S1001: Smart glasses acquire images captured by the camera (as an example of a third image).
[0273] In this embodiment, when a user triggers an AI analysis command, an image within the user's field of view can be captured by the camera. This image may include all objects within the camera's FOV2 range (as an example of a second field of view). The AI analysis command can be a voice command, or it can be triggered by clicking a touchpad, pressing a button, or rotating a knob on the smart glasses, etc. This embodiment does not specifically limit the method by which the user triggers the AI analysis command. For example, when performing AI object recognition on an image captured by the smart glasses' camera, if there are many objects in the frame and the user needs the smart glasses to automatically analyze the objects, they can say, "Please analyze the objects in the image." The smart glasses can acquire this voice message and analyze it to determine that the user has triggered the AI analysis command.
[0274] In some embodiments, the image can be an image captured by a camera in the object selection method provided in the embodiments shown in Figures 6 to 8e. That is, the user does not need to trigger an AI analysis command, and the embodiments of this application can be performed after each step of the object selection method provided in the embodiments shown in Figures 6 to 8e, for example, when the smart glasses do not recognize the object selected by selection box 1.
[0275] S1002: The smart glasses identify the image captured by the camera and obtain the image region where the object of interest is located.
[0276] In this embodiment, image analysis can be performed on the image to automatically identify objects of potential interest to the user and determine the location and pixel range of the objects of interest in the image. For example, various AI real-time analysis methods can be used to identify the image and obtain the objects of interest and the image region they occupy. This embodiment does not specifically limit the image analysis method used, and the specific image analysis process will not be described in detail here.
[0277] For example, referring to Figure 10b of the specification, suppose the image captured by the camera includes two objects: a truck and a running person. Recognizing this image reveals that the object of interest is the truck, and determines the image region where the truck is located and the center position of that region (i.e., the center point of the truck).
[0278] S1003: The smart glasses determine selection box 2 (as an example of a second selection box) based on the image region where the object of interest is located.
[0279] In this embodiment, when the selection box 2 is displayed on the lens of the smart glasses, it can form a field of view range with the human eye position (as an example of a first preset position) as the reference point, and the field of view range can cover at least a portion of the field of view range corresponding to the image area, that is, the field of view range formed by the selection box 2 at least partially overlaps with the field of view range corresponding to the image area. In some embodiments, the field of view range formed by the selection box 2 completely overlaps with the field of view range corresponding to the image area.
[0280] In this embodiment, the smart glasses can determine the selection box parameters of selection box 2 based on the image region where the object of interest is located. These selection box parameters may include, but are not limited to, the size and position of selection box 2. The specific method for determining selection box 2 will be described in detail later.
[0281] S1004: The smart glasses display a selection box 2 on the lenses of the smart glasses.
[0282] In this embodiment, the smart glasses can display the selection box 2 on the lens of the smart glasses based on the selection box parameters of the selection box 2. Specifically, the selection box 2 can be displayed on the lens of the smart glasses by controlling the electrochromic film through the driver IC, based on parameters such as the size, position, and shape of the selection box 2. The method of displaying the selection box 2 is similar to the method of displaying the selection box 1 in the embodiments shown in Figures 6 to 8e, and will not be described again in this embodiment.
[0283] In some embodiments, before presenting selection box 2, the smart glasses may output a query message to the user, and present selection box 2 upon receiving confirmation from the user based on the query message. This query message instructs the user to confirm the selection of the identified object of interest. For example, the smart glasses may output a voice prompt asking, "Would you like to select the item on the left in front?" If the user replies "Yes," it confirms that the user's confirmation has been received, and selection box 2 can then be presented on the lenses of the smart glasses.
[0284] It should be noted that other contents in steps S1001 to S1004 in the embodiments of this application can be referred to the relevant contents in the embodiments shown in Figures 6 to 8e, and will not be repeated here.
[0285] In summary, in this embodiment, the smart glasses can perform image analysis on images captured by the camera, automatically identify objects of interest to the user, determine a selection box on the lens based on the image area where the object of interest is located, and finally present it. In this way, the smart glasses can proactively label objects that the user may be interested in and recommend them to the user. The user can quickly select the object based on the recommendation, thus improving the user experience.
[0286] The specific method for determining selection box 2 is described in detail below.
[0287] In some embodiments, the smart glasses determine the selection box 2 based on the image region where the object of interest is located, which may include: the smart glasses determining the FOV3 range corresponding to the image region (as an example of the third field of view range) based on the image region where the object of interest is located and the FOV2 range of the camera; the smart glasses determining the selection box parameters of the selection box 2 based at least on the FOV3 range.
[0288] In this embodiment, after determining the image region where the object of interest is located, the FOV3 range corresponding to the object of interest can be obtained within the FOV2 range of the camera, based on the center position and range of the image region where the object of interest is located. Specifically, the proportional relationship between the FOV3 range and the FOV2 range can be determined first based on the proportional relationship between the range of the image region where the object of interest is located and the entire image, thereby determining the size of the FOV3 (HFOV3, VFOV3) range. Alternatively, the position of the FOV3 range within the FOV2 range can be directly determined based on the center position of the image region where the object of interest is located.
[0289] For example, assuming the image captured by the camera is 1920*1920 pixels, the image region containing the object of interest is 640*640 pixels, and the FOV2 range is a rectangular area with HFOV2 = 60° and VFOV2 = 60°, then the FOV3 range can be determined to be a rectangular area with HFOV1 = 20° and VFOV1 = 20°. The position (x3, y3) of the FOV3 range within the FOV2 range can also be determined based on the center position of the image region containing the object of interest, where (x3, y3) can be coordinates in a coordinate system with the center point of the FOV2 range as the origin.
[0290] In this embodiment, the selection parameters of selection box 2 may include the size and position of selection box 2. When selection box 2 is formed on the lens of smart glasses, it can form a field of view range with a preset position as a reference point, and the field of view range can cover at least a portion of the FOV3 range, that is, the field of view range formed by selection box 2 at least partially overlaps with the FOV3 range. In some embodiments, the field of view range formed by selection box 2 completely overlaps with the FOV3 range.
[0291] In some embodiments, the smart glasses determine the selection box parameters of the selection box 2 based at least on the FOV3 range, which may include: obtaining a first distance, the first distance being used to characterize the distance between the human eye position (as an example of a first preset position) and the plane where the lens of the smart glasses is located; determining the selection box size of the selection box 2 based on the FOV3 range and the first distance; and determining the selection box position of the selection box 2 based on the position of the FOV3 range within the FOV2 range.
[0292] In this embodiment, the vertical height *a* and horizontal width *b* of the selection box 2 can be calculated using trigonometric functions based on the values of VFOV3 and HFOV3 within the FOV3 range, and the first distance between the preset position and the plane where the lens is located. The position coordinates (x3, y3) of the FOV3 range within the FOV2 range are multiplied by a second scaling factor to obtain the selection box position (x3', y3') of the selection box 2. Referring to Figure 10c of the specification, the selection box position (x3', y3') of the selection box 2 can be the coordinates in a coordinate system with the camera position as the origin in the plane where the lens is located. The second scaling factor is the reciprocal of the first scaling factor, and the second scaling factor can also be predetermined based on the camera's hardware parameters; this embodiment does not specifically limit this.
[0293] It is understood that the embodiments of this application determine the corresponding field of view range based on the image region where the object of interest is located, and calculate the selection box parameters presented on the lens by the selection box forming the field of view range, thereby accurately determining the selection box corresponding to the image region where the object of interest is located.
[0294] In some embodiments, the selection parameters of selection box 2 may further include the selection shape of selection box 2.
[0295] In some embodiments, the smart glasses determine the selection box parameters of the selection box 2 based at least on the FOV3 range, which may include: determining the shape feature information of the FOV3 range; and determining the selection box shape of the selection box 2 based on the shape feature information of the FOV3 range.
[0296] In this embodiment, a suitable selection box shape can be selected based on the shape of the FOV3 range. The shape of the FOV3 range can be determined based on the shape of the image region where the identified object of interest is located. For example, if the image region where the identified object of interest is located can be a rectangle, then the shape of the FOV3 range can also be a rectangle, and correspondingly, the selection box shape of selection box 2 can also be selected as a rectangle.
[0297] In some embodiments, the smart glasses determine the selection box parameters of the selection box 2 based at least on the FOV3 range, which may include: determining the shape feature information of the object of interest; and determining the selection box shape of the selection box 2 based on the shape feature information of the object of interest.
[0298] In this embodiment, a suitable selection box shape can be automatically selected based on the actual shape of the identified object of interest, thereby presenting the selection box 2 in the most appropriate shape. For example, assuming the object of interest is a truck that is horizontal in the field of view, a horizontal rectangle or horizontal bar shape can be selected. Assuming the object of interest is a person or a tree that is vertical in the field of view, a vertical rectangle or vertical bar shape can be selected.
[0299] In some embodiments, the shape of the selection box 2 can also be directly selected as a preset shape. The preset shape can be determined in advance according to actual needs, such as a rectangle, etc., and this application embodiment does not specifically limit it.
[0300] It is understood that the embodiments of this application determine the shape of the selection box presented on the lens based on the shape characteristics of the object of interest or the shape characteristics of the field of view range formed by the selection box, thereby enabling the selection box to be presented in the most suitable shape and further improving the user experience.
[0301] The above describes the relevant methods for real-world recognition using smart glasses. Below, based on the scenario of smart glasses working in conjunction with a mobile phone, another embodiment of this application provides a detailed description of the selection box presentation method.
[0302] The following describes in detail a selection box presentation method provided in one embodiment of this application, using a mobile phone as an example. A camera is located on one side of the mobile phone's display screen. This camera can be of various types, and this embodiment does not specifically limit the camera located on the mobile phone. Referring to FIG11a, FIG11a shows a flowchart of a selection box presentation method provided in one embodiment of this application. This method can be executed by a mobile phone and may include steps S1101-S1104.
[0303] S1101: The phone displays the first screen.
[0304] In this embodiment of the application, when the mobile phone interacts with the smart glasses, it can display a first interface on the screen. The first interface can be of various types, such as including but not limited to the system desktop, application page, etc. This embodiment of the application does not specifically limit it.
[0305] S1102: The phone displays selection box 4 on the first screen (as an example of the fourth selection box).
[0306] In this embodiment, after the user puts on the smart glasses and opens the selection box, the smart glasses can display selection box 3 (as an example of the third selection box) on the lens. The specific details of the position, size, shape, and presentation method of the selection box 3 can be referred to the relevant content of selection box 110 in the embodiments shown in Figures 2a to 3b, and will not be repeated here in this embodiment.
[0307] In this embodiment, after the smart glasses open the selection box, a selection box activation request can be sent to the mobile phone. The mobile phone can respond to the selection box activation request and display selection box 4 on the first interface. The size, shape, and presentation method of the selection box 4 can be referred to the relevant content of selection box 220 in the embodiments shown in Figures 4a and 4b, and will not be repeated here.
[0308] In some embodiments, displaying the selection box 4 on the first interface of the mobile phone may include displaying the selection box 4 at the center position of the mobile phone's display screen (as an example of a second preset position).
[0309] The second preset position can be preset according to actual needs, and this application embodiment does not specifically limit it. The above implementation method of using the center position of the display screen as the second preset position is only an example. In some embodiments, the second preset position can also be set to the upper left corner, upper right corner, lower left corner, lower right corner, left middle or right center of the display screen. This application embodiment does not specifically limit it.
[0310] In some embodiments, displaying a selection box 4 on a first interface on a mobile phone may include: the mobile phone acquiring an image captured by a camera (as an example of a fourth image), the image captured by the camera including the selection box 3 presented on the lens of the smart glasses; determining a first position of the selection box 3 relative to the mobile phone based on the image captured by the camera; determining a third position of the selection box 4 on the display screen of the mobile phone based on the first position of the selection box 3; and displaying the selection box 4 at the third position on the display screen of the mobile phone.
[0311] In this embodiment, selection box 4 can be displayed corresponding to selection box 3. The camera on the mobile phone can capture images of selection box 3 displayed on at least the lenses of the smart glasses, so as to track selection box 3 in real time, and map the position of selection box 3 to the plane where the mobile phone display screen is located through coordinate mapping, thereby determining the display position of selection box 4.
[0312] The number of selection boxes 3 can be one or more, and this embodiment does not specifically limit the number. For example, when there is only one selection box 3, it can be a selection box on the left lens of the smart glasses. When there are multiple selection boxes 3, they can include at least one selection box on the left lens and at least one selection box on the right lens of the smart glasses.
[0313] In this embodiment, the mobile phone can calculate the center point position of the selection box 3 on the smart glasses in real time based on the image captured by the camera and the Simultaneous Localization and Mapping (SLAM) algorithm, which is used as the first position of the selection box 3 relative to the mobile phone.
[0314] In this embodiment, the mobile phone can map the first position of the selection box 3 to the plane where the mobile phone's display screen is located based on a preset mapping relationship, so as to determine the third position of the selection box 4 on the mobile phone's display screen, and thus display the selection box 4 at the third position. This mapping relationship can be preset according to actual needs, for example, it can be preset according to the relative pose between the user's head and the mobile phone. This embodiment does not specifically limit this.
[0315] S1103: The mobile phone detects that the selection box 3 displayed on the lens of the smart glasses has changed from a first position to a second position relative to the mobile phone.
[0316] In this embodiment, the camera on the mobile phone can capture images in real time, including at least the selection box 3 displayed on the lens of the smart glasses, to track the selection box 3 in real time and determine whether the position of the selection box 3 relative to the mobile phone has changed. When it is determined that the position of the selection box 3 relative to the mobile phone at the current moment is a second position, and it is determined that the position of the selection box 3 relative to the mobile phone at the previous moment was a first position, it can be determined that the selection box 3 displayed on the lens of the smart glasses has changed from the first position to the second position relative to the mobile phone.
[0317] S1104: The mobile phone adjusts the display position of the selection box 4 based on the change in the position of the selection box 3.
[0318] In this embodiment of the application, the mobile phone can use coordinate mapping to map the position change of the selection box 3 to the plane where the mobile phone display screen is located, thereby adjusting the position of the selection box 4 on the mobile phone display screen and determining the object selected by the selection box 4.
[0319] In some embodiments, the mobile phone adjusts the display position of the selection box 4 based on the position change of the selection box 3, which may include: determining a first position change amount of the selection box 3 from a first position to a second position; determining a second position change amount of the selection box 4 on the mobile phone display screen based on the first position change amount; determining a fourth position of the selection box 4 on the mobile phone display screen based on the second position change amount; and adjusting the selection box 4 to the fourth position.
[0320] In this embodiment of the application, referring to Figures 11b and 11c of the specification, the plane where the lenses of the smart glasses are located can be denoted as plane Z1. The plane where the phone's display screen is located can be denoted as plane Z2 = 0, and the position of the phone's camera can be used as the origin (0, 0, 0) of the coordinates in plane Z2.
[0321] In this embodiment, when the smart glasses move with the user's head, the change in the selection box 3 displayed on the lenses of the smart glasses relative to the phone from a first position (x4, y4, z4) to a second position (x4', y4', z4') can be detected. The change in the first position of the selection box 3 can be directly determined based on the first position (x4, y4, z4) and the second position (x4', y4', z4'). For example, the change in the selection box 3 along the z-axis can be ignored, thus determining that the second position (x4', y4', z4') of the selection box 3 is still located on the Z1 plane. At this time, the change in the position of the selection box 3 relative to the phone can be determined to be the change in the position of the selection box 3 on the Z1 plane. That is, if the position of the selection box 3 relative to the phone changes from (x4, y4) to (x4', y4'), then the change in the first position of the selection box 3 can be determined as Δ = (x4' - x4, y4' - y4).
[0322] In this embodiment, the mobile phone can determine the position change of the selection box 4 based on the position change of the selection box 3. The position changes of the selection box 3 and the selection box 4 can change proportionally or unequally, and this embodiment does not specifically limit this. For example, when the pose tracking accuracy is high, the position changes of the selection box 3 and the selection box 4 can change proportionally; when the pose tracking accuracy is poor, the position changes of the selection box 3 and the selection box 4 can change unequally.
[0323] In some embodiments, the mobile phone determines the second position change of the selection box 4 on the mobile phone's display screen based on the first position change, which may include: the mobile phone multiplying the first position change by a preset scaling factor to obtain the second position change.
[0324] In this embodiment, the mobile phone can directly use the product of the x-axis position change Δx = x4' - x4 of the selection box 3 and a preset scaling factor as the x-axis position change Δx′ of the selection box 4, and the product of the y-axis position change Δy = y4' - y4 of the selection box 3 and a preset scaling factor as the y-axis position change Δy′ of the selection box 4. This preset scaling factor can be predetermined based on actual conditions, such as the phone's pose tracking accuracy for the selection box 3. This embodiment does not specifically limit this.
[0325] In some embodiments, when there are multiple selection boxes 3, the mobile phone can obtain multiple first position changes. In this case, the mobile phone can first perform a weighted summation of these multiple first position changes, and then multiply the weighted position changes by a preset proportional coefficient to obtain a second position change. The weighting coefficient used in the weighted summation process can also be predetermined according to actual conditions; this embodiment does not specifically limit this.
[0326] In this embodiment, the mobile phone can first obtain the position of the selection box 4 in the Z2 plane at the previous moment, assuming it is the third position (x5, y5). Based on the third position (x5, y5) at the previous moment and the change in the second position, the fourth position (x5', y5') of the selection box 4 in the Z2 plane at the current moment is determined. Specifically, the mobile phone can add the x-axis coordinate x5 at the previous moment to the x-axis position change Δx' to obtain the x-axis coordinate x5' at the current moment, and add the y-axis coordinate y5 at the previous moment to the y-axis position change Δy' to obtain the y-axis coordinate y5' at the current moment.
[0327] In this embodiment of the application, after the mobile phone determines that the selection box 4 is currently at the fourth position (x5', y5') in the Z2 plane, it can adjust the selection box 4 from the third position (x5, y5) to the fourth position (x5', y5') to obtain the updated selection box 4.
[0328] In some embodiments, the second position of the selection box 3 can be directly mapped to the plane where the mobile phone display screen is located to determine the fourth position of the selection box 4 on the mobile phone display screen; the selection box 4 is then adjusted to the fourth position. The specific mapping method is the same as the method of mapping the first position of the selection box 3 to the plane where the mobile phone display screen is located to determine the third position of the selection box 4 on the mobile phone display screen, and will not be described again in the embodiments of this application.
[0329] In some embodiments, the method may further include: determining the object selected by the selection box 4; and performing operations on the object selected by the selection box 4.
[0330] In this embodiment, the mobile phone can also determine the object displayed at the location of the selection box 4 from the page content displayed on the screen, and use this object as the object selected by the selection box 4. The object selected by the selection box 4 may include, but is not limited to, text, icons, and other objects on the screen. For example, referring to Figures 12a to 12c of the specification, after the mobile phone updates the selection box 4, it can determine that the object selected by the selection box 4 may be text in the application page of a reading application, an application icon on the system desktop, or an icon in the application page of a video application.
[0331] In this embodiment, different operations can be performed on different objects. These operations may include, but are not limited to, one or more of the following: copy, paste, click, select, and play. For example, when the object selected by selection box 4 is text on the application page of a reading application, the text can be copied. When the object selected by selection box 4 is an application icon on the system desktop, the application icon can be selected. When the object selected by selection box 4 is an icon on the application page of a video application, the icon can be clicked, and so on.
[0332] In some embodiments, the position of the selection box 4 displayed on the mobile phone screen can also dynamically and automatically change according to the content displayed on the screen. When within an application, the position of the selection box 4 can be adaptively adjusted according to the UI design of the application. For example, when it is determined that the distance between the fourth position (x5', y5') of the selection box 4 and the center position of an icon is less than a preset threshold, the position of the selection box 4 can be directly adjusted to the center position of the icon so that the selection box 4 is displayed at that position. The preset threshold can be preset according to actual needs, for example, it can be set to a distance of 10 pixels, and this application embodiment does not specifically limit it.
[0333] In some embodiments, the size and shape of the selection box 4 displayed on the phone's screen can dynamically and automatically change according to the content displayed on the screen. When within an application, the size and shape of the selection box 4 can be adaptively adjusted according to the application's UI design. For example, as shown in Figure 12a, when a user opens a reading application to read an e-book, the object selected by the selection box 4 can be a piece of text 1211 on the application page 1210 of the reading application. In this case, the selection box 4 can be a horizontal rectangle, and its size can be the size of three lines of text selected. As shown in Figure 12b, when the phone screen displays the system desktop, the object selected by the selection box 4 can be an application icon 1221 on the system desktop 1220. In this case, the selection box 4 can be a square, and its size can be larger than the size of an application icon. As shown in Figure 12c, when a user opens a video application to watch a video, the object selected by the selection box 4 can be an icon 1231 on the application page 1230 of the video application used to trigger video playback. In this case, the selection box 4 can be a horizontal rectangle, and its size can be larger than the size of the entire icon.
[0334] It should be noted that other contents in steps S1101 to S1104 in the embodiments of this application can be referred to the relevant contents in the embodiments shown in Figures 6 to 8e, and will not be repeated here.
[0335] In summary, this application embodiment utilizes a camera to track the selection box on the smart glasses in real time, and through coordinate mapping, maps the positional change of the selection box on the smart glasses onto the plane of the mobile phone display screen, thereby adjusting the selection box on the mobile phone display screen and changing the object selected by the selection box on the display screen. This interaction method is more intuitive and does not require the user's hands to operate, making it suitable for accessible scenarios.
[0336] Furthermore, by directly tracking the selection boxes on the smart glasses lenses, this tracking method offers higher accuracy and greater adaptability compared to direct eye tracking due to the obvious feature points of the selection boxes. It eliminates the need for eye calibration and adaptation for different user groups. In some implementations, selection boxes on both lenses can be tracked simultaneously, thereby improving tracking accuracy and stability, and further enhancing the accuracy and stability of selecting content on the display screen using the selection boxes.
[0337] This application also provides an electronic device, including:
[0338] Memory, used to store instructions executed by one or more processors of an electronic device, and
[0339] When the processor executes instructions in the memory, it can cause the electronic device to perform the object selection method shown in Figures 6 to 9e in the above embodiments, the selection box presentation method shown in Figures 10a to 10c in the above embodiments, or the selection box presentation method shown in Figures 11a to 12c in the above embodiments.
[0340] This application also provides a computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the object selection method shown in Figures 6 to 9e in the above embodiments, the selection box presentation method shown in Figures 10a to 10c in the above embodiments, or the selection box presentation method shown in Figures 11a to 12c in the above embodiments.
[0341] This application also provides a computer program product containing instructions that, when the computer program product is run on an electronic device, cause the processor to execute the object selection method shown in Figures 6 to 9e in the above embodiments, execute the selection box presentation method shown in Figures 10a to 10c in the above embodiments, or execute the selection box presentation method shown in Figures 11a to 12c in the above embodiments.
[0342] Referring now to FIG13, a block diagram of an electronic device 1300 according to an embodiment of the present application is shown. The electronic device 1300 may include one or more processors 1301 coupled to a controller hub 1303. In at least one embodiment, the controller hub 1303 communicates with the processor 1301 via a multi-branch bus such as a Front Side Bus (FSB), a point-to-point interface such as a Quick Path Interconnect (QPI), or a similar connection 1306. The processor 1301 executes instructions controlling general-type data processing operations. In one embodiment, the controller hub 1303 includes, but is not limited to, a Graphics Memory Controller Hub (GMCH) (not shown) and an Input / Output Hub (IOH) (which may be on a separate chip) (not shown), wherein the GMCH includes memory and a graphics controller and is coupled to the IOH.
[0343] Electronic device 1300 may also include a coprocessor 1302 and a memory 1304 coupled to a controller hub 1303. Alternatively, one or both of the memory and GMCH may be integrated within the processor (as described in this application), with memory 1304 and coprocessor 1302 directly coupled to processor 1301 and controller hub 1303, which is on a single chip with IOH.
[0344] Memory 1304 may be, for example, Dynamic Random Access Memory (DRAM), Phase Change Memory (PCM), or a combination of both. As a computer-readable storage medium, memory 1304 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. For example, memory 1304 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as one or more hard-disk drives (HDD(s)), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives.
[0345] According to some embodiments of this application, the memory 1304, which is a computer-readable storage medium, stores instructions that, when executed on a computer, cause the system 1300 to perform an object selection method or a selection box presentation method according to the above embodiments. Specifically, refer to the methods shown in Figures 6 to 12c in the above embodiments, which will not be described again here.
[0346] In one embodiment, the coprocessor 1302 is a dedicated processor, such as, for example, a high-throughput many-integerated core (MIC) processor, a network or communication processor, a compression engine, a graphics processor, a general-purpose graphics processor (GPGPU), or an embedded processor, etc. Optional properties of the coprocessor 1302 are indicated by dashed lines in Figure 13.
[0347] In one embodiment, electronic device 1300 may further include a Network Interface Controller (NIC) 1306. The network interface 1306 may include a transceiver for providing a radio interface for electronic device 1300 to communicate with any other suitable device (such as a front-end module, antenna, etc.). In various embodiments, the network interface 1306 may be integrated with other components of electronic device 1300. The network interface 1306 can implement the functions of the communication unit in the above embodiments.
[0348] Electronic device 1300 may further include input / output (I / O) device 1305. I / O 1305 may include: a user interface designed to enable a user to interact with electronic device 1300; a peripheral component interface designed to enable peripheral components to also interact with electronic device 1300; and / or sensors designed to determine environmental conditions and / or location information related to electronic device 1300.
[0349] It is worth noting that Figure 13 is merely exemplary. That is, although Figure 13 shows an electronic device 1300 including multiple devices such as a processor 1301, a controller hub 1303, and a memory 1304, in actual applications, devices using the methods of this application may include only a subset of the devices in the electronic device 1300; for example, it may only include the processor 1301 and the network interface 1306. The nature of optional devices in Figure 13 is shown with dashed lines.
[0350] Referring now to FIG. 14, a block diagram of a SoC (System on Chip) 1400 according to an embodiment of this application is shown. In FIG. 14, similar components have the same reference numerals. Additionally, dashed boxes represent optional features of more advanced SoCs. In FIG. 14, the SoC 1400 includes: an interconnect unit 1450 coupled to a processor 1410; a system proxy unit 1470; a bus controller unit 1480; an integrated memory controller unit 1440; a group or one or more coprocessors 1420, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 1430; and a direct memory access (DMA) unit 1460. In one embodiment, the coprocessor 1420 includes a dedicated processor, such as, for example, a network or communication processor, a compression engine, a GPGPU, a high-throughput MIC processor, or an embedded processor.
[0351] The static random access memory (SRAM) cell 1430 may include one or more tangible, non-transitory computer-readable storage media for storing data and / or instructions. The computer-readable storage media may store instructions, specifically, temporary and permanent copies of those instructions. These instructions may include, when executed by at least one unit in the processor, causing the SoC 1400 to perform an object selection method or a selection box presentation method according to the above embodiments, specifically referring to the methods shown in Figures 6 to 12c of the above embodiments, which will not be repeated here.
[0352] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0353] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.
[0354] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0355] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, compact disc read-only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagated signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0356] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the accompanying drawings. Furthermore, including structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0357] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.
[0358] It should be noted that in the examples and description of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0359] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.
Claims
1. A method for object selection, applied to a smart head-mounted device, comprising: The smart head-mounted device is equipped with lenses, and the method includes: Upon receiving a first instruction, a first selection box is displayed on the lens of the smart head-mounted device; the lens within the first selection box displays a first transmittance, and the lens outside the first selection box displays a second transmittance, wherein the first transmittance and the second transmittance are different; In response to a second instruction from the user, the object selected by the first selection box is determined.
2. The method of claim 1, wherein, The first transmittance is greater than the second transmittance.
3. The method of claim 1, wherein, The first selection box is presented at the center of the lens, and the shape of the first selection box is rectangular, circular, elliptical, horizontal or vertical.
4. The method according to claim 1 or 3, characterized in that, The method further includes: In response to a third instruction from the user, at least one of the position of the first selection box on the lens, the shape of the first selection box, and the size of the first selection box is adjusted.
5. The method of claim 1, wherein, The object selected by the first selection box is the object that the user sees through the lens of the smart head-mounted device when looking outwards.
6. The method of claim 1, wherein, The smart head-mounted device is equipped with a camera; Determining the object selected by the first selection box includes: Determine the first field of view range formed by the first selection box with the first preset position as the reference point and the second field of view range of the camera; Determine the mapping relationship between the first field of view range and the second field of view range; Based on the mapping relationship, a second image corresponding to the object selected by the first selection box is determined from the first image captured by the camera.
7. The method of claim 6, wherein, The first preset position is used to identify the position of the user's eyes when the user wears the smart head-mounted device; the first field of view is the range that the user sees through the first selection box when looking outward through the lenses of the smart head-mounted device.
8. The method of claim 6, wherein, Determining the first field of view range formed by the first selection box with the first preset position as the reference point includes: Obtain the first selection box parameters of the first selection box, the first selection box parameters including the size of the first selection box and the shape of the first selection box; A first distance is obtained, which is used to characterize the distance between the first preset position and the first plane where the lens of the smart head-mounted device is located; The first field of view range is determined based on the first selection box parameters and the first distance.
9. The method of claim 6, wherein, Determining the mapping relationship between the first field of view range and the second field of view range includes: The first field of view range is mapped to the second field of view range using a center-aligned method to obtain the mapping relationship; or, Obtain the position of the first selection box and the position of the camera; Based on the position of the first selection box and the position of the camera, determine the relative position information between the first selection box and the camera; Based on the relative position information, determine the mapping position of the first field of view range in the second field of view range; The mapping relationship is obtained by mapping the first field of view range to the mapping position in the second field of view range.
10. The method of claim 1, wherein, The method further includes: Processing the object selected by the first selection box includes: The object selected by the first selection box is identified, the type of the object is obtained and output; and / or, The object selected by the first selection box is analyzed to obtain and output its feature information; and / or, The second instruction is responded to based on the object selected by the first selection box. 11.A method for presenting a selection frame, applied to a smart head-mounted device, the method comprising: The smart head-mounted device is equipped with lenses and a camera, and the method includes: Acquire the third image captured by the camera; The third image is identified to obtain the image region where the object of interest is located in the third image; A second selection box is determined based on the image region where the object of interest is located; wherein, when the second selection box is presented on the lens of the smart head-mounted device, the field of view range formed with the first preset position as the reference point can cover at least a portion of the field of view range corresponding to the image region; The second selection box is displayed on the lens of the smart head-mounted device.
12. The method of claim 11, wherein, Determining the second selection box based on the image region where the object of interest is located includes: Based on the image region where the object of interest is located and the second field of view range of the camera, determine the third field of view range corresponding to the image region; The second selection box parameters of the second selection box are determined based at least on the third field of view range.
13. The method of claim 12, wherein, The second selection box parameters include the size of the second selection box and the position of the second selection box; The determination of the second selection box parameters based at least on the third field of view range includes: A first distance is obtained, which is used to characterize the distance between the first preset position and the first plane where the lens of the smart head-mounted device is located; The size of the second selection box is determined based on the third field of view and the first distance; The position of the second selection box is determined based on the position of the third field of view within the second field of view.
14. The method of claim 13, wherein, The second selection box parameter also includes the shape of the second selection box; The determination of the second selection box parameters based at least on the third field of view range includes: Determine the shape feature information of the third field of view range; The shape of the second selection box is determined based on the shape feature information of the third field of view range; or, Determine the shape feature information of the object of interest; The shape of the second selection box is determined based on the shape feature information of the object of interest; or, The shape of the second selection box is determined to be a preset shape.
15. A marquee selection presentation method applied to an electronic device, comprising: The method includes: Display the first interface; A fourth selection box is displayed in the first interface; The third selection box displayed on the lens of the smart head-mounted device was detected to have changed from a first position to a second position relative to the electronic device; Based on the change in the position of the third selection box, the display position of the fourth selection box is adjusted.
16. The method of claim 15, wherein, Displaying the fourth selection box in the first interface includes: The fourth selection box is displayed at a second preset position on the display screen of the electronic device.
17. The method of claim 15, wherein, The fourth selection box is displayed in correspondence with the third selection box; Displaying the fourth selection box in the first interface includes: Acquire a fourth image, the fourth image including a third selection box presented on the lens of the smart head-mounted device; The first position of the third selection box relative to the electronic device is determined based on the fourth image; The third position of the fourth selection box on the display screen of the electronic device is determined based on the first position of the third selection box; The fourth selection box is displayed at the third position on the display screen of the electronic device.
18. The method of claim 15, wherein, The adjustment of the display position of the fourth selection box based on the position change of the third selection box includes: Determine the first position change amount of the third selection box from the first position to the second position; The second position change of the fourth selection box on the display screen of the electronic device is determined based on the first position change. The fourth position of the fourth selection box on the display screen of the electronic device is determined based on the second position change amount; Adjust the fourth selection box to the fourth position.
19. The method of claim 18, wherein, Determining the second position change of the fourth selection box on the display screen of the electronic device based on the first position change includes: The second position change is obtained by multiplying the first position change by a preset scaling factor.
20. The method of claim 15, wherein, The method further includes: Determine the object selected by the fourth selection box; Perform operations on the object selected by the fourth selection box.
21. An electronic device, comprising: include: A memory for storing instructions executed by one or more processors of the electronic device; The processor, when executing the instructions in the memory, can cause the electronic device to perform the object selection method according to any one of claims 1 to 10, the selection box presentation method according to any one of claims 11 to 14, or the selection box presentation method according to any one of claims 15 to 20.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the object selection method according to any one of claims 1 to 10, the selection box presentation method according to any one of claims 11 to 14, or the selection box presentation method according to any one of claims 15 to 20.
Citation Information
Patent Citations
Method for enhancing visual ability of human eyes based on MR glasses, system and equipment
CN113709410A
Target selection method and device based on intelligent glasses and electronic equipment
CN116301361A
Interaction method and system based on handheld intelligent device and head-mounted display device
CN116991234A
Image information acquisition method and device, intelligent glasses and storage medium
CN119946418A
Wearable smart glasses
EP3096517A1