Human-computer interaction method and smart glasses
By recognizing user gestures through the smart glasses' camera, the problem of complex voice descriptions in existing AI glasses is solved, resulting in a more convenient interactive experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-07-23
Smart Images

Figure CN2025147071_23072026_PF_FP_ABST
Abstract
Description
A human-computer interaction method and smart glasses
[0001] This application claims priority to Chinese patent application No. 202510072304.7, filed on January 16, 2025, entitled "A Human-Computer Interaction Method and Smart Glasses", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of information technology (IT) technology, and more particularly to a human-computer interaction method and smart glasses. Background Technology
[0003] A new type of artificial intelligence (AI) glasses has recently emerged on the market. Its design is closer to ordinary glasses, offering a superior wearing experience compared to augmented reality (AR) glasses, and its form factor is simpler. These AI glasses have built-in AI programs; users simply need to activate the program and describe the object in detail via voice to initiate a human-computer dialogue. However, specifying a target object requires detailed description, which leads to two main problems: first, long and complex sentences increase the difficulty of speech recognition; second, users must carefully organize their language to ensure accurate communication of their intent, raising the barrier to entry. These issues limit the smoothness of the user experience, making the interaction less convenient and negatively impacting the overall user experience. Summary of the Invention
[0004] This application provides a human-computer interaction method, smart glasses, a computer storage medium, and a computer product, which can improve the convenience of the interaction process on smart glasses and enhance the user experience.
[0005] In a first aspect, this application provides a human-computer interaction method applied to smart glasses, comprising: acquiring image data of the physical world through a camera on the smart glasses when the artificial intelligence (AI) program in the smart glasses is in an awakened state; determining a target object based on the preset gesture and the image data when the image data contains a preset gesture made by the user, wherein the target object is the object in the physical world that the user inquires about; generating response content based on the target object and a first voice input by the user; and presenting the response content in at least one modality.
[0006] In this way, during the interaction between the user and the smart glasses, the user's gestures contained in the image data captured by the camera on the smart glasses can be used to identify the target object. This eliminates the need for the user to use verbal descriptions to identify the target object, reducing the difficulty of interaction and improving the user experience.
[0007] In one possible implementation, the recognition elements of the preset gesture include one or more of the following: three fingers are folded together and the index finger is extended; or, the tip of the index finger moves in an arc, and the central angle of the arc is greater than a preset degree; or, the thumb is placed against the base of the index finger. In this way, it is possible to identify whether the image data contains the preset gesture through these elements.
[0008] In one possible implementation, the preset gesture is: the index finger is extended, the thumb is placed against the base of the index finger, and the thumb is placed against the base of the index finger for a duration exceeding a preset duration.
[0009] In one possible implementation, determining the target object based on a preset gesture and image data includes: identifying objects contained in a first image, wherein the first image is any frame of the image data containing the preset gesture; identifying the target position of the index fingertip in the preset gesture in the first image; and determining a first duration for which the index fingertip is at the target position; if the first duration is longer than a preset duration, identifying objects in the first image that overlap with the target position as the target object. Thus, the target object can be determined from the image data based on the position pointed to by the index finger.
[0010] In one possible implementation, the method further includes: if no object overlaps with the target location among the objects contained in the first image, and the first duration is longer than a preset duration, then the object closest to the target location among the objects contained in the first image is taken as the target object. This allows for accurate identification of the target object from the image data even when there is a discrepancy between the user's index finger and the object being pointed to.
[0011] In one possible implementation, the object closest to the target position among the objects contained in the first image is selected as the target object. This includes: determining the target direction pointed to by the index finger in the preset gesture; and selecting the object located on one side of the target direction and closest to the target position among the objects contained in the first image as the target object. This allows for accurate identification of the target object from the image data even when there is a deviation between the user's index finger and the object being pointed to.
[0012] In one possible implementation, the preset gesture is: the index finger is extended, the fingertip of the index finger moves, and the movement trajectory is an arc, the central angle of the arc is greater than a preset degree.
[0013] In one possible implementation, determining the target object based on a preset gesture and image data includes: determining the movement trajectory of the preset gesture in multiple consecutive images contained in the image data, and determining the trajectory position of the movement trajectory in any one of the multiple consecutive images; identifying objects contained in a first image, wherein the first image is any one of the frames in the image data; and identifying the object in the first image that overlaps with the trajectory position as the target object. Thus, the target object can be determined from the image data by the position circled by the index finger.
[0014] In one possible implementation, after identifying the target object, the process further includes: extracting the object features of the target object from the target image containing the target object in the image data; and, based on the object features, determining the target object in images acquired after the target image in the image data. This way, the user doesn't need to constantly maintain a preset gesture, improving the user experience.
[0015] In one possible implementation, there are multiple target objects. That is, multiple target objects are identified.
[0016] In one possible implementation, the response content is presented in at least one modality, including displaying the response content on the lens screen of the smart glasses. In this way, the user can observe the response content through the lens screen.
[0017] In one possible implementation, after displaying the response content on the lens screen of the smart glasses, the method further includes: detecting a user's swipe operation on an input component associated with the smart glasses; and, in response to the swipe operation, changing at least a portion of the content displayed on the lens screen, wherein the changes to the content displayed on the lens screen include one or more of the following: zooming in / out of text / text boxes, highlighting / unhighlighting special text, displaying special text annotations, displaying additional annotation information, and displaying a summary of the response content. In this way, after the smart glasses display the response content, the user can interact with the smart glasses to view information associated with the response content, thus improving the user experience.
[0018] In one possible implementation, in response to a swipe operation, at least a portion of the content displayed on the lens screen is modified, including: enlarging the text / text box when the swipe operation is towards the lens screen; and shrinking the text / text box when the swipe operation is away from the lens screen. This allows users to choose to enlarge or shrink the text / text box through specific actions, improving the user experience.
[0019] In one possible implementation, in response to a swipe operation, at least a portion of the content displayed on the lens screen is changed, including: highlighting special text in the response content when the swipe operation is towards the lens screen and the swipe distance is within a first distance range, or when the swipe operation is towards the lens screen and the swipe operation is performed multiple times consecutively, and the interval between adjacent swipes is less than a preset duration; and de-highlighting the special text in the response content when the swipe operation is away from the lens screen and the swipe distance is within a second distance range, or when the swipe operation is away from the lens screen and the swipe operation is performed multiple times consecutively, and the interval between adjacent swipes is less than a preset duration. This allows users to select whether to highlight or de-highlight special text in the response content through a specific operation, improving the user experience.
[0020] In one possible implementation, after highlighting specific text in the response content, the method further includes: detecting a swipe gesture on the input component, where the swipe gesture is either an up swipe or a down swipe; if the swipe gesture is the first down swipe, adding a focus box to the first specific text in the response content; and adjusting the position of the focus box based on subsequent swipe gestures. This allows users to select the content they want to focus on through specific actions, improving the user experience.
[0021] In one possible implementation, after adding a focus box to the target specific text in the response content, the method further includes: detecting a click operation on the input component; and in response to the click operation, displaying a target card on the lens screen, the target card carrying an explanation of the target specific text. This allows users to view an explanation of specific text in the response content through a specific action, improving the user experience.
[0022] In one possible implementation, after displaying the explanation of the target specific text on the lens screen, the method further includes: detecting a second click operation on the input component; and removing the target card from the lens screen in response to the second click operation. This way, after viewing the explanation of specific text in the response content, the user can close the explanation through a specific action, improving the user experience.
[0023] In one possible implementation, after highlighting the special text in the response content, the method further includes: detecting a swipe gesture on the input component towards the lens screen, where the swipe distance is within a second distance range; or detecting multiple consecutive swipes on the input component with adjacent swipe intervals less than a preset value; and adding annotation information to the special text in the response content. This allows users to select detailed explanations of the content they wish to focus on through specific actions, improving the user experience.
[0024] In one possible implementation, changing at least a portion of the content displayed on the lens screen in response to a swipe operation further includes: displaying a summary of the response content and de-displaying the response content when the swipe operation is a swipe away from the lens screen and the swipe distance is greater than a preset distance, or when the swipe operation is a swipe away from the lens screen and is performed multiple times consecutively with adjacent swipe intervals less than a preset time. This allows users to select and view a summary of the response content through a specific operation, improving the user experience.
[0025] In one possible implementation, the preset gesture can be a set of consecutive gestures; determining the target object based on the preset gestures and image data includes: determining multiple target objects from the image data based on the consecutive gestures and image data, wherein one gesture included in the consecutive gestures is used to specify an object.
[0026] Secondly, this application provides a human-computer interaction method applied to smart glasses, comprising: acquiring image data of the physical world through a camera on the smart glasses while the artificial intelligence (AI) program in the smart glasses is in an awakened state; identifying objects in the physical world from the image data and displaying object identifiers of the identified objects on the lens screen of the smart glasses; detecting an object selection operation by the user on an input component related to the smart glasses, wherein the object selection operation is used to select an object from the object identifiers displayed on the lens screen; generating response content based on the target object selected by the object selection operation and a first voice input by the user; and presenting the response content in at least one modality.
[0027] In this way, during the interaction between the user and the smart glasses, the image data collected by the camera on the smart glasses can be used to identify objects in the physical world and detect the user's input of object selection operations to determine the target object. This eliminates the need for users to use verbal descriptions to determine the target object, reducing the difficulty of interaction and improving the user experience.
[0028] In one possible implementation, the input component is configured on the smart glasses. For example, the input component could be configured on the temples of the smart glasses.
[0029] In one possible implementation, the input component is configured on a device other than the smart glasses. For example, the device other than the smart glasses could be, but is not limited to, a smartphone, a smartwatch, etc.
[0030] Thirdly, this application provides a smart glasses, comprising: at least one memory for storing a program; at least one processor for executing the program stored in the memory; wherein, when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation thereof, or to execute the method described in the second aspect or any possible implementation thereof.
[0031] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to perform the method described in the first aspect or any possible implementation thereof, or to perform the method described in the second aspect or any possible implementation thereof.
[0032] Fifthly, this application provides a computer program product that, when run on a processor, causes the processor to execute the method described in the first aspect or any possible implementation thereof, or to execute the method described in the second aspect or any possible implementation thereof.
[0033] It is understood that the beneficial effects of the third to fifth aspects mentioned above can be found in the relevant descriptions of the first to third aspects mentioned above, and will not be repeated here. Attached Figure Description
[0034] Figure 1 is a structural schematic diagram of a smart glasses provided in an embodiment of this application;
[0035] Figure 2 is a schematic diagram of the appearance of a smart glasses provided in an embodiment of this application;
[0036] Figure 3 is a flowchart illustrating a human-computer interaction method provided in an embodiment of this application;
[0037] Figure 4 is a schematic diagram of a gesture provided in an embodiment of this application;
[0038] Figure 5 is a schematic diagram of selecting a target object according to an embodiment of this application;
[0039] Figure 6 is a schematic diagram of selecting a target object based on different finger directions according to an embodiment of this application;
[0040] Figure 7 is a schematic diagram of selecting multiple target objects according to an embodiment of this application;
[0041] Figure 8 is a schematic diagram of a usage scenario for smart glasses provided in an embodiment of this application;
[0042] Figure 9 is a flowchart illustrating another human-computer interaction method provided in an embodiment of this application;
[0043] Figure 10 is a flowchart illustrating another human-computer interaction method provided in an embodiment of this application;
[0044] Figure 11 is a schematic diagram showing the change of the response content displayed on a lens according to an embodiment of this application;
[0045] Figure 12 is a schematic diagram of a sliding device on the temple of a pair of glasses provided in an embodiment of this application;
[0046] Figure 13 is a schematic diagram showing the change of the response content displayed on a lens according to an embodiment of this application;
[0047] Figure 14 is a schematic diagram showing the change of the response content displayed on a lens according to an embodiment of this application;
[0048] Figure 15 is a schematic diagram showing the change of the response content displayed on a lens according to an embodiment of this application. Detailed Implementation
[0049] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.
[0050] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. For example, "first response message" and "second response message," etc., are used to distinguish different response messages, not to describe a specific order of response messages.
[0051] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0052] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.
[0053] For example, Figure 1 shows a schematic diagram of the structure of a smart glasses according to an embodiment of this application. As shown in Figure 1, the smart glasses 100 may include: a processor 110, a memory 120, a camera 130, a communication module 140, a power supply 150, an input component 160, an audio module 170, a speaker 180, and a microphone 190.
[0054] The processor 110 is the computing and control core of the smart glasses 100. The processor 110 may include one or more processing units. For example, the processor 110 may include one or more of the following: application processor (AP), modem, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. In some embodiments, the processor 110 can extract and recognize user gestures from image data of the physical world captured by the camera 130, and determine the target object inquired about by the user based on the recognized gestures and image data. In other embodiments, the processor 110 can identify objects contained in the physical world from the image data of the physical world captured by the camera 130. In still other embodiments, the processor 110 may also generate response content based on the target object and the user's voice input, and present the response content in at least one modality.
[0055] The memory 120 may store programs that can be executed by the processor 110, enabling the processor 110 to perform at least some or all of the steps in the methods provided in this embodiment. The memory 120 may also store data. The processor 110 can read the data stored in the memory 120. Additionally, the memory 120 may also store an operating system. The memory 120 and the processor 110 can be configured separately. Alternatively, the memory 120 can be integrated into the processor 110. In this embodiment, the memory 120 may store AI programs. An AI program refers to a software system written using AI technology, which enables the smart glasses 100 to possess intelligent perception, learning, reasoning, and decision-making capabilities. These programs typically run on the processor of the smart glasses, processing data from various sensors and analyzing it according to preset algorithms and models to provide users with various intelligent services.
[0056] Camera 130 can be used to capture still images or videos, such as collecting image data from the physical world. Camera 130 may include a lens, a photosensitive element, an ISP, etc. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element may be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, and other formats. For example, as shown in FIG2, camera 130 may be configured on the frame of smart glasses 100. In this embodiment, when camera 130 is working, it can capture the physical world directly in front of the user as image data, which is then processed by processor 110 for image analysis and other processing. In some embodiments, smart device 100 may include one or N cameras 130, where N is a positive integer greater than 1.
[0057] The communication module 140 may include, but is not limited to, a wireless communication module. The communication module 140 can be applied to the smart glasses 100 using wireless local area networks (WLANs) (such as Wi-Fi), Bluetooth, GNSS, frequency modulation (FM), near field communication (NFC), infrared (IR), and other wireless communication solutions. For example, the communication module 140 can be used for the smart glasses 100 to communicate with other devices (such as smartphones, smartwatches, etc.) to complete data interaction. In some embodiments, the communication module 140 can be integrated into the processor 110 or configured separately from the processor 110.
[0058] The power source 150 can be used to power various components in the smart glasses 100. In some embodiments, the power source 150 can be a battery, such as a rechargeable battery.
[0059] The input component 160 can be a device for exchanging information between the user and the smart glasses 100. For example, the input component 160 can be, but is not limited to, a touch recognition module, a light recognition module, etc., capable of recognizing gesture types performed by the user, such as swiping, tapping, and long-pressing. For example, as shown in Figure 2, the input component 160 can be configured on the temple of the smart glasses 100. Of course, the input component 160 can also be configured on other devices besides the smart glasses 100; in this case, the other device can communicate with the smart glasses 100 via a network or other means.
[0060] The audio module 170 can be used to convert analog audio electrical signals into digital audio signals, and also to convert digital audio signals into analog audio electrical signals for output. The audio module 170 can, but is not limited to, transmit audio signals with the communication module 140. The audio module 170 can be used to encode and decode audio signals. In some embodiments, the audio module 170 can be located in the processor 110, or some functional modules of the audio module 170 can be located in the processor 110.
[0061] The loudspeaker 180, also known as a "horn", is used to convert analog audio electrical signals into sound signals.
[0062] The microphone 190, also known as a "microphone" or "voice transducer," is used to convert sound signals into analog audio electrical signals. The smart glasses 100 may include at least one microphone 190. The microphone 190 can collect sound from the environment in which the smart glasses 100 is located, such as the user's voice, and convert the collected sound into analog audio electrical signals. In some embodiments, the microphone 190 may be a microphone or a microphone array.
[0063] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the smart glasses 100. In other embodiments of this application, the smart glasses 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0064] The following is an introduction to the usage process of the Smart Glasses 100.
[0065] For example, Figure 3 shows a flowchart of a human-computer interaction method provided in an embodiment of this application. This method can be applied to smart glasses. In this human-computer interaction method, the smart glasses may or may not have a lens screen, depending on the actual situation, and is not limited here. The lens screen is a lens with display capabilities, which combines a micro-display screen with an optical lens, allowing the user to see the displayed content through the lens. As shown in Figure 3, the human-computer interaction method includes the following steps:
[0066] S301. When the AI program in the smart glasses is in a woke-up state, image data of the physical world is acquired through the camera on the smart glasses.
[0067] In this embodiment, users can wake up the AI program in the smart glasses in various ways. For example, they can wake it up using a specified wake word, or, as shown in Figure 8, they can wake it up by performing a preset action (such as a long press) on the input component of the smart glasses. When the AI program in the smart glasses is awake, the smart glasses can acquire image data of the physical world through its camera. Furthermore, when the AI program is awake, the smart glasses can remind the user that the AI program has entered a ready-to-work state through voice prompts, lens screen displays, etc. After being awakened, the AI program receives input in different ways, referred to as different modal inputs. For example, the AI program can have a dialogue modality and a visual modality. In the dialogue modality, the user provides input to the smart glasses through voice, and the AI program acquires the user's input by calling a microphone or other sound-receiving device. In the visual modality, the AI program acquires input information by calling the camera; that is, it uses the image captured by the camera as the input to the AI program. In one design, after the AI program on the smart glasses is awakened, the smart glasses can default to the visual modality and automatically acquire image data of the physical world through the camera. For example, referring to Figure 8, the smart glasses can acquire image data of the physical world 002 as observed by the user through the camera. In another design, after the AI program on the smart glasses is activated, the smart glasses can first enter a conversational mode. Then, the user can control the smart glasses to switch from the conversational mode to the visual mode by performing preset actions. Further, in the visual mode, the smart glasses begin to acquire image data of the physical world in front of the user through the camera.
[0068] S302. Extract and recognize user gestures from the image data acquired by the camera, and determine the target object based on the user gestures and image data.
[0069] In this embodiment, after acquiring image data, the smart glasses can perform two actions to determine the target object. For example, the target object is the object the user inquires about, or an object existing in the physical world. For instance, referring to Figure 8, the target object could be 003. The first action performed by the smart glasses is to recognize a preset user gesture (or "preset gesture") from the image data. This gesture recognition based on image data—that is, extracting and recognizing gestures from the graphic data acquired by the camera—works very naturally with the smart glasses when determining the target object, conforming to natural human habits and greatly reducing the perception of deliberate interaction during use. The second action performed by the smart glasses is to determine the target object from the image data. These two actions are described below.
[0070] 1) Recognize preset user gestures from image data
[0071] The preset user gesture can be a static gesture or a dynamic gesture. Static gestures can be recognized using a single frame image, while dynamic gestures, due to their duration, require recognition based on multiple consecutive frames. In this embodiment, the recognition elements of the preset gesture may include: ① three fingers (middle, ring, and little fingers) are folded together, and the index finger is extended; ② the index fingertip moves, and the trajectory is an arc, with the central angle of the arc greater than a preset value (e.g., 210 degrees); ③ the thumb is placed against the base of the index finger. Element ③ can also include a duration element recognition, making the holding time exceed a preset value (e.g., 1.5 seconds). In this embodiment, when elements ① and ② are simultaneously satisfied, or elements ① and ③ are satisfied, or the duration of element ① exceeds the preset value, the preset gesture is considered recognized. For example, as shown in Figure 4, which illustrates two gestures, both of which can be used to determine the target object 41 and can be used for target object determination at different granularities. The gesture shown in Figure 4(A) primarily uses finger pointing to determine the target object, and can be used for large-granularity target identification, such as distant buildings, paintings on walls, and books on tables. The gesture shown in Figure 4(B), where the index finger performs a selection-like action, is used to identify smaller-granularity target objects, such as an element within a painting on a wall or an equation in a book. Since the actual trajectory of the user's selected object is usually not a standard circular curve, the central angle can be determined by processing the index finger's movement trajectory to obtain a standard circle that most closely approximates the user's actual movement trajectory. The central angle can then be determined based on this standard circle.
[0072] For example, when recognizing a preset user gesture, algorithms such as contour recognition can be used to identify and extract images containing the user's hand (i.e., hand images) from image data. Then, gesture element recognition is performed on the hand image to determine whether the image data contains the preset user gesture. For instance, as shown in Figure 5, which illustrates a frame captured by the smart glasses' camera, recognizing the gesture in this image first involves identifying and extracting the hand image based on algorithms such as contour recognition. Then, gesture element recognition is performed on the hand image, identifying the following elements: ① three fingers are folded together, and the index finger is extended; ② the thumb is placed against the base of the index finger. These two elements match the elements of the preset gesture, thus recognizing the preset gesture and proceeding to the determination of the target object. In Figure 5, if duration recognition is to be added, assuming that the image frame shown in Figure 5 is the first recognition of the thumb touching the base of the index finger, the hand contour extraction and element ② recognition are performed on the image frames within the next 1.5 seconds respectively. If the image frames within the next 1.5 seconds that recognize element ② occupy the preset value (such as 90%) of all image frames, then the element recognition of duration is considered to be satisfied, that is, the preset gesture is considered to be recognized.
[0073] 2) Determine the target object from image data
[0074] After recognizing a preset user gesture, the smart glasses can determine the target object based on any frame of the recognized gesture and extract its image features. This way, even if the user subsequently releases the gesture, the target object can still be identified from the image after the gesture is released, based on its image features. Since performing a gesture takes only a short time, but speaking a question may take longer or occur after the gesture, recording the target object's image features after gesture recognition allows the user to return to a natural state after performing the gesture, without needing to maintain the gesture continuously, thus aligning with the body's natural habits. For example, assuming t0 < t1 < t2, at time t0 the user performs a gesture and asks a question, at time t1 the gesture is completed and the user returns to a natural state, and at time t2 the user ends the question. At this time, the target object can be determined by combining the user's gesture and any frame of image between time t0 and t1, and the features of the target object can be extracted from the image. Between time t1 and t2, the features of each object contained in the images during this period can be extracted and compared with the features of the target object extracted before time t1, so as to determine the target object from the images between time t1 and t2. Furthermore, assuming t0 < t1 < t2, if the user performs a gesture between t0 and t1, and the gesture is completed and the user returns to a natural state at t1, and the user asks a question between t1 and t2, then the target object can be identified by combining the user's gesture with any frame of the image between t0 and t1, and the features of the target object can be extracted from the image. Between t1 and t2, the features of each object contained in the images during this period can be extracted and compared with the features of the target object extracted before t1, so as to identify the target object from the images between t1 and t2, thus facilitating the answer to the user's question.
[0075] If the gesture is static (e.g., index finger extended, thumb pressed against the base of the index finger, and the thumb pressed against the base of the index finger for a duration exceeding a preset time), the target object can be determined based on any frame of the recognized static gesture, identifying the object pointed to by the index finger as the target object. For example, the smart glasses can first identify the object contained in any image frame of the static gesture (this can be done simultaneously with hand image recognition based on contour recognition); then, it identifies the target position of the index finger's tip in the static gesture to determine objects near the tip (free-moving end) of the index finger, checking if any object overlaps with the tip of the index finger, i.e., checking if an object exists at the target position where the index finger's tip is located. Since the smart glasses' camera is at roughly the same viewing angle as the human eye, any object overlapping with the tip of the index finger in the image data is the target object. For example, referring to Figure 5, the bullet-shaped building in Figure 5 overlaps with the tip of the index finger, indicating that the user is pointing to the bullet-shaped building, which is the target object. For example, to improve the accuracy of object recognition, the first duration for which the index fingertip is in the target position during a static gesture can be identified. When the first duration exceeds a preset duration, objects overlapping with the target position within the image can be identified as the target object. Alternatively, if no object overlaps with the index fingertip, objects near the index fingertip in the image data can be identified as the target object. For instance, in an image containing a static gesture, the object closest to the index fingertip is identified as the target object based on the distance between objects in the image. Of course, the duration for which the index fingertip is in the target position during the static gesture can also be considered. In some designs, the direction the index finger points can also be considered, and the target object can be determined based on this direction. For example, in image recognition, the direction the index finger points can be determined based on its length extension, and then the target object is identified from objects near the index fingertip based on whether the direction is shifted to the left or right and the distance of the object from the index fingertip. Specifically, if the direction the index finger points is to the left, the object closest to the tip of the index finger on the left side of the image is identified as the target object; if it points to the right, the object closest to the tip of the index finger on the right side of the image is identified as the target object. In other words, the target direction of the index finger in the static gesture can be determined first, and then the object in a specific image of the static gesture within the image data that is on the side of the target direction and closest to the tip of the index finger is identified as the target object. For example, as shown in Figure 6, in Figure 6(A), there is no object at the tip of the index finger. Based on the fact that the index finger can be clearly pointed, Figure 6(A) points to the left (left-leaning), and Figure 6(B) points to the right (right-leaning). In this case, the bullet-shaped building to the left of the index finger in Figure 6(A) can be identified as the target object, and the cuboid building to the right of the index finger in Figure 6(B) can be identified as the target object.It's understandable that when identifying nearby objects as target objects, a maximum error value can be set. This means only objects whose distance from the fingertip is less than the maximum error value will be considered as target objects; objects outside the maximum error value will be ignored. For example, the distance determination could be the minimum distance from any point on the edge of an object to the fingertip. Furthermore, if no overlapping objects exist, it can be considered that no target object has been identified, and no further response is given.
[0076] If the gesture is dynamic (e.g., the index finger extends, the fingertip moves, and the trajectory is an arc, the central angle of which is greater than a preset degree, etc.), the movement trajectory of the gesture is first determined based on multiple image frames of the continuous gesture. Then, the movement trajectory of the gesture is superimposed onto any one frame of the multiple image frames of the continuous gesture to determine the position of the movement trajectory in that frame. Finally, objects contained in a certain image of the image data can be identified, and the objects that overlap with the movement trajectory of the gesture are taken as the target objects, that is, the target objects are determined from the image frames based on the gesture trajectory. For example, when the index finger performs a selection, the selection trajectory of the index finger relative to the camera is determined, and then this selection trajectory is superimposed onto the image frame; the objects located within the selection trajectory are the target objects. It is understood that the target objects may not be completely within the selection trajectory.
[0077] The above scheme only describes identifying a single target object based on a single gesture. In some scenarios, users may need to specify multiple objects and ask questions to the smart glasses based on these objects. Therefore, in some designs, multiple target objects can be identified based on user gestures. Based on this, user gesture recognition is not merely the recognition of a single preset gesture, but rather the recognition of a series of consecutive gestures. For example, the recognition of consecutive gestures can be broadly categorized as follows:
[0078] a) Determine the image frames corresponding to a set of consecutive gestures: Consecutive gestures are a complete set of actions performed by the user. This set of actions is captured completely by the camera, and a set of image frames corresponding to the user raising their hand (the hand image is recognized in the image frame) to lowering their hand (the hand image disappears in the image frame) is determined.
[0079] b) Recognizing preset gestures based on image frames: This is basically the same as recognizing a single preset gesture as described above, that is, recognizing whether a preset gesture exists in each image frame.
[0080] c) Target Object Determination: The determination principle is basically the same as that for determining the target object when using a single gesture, as described above. The difference lies in the following: If the preset gesture uses the recognition of elements ① and ③, the object can be directly determined based on the gesture since the index finger has selected the preset object. However, if element ① or a combination of elements ① and ② is used, the user will maintain the preset gesture while moving towards multiple different objects. Since the user does not actually intend to determine the target object during this movement, gestures that interfere with the recognition of the preset gesture during movement need to be excluded. In one design, a duration-based recognition element can be added when recognizing the target object; that is, the target object will only be recognized based on the preset gesture if it is maintained for a preset time. In another design, two target objects are determined based on the target object recognized in the first image frame where the preset gesture is recognized, and the target object recognized in the last image frame where the preset gesture is recognized. Alternatively, the two designs can be combined: based on the target object recognized in the first image frame where the preset gesture is recognized, and the target object recognized in the last image frame where the preset gesture is recognized, a duration-based recognition element is added for the recognition of the target object during intermediate movement. It is understandable that during a user's continuous gestures, head movements may occur, thus changing the corresponding image frame. For example, as shown in Figure 7, a continuous gesture with a movement trajectory as indicated by the arrows shows the user maintaining the gesture. However, if the gesture is held for more than a preset time of 1.5 seconds at each of the three locations (buildings A, B, and C) indicated in Figure 7, then buildings A, B, and C are all identified as target objects. If the gesture is held for more than 1.5 seconds only at locations A and C, then buildings A and C are identified as target objects. For example, when there are multiple target objects, one target object can be identified based on a single preset gesture.
[0081] S303, Receive voice input from the user.
[0082] In this embodiment, the user can perform voice input during the gesture execution process. The voice input time can be shorter or longer than the gesture execution time. Furthermore, the user can also perform voice input first, followed by the gesture, depending on the actual situation; no specific limitation is made here. That is, there is no specific order in which S303 and S02 are executed. For example, in Figure 5, the user's voice input could be: "Introduce this building"; in Figure 7, the user's voice input could be: "What are these three buildings, and which well-known companies are located inside?". For example, if the voice input time is shorter than the gesture execution time—that is, when the user issues a command, and the gesture has not yet ended after the user's voice command, the smart glasses can still recognize the hand image in the received image frame. In this case, the smart glasses can wait for the gesture execution to finish before performing the next action based on the target object recognized by the gesture. For example, the criteria for determining the end of a gesture may be: ① no hand image / preset gesture is detected; ② the preset gesture is maintained for more than a preset time, and this preset time is longer than the time required to recognize the preset gesture. For example, if the duration of recognizing the preset gesture is determined to be 1.5s, then the preset time for the gesture to end may be 5s.
[0083] S304. Based on the target object and the user's voice input, generate response content and present the response content in at least one modality.
[0084] In this embodiment, when both the target object and the user's input voice are present, the smart glasses can use both as input to a neural network model. This allows the neural network model to perform semantic understanding based on the target object and user's input voice and output a response. The smart glasses can present the response in at least one modality, such as through voice broadcasting. Alternatively, when the smart glasses are equipped with lenses, the response can be displayed on a lens screen.
[0085] In this way, during the interaction between the user and the smart glasses, the user's gestures contained in the image data captured by the camera on the smart glasses can be used to identify the target object. This eliminates the need for the user to use verbal descriptions to identify the target object, reducing the difficulty of interaction and improving the user experience.
[0086] In the human-computer interaction method described in Figure 3 above, the target object is determined by gestures recognized from image data. This method, when used with smart glasses, is very natural in identifying the target object, conforming to natural human habits and greatly reducing the perception of deliberate interaction during use. Besides this method of determining the target object, users can also determine the target object through input components associated with the smart glasses. These input components can be configured on the smart glasses themselves. Alternatively, they can be configured on other devices (such as smartphones, smartwatches, etc.). In this case, the smart glasses and the device with the input component can communicate via a network, allowing the device to transmit user actions to the smart glasses. For example, when the device with the input component is a smartphone, the user can select the object by swiping or pinching on the smartphone screen. When the device with the input component is a smartwatch, the user can select the object by rotating the smartwatch crown. This human-computer interaction method will be described below.
[0087] For example, Figure 9 shows a flowchart of another human-computer interaction method provided in an embodiment of this application. This method can be applied to smart glasses. In this human-computer interaction method, the smart glasses are equipped with a lens screen. S901, S904, and S905 in Figure 9 can be found in the relevant descriptions of S301, S303, and S304 in Figure 3 above, and will not be repeated here. As shown in Figure 9, the human-computer interaction method includes the following steps:
[0088] S901. When the AI program in the smart glasses is in a woke-up state, image data of the physical world in front of the user is acquired through the camera on the smart glasses.
[0089] S902. Identify objects in the physical world from image data acquired by the camera and display the identifiers of the identified objects on the lens screen.
[0090] In this embodiment, after acquiring image data, the smart glasses can identify objects contained in the image data through neural network models to identify objects in the physical world. After identifying objects in the physical world, the smart glasses can display the identifiers of the identified objects on the lens screen. For example, the display position of an object's identifier on the lens screen can overlap with the object as seen by the user's naked eye within the user's field of vision, allowing the user to select an object by selecting its identifier. In other words, when the smart glasses enter visual mode, the lens screen can display the identifiers of objects in the physical world. Of course, image data can also be displayed synchronously on the lens screen, depending on the actual situation, and is not limited here. For example, when image data is not displayed on the lens screen, it is necessary to determine the specific display position of each object when the image data is displayed on the lens screen based on the timeline correspondence, and display the identifiers of each object at the corresponding positions on the lens screen. There is a correlation between the position of the object identifier on the lens screen and the position of the object associated with the object identifier in the physical world.
[0091] S903: Detect the user's object selection operation on the input component related to the smart glasses, and determine the target object based on the object selection operation.
[0092] In this embodiment, users can interact with the input components associated with the smart glasses to achieve gesture input. Common gestures include tapping, double-tapping, and swiping. Tapping and double-tapping are typically used to select / select a displayed object. For example, a user can swipe on the input component associated with the smart glasses to move the focus box between multiple object icons to select the desired object. For example, the user's object selection operation can be called an "object selection operation," which can be used to select an object from the object icons displayed on the lens screen. After selecting an object, the user can confirm the object as the target object by using a confirmation gesture (such as double-tapping) on the input component. Alternatively, the smart glasses can also determine the target object by the hovering time of the focus box. For example, if the focus box hovers over an object for more than a preset time, that object can be designated as the target object. In addition to selecting objects through the focus box, objects can also be selected using the cursor. In this case, the user can move the cursor displayed on the lens screen using the input component, and the smart glasses can directly determine the target object based on the cursor position and the object's position in the image data.
[0093] S904, Receive voice input from the user.
[0094] S905. Based on the target object and the user's voice input, generate response content and present the response content in at least one modality.
[0095] In this way, during the interaction between the user and the smart glasses, the image data collected by the camera on the smart glasses can be used to identify objects in the physical world and detect the user's input of object selection operations to determine the target object. This eliminates the need for users to use verbal descriptions to determine the target object, reducing the difficulty of interaction and improving the user experience.
[0096] In the human-computer interaction method described in Figure 3 or Figure 9 above, when the smart glasses display the response content on their lens screen, the user can also interact with the response content. This human-computer interaction method is described below.
[0097] For example, Figure 10 shows a flowchart of another human-computer interaction method provided in an embodiment of this application. This method can be applied to smart glasses. In this human-computer interaction method, the smart glasses are equipped with a lens screen. As shown in Figure 10, the human-computer interaction method includes the following steps:
[0098] S1001. Display the response content on the lens screen.
[0099] In this embodiment, the smart glasses can display response content on their lens screen. The response content can be text, and in some cases, it can include images / links. However, since smart glasses are a minimalist device, text content is typically displayed when showing the response; the specific display depends on the actual situation and is not limited here. For example, the response content can be displayed in a blank space, meaning a space that does not obscure the target object. For instance, as shown in Figure 11, Figure 11(A) shows a simplified schematic diagram of the response content displayed on the lens screen. Area 1101 shows the response content displayed on the lens screen, and its display position can be calculated according to a preset algorithm, which can be based on the image captured by the camera. As can be seen from Figure 11(A), the response content displayed in area 1101 does not obscure the target object 1102. It should be understood that the response content in this step can be the response content generated in Figure 3 or Figure 9.
[0100] S1002, The user's swipe operation on the input component related to the smart glasses is detected.
[0101] In this embodiment, the user can perform swipe operations on the input components associated with the smart glasses. Different swipe operations can be used to perform different functions, resulting in different visual presentations of the response or the addition of extra annotations, thereby enabling the user to obtain more information. For example, different swipe operations have different swipe data. Different swipe data can correspond to different functions. The swipe data can include one or more of the following: swipe direction, swipe distance, and number of swipes.
[0102] S1003. In response to the swipe operation, at least a portion of the content displayed on the lens screen is changed, wherein the changes to the content displayed on the lens screen include one or more of the following: zooming in / out of text / text boxes, highlighting / unhighlighting special text, displaying special text annotations, displaying additional annotation information, and displaying a summary of the response content.
[0103] In this embodiment, different sliding operations correspond to different functions. These functions may include one or more of the following: zooming in / out of text / text boxes, highlighting / unhighlighting special text, displaying special text annotations, displaying additional annotation information, and displaying a summary of the response content. These functions are described in detail below.
[0104] 1) Zooming in / out of text / text boxes
[0105] When the input component is configured on smart glasses, due to the limited space in the temples, only very simple interactions can be supported. Therefore, different functions can be defined by different sliding distances or by multiple slidings within a preset time. Specifically, when defining functions by different sliding distances, if the sliding distance is greater than S∈(0,S1], the text box is zoomed in / out. When defining different functions by multiple slidings within a preset time, the first sliding action zooms in / out. When the input component is configured on smart glasses, the direction closest to the lens screen is considered forward. Sliding forward (towards the lens screen) zooms in on the text / text box, and sliding backward (away from the lens screen) zooms out. Enlarging the text means increasing the font size. Zooming in / out refers to the area of the text box increasing / decreasing, specifically in both length and width. To avoid obscuring the target object as much as possible, only the length or only the width may change. The shape / aspect ratio of the text box can be changed. For example, referring to Figures 11 and 12, based on Figure 11(A), a forward swipe is performed on the input component 1201 configured on the smart glasses in Figure 12, or when the forward swipe distance is less than S1, the response content in area 1101 of Figure 11(A) can be switched to the response content in area 1101 of Figure 11(B). The font of the response content displayed in area 1101 of Figure 11(B) becomes larger, which in turn increases the area of the text box. For example, the mapping relationship between the forward swipe distance and the font can be defined by a function. For instance, the initially displayed response content is in 10-point font, and the font size increases / decreases by one size after each swipe distance.
[0106] 2) Highlight / Unhighlight special text
[0107] Within the response content, smart glasses can further explain certain text, such as by adding annotations or selecting text as new input to provide more information to the user. This text can be referred to as special text. For example, when performing a second swipe in a series of swipes, and the interval between the two swipes is less than a preset time, the special text can be highlighted. If special text is already highlighted, it is dehighlighted, depending on the direction of the swipe. Alternatively, when the single swipe distance S∈(S1,S2], the special text is highlighted / dehighlighted, depending on the direction of the swipe. In some designs, the smart glasses respond in real-time during a single swipe; that is, the corresponding function is executed as soon as the swipe reaches the threshold, rather than responding only after the gesture is released. Therefore, the text / text box zoom-in / zoom-out function may have already been executed before highlighting the special text. In other designs, the gesture command is only generated after the gesture is released. Therefore, if the single swipe distance is greater than S1 and less than S2, and the swipe is released before it is less than S2, the special text will be highlighted. The gesture then generates instructions for highlighting special text. For example, swiping forward towards the lens screen can highlight special text; if highlighted text is already displayed, the response stops. Swiping backward cancels the highlighting. Of course, both forward and backward swipes can have a set swipe distance range, or a combination of swipe count and swipe interval duration, etc. Different swipe directions, distance ranges, counts, and intervals can all correspond to different functions. For example, when swiping towards the lens screen, and... When the sliding distance is within a first distance range, or when sliding towards the lens screen and the sliding operation is multiple consecutive slidings, and the interval between adjacent slidings is less than a preset duration, special text in the response content is highlighted; when sliding away from the lens screen and the sliding distance is within a second distance range, or when sliding away from the lens screen and the sliding operation is multiple consecutive slidings, and the interval between adjacent slidings is less than a preset duration, the special text in the response content is dehighlighted. For example, as shown in Figure 13, Figure 13(A) shows the real-time response interface when the sliding distance is (S1, S2], responding to this... Using a swipe gesture, the smart glasses recognize special text based on the response content and highlight it. Highlighting can be done by adding shading, changing font size or color, or adding underlines to differentiate it from the original response content. For example, "Shenzhen Bay Sports Center" and "Shenzhen Bay Port" are highlighted. Figure 13(A) shows an example of adding a shading effect to indicate highlighting. In some designs, special text also has different types, each with different identifiers, such as different shading colors, allowing users to quickly identify the type of special text.Special text types can be predefined, such as proper nouns, nouns (including personal names, place names, biological names, etc.), special expressions (such as slang and idioms), and unclear expressions (which require interpretation in context).
[0108] For example, when highlighting specific text, different types of specific text can be selected using a swipe up or down gesture.
[0109] The first swipe down adds a focus box to the first special text to indicate that it is selected; a second swipe down adds a focus box to the second special text. In other words, after highlighting special text in the response content, a swipe gesture on the input component can be detected (either an up or down swipe). If the swipe gesture is a down swipe, a focus box is added to the first special text in the response content, and the position of the focus box is adjusted based on subsequent swipe gestures. If, after adding a focus box to a special text, the user performs a click (e.g., a single tap or double tap) on the input component, the smart glasses detect this click and display a target card on the lens screen. This target card can contain an explanation of the selected special text. Furthermore, the user can perform a second click on the input component; the smart glasses detect this click and remove the displayed target card from the lens screen. For example, continuing to refer to Figure 13, based on Figure 13(A), performing a swipe down gesture will add a focus box to the first special text "Shenzhen Bay Sports Center". Clicking it again will display a card 1302 below, explaining the Shenzhen Bay Sports Center. Double-clicking again will make the card below disappear.
[0110] 3) Display additional annotation information
[0111] The additional annotation information refers to a brief explanation of the special text. Continuing to slide forward (i.e., sliding towards the lens screen) until the sliding distance is greater than S2, or performing a third forward slide, or performing multiple consecutive slides with adjacent slide intervals less than a preset value, adds additional annotation information to the special text in the response content. This is equivalent to incorporating a summary of the special text into the existing response content, displayed directly in parentheses after the special text as an annotation. In other words, when the smart glasses detect a sliding gesture towards the lens screen on the input component, and the sliding distance is within the second distance range, or when they detect multiple consecutive slides on the input component with adjacent slide intervals less than a preset value, they can add annotation information to the special text in the response content. For example, continuing to refer to Figure 13, based on Figure 13(B), if the user's sliding distance on the input component is greater than S2, or after three consecutive slides (with adjacent slide intervals less than a preset value), the content shown in Figure 13(C) can be displayed. In Figure 13(C), the summary from card 1302 in Figure 13(B) is incorporated into the existing response content and displayed directly in parentheses after the special text as a note.
[0112] 4) Display a summary of the response content
[0113] If the sliding continues backward until the sliding distance is greater than S2, or a third backward sliding is performed, a summary of the response content is extracted and displayed. For example, if the smart glasses detect that the sliding operation is a slide away from the lens screen and the sliding distance is greater than a preset distance, or if the sliding operation is a slide away from the lens screen and is performed multiple times consecutively, and the interval between adjacent slides is less than a preset time, the summary of the response content can be displayed, and the display of the response content can be canceled. For example, as shown in Figure 14, after displaying the response content in Figure 14(A), after detecting a specific sliding operation, the interface shown in Figure 14(B) can be displayed. Comparing Figure 14(A) and (B), it can be seen that the response content displayed in area 1401 of Figure 14(A) has been switched to the content displayed in area 1402 of Figure 14(B). For example, the summary content can contain multiple shorter sentences. The summary sentences can be simplified from one or more sentences by removing secondary information; therefore, it can be said that one sentence in the summary content is related to multiple parts of the original response content. In some designs, multiple short phrases in the summary design can be given focus controls, which can be selected by swiping up or down. Selecting a phrase and then clicking or double-clicking it displays the original response content associated with that phrase. The display can be within a card of the response content or on an additional card. For example, as shown in Figure 15, which illustrates the interaction before and after expanding the associated response content of the first summary, the summary is automatically segmented, and each phrase is displayed separately with a focus box added. Figure 15(A) shows three summary phrases, and each phrase in area 1501 corresponds to a focus control. Users can select among the three summary phrases by swiping up or down. In Figure 15(A), the user can select the focus control corresponding to the phrase "China Resources Building, also known as Spring Bamboo Shoots," and after double-clicking, as shown in Figure 15(B), the original response content associated with the selected summary phrase is displayed below in a card 1502.
[0114] It should be understood that the swipe operations corresponding to the above-mentioned functions can be set according to the actual situation, and are not limited here. In addition, when the input component is configured on a device other than smart glasses (such as a smartphone, smartwatch, etc.), the user can perform the corresponding swipe operations on the relevant device. In this case, the functions of different swipe operations can also be determined according to the actual situation, and are not limited here.
[0115] In this way, after the smart glasses display the response content, users can interact with the smart glasses to view information related to the response content, thus improving the user experience.
[0116] It is understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. Furthermore, in some possible implementations, each step in the above embodiments may be selectively executed according to actual circumstances; it may be partially or fully executed, and no limitation is made here. In addition, steps in different embodiments can also be combined according to actual circumstances, and the combined solution is still within the protection scope of this application.
[0117] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0118] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0119] It is understood that the processor in the embodiments of this application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor may be a microprocessor or any conventional processor.
[0120] The method steps in the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0121] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0122] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A human-computer interaction method, characterized in that, Applications in smart glasses include: When the artificial intelligence (AI) program in the smart glasses is in a wake-up state, it acquires image data of the physical world through the camera on the smart glasses. If the image data contains a preset gesture made by the user, a target object is determined based on the preset gesture and the image data, wherein the target object is the object that the user inquired about in the physical world; Based on the target object and the first voice input by the user, generate response content; The response content is presented in at least one modality.
2. The method according to claim 1, characterized in that, The recognition elements of the preset gesture include one or more of the following: Bring the three fingers together and extend the index finger; Alternatively, the fingertip of the index finger moves, and the movement trajectory is an arc, the central angle of the arc being greater than a preset degree; Alternatively, place your thumb against the base of your index finger.
3. The method according to claim 1 or 2, characterized in that, The preset gesture is as follows: the index finger is extended, the thumb is placed against the base of the index finger, and the duration of the thumb being placed against the base of the index finger exceeds the preset duration.
4. The method according to claim 3, characterized in that, The step of determining the target object based on the preset gesture and the image data includes: Identify objects contained in a first image, wherein the first image is any frame of the image data that contains the preset gesture; Identify the target position of the index finger tip in the preset gesture in the first image, and determine the first duration for which the index finger tip in the preset gesture is at the target position; If the first duration is longer than a preset duration, the object that overlaps with the target location among the objects contained in the first image is taken as the target object.
5. The method according to claim 4, characterized in that, The method further includes: If there are no objects in the first image that overlap with the target location, and the first duration is longer than the preset duration, the object in the first image that is closest to the target location is taken as the target object.
6. The method according to claim 5, characterized in that, The step of selecting the object in the first image that is closest to the target location as the target object includes: Determine the target direction pointed to by the index finger in the preset gesture; The object that is located on one side of the target direction and is closest to the target position among the objects contained in the first image is taken as the target object.
7. The method according to claim 1 or 2, characterized in that, The preset gesture is: the index finger is extended, the fingertip of the index finger moves, and the movement trajectory is an arc, the central angle of the arc is greater than a preset degree.
8. The method according to claim 7, characterized in that, The step of determining the target object based on the preset gesture and the image data includes: The movement trajectory of the preset gesture is determined in the multiple consecutive images contained in the image data, and the trajectory position of the movement trajectory in any one of the multiple consecutive images is determined. Identify objects contained in a first image, wherein the first image is any frame of the image data; The object that overlaps with the trajectory position among the objects contained in the first image is taken as the target object.
9. The method according to any one of claims 3-8, characterized in that, After identifying the target object, the process also includes: Extract the object features of the target object from the target image containing the target object from the image data; Based on the object characteristics, the target object is determined from images acquired after the target image in the image data.
10. The method according to any one of claims 1-9, characterized in that, There are multiple target objects.
11. The method according to any one of claims 1-10, characterized in that, Presenting the response content in at least one modality includes: The response content is displayed on the lens screen of the smart glasses.
12. The method according to claim 11, characterized in that, After displaying the response content on the lens screen of the smart glasses, the following is also included: The user's swipe operation on the input component associated with the smart glasses was detected; In response to the swipe operation, at least a portion of the content displayed on the lens screen is changed, wherein the changes to the content displayed on the lens screen include one or more of the following: zooming in / out of text / text boxes, highlighting / unhighlighting special text, displaying special text annotations, displaying additional annotation information, and displaying a summary of the response content.
13. The method according to claim 12, characterized in that, The response to the swipe operation, changing at least a portion of the content displayed on the lens screen, includes: When the sliding operation is a sliding motion toward the lens screen, the text / text box is enlarged; When the sliding operation is a sliding motion away from the lens screen, the text / text box shrinks.
14. The method according to claim 12, characterized in that, The response to the swipe operation, changing at least a portion of the content displayed on the lens screen, includes: When the sliding operation is a slide toward the lens screen and the sliding distance is within a first distance range, or when the sliding operation is a slide toward the lens screen and the sliding operation is a series of consecutive slides, and the interval between adjacent slides is less than a preset time, the special text in the response content is highlighted. If the sliding operation is a sliding away from the lens screen and the sliding distance is within the second distance range, or if the sliding operation is a sliding away from the lens screen and the sliding operation is a continuous sliding multiple times, and the interval between adjacent sliding is less than a preset time, then the special text in the response content will be de-highlighted.
15. The method according to claim 14, characterized in that, Following the highlighting of specific text within the response content, the following is also included: A swipe gesture is detected on the input component, the swipe gesture being either an up swipe gesture or a down swipe gesture; If the swipe gesture is the first downward swipe gesture, add a focus box to the first special text in the response content, and adjust the position of the focus box based on the swipe gestures following the first downward swipe gesture.
16. The method according to claim 15, characterized in that, After adding a focus box to the target special text in the response content, the following is also included: A click operation was detected on the input component; In response to the click operation, a target card is displayed on the lens screen, the target card being used to carry an explanation of the target-specific text.
17. The method according to claim 16, characterized in that, After displaying the explanation of the target special text on the lens screen, the method further includes: A second click operation was detected on the input component; In response to the second click operation, the target card is removed from the lens screen.
18. The method according to claim 14, characterized in that, Following the highlighting of specific text within the response content, the following is also included: A swipe gesture toward the lens screen is detected on the input component, and the swipe distance is within a second distance range; or, multiple consecutive swipes are detected on the input component, and the interval between adjacent swipes is less than a preset value. Add annotation information to special text in the response content.
19. The method according to claim 12, characterized in that, The method of changing at least a portion of the content displayed on the lens screen in response to the sliding operation further includes: If the sliding operation is a sliding away from the lens screen and the sliding distance is greater than a preset distance, or if the sliding operation is a sliding away from the lens screen and is a continuous sliding multiple times, and the interval between adjacent sliding is less than a preset time, then a summary of the response content is displayed, and the display of the response content is canceled.
20. A type of smart glasses, characterized in that, include: At least one memory for storing programs; At least one processor for executing the program stored in the memory; When the program stored in the memory is executed, the processor is used to execute the method as described in any one of claims 1-19.
21. A computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to perform the method as described in any one of claims 1-19.
22. A computer program product, characterized in that, When the computer program product is run on a processor, the processor causes the processor to perform the method as described in any one of claims 1-19.