System and method for identifying page turns using a portable device

CN122804207APending Publication Date: 2026-09-22KIBIM LEARNING INC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202480057375.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-06
Filing Date
2024-07-05
Publication Date
2026-09-22

Smart Images

  • Figure CN122804207A_ABST
    Figure CN122804207A_ABST
Patent Text Reader

Abstract

Systems and methods are described in which a light beam generated by a handheld device is used to point to a printed page containing visual content, e.g., displayed in a book or magazine. The page being pointed to can be identified using computer vision applied to an image acquired by a camera of the device pointed in the same direction as the light beam. Upon identifying that the user has turned to a new page or page section, the device can perform one or more actions. Such actions can augment the visual content of the printed book to include auditory elements, additional visual components, and / or haptic stimuli implemented by the handheld device. Systems and methods can provide a simple and intuitive method for human-computer interaction that can be particularly suitable for children and other learners.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application is a continuation in part of co-pending application serial number 18 / 579,885, filed March 6, 2024, which is a continuation of application serial number 18 / 220,738 (now U.S. Patent No. 11,989,357), filed July 11, 2023; and is a continuation in part of co-pending application serial number 18 / 382,456, filed October 20, 2023, which is a continuation in part of application serial number 18 / 220,738 (now U.S. Patent No. 11,989,357), the entire disclosure of which is expressly incorporated herein by reference. Technical Field

[0002] This application relates in its entirety to systems and methods for an individual to identify (e.g., in a book) a displayed object or location within a page using a beam of light projected from a handheld electronic device. Additionally, by monitoring images acquired by the device's camera, the device can identify and automatically perform one or more interactive actions whenever the user turns to a new page. While handheld devices can be used by anyone, they are particularly well-suited for use by young children or learners using simple interactive signals that lack the required precise manual dexterity and / or understanding of screen-based interactive sequences. The systems and methods described herein employ techniques from the fields of mechanical design, electronic design, firmware design, computer programming, optics, computer vision (CV), ergonomics (including child safety), human motion control, and human-computer interaction. The systems and methods can provide users (especially children or learners) with a familiar machine interface to instinctively and / or confidently indicate locations within other displays of a book or printed content, and automatically perform one or more "new page" actions whenever the user turns to a new page. Background Technology

[0003] In recent years, the world has become increasingly reliant on portable electronic devices, which have become more powerful, sophisticated, and useful to a wide range of users. However, while children may quickly adopt certain aspects of electronic devices designed for more experienced users, younger children may benefit from access to small, lightweight, colorful, fun, information-rich, ergonomically designed for children (including child safety), and easy-to-use interactive electronic devices. The systems and methods disclosed herein utilize advances in the field of optics, including visible light (i.e., often referred to as “lasers”) pointers, motion sound and / or vibration generation (using miniature coils, piezoelectric elements, and / or tactile units), portable displays, inertial measurement units (sometimes also referred to as inertial motion units), and telemetry.

[0004] Visible light pointers (also known as "laser pens") commonly used in commercial and educational environments typically generate beams from laser diodes (i.e., PIN diodes) that have undoped intrinsic (I) semiconductor regions between p (P) and n (N) type semiconductor regions. Such coherent and collimated light sources are generally considered safe within specified power levels and when operated correctly. Furthermore, if pointed at the eye, the corneal reflex (also known as a blink or eyelid reflex) ensures an unconscious aversion to the bright light (and foreign object).

[0005] However, incoherent light-emitting diode (LED) sources can be used to achieve further eye safety. Such incoherent sources (so-called “point source” LEDs) can be collimated using sophisticated optics (e.g., including so-called “pre-collimation”) to produce a beam with minimal and / or controlled divergence. If desired, point source LEDs can also generate a beam consisting of a range of spectral frequencies (i.e., compared to the dominant monochromatic light produced by a single laser).

[0006] Loudspeakers associated with televisions, theaters, and other fixed locations typically employ one or more electromagnetic moving coils. In handheld and / or mobile devices, similar electromagnetic coil methods and / or piezoelectric (sometimes referred to as "buzzers") designs can be used to generate vibrations in miniature loudspeakers. Vibrations (e.g., those associated with alarms) can also be generated by haptic units (also known as kinematic communication). Haptic units typically employ eccentric (i.e., unbalanced) rotating masses or piezoelectric actuators to produce audible and / or perceptible vibrations (especially at the lower end of the audio spectrum).

[0007] Visual displays or indicators can consist of any number of monochromatic or multicolor, addressable light sources or pixels. Displays range from those with a single light source (e.g., illuminating a sphere via waveguide transmission) to those capable of displaying a single digit (e.g., a seven-segment display) or alphanumeric characters (e.g., a five-pixel by eight-pixel array), to high-resolution screens with tens of millions of pixels. Regardless of scale, displays are typically implemented as: 1) a two-dimensional array of light sources (most commonly some form of light-emitting diode (LED), including organic LEDs (OLEDs)), or 2) two polarizing glass plates sandwiching liquid crystal material (i.e., forming a liquid crystal display, LCD), which responds to an electric current to allow light of different wavelengths from one or more illumination sources (i.e., backlight) to pass through.

[0008] Inertial measurement unit (IMU), accelerometer and / or magnetometer tracking may combine any or all of the following: 1) a linear accelerometer that measures the force generated during movement (i.e., governed by Newton's second law of motion) on up to three axes or dimensions, 2) gyroscope-based sensing of rotational rate or velocity on up to three rotational axes, 3) a magnetometer that measures magnetic fields (i.e. magnetic dipole moments) including fields generated by the Earth, and / or 4) the Earth's gravity (including gravitational orientation) by measuring forces on its internal mass.

[0009] Advances in electronic devices (i.e., hardware), standardized communication protocols, and the allocation of dedicated frequencies within the electromagnetic spectrum have led to the development of a wide range of portable devices capable of wireless communication with other nearby devices and large-scale communication systems, including the World Wide Web and the metaverse. Considerations for which protocols (or combinations of available protocols) to employ in such portable devices include power consumption, communication range (e.g., from a few centimeters to hundreds of meters and beyond), and available bandwidth.

[0010] Currently, Wi-Fi (e.g., based on the IEEE 802.11 series of standards) and Bluetooth (managed by the Bluetooth Special Interest Group) are used in many portable devices. Less common and / or older communication protocols in portable devices in home environments include Zigbee, Z-wave, and cellular or mobile phone-based networks. Generally speaking (i.e., with many exceptions, especially considering newer standards), Wi-Fi offers greater range, greater bandwidth, and a more direct path to the internet compared to Bluetooth. Bluetooth, on the other hand, (including Bluetooth Low Energy (BLE)) offers lower power consumption, a shorter operating range (which can be advantageous in some applications), and less complex circuitry to support communication.

[0011] Advances in miniaturization, power reduction, and increased complexity of electronic devices—including those applied to displays, microelectromechanical systems (MEMS) including inertial measurement units (IMUs), and remote communication—have revolutionized the mobile device industry. These portable devices have become increasingly sophisticated, allowing users to simultaneously communicate, interact, geolocate, monitor exercise, track health, be alerted to danger, capture video, and execute financial transactions. Systems and methods that facilitate simple and intuitive interaction with handheld (or other body-operated) pointing devices could be useful. Summary of the Invention

[0012] In light of the foregoing, this paper provides systems and methods for describing lightweight, easy-to-use, and intuitive handheld devices that are particularly well-suited for machine-based interaction by children or other learners. While the devices may be partially accepted as toys or “friends,” the computational flexibility embedded within them allows them to be used as a means for embodied learning, emotional support, cognitive development, facilitating communication, expressing creativity, play, developing mindfulness, and / or enhancing imagination.

[0013] Throughout this description, the portable device is referred to as "handheld." While the device can be manipulated by the user's hands, alternatively or additionally (e.g., at different times), it can be manipulated by other parts of the user's body, including the head, wrists, arms, shoulders, legs, or chest. Attachment of the device to a body part can be aided by one or more support components, such as a headband, wristband, or chest strap. Manipulation by other parts of the body may, for example, free the user's hands to perform other tasks such as turning pages (e.g., using one or both hands), or while holding a young child while manipulating a book.

[0014] For example, handheld devices can assist in areas related to basic reading, literacy, conversational reading, CROWD (Complete, Recall, Open-ended, WH-prompt [where, when, why, what, who], and Extend] questioning, mathematics, and understanding of science and technology. Additionally, devices can help bridge learners' zone of proximal development (ZPD) by providing ubiquitous (machine-based) educational support and / or guidance. Transforming the content of pages (e.g., within books or magazines) into visual, auditory, and / or tactile experiences can significantly aid in the acquisition of new knowledge, skills, and memory. Furthermore, portable, lightweight, and “fun” handheld devices can stimulate physical movement in children (and adults), including motor and kinesthetic activities.

[0015] According to one aspect, systems and methods are provided for an individual to select an object or location within viewable content on a page, such as a book or magazine. As described in more detail below, within the scope of this description, the term "page" refers to any substantially two-dimensional surface capable of displaying viewable content. Page content may include any combination of text, symbols, drawings, and / or images; and may be displayed in color, grayscale, or black and white.

[0016] An individual may use a beam of light emitted from a handheld device to select (i.e., point, identify, and / or indicate) objects and / or locations within a page. A camera within the handheld device, pointed in the same direction as the beam, can acquire an image of the area of ​​the page being pointed at. The image may include reflections of the beam (e.g., containing incident light reflected from the page), or the beam may be turned off (e.g., momentarily) when acquiring the image (e.g., thus allowing the page content to be imaged without interference from beam reflections).

[0017] In either case, the position of the light beam within the image acquired by the camera can be determined based on the close co-location (and joint movement) of the camera and the light beam within the body of the handheld device and their pointing in the same direction. For example, the position of the light beam within the image acquired by the camera can be calculated based on the beam pointing and the geometry and / or calibration process of the camera imaging to empirically identify the location of the light beam reflection within the image acquired by the camera.

[0018] The selection made by the handheld device user can be signaled using any of a variety of indication methods employing one or more sensing elements of the handheld device. A device switch (such as a button, contact switch, or proximity sensor) can be used to indicate the selection (and optionally control the beam, such as turning it on or off). Alternatively, keywords, phrases, or sounds generated by the user and sensed by the device's microphone can be identified to indicate the selection (and beam control).

[0019] Alternatively or additionally, user controls of the orientation and / or movement of the handheld device sensed by the embedded IMU (e.g., movement gestures, tapping the handheld device, or tapping the handheld device against another object) can be used to indicate selection. Within the signaling mechanism that generates movement of the handheld device, images captured by a camera just before any movement can be used during the process to identify the selected object and / or location within the page.

[0020] In another aspect of the system and method, a processor within the handheld device can acquire a predetermined interactive page layout (which may include one or more object templates) of a set of pages that an individual might view (e.g., a book, a magazine). Using computer vision (CV) methods (e.g., neural networks, machine learning, transformers, generative artificial intelligence (AI), and / or template matching), a match can be determined between an image acquired by a camera and the page layout or one or more object templates. CV-based matching can identify the page that the device user is viewing (e.g., within a book or magazine) and the location, object, word, or target within the page pointed to by the beam.

[0021] Optionally, page and content selection can be performed in two stages, where a context object (e.g., book cover, title, word or phrase, printed object, drawing, real-world object) can be selected first to assign a "context" (e.g., specific book or book chapter, type of book or magazine, subject area, skill requirements) to subsequent selections. A page layout including object templates relevant only to the identified context is then considered for CV-based matching with images acquired by the camera.

[0022] Compared to global classification methods used to identify objects within an image, CV and / or AI methods based on matching images acquired by the camera with a finite set of predetermined page layouts and / or object templates can reduce the computational resources required (e.g., on portable handheld devices) and promote greater accuracy when identifying content pointed to using a beam of light.

[0023] A pre-defined interactive page layout may include (or refer to an additional dataset including the following) interactive actions and / or additional content that may be implemented as a result of selecting an object, location, or area within the page. Examples of such interactive actions include word pronunciation, sound effects, displaying the spelling of an object (e.g., text, not shown), broadcasting a speech element of the selected object, additional story content, questions about the page content, rhythmic features associated with the selected object (e.g., which may form the basis for tactile or vibratory stimulation of the device user's hand), rewarding or comforting auditory feedback when a specific choice is made, related additional or sequential content, etc.

[0024] Based on interactive content within a predefined database of interactive page layouts, one or more visual, acoustic, and / or tactile actions can be performed on the handheld device itself. Alternatively or additionally, selections, attributes, and timing of selections can be communicated to one or more external processors, which can then record interactions (e.g., for educational and / or parental monitoring) and / or perform additional and / or supplementary interactions involving the handheld device user. Automatic recording of reading engagement, content, progress, and / or comprehension (e.g., including comparisons with other children of similar ages) can provide insights into a child's emotional and cognitive development.

[0025] In another aspect of the device and method, during interactions involving books or magazines, the device user can periodically turn pages (and guide the device's beam) to focus attention on the content of a new page, a new page area, or even a new page within a different book. The device can determine this shift in user focus by detecting the transition to the new page within one or more images captured by the device's camera (aligned with the beam). The device can be handheld and manipulated by the user's hand, or attached to other parts of the user's body, such as the wrist, arm, head, or chest. By continuously monitoring images captured by the camera, the user's intention to transition to the new page can be determined in real time without other input patterns (e.g., pressing a button) or user signaling instructions.

[0026] Additionally, the device may automatically perform one or more actions as a result of detecting such a shift in user focus. These actions may include: turning off the device beam; indicating that a new page has been detected via one or more auditory, visual, and / or haptic cues; playing a page title, page text, one or more descriptions of page objects, and / or a new page number on the device speaker; and so on. Once the "new page" introduction action is complete, the beam may be turned back on (if necessary), and the user may continue to interact with the content of the new page in exploration mode.

[0027] Alternatively or additionally, buttons (or other indication methods, such as voice commands sensed by the device's microphone or device gestures sensed by the IMU) may be available (e.g., at any time) to indicate the user's intention to shift focus to a new page. Continuously determining user intent based on images acquired by the camera and / or having a always-available "new page" indicator (e.g., a button) and automatically executing one or more actions associated with any new page can simplify and accelerate device interaction. Automatically initiating one or more actions after determining that the user's focus has shifted to a new page can further reduce the number of guiding prompts (which can sometimes become repetitive) issued by the device.

[0028] In summary, once the beam of light is directed at an object or location within the page, interaction facilitated by a handheld device can help bring printed or displayed content to life by adding auditory elements, additional visual components, and / or tactile stimulation felt by the hand of the handheld device user. Enhancing printed content with interactive sequences that include real-time feedback relevant to the content not only provides ubiquitous, machine-based guidance while reading but can also be "fun" and / or help maintain emotional engagement (rather than overwhelming) within a learning environment.

[0029] According to an example, a method is provided for performing an action based on a specified page location selected by a person using a handheld device, the handheld device comprising: a device processor; a device light source configured to generate a projected beam that produces one or more light reflections from one or more visible objects viewable by the person; and a device camera aligned such that its field of view includes the reflection locations of the one or more beam reflections and operatively coupled to the device processor, the method comprising: acquiring one or more predetermined interactive page layouts by the device processor; acquiring a camera image by the device camera when the handheld device is manipulated such that the projected beam is directed from the device light source to the specified page location; calculating a positioned image by the device processor based on a match between the camera image and the one or more predetermined interactive page layouts; identifying the specified page location by the device processor based on the reflection locations within the positioned image; and performing the action by one or both of the device processor and a remotely connected processor at least partially based on the specified page location within the one or more predetermined interactive page layouts.

[0030] According to another example, a method is provided for performing an action based on a specified page location within an identified context selected by a person using a handheld device, the handheld device comprising: a device processor; a device light source configured to generate a projected beam that produces one or more light reflections from one or more visible objects that can be viewed by the person; and a device camera aligned such that its field of view includes the reflection locations of the one or more light reflections and operatively coupled to the device processor, the method comprising: acquiring one or more predetermined context templates by the device processor; acquiring a first camera image by the device camera when the handheld device is manipulated such that the projected beam is directed from the device light source to the context object; and performing an action based on the specified page location within an identified context selected by a person using a handheld device, the handheld device comprising: acquiring one or more predetermined context templates by the device processor; acquiring a first camera image by the device camera when the handheld device is manipulated such that the projected beam is directed from the device light source to the context object; and performing an action by the device processor ... The process involves: calculating the identified context by matching a first camera image with one or more predetermined context templates; obtaining one or more predetermined interactive page layouts associated with the identified context by the device processor; acquiring a second camera image by the device camera when the handheld device is manipulated such that a projected beam is directed from the device light source to the designated page location; calculating a localized image by the device processor based on the matching of the second camera image with the one or more predetermined interactive page layouts; identifying the designated page location by the device processor based on the reflection position within the localized image; and performing the action by one or both of the device processor and a remotely connected processor, at least partially based on the designated page location within the one or more predetermined interactive page layouts.

[0031] According to another example, a method is provided for performing an action based on a specified page location selected by a person using a handheld device, the handheld device comprising: a device processor; a device light source configured to generate a projected light beam that produces one or more light beam reflections from one or more visible objects that can be viewed by the person; and a device camera aligned such that the camera's field of view includes the reflection locations of the one or more light reflections and operatively coupled to the device processor, the method comprising: acquiring one or more predetermined interactive page layouts by the device processor; and when the handheld device is manipulated such that the projected light beam is directed from the device light source to the specified page location. The device acquires two or more camera images; the device processor determines that the image movement measured in the two or more camera images is less than a predetermined movement threshold, and the two or more camera images are acquired over an acquisition time greater than a predetermined dwell time threshold; the device processor calculates a positioned image based on the matching of the camera images with one or more predetermined interactive page layouts; the device processor identifies the designated page position based on the reflection position within the positioned image; and one or both of the device processor and a remotely connected processor perform the action at least partially based on the designated page position within the one or more predetermined interactive page layouts.

[0032] According to yet another example, a method is provided for performing an action based on a specified page location selected by a person using a handheld device, the handheld device comprising: a device processor; a device light source configured to generate a light beam that produces one or more light reflections from one or more visible objects that can be viewed by the person; a device camera aligned such that the camera's field of view includes the reflection locations of the one or more light reflections and operatively coupled to the device processor; and a device switch operatively coupled to the device processor, the method comprising: obtaining one or more predetermined interactive page layouts by the device processor; and determining by the device processor... The device switch is set to a first state; when the handheld device is manipulated such that the projected beam is directed from the device light source to the designated page position, the device processor determines that the device switch is in a second state; the device camera acquires a camera image; the device processor calculates a positioned image based on the matching of the camera image with one or more predetermined interactive page layouts; the device processor identifies the designated page position based on the reflection position within the positioned image; and one or both of the device processor and a remotely connected processor perform the action at least partially based on the designated page position within the one or more predetermined interactive page layouts.

[0033] According to yet another example, a method is provided for performing an action based on a specified page location selected by a person using a handheld device, the handheld device comprising: a device processor; a device light source configured to generate a light beam that produces one or more light reflections from one or more visible objects that can be viewed by the person; a device camera aligned such that the camera's field of view includes the reflection locations of the one or more light reflections and operatively coupled to the device processor; and a device switch operatively coupled to the device processor, the method comprising: acquiring one or more predetermined interactive page layouts by the device processor; determining by the device processor that the device switch is in a first state; and when the handheld device is in a first state... When the device is manipulated such that a projected beam is directed from the light source to the designated page position, the device camera acquires one or more camera images; the device processor determines that the device switch is in a second state; the device processor isolates the most recently stabilized camera image from the one or more camera images; the device processor calculates a positioned image based on the matching of the most recently stabilized camera image with the one or more predetermined interactive page layouts; the device processor identifies the designated page position based on a predetermined beam pointing position within the positioned image; and one or both of the device processor and a remotely connected processor perform the actions at least partially based on the designated page position within the one or more predetermined interactive page layouts.

[0034] According to an example, a handheld device is provided, the handheld device comprising: a device body configured to be operated by a device user's hand; electronic circuitry within the device body, the electronic circuitry including a device processor; a device beam source configured to emit a beam of light away from the device body to generate a beam image reflection on a viewable surface, the beam image reflection being focused at a predetermined distance from the viewable surface; and a device camera operatively coupled to the device processor and aligned such that the field of view of the device camera includes the beam image reflection, wherein the handheld device is configured to: generate the beam image reflection on the viewable surface, the beam image reflection being focused at the predetermined distance to encourage the device user to position the handheld device at the predetermined distance from the viewable surface; and acquire camera images.

[0035] According to another example, a method is provided for encouraging a device user to position a handheld device at a predetermined distance from a viewable surface, wherein the handheld device includes: a device body configured to be manipulated by the user's hand; electronic circuitry within the device body, the electronic circuitry including a device processor; a device beam source configured to generate a beam image reflection, the beam image reflection being focused at the predetermined distance; and a device camera operatively coupled to the device processor and aligned such that the device camera's field of view includes the beam image reflection, the method comprising: generating the beam image reflection on the viewable surface by the device beam source, the beam image reflection being focused at the predetermined distance to encourage the user to position the handheld device at the predetermined distance from the viewable surface; and acquiring a camera image by the device camera.

[0036] According to yet another example, a method is provided for encouraging a device user to position a handheld device at a predetermined distance from a viewable surface, wherein the handheld device includes: a device body configured to be manipulated by the user's hand; electronic circuitry within the device body, the electronic circuitry including a device processor; a device beam source configured to generate a beam image reflection on the viewable surface; one or more movable reflective surfaces within a device beam path originating from the device beam source, each operatively coupled to the device processor; and a device camera operatively coupled to the device processor and aligned such that the device camera's field of view includes the beam image reflection, the method comprising: generating the beam image reflection on the viewable surface by the device beam source; moving the one or more movable reflective surfaces by the device processor to focus the beam image reflection at the predetermined distance to encourage the user to position the handheld device at the predetermined distance from the viewable surface; and acquiring a camera image by the device camera.

[0037] According to an example, a device is provided, the device comprising: a device body configured to be operated by a device user; electronic circuitry within the device body, the electronic circuitry including a device processor; a device light source fixed to the device body and configured to emit projected light away from the device body to generate light reflections on a viewable page that can be viewed by the device user; and a device camera fixed to the device body, the device camera including a field of view encompassing some or all of the light reflections, wherein the device processor is configured to: acquire a first camera image via the device camera; identify a first page layout based on a match between the first camera image and one of two or more predetermined page layouts; acquire a second camera image via the device camera after acquiring the first camera image; identify a second page layout based on a match between the second camera image and one of the two or more predetermined page layouts; determine that the second page layout is different from the first page layout; and perform an action.

[0038] According to another example, a device is provided, the device comprising: a device body configured to be operated by a device user; electronic circuitry within the device body, the electronic circuitry including a device processor; a device light source fixed to the device body and configured to emit projected light away from the device body to generate light reflection on a viewable page that can be viewed by the device user; and a device camera fixed to the device body, the device camera including a field of view encompassing some or all of the light reflection, wherein the device processor is configured to: acquire a first camera image via the device camera; identify a first page within the first camera image; acquire a second camera image via the device camera after acquiring the first camera image; identify a second page within the second camera image; determine that the second page is different from the first page; and perform an action.

[0039] According to another example, a device is provided, the device comprising: a device body configured to be operated by a device user; electronic circuitry within the device body, the electronic circuitry including a device processor; a device switch fixed to the device body; a device light source fixed to the device body and configured to emit projected light away from the device body to generate light reflection on a viewable page that can be viewed by the device user; and a device camera fixed to the device body, the device camera including a field of view encompassing some or all of the light reflection, wherein the device processor is configured to: acquire a first camera image via the device camera; identify a first page within the first camera image; acquire a first switch state via the device switch; acquire a second switch state via the device switch after acquiring the first switch state; determine that the second switch state is different from the first switch state; acquire a second camera image via the device camera; identify a second page within the second camera image; determine that the second page is different from the first page; and perform an action.

[0040] Other aspects and features, including the needs and uses of the invention, will become apparent from the following description taken in conjunction with the accompanying drawings. Attached Figure Description

[0041] A more complete understanding can be obtained by referring to the specific embodiments when considered in conjunction with the following illustrative drawings. In the drawings, the same reference numerals denote the same elements or actions throughout. The presented examples are illustrated in the drawings, wherein: Figure 1 An exemplary manipulation of a handheld device by a child is illustrated, which is used to identify and select a dog in a scene from a picture book (which includes several additional characters and objects) by pointing a beam of light emitted from the device at the shape of a dog.

[0042] Figure 2 It shows in Figure 1 In the exemplary scene presented, an overlay of images acquired by a camera that has been matched with the page layout using computer vision is shown, wherein the beam pointing position (indicated by the crosshair target) may be known within the camera's field of view.

[0043] Figure 3 Examples of parameters that define the positioning of a camera's field of view within a page layout are illustrated, including horizontal and vertical references (e.g., corners) of the image acquired by the camera, positioning (x, y), orientation (θ), and magnification (x, y). m ).

[0044] Figure 4This is an exemplary interconnection layout of components within a handheld device (where some components may not be used during some applications), illustrating the primary direction of information flow relative to the bus structure that forms the backbone of the electronic circuitry.

[0045] Figure 5 These are electronic schematic diagrams and light diagrams illustrating exemplary elements for the generation, light path, and detection of a feline object selected on a page of a book using a light beam emitted from a handheld device via a camera.

[0046] Figure 6 This is an exploded view of an exemplary handheld device showing the position, pointing direction, and relative size of the beam source and camera.

[0047] Figure 7 The flowchart illustrates exemplary steps in which an image of a shoe is pointed to using a beam of light generated by a handheld device, which is located within an interactive page layout that then provides auditory feedback to the device user.

[0048] Figure 8 The flowchart illustrates an exemplary two-stage process, in which a beam is first used to select a context (in this case, the word "cat" in the animal list), and subsequent handheld device interactions are limited to identifying the location of the beam within the page in relation to the selected context.

[0049] Figure 9 This is an exemplary flowchart in which buttons turn the light beam on and off, and signal the selection when a selection is made from a collection of clothing, thereby generating auditory feedback related to shoes.

[0050] Figure 10 This is an exemplary diagram illustrating the formation of a structured light pattern that produces an image of a reflected beam of light that is perceived by the device user as focused at approximately a target working distance between the handheld device and the reflective surface (i.e., forming a smiley face), and elsewhere as a blurred image (i.e., out of focus).

[0051] Figure 11 This is a flowchart illustrating exemplary steps for performing one or more actions based on a machine determining whether a beam of light pointed at the device is focused and based on whether the beam is focused (i.e., the handheld device is approximately at a desired distance from the viewable surface). Detailed Implementation

[0052] Before describing the examples, it should be understood that the invention is not limited to the specific examples described herein, as these examples are of course subject to variation. It should also be understood that the terminology used herein is for the purpose of describing specific examples only and is not intended to be restrictive, as the scope of the invention will be limited only by the appended claims.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. It must be noted that, unless the context clearly specifies otherwise, the singular forms “a,” “an,” and “described” as used herein and in the appended claims include plural indicators. Thus, for example, reference to “compound” includes a variety of such compounds, and reference to “polymer” includes reference to one or more polymers and their equivalents known to those skilled in the art, etc.

[0054] Where a range of values ​​is provided, it should be understood that, unless the context explicitly specifies otherwise, each intermediate value between the upper and lower limits of the range, down to one-tenth of the lower limit unit, is also specifically disclosed. Every smaller range between any stated value or intermediate value within the range and any other stated value or intermediate value within the range is included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range, and each range in which any limit, no limit, or both limits are included is also covered in this invention, subject to any specifically excluded limit within the range. Where the range includes one or two limits, ranges excluding one or both of those included limits are also included in this invention.

[0055] This document provides certain ranges in which numerical values ​​are preceded by the term "approximately". The term "approximately" is used in this document to provide literal support for the exact number that follows it, as well as numbers that are close to or approximate to the number following the term. In determining whether a number is close to or approximates a specific stated number, a close to or approximate unstated number may be a number that is substantially equivalent to the specific stated number in the context in which it is presented.

[0056] According to one aspect, systems and methods are provided for individuals to select (point to, identify, and / or indicate) a location or object within the visual content of a page of a book (e.g., selecting from a plurality of viewable objects) and subsequently interact with the page content based on the selection made using a handheld device. Within the description herein, the term "page" refers to any substantially two-dimensional surface capable of displaying viewable content. A page may be constructed of one or more materials, including paper, cardboard, film, cloth, wood, plastic, glass, painted surfaces, printed surfaces, textured surfaces, surfaces reinforced with three-dimensional elements, flexible surfaces, electronic displays, etc.

[0057] Following a similar line of thought, the term "book" is used herein to refer to any collection of one or more pages. A book can include printed materials such as traditional (e.g., bound) books, magazines, pamphlets, newspapers, handwritten notes and / or drawings, book covers, chapters, tattoos, boxes, logos, posters, scrapbooks, collections of photographs or drawings, etc. Pages can also be displayed on electronic screens such as tablets, e-readers, televisions, light-projecting surfaces, mobile phones, or other screen-based devices.

[0058] Page content is limited only by the author's imagination (one or more). Typically, for example, children's books may contain a combination of text and illustrations. More generally, page content may include any combination of text, symbols (e.g., including a range of symbols available in different languages), logos, specifications, drawings, and / or images; and may be displayed in color, grayscale, or black and white. Content may depict real or fictional scenes, or a mixture of both.

[0059] According to another aspect, devices, systems, and methods are provided, wherein one or more predetermined page layouts, including one or more page object templates, are known to one or more processors, including a handheld device processor. A handheld device user can use a beam of light generated by the device to point to a selected (i.e., user-selected) location within the page. As described in more detail below, a laser diode or light-emitting diode (LED) with beam-shaping optics can be used to generate the beam.

[0060] A camera within a handheld device, pointed in the same direction as the light beam, can acquire one or more images of a location or area within a page pointed to by the user of the handheld device. When the light beam is on, and provided the surface is sufficiently reflective, the image acquired by the camera can include reflections of the light beam (e.g., incident light reflected from the page).

[0061] Alternatively, the beam can be turned off during image acquisition. Turning off the beam (e.g., at least temporarily during camera-based image acquisition) allows page content to be processed without beam reflection interference (e.g., using CV methods). Even when the beam is turned off, because both the beam and the camera move simultaneously (i.e., fixed to or embedded within the handheld device body), the location or area the beam points to within the camera image can be known regardless of the physical location, pointing direction, or overall orientation of the handheld device in (3D) space.

[0062] The construction of a handheld device can strive to place the reflected light beam at the center of the camera image. However, given a small separation between the light source and the camera sensor (e.g., due to physical construction constraints, such as...), Figure 6 As shown, the beam may not appear in the center of the camera image at all working distances.

[0063] The pointing directions of the light beam and the camera field of view can be aligned to converge at a preferred working distance. In this case, the light beam can be positioned approximately at the center of the camera image (or some other chosen camera image location) over a range of working distances. In this configuration, the position of the light beam can vary within a limited range as the distance from the handheld device to the reflective surface changes (typically, in one dimension along an axis in the image plane, in a direction defined by a line passing through the center of the camera field of view and the center of the light beam source).

[0064] At a specific working distance (i.e., from the handheld device to the reflective surface), given the direction of the beam pointing, the direction of the camera image acquisition, and the physical separation between the two, the location of the reflection can be calculated using geometry (similar to the geometry describing parallax) (see, for example...). Figure 6 By keeping the physical separation between the beam and the camera small, the beam pointing area within the camera image can be kept small within the working distance range used during typical applications.

[0065] Alternatively, the beam and camera can be aligned to project and capture parallel (i.e., non-converging) light rays. In this case, the reflection can be offset from the center of the camera's field of view by an amount that varies with the working distance. The separation between the center of the camera image and the center of the beam decreases with increasing working distance (e.g., approaching zero distance at infinity). By keeping the physical distance separating the beam and camera small, the separation can be similarly kept small.

[0066] Regardless of the alignment configuration, the position of the beam within the image acquired by the camera (including when the beam is off) can be determined based on the camera and beam being co-located within the body of the handheld device, pointing in the same direction, and moving together. The beam position within the image acquired by the camera can be calculated based on the separation and pointing geometry of the beam and camera, and / or determined empirically via a calibration process (e.g., prior to deployment), for example using the following steps: 1) Acquire a baseline image (e.g., from a featureless surface) that does not include reflections from the projected beam; 2) Acquire a beam reflection image, which includes one or more reflections produced by the projected beam; 3) Calculate the subtracted pixel intensity image based on subtracting the baseline image from the beam reflection image; and 4) Calculate the beam pointing position based on the pixels in the subtracted pixel intensity image that exceed a predetermined light intensity threshold.

[0067] The beam pointing region can be identified based on the location of pixels exceeding an intensity threshold, and / or the location of a single beam pointing position can be determined based on the calculated center location of pixels exceeding the threshold intensity (e.g., the two-dimensional median, mean, or center of a beam spread function fitted to the intensity distribution). Calibration can be performed at different working distances to map the full range of the beam pointing region.

[0068] In the additional examples presented herein, when the beam is on, identifying the pointed object or location within an image acquired by the camera can be based on identifying the beam reflection as a high-brightness area (e.g., high intensity within a pixel location region). Identifying the location of such beam reflections may also take into account the color of the beam (i.e., identifying only higher intensities within one or more colors associated with the beam's spectrum). Knowing the relationship between the pointing location and the working distance further allows for calculating an estimate of the distance from the handheld device to the reflective surface when an image acquired by a camera with beam reflections is available.

[0069] In another example, based on the typical spectral sensitivity of the human eye, beams of light in the green portion of the visible spectrum are most easily sensed by most individuals. Many so-called RGB (i.e., red, green, blue) cameras contain twice as many green sensor elements as red or blue elements. Utilizing beams of light in the middle range of the visible spectrum (e.g., green) allows the beam intensity to be kept low but easily detectable (e.g., by both humans and cameras), thereby improving the reliability of detecting reflections within camera-based images and increasing overall eye safety.

[0070] According to other aspects of the devices, systems, and methods described herein, a predetermined interactive page layout may include object templates, object locations, object orientations, page specifications, text or symbol formats, and / or additional attributes of visual properties associated with the page content. Using CV methods (e.g., template matching, neural network classification, machine learning, transformer models) based on matching images acquired by one or more cameras with the predetermined page layout or object template, the location of an image within the page layout can be determined. As a result of this location, a handheld device processor can identify selected (i.e., indicated by a beam of light) locations, regions, and / or objects within the page.

[0071] Within the scope of this description, the term "positioning" (and related terms such as "locating") is generally used to describe the process associated with identifying the matching between a camera-acquired image and an interactive page layout. As described in more detail below, the positioning process may include determining multiple measurements of the camera-acquired image within the page layout, such as horizontal and vertical references (e.g., corners), positioning (x, y), orientation (θ), and magnification (x, y). mFollowing a similar line of thought, the term "located" (and related terms such as "position") is often used to describe the process of identifying the position of a beam of light within a camera image (and therefore within a page layout, once the camera image is positioned within the layout) and / or the position of objects within a page.

[0072] Within a page layout, object templates and / or attributes may include, for example, the position of objects or object elements, as well as object shape, size, outline, pattern, color, and / or texture. These can be stored in various formats, including bitmaps, vector-based scalable architectures, portable images, object element measurements, two-dimensional or three-dimensional computer-aided design (CAD) datasets, text-based descriptions and / or specifications, etc.

[0073] The interactive page layout dataset of objects may additionally include or point to a series of actions, properties, and / or functions associated with each page location (or page area). Therefore, the object or location being pointed to (and optionally, any other viewable object within a region of the camera image) may point to an operable dataset, which may, for example, contain sounds emitted by the object, sounds associated with the object's functions, pronunciation and / or speech elements of the object's name or description, spelling of the object's name, etc.

[0074] Other examples of datasets associated with page locations or objects include: cues or nudges from auditory or visual displays associated with the location the beam is pointing to; words or phrases describing the object; questions and / or hints associated with one or more objects in the pointed area; the spelling and / or phonemes of one or more of the selected object; questions about a story associated with the object or location; additional story content; rhythmic features associated with the selected object (e.g., which may form the basis of tactile or vibratory stimulation of the device user's hand); auditory feedback that rewards or comforts after a particular selection is made; pointers to associated additional or sequential content; actions performed if a selection is not made within a predetermined time; and / or one or more actions performed as a component of an interactive sequence after a successful selection of a page location or object (e.g., by a handheld device or by a remote processor), as described in more detail below.

[0075] In another example, page and content selection can be performed in two phases, where a “context” (e.g., a specific book, type of book or magazine, subject area) can be selected first. Therefore, page layouts and / or templates that are only relevant to the identified context can then be considered for CV-based matching (i.e., selection via beam pointing) with the image acquired by the camera during the second phase.

[0076] Predefined context templates may include one or more book covers, magazine covers, book chapters, toys, characters, anatomical elements, animals, pictures, drawings, words, phrases, household items, classroom items, tools, cars, clothing, etc. As an example, pointing to a book cover can cause an interactive page layout of all pages of the identified book to be retrieved and subsequently used to identify beam-based selections. Context objects (i.e., those used to identify context by pointing to a beam) can be "virtual" (e.g., images of real objects printed on a page or displayed on a screen) or "real-world" objects (e.g., cats or clothing in the device user's environment).

[0077] In some cases, it may not be necessary to point the beam at a specific item within the image captured by the camera to evaluate the context. For example, when pointing at a book cover, the cover can be identified without specifically pointing to one or more words in the title or author's name. CV methods can determine a match between the entire image captured by the camera or any part of the image and the entire context template or any part of the context template. In other cases, such as when pointing at a real-world object, when determining the existence of a match with a predetermined database of context object templates, a bounded region at the area where the beam is pointing within the image captured by the camera can be considered.

[0078] In addition to the visual properties within the object template (as described above), each context template dataset may point to all available interactive pages (or subsets thereof) associated with the context object, and / or provide applicable search terms or methods to identify such interactive page layouts. Once a context has been identified, only the context-related interactive page layouts may be considered during subsequent interactions until the end of the context sequence is signaled and / or another context is identified.

[0079] Signaling the end of a context sequence can originate from an expectation of a change in topic from the device user, such as via a movement gesture (sensed by an IMU), a pause during movement or other anticipated response, spoken words (sensed by a microphone), or pressing or releasing a device button. Similarly, limiting CV comparisons to the context page layout can stop at the end of a story or other book element.

[0080] If no contextual selection (e.g., a book cover) is identified using CV or other methods utilizing images acquired by the handheld device's camera, the device can (e.g., automatically) query a central (e.g., remote) repository of predefined contextual templates. If one or more contextual matches are found within the repository dataset (e.g., using CV or other methods), one or more contextual templates and the associated dataset can be downloaded (e.g., automatically, wirelessly) to the handheld device, allowing contextual interaction to continue. If no contextual match is found (e.g., as a result of imaging an image of an unknown book cover), the contextual query can be placed in a database for use in developing future interactive book content.

[0081] Regardless of whether a context-pointing step is included, CV and AI methods based on matching images acquired by the camera with a limited set of page layouts, object templates, and / or context templates can significantly reduce the required computational resources (e.g., compared to global classification methods). As an example, template matching can consider only page layouts relevant to the pages of a book provided to children. CV and AI methods utilizing (e.g., convolutional) neural networks can be trained on the pages of such provided books and only consider those pages.

[0082] In another example, network training (and other programming in handheld device CV and AI processes) can leverage a wide range of distributed machine learning resources (e.g., TensorFlow, SageMaker, Watson Studio). The constrained nature of comparing camera-based images with layout or template databases can greatly simplify the training (and classification) process to determine if a match exists with a page layout and / or object template. Training of classification networks can be restricted to a predetermined layout and / or template database (a library of children's books), a subset of data such as a specific set of books or magazines, or even a single book or chapter (e.g., when the context is known).

[0083] Such constrained datasets also allow for the implementation of relatively simple classification networks and / or decision trees. Optionally, classification can be performed entirely on a handheld device (with limited computing resources) and / or without being transferred to a remote device (e.g., to access a much larger amount of computing resources). Such classification can be performed using neural networks (or other CV and AI methods) with hardware typically found on mobile devices. As an example, MobileNet and EfficientNet Lite are platforms designed for mobile devices that have sufficient computing power to determine the location of images captured by a camera within a book page.

[0084] Classification based on known (e.g., relatively small compared to global classification methods used to identify any object) layouts and / or template datasets can also promote higher accuracy when identifying content pointed to by a beam of light. Such restricted classification may be more robust because: 1) images are compared only to a database of discoverable pages (e.g., not to all possible objects in the world), 2) discoverable pages can be used to perform training, and 3) the CV matching threshold can be adjusted to help ensure accurate reflection of the device user's intent.

[0085] The threshold used to locate camera-based images within a page can be adjusted based on factors such as hardware (e.g., camera resolution), environment (e.g., lighting, object size), specific application (e.g., the presence of multiple objects that look similar), user category (e.g., young vs. old, experienced vs. novice), or specific user (e.g., considering the previously successful object selection rate)).

[0086] According to another aspect of the system and method described herein, the predetermined dataset for the page layout may include and / or point to a dataset describing one or more actions to be performed by the handheld device and / or one or more connected processors during and / or after selection. Actions performed on the handheld device itself may include playing one or more sounds on the device speaker, displaying one or more lighting patterns on one or more device displays, and / or activating the device haptic unit (where the device component may be operatively coupled to the device processor).

[0087] Steps for initiating actions on one or more external devices may include transmitting to one or more remote processors an interactive page layout, any context and / or object templates, any or all camera images (particularly those used during object selection), the acquisition time when acquiring camera images, the positioned image (e.g., determined positioning, orientation, and magnification parameters), the beam pointing position, the positions of one or more specified objects within the page layout, one or more actions within the page layout dataset, and feedback elements generated by the handheld device. Failure to make any selection within a specified time or indication that does not produce a match with any acquired template may also be communicated to the external processor.

[0088] When a dwell time is used to indicate or identify when a location is specified by the device user (see below), additional dwell time thresholds and measurements may be included in the transmitted data. These include two or more camera images for measuring movement, measured image movement, IMU-based measurements of movement (e.g., gestures or taps), and / or predetermined dwell time amplitude and / or time thresholds. Data related to other signaling mechanisms used to identify when a selection is made by the device user may also be transmitted, including voice or audio indications (e.g., sensed by the device microphone) made by the device user, timing and identification of device switch (or other handheld device sensor) activation or release, etc.

[0089] Interaction facilitated by handheld devices can help bring printed or displayed content to life by adding audio, additional visual elements, and / or vibrational stimulation felt by the user's hand (or other body parts). Printed content enhanced with real-time interactive sequences that include content-related feedback can not only provide machine-based guidance while reading but also be "fun," thus helping to maintain emotional engagement, especially when read by and / or to children. For example, reading a book can be enhanced by adding queries, questions (for parents or guardians and / or children), additional relevant information, sounds, sound effects, audiovisual representations of relevant objects, real-time feedback after discovery, and so on.

[0090] Interaction with objects in a book may be a shared experience with parents, friends, guardians, or teachers. The use of a handheld device to control the delivery of serial content is more fully described in co-pending U.S. Application Serial No. 18 / 091,274, filed December 29, 2022, the entire disclosure of which is expressly incorporated herein by reference. Shared control for navigating to new pages or panels to select objects while viewing a book or magazine is more fully described in U.S. Patent No. 11,652,654, filed November 22, 2021, the entire disclosure of which is expressly incorporated herein by reference.

[0091] As mentioned in the invention description, parents, peers, guardians, or teachers can help learners transition through their zone of proximal development (ZPD). ZPD is a framework in educational psychology that distinguishes what learners can do without help relative to what they can do with guidance (and also relative to what learners cannot do even with guidance). Making books interactive (especially challenging ones) can provide always-available support and guidance to help learners transition through their ZPD in different subject areas without the need for assistance from individuals (e.g., teachers, family members, peers).

[0092] The ubiquitous tools available for this support and guidance not only reduce the need for the presence of individuals with sufficient literacy and / or skills, but also allow for machine-based transitions via ZPD at a time, place, comfortable environment, and pace chosen by the learner. Furthermore, personalized devices that appear tireless and have been previously used (e.g., with repetitive reward feedback) (e.g., in appearance, spoken dialect, knowledge level based on the learner's individual background) further aid learners' acceptance, independence, self-motivation, and confidence in using handheld devices. Maintaining a challenging environment by bringing books to life avoids boredom and / or loss of interest (and, additionally, avoids an interactive pace that becomes overwhelming).

[0093] Alternatively, the handheld device and / or remote processor may simultaneously perform continuous, real-time assessments of engagement, language skills, reading ability, and / or comprehension. Assessment metrics may include, for example, the time a child spends interacting, the success rate of finding objects based on queries (particularly within different thematic areas, including areas of identified interest such as sports, science, or art), the time required to make direction-based choices (e.g., typically related to attention and / or interest), and the overall progress rate in “discovering” new objects within pages of a series of content (such as books or magazines).

[0094] Such assessments can be compared with previous interactions by the same individuals (e.g., to determine progress in a specific subject area), interactions by others (e.g., at the same age, cultural background, or educational level) using the same or similar sequences of interactions, and / or performance between different groups (e.g., comparing geographical, economic, and / or social clusters).

[0095] Demonstrating milestone responses across various aspects of cognitive processing (e.g., initial indications involving distinguishing colors, phonemes and / or words, understanding the number of objects, performing simple mathematical operations, and gesture responses requiring controlled motor function) can be particularly useful in monitoring child development, learning rates in older users, assessing the possibility of presenting more challenging storylines, and / or enhancing engagement. Auditory, tactile, and / or visual acuity can also be monitored continuously by handheld devices.

[0096] Handheld devices and / or external processors can record interactions (e.g., for educational and / or parental monitoring). Automated recording of interactions, reading engagement, content, progress, vocabulary acquisition, reading fluency, and / or comprehension (e.g., including comparisons with other children of similar ages) can provide insights into a child's emotional and cognitive development.

[0097] In another example from this paper, a range of individual considerations may be taken into account during the development of handheld device interactions, including the age appropriateness and / or education level of the device user, individual preferences and / or interests, the educational and / or entertainment value of the page content, and whether the object is expected to be the next discovery within a series of sequences (e.g., a storyline, a sequence of letters or numbers). For example, children's books suitable for five to nine years old (i.e., early readers) typically contain lightly themed text (usually no more than 2,000 words) and illustrations. Children not only learn how to pronounce words but also recognize sounds often associated with selected illustrations. Optionally, such considerations may be known when developing page layouts and interactive content.

[0098] Interactive components (e.g., obtained from a page layout dataset) may be generated by a handheld device to initiate and / or guide interactions toward viewable objects (e.g., connections within a storyline, the next object within a logical sequence, the introduction of new objects and / or concepts). Hints may include AI-generated and / or scripted (i.e., pre-established) sequences comprising a combination of visual displays, auditory sounds, and / or tactile vibrations.

[0099] The scripted sequence may additionally include conditional dependencies (i.e., selection from two or more interaction scenarios) determined based on real-time conditions (e.g., success during a previous selection period, time of day, user age). Using the symbol “[hint]” to denote one or more hints or cues associated with an object, exemplary scripted hints (visual, auditory, and / or tactile presentations) for handheld devices include: Can you find the [hint]? Show me the biggest [hint].

[0100] What letter follows the displayed [prompt]? What animal makes this sound? [Hint] Which instrument produces the [hint] beat you feel? Who is viewing the [tips] on this page? Find all the tips on the page by pointing to each tip.

[0101] Furthermore, the handheld device processor may include an AI-driven “personality” (i.e., an artificial intelligence personality, AIP), a transformer model, and / or a large language model (e.g., ChatGPT, Cohere, GooseAI). An AIP instantiated within the handheld device can enhance user interaction by including a familiar appearance, interactive format, physical form, and / or voice that may further include personal insights about the user (e.g., likes, dislikes, preferences).

[0102] Human-computer interaction enhanced by AIP is more fully described in U.S. Patent No. 10,915,814, filed June 15, 2020, and U.S. Patent No. 10,963,816, filed October 23, 2020, the entire disclosure of which is expressly incorporated herein by reference. A more fully described method for determining context based on audiovisual content and subsequently generating a session by a virtual agent based on one or more such contexts is also described herein by reference.

[0103] Whether used alone or as part of a larger system, a handheld device familiar to an individual (e.g., a child) can be a particularly persuasive element of auditory, tactile, and / or visual reward as a result of object selection (or conversely, as a storyline component informing the user that the selection may not be the correct one). The handheld device can even be colored and / or decorated as a child's unique property. Following a similar line of thought, auditory feedback (voice, one or more languages, alarm tone, overall volume) and / or visual feedback (letters, symbols, one or more languages, visual object size settings) can be pre-selected to suit the individual user's preferences, adaptations (e.g., hearing), skills, and / or other abilities.

[0104] When used alone (e.g., while reading a book), interactions with a handheld device eliminate the need for accessories or other devices such as computer screens, computer mice, trackballs, styluses, tablets, or mobile devices when selecting objects and performing activities. Eliminating such accessories (often designed for older or adult users) also eliminates the need for younger users to understand the interactive sequences involving such devices or pointing mechanisms. When using a handheld device without a computer screen, interacting with images in a book and / or objects in the real world (given the relative richness of such interactions, approaching the richness of screen-based interactions) can be symbolically described as using the device to “make the world your screen without a screen.”

[0105] Feedback provided by a handheld device based on a selected object or location within the layout of an identified page can be conveyed through the device speaker, one or more device displays, and / or haptic units. Auditory interactions or actions may include: sounds or sound effects typically produced by the selected object; sounds associated with a description of an activity using the object; the pronounced name of the object (including proper nouns); congratulatory phrases or sentences; a part or even a single letter of a name associated with the object (e.g., beginning with a letter); statements about the object and / or its functions; verbal descriptions of one or more object attributes; questions about functions and / or object attributes; musical scores associated with the object; chiming indicators; references or statements about the object; verbal quizzes in which the selected object (or the next object to be selected) is the answer; and so on.

[0106] Visual interactions or actions may include: displaying the name of a selected object or object category, an image or drawing of another object that looks similar, an image of an object within a class or category of the object, the outline of the object, a cartoon of the object, a part of a word or letter associated with the spelling of the object (e.g., the first letter), one or more colors of the object (e.g., displaying a sample of one or more words that have actual colors and / or describe colors), the size of the object (e.g., particularly relative to other viewable objects), a statement or question about the object, a mathematical problem in which the object (numerical or symbolic) is the solution, the next object in a sequence of objects (e.g., a letter in the alphabet), a phrase in which the object is a missing element, etc.

[0107] Tactile feedback during selection may include simply confirming success with a vibration pattern, generating vibrations at frequencies associated with the motion and / or sound produced by the selected object, etc. Tactile actions may also include vibrations that are synchronized (or at least generated at similar frequencies) with the motion or sound typically produced by the selected object. As an example, a tactile vibration that pulses approximately once per second may be associated with an image pointing to the heart, or a vibration and / or sound generated approximately every thirty to forty milliseconds for approximately ten to fifteen milliseconds may simulate a cat's purring when pointed at.

[0108] Combinations of visual, auditory, and / or tactile actions can be generated by a handheld device and / or via a device operatively coupled to a remote processor. The combinations of actions can be generated simultaneously or as a tightly timed sequence.

[0109] According to another aspect of the system and method described herein, a light beam emitted by a handheld device can be generated using one or more laser diodes (such as those manufactured by OSRAM and ROHM Semiconductor). Laser diodes (and lasers in general) produce coherent, collimated, and monochromatic light sources.

[0110] Given the portability of handheld devices, where the near-field light source can be pointed in any direction, using incoherent sources such as (non-laser) light-emitting diodes (LEDs) provides increased eye safety (especially during use by children in environments often considered "uncontrolled" from a safety perspective). LEDs, so-called point light sources (such as those manufactured by Jenoptik and Marktech Optoelectronics), produce incoherent, collimated, and (optionally) multicolor light sources.

[0111] Optical components associated with the LED point source can control the beam divergence, which in turn can guide the size of the reflective spot at a typical working distance (e.g., approximately 0.05 to 1.0 meters when a handheld device is used to point at an object within a page of a book or a real-world object). Figure 5 The desired spot size may vary depending on the application environment and / or the user. For example, younger children may prefer to point to larger objects closer to the page in a children's book, while older children or adults may prefer to point to objects that are as small as individual words or symbols on the page or screen (i.e., using smaller and / or less divergent beams).

[0112] In the additional example, a broad-spectrum (at least compared to a laser) and / or multi-color light source generated by a (non-laser) LED can help those who may be colorblind in areas of the visible spectrum to see the reflected beam. Multi-color light sources can also be seen more consistently by all users when reflected from various surfaces. As an example, a pure green light source may not be easily seen when reflected from a pure red surface (e.g., an area of ​​a page). Multi-color light sources (especially in the red-green portion of the visible spectrum) can help mitigate this problem. Higher-energy photons in the deep blue end of the visible spectrum can be avoided for eye safety reasons.

[0113] In another example, the device beam source can be operatively coupled to a device processor, allowing control over the intensity of the beam source, including turning the beam on and off. For example, turning on the beam via a handheld device can serve as a cue indicating the selection of an expected object (e.g., after a hearing problem).

[0114] Optionally, the light beam can be turned off during the acquisition of camera images, for example, to avoid beam reflection and / or pixel saturation at or near the reflection (e.g., pixel "bleeding" due to saturation). The absence of reflection from the pointed object during CV processing (where reflection can be considered "noise" when identifying the object) reduces computational requirements and increases accuracy.

[0115] The beam can also be turned off after a match is confirmed as a component to signal success to the user. Conversely, turning the beam on during interaction can indicate to the device user that further searching of page objects or locations is desired. The use of a beam emitted from a handheld device as a pointing indicator in response to inquiries and / or suggestions is further described in co-pending application serial number 18 / 201,094, filed May 23, 2023, the entire disclosure of which is expressly incorporated herein by reference.

[0116] As described in more detail below, device users can use various methods (e.g., dwell time, voice commands) to indicate (i.e., by the user to the handheld device) the beam pointing to a desired selection. The method may involve multiple times of turning the beam on and off. For example, if the dwell time is measured based on not seeing movement within an image acquired by the camera, the beam can be turned off when each image is acquired, but kept on at other times to keep the device user engaged in the pointing process. The beam can be turned off once a predetermined dwell time has been exceeded, or once the handheld device has detected another means of indicating selection.

[0117] Alternatively, the beam intensity can be modulated, for example, based on measurements of one or more reflections within an image acquired by a handheld device camera. The reflection intensity can be clearly discerned by the user against the background (e.g., taking into account ambient lighting conditions, adapting to the reflectivity of different surfaces, and / or adapting to visual impairment), but is not indiscriminate (e.g., based on user preference). The beam intensity can be modulated by a variety of methods known in the art, including adjusting the magnitude of the beam drive current (e.g., utilizing transistor-based circuitry) and / or using pulse width modulation (i.e., PWM).

[0118] As another aspect of the system and method described herein, one or more illumination patterns may be projected within a device beam (and / or displayed on a device display). Using a beam to project one or more illumination patterns effectively combines the function of one or more individual displays on a handheld device with the direction of the beam. The illumination patterns projected within the beam may be formed, for example, using a micro-LED array, LCD filtering, or DLP (i.e., digital light processing using an array of micromirrors) methods.

[0119] Such lighting patterns can range from simple images (e.g., one or more dots or lines) to complex, rich, high-resolution images. The lighting patterns can be animated, generated to enhance or expand the objects, images, or text projected onto them. They can move proportionally to the movement of the handheld device from which they are projected, or they can be dynamically stabilized using images seen by an IMU and / or a camera, so that they appear stable even when the user moves the handheld device around.

[0120] In another example, during an interaction performed by a handheld device, one or more page objects identified within an image captured by the handheld device's camera can be "enlarged" by beam projection. For instance, if the beam is pointed at a squirrel while searching for a bird in the context of a story (e.g., being listened to), the beam can project a pair of wings superimposed on a printed image of the squirrel as components of a (silly) interactive query asking whether the pointed object is a bird. As yet another example, the apparent color of one or more components of a printed object can be altered by illuminating the overall shape of the object (or individual object components) with a selected color within the beam.

[0121] Lighting patterns can be synchronized with other interactive modes during any activity. For example, each word read aloud by the device can be independently and dynamically illuminated, bounded, underlined, or otherwise annotated. Following a similar line of thought, static images (e.g., of animals) can be expanded with projected images, which can then jump around the printed page.

[0122] When images and / or symbols (e.g., letters forming words and phrases) are too long and / or too complex to be displayed at once, messages and / or patterns within the beam can be "scrolled." Scrolling text or graphics can display a segment at a time in a predetermined direction (e.g., up, down, horizontal) (e.g., providing an appearance of movement). During and after the process of using the beam to select objects (i.e., when attention can be focused on the beam), the message embedded within the beam (e.g., the name of the identified object) may be particularly noticeable, effective, and / or meaningful.

[0123] The illumination pattern generated by the beam source can be used to enhance pointing functionality, including controlling the size and / or shape of the beam that can be viewed by the device user. Within the illumination pattern, the beam size (e.g., related to the number of illuminated pixels) and relative position (the location of the illuminated pixels) can be controlled by the handheld device. For example, different beam sizes can be used during different applications (such as pointing to letters within text (using a narrow beam) versus pointing to relatively large cartoon characters (using a larger beam)). The position of the beam can be "advanced," or the pointing direction can be indicated by the handheld device to help guide the user's attention to a specific (e.g., nearby) object within the camera's field of view.

[0124] As another aspect of the system and method described herein, viewable information and / or symbols within a beam projection with similar optical performance to a device camera (e.g., common depth of field, minimal distortion when viewed perpendicular to a reflective surface) tend to encourage and / or mentally advance the handheld device user to orient and / or position the handheld device such that the information and / or symbols are most viewable by both the device user and the device camera (e.g., in focus, rather than skewed). Therefore, images acquired by a well-positioned and oriented camera (i.e., at working distances and angles easily viewable by both the user and the camera) facilitate computer vision processing (e.g., improving classification reliability and accuracy).

[0125] As an example, symbols or patterns including "smiley faces" are easily (and perhaps even inherently) recognized by young children. Children can naturally position the device at a distance and / or angle with minimal guidance to obtain a clear, undistorted image of the smiley face (given the child's innate recognition of this universally recognizable image). Furthermore, the lens and filter structure of the smiley face beam can be designed to present a focused image only at a desired "optimal spot" distance from the reflective surface. The ability to easily view and / or identify the pattern of the projected beam, which the user may not be aware of, also enhances the image quality used for camera-based image processing.

[0126] Furthermore, if the projected image has a directional or typical viewing orientation (e.g., text, an image of a standing person), most users are likely to manipulate the handheld device so that the projected image is oriented in the object's typical viewing orientation. Alternatively, a specific orientation of the projected pattern or image can be maintained (e.g., using Earth's gravity as a reference) by measuring the orientation of the handheld device (e.g., using an IMU). The image can be projected in a manner that maintains a preferred orientation, for example, relative to other objects in the handheld device user's environment.

[0127] As another aspect of the system and method described herein, a handheld device user can (i.e., instruct the handheld device) to point a beam at a selected object and to make a selection. Such instructions can be given via a range of interactive methods, such as: 1. Press or release a switch (e.g., a button or other contact or proximity sensor) that is a component of a handheld device. 2. Provide verbal instructions (e.g., saying "now") that are sensed by the microphone of a handheld device and identified by the device processor or a remote processor (e.g., using natural language processing for classification). 3. Point the beam at the object (i.e., without any substantial movement) and "stay" for a predetermined (e.g., based on user preference) time. 4. Orient the device in a predetermined direction sensed by the handheld device's IMU (e.g., vertically relative to Earth's gravity, tilting the device forward), or 5. Gestures or taps on the handheld device are also sensed by the handheld device's IMU.

[0128] In these latter exemplary cases, where signaling to the user that movement of the handheld device (e.g., gesture, tap) can produce motion within the camera's field of view, still images can be isolated prior to any process that might produce motion (e.g., from a series of continuously sampled images). Images acquired by the camera prior to any motion-based signaling can be used to identify the viewable object or location being pointed at.

[0129] In another example, one approach to implementing a dwell-based method involves ensuring a sufficient number of consecutive images (e.g., calculated from the desired dwell time divided by the frame rate) to reveal substantially stationary viewable objects and / or light beam reflections. CV techniques (such as template matching, computer vision, or neural network classification) can be used to calculate one or more spatial offsets by comparing consecutively acquired pairs of camera images. Image movement (e.g., compared to a dwell movement threshold) can be calculated based on one or more spatial offsets or the sum of offsets over a selected time period.

[0130] When determining rapid and / or accurate dwell time, motion measurements based on camera images require high frame rates and the resulting computational and / or power demands. Alternative methods for determining whether sufficient dwell time has elapsed include using an IMU to assess whether a handheld device has remained substantially stationary for a predetermined period of time.

[0131] Converting analog IMU data into a digital form suitable for processing can be done using analog-to-digital (A / D) conversion techniques known in the art. IMU sampling rates typically range from approximately ten (10) samples / second (or even lower rates if desired) to approximately ten thousand (10,000) samples / second, with higher IMU sampling rates involving trade-offs related to signal noise, cost, power consumption, and / or circuit complexity. Motion and / or dwell time thresholds can be based on the device user's preferences.

[0132] Gesture-based selection instructions may include translation, rotation, lack of movement, tapping the device, and / or device orientation. One or more user intentions may be signaled, for example, through the following: 1. Any kind of motion (e.g., above IMU noise level). 2. Movement in a specific direction, 3. Speeds exceeding a threshold (e.g., in any direction). 4. Gestures using handheld devices (e.g., known mobile modes), 5. Point the handheld device in the predetermined direction. 6. Tap the handheld device lightly with the fingers of your other hand. 7. Tap the handheld device lightly with an impact object (e.g., a stylus). 8. Striking a handheld device against a solid object (e.g., a table), and / or 9. Impact the attached handheld device with the handheld device.

[0133] In further examples of using IMU data streams to determine user intent, a “tap” on a handheld device can be identified as the result of intentionally moving and subsequently causing an object (i.e., the “impact object”) to strike a location (i.e., the “tap location”) on the surface of the handheld device that the user aims at. The calculated tap location on the handheld device can be used by the device user to convey additional information about intent (i.e., in addition to making an object selection). As an example, the confidence level of a user in making a selection, indicating the first or last selection as part of a set of objects, or expecting to “skip forward” during a sequential interactive sequence can be signaled based on one or more directional movements and / or one or more tap locations on the handheld device.

[0134] The characteristics of a tap can be determined when a stationary handheld device is struck by a moving object (e.g., fingers of a hand opposite the hand holding the device), when the handheld device itself is moved to strike another object (e.g., a table, another handheld device), or when both the striking object and the handheld device move simultaneously before contact. IMU data streams before and after the tap can help determine whether a striking object was used to tap the stationary device, whether the device was forcibly moved toward another object, or whether both processes occurred simultaneously.

[0135] The tap location can be determined using unique “features” or waveform patterns (e.g., peak force, acceleration direction) within the IMU data stream (i.e., particularly accelerometer and gyroscope data) that vary depending on the tap location. The determination of tap locations on the surface of a handheld device based on inertial (i.e., IMU) measurements and subsequent activity control is described more fully in U.S. Patent No. 11,614,781, filed July 26, 2022, the entire disclosure of which is expressly incorporated herein by reference.

[0136] As another aspect of the apparatus and method described herein, the handheld device processor can perform various actions based on indicated selections. After determining that a camera image does not match any page layout, the handheld device can simply acquire subsequent camera images to continue monitoring whether a match has been found. Alternatively or additionally, prompts and / or hints may be provided, including repeating previous interactions, keeping the beam on, and / or presenting new interactions (e.g., from a template database) to accelerate and / or enhance the selection process.

[0137] After determining the page layout position and / or selecting objects, actions performed by the processor within the handheld device can be performed within the device itself. As described more fully above, this may include, for example, auditory, tactile, and / or visual actions that indicate a successful page layout match and / or enhance the storyline within the page.

[0138] Actions performed by the handheld device processor may include transmitting available information related to the selection process to a remote device, where further actions (one or more) may be performed. The transmitted information may include camera images (one or more), the acquisition time of acquiring (one or more) camera images, a predetermined area to which the camera image beam is pointed, a template for the selected page, any contextual information, and may include one or more selected objects or locations in the transmitted dataset.

[0139] Especially when used in recreational, educational, and / or collaborative settings, the ability to transmit the results of object discovery allows handheld devices to become components of larger systems. For example, when used by a child or learner, the experience can be shared, registered, assessed, and / or simply enjoyed (e.g., successfully or unsuccessfully selecting an object) with connected parents, relatives, friends, and / or guardians. Real-time and largely continuous assessment of metrics such as literacy, reading comprehension, overall reading interest, and skill development can help identify children who may benefit from additional support or resources (at an early stage, when intervention is most beneficial). As an example, early intervention support could include adaptations for dyslexia, calculia, autism, or identifying users who may benefit from gifted programs.

[0140] Optionally, the handheld device may additionally include one or more photodiodes, optical blood sensors, and / or electrocardiogram sensors, each operatively coupled to the device processor. These handheld device components can provide additional elements (i.e., inputs) to help monitor and determine user interactions. For example, a data stream from a heart rate monitor can indicate stress or pressure during the selection process. Based on the detected level, previous interactions, and / or predefined user preferences, interactions involving object selection can be limited, delayed, or abandoned.

[0141] In the additional example, while not strictly "handheld," such portable electronic devices can be fixed and / or manipulated by other parts of the body. For example, a device that interacts with a user to point a beam of light at an object can be fixed to an arm, leg, foot, or head. This positioning can be used to address accessibility issues for individuals with limited upper limb and / or hand movement, individuals lacking sufficient hand dexterity to convey intent, individuals without hands, and / or during situations where hands may be needed for other activities.

[0142] Interactions using handheld devices can further consider accessibility-related factors. For example, specific colors and / or color patterns can be avoided in visual interactions when individuals with different forms of color blindness use the device. The size and / or intensity of symbols or images broadcast on one or more handheld device displays and / or within the beam can be adapted to the visually impaired individual. Media containing selectable objects can be Braille-enhanced (e.g., containing both Braille and images), and / or contain patterns and / or textures with raised edges. Beam intensity can be enhanced and / or the handheld device camera can track finger pointing (e.g., within the area containing Braille) to supplement the pointing of the beam.

[0143] Following a similar line of thought, if an individual has hearing loss in one or more audio frequency ranges, those frequencies can be avoided or their intensity amplified within the audio interaction generated by a handheld device (e.g., depending on the type of hearing loss). Tactile interaction can also be modulated to take into account an individual's increased or suppressed tactile sensitivity.

[0144] During activities involving young children or individuals with cognitive impairments, interactions may involve significant "guessing" and / or require guidance from the device user. Assisting the user and / or relaxing pointing precision during interaction can be considered a form of "interpretive control." Interpretive control may include "propelling" a response or reaction toward one or more targets (e.g., providing intermediate cues). For example, young children may not fully understand how to manipulate a handheld device. During such interactions, auditory instructions may accompany the interaction (e.g., broadcasting "Keep the wand straight up"), thereby guiding the individual toward a choice.

[0145] Similarly, as a user approaches a selectable object and / or anticipates a response, a flashing display and / or a repetitive sound may be broadcast (where the frequency may be related to the proximity of a particular selection to a cue or attribute). On the other hand, a response where there is no apparent attempt to point at the beam may be accompanied by an "inquiry" instruction (e.g., haptic feedback and / or a buzzing sound) as a prompt to encourage alternative considerations. Further aspects of interpretive control are described more fully in U.S. Patent No. 11,334,178, filed August 6, 2021, and U.S. Patent No. 11,409,359, filed November 19, 2021, the entire disclosure of which is expressly incorporated herein by reference.

[0146] Figure 1An exemplary scene is shown in which a child 11 uses their right hand 14 to manipulate a beam of light 10a generated by a handheld device 15 to point at (and select) a picture of a dog 12b. The dog 12b is one of several characters in a cartoon scene 12a within two pages 13a, 13b of a children's book. The child 11 is able to see a reflection 10b generated by the beam of light 10a at the location of the dog 12b on the far right page 13b of the printed book. The child can identify and select the dog 10b during an interactive sequence using the handheld device 15, resulting in auditory feedback (e.g., barking sounds) played on the device speaker 16 and / or visual instructions (e.g., spelling the word "dog") presented on one or more device displays 17a, 17b, 17c and / or within the projected beam of light 10a.

[0147] Optionally, pressing button 18 (or other signaling mechanism, such as a voice prompt like saying "OK" detected by the device microphone, or a hand gesture or orientation of the handheld device sensed by the IMU relative to Earth's gravity) can be used to turn on beam 10a. Releasing button 18 (or other signaling mechanism) can be used to indicate that a selection has been made (i.e., to dog 10b), and optionally, beam 10a can be turned off (e.g., until another selection is made).

[0148] A handheld device camera pointing at the page in the same direction as beam 10a (in... Figure 1 (Not visible within the viewing angle) One or more images of the area pointed to by the beam can be acquired. The co-location of the beam-pointing area and the camera's field of view allows the determination of objects and / or locations selected using the beam within the template of an interactive book page.

[0149] Figure 2 continue Figure 1 The exemplary scenario illustrated here shows the field of view of the handheld device's camera at position 21 within the two book pages 23a and 23b. Figure 1 As shown, the handheld device 25 is manipulated by the child's right hand 24 to point at a selected object (e.g., a dog 22b) or location within the cartoon scene 22a.

[0150] During camera image acquisition, the light beam may remain on (e.g., typically generating a visible reflection from page 23b), or alternatively, the light beam may be momentarily (i.e., during camera acquisition) switched off, such as... Figure 2The beam path 20a (if it is on) is indicated by a dashed line. Because the beam and the camera's optical path originate within the handheld device 25 and point in the same direction, the position indicated by the beam (indicated by the crosshair pattern 20b) can be known within the camera image even when the beam is off. Turning off the beam during camera-based acquisition helps the CV process match the camera image to its position within the page layout by avoiding image distortion caused by beam reflection.

[0151] When a predetermined interactive page layout or template, or part thereof, that matches the field of view of a handheld device camera image is found, the image can be overlaid on page 21 (e.g., ...). Figure 2 (As illustrated). This allows the beam position (i.e., known within the camera image) to be determined within the page layout (and / or its associated database).

[0152] Knowing the location pointed to by the device user triggers access to a predefined dataset associated with that location (or region) within the interactive page layout. Figure 2 In the exemplified dog case, information about the specific dog being referred to, dogs in general, and / or the role of dogs within the story context can be broadcast.

[0153] For example, three spherical displays 27a, 27b, and 27c on the handheld device can spell the word "dog" (i.e., after pointing to dog 22b). Alternatively or additionally, the device speaker 26a can announce the word "dog" at 26b, produce a barking sound, and / or provide the correct name of the dog and general information about the dog. As another example of reward and / or feedback from the handheld device, tactile stimulation generated by the handheld device can be accompanied by a barking sound and / or confirmation (using vibration) that an expected choice has been made (e.g., in response to a query from the handheld device).

[0154] Additionally, the occurrence of user selection, the location of the selection within the template, the associated layout dataset, timing, and the identity of any (or, if, any) selected object 22b can subsequently control further actions, which are directly performed by the handheld device 25 and / or communicated to one or more remote devices (not shown) to trigger further activities.

[0155] Figure 3Exemplary parameters are illustrated that can be computed during a CV-based method to match an image 31 acquired by a camera with a predetermined page layout or template 33. In this example, the page contains a symbol of a star 32a, a cartoon drawing of a unicorn 32b, a drawing of a cat 32c, some text about the cat's name 32d, and page numbers 32e. In addition to the page layout 33, page attributes may point to one or more datasets containing additional information (e.g., text, audio clips, queries, sound effects, follow-up prompts) that can be used during subsequent interactions involving a handheld device user.

[0156] The method used to locate images acquired by the camera within the dataset of the layout and / or template can employ computer vision techniques such as convolutional neural networks, machine learning, deep learning networks, transformer models, and / or template matching. Position parameters may include a reference location of the camera's field of view 31 within the template 33. Horizontal and vertical coordinates (e.g., using a Cartesian coordinate system) can be specified, for example, by determining the positioning of the lower left corner (typically denoted as (x, y)) of the camera's field of view at 35a relative to the coordinate system of the page layout, where the origin (typically denoted as (0, 0)) is located at its lower left corner.

[0157] Because device users can tilt (i.e., orient) their handheld devices when pointing at a beam of light (e.g., relative to the orientation of the page), the camera's field of view may not be aligned with the coordinate system of the page layout. To account for this, an orientation angle, typically denoted as θ, can be calculated based on the image orientation.

[0158] Following a similar line of thought, users can move their handheld devices around at varying distances from the page surface. Under these conditions, using a device camera with fixed optics (i.e., without optical zoom capabilities), the area of ​​the page covered by the camera's field of view varies (i.e., a larger area is covered when the handheld device is held further away from the page). Therefore, the area typically represented as [missing information] can be calculated based on the size of the page covered. m The magnification factor. In summary, (x,y), θ and m It can be used to calculate (i.e., overlay) any location within an image captured by the camera (e.g., the location of a light beam) onto the page layout.

[0159] Figure 4This is an exemplary electronic interconnection diagram of a handheld device 45, illustrating components at 42a, 42b, 42c, 42d, 42e, 42f, 42g, 42h, 42i, 42j, 43, and 44, and the main direction of information flow during use (i.e., indicated by the direction of arrows relative to the electronic bus structure 40 that forms the backbone of the device's circuitry). All electronic components can communicate with one or more processors 43 via this electronic bus 40 and / or via direct paths (not shown). Some components may be unnecessary or not used during a particular application.

[0160] The core of a portable handheld device can be one or more processors (including microcomputers, microcontrollers, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.) 43 powered by one or more (typically rechargeable or replaceable) batteries 44. For example... Figure 6 As shown, the handheld device components also include a beam generating assembly 42c (e.g., typically a light-emitting diode) and a camera 42d for detecting objects in the beam region (which may include reflections generated by the beam). If embedded within the core of the handheld device 45, both the beam source 42c and the camera 42d may need to pass through one or more optical apertures and / or optical transparency (41b and 41c, respectively) through any handheld device housing 45 or other structure(s).

[0161] During applications involving acoustic cues or feedback, a speaker 42f (e.g., an electromagnetic coil or piezoelectric-based) 42f may be utilized. Similarly, during applications that may involve audio-based user interaction, a microphone 42e may acquire sound from the environment of the handheld device. If embedded within the handheld device 45, the operation of both the speaker 42f and the microphone 42e may be aided by acoustic transparency through the handheld device housing 45 or other (one or more) structures, for example by tight coupling to the device housing and / or including multiple perforations 41d (e.g., as shown in the image). Figure 2 (As further illustrated in 26a).

[0162] In applications that include, for example, vibration feedback and / or alerting the user to a possible selection, a haptic unit 42a (e.g., an eccentric rotating mass or a piezoelectric actuator) may be employed. One or more haptic units may be mechanically coupled to a location on the device housing (e.g., to be felt at a specific location on the device) or may be fixed to an internal support structure (e.g., designed to be felt more generally across the entire surface of the device).

[0163] Similarly, during applications that include visual feedback or response following object selection, one or more displays 42b may be used to display, for example, letters (as illustrated), words, images, and / or pictures associated with the selected object (or to indicate that the object has been incorrectly pointed to). Such one or more displays may be fixed to the main handheld device body and / or external to the main handheld device body (as shown in 42b), and / or may be optically transparent within the device housing (as indicated in 41a).

[0164] During a typical interaction, the user can signal to the handheld device at various times, such as when ready to select another object, pointing to a new object, or agreeing to a previously selected object. User signaling can be indicated by verbal feedback sensed by the microphone and by handheld device movement gestures or physical orientation sensed by the IMU 42g. Although exemplified as a single device at 42g, different implementations may involve distributed sub-components, which, for example, sense acceleration, gyroscopic motion, magnetic orientation, and gravity separately. Furthermore, the sub-components may be located in different regions of the device structure (e.g., the distal arm, the electrically quiet region) to enhance the signal-to-noise ratio, for example, during sensed motion.

[0165] One or more switching devices may also be used to indicate user signals, including one or more buttons, toggle switches, contact switches, capacitive switches, proximity switches, etc. Such switch-based sensors may require structural components at or near the surface 41e of the handheld device to transmit force and / or movement to circuitry located more internally.

[0166] Remote communication to and from the handheld device 45 can be achieved using Wi-Fi 42i and / or Bluetooth 42j hardware and protocols (e.g., each using different regions of the electromagnetic spectrum). In an exemplary scenario employing both protocols, the shorter-range Bluetooth 42j can be used, for example, to register the handheld device using a mobile phone or tablet (e.g., identifying a Wi-Fi network and entering a password). Subsequently, the Wi-Fi protocol can be employed to allow the activated handheld device to communicate directly with other more distant devices and / or the World Wide Web.

[0167] Figure 5The components of the electronic schematic and ray diagram are shown, illustrating exemplary elements 51a, 51b, 51c, 51d for beam generation, beam illumination 50a and reflective light path 50c, and a picture of a cat 59 being pointed at on pages 53a, 53b of book 57 detected by camera 56a. The picture of the cat 59 can be selected within a page sketch including a cartoon character 52a on the leftmost page 53a, a second cartoon character 52c on the rightmost page 53b, and a unicorn 52b. The electronic circuits 51a, 51b, 51c, 51d, and camera 56a and its associated optics 54 can all be integrated into the body of a handheld device (not shown).

[0168] Components of the beam generating circuit may include: 1) a power supply 51a, which typically includes a rechargeable or replaceable battery within a portable handheld device; 2) optionally, a switch 51b or other electronic control device (e.g., a button, relay, transistor) that can be controlled by the handheld device processor and / or the device user to turn the beam on and off; 3) a resistor 51c (and / or a transistor that can adjust the beam intensity) that limits the current delivered to the beam source, since the LED is typically configured in a forward bias (i.e., lower resistance) direction; and 4) a beam source that typically includes a laser or a light-emitting diode 51d.

[0169] Precision optics (which may include multiple optical elements, including some encapsulated within a diode-based light source, not shown) can largely collimate the beam 50a and optionally also provide a small (i.e., designed) beam divergence. Therefore, the beam size emitted from the light source 58a can be smaller than the beam size at a certain distance along the optical path 58b. The divergence of the illumination beam 50a and the distance between the light source 51d and the selected object 59 largely control the size of the reflected beam spot 50b.

[0170] Furthermore, the beam reflected from the selected object 50c can continue to diverge, for example in... Figure 5 In this context, the reflected beam 58c is larger than the illumination beam 58b. The size and shape of the illumination spot 50b (and its reflection) can also be influenced by the position of the beam relative to the reflecting surface (e.g., the angle relative to the normal of the reflecting surface) and / or the shape of one or more reflecting surfaces. For example, the horizontal dimension of the leftmost page 53a is convexly curved relative to the incident beam 50a, thereby causing the illumination beam (e.g., Gaussian profile and / or circular) to generate a reflection spot 50b that may be inherently elliptical (i.e., wider in the horizontal dimension).

[0171] Light from the field of view of camera 56a can be collected by camera optics 54, which focuses an image represented by ray 55a (including a focused spot of reflected beam 55b) onto the light-sensing component of camera 56b. Such sensed images can then be digitized using techniques known in the art and subsequently processed (e.g., using CV techniques) to identify the pointed location within the camera's field of view.

[0172] Figure 6 This is an exploded view of the handheld device 65, showing exemplary locations of the beam source 61a and camera 66a. Such components may be internalized within the handheld device 65 during final assembly. This view of the handheld device 65 also shows the back side of three spherical displays 67a, 67b, 67c attached to the body of the handheld device 65.

[0173] The beam source may include a laser or non-laser light-emitting diode 61a, which may also include embedded and / or external optical components. Figure 6 (Not visible in the image) to form, structure, and / or collimate the beam 60. Beam generating electronics and optics may be housed in a subassembly 61b, which provides electrical contacts for the beam source and provides precise control over beam aiming.

[0174] Following a similar line of thought, the image acquisition process is achieved through a focusing optics 64a incorporated within a threaded housing 64b, which allows additional (optional) optics to be included in the optical path for amplification and / or optical filtering (e.g., to reject reflected light from the beam). The optical components are attached to a camera assembly 66a (i.e., including an image sensing surface), which is then housed in a sub-assembly that provides electrical contacts for the camera and precise control over the image detection direction.

[0175] Figure 6 One aspect of the exemplary configuration shown includes a light beam 60 and image-acquiring optics for a camera 64a, both pointing in the same direction 62. Therefore, reflections of the light beam from any viewable object occur within approximately the same area of ​​the image acquired by the camera, regardless of the overall pointing direction and / or orientation of the handheld device. Depending on the relative alignment and separation of the beam source and camera 63, the position of the beam reflection may be centered (at a typical working distance) or slightly off-center from the center of the image acquired by the camera. Additionally, due to the (designed to be small) separation at 63 between the beam source 61a and the camera 66a, small differences in the beam position may occur at different distances from the handheld device to the reflective surface.

[0176] Such differences can be estimated using mathematical techniques (similar to those used to describe parallax). Therefore, even if the beam is turned off (i.e., there is no beam reflection within the camera image), the beam's position can be determined based on where the beam optics are pointed within the camera's field of view. Conversely, any measured offset of the position of the center of beam reflection (or any other reference) within the camera image can be used to estimate the distance from the handheld device (more specifically, the device camera) to the viewable object based on geometry.

[0177] Figure 7 This is an exemplary flowchart illustrating the steps of selecting an image of a shoe 72c within an interactive page 74a, directed by a beam of light 72b generated by a handheld device 72d. The shoe 72c is one selection within a collection of fashion accessories including a t-shirt 72a. Once within the page layout, predetermined auditory feedback associated with the shoe selection (e.g., type, features) within the interactive page can be played on the device speaker 78b. The steps in this selection process include: 1) At 70a, the processor of the handheld device obtains a predetermined interactive page layout 71a, which includes a page 71b containing clothing accessories, including baseball caps, t-shirts and shoes; 2) At 70b, the user points the beam 72b generated by the handheld device 72d at the shoe 72c located near the image of the t-shirt 72a; 3) Optionally (indicated by the dashed outline), at 70c, the light beam emitted from the handheld device 73b (pointing to the image of the shoe 73a) can be turned off to avoid beam reflection before acquiring the camera-based image; 4) At 70d, an image is acquired using a camera (invisible) and focusing optics 74d (depicted separately from the main body of the handheld device for illustrative purposes only), wherein the field of view 74a of the camera includes the lower half of the t-shirt 74b and the shoe 75c. 5) At 70e, compare the page layout and / or template previously obtained at step 70a with the camera-based image 75a (containing the lower half of the t-shirt 75b and the shoe 75c) to determine the best match and alignment within the page layout; 6) At 70f, isolate the area of ​​the camera image 76a, which includes the location or area of ​​the beam reflection 76b of the shoe segment 76c as indicated by the device user (regardless of whether the beam is on); 7) At 70g, based on the selected position or area within the page layout previously obtained in step 70a, identify the selected object (i.e., as a shoe) and associated interactive elements from the page layout database; and 8) At 70h, perform auditory (using a handheld device speaker 78b) and / or visual (using one or more device displays 78a) interactive actions within the page layout database based on the shoe selection.

[0178] Figure 8 Based on Figure 7 The flowchart illustrates a two-stage process for users to interact with cat-related images and content by first specifying (i.e., selecting) a feline context and then interacting with one or more interactive pages containing content about cats. The first stage involves selecting the word "cat (CAT)" from a cluster of text options (a list of animals 82a, 82b, 82c) using a beam 82e emitted from a handheld device 82e. In this case, context identification includes a step at 80e that uses the beam position to isolate specific contextual objects within the image acquired by the camera (an optional step that might not be necessary in other cases, such as identifying a book cover).

[0179] Once the device user has established a feline context, the device processor identifies and / or retrieves a cat-related page layout (e.g., at 85a, a book, poster, or article about cats). During this second phase, the user can interact with the page content (including both auditory and visual information and / or feedback) (again, using the beam at 86c). The steps in this two-phase process include: 1) At 80a, the processor of the handheld device obtains a page layout and / or object template 81a, which includes a list of animals, including images and text related to cats 81b; 2) At 80b, using a beam 82d generated by a handheld device 82e, the user points to the word "cat" 82c from a list of animals including "cow" 82a and "dog" 82b; 3) Optionally (indicated by the dashed outline), at 80c, the beam emitted from the handheld device 82f (pointing to the word "cat (CAT)" 82c) can be turned off to avoid interference from beam reflections within the image when identifying text content; 4) At 80d, an image is acquired using a camera (invisible) and focusing optics 83d (depicted separately from the main body of the handheld device for illustrative purposes only), wherein the field of view 83a of the camera includes the words “dog (DOG)” 83b and “cat (CAT)” 83c; 5) At 80e (where step 80e may not be performed in all cases, as indicated by the dashed outline), isolate the area of ​​the camera image 83e, which includes the location or area of ​​the beam reflection 83f (whether the beam is on or off) of the segment indicated by the device user. 6) At 80f, in this exemplary case, the template matching method at 84a (using the template obtained at step 80a) is used to identify the selection pointed to by the beam; 7) At 80g, template matching caused the word "cat (CAT)" 84b to be identified as context for future interactions; 8) At 80h, the processor of the handheld device acquires multiple page layouts related to cat 85a, including pages containing images and text about a cat 85b named Sam; 9) At 80i, the beam 86c generated by the handheld device 86d is used again, and the user points to the picture of the selected cat 86b near the picture of the unicorn 86a (e.g., on the same page); 10) Optionally (indicated by the dashed outline), at 80j, the beam emitted from the handheld device 86e (pointing to the cat 86c) can be turned off to avoid beam reflection when acquiring camera-based images; 11) At 80k, using a handheld camera and focusing optics 87d (depicted separately from the handheld device for illustrative purposes only), an image 87a is obtained, including a picture of a cat 87c and a nearby unicorn 87b; 12) At 80l, the interactive page layout about the cat previously obtained in step 80h is compared with the camera-based image 87d (containing a picture of a unicorn 87e and a cat 87f) to determine the match and optimal alignment within the layout; 13) At 80m, isolate the area of ​​the camera image 88a, which includes the location or area of ​​the beam reflection 88b (whether the beam is on or off) of the segment in the area of ​​the cat 88c as indicated by the device user; 14) At 80n, based on the selected location or region within the page layout previously obtained at step 80h, identify the selected object (i.e., the cat named "Sam") and interactive elements (including Sam's ability to purr at 89b) from the page layout dataset; and 15) At 80°, based on the handheld device selection, perform an auditory (a gurgling sound from the handheld device speaker 89f) and visual (spelling the name "SAM" on the three device displays 89e, 89d, and 89c) interactive action.

[0180] Figure 9This is a flowchart where a handheld device button 91c is used to turn the beam 92c on and off and to indicate to the device user that a selection has been made. In this exemplary case, selecting shoes 92b from clothing items results in an auditory description (e.g., related to shoe performance) at 98c. Additionally, isolating camera-based images collected just before the release button avoids motion-based distortion of the images due to the release button. Exemplary steps in the process include: 1) At 90a, the processor of the handheld device acquires the page layout 91a, which includes clothing selection; 2) At 90b, using the thumb 91d, the user presses button 91c on the handheld device 91b to turn on the pointing beam (if it is not already turned on and / or to indicate to the device that a selection is about to be made). 3) At 90c, the user manipulates the handheld device 92d to point the beam 92c at the shoe 92b (located on the page adjacent to the baseball cap 92a). 4) At 90d, a handheld device camera and optics 93c collect images 93a, wherein the camera's field of view includes the shoe 93b pointed to by the device user; 5) At 90e, the processor within the handheld device 94a obtains the state of button 94b to determine whether it has been released by the user's thumb 94c; 6) At 90f, if the button has not been released, return to step 90d at 99a to wait for the button to be released; otherwise, continue at 99b to determine the page area being pointed to. 7) Optionally (indicated by the dashed outline), at 90g, turn off beam 95c (pointing to the area of ​​the interactive page containing baseball cap 95a and shoe 95b). 8) At 90h, to avoid the consequences of movement within the camera image when the button is released, the most recently acquired image 96b (i.e., before the button is released) is isolated from a series of acquired images 96a; 9) At 90i, the page layout previously obtained at step 90a is compared with the camera field of view 97b (including a small portion of the cap 97a and the shoe 97c) to determine the matching and optimal alignment within the page layout; 9) At 90j, based on the selected beam position within page layout 98a, identify the selected object (i.e., as shoe 98b) and retrieve the interactive element from the page layout dataset; and 10) At 90k, use a handheld device speaker 78b to play an auditory description related to the selected shoe (obtained from the selected page dataset).

[0181] As described above in the specific embodiments, the method may include encouraging or "mentally prompting" the device user to orient the handheld device toward a target distance from a viewable surface. The viewable surface may include text, symbols, pictures, drawings, or other content, which may be identified and / or selected by the device user based on a pointing beam and an image acquired by a handheld device camera aligned with the beam. Because the user can perceive and / or recognize the pattern of the focused beam reflected from the viewable surface, encouragement can be generated to orient the handheld device at approximately the target distance.

[0182] Specifically, if the visual pattern is perceived as pleasing (e.g., a smiley face, the outline of a favorite toy, etc.) and / or useful (e.g., a pointing arrow, the outline of a finger, etc.), then the user of the handheld device may tend to keep the image projected by the beam in focus, and thus keep the handheld device approximately a target working distance away from the viewable surface. As described in more detail below, this involves maintaining the handheld device at approximately a target distance from the viewable surface while controlling the field of view of the co-aligned device camera.

[0183] When used by young children, some patterns (e.g., smiley faces) can evoke positive or desired responses that are inherent and / or reinforced by interactions within the child's environment (e.g., viewing familiar forms and / or faces). In another example, eye-catching (e.g., containing bright colors) and / or unappealing (e.g., angry faces or frightening forms) beam patterns may be of particular interest or attention to some device users. Beam image patterns may include human faces, animal faces (e.g., looking cute or fierce), animal forms, recognizable shapes, cartoon characters, toys, circles, rectangles, polygons, arrows, crosses, fingers, hands, (one or more) characters, (one or more) numbers, (one or more) other symbols, and emojis that project any of a range of expressions. Additionally, the projected beam image may be animated, distorted (e.g., in color, intensity, and / or form), projected as a series of images, and / or made to flash.

[0184] The viewable surface from which the beam image can be visualized can be made of paper, cardboard, film, cloth, wood, plastic, or glass. The surface can be painted, printed, textured, reinforced with three-dimensional elements, and / or flexible. The viewable surface can also be some form of electronic display, including e-readers, tablets, signs, or mobile devices. For example, the viewable surface can also be a component of a book, including book covers, brochures, boxes, signs, newspapers, magazines, printable surfaces, tattoos, or displays.

[0185] In the additional example, the beam image can be most identifiable when the handheld device is positioned by the device user perpendicular to the viewable surface (i.e., apart from focusing). When approximately perpendicular to the viewable surface, the beam image is likely to appear most consistent with the structured light pattern within the handheld device's beam. If the projected beam is directed toward a location on the viewable surface that is away from the surface normal and / or at an acute angle, the beam image reflection will appear skewed (e.g., elongated away from the surface normal along the axis of the handheld device) and / or increasingly distorted due to any surface curvature and / or imperfections.

[0186] In another example of this document, the result of holding a handheld device at approximately the target working distance and generally perpendicular to the surface direction may include: 1) maintaining a predetermined target field of view by a camera pointed in the same direction as the beam, 2) avoiding images acquired by a camera from a reflective surface that is skewed due to viewing from an angle not approximately perpendicular to the surface, and / or 3) ensuring that images acquired by a camera of objects on the surface are focused by camera optics having a depth of field that includes any target distance (e.g., which may be controlled by the device processor, as further described below).

[0187] In camera systems that do not incorporate scaling capabilities (or other methods for adjusting focusing optics), the camera's field of view is largely determined by the distance from the camera's light sensor array to the viewable surface. Therefore, for example, if the camera acquires an image of an area containing small objects far from a handheld device (or more specifically, far from the camera sensor array), the objects may appear as small dots (i.e., with little or no structure, defined by a small number of camera pixels). Conversely, if the device camera is held too close to the viewable surface, only a small portion of the object may be visible in the image acquired by the camera. Identifying such objects, or even determining the presence of a target object, can be challenging for computer vision (CV)-based processing.

[0188] Even when applying CV to less extreme cases, greater accuracy and robustness in identifying objects within an image acquired by a camera can be achieved when the size range of the object contour is within a target range that ensures a sufficient number of pixels are involved to identify object details but has a sufficient field of view to identify the entire object form. Limiting CV-based methods to a size range approximately at the focusing distance of the beam simplifies, for example, neural network training, object template definition, and / or other CV-based algorithmic approaches.

[0189] In another example in this paper, the ability to change the focus distance of the projected beam image allows a handheld device to influence the working distance by encouraging the device user to follow (i.e., by positioning to keep the beam image in focus) changes in the beam focus distance. Therefore, the field of view and size of objects within an image acquired by a co-aligned camera can be (indirectly) controlled via the control of the beam focus distance.

[0190] Applications involving interaction with viewable objects of varying sizes can benefit from control over the working distance implemented by controlling the beam focusing distance. As an example, when identifying text printed in small fonts, CV-based processing of the camera-acquired images (e.g., optical character recognition) can benefit from a reduced working distance (and therefore a smaller camera field of view, where the object is thus imaged by more pixels). Conversely, if pages within a book contain large images of objects with little detail (e.g., cartoon drawings or characters), CV processing of the camera-acquired images can benefit from a greater working distance (i.e., beam focusing distance).

[0191] Additionally, if the focusing distance of the beam is changed under the control of the device processor, the area covered by the camera's field of view and thus the size of the object within the camera's sensor array (i.e., the number of pixels affected by each object) can be estimated and / or known during the CV-based method used to identify such objects.

[0192] In miniature beamforming optics configurations for handheld devices, the focusing distance can be controlled by altering the distance the beam travels before the beamforming optics or by changing the beamforming optics themselves (e.g., changing the shape or positioning of optical elements). One or more of these strategies can be implemented at one or more locations along the beam's optical path. The optical path (e.g., to change the beam path) can be altered by inserting and / or moving one or more reflective surfaces or refractive elements.

[0193] The reflective surface can be, for example, a component of a microelectromechanical system (MEMS) device, which actuates one or more mirrors to redirect the entire light beam or a beam element (individually). Alternatively or additionally, one or more faceted prisms (typically using internal reflection) can be attached to an actuator to alter the optical path. In another example, one or more miniature deformable lenses can be included in the optical path. The miniature actuator for altering the reflecting or refractive element can employ a piezoelectric, electrostatic, and / or electromagnetic mechanism operatively coupled to a device processor.

[0194] Multiple reflective surfaces (e.g., ranging from two to ten or more) can amplify the effect of small movements of a light guide element that alters the focusing distance. Within a miniature beam-shaping element (e.g., having an optical path in the range of less than a few millimeters) designed to be held in a hand, a change in the optical path in the sub-millimeter range can produce a change in the focusing distance of the beam in the centimeter range.

[0195] The target beam focusing distance can be dynamically adjusted based on factors such as the interactive visual environment, the type of interaction being performed, individual user preferences, and / or the characteristics of the viewable content (particularly the selectable object size and level of detail required for object identification). A desktop environment where the user is pointing at an object within a handheld book benefits from a shorter beam focusing distance compared to, for example, an electronic screen pointing at a large desktop. Rapid back-and-forth interactive sequences also benefit from shorter target beam distances. User preferences can take into account an individual's (particularly children's) visual acuity, age, and / or motor skills. When a device user selects different books or magazines with different fonts and / or image sizes, the beam focusing distance can be adjusted to accommodate this change in content (and to suit the user's convenient viewing).

[0196] In another example, the beam image size can also be configured statically (e.g., using a light-blocking filter) or changed dynamically. Dynamic control of the beam image size can use techniques to similarly modify the focusing distance (i.e., move or insert reflective and / or refractive optics) to produce a range of light covering an area from a small focused beam (e.g., the size of an alphanumeric character) to the size of a book page or larger. In the latter case, the light source of a handheld device can exhibit functions more resembling a flashlight or searchlight compared to, for example, a so-called "laser pointer." Similar to the control of the focusing distance, the beam image size can take into account the type of interaction being implemented, the user's (especially children's) visual acuity, age, and / or cognitive abilities.

[0197] As described above, although not strictly "handheld," the device can be fixed or attached to and manipulated by a body part other than the hand (e.g., temporarily or for extended periods). For example, the device can be fixed or attached to a device worn on the user's head. In this case, the illumination beam not only indicates to the device wearer where the device's camera is pointing in common alignment, but also informs others who can see the viewable surface where the center of attention for the viewable content may be located. In some applications, a relatively large beam (e.g., comparable to the flashlight or searchlight described above) can effectively inform other nearby individuals which page or areas of pages are being viewed and any information contained within the beam (in addition to encouraging the device user to position the device approximately at a target distance from the viewable surface).

[0198] Within the additional examples herein, methods for generating structured light patterns that produce recognizable beam image reflections on a viewable surface may include generating patterns from multiple addressable (i.e., by a device processor) light sources or reflective surfaces, and / or blocking or deflecting selected light. Exemplary methods include: 1. One or more light-blocking filters (e.g., films or masks) may be inserted within the beam path of a light beam, blocking all or selected wavelengths of light. This method typically requires optical elements to project the light beam and control the focal plane for light blocking and display.

[0199] 2. Digital light processing (DLP) projector technology (e.g., Texas Instruments' DLP microsystems) combines a light-generating element with operable micromirrors that control the projection of light. Control of the micromirror array by the device processor facilitates the dynamic display of patterns.

[0200] 3. A liquid crystal (LC) filter or projection array (e.g., similar to those used in displays manufactured by Epson and Sony) can be a component within the optical path of the light beam. Addressable LC arrays can be controlled by a device processor, thereby allowing dynamic display of patterns.

[0201] 4. Multiple beam sources, each coupled to the device processor (e.g., LEDs, including so-called micro-LEDs manufactured by Mojo Vision), can facilitate dynamic beam imaging. The LED array can contain a relatively small number of sources (e.g., a five-by-seven grid capable of producing alphanumeric characters and symbols) or a much larger number of sources capable of producing image details that include complex animations.

[0202] The device's light source may include one or more light-emitting diodes (including those incorporating quantum dots, micro-LEDs, and / or organic LEDs) or laser diodes, and may be monochromatic or multicolor. The light intensity (including turning the beam on or off) may be adjusted using various methods, including pulse-width modulation (PWM) to control the beam drive current and / or the drive current, (e.g., by an operatively coupled device processor). Modulating the light intensity or color (e.g., by selecting from different light sources) may be used to attract the user's attention during a selection process anticipated by the user (e.g., by turning on the beam or causing it to flash), or to provide a "visual reward" (e.g., after a selection is made).

[0203] The beam intensity and / or (one or more) wavelengths can also be modulated to take into account environmental conditions during use. If used in a dark environment, for example, in an ambient light region (e.g., not containing the beam image) sensed by the device's photodetector and / or captured by the camera, the beam intensity can be reduced to save power and / or prevent overwhelming the visual sensitivity of the handheld device user. Conversely, in bright or visually noisy environments, the beam intensity can be increased.

[0204] The beam intensity and / or (one or more) wavelengths can also be modulated to account for the reflectivity (including color) of visible surfaces in the area being pointed at. For example, when aiming at an area of ​​a drawing primarily containing red pigment, a beam consisting mainly of green wavelengths may not be well reflected or even visible. The overall color of the reflective surface can be estimated within images acquired from a camera in a region near the beam's image area. The beam color can then be adjusted to make the reflection more apparent to both the device camera and the user. Similar strategies can be used to compensate for individual visual acuity and / or to account for beam intensity preferences.

[0205] As another example, if the viewable surface includes an electronic display (e.g., an e-reader, a tablet), the intensity of beam reflection (including from the subsurface structure) can be reduced due to the typically low reflectivity of such surfaces. The beam color and / or intensity can be adjusted (e.g., dynamically) to help ensure the beam image has sufficient intensity to be visualized by the device user.

[0206] Additionally, handheld devices can use the low reflectivity of such surfaces, measured within a beam image region in an image acquired by a camera, to determine the presence of some form of electronic display. A reduction or absence of the beam image can be coupled with the determination of the presence of dynamic changes in individual objects within the viewable surface. Movement or changes in the appearance of viewable objects, where the rest of the viewable surface does not move concurrently (e.g., indicating movement of the handheld device camera and / or the entire viewable surface), can help indicate the presence of an electronic display surface (e.g., containing dynamic content).

[0207] Within beam illumination methods that allow for alteration of the structured light pattern of the beam (e.g., LC, DLP, LED arrays), the beam pattern can be dynamically changed, for example, as a cue mechanism during the selection process (e.g., changing facial expressions, components that magnify the beam image) and / or as a reward after selection (e.g., projecting a starburst pattern). The altered light pattern can be superimposed on dynamic changes in beam intensity and / or wavelength.

[0208] Changing the beam projection pattern, intensity, and / or color can also be used by handheld devices as feedback to indicate to the device user whether the device is maintaining approximately the target distance from the reflective surface. If CV-based analysis of the beam projection area within the image acquired by the camera determines that the beam image is out of focus (i.e., the handheld device is not maintaining approximately the target distance from the reflective surface), changes in the beam pattern, intensity, or color can be used to prompt the user to correct the device positioning.

[0209] Such prompts may include, for example, removing details within the projected image so that, if not too far out of focus, the user can roughly visualize the outline of the beam pattern, and / or altering the beam intensity and / or color that might be perceived even within a severely out-of-focus beam image. As another example, a smiling face that is clearly visible within the beam image when focused (i.e., pleasing to most viewers) may shift towards a frowning or sad face (e.g., possibly still recognizable) as the handheld device moves away from the target distance.

[0210] Alternatively or additionally, handheld device positioning feedback may include other modes. The use of other feedback modes may be useful when the beam image is excessively out of focus (i.e., the beam image may be indistinguishable) and / or when there is little or no device processor controlling the device on the beam image (e.g., using a simple light-blocking filter in the optical path to form the beam).

[0211] Positional feedback can be provided to the device user through auditory, other visual, or tactile means. For example, if a speaker is included within a handheld device and operatively coupled to the device processor, it can provide auditory cues or hints when the light beam is determined to be focused (within an image captured by a camera) or conversely, unfocused. A display or other light source operatively coupled to the device processor (i.e., not associated with the light beam) can provide similar feedback by changing the display intensity, color, or content. Following a similar line of thought, a tactile unit can be operatively coupled to the device processor and activated when the light beam is unfocused (or, alternatively, when focused).

[0212] Within each of these alarm or notification modalities, a measure of focus can be reflected in the feedback. The amplitude, pitch, and / or content (including words, phrases, or entire sentences) of a sound played from a speaker can reflect focus. Similarly, the brightness, color, and / or content of a visual cue or prompt can be modulated based on a measure of focus. The frequency, amplitude, and / or pattern of tactile stimuli can similarly indicate to the device user how close the handheld device is to a target on a viewable surface.

[0213] As just described, a measure of the degree of focus within a beam of image data acquired by the camera can be used as the basis for user feedback (and / or for performing one or more other actions). Many CV-based methods are known in the art for determining the degree of focus within the entire image or a sub-region of the image. At the simple end of a range of methods (e.g., easier to implement), increased focus is generally associated with increased contrast (especially if the image shape does not change). Therefore, a measure of contrast within an image region can be used as an indicator of the degree of focus.

[0214] Additionally, multiple kernel-based operators (two-dimensional in the case of images) including the calculation of Laplacian can be used to estimate edge sharpness and associated focus. An image or image region can be transformed to frequency space via Fourier transform methods, where the focused image can be determined within the image containing high-frequency components with larger values.

[0215] Artificial neural networks (ANNs) trained to measure focus within an image can also provide a measure of focus. This is especially useful if the image region is small and well-defined (see example...). Figure 10 Furthermore, since the image profile is based on a small, predefined dataset of potential beam projection images, such trained networks can be trained relatively quickly and / or kept small (e.g., by reducing the number of nodes and / or layers).

[0216] In another example in this paper, operational issues that may arise when performing one or more actions during an interactive sequence of using a beam of light to point at a selectable area on a viewable surface include: how can the user be informed of the boundaries of the selectable area, when should images acquired by the camera be processed to identify the content within such areas, and when should the resulting device-generated actions be performed? For example, if the device processor were to apply CV analysis to every image acquired by a camera available on a handheld device, the user could quickly become overwhelmed by unintentional actions and might not be able to effectively communicate intent.

[0217] Addressing these issues may include methods for informing users of the boundaries of each area associated with a selectable action. When a handheld device brings a book or magazine page to life, the size and / or number of objects that may be included in a particular area that generates an action when selected may not be immediately apparent. For example, within a section of a page containing text, in some cases (e.g., when learning to spell), individual letters may be selectable or operable objects. In other cases, words, phrases, sentences, paragraphs, cartoon bubbles, descriptions, and associated pictures may each (i.e., not readily apparent to the device user) reside within individual areas that generate an action when selected.

[0218] One method of providing feedback to notify a user of a selected area or operable boundary includes providing one or more audible cues when moving from one area to another and / or when hovering over an operable area for a short period of time. Since a user can easily point to or quickly sweep a beam of light across multiple operable areas within a short timeframe, the audio cues or prompts (e.g., ping-pong, ding-ding, ringing, clanging, clattering) can typically be brief (e.g., lasting less than one second). Furthermore, if multiple operable areas are detected rapidly and consecutively, the continuous prompts can be quickly stopped and / or one or more queued prompts can be eliminated to avoid overwhelming the device user.

[0219] Depending on the identity or category within the "flyover" or "hover" area, different sounds or cues may be generated. For example, an area that, if selected, might lead to an interactive query originating from a handheld device, might generate a more alert "questioning" cue; while an area that merely leads to an interactive exchange (i.e., no questioning) might generate a different (e.g., less alert) informational cues. In another example, a short sound effect associated with the identified object may be used as a cue. For example, a cue for an area containing a cat might include a meow. More generally, cues indicating an area pointed to by a beam of light may depend on the predetermined page layout characteristics of that area and / or objects within the camera-captured image of that area (e.g., using a CV-based method for identification).

[0220] Alternatively or additionally, display elements on a handheld device (e.g., indicator LEDs, orbs, display panels) and / or any one or both of a pointing beam of light can be used to provide one or more visual cues. Similar to the auditory cues just described, visual cues indicating the movement of the beam across different operable areas can be short and designed not to overwhelm the device user. As an example, visual cues may include changes in the brightness (including flashing), color(s), and / or pattern of the displayed content. Similar to the auditory methods just described, visual cues may depend on information associated with the page layout of the identified area and / or one or more objects identified within a camera-captured image of that area.

[0221] Device users with hearing impairments may use visual cues. Conversely, visually impaired individuals may prefer auditory cues. Alternatively or additionally, tactile cues or cues may be generated via a tactile unit operatively coupled to the device processor. The device user's hand may perceive the tactile cue when a beam of light moves into or remains within an operable area. The duration and / or intensity of the tactile cue may depend on page layout information associated with the area and / or on one or more objects identified within an image captured by the camera. For example, a prolonged (e.g., longer than one second) strong tactile cue may be implemented when the beam of light is directed at an identified area (where a query to the device user can be generated upon selection of the area).

[0222] Once a user becomes aware (e.g., via auditory, visual, and / or tactile means) that a beam of light is directed at an operable area, the user can then indicate (i.e., to the handheld device) that a selection is being made. Such indication can be made via a switching component (e.g., a button, toggle switch, contact sensor, proximity sensor) operably coupled to the device processor and controlled by the device user. Following a similar line of thought, voice control (e.g., identified words, phrases, or sounds) sensed by a microphone operably coupled to the device processor can be used to indicate a selection. Alternatively or additionally, when the handheld device is manipulated by the device user, the movement (e.g., shaking) and / or orientation (e.g., relative to gravity) of the handheld device can be sensed by an IMU operably coupled to the device processor and used to indicate that a selection is being made.

[0223] Additionally, the indication method (e.g., using voice or a switch) can be associated with one or more auditory, visual, and / or tactile cues just described. During the selection interaction, cues can be used to signal when an indication and / or indication modality is expected. This interaction strategy helps avoid the user unnecessarily generating indications at unexpected times, using expected modalities (e.g., voice relative to a switch) to generate indications, and / or providing further input (i.e., to the handheld device processor) during the selection process. As an example of the latter, visual cues can be provided, where the displayed color may correspond to the color of one of several distinctly colored button switches (e.g., where the contact surfaces may also differ in size and / or texture). Pressing a button of a similar color to the displayed color indicates agreement to the interactive content; while pressing another button may signal the device user's disagreement, alternative choice, or unexpected reaction.

[0224] Figure 10 This is an exemplary diagram (with elements resembling a ray diagram) illustrating the formation of a structured light pattern using a light-blocking filter 102 and the resulting focused reflected beam image 103b at a preferred or target working distance between the handheld device beam source 105 and the viewable surface 106b of the reflected beam. The focused reflection encourages the device user (e.g., especially if it is pleasing to the user) to manipulate the handheld device at the target working distance. At distances less than (e.g., at 106a) or greater than (e.g., at 106c) approximately 106b, a blurred image (i.e., unfocused) (e.g., 103a and 103c) can be observed, thus hindering handheld device manipulation and / or operation in these areas.

[0225] In this example, a light source (e.g., an LED at 100) that allows light to pass through one or more optical elements 101a is used to form a structured light pattern, thereby allowing the blocking filter 102 to then structure the light pattern. Depending on the image forming characteristics of the light source (e.g., whether a directional array and / or collimating optics are included within the source), some beamforming configurations may not require the optical elements depicted at 101a.

[0226] A second set of one or more optical elements 101b allows the projection of a structured light pattern as a beam (where the light elements are illustrated at 104a and 104b). The beam image (smiley face 103b) appears focused at a target working distance 106b from the device light source 105. If the handheld device is held too close to the reflective surface, the beam image becomes out of focus 103a. If the beam is configured to diverge slightly within a typical working distance (e.g., up to one meter), the out-of-focus reflective image may also appear smaller (e.g., compared to a focused or any further beam image). Similarly, if the handheld device is held too far from the reflective surface at 106c, the beam image also appears out of focus 103c to the device user (and larger in the case of a diverging beam).

[0227] Optionally, the optical elements may additionally include one or more movable reflective surfaces (e.g., at 107) operatively coupled to the device processor. Such surfaces may be components of, for example, MEMS devices and / or faceted prisms, which can alter the optical path distance from the light source 100 to the focusing and / or collimating optics 101b. The optical components depicted at 101b may include one or more optical elements that guide light into and / or out of the reflective surface region 107.

[0228] exist Figure 10 In the exemplary beam image pattern shown, to illustrate the overall circular beam structure, the bright (i.e., illuminated) elements of the beam image are drawn as dark components 103b (i.e., against a white graphic background). Following a similar line of thought, the elements allowing light to pass through within the light-blocking filter at 102 have been rendered white (again, allowing the circular pattern of the entire beam to be visualized). Typically, the blocking elements of structured light can block selected wavelengths or all wavelengths (e.g., using film or LCD filters) and introduce a wide range of image complexity, as long as the diffraction limit of light (at a specific wavelength) is not violated. Systems that directly generate structured light patterns (e.g., arrays of LEDs or microLEDs, DLP projection) can produce beam images with similar complexity and optical limitations.

[0229] Figure 11This is a flowchart illustrating exemplary steps, wherein, in addition to the device user viewing the beam reflection to determine whether it is focused (i.e., at approximately a desired distance), beam focus can also be calculated from the beam pointing region within images 113a acquired by one or more cameras. Positioning feedback and whether the user 116 is making a feasible beam pointing selection (i.e., within the target camera's field of view) can be provided to the user using periodic or continuous machine-based evaluations of whether the handheld device is at approximately a desired distance from the viewable surface. Such feedback to the user indicating whether the handheld device is being held at approximately a desired distance may include visual cues 117a, auditory cues 117b, and / or tactile cues. Exemplary steps in the process include: 1) At 110a, the user directs a structured beam emitted from the handheld device 111d at 111a onto a selected object on a viewable surface (in this case, the cat at 111c), thereby generating a beam reflection at 111b. 2) At 110b, an image is acquired using a camera (invisible) and focusing optics 112c (depicted separately from the main body of the handheld device for illustrative purposes only), wherein the field of view 112a of the camera includes the beam reflection at 112b; 3) At 110c, based on the knowledge of the co-aligned beam and the beam pointing area within the camera image, and / or using a computer vision method to isolate the area within the image 113a acquired by the camera, the computer vision method may optionally (indicated by the dashed outline) include comparing one or more areas of the camera field of view with a predetermined template of the (focused) beam image at 113c. 4) At 110d, the focus within the image acquired by the camera at the beam pointing region is measured (by the device processor) using a kernel operator (as depicted in 114) and / or any of a number of additional computer vision techniques for determining focus. 5) At 110e, determine whether the beam image is sufficiently focused to be within approximately the desired distance, and if so, continue processing the image acquired by the camera at 115b; otherwise (i.e., if not, at 115a), allow the user to continue manipulating the handheld device at 110a. 6) At 110f, continued image processing may include identifying text, symbols, and / or one or more objects near the beam of light (determined to be in focus), wherein the feline form is depicted at 116; 7) At 110g, optionally (indicated by the dashed outline), visual (e.g., via one or more device displays at 117a, and / or by modulating the beam intensity, including turning it off as depicted), auditory (using a device speaker at 117b) and / or tactile feedback is provided to the device user, indicating that the handheld device has been correctly positioned and selected, thereby triggering one or more actions by the device processor.

[0230] In another aspect of the device and method, during an interactive sequence involving book content, identifying when a user has turned their attention to a new page can be detected by the handheld device and automatically trigger one or more actions (e.g., related to the new page). As described in more detail above, the term "book" is used herein to refer to any collection of pages, including, for example, bound books, magazines, newspapers, scrapbooks, etc.

[0231] Additionally, the reference to "handheld" devices in this document helps visualize typical use of the device. However, the device (including the beam source and the co-aligned camera) can be attached to or connected to any other part of the human body and / or manipulated by any other part of the human body, including the user's head, wrist, arm, shoulder, leg, or chest. Manipulation of other parts of the body can, for example, free the user's hands to perform other tasks, such as holding an infant or young child, using a finger or stylus to point at an object (e.g., text, drawing), drawing on a whiteboard, generating a signature, or other gestures.

[0232] Physical attachment of the device to a body part can be aided by one or more support structures, such as headbands, wrist straps, or shoulder or chest straps. Attachment of the portable device to the support structure can be aided by configurations that allow for quick and easy attachment and detachment. For example, one or more attachment points can be magnetically held using simple latching mechanisms and / or hook-and-loop fastening systems (e.g., manufactured by Velcro). Quick and easy attachment and detachment facilitates the use of different body parts or the adoption (and purchase) of a single portable device by different users at different times. Furthermore, different devices can be specifically designed (e.g., with different device body shapes and / or optical working distances) to facilitate manipulation using different body parts.

[0233] As described above, the light beam from a light source co-aligned with the field of view of the device camera can be collimated or have controlled divergence. Additionally, the light can be coherent (e.g., generated by one or more laser diodes) or incoherent (e.g., generated by one or more LEDs). The light beam can also be structured to produce a reflected image that can be recognized by the device user. In another example, beam-shaping optics can be configured to produce a focused reflected image over a distance range from the device to the reflective surface, wherein if the device user positions the device to view the focused image reflection, the field of view of the co-aligned camera (with a similar focal length range) can acquire and process the focused image with sufficient resolution to identify the page object (e.g., using a CV method).

[0234] The focal length (or focal length range) of the light source and the co-aligned camera can also be configured to match different device configurations and / or optical path geometry within an application. For example, when manipulated by the device user's head to guide the beam at a page of a book, the focal length can be extended (e.g., up to about two meters). When manipulated by the device user's hand or arm, the focal length can typically be reduced (e.g., up to about one meter) due to reaching the page with the device via the hand or arm. Other aspects of generating structured beams to produce reflected images (e.g., smiley faces) focused within the working distance of a camera are described in co-pending application serial number 18 / 382,456, filed October 20, 2023, the entire disclosure of which is expressly incorporated herein by reference.

[0235] Device users can move interactive focus from the initial page to a new page in the following ways: 1) directing light emitted by the device to a new section or page within a book spread, or 2) flipping one or more pages to open the book to a new book spread and directing light emitted by the device to the new page or spread. As is commonly used in the book manufacturing industry, a book spread comprises two adjacent pages that can be viewed together when the book is opened. During the interactive sequence, a new page may also include a front or back cover, and / or components of a separate book or other printed material (e.g., a poster, sticker).

[0236] Furthermore, such books can be organized according to natural or identifiable (i.e., for device users) sequences, such as numbered pages in a typical book. For example, sequences can also be identified based on alphabetical words or identifiable objects, a series of images illustrating causal actions, a series of related objects, the chronological order of events depicted, ordered mathematical relationships (e.g., Fibonacci series), etc.

[0237] Switching user focus by pointing the device's beam at a new page can be considered an input to a handheld device that indicates user intent. Determining user intent based on one or more pages of the manipulated book detected within an image acquired by the device's camera reduces reliance on other forms of device input. In other words, typical signaling indications used to trigger interaction by pressing buttons or providing voice commands can be reduced or eliminated when user intent is determined from the image acquired by the camera and / or when one or more actions triggered after a new page is detected are automatically performed.

[0238] Determining whether a user turns to the next page in a natural (e.g., numbered, alphabetical) sequence can provide additional information about the device user's intent and influence one or more actions triggered by the device. For example, if the new page is not the next page in an ordered sequence, the device can generate audio, visual, or haptic cues to alert and / or warn the user that the new page is out of order. Following a similar line of thought, when the new page is an element of a new book, unique audio, visual, or haptic "new book" cues can be provided.

[0239] As described in more detail above, book pages can be identified (and linked to interactive content) by matching images captured by the page's camera with page templates or layouts within a page layout database. In this case, the process of determining whether a new page has been encountered can be performed by comparing the page identity determined within the page layout database. If the page identity determined from a newly acquired page image differs from the page identity of a previously acquired page image, an interactive sequence for the new page can be implemented. To reduce the occurrence of erroneous determinations of the presence of new pages (i.e., false positives), the interactive sequence for the new page can be implemented only after a predetermined number of images (e.g., two or more) of the newly identified page have been acquired and identified.

[0240] In another example, the transition to a new page can also be determined without involving a database of page layouts and / or templates. An initial camera-based image can be acquired during the interactive session. Additional camera-based images can then be acquired and compared with the initial camera image using computer vision methods (e.g., convolutional neural networks). If the comparison leads to the determination that there is no difference in page image or page identity (if determined) between the two images, the interaction using the device can continue uninterrupted.

[0241] However, if comparisons of images acquired by the camera result in the identification of different pages, an interactive sequence can be implemented for the new page. For example, unique page elements can be determined based on differences in text, background, drawings, images, and / or page numbers. Similar to the process just described involving the identification of pages within a page layout database, false positives for new pages can be reduced by ensuring that a predetermined number of images (i.e., two or more) are identified as different from the initial camera-based images.

[0242] Alternatively or additionally, under certain conditions, users may expect to intentionally signal the end of an interaction on a page and then look for the next page during the interactive sequence. As an example, a user might want to leave a page but want to explore or "look around" for a new page of interest without repeatedly triggering unintended interactions involving the current page or initiating any automated actions related to an unintended new page (i.e., until a new page selection has been indicated).

[0243] The intention to move to a new page can be signaled using a switch or other signaling methods, such as voice commands (sensed by the device's microphone) or handheld device gestures (sensed by the device's IMU). For example, if the handheld device has more than one available button (or other signaling mechanism), it is preferable (i.e., from the perspective of interaction simplicity) to assign specific switches or buttons (e.g., identified by size and / or color) to be regularly available as a "new page" button. In this way, the ability to signal movement to a new page can always be consistently available to the user. Switches within or on the handheld device body can be, for example, buttons, rocker switches, contact switches, or proximity-sensitive switches.

[0244] Once the initial "New Page" button has been pressed and the process of selecting a new page has been completed, the user can trigger a "New Page" sequence of one or more actions: 1) pointing the device's light source at the new page, or 2) pressing the "New Page" button a second time when ready to point the handheld device camera at the new page. Device users can interchangeably use either camera-based page identification based on turning to the new page or signaling that senses the transition to the new page using a "New Page" switch.

[0245] Determining a user's intent to navigate to a new page can automatically trigger one or more actions to: 1) complete any interactive elements related to the previous page, and / or 2) introduce the new page to the device user. For example, actions might include: 1) Stop the audio, visual, tactile and / or other interactive elements of a continuously interactive sequence; 2) If a new page is out of order in its natural or identifiable order, generate an audio, visual, or tactile warning or indicator to indicate the page is out of order. 3) Generate one or more audio cues or “new page” indicators (e.g., ideally short sounds or sound effects familiar to the user) that signal that interactions involving the new page will follow; 4) Generate haptic cues or a series of vibrations (e.g., continuous vibrations for one or more predetermined time periods separated by short intervals) that signal an expected interaction involving a new page; 5) Turn on one or more visual indicators (e.g., one or more LEDs) on the device for one or more predetermined times (e.g., make one or more indicators appear to flash or blink); 6) Generate one or more changes in the intensity and / or color of one or more visual indicators (e.g., LEDs) on the device, thereby visually signaling the start of a new page interactive sequence; 7) Play on the device speaker (i.e., in the form of words and / or sound) the page title (if present), page number (if determined), one or more descriptions of one or more objects (e.g., images, drawings) identified on the new page, and / or text (and / or related comments) identified on the new page and / or within a page layout database (if available) (e.g., using optical character recognition); and / or 8) Turn off the device light source (especially during the display of text and / or new page images on the device speaker when it is not expected to use the light source for pointing).

[0246] Once the introductory action that initiates user interaction involving a new page has been performed (e.g., a new page prompt, playing text related to the new page on the device's speakers), the user can then explore any or all interactive elements within the new page. For example, exploration might include reactivating the device's light source (if needed) and using the light to point at any number of text or image elements within the new page. This exploration mode can continue until the device user instructs them to move to the next new page (e.g., by turning a page and pointing the device's light source at the next page, or by pressing a "New Page" button).

[0247] In other aspects of the device and method, during a typical interactive sequence, when the user encounters a delay after the expected response, guiding cues can be played on the device's speakers to help guide the user's interaction. Auditory cues can be helpful, especially for young children or learners, and / or during the initial stages of device use. However, cues can become repetitive or even distracting when too many are given and / or played at times that may not be helpful.

[0248] Based on the aspects of the method just described, the number of prompts can be reduced by: 1) minimizing the number of button presses (or other user signaling mechanisms) by recognizing the transition to a new page within camera-based images; 2) continuously monitoring camera-acquired images to quickly detect focus shifts to the new page; and 3) simplifying the interactive sequence by automating interactive steps to introduce each new page. In summary, a simple and rapid interactive process that automatically identifies and introduces new pages and reduces the number of user prompts can help maintain user engagement.

[0249] The foregoing disclosure of embodiments has been presented for purposes of illustration and description. This invention is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many variations and modifications of the examples described herein will be apparent to those skilled in the art based on the foregoing disclosure. It should be understood that, depending on the intended use of the examples, various components and features described with respect to particular examples may be added, removed, and / or replaced with other examples.

[0250] Furthermore, in describing representative examples, the specification may have presented methods and / or processes as a specific sequence of steps. However, the method or process should not be limited to the specific sequence of steps described herein, to the extent that it does not depend on the specific order of the steps set forth herein. Other sequences of steps are also possible, as will be understood by those skilled in the art. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims.

[0251] While the invention is readily adaptable to various modifications and alternatives, specific examples have been shown in the accompanying drawings and described in detail herein. It should be understood that the invention is not limited to the particular forms or methods disclosed; rather, it encompasses all modifications, equivalents, and alternatives falling within the scope of the appended claims.

Claims

1. A handheld device, the handheld device comprising: The main body of the device is configured to be manually operated by a user of the device. The electronic circuitry within the main body of the device includes a device processor. A device beam source configured to emit a beam of light away from the device body to generate a beam image reflection on a viewable surface that creates an image, the image being focused at a predetermined distance from the viewable surface; as well as A device camera, operatively coupled to the device processor and aligned such that the field of view of the device camera includes the reflection of the light beam image, wherein the handheld device is configured to: The beam image is generated on the viewable surface so that the image can be viewed by the device user and focused at the predetermined distance to encourage the device user to position the handheld device at the predetermined distance from the viewable surface. as well as Acquire camera images.

2. The handheld device of claim 1, wherein the image includes one of an object, symbol, or pattern on the viewable surface.

3. The handheld device of claim 1, wherein the device light source is one of: one or more light-emitting diodes, one or more micro light-emitting diodes, one or more laser diodes, and one or more digital light processing projector elements, each operatively coupled to the device processor.

4. The handheld device of claim 1, wherein the reflected light beam image is presented to the device user as one of the following: a smiley face, a human face, an animal face, a toy, a cartoon character, a circle, a rectangle, a polygon, an arrow, a cross, a finger, a hand, one or more characters, one or more numbers, one or more symbols, one or more recognizable shapes, and an animation.

5. The handheld device of claim 1, wherein the structured light pattern that generates the reflected light beam image is generated by one of the following: one or more light-blocking filters in the optical path of the device beam source; liquid crystal filter elements in the optical path of the beam source, each operatively coupled to the device processor; digital light processing projector elements, each operatively coupled to the device processor; and a plurality of beam sources, each operatively coupled to the device processor.

6. The handheld device of claim 5, wherein the device processor modifies the structured light pattern based at least in part on content determined by the device processor within the camera image.

7. The handheld device of claim 1, wherein the beam source is operatively coupled to the device processor and is further configured to control the beam intensity by adjusting one or both of a beam drive current and a pulse width modulation of the beam drive current.

8. The handheld device according to claim 7, wherein the device processor is further configured to: The image beam reflection in the camera image is identified; and The beam intensity is controlled based on the detected light intensity measured in the camera image within one or both of the beam image reflection region and the ambient light image region that does not contain the beam image reflection region.

9. A method for encouraging a device user to position a handheld device at a predetermined distance from a viewable surface, wherein the handheld device comprises: The main body of the device is configured to be operated by the hand of the user of the device; The electronic circuitry within the main body of the device includes a device processor; a device beam source configured to emit a beam of light to generate a beam image reflection that creates an image, the image being focused at the predetermined distance. The method includes: a device camera operatively coupled to the device processor and aligned such that the device camera's field of view includes the reflection of the light beam image. The device's beam source generates a beam image reflection on the viewable surface, such that the image can be viewed by the device user and focused at the predetermined distance to encourage the device user to position the handheld device at the predetermined distance from the viewable surface; as well as The camera of the device acquires camera images.

10. The method according to claim 9, further comprising: The device processor identifies the reflection of the light beam in the camera image; as well as The device processor determines the state of the beam reflected from the beam image within the camera image as either focused or out of focus, and the device processor performs an action based at least in part on the beam state.

11. The method of claim 10, wherein determining the beam image state comprises calculating one or more of the following by the device processor applied to the camera image: a trained neural network, one or more kernel operations, one or more image contrast measurements, and a Fourier transform.

12. The method of claim 10, wherein the beam source is operatively coupled to the device processor, and the action includes changing one or more of the beam intensity, beam color, and beam structured light pattern.

13. The method of claim 10, wherein the haptic unit within the handheld device is operatively coupled to the device processor, and the action includes activating the haptic unit.

14. The method of claim 10, wherein a device speaker within the handheld device is operatively coupled to the device processor, and the action includes playing an instruction sound on the device speaker.

15. The method of claim 9, wherein the viewable surface comprises one or more of paper, cardboard, film, cloth, wood, plastic, glass, painted surface, printed surface, textured surface, reinforced surface enhanced with three-dimensional elements, flexible surface, and electronic display.

16. The method of claim 9, wherein the viewable surface is one or more of a book, book cover, brochure, box, sign, newspaper, magazine, tablet computer, printable surface, tattoo, mobile device, e-book reader, and display screen.

17. The method of claim 9, wherein the image comprises one of an object, symbol, or pattern that can be viewed on the viewable surface.

18. A method for encouraging a device user to position a handheld device at a predetermined distance from a viewable surface, wherein the handheld device comprises: The main body of the device is configured to be operated by the hand of the user of the device; Electronic circuitry within the device body, the electronic circuitry including a device processor; a device beam source configured to emit a beam to generate a beam image reflection that creates an image viewable on the viewable surface; and one or more of one or more reflective surfaces and one or more refractive elements within a device beam path originating from the device beam source, each operatively coupled to the device processor. The method includes a device camera operatively coupled to the device processor and aligned such that the device camera's field of view includes the reflection of the light beam image. The beam image is generated by the beam source of the device on the viewable surface and reflected. The device processor moves one or more of the one or more reflective surfaces and one or more refractive elements to reflect and focus the beam image of the image at the predetermined distance, thereby encouraging the device user to position the handheld device at the predetermined distance from the viewable surface; and The camera of the device acquires camera images.

19. The method of claim 18, wherein each of the one or more reflective surfaces comprises one or more movable mirrors, one or more microelectromechanical systems (MEMS), and one or more movable prisms.

20. The method of claim 18, wherein moving one or more of the one or more reflective surfaces and the one or more refractive elements alters one or both of the beam path distance and the optical path direction of the beam traveled by the device beam.

21. A handheld device, the handheld device comprising: The main body of the device is configured to be manually operated by a user of the device. The electronic circuitry within the main body of the device includes a device processor. A device beam source, the device beam source being configured to project a beam away from the device body to generate a beam reflection at a beam reflection location on a surface that can be viewed by the user of the device; as well as A device camera, operatively coupled to the device processor and aligned such that the field of view of the device camera includes the beam reflection location, wherein the device processor is configured to: Acquire camera images; The selected page layout is identified based on the matching of the camera image with one or more predetermined interactive page layouts; as well as The action is performed based at least in part on the selected page layout.

22. The handheld device of claim 21, wherein the device processor is further configured to determine a selected page layout beam position based on the beam reflection position within the camera image matching the selected page layout, and to perform the action based at least in part on the selected page layout beam position.

23. The handheld device of claim 22, wherein the processor is further configured to determine the match between the camera image and one of the one or more predetermined interactive page layouts based at least in part on one or more of template matching, computer vision, machine learning, transformer models, and neural network classification for calculating the matching alignment between the camera image and the one or more predetermined interactive page layouts.

24. The handheld device of claim 22, wherein the device processor is further configured to determine the match between the camera image and one of the one or more predetermined interactive page layouts based at least in part on one of machine learning and neural network classification for calculating the matching alignment between the camera image and the one or more predetermined interactive page layouts, and wherein the one or more predetermined interactive page layouts are used to train the one or both of the machine learning and neural network classification.

25. The handheld device of claim 21, wherein the device processor is further configured to: Acquire one or more additional camera images; and It is determined that the image movement measured in the camera image and the one or more additional camera images is less than a predetermined movement threshold in the image acquisition time that is greater than a predetermined threshold time.

26. The handheld device of claim 22, wherein the device processor is further configured to perform the action, the action comprising one or more of the following: Transmit the camera image to one or more remote processors, acquire the image acquisition time of the camera image, the selected page layout, and the beam position of the selected page layout; Play one or more sounds on a device speaker operatively coupled to the device processor; Display one or more lighting patterns on one or more device displays operatively coupled to the device processor; as well as Activate the device haptic unit operatively coupled to the device processor.

27. The handheld device according to claim 21, wherein the device light source is one of a light-emitting diode and a laser diode.

28. The handheld device of claim 21, wherein the projected beam is one or more of collimated, incoherent, divergent, and patterned.

29. The handheld device of claim 21, wherein the device light source is operatively coupled to the device processor, and wherein the device processor is further configured to control the intensity of the device light source by adjusting one or more of an optical drive current and a pulse width modulation of the optical drive current.

30. The handheld device of claim 29, wherein the device processor is further configured to control the beam intensity based on detected light intensity measured in a camera image within one or both of a beam image region containing the beam reflection location and an ambient light image region not containing the beam image reflection.

31. The handheld device of claim 29, wherein the device processor is further configured to perform one or more of the following: The light source is turned off before the camera image is acquired, and The light source is turned on after the camera image is acquired.

32. The handheld device of claim 29, wherein the device processor is further configured to: Baseline images are acquired when the device's beam source is turned off; Acquire a beam reflection image including the beam reflection; The subtracted pixel intensity image is calculated by subtracting the baseline image from the beam reflection image; and The beam reflection position is calculated based on the number of beam sensing pixels in the subtracted pixel intensity image that exceed a predetermined light intensity threshold.

33. The handheld device of claim 21, wherein each of the one or more predetermined interactive page layouts comprises one or more of the following: one or more page positions of one or more page objects, one or more shapes of the one or more page objects, one or more camera images of the one or more page objects, one or more sizes of the one or more page objects, one or more colors of the one or more page objects, one or more textures of the one or more page objects, one or more patterns within the one or more page objects, one or more sounds generally generated by the one or more page objects, one or more names of the one or more page objects, one or more descriptions of the one or more page objects, one or more functions of the one or more page objects, one or more questions associated with the one or more page objects, one or more predetermined actions associated with one or more page areas, and one or more related objects associated with the one or more page objects.

34. The handheld device of claim 21, wherein the one or more predetermined interactive page layouts are displayed on one or more of the following: a book, a book cover, a brochure, a box, a sign, a newspaper, a magazine, a poster, a tablet computer, a printable surface, a painted surface, a textured surface, a flexible surface, an enhanced surface having one or more three-dimensional elements, a sign, a tattoo, a mobile device, a tablet computer, an e-book reader, a television, and a display screen.

35. A handheld device, the handheld device comprising: The main body of the device is configured to be manually operated by a user of the device. The electronic circuitry within the main body of the device includes a device processor. A device beam source, the device beam source being configured to project a beam away from the device body to generate a beam reflection at a beam reflection location on a viewable surface that can be viewed by the user of the device; A device camera, operatively coupled to the device processor and aligned such that the field of view of the device camera includes the location of the beam reflection. A device switch, operatively coupled to the device processor, wherein the device processor is configured to: Determine that the currently acquired switch state of the device switch is different from the previously acquired switch state of the device switch; Acquire camera images; The selected page layout is identified based on the matching of the camera image with one or more predetermined interactive page layouts; as well as The action is performed based at least in part on the selected page layout.

36. The handheld device of claim 35, wherein the device processor is further configured to determine a selected page layout beam position based on the beam reflection position within the camera image matching the selected page layout, and to perform the action based at least in part on the selected page layout beam position.

37. The handheld device of claim 35, wherein the handheld device further includes a device speaker operatively coupled to the device processor, and wherein the device processor is configured to play one or more indicator sounds on the device speaker when it is determined that the currently acquired switch state of the device switch is different from the previously acquired switch state of the device switch.

38. The handheld device of claim 35, wherein the viewable surface is one or more of a book, book cover, brochure, box, sign, newspaper, magazine, tablet computer, printable surface, tattoo, mobile device, e-book reader, and display screen.

39. The handheld device of claim 35, wherein the viewable surface is composed of one or more of paper, cardboard, film, cloth, wood, plastic, glass, painted surface, printed surface, textured surface, reinforced surface with three-dimensional elements, flexible surface, and electronic display.

40. A handheld device, the handheld device comprising: The main body of the device is configured to be manually operated by a user of the device. The electronic circuitry within the main body of the device includes a device processor. A device beam source, the device beam source being configured to project a beam away from the device body to generate a beam reflection at a beam reflection location on a viewable surface that can be viewed by the user of the device; A device camera, operatively coupled to the device processor and aligned such that the field of view of the device camera includes the location of the beam reflection. At least one device inertial measurement unit, said at least one device inertial measurement unit being operatively coupled to said device processor, said device processor being configured to: Inertial measurement data is acquired from the at least one device inertial measurement unit; The inertial measurement data is used to determine the duration for which the handheld device's movement remains below a predetermined movement threshold. Acquire camera images; The selected page layout is identified based on the matching of the camera image with one or more predetermined interactive page layouts; as well as The action is performed based at least in part on the selected page layout.

41. An apparatus, the apparatus comprising: The main body of the device is configured to be operated by a device user; The electronic circuitry within the main body of the device includes a device processor. A device light source, which is fixed to the device body and configured to emit projected light away from the device body to generate light reflection on a viewable page that can be viewed by the device user; as well as A device camera fixed to the main body of the device, the device camera including some or all of the field of view encompassing the light reflection. The device processor is configured to: The device acquires a first camera image via its camera; The first page layout is identified based on the matching of the first camera image with one of two or more predetermined page layouts; After acquiring the first camera image, a second camera image is acquired via the device camera; The second page layout is identified based on the matching of the second camera image with one of the two or more predetermined page layouts; Determine that the layout of the second page is different from the layout of the first page; and Perform the action.

42. The device of claim 41, wherein the device light source is operatively coupled to the device processor, and wherein the action includes turning off the device light source.

43. The device of claim 42, wherein the device further comprises a device switch operatively coupled to the device processor, and wherein the device processor is further configured to turn on the device light source when it is determined that a currently acquired switch state of the device switch is different from a previously acquired switch state of the device switch.

44. The device of claim 41, wherein the device further comprises a device speaker operatively coupled to the device processor, and wherein the action comprises playing one or more of the following on the device speaker: one or more new page prompt sounds, one or more out-of-order page prompts, one or more words obtained from the second page layout, and one or more sounds obtained from the second page layout.

45. The device of claim 44, wherein the device further comprises a device switch operatively coupled to the device processor, and wherein the device processor stops the playback on the device speaker when it is determined that a currently acquired switch state of the device switch is different from a previously acquired switch state of the device switch.

46. ​​The device of claim 41, wherein the device further comprises one or more device indicator lights operatively coupled to the device processor, and wherein the action comprises one or more of the following: turning on the one or more device indicator lights for one or more predetermined times, changing the intensity of the one or more device indicator lights, and changing the color of the one or more device indicator lights.

47. An apparatus, the apparatus comprising: The main body of the device is configured to be operated by a device user; The electronic circuitry within the main body of the device includes a device processor. A device light source, which is fixed to the device body and configured to emit projected light away from the device body to generate light reflection on a viewable page that can be viewed by the device user; as well as A device camera fixed to the main body of the device, the device camera including some or all of the field of view encompassing the light reflection. The device processor is configured to: The device acquires a first camera image via its camera; Identify the first page within the first camera image; After acquiring the first camera image, a second camera image is acquired via the device camera; Identify the second page within the second camera image; It is determined that the second page is different from the first page; as well as Perform the action.

48. The device of claim 47, wherein the device is configured to be operated by one of the device user's hand, the device user's head, the device user's wrist, the device user's arm, the device user's shoulder, the device user's leg, and the device user's chest.

49. The device of claim 47, wherein the device light source is operatively coupled to the device processor, and wherein the action includes turning off the device light source.

50. The device of claim 49, wherein the device further comprises a device switch operatively coupled to the device processor, and wherein the device processor is further configured to turn on the device light source when it is determined that a currently acquired switch state of the device switch is different from a previously acquired switch state of the device switch.

51. The device of claim 49, wherein the device further comprises a device speaker operatively coupled to the device processor, and wherein the action comprises playing one or more of the following on the device speaker: one or more new page prompts, one or more out-of-order page prompts, one or more text words identified in the second camera image, and one or more descriptive words relating to one or more objects identified in the second camera image.

52. The device of claim 51, wherein the device further comprises a device switch operatively coupled to the device processor, and wherein the device processor stops the playback on the device speaker when it is determined that a currently acquired switch state of the device switch is different from a previously acquired switch state of the device switch.

53. The device of claim 51, wherein the device light source on the device speaker is turned on after the playback is completed.

54. The device of claim 47, wherein the device further comprises a device haptic unit operatively coupled to the device processor, and wherein the action comprises activating the device haptic unit for one or more predetermined times.

55. The device of claim 47, wherein the device further comprises one or more device indicator lights operatively coupled to the device processor, and wherein the action comprises one or more of the following: turning on the one or more indicator lights for one or more predetermined times, changing the intensity of the one or more device indicator lights, and changing the color of the one or more device indicator lights.

56. An apparatus, the apparatus comprising: The main body of the device is configured to be operated by a device user; The electronic circuitry within the main body of the device includes a device processor. A device switch, wherein the device switch is fixed to the device body; A device light source, which is fixed to the device body and configured to emit projected light away from the device body to generate light reflection on a viewable page that can be viewed by the device user; as well as A device camera fixed to the main body of the device, the device camera including some or all of the field of view encompassing the light reflection. The device processor is configured to: The device acquires a first camera image via its camera; Identify the first page within the first camera image; The first switch state is obtained via the device switch; After obtaining the first switch state, the second switch state is obtained via the device switch; It is determined that the second switch state is different from the first switch state; The second camera image is acquired via the device's camera; Identify the second page within the second camera image; It is determined that the second page is different from the first page; as well as Perform the action.

57. The device of claim 56, wherein the device switch comprises one of a push-button switch, a rocker switch, a contact switch, and a proximity-sensitive switch.

58. The device of claim 56, wherein the device light source is operatively coupled to the device processor, and wherein the action includes turning off the device light source.

59. The device of claim 56, wherein the device further comprises a device speaker operatively coupled to the device processor, and wherein the action comprises playing one or more of the following on the device speaker: one or more new page prompts, one or more out-of-order page prompts, one or more text words identified in the second camera image, and one or more descriptive words relating to one or more objects identified in the second camera image.

60. The device of claim 59, wherein when it is determined that the currently acquired switch state of the device switch is different from the previously acquired switch state of the device switch, the device processor stops the playback on the device speaker.

61. A method for performing an action based on a specified page location selected by a person using a handheld device, the handheld device comprising: Device processor; A device light source configured to generate a projected beam of light that produces one or more light reflections from one or more visible objects that can be viewed by the person. The method includes a device camera, the device camera being aligned such that the camera's field of view includes the reflection locations of the one or more light reflections, and being operatively coupled to the device processor. The device processor obtains one or more predetermined interactive page layouts; When the handheld device is manipulated such that the projected beam is directed from the device's light source to the designated page location, the device's camera acquires a camera image; The device processor calculates the localized image based on the matching of the camera image with one or more predetermined interactive page layouts; The device processor identifies the specified page location based on the reflection position within the located image; and The action is performed by one or both of the device processor and the remotely connected processor, at least in part, based on the specified page position within the one or more predetermined interactive page layouts.

62. The method of claim 61, wherein the device light source is one of a light-emitting diode and a laser diode.

63. The method of claim 61, wherein the device light source is operatively coupled to the device processor, the method further comprising controlling the intensity of the device light source by adjusting one or more of an optical drive current and a pulse width modulation of the optical drive current.

64. The method of claim 61, wherein the projected beam is one or more of collimated, incoherent, divergent, and patterned.

65. The method of claim 61, wherein the one or more predetermined interactive page layouts include one or more of the following: one or more locations of one or more page objects, one or more shapes of the one or more page objects, one or more camera images of the one or more page objects, one or more sizes of the one or more page objects, one or more colors of the one or more page objects, one or more textures of the one or more page objects, one or more patterns within the one or more page objects, one or more sounds generally generated by the one or more page objects, one or more names of the one or more page objects, one or more descriptions of the one or more page objects, one or more functions of the one or more page objects, one or more questions associated with the one or more page objects, and one or more related objects associated with the one or more page objects.

66. The method of claim 61, wherein the one or more predetermined interactive page layouts are displayed on one or more of the following: a book, a book cover, a brochure, a box, a sign, a newspaper, a magazine, a poster, a tablet computer, a printable surface, a painted surface, a textured surface, a flexible surface, an enhanced surface having one or more three-dimensional elements, a sign, a tattoo, a mobile device, a tablet computer, an e-book reader, a television, and a display screen.

67. The method of claim 61, wherein calculating the localized image comprises using one or more of template matching, computer vision, machine learning, transformer models, and neural network classification to calculate the matching position of the camera image with the one or more predetermined interactive page layouts.

68. The method of claim 67, wherein one or both of the machine learning and the neural network classification are trained on the one or more predetermined interactive page layouts.

69. The method of claim 61, wherein calculating the image of the positioning includes determining a horizontal position of the camera image, a vertical position of the camera image, a magnification of the camera image, and a camera image orientation that generate the matching position of the camera image with one or more components of the predetermined interactive page layout.

70. The method of claim 61, wherein the predetermined beam pointing position within the positioned image is determined by: The device camera acquires a baseline image that excludes reflections from the projected beam; The device camera acquires a beam reflection image, the beam reflection image including one or more reflections generated by the projected beam; The device processor calculates the subtracted pixel intensity image based on subtracting the baseline image from the beam reflection image; as well as The device processor assigns the center position of the light beam sensing pixel in the subtracted pixel intensity image that exceeds a predetermined light intensity threshold to the predetermined light beam pointing position.

71. The method of claim 61, wherein the light source is operatively coupled to the device processor, the method further comprising one or more of the following: turning on the light source after acquiring the predetermined page layout, turning off the light source before acquiring the camera image, and turning off the light source after acquiring the camera image.

72. The method of claim 61, wherein the action comprises one or more of the following: Transmit the camera image to one or more remote processors, acquire the acquisition time of the camera image, the predetermined beam pointing position, one or more predetermined interactive page layouts, and the specified page position; Play one or more sounds on a device speaker operatively coupled to the device processor; Display one or more lighting patterns on one or more device displays operatively coupled to the device processor; as well as Activate the device haptic unit operatively coupled to the device processor.

73. A method for performing an action based on a specified page location within an identified context selected by a person using a handheld device, the handheld device comprising: Device processor; A device light source configured to generate a projected beam of light that produces one or more light reflections from one or more visible objects that can be viewed by the person. The method includes a device camera, the device camera being aligned such that the camera's field of view includes the reflection locations of the one or more light reflections, and being operatively coupled to the device processor. The device processor obtains one or more predetermined context templates; When the handheld device is manipulated such that the projected beam is directed from the device light source toward the context object, a first camera image is acquired by the device camera; The device processor classifies the identified context based on the matching of the first camera image with one or more predetermined context templates; The device processor obtains one or more predetermined interactive page layouts associated with the identified context; When the handheld device is manipulated such that the projected beam is directed from the device's light source to the designated page position, the device's camera acquires a second camera image; The device processor calculates the localized image based on the matching of the second camera image with one or more predetermined interactive page layouts; The device processor identifies the specified page location based on the reflection position within the located image; and The action is performed by one or both of the device processor and the remotely connected processor, at least in part, based on the specified page position within the one or more predetermined interactive page layouts.

74. The method of claim 73, wherein calculating the identified context further comprises the device processor isolating the first camera image to the region in the first camera image pointed to by the projected beam.

75. The method of claim 73, wherein the one or more predetermined context templates include attributes associated with one or more book covers, magazine covers, book chapters, pages, toys, people, anatomical elements, animals, pictures, drawings, words, phrases, household items, classroom items, tools, cars, and clothing.

76. The method of claim 73, wherein the action comprises one or more of the following: Transmit one or more of the following to one or more remote processors: the one or more predetermined context templates, the first camera image, a first acquisition time for acquiring the first camera image, the second camera image, a second acquisition time for acquiring the second camera image, the one or more predetermined interactive page layouts, the positioned image, the predetermined beam pointing position, and the specified page position; Play one or more sounds on a device speaker operatively coupled to the device processor; Display one or more lighting patterns on one or more device displays operatively coupled to the device processor; as well as Activate the device haptic unit operatively coupled to the device processor.

77. A method for performing an action based on a specified page location selected by a person using a handheld device, the handheld device comprising: Device processor; A device light source configured to generate a projected beam of light that produces one or more light reflections from one or more visible objects that can be viewed by the person. The method includes a device camera, the device camera being aligned such that the camera's field of view includes the reflection locations of the one or more light beams, and being operatively coupled to the device processor. The device processor obtains one or more predetermined interactive page layouts; When the handheld device is manipulated such that the projected beam is directed from the device's light source to the designated page location, two or more camera images are acquired by the device's camera; The device processor determines that the image movement measured in the two or more camera images is less than a predetermined movement threshold, and the two or more camera images were acquired over an acquisition time greater than a predetermined dwell time threshold; The device processor calculates the localized image based on the matching of the camera image with one or more predetermined interactive page layouts; The device processor identifies the specified page location based on the reflection position within the located image; and The action is performed by one or both of the device processor and the remotely connected processor, at least in part, based on the specified page position within the one or more predetermined interactive page layouts.

78. The method of claim 77, wherein one or more of template matching, computer vision, and neural network classification are used to calculate one or more spatial offsets of a pair of consecutively acquired camera images, and the image movement is calculated based on the sum of the one or more spatial offsets.

79. The method of claim 77, wherein the light source is operatively coupled to the device processor, the method further comprising one or more of the following: turning on the light source after acquiring the predetermined page layout, turning off the light source before acquiring each of the two or more camera images, turning on the light source after acquiring each of the two or more camera images, and turning off the light source after determining that the image movement is less than the predetermined movement threshold and the acquisition time is greater than the predetermined dwell time threshold.

80. The method of claim 77, wherein the action comprises one or more of the following: Transmit one or more of the following to one or more remote processors: the two or more camera images, the image movement, the predetermined dwell time threshold, the acquisition time of the two or more camera images, the predetermined beam pointing position, the one or more predetermined interactive page layouts, and the specified page position; Play one or more sounds on a device speaker operatively coupled to the device processor; Display one or more lighting patterns on one or more device displays operatively coupled to the device processor; as well as Activate the device haptic unit operatively coupled to the device processor.

Citation Information

Patent Citations

  • Systems and methods for time-sharing interactions using a shared artificial intelligence personality

    US10915814B2

  • Systems and methods for time-shifting interactions using a shared artificial intelligence personality

    US10963816B1

  • Systems and methods for bimanual control of virtual objects

    US11334178B1

  • Systems and methods to enhance interactive engagement with shared content by a contextual virtual agent

    US11366997B2

  • Systems and methods for collective control of virtual objects

    US11409359B1