A system and method for interactively specifying page locations by interacting with discoverable objects and pointing a light beam using a portable device.

A portable device with a light beam and computer vision enhances interaction by providing auditory, tactile, and visual prompts, addressing the lack of intuitive methods in existing devices, improving educational effectiveness and engagement.

JP2026510211APending Publication Date: 2026-04-02KIBEAM LEARNING INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing electronic devices, particularly for children, lack intuitive and safe interaction methods that do not require precise manual dexterity or understanding of screen-based operations, limiting their usability and educational effectiveness.

Method used

A portable device using a light beam emitted from a laser pointer or LED, combined with computer vision and inertial measurement units, provides auditory, tactile, and visual prompts to help users identify and interact with real or virtual objects, enhancing learning and engagement through interactive page navigation.

Benefits of technology

The system enables intuitive object selection and location identification, providing continuous educational support, enhancing literacy, mathematics, and science comprehension, while motivating physical activity and emotional engagement without overwhelming children.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510211000001_ABST
    Figure 2026510211000001_ABST
Patent Text Reader

Abstract

A system and method are described for communicating one or more attributes or cues related to a detectable object to a user of a mobile device by auditory, tactile, and / or visual means. Using a light beam generated by the mobile device and directed in the same direction as the device's camera, the user can point to one or more objects in their environment that are thought to be related to one or more attributes or cues. Based on the pointed-out area in the camera image, the mobile device classifies the objects in the image and determines whether they match a template or classification of a detectable object. The match or mismatch of the pointed-out object may determine subsequent actions. This system and method provides a simple and intuitive approach for human-machine interaction, particularly suitable for learners.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related application data This application claims the benefit of priority of co-pending provisional application serial number 63 / 441,731, filed January 27, 2023, and is a partial continuation of co-pending U.S. application serial numbers 18 / 201,094, filed May 23, 2023, 18 / 220,738 (filed July 11, 2023), and 18 / 382,456 (filed October 20, 2023). The disclosures of these applications are hereby incorporated by reference in their entirety.

[0002] Technical field This application generally relates to systems and methods for identifying physical or virtual objects and / or objects or locations displayed within a page (e.g., a page of a book) using a light beam emitted from a portable electronic device by an individual. This portable device can be used by anyone, but is particularly suitable for young children and learners. This is to utilize simple interactive signals that do not require precise manual dexterity or an understanding of screen-based interactive operation sequences. The systems and methods described herein employ technologies in the fields of mechanical design, electronic design, firmware design, computer programming, inertial measurement units (IMUs), optics, computer vision (CV), ergonomics (including safety design for children), human motion control, and human-machine interaction. The systems and methods can provide a user-friendly mechanical interface for intuitively and confidently indicating the selection of a discoverable object from among a plurality of visible objects, particularly for children and learners.

Background Art

[0003] background In recent years, the world has become increasingly reliant on portable electronic devices, which have become more powerful, sophisticated, and useful for a wide range of users. However, while children may quickly master certain functions of electronic devices designed for experienced users, younger children would benefit from access to small, lightweight, colorful, playful, informative, ergonomically designed (including child safety features), and easy-to-use interactive electronic devices. The systems and methods disclosed herein leverage advances in the field of optics, such as visible light (so-called "laser") pointers, mobile acoustic and / or vibration generators (employing miniature coils, piezoelectric elements, and / or tactile units), portable displays, inertial measurement devices (also known as inertial motion units), and communication technologies.

[0004] The beam of a visible light pointer (also called a "laser pen") is commonly used in business and educational settings and is typically generated by a laser diode (i.e., a PIN diode) that has a pure (I) semiconductor without impurity doping between the p-type (P) and n-type (N) semiconductor regions. When operated properly within a given power level, such a coherent, parallel light source is generally considered safe. Furthermore, if the light is directed towards the eye, the corneal reflex (also called the blink reflex or eyelid reflex) ensures an unconscious avoidance response to bright light (and foreign objects).

[0005] However, using non-coherent light-emitting diode (LED) light sources may offer further eye safety. Such non-coherent light sources (so-called "point source" LEDs) can be collimated using precision optics (e.g., including so-called "pre-collimation") to produce a light beam with minimal and / or controlled divergence angles. If necessary, point source LEDs can also produce beams consisting of a wide spectral frequency range compared to the predominantly monochromatic light produced by a single laser.

[0006] Speakers installed in televisions, theaters, and other fixed installations typically employ one or more electromagnetic induction coils. Within portable and / or mobile devices, similar electromagnetic coil schemes and / or piezoelectric (also known as "buzzers") designs may be used to generate vibrations in small speakers. Vibrations (particularly those associated with alerts) can also be generated by haptic units (also known as kinesthetic communication). Haptic units typically use eccentric (i.e., unbalanced) rotating masses or piezoelectric actuators to generate vibrations (particularly in the low frequencies of the audio spectrum) that are perceptible by hearing and / or touch.

[0007] A visual display device or indicator can consist of any number of addressable light sources or pixels, either monochromatic or multicolored. Displays range from single light sources (e.g., illuminating a sphere transmitted through a waveguide) to those capable of displaying a single digit (e.g., a 7-segment display) or alphanumeric characters (e.g., a 5x8 pixel array) to high-resolution screens with tens of millions of pixels. Regardless of scale, displays are typically implemented as follows: 1) a two-dimensional array of light sources (the most common form being light-emitting diodes (LEDs), including organic light-emitting diodes (OLEDs)); or 2) two polarized glass plates sandwiching a liquid crystal material that reacts to an electrical signal to allow light of different wavelengths from multiple light sources (backlights) to pass through (i.e., forming a liquid crystal display (LCD)).

[0008] Tracking using inertial measuring units (IMUs), accelerometers, and / or magnetometers may incorporate any or all combinations of the following: 1) linear accelerometers that measure forces (i.e., governed by Newton's second law of motion) generated during motion in up to three axes or three dimensions; 2) gyroscope-based sensing that measures rotational velocity or acceleration in up to three rotational axes; 3) magnetometers that measure magnetic fields (i.e., magnetic dipole moments), including the magnetic field generated by the Earth; and / or 4) magnetometers that measure the Earth's gravity (including gravitational direction) by measuring the force acting on internal mass. The accuracy of IMUs, accelerometers, and magnetometers varies considerably depending on size, operating range, compensation hardware that may be used to correct the measurement (impacting cost), environmental factors including temperature gradients, whether or not individual device calibration is performed, and the time required for measurement (including integration time in some measurement types).

[0009] Advances in electronic devices (i.e., hardware), standardized communication protocols, and dedicated frequency allocations within the electromagnetic spectrum have led to the development of a diverse range of portable devices capable of wirelessly communicating with other nearby devices and large-scale communication systems, including the World Wide Web and the metaverse. When selecting the protocols (or combinations of available protocols) to employ in these portable devices, factors such as power consumption, communication range (e.g., from a few centimeters to hundreds of meters or more), and available bandwidth are considered.

[0010] Currently, Wi-Fi (e.g., based on the IEEE 802.11 standard family) and Bluetooth (managed by the Bluetooth Special Interest Group) are used in many mobile devices. In home environments, less common / older communication protocols also exist for mobile devices, such as Zigbee, Zwave, and cellular / mobile communication-based networks. Generally (with many exceptions, especially when considering newer standards), Wi-Fi offers a wider range, greater bandwidth, and a more direct path to the internet compared to Bluetooth. Bluetooth (including Bluetooth Low Energy (BLE)), on the other hand, offers lower power consumption, a shorter operating range (which can be advantageous in some applications), and a simpler circuit configuration to support communication.

[0011] Advances in miniaturization technology, reduced power consumption, and the sophistication of electronic devices, including displays, IMUs, and communication equipment, have revolutionized the mobile device industry. These portable devices continue to evolve, enabling users to simultaneously communicate, operate, locate, monitor exercise, manage health, receive warnings, record videos, and conduct financial transactions. Systems and methods that enable simple and intuitive operation with handheld pointing devices are invaluable. [Overview of the project]

[0012] overview From the above perspective, this specification provides a system and method for describing a lightweight, easy-to-use, and intuitive handheld device particularly suitable for machine-based interaction by children and other learners. While the device may be accepted by children in part as a toy, the flexibility of the computational processing built into the device allows it to be used as a means for play, embodied learning, emotional support, cognitive development, communication, expression of creativity, cultivation of mindfulness, and / or enhancement of imagination.

[0013] Mobile devices may support areas related to literacy, mathematics, science and technology comprehension, basic reading comprehension, interactive reading comprehension, and CROWD (completion, recall, open-ended, WH-questions, and distance-finding) question formats. By providing continuous (machine-based) educational support, mobile devices may enhance learners' Zone of Proximal Development (ZPD). Furthermore, portable, lightweight, and "fun" mobile devices may motivate children (and adults) to engage in physical activity, including motor and sensorimotor activities.

[0014] According to one embodiment, a system and method are provided for an individual to identify a real or virtual object from among several visible objects based on one or more auditory, tactile, and / or visual prompts and / or cues generated by a portable electronic device. The prompts and / or cues relating to the discoverable object may be based on one or more features, attributes, and / or relevances associated with the object. Once the prompts and / or cues are provided, the user of the portable device can direct a beam of light emitted from the device towards a visible object in their environment that is considered most relevant to the prompts and / or cues from the perspective of the device user.

[0015] A camera on a mobile device, pointed in the same direction as the light beam source, acquires one or more images of objects illuminated by the light beam. The classification(s) of the objects illuminated in the camera images are compared to one or more templates and / or classifications associated with discoverable objects to determine whether there is a match with prompts and / or cues.

[0016] One or more auditory, tactile, and / or visual prompts or cues may include aspects that a particular individual may know (e.g., proper names and / or birthdays of family members or pets) or aspects that may be more generally associated with the discoverable object (e.g., common name, action, sound, function, combination with other objects, typical color). In the case of children, special attention may be paid to age appropriateness and educational level when presenting prompts or cues via mobile devices.

[0017] Examples of audio prompts and cues include: sounds and sound effects typically produced by discoverable objects; linguistic descriptions of activities involving the object; clear names of the object (including proper nouns); descriptions of the object; questions about the object's function; linguistic descriptions of the object's attributes; and linguistic quizzes where the discoverable object is the answer. Tactile attributes or cues may include vibrations synchronized with (or at least at a similar frequency to) the actions and sounds typically produced by discoverable objects.

[0018] Visual prompts or cues may include: displaying the name of a discoverable object, an image of the object, an outline of the object, a portrait of the object, the name of an object category, one or more colors of the object (e.g., displaying an actual color swatch and / or one or more words describing the color), the size of the object (e.g., its size relative to other displayed objects), a sentence or question about the object, an object missing from a known list of objects (e.g., a letter in the alphabet), or a fill-in-the-blank phrase in which the object is the missing element. Visual prompts or cues may be displayed on one or more mobile device displays, projected (including scrolling) within a mobile device beam, or displayed on one or more separate (digital) screens.

[0019] A mobile device can acquire an image that includes a visible object pointed to by the mobile device user. The visible object is located in the image based on the presence of (spot or regional) reflections from the visible object generated by a light beam.

[0020] Alternatively, the position of a selected object within the camera image may be determined based on the recognition of the beam position, which results from the light beam source and camera being aligned. In other words, even when the light beam is off, the beam position can be calculated based on geometry (i.e., knowing the origin and direction of both the light beam source and the camera). In this case, the beam simply serves as an indicator that visually shows the device user the area within the field of view of the portable device camera, where the pointed-to object is compared to a template of detectable objects.

[0021] Subsequently, the selected objects are classified using computer vision techniques (e.g., neural networks, template matching). A match with a detectable object is declared if the boundary region of the selected object (and optionally other objects visible in the camera image) correlates well with the template and / or classification associated with the detectable object.

[0022] The presence or absence of a match may be indicated to the user of the mobile device using visual (e.g., by using one or more device display devices and / or by modulating a light beam source), auditory (e.g., by using a device speaker), and / or vibratory (e.g., by using a device haptic unit) feedback. Based on the match (or mismatch) with a detectable object, the mobile device may perform additional actions, including sending one or more prompts and / or queuings, timing and identification information (optionally including acquired images) of the detected object, or a non-detectable state, to one or more remote processors that may perform further actions.

[0023] According to one example, a method for a human to indicate a detectable object using a mobile terminal is provided. The mobile terminal includes a terminal processor, a terminal light beam source configured to generate a light beam that generates one or more light beam reflections from one or more visible objects visible to a human, a device camera arranged such that a camera field of view includes a beam position area of one or more light beam reflections and operatively connected to a device processor, and a device speaker operatively connected to the device processor. The method includes the following: playback of one or more audible cues related to a detectable object by the device speaker; a human operating the mobile terminal and the device camera capturing a camera image when a light beam emitted from the device light beam source is directed at one or more visible objects; the device processor identifying one or more indicated objects in the beam position area within the camera image; and the device processor determining whether one or more of the one or more indicated objects match a predetermined template of a detectable object.

[0024] According to another example, a method for a human to indicate a detectable object using a mobile terminal is provided. The mobile terminal includes a terminal processor, a terminal light beam source configured to generate a light beam that generates one or more light beam reflections from one or more visible objects visible to a human, a device camera arranged such that a camera field of view includes a beam position area of one or more light beam reflections and operatively connected to a device processor, and a device speaker operatively connected to the device processor. The method includes the following: playback of one or more audible cues related to a detectable object by the device speaker; a human operating the mobile device and the device camera acquiring a camera image when a projection light beam is irradiated from the device light beam source toward one or more visible objects; the device processor separating one or more indicated objects in the beam position area within the camera image; and the device processor determining whether one or more of the one or more indicated objects match a predetermined template of a detectable object.

[0025] According to another embodiment, a method for a human to use a mobile terminal to indicate a detectable object is provided. The mobile terminal includes a terminal processor, a terminal light beam source configured to generate a light beam that generates one or more light beam reflections from one or more visible objects visible to a human, a terminal camera disposed to include a beam position region of the one or more light beam reflections in a field of view and operatively connected to the terminal processor, and one or more device displays operatively connected to the device processor. The method includes: displaying, by the one or more device displays, one or more visual cues related to a detectable object; when a human operates the mobile terminal and the irradiated light beam is irradiated from the terminal light beam source toward one or more visible objects, taking a camera image by the terminal camera; separating, by the terminal processor, one or more indicated objects in the beam position region in the camera image; and determining, by the terminal processor, whether one or more of the one or more indicated objects match a predetermined template of a detectable object.

[0026] According to yet another example, a method for a human to use a mobile terminal to indicate a detectable object is provided. The mobile terminal includes a terminal processor, a terminal light beam source configured to generate a light beam that generates one or more light beam reflections from one or more visible objects visible to a human, a device camera (disposed such that a camera field of view includes a beam position region of the one or more light beam reflections and operatively connected to the device processor), and a device tactile unit operatively connected to the device processor. The method includes: generating, by the device tactile unit, a perceptible tactile vibration at one or more tactile frequencies related to one or more actions and sounds related to a detectable object; taking, by the terminal camera, a camera image when the mobile terminal is operated and the projected light beam is irradiated from the terminal light beam source toward one or more visible objects; separating, by the terminal processor, one or more indicated objects in the beam position region in the camera image; and determining, by the terminal processor, whether one or more of the one or more indicated objects match a predetermined template of a detectable object.

[0027] Another example provides a method for a human to indicate a detectable object using a mobile device. The mobile device comprises a terminal processor, a terminal light beam source configured to generate a light beam that produces one or more light beam reflections from one or more visible objects that are visible to a human, a terminal camera operationally connected to the terminal processor and positioned to include the beam position regions of one or more light beam reflections in the camera's field of view, a device speaker operationally connected to a device processor, and a device inertial measurement unit operationally connected to the device processor. This method includes: playback of one or more audible cues related to discoverable objects by a device speaker; acquisition of one or more camera images by a device camera when a handheld device is operated and a projection beam is directed from a device light beam source toward one or more visible objects; acquisition of inertial measurement data by a device inertial measurement unit; determination by a device processor that the inertial measurement data includes any of the following: a device operation amount less than a predetermined maximum operation threshold at a predetermined minimum time threshold, a predetermined handheld device gesture operation, or a predetermined handheld device orientation; separation of the most recent camera image from one or more camera images by the device processor; separation of one or more indicated objects in the beam position region within the most recent camera image by the device processor; and determination by the device processor whether one or more of the indicated objects match a predetermined template of discoverable objects.

[0028] Another example involves providing portable devices that can support areas related to basic reading comprehension, literacy, interactive reading, CROWD (Completion, Recall, Open-ended, WH-prompt [where, when, why, what, who], and Distancing) questioning, mathematics, and science and technology understanding. Furthermore, these devices may help traverse learners' zones of proximal development (i.e., the difference between what learners can achieve without assistance and what they can achieve with instruction, the ZPD) by providing ubiquitous (machine-based) educational support and / or guidance. Transforming page content (e.g., within books and magazines) into visual, auditory, and tactile experiences can significantly accelerate the acquisition of new knowledge, skills, and memories. In addition, portable, lightweight, and "fun" handheld devices can motivate children (and adults) to engage in physical activity, including motor and sensorimotor activities.

[0029] According to one embodiment, a system and method are provided for an individual to select objects and locations from visible content within, for example, the pages of a book or magazine. As will be detailed in the detailed description below, in this specification, “page” means a substantially two-dimensional surface on which visible content can be displayed. Page content may include any combination of text, symbols, drawings, and images, and may be displayed in color, grayscale, or black and white.

[0030] An individual can use a light beam emitted from a mobile device to select (i.e., point to, identify, and / or indicate) objects and / or locations within a page. The camera in the mobile device is directed in the same direction as the light beam and can acquire an image of the pointed-to area of ​​the page. The image may include reflections of the light ray (e.g., incident light reflected from the page). Alternatively, the light ray may be (temporarily) switched off during image acquisition (e.g., to allow capturing page content without interference from light reflections).

[0031] In either case, the position of the light beam in the camera-acquired image can be determined based on the fact that the camera and beam are positioned close together (and move in conjunction) within the mobile device body and are pointing in the same direction. The beam position in the camera-acquired image can be calculated, for example, by calculations based on the geometric relationship between beam directionality and camera imaging, and / or by calibration processing that determines the beam reflection position in the camera-acquired image through actual measurement.

[0032] The selection made by the mobile device user is notified using various display methods that utilize one or more mobile device detection elements. The selection (and optionally control of the light beam, e.g., on / off) is indicated using device switches such as push buttons, contact switches, and proximity sensors. Alternatively, the selection (and beam control) can be indicated by identifying keywords, phrases, or voices spoken by the user and detected by the device microphone.

[0033] Alternatively, the user can indicate a selection by controlling the orientation and movement of the mobile device (e.g., motion gestures, taps on the device, taps on other objects) detected by the built-in IMU. In the signal generation mechanism that moves the mobile device, images acquired by the camera immediately before the action can be used to identify the selected object and its location on the page.

[0034] In further aspects of the system and method, the processor within the mobile device can acquire a predetermined interactive page layout (which may include one or more object templates) of a set of pages that an individual might view (e.g., books, magazines). Computer vision (CV) techniques (e.g., neural networks, machine learning, transformers, generative artificial intelligence (AI), and / or template matching) can be used to determine the match between the image acquired by the camera and the page layout or one or more object templates. CV-based matching can identify both the page the device user is viewing (e.g., within a book or magazine) and the location, object, word, or target within the page pointed to by the light beam.

[0035] Optionally, page and content selection can be performed in two stages: first, contextual objects (e.g., book cover, title, word or phrase, printed material, drawing, real-world object) are selected and a "context" (e.g., specific book or chapter, book or magazine type, topic area, skill requirements) is assigned, preparing for subsequent selections. Then, page layouts containing only object templates relevant to the identified context are considered for CV-based matching with images acquired by the camera.

[0036] Compared to global classification methods that identify objects within an image, cross-validation and / or AI methods that match camera-acquired images against a finite set of predefined page layouts and / or object templates may reduce the computational resources required (e.g., on portable handheld devices) and contribute to improved accuracy in identifying content pointed to by a light beam.

[0037] Predefined interactive page layouts include (or refer to) interactive actions and additional content that can be performed as a result of selecting objects, locations, or areas within the page. Examples include: word pronunciation, sound effects, display of object spelling (e.g., displaying as text), playback of audio elements of selected objects, additional story content, questions about page content, rhythmic features associated with selected objects (e.g., those that may form the basis of tactile or vibratory stimulation to the device user's hand), audio feedback of reward or comfort for making a particular selection, and related additional or sequential content.

[0038] Based on interactive content within a designated interactive page layout database, one or more visual, auditory, and / or haptic actions can be performed on the mobile device itself. Alternatively, selections, attributes, and timing of selections can be transmitted to one or more external processors, which can then record interactions (e.g., for educational and / or parental monitoring) and / or perform additional and / or complementary interactions involving the mobile device user. Automatic recording of reading immersion, content, progress, and / or comprehension (e.g., comparison with other children of the same age) can provide insights into the child's emotional and cognitive development.

[0039] In summary, by directing light rays towards objects and locations on a page, mobile device-based interaction, along with audio elements, additional visual components, and / or tactile stimuli transmitted to the user's hands, can help "bring printed or displayed content to life." Extending printed content with interactive sequences that include real-time feedback relevant to the content can not only provide constant machine-based guidance during reading, but also help deliver "fun" in a learning environment and maintain emotional engagement (without being overwhelming).

[0040] The present invention provides a method for performing an action based on a specified page location selected by a human using a mobile device, according to one example. The mobile device comprises a terminal processor, a terminal light source configured to generate a projected light beam that produces one or more light reflections from one or more visible objects visible to a human, and a device camera that is aligned such that its camera field of view includes the reflection locations of one or more light beam reflections and is operationally connected to the device processor. The method includes: the device processor acquiring one or more predetermined interactive page layouts; the device camera acquiring a camera image with the mobile device operated and the projected light beam illuminating the specified page location from the device light source; the device processor calculating a positioned image based on the matching of the camera image with one or more predetermined interactive page layouts; the device processor identifying a specified page location based on the reflection locations in the positioned image; and the device processor and / or remotely connected processor performing an action at least partially based on the specified page location in one or more predetermined interactive page layouts.

[0041] A method is provided to perform an action based on a specified page position within a identified context, as in another example. This context is selected by a human using a mobile device, which includes a device processor, a device light source (configured to generate a projected light beam that produces one or more light reflections from one or more visible objects visible to a human), and a device camera (the camera's field of view is adjusted to include the reflection locations of one or more light reflections and is operationally connected to the device processor). The device camera is positioned so that its field of view includes the reflection locations of one or more light reflections and is operationally connected to the device processor. This method includes: obtaining one or more predetermined context templates by the device processor; obtaining a first camera image by the device camera when the mobile terminal is operated and a projected light beam is directed from the device light source toward a context object; calculating a specified context by the device processor based on the match between the first camera image and one or more predetermined context templates; obtaining one or more predetermined interactive page layouts associated with the specified context by the device processor; obtaining a second camera image by the terminal camera when the mobile terminal is operated so that the projected light beam is directed from the terminal light source toward a specified page position; calculating a positioned image by the device processor based on the match between the second camera image and one or more predetermined interactive page layouts; identifying a specified page position based on the reflection position in the positioned image by the device processor; and performing actions by the device processor and / or remote connection processor at least partially based on the specified page position in one or more predetermined interactive page layouts.

[0042] Another example provides a method for performing an action based on a specified page location selected by a human using a mobile device. The mobile device comprises a terminal processor, a terminal light source configured to generate a projection light beam that produces one or more light beam reflections from one or more visible objects visible to a human, and a device camera positioned such that its camera field of view includes the location of one or more light reflections and operationally connected to the device processor. The method includes: the device processor acquiring one or more predetermined interactive page layouts; the mobile device being operated and the device camera acquiring two or more camera images with the projection light beam directed from the device light source at the specified page location; the device processor determining that the amount of image displacement measured in the two or more camera images is less than a predetermined displacement threshold and that the two or more camera images were acquired for an acquisition time exceeding a predetermined dwell time threshold; the device processor calculating a positioning image based on the match between the camera images and one or more predetermined interactive page layouts; the device processor identifying the specified page location based on the reflection locations in the positioning image; and either or both of the device processor and a remotely connected processor performing an action at least partially based on the specified page location in one or more predetermined interactive page layouts.

[0043] Further, according to another example, a method is provided for performing an action based on a specified page position selected by a human using a portable device, the portable device comprising a device processor, a device light source configured to generate a light beam that produces one or more light reflections from one or more visible objects visible to a human, a device camera positioned to include the reflection positions of one or more light reflections in the camera's field of view and operationally connected to the device processor, and a device switch operationally connected to the device processor. The method comprises: the device processor obtaining one or more predetermined interactive page layouts; the device processor determining that the device switch is in a first state; the device processor determining that the device switch is in a second state when the portable device has been operated and the light beam emitted from the device light source is pointing to a specified page position; the device camera obtaining a camera image; the device processor calculating a positioned image based on the matching of the camera image with one or more predetermined interactive page layouts; the device processor identifying a specified page position based on the reflection positions in the positioned image; and the device processor and / or remotely connected processor performing an action at least partially based on the specified page position in one or more predetermined interactive page layouts.

[0044] Another example provides a method for performing an action based on a specified page location selected by a human using a mobile device. The mobile device comprises a terminal processor, a terminal light source configured to generate a light beam that produces one or more light reflections from one or more visible objects visible to a human, a device camera (positioned so that the camera field of view includes the reflection locations of one or more light reflections and is operationally connected to the device processor), and a device switch operationally connected to the device processor. This method includes: obtaining one or more predetermined interactive page layouts by the device processor; determining by the device processor that the device switch is in a first state; obtaining one or more camera images by the device camera when the mobile terminal is operated and a light beam is directed from a light source towards a specified page position; determining by the device processor that the device switch is in a second state; separating the most recent stable camera image from one or more camera images by the device processor; calculating a positioned image by the device processor based on the match between the most recent stable camera image and one or more predetermined interactive page layouts; identifying a specified page position by the device processor based on a predetermined beam-direction position in the positioned image; and performing actions by the device processor and / or remote connection processor at least partially based on the specified page position in one or more predetermined interactive page layouts.

[0045] A portable terminal is provided according to the example. The portable terminal includes: a terminal body configured to be operated by the terminal user; electronic circuitry located within the terminal body, including a terminal processor; a terminal light beam source configured to emit light from the terminal body to produce a beam image reflection on a visible surface, which is focused at a predetermined distance from the visible surface; and a device camera operably connected to the device processor and aligned such that the field of view of the device camera includes the beam image reflection, wherein the portable device is configured to: produce a beam image reflection on a visible surface that is focused at a predetermined distance, prompting the device user to position the portable device at a predetermined distance from the visible surface; and acquire a camera image.

[0046] Another example provides a method for prompting a device user to position a handheld device at a predetermined distance from the visible surface. The handheld device comprises a device body configured to be operated in the hand of the device user; electronic circuitry located within the device body, including a device processor; a device light beam source configured to produce a beam image reflection that is focused at a predetermined distance; and a device camera operationally connected to the device processor and aligned so that the device camera's field of view includes the beam image reflection. The method includes: the device light beam source generating a beam image reflection that is focused at a predetermined distance on a visible surface, prompting the device user to position a handheld device at a predetermined distance from the visible surface; and the device camera acquiring a camera image.

[0047] Another example provides a method for prompting a device user to position a handheld device at a predetermined distance from a visible surface. Here, the handheld device includes a device body configured to be operated by the device user's hand. The electronic circuitry within the device body includes a device processor. A device light beam source is configured to generate a beam image reflection on a visible surface. The method includes one or more movable reflective surfaces positioned within the device light beam path emitted from the device light beam source and each operationally connected to the device processor; and a device camera operationally connected to the device processor and aligned so that the device camera's field of view includes the beam image reflection. The method includes: the device light beam source generates a beam image reflection on a visible surface; the device processor moves one or more movable reflective surfaces to focus the beam image reflection at a predetermined distance, prompting the device user to position a handheld device at a predetermined distance from a visible surface; and the device camera captures a camera image.

[0048] Other aspects and features, including the necessity and use of the present invention, will become clear from the following description in conjunction with the accompanying drawings. [Brief explanation of the drawing]

[0049] For a more complete understanding, it is advisable to refer to the following drawings and examine the detailed explanation. In the drawings, the same reference numerals indicate the same components or operations throughout the drawings. The embodiments shown in the attached drawings are as follows: [Figure 1] Figure 1 shows an example of a child operating a mobile device to identify and select a dog in a book scene (including multiple additional characters and objects) based on audio cues. Specifically, the child directs a beam of light emitted from the mobile device towards the shape of a dog. [Figure 2] Figure 2 shows an example of an operation in which a child operates a mobile device and identifies a cat based on a visual cue by shining a beam of light towards the shape of a cat in an image displayed on a tablet device. [Figure 3]Figure 3 is a flowchart illustrating an exemplary procedure for detecting a cat (e.g., a stuffed animal, an image on a screen, or a real cat) by displaying the word "CAT" on a mobile device and shining a light beam towards it. [Figure 4] Figure 4 is a flowchart similar to Figure 3, illustrating an exemplary procedure using a portable device to locate a hammer in a set of construction tools by shining a light beam onto the hammer based on audio cues. [Figure 5] Figure 5 shows an electronic circuit diagram and ray diagram illustrating typical components for generating an object hidden (camouflaged) inside a book (e.g., a candy cane hidden in the shape of a snake) using a light ray emitted from a mobile device, and for detecting it using a camera and light path. [Figure 6] Figure 6 is an exploded view of an exemplary portable device, showing the position, irradiation direction, and relative size of the light beam source and camera. [Figure 7] Figure 7 shows an exemplary interconnection layout of components within a portable device, illustrating the main directions of information flow to the electronic bus structure that forms the backbone of the electronic circuit (some components may not be used in some applications). [Figure 8] Figure 8 is a flowchart illustrating exemplary steps following the display of a text-based cue (i.e., the word "CAT"), where beam dwell time measured using an IMU is used to indicate selection based on camera images acquired before significant movement of the portable device. [Figure 9] Figure 9 is a flowchart illustrating exemplary steps following an audio prompt to select clothing, using the orientation of the mobile device when pointing to a light beam to show the selection of a long-sleeved shirt. [Figure 10] Figure 10 shows an example of a child operating a mobile device and shining a beam of light emitted from the device towards the shape of a dog to identify and select a dog from a picture book scene that includes multiple additional characters and objects. [Figure 11]Figure 11 shows the superposition of camera-acquired images matched with the page layout using computer vision in the illustrative scenario shown in Figure 10. In this case, the direction of the light beam (indicated by the crosshair target) can be identified within the camera's field of view. [Figure 12] Figure 12 shows exemplary parameters that can define the camera's field of view within a page layout, including horizontal and vertical reference positions (e.g., corner positions (x,y)), orientation (q), and magnification (m) of the image acquired by the camera. [Figure 13] Figure 13 shows an exemplary interconnection layout of components within a mobile device, illustrating the main directions of information flow to the bus structure that forms the backbone of the electronic circuit (some components may not be used in some applications). [Figure 13] Figure 14 shows an electronic circuit diagram and ray diagram illustrating the process by which a camera generates, traces, and detects a selected feline object on a book page using a light beam emitted from a mobile device. [Figure 15] Figure 15 is an exploded view of an exemplary portable device, showing the position, direction of illumination, and relative size of the light beam source and camera. [Figure 16] Figure 16 is a flowchart illustrating an exemplary procedure in which an image of a shoe, targeted using a light beam generated by a mobile device, is localized within an interactive page layout, and this provides audio feedback to the device user. [Figure 17] Figure 17 is a flowchart illustrating an exemplary two-step process in which a context (in this case, the word "CAT" in the animal list) is first selected with a light beam, and subsequent mobile device operations are limited to identifying the beam position within the page related to the selected context. [Figure 18] Figure 18 is an illustrative flowchart in which a push button switches a light beam on and off, notifies the user of their selection when choosing from a clothing collection, and generates audio feedback related to shoes. [Figure 19]Figure 19 is an illustrative diagram showing the formation of a structured light pattern, which generates a reflected beam image (i.e., the formation of a smiley face). This image is perceived as in focus by the terminal user at the target working distance between the mobile device and the reflective surface. Otherwise, the image is blurred (i.e., out of focus). [Figure 20] Figure 20 is a flowchart illustrating exemplary steps for mechanically determining whether a device-directed beam is in focus and for performing one or more actions based on whether the beam is in focus (i.e., performing actions based on whether the portable device is at a desired distance from the visible surface). [Modes for carrying out the invention]

[0050] Detailed explanation Before describing the examples, it should be understood that the present invention can naturally take various forms and is not limited to the specific examples described herein. It should also be understood that the terms used herein are for describing specific examples and are not intended to be limiting, for the scope of the present invention is limited only by the appended claims.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. Note that the singular forms “a,” “an,” and “the” used herein and in the appended claims include the plural form unless the context clearly indicates a plural. Therefore, for example, a reference to “compound” includes multiple compounds, and a reference to “polymer” includes one or more polymers and their equivalents that are well known to those skilled in the art.

[0052] Where a range of values ​​is provided, unless the context clearly indicates otherwise, each intermediate value (to the nearest tenth of a unit) between the upper and lower limits of that range is also expressly disclosed. Each subrange between any value or intermediate value within a stated range and any other stated value or intermediate value within that range is also included in the present invention. The upper and lower limits of these smaller ranges may be independently included in or excluded from the range, and ranges that include either one, both, or both limits are also included in the present invention unless there are limits that are specifically excluded within the stated range. Where a stated range includes either one or both limits, ranges that exclude either one or both of those included limits are also included in the present invention.

[0053] In this specification, the term "approximately" may be used to indicate a range of a number. The term "approximately" is used to provide literal support for the exact number that precedes it, as well as for a number that approximates or roughly matches the preceding number. In determining whether a number approximates or roughly matches a specifically stated number, an unstated number that approximates or roughly matches may, in the context in which it is presented, provide a substantially equivalent to the specifically stated number.

[0054] According to one embodiment, a device, system, and method are provided for an individual to select (i.e., point to, identify, and / or indicate) a real or virtual object from among several objects visible in the individual's environment, using a light beam generated by a portable electronic device, in response to one or more auditory, tactile, and / or visual prompts and / or cues generated by the portable electronic device. The prompts or cues may include aspects that the particular individual may know (e.g., proper nouns of family members or colleagues) or aspects that may be more generally associated with the identification, characteristics, functions, related objects, and other attributes of the discoverable object.

[0055] In addition to physical objects in the user's environment, discoverable objects may appear on paper (e.g., books, pamphlets, newspapers, handwritten notes, magazines), one or more other surfaces (e.g., book covers, boxes, signs, posters), or on electronic displays (e.g., tablets, televisions, mobile phones, and other screen-based devices). Such discoverable objects (and their display media) can be made "interactive" using mobile devices, especially when children are reading and / or being read to. For example, reading a book can be made more engaging by adding questions (for parents or guardians and children), additional relevant information, audio, sound effects, audiovisual presentations of relevant objects, and real-time feedback after discovery.

[0056] When generating one or more attributes and / or cues transmitted by a mobile device, various individual considerations may be taken into account, such as the age suitability and / or educational level of the device user, personal hobbies and / or interests, the educational and / or entertainment value of discoverable objects, and whether the object is expected to be the next discovery in a sequential order (e.g., alphabetically or numerically). These considerations may be taken into account by the mobile device and / or one or more remote processors that formulate prompts and / or cues transmitted to the mobile device.

[0057] In a further aspect, mobile devices and / or remote processors can simultaneously provide continuous, real-time assessments of engagement, language ability, reading comprehension, and / or comprehension. Assessment metrics may include, for example, the measured time a child spends interacting, the success rate of object discovery based on attributes and cues (particularly within different topical areas and / or areas of potential interest, such as sports, science, and art), the time required for pointing-based discovery (often related to attention and interest), and the overall progress rate in "discovering" objects within continuous content such as books and magazines.

[0058] These assessments can be compared to past interactions with the same individual (e.g., to determine progress in a specific topic area), interactions with others using the same or similar cues (e.g., same age, cultural background, education level), and / or cognitive processing performance across different groups (e.g., comparison of geographical, economic, and / or social clusters). Milestone responses demonstrating various aspects of cognitive processing (e.g., color discrimination, phoneme and word distinction, understanding the number of objects, performing simple arithmetic operations, gesture responses requiring controlled motor skills) can be particularly useful in monitoring a child's development, assessing their ability to present more challenging cues (and discoverable objects), and enhancing engagement. Auditory, tactile, and visual acuity can also be continuously monitored.

[0059] Specific examples of auditory attributes and cues include: sounds and sound effects that the discoverable object typically makes; sounds related to a description of an activity using the object; a clear name of the object (including proper nouns); a part or single letter of a name related to the object (e.g., beginning with a specific letter); a description of the object and / or its function; a linguistic description of one or more object attributes; a question about the function and / or attributes of the object; musical scores related to the object; quotes or proverbs that are complete with the object; and oral quizzes in which the discoverable object is the answer.

[0060] Visual attributes or cues may include: the display of the name of a discoverable object; an image or drawing of a discoverable object or an object with a similar appearance; an image of an object within a class or category of objects; an outline of an object; a portrait of an object; the name of an object or a category of objects; a word or part of a letter related to the spelling of an object (e.g., initial); one or more colors of an object (e.g., a display of an actual color swatch or a word describing a color); the size of an object (e.g., its size, especially in relation to other visible objects); sentences or questions about an object; mathematical problems in which an object (a number or symbol) is the answer; objects missing from a known sequence of objects (e.g., letters of the alphabet); fill-in-the-blank phrases in which an object is the missing element; and questions or instructions related to prompts taught in the acronym CROWD in interactive reading (completion prompts, recall prompts, open-ended prompts, WH prompts [where, when, why, what, who], distance prompts). Cues, prompts, and questions can be generated according to Vygotsky's principle of zone of proximal development to optimize user learning.

[0061] Tactile attributes or cues may include vibrations that are synchronized with (or at least at a similar frequency to) the actions or sounds that the detectable object would normally produce. For example, an image of a heart could be associated with a pulse vibration occurring approximately once per second. Furthermore, any combination of visual, auditory, and / or tactile cues may be generated by the mobile device or provided simultaneously or as a tightly timed sequence (if generated, for example, by a remote processor).

[0062] In further examples herein, a mobile device acquires a camera-based image containing a visible object pointed to by an individual. The visible object is located in the image based on the presence of reflections from the visible object generated by a light beam. Alternatively, the location of a selected object in the camera image can be determined by knowing the projection position of the light beam when the light source and camera are facing the same direction. In other words, even when the light beam is off, the position of the light beam in the camera image can be mathematically calculated based on a geometry similar to parallax (i.e., knowing the positions of both the light beam source and the camera).

[0063] Subsequently, the selected objects (and optionally other visible objects in the area of ​​the camera image) can be separated using computer vision techniques (e.g., template matching, neural network classification). A given template for a detectable object may include visual characteristics such as shape, size, profile, pattern, color, and / or texture. Detection may be declared if one or more classifications of the selected object (and optionally other visible objects in the field of view of the camera image) match one or more templates or classifications of the detectable object.

[0064] A database containing templates for discoverable objects may include one or more prompts and / or queues (as described above) associated with each object, one or more alternative (e.g., backup) prompts or cues for when an object is not found quickly, one or more actions to be taken if discovery is not achieved within a given period, and / or one or more actions to be taken when a discoverable object is successfully identified (e.g., by a mobile device or remote processor). These will be described in detail below.

[0065] In further examples herein, prompts and / or cues for initiating and / or guiding interaction with a mobile device toward discovery may include scripted (i.e., pre-established) sequences, including a combination of visual displays, audio sequences, and / or haptic vibrations. Such sequences may additionally include conditional dependencies (i.e., selection from two or more interaction scenarios) determined based on real-time conditions (e.g., past discovery success, time of day, user's age).

[0066] Typical scripted prompts (presented visually, audibly, and / or haptically) on mobile devices include the following, using the symbol "[prompt]" to represent one or more prompts, characteristics, or clues related to discoverable objects: Can you find the [prompt]? Please show me the largest [prompt]. What character comes after [prompt]? What animal makes this sound [prompt]? What instrument plays the rhythm of the [prompt] you're feeling?

[0067] Furthermore, the processor in a mobile device may include a “personality” driven by AI (Artificial Intelligence Personality, AIP), transformer models, and / or large-scale language models (e.g., ChatGPT, Cohere, GooseAI). The AIP enhances the interaction between the user and the mobile device by incorporating a friendly appearance, conversational sequences, physical form, and / or voice, which may include personal insights about the device user (e.g., likes, dislikes, preferences).

[0068] Human-machine interactions enhanced by AIP are described in detail in U.S. Patent No. 10,915,814 filed June 15, 2020, and U.S. Patent No. 10,963,816 filed October 23, 2020, all of which are expressly incorporated herein by reference. A method for determining context from audiovisual content and for a virtual agent to subsequently generate conversation based on that context is described in detail in U.S. Patent No. 11,366,997 filed April 17, 2021, all of which are expressly incorporated herein by reference.

[0069] Further aspects of the system and methodology allow for network training and other programming of handheld device CV schemes to leverage global computing resources (e.g., TensorFlow). The limited nature of discoverable objects within a template database can significantly simplify both the training and classification processes for determining the presence or absence of a match (i.e., binary determination). Training of the classification network may be limited to a predefined template database, or even to a subset of the database if the context during the discovery process is known.

[0070] Similarly, it is possible to identify matches using relatively simple classification networks or decision trees within the mobile device. Classification may be performed on devices with limited computing resources, or without sending data to remote devices (avoiding access to larger computing resources). Such classifications can be performed using neural networks (or other computer vision techniques) with hardware commonly found in mobile devices. For example, MobileNet and EfficientNet Lite are sufficient platforms for determining the match (or mismatch) between the contents of a camera image and detectable objects.

[0071] Furthermore, compared to general-purpose CV classification schemes, such restricted classification can be significantly more robust because: 1) images are compared only against a database of detectable objects (e.g., not just any object in the world); and 2) binary determination allows for adjustment of the match threshold, which helps to accurately measure the intent of the device user. The match threshold for objects pointed to in camera images may be adjusted based on factors such as hardware (e.g., camera resolution), environment (e.g., lighting, object size), specific application (e.g., inspection-based questions, presence of multiple objects with similar appearances), user category (e.g., experienced vs. novice), or specific user (e.g., young vs. older, considering past detection success rates).

[0072] The presence or absence of a match may be indicated to the user of the mobile device using visual (e.g., using one or more device displays), auditory (e.g., using device speakers), and / or vibratory (e.g., using device haptic units) feedback. In addition, or instead, the mobile device may perform other actions based on the match (or non-match) of detectable objects, such as sending one or more cues, selection occurrence, timing, and identification information (optionally including camera images) of detected (or undetected) objects to one or more remote processors.

[0073] Especially when used in entertainment, educational, and / or collaborative work environments, the ability to transmit discoveries of detectable objects makes mobile devices part of a larger system. For example, when used by children, they can share, record, evaluate, and / or simply enjoy their experiences (success / failure in discovering objects) with parents, relatives, friends, and / or guardians.

[0074] Shared experiences (e.g., with parents) include object discovery within books and magazines. A method for controlling the delivery of such serialized content on mobile devices is described in detail in U.S. Joint Application No. 18 / 091,274, filed December 29, 2022 (expressly incorporated herein by reference). A technology for sharing the process of moving to new pages or panels to search for discoverable objects when browsing books and other media is described in detail in U.S. Patent No. 11,652,654, filed November 22, 2021, all of which are expressly incorporated herein by reference.

[0075] When used independently, mobile device-based interactions eliminate the need for computer screens, computer mice, trackballs, styluses, tablets, mobile devices, and other accessories or devices when selecting objects or performing activities, thereby eliminating the need for users to understand interactions involving such devices or pointing mechanisms.

[0076] Whether used independently or as part of a larger system, a mobile device familiar to an individual (e.g., a child) can be a particularly compelling element of auditory, tactile, and / or visual rewards as a result of discovering discoverable objects (or conversely, notifying the user that no matching objects were found). The mobile device may also be colored and / or decorated to become the child's own personal possession. Similarly, audio cues (voices, one or more languages, alert sounds, overall volume) and / or visual cues (letters, symbols, one or more languages, visual object sizes) may be pre-selected to suit the individual device user's preferences, considerations, skills, and / or abilities.

[0077] In accordance with further aspects of this system and method, the light beam emitted from the mobile device is generated using one or more laser diodes, such as those manufactured by OSRAM or ROHM Semiconductor. The laser diodes (and general lasers) generate a coherent, parallel, and monochromatic light source.

[0078] Portable devices offer excellent portability and the ability to direct the proximity light source in any direction, thus enhancing eye safety, especially when used by children in environments often considered "uncontrollable" from a safety perspective, by utilizing non-coherent light sources (e.g., non-oscillating light-emitting diodes (LEDs)). LED point sources from companies such as Jenoptik and Marktech Optoelectronics can generate non-coherent, highly parallelized (optionally) polychromatic light sources.

[0079] Optical components associated with an LED point light source control the beam divergence angle, thereby allowing for the determination of the reflective spot size (see Figure 5) at a typical working distance (e.g., 0.05 to 1.0 meters when using a mobile device to point to objects on a book page). The desired spot size may vary depending on the application environment and the user. For example, a young child might only want to point to large objects on the pages of a children's book, while older children or adults might prefer to point to smaller objects such as individual words or symbols on a page or screen (i.e., using a smaller / more focused beam).

[0080] As an additional example, broad-spectrum (at least compared to lasers) and / or multi-color light sources produced by (non-laser) LEDs may help people with color blindness in one or more regions of the visible spectrum to perceive beam reflections. Multi-color light sources may also be more consistently visible to all device users when reflected off different surfaces. For example, a pure green light source may be difficult to see when reflected off a pure red surface (such as part of a book page). Multi-color light sources in the red-green region of the visible spectrum may help mitigate this problem. High-energy photons in the deep blue region of the visible spectrum may be avoided for eye safety reasons.

[0081] As a further example, the device beam source is operationally coupled with the device processor, enabling intensity control, including turning the beam on and off. For instance, beam activation via a mobile device can be incorporated as a prompt to the user (after providing a prompt or cue) indicating that directing towards a detectable object is expected.

[0082] Subsequently, by turning off the beam while acquiring camera images, beam reflection can be avoided, and pixel saturation near beam reflection points (such as "bleed" due to camera pixel saturation) can be prevented. If there are no reflections from the object (reflections that may be considered "noise" during object identification), the requirements for both machine learning-based training and classification processes may be reduced and accuracy may improve.

[0083] The beam can also be turned off when a match is found (e.g., as part of a success notification to the user). Conversely, keeping the beam on during interaction may indicate to the device user that further searching for discoverable objects is expected.

[0084] The beam intensity can also be modulated based on measuring one or more reflections in an image acquired, for example, by a mobile device camera. The reflection intensity can be set so that it is clearly distinguishable to the user against the background (e.g., to take into account ambient lighting conditions, to accommodate the reflectivity of different surfaces, and / or to accommodate visual impairments), but not so overwhelming (e.g., based on user preference). The beam intensity can be modulated by several means known in the art. These include adjusting the magnitude of the optical beam drive current (e.g., using transistor-based circuits), and / or using pulse width modulation (PWM) of the drive circuit.

[0085] As a further aspect of this system and method, it is possible to display one or more illumination patterns on a device display device and / or project them into the device beam. By projecting one or more illumination patterns using the beam, the role of an independent display device on a handheld device is effectively integrated with beam directivity. The illumination patterns projected into the beam can be formed, for example, using micro-LED arrays, LCD filtering, or DLP (i.e., digital light processing with a micro-mirror array) techniques known in the art.

[0086] When images or symbols (e.g., characters that make up words or phrases) are too long or too complex to display all at once, messages or patterns within the beam can be "scrolled." Scrolled text or graphics are displayed in a predetermined direction (e.g., up, down, horizontal) and in sections at a time (e.g., providing a dynamic appearance). During and after the object detection process using the light beam (i.e., when attention is focused on the beam), messages embedded within the beam (e.g., the names of identified objects) may be recognized and be effective and / or meaningful.

[0087] The illumination patterns generated by the beam light source are used to enhance pointing functionality, including control over the size and / or shape of the beam visible to the device user. Within the illumination pattern, the size (e.g., related to the number of illuminated pixels) and relative position (the position of the illuminated pixels) are controlled by the mobile device. Different beam sizes can be used for different applications, such as pointing to characters in text (using a thin beam) and pointing to large cartoon characters (using a larger beam). The beam position may be "fine-tuned" by the mobile device to direct the user's attention to a specific (e.g., nearby) object within the camera's field of view.

[0088] The lighting patterns generated within a beam can also be used to "enhance," add to, or enhance printed materials and other external (i.e., mobile device) content. One or more reflective objects identified in an image captured by a mobile device's camera may be enhanced by the projection of a light beam. For example, if you are looking for a bird as a discoverable object and point the beam at a squirrel, the beam will project a pair of wings onto a printed image of the squirrel, serving as components of the (funny) question of whether the pointed object is a bird. Another example is changing the apparent color of certain components in a printed material by illuminating the entire object (or individual components) with a selected color within the beam. Such enhancement of visual content (along with additional audio content) can help "bring to life" the static content of books and other printed materials.

[0089] Furthermore, visible information and / or symbols within a beam projection with optical performance equivalent to that of the device camera (e.g., common depth of field, no distortion when viewed perpendicular to a reflective surface) tend to prompt (or psychologically guide) mobile device users to adjust the orientation and position of their device so that the information and / or symbols are most easily visible to both the user and the device camera (e.g., in focus and undistorted). As a result, properly positioned and oriented camera-acquired images (i.e., images easily visible to the user) can facilitate computer vision processing (e.g., improved classification reliability and accuracy). Users may not realize that their ability to easily see and identify the projected beam pattern also contributes to improved image quality in camera-based processing.

[0090] As a further aspect of this system and method, the structure of the portable device may include orienting the device camera in the same direction as the light beam. This makes it possible to identify the beam's illumination location (at least a limited area) in the camera image (e.g., even when the beam is off). Because the beam and camera move simultaneously (i.e., both are fixed to or built into the portable device), the location (or area) to which the beam is directed in the camera image can be identified regardless of the physical position, direction of orientation, or overall orientation of the portable device in (three-dimensional) space.

[0091] Ideally, the structure of the handheld device would position the beam reflection at the center of the camera image. However, if there is a slight gap between the beam source and the camera sensor due to physical structural constraints, the beam may not appear at the center of the camera image at all working distances. Given the beam direction, camera image acquisition direction, and the physical distance between them at a specific working distance, the reflection position can be calculated using simple geometry (similar to the geometry that describes parallax) (see Figure 6).

[0092] For example, the beam and camera can be adjusted to project and acquire parallel (i.e., non-converging) rays. In this case, the reflected image is offset from the center of the camera image by an amount that depends on the working distance (i.e., the distance from the mobile device to the reflective surface). The separation distance between the center of the camera image and the center of the beam decreases as the working distance increases (for example, it approaches zero at infinity). The separation distance can be kept small by keeping the physical distance between the beam and the camera small.

[0093] Alternatively, the beam and camera direction can be adjusted to converge them at a predetermined working distance. In this case, the beam can be set to appear at the center of the camera image (or a selected image position) at a predetermined working distance. As the distance from the handheld device to the reflective surface changes, the beam position fluctuates within a limited range (generally one-dimensional, related to the axis defined by the camera and light source). Again, keeping the physical separation distance between the beam and camera small helps to keep the beam-directed area in the camera image small within the working distance range used in typical applications.

[0094] The calibration process for determining the beam region and / or position within the camera image includes the following: Obtain a baseline camera image that does not include reflections from the projection light beam (e.g., from a featureless surface). Acquire a light beam reflection image (when the beam is lit) that includes one or more reflections from the projected light beam. Calculate a difference pixel intensity image by subtracting the baseline image from the light beam reflection image, and Pixels in the subtracted pixel intensity image that detect a light beam exceeding a predetermined light intensity threshold are assigned to the beam position region.

[0095] The beam-direction region can be identified based on measurements of pixels exceeding a threshold. Alternatively, a single beam-direction position can be determined from the center position (e.g., two-dimensional median or mean) of all pixels exceeding the threshold intensity. Calibration can be performed at different working distances to map the entire range of the beam-direction region.

[0096] In additional examples described herein, identifying an object pointed to in a camera-based image (when the beam is on) may be based on repeatedly identifying beam reflections as high-luminosity regions (e.g., high intensity within a region of pixel location). Such beam reflection localization may also take into account the color of the light beam (i.e., identifying high luminosity only within one or more color gamuts related to the light beam spectrum). By understanding the relationship between the pointed-to location and the working distance, the distance from the handheld device to the reflective surface can be estimated when a camera-based image of the beam reflection is available.

[0097] As a further example, based on the spectral sensitivity of the typical human eye, a light beam in the green region of the visible spectrum is likely to be most easily perceived by the vast majority of individuals. Many so-called RGB (red, green, blue) cameras have twice as many green sensor elements as red or blue sensor elements. Utilizing a light beam in the mid-band of the visible light spectrum (e.g., green) allows for easier detection (e.g., by both humans and cameras) while keeping beam intensity low, improving the reliability of reflectance detection in camera-based images and enhancing overall eye safety.

[0098] As a further aspect of this system and method, the mobile device user can instruct (the mobile device) that the beam is directed at a selected object. Such instruction is given through various means of interaction, including: Pressing or releasing a push button (or other contact sensor or proximity sensor) that is a component of a mobile device. Providing voice commands (e.g., "now") that are detected by the microphone of a mobile device and identified by the device processor or remote processor. To direct a beam towards an object (i.e., without significant movement) and maintain it for a predetermined "dwell" time (e.g., based on user preference), The orientation of the handheld device to a predetermined direction (e.g., perpendicular to Earth's gravity) detected by the IMU of the handheld device, or Gestures or taps on the handheld device (e.g., tilting the device forward) are also detected by the handheld device's IMU.

[0099] In the latter exemplary case, since user actions on the mobile device (e.g., gestures or taps) may affect the image stability within the camera's field of view, the indicated visible object can be identified using a still image (e.g., one that does not show substantial movement compared to one or more previously acquired images) that has been separated before the action (e.g., from a series of sequentially sampled images).

[0100] As yet another example, one way to implement a dwell time-based method involves obtaining a number of consecutive images calculated by dividing a given dwell time by the frame rate, thereby revealing substantially stationary visible objects and / or beam reflections. Due to the rapid and / or precise dwell time requirements, this method requires a high frame rate for the camera and the associated computational load and / or power consumption. As an alternative method for determining whether sufficient dwell time has elapsed, an IMU can be used to evaluate whether a handheld device has remained substantially stationary for a given period. Generally, IMU data for performance evaluation can be acquired with higher temporal resolution compared to processes involving full-frame camera image acquisition.

[0101] To convert analog IMU data into a digital format suitable for processing, analog-to-digital (A / D) conversion techniques well-known in this field can be used. IMU sampling rates generally range from approximately 10 samples / second to approximately 10,000 samples / second. Here, (as explained in the background section above) increasing the IMU sampling rate results in trade-offs regarding signal noise, cost, power consumption, and / or circuit complexity. IMU data streams may include one or more of the following: Accelerometer data of up to 3 channels (i.e., representing 3 orthogonal spatial dimensions), Gyroscope rotation speeds of up to 3 channels (i.e., representing rotation around 3 axes, often referred to as pitch, roll, and yaw), Magnetometer data (representing magnetic forces, including Earth's magnetic attraction) for up to 3 channels (i.e., attitude in 3 orthogonal dimensions), and Inertial force data for internal mass (which may include Earth's gravity) across up to three channels (i.e., representing attitude in three orthogonal dimensions).

[0102] Data from a triaxial accelerometer can be considered as time-varying vectors in three-dimensional space, where each axis can be denoted as X, Y, and Z. When accelerometer data is treated as a vector, the absolute value of acceleration |A| can be calculated according to equation (1). (Formula 1) TIFF2026510211000002.tif11123 Here, X i , Y i , Z i The sample values ​​from the accelerometer in each 3D (where "i" represents the sample index) are X b , Y b , Z b These represent the so-called "baseline" values ​​in the same three dimensions.

[0103] Reference values ​​can take into account factors such as electronic offset and may be determined during a motion-free "calibration" period (e.g., by calculating average values ​​or reducing the effects of noise). Three-dimensional acceleration directions (e.g., using spherical, Cartesian, and / or polar coordinate systems) may also be calculated from such data streams. A similar approach can also be implemented by calculating the orientation vector relative to Earth's gravity and / or magnetic attraction based on a multidimensional IMU gyroscope data stream.

[0104] User interactions on mobile devices include translational motion, rotation, tapping the device, and / or orientation of the device. User intent is indicated, for example, by: All kinds of operations (e.g., absolute values ​​exceeding the IMU noise level), Movement in a specific direction, Speed ​​exceeding a threshold (e.g., in any direction), Gestures using mobile devices (known movement patterns, etc.), Orientation of the mobile device in a predetermined direction, Lightly tap the mobile device with the fingers of your other hand. Smashing a mobile device with an object (e.g., a stylus), Slamming a mobile device against a solid object (e.g., a desk), and / or The act of slamming one mobile device against another.

[0105] As a further example of user intent determination based on IMU data streams, a "tap" on a handheld device may be identified as the result of a user-intentionally moved object (i.e., "object") making contact with a target location on the handheld device's surface (i.e., "tap location"). The tap location calculated on the handheld device can be used to convey additional information about the device user's intent (i.e., in addition to object selection). For example, the user's confidence in a selection, indicating the first or last selection among a group of objects, or a desire to "postpone" during a sequence of interactions, can each be signaled based on directional movement and / or tap location on the handheld device.

[0106] A tap is determined when a stationary mobile device collides with a moving object (e.g., fingers of the hand opposite to the one holding the device), when the mobile device itself moves and collides with another object (e.g., a table), or when the colliding object and the mobile device move simultaneously before contact. The IMU data stream before and after the tap helps determine whether an object was used to strike the stationary device, whether the device was forcibly moved to another object, or whether both processes occurred simultaneously.

[0107] Tap locations can be identified using characteristic "signatures" or waveform patterns (e.g., peak force, acceleration direction) within the IMU data stream (particularly accelerometer and gyroscope data), which depend on the tap location. Identification of tap locations on the surface of a mobile device based on inertial (i.e., IMU) measurements, and subsequent activity control based on the tap location, are described in detail in U.S. Patent No. 11,614,781 filed July 26, 2022, all of which are expressly incorporated herein by reference.

[0108] As a further aspect of the apparatus and methods described herein, the mobile terminal processor may perform one or more actions based on determining a match (or mismatch) with the characteristics of a detectable object. If a mismatch of one or more objects in the camera image is determined, the mobile terminal may simply acquire one or more subsequent camera images and continue to monitor whether a match has been detected. Alternatively, it may provide additional prompts and / or cues. Such prompts and / or cues may include repetition of previous prompts and / or cues, or presentation of new prompts and / or cues (e.g., those obtained from a template database) to expedite and / or enhance the discovery process.

[0109] If a match is detected, actions performed by the processor in the mobile device include sending available information related to the discovery process to the remote device. For example, the remote device may perform further actions. This information may include auditory, tactile, and / or visual cues used to initiate the discovery process. Furthermore, the transmitted dataset may include camera images, the acquisition time of the acquired camera images, a predetermined camera image light beam directional region, a predetermined template of the discoverable object, and one or more indicator objects.

[0110] Alternatively, or in addition, user-involved actions may be performed within the handheld device itself. For example, one or more sounds may be played through the device's speaker. These may include a celebratory phrase or sentence, the name of the discovered object, a sound emitted by the discovered object, a celebratory chime, or a sound emitted by another object associated with the discoverable object.

[0111] Lighting patterns (e.g., displays on a screen or projections within a beam) may include congratulatory phrases or sentences, display names of the discovered object, portraits or diagrams of the discovered object, and images of other objects related to the discovered object (e.g., letters of the alphabet, the next object in a series of objects). Haptic feedback upon discovery may include simply recognizing success with a vibration pattern, or generating vibrations at frequencies related to the movement or sound emitted by the discovered object.

[0112] The mobile device may further include one or more photodiodes, an optical blood sensor, and / or an electrical heart rate sensor. Each of these is operationally connected to the device processor. These mobile device components may provide additional elements (i.e., inputs) that help monitor and determine user interactions. For example, a data stream from a heart rate monitor may indicate stress or distress during the discovery process. Based on detected levels, past interactions, and / or predefined user settings, object discovery may be limited, delayed, or aborted.

[0113] As an additional example, although not strictly "portable," such portable electronic devices may be attached to and / or operated from other parts of the human body. For example, a user-interacting device that directs a beam of light towards detectable objects may be attached to the arm, leg, foot, or head. Such arrangements are used to address accessibility issues in individuals with limited upper limb and / or hand movement, individuals lacking sufficient manual dexterity to communicate intentions, individuals with missing hands, and / or situations where hands are required for other activities.

[0114] Accessibility-related factors are also considered when using handheld devices. For example, when an individual with color blindness uses a device, it is possible to avoid certain colors or color patterns in visual cues. Individuals with visual impairments can be accommodated by adjusting the size and intensity of cues displayed on one or more handheld device screens and / or within the beam. Media containing detectable objects can be braille-enhanced (e.g., including both braille and images) or include patterns and textures with raised edges. Pointing using a light beam can be complemented by enhancing beam intensity or by having the handheld device's camera track finger pointing (e.g., within areas containing braille).

[0115] Similarly, if an individual has hearing impairment in one or more speech frequency bands, it is possible to avoid or amplify those frequencies in the voice guidance generated by the mobile device (e.g., depending on the type of hearing impairment). Tactile interactions can also be adjusted to take into account the individual's heightened or suppressed tactile sensitivity.

[0116] For example, activities involving young children or individuals with cognitive impairments may require considerable "guessing" and guidance from the device user. User assistance during operation and relaxation of expected response precision can be considered a form of "interpretive control." Interpretive control includes "guiding" to one or more target responses or reactions (e.g., providing intermediate hints). For example, a young child may not fully understand how to operate a mobile device to effectively shine a beam of light. During such interactions, voice instructions accompany the interactive process (e.g., broadcasting "Raise the wand straight up") and guide the individual to make a choice.

[0117] Similarly, when the user approaches a discoverable object or a predictable response, flashing indicators or repeating sounds (the frequency of which may be related to how close the cue or attribute is to a particular selection) may be emitted. On the other hand, responses that do not involve intentionally pointing the beam may be accompanied by “questioning” indicators (e.g., haptic feedback and / or buzzing sounds), which serve as a prompt to consider alternatives. Further aspects of interpretation control are described in detail in U.S. Patent No. 11,334,178 filed August 6, 2021, and U.S. Patent No. 11,409,359 filed November 19, 2021, all of which are expressly incorporated herein by reference.

[0118] Figure 1 shows an exemplary scenario in which a child 11 finds a visible object based on a barking sound 16b played from the speaker 16a of the mobile device 15. The mobile device 15 can additionally play instructions 16b, such as "Let's find the dog!", through its speaker 16a. As a further example of a dog signal, the mobile device 15 can project the word "dog" or other descriptive phrases or symbols commonly associated with dogs onto one or more of its terminal displays 17a, 17b, 17c, or by scrolling them using the terminal beam.

[0119] Furthermore, an image of the dog 12b in the comic scene 12a (e.g., one previously taken with the mobile device camera) and / or one or more drawings or other representations of a dog can be projected onto one or more displays 17a, 17b, 17c. The barking may be accompanied by vibrations generated using a tactile unit (not shown) built into the mobile device 15 (i.e., perceived audibly and / or tactilely by the holding hand 14 of the device user 11), thereby informing the user 11 that the discovery of a visible object is expected.

[0120] Child 11 can use their right hand 14 to manipulate a ray of light 10a emitted from a mobile device 15 and direct it towards a picture of a dog 12b in a comic strip scene 12a, which includes multiple additional displays placed across two pages of a magazine 13a and 13b. The ray 10a generates a reflection of light 10b at the position 12b of the dog (i.e., the selected object) on the rightmost page 13b. Images acquired by a camera (not shown in Figure 1, pointed at the page in the same direction as the ray 10a) may classify the object 12b at the ray position 10b as a dog (regardless of whether the ray reflection is present in the camera image). If the child correctly points to the dog 12b, they are provided with an acoustic, tactile, and / or visual reward.

[0121] If the classification of the pointed-out object is determined not to belong to a dog breed, the entire process may be repeated by rebroadcasting acoustic, tactile, or visual cues related to discoverable objects, broadcasting additional or alternative cues, or broadcasting acoustic, tactile, or visual cues indicating that the provided cues are not related to the pointed-out object. Furthermore, the selection, timing, and identity (or lack thereof) of the discovered object 12b may control actions subsequently performed directly by the handheld device 15 and / or actions transmitted to one or more remote devices (not shown), and may coordinate further activities.

[0122] Figure 2 shows another illustrative scenario in which three spherical displays 27a, 27b, and 27c on a mobile device spell out the word "CAT". Additional cues related to the discoverable object include playing a sound and / or pronouncing the word "cat" or other related terms using the mobile device's speaker 26 (i.e., distinct from the sound 23b emitted from tablet 23a).

[0123] The child at position 21 can use their right hand at position 24a to direct the ray of light at position 20a, projected from the mobile device at position 25, towards the cartoon cat at position 22a. The cat at position 22a and the unicorn at position 22b are components of a presentation combining audio 23b and video 23a. The ray of light 20a may reflect off the tablet screen 20b in the area of ​​the cat 22a during the audiovisual sequence.

[0124] An instruction indicating that the target being illuminated by 20a has been selected is communicated by the user 21 by one or more of the following: 1) uttering an instruction (e.g., "Now" or "OK") that is detected by a microphone (not shown) embedded in the mobile terminal 25; 2) pressing one of several push buttons (e.g., 25a) located on devices 25a and 25b with the thumb 24b (or another finger); 3) making a gesture 29a, 29b with a predetermined motion pattern detected by an IMU (not shown) incorporated in the device 25, and / or 4) pointing the handheld device 25 in a predetermined direction (e.g., relative to Earth's gravity) to indicate that an image processing step should be initiated.

[0125] By analyzing images acquired by a mobile device camera (not shown in Figure 2, pointing in the same direction as the light beam 20a), it is possible to classify the object 22a at the beam reflection position 20b as a feline, and subsequently generate auditory, tactile, and / or visual feedback related to the discovery by the child 21. The images captured by the camera may or may not have the beam reflection section 20b present. For example, the beam can be turned off during the period of image acquisition by the camera (e.g., to ensure object identification without interference from beam reflection). The selection, timing, and identification of the discovered cat 22a (including predefined behaviors in the discoverable object database) may then be used to control actions on the mobile device 25, tablet 23a, and / or other remote processors (not shown).

[0126] Figure 3 illustrates the procedure for detecting a cat (e.g., a stuffed animal, a cat image in a book or on a screen, or a real cat) in a group of visible objects, including a unicorn (32a), using visual cues (displayed on the mobile device displays 31c, 31b, and 31a) and a light beam 32c generated by the mobile device 32d. These procedures include: In 30a, as a visual cue related to a discoverable object, the three display devices that are components of the mobile terminal 31d display "C" on 31c, "A" on 31b, and "T" on 31a; In 30b, the mobile device 32d is moved and / or turned, and the light beam 32c is manually directed towards the visible cat 32b; In 30c, an image including the cat 33a within the beam reflection region 33b (the beam itself is not shown in 30c for clarity) is acquired using a camera (not shown) and a focusing optical system 33c (shown separately from the main body of the portable device for illustrative purposes) located within the portable device 33d; In 30d, a neural network 34, template matching, and / or other classification schemes are used to determine whether one or more objects in the camera image 33a match one or more objects in the detectable object database (e.g., detectable objects associated with the “CAT” visual cues 31c, 31b, and 31a in particular); In 30e, it is determined whether there is a match between the object classified in the camera image 33a and one or more predefined templates or classifications of the discoverable object 37. If there is no match, it returns to 38 and retransmits and / or generates a new discoverable object queue 30a; or In step 30f, it is determined that the object pointed to by the device user is a detectable object (e.g., the cat in step 35); and Optionally, the light beam used to point to visible objects 36a, 36b, as indicated by the dotted rectangular frame 30g, may be turned off, and / or the successful pointing of discoverable objects may be rewarded with haptic vibration and / or other acoustic / visual cues generated by the mobile terminal 36c.

[0127] Figure 4 is a flowchart illustrating an exemplary procedure in which a mobile device 42d finds a hammer 42b among a set of construction tools (including a saw 42a) using one or more acoustic cues 41c played by the mobile device 42a. In this case, in 40a, one or more acoustic cues emitted from the mobile device speaker 41b may include the sound of hammering a nail. Such rhythmic sounds may (optionally) be accompanied by vibrations generated by a haptic unit (not shown) built into the mobile device, which can be felt in the hand 42d of the device user. Alternatively, the terminal speaker 41b may also play the word "hammer" or a question such as "What is the tool used to drive or remove a nail?" in 41c.

[0128] Similar to Figure 3, the exemplary steps of the sequence shown in Figure 4 include, in 40b, manually operating the handheld device 42d to direct the light beam 42c toward the hammer 42b (i.e., taking care not to let the beam linger on other visible objects such as the saw 42a). In 40c, an image 43a containing the hammer 43b (with or without beam reflection) is acquired using a camera in the handheld terminal 43d (not shown, facing the same direction as the beam of 42c) and a focusing optical system 43c. In 40c, a neural network, template matching, and / or other classification scheme (i.e., the CV method of 44) can be used to classify the selected object 43b in the camera image 43a using a detectable object database.

[0129] In this exemplary case, the description of the detectable object and / or a specific hammer type in the template (i.e., 47, the pawn hammer) may not exactly match the hammer type indicated in the camera image (i.e., 42b, the claw hammer). The classification match in 42e is based on one or more predetermined thresholds for classification match, allowing any object 42b that is generally considered a hammer to generate a reward for finding the detectable object.

[0130] If, in step 40e, the user's intent is classified as "uncertain" or "does not match" the cues provided by the mobile device in step 40a, the process returns to step 48, where a process is performed that allows the user to hear the original cues and / or additional acoustic cues regarding discoverable objects again. The user can then continue aiming the light beam 42c at (the same or) other visible objects.

[0131] Meanwhile, in 40f, if it is determined that the user's instructions match the auditory cue provided by the mobile device, the identified object (i.e., the hammer with claws at position 45) is designated as the object to be found. In 40g, optionally (as indicated by the dashed rectangular frame), the light beam generated by the mobile device 46d to point to visible objects 46b, 46c can be turned off (e.g., to transition to another activity). It is also possible to acoustically notify 46a of the successful pointing of a findable object (e.g., "Found!") and / or other visual or tactile methods. Subsequently, the object image 43a, classified identification information, and / or the timing of object selection by the device user are transmitted to one or more remote processors (not shown) to control the operation of one or more mobile devices and / or for further processing.

[0132] Figure 5 shows the components of the electronic circuit diagram and ray diagram illustrating the optical paths of beam generation (51a, 51b, 51c, 51d), beam irradiation (50a), and reflected light (50c), as well as the operation by which camera 56a detects the state in which it points to the candy cane pattern 59 on pages 53a and 53b. The electronic circuits 51a, 51b, 51c, and 51d, and the camera and associated optical system 54 of 56a, can all be incorporated into a portable device body (not shown).

[0133] The candy cane figure 59 may be found within a sketch that includes a cartoon character 52a on the leftmost page 53a, another cartoon character 52d on the rightmost page 53b, and a cat 52c. In this exemplary case, the candy cane figure 59 is intentionally incorporated (i.e., embedded) within the snake figure 52b. Camouflaging such images can be a challenge to encourage children (and others) to look more carefully at the contents of a book. Such playful strategies can make book-related activities more enjoyable, educational, and beneficial for children (and adults).

[0134] The components of the beam generation circuit include: 1) a power supply 51a consisting of a rechargeable or replaceable battery, typically built into a portable handheld device; 2) optionally a switch 51b or other electronic control device (e.g., push button, relay, transistor) controlled by the handheld device processor and / or device user to turn the pointing beam on and off; 3) a resistor 51c (and / or a transistor with adjustable beam intensity) that limits the current to the beam source; 4) the beam source, typically consisting of a light-emitting diode (LED) or laser diode 51d, since LEDs are generally configured in the forward bias direction (i.e., low resistance direction).

[0135] A precision optical system (which may consist of multiple optical elements, including those enclosed within a diode-based light source, but not shown) makes the beam 50a nearly parallel and provides a small beam divergence angle as needed. As a result, the beam dimension emitted from the light source 58a may be smaller than that of a point at a distance along the optical path 58b. The divergence angle of the illumination beam 50a and the distance between the light source 51d and the selected object 59 primarily determine the reflected beam spot size 50b.

[0136] Furthermore, the beam reflected by the selected object 50c may continue to diffuse, and for example in Figure 5, the size of the reflected beam 58c is larger than that of the irradiated beam 58b. The size and shape of the irradiated spot 50b (and its reflection) may also be affected by the position of the beam relative to the reflective surface (e.g., the angle with respect to the normal of the reflective surface) and / or the shape of the reflective surface. Because the horizontal dimension of page 53a on the far left is convexly curved relative to the incident beam 50a, an illumination beam (e.g., Gaussian distribution and / or circular) will produce a reflected spot 50b that is essentially elliptical (i.e., has a wide horizontal dimension).

[0137] Light from the field of view of camera 56a is collected by camera optics 54, and a camera image 55a, including a focused reflected beam spot 55b, is focused onto the camera's light-sensing component 56b. Such detected images are digitized using techniques well known in the art and then processed (e.g., using CV techniques) to identify objects targeted within the camera's field of view.

[0138] Figure 6 is an exploded view of the portable device 65, showing the exemplary placement of the light beam source 61a and the camera 66a. These components may be incorporated inside the portable device 65 during final assembly. This figure of the portable device 65 also shows the back surfaces of the three spherical displays 67a, 67b, and 67c attached to the main body 65.

[0139] The light beam generator consists of laser or non-laser light-emitting diodes (61a) and may include built-in and / or external optical components (not visible in Figure 6) for forming, structuring, and / or parallelizing the light beam (60). The beam-generating electronics and optics are housed in a subassembly (61b) that provides electrical connections to the beam generator and precise control of beam aiming.

[0140] Similarly, the image acquisition process is realized by a focusing optical system 64a incorporated within a screw-in housing 64b. This housing may include additional (optional) optics in the optical path for magnification and / or optical filtering (e.g., to remove reflected light emanating from the beam). The optical elements are mounted on a camera assembly (i.e., including the image sensor surface) 66a, which is housed in a subassembly that provides precise control of the camera's electrical contacts and imaging direction.

[0141] One aspect of the exemplary configuration shown in Figure 6 includes the fact that the optical beam 60 and the camera's image acquisition optics 64a are pointed in the same direction 62. As a result, beam reflection from visible objects occurs in almost the same area within the camera image, regardless of the overall direction of the handheld device. Depending on the relative alignment and separation distance (i.e., the positional relationship between the beam source and the camera 63), the position of the beam reflection may be slightly off from the center of the image. Furthermore, due to the (design-small) separation distance 63 between the beam source 61a and the camera 66a, slight differences in beam position may occur if the distance from the mobile device to the reflective surface varies.

[0142] Such differences can be estimated using mathematical methods similar to those used to describe parallax. As a result, even when the targeting beam is off (i.e., when there are no beam reflections in the camera image), one or more target objects can be identified based on the position of the beam optics within the camera's field of view. Conversely, the measured shift to the center position (or other reference point) of the light beam reflection in the camera image can be used geometrically to estimate the distance from the mobile device (more specifically the device camera) to the object being viewed.

[0143] Figure 7 is an example of an electronic interconnection diagram of a mobile terminal 75, showing components 72a, 72b, 72c, 72d, 72e, 72f, 72g, 72h, 72i, 72j, 73, 74, and the main directions of information flow during use (i.e., indicated by the direction of the arrows relative to the electronic bus structure 70 that forms the backbone of the device circuit). All electronic components can communicate with one or more processors located at position 73 via this electronic bus 70 and / or by a direct path (not shown). In certain applications, some components may not be required or used.

[0144] The core of a portable handheld device is one or more processors (including microcomputers, microcontrollers, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.) 73, powered by one or more (usually rechargeable or replaceable) batteries 74. As shown in Figure 6, the components of the handheld device also include a beam-generating component (e.g., typically a laser oscillator or light-emitting diode) 72c, and a camera 72d for detecting objects within the beam's illumination area (which may include reflections from the beam). When integrated into the handheld device body 75, both the beam source 72c and the camera 72d may require one or more optical apertures and / or optical transmissions (71b and 71c, respectively) passing through the handheld device housing 75 or other structure.

[0145] In applications involving acoustic cues, a speaker (e.g., electromagnetic coil or piezoelectric) 72f may be used. Similarly, in applications that may involve voice-based user interaction, a microphone 72e captures ambient sounds from the mobile device. When integrated within a handheld device 75, the operation of both the speaker 72f and microphone 72e may be assisted by acoustic transmission through the handheld device housing 75 or other structures, for example by making them tightly attached to the device housing and / or by providing multiple perforations in 71d (as further shown, for example, in 16a of Figure 1).

[0146] In applications involving vibration and / or rhythmic prompts or cues, and / or to inform the user that a reward (or mismatch warning) is expected, tactile units (e.g., eccentric rotating masses or piezoelectric actuators) positioned at 72a may be used. One or more tactile units may be mechanically coupled to positions on the device housing (e.g., to be felt at specific locations on the device) or fixed to an internal support structure (e.g., designed to be felt more generally across the entire surface of the device).

[0147] Similarly, applications that include visual cues about detectable objects may utilize one or more displays located at position 72b to display, for example, letters, words, images, and / or shapes related to the detectable object, as shown in the illustration, or to provide post-detection feedback (or indicate that the object was pointed to incorrectly). Such one or more displays may be fixed to the main portable device body (as shown in 72b) and / or located externally and / or may employ optical transparency within the device housing as shown in 71a.

[0148] During typical operation, the user may send signals to the mobile device at various times, such as when ready to detect another object, when pointing to a new object, or when agreeing on a previously detected object. User signals are indicated by voice feedback (detected by the microphone as described above), motion gestures detected by the IMU (72g), and the physical orientation of the mobile device. Although Figure 72g is shown as a single device, different implementations may employ distributed subcomponents. For example, configurations that individually detect acceleration, gyroscopic motion, magnetic orientation, and gravitational force are conceivable. Furthermore, it is possible to improve the signal-to-noise ratio during detection by placing subcomponents in different areas of the device structure (e.g., distal arm, electrically quiet areas).

[0149] User signals can also be indicated using one or more switch mechanisms, including push buttons, toggle switches, contact switches, capacitive switches, and proximity switches. Such switch-based sensors may require structural components on or near the surface of the mobile device to transmit force or transmit operation to circuits located further inside (see 71e).

[0150] Communication with the mobile device 75 is implemented using Wi-Fi 72i and / or Bluetooth 72j hardware and protocols (each using different electromagnetic spectrum regions). In an exemplary scenario where both protocols are used together, Bluetooth 72j for short-range communication may be used, for example, when registering the mobile device using a mobile phone or tablet (e.g., identifying a Wi-Fi network and entering a password). Subsequently, the Wi-Fi protocol is adopted, allowing the activated mobile device to communicate directly with other more distant devices or the World Wide Web.

[0151] Figure 8 is a flowchart illustrating exemplary steps following a text-based cue (i.e., three letters representing the word "CAT" as shown in 81c, 81b, and 81a, respectively), where the dwell time (i.e., lack of significant terminal movement) pointing to the mobile terminal beam, measured using an embedded IMU, is used to indicate the user's selection of a visible object. After selection, the mobile terminal processor (or remote processor) determines whether the selected object (82b) is a feline (e.g., a match against a given detectable object template) and performs an action based on the result.

[0152] The steps in this process include: In 80a, the letters "C" (81c), "A" (81b), and "T" (81a) are displayed on the three displays (81d) that make up the mobile terminal as a visual cue related to a detectable object. In 80b, the user can use the light beam (82d) of the mobile device to point to a selected object (i.e., a cat (82b)) within the camera's field of view, including other objects such as a unicorn (82a); In 80c, one or more camera images 83a of the scene in the same direction as the beam are acquired using the camera and camera optical system 83b (for clarity, the optical system 83b is shown separately from the terminal), which are components of the mobile terminal 83c; In 80d, IMU data is acquired to monitor the movement of the handheld device (including acceleration and rotation). This is represented as one or more vectors in three-dimensional space and labeled X in 84a, Y in 84b, and Z in 84c; At 80e, it is determined when the operating amount (|A|) (including acceleration and / or rotation) of the mobile device falls below a predetermined operating threshold, indicated by a dashed horizontal line at 85c, and this state continues for a sufficient duration or del time at 85a. This determination begins when |A| falls below the threshold 85c at the time indicated by a dashed vertical line at 85d. In 80f, when IMU data is sampled, if the waiting time (i.e., the state in which the operation is less than the operating threshold of 85c) has not reached a predetermined waiting time threshold, 89a returns to acquiring additional camera images in 80c and acquiring IMU data in 80d; otherwise, if a sufficient amount of time has elapsed to exceed the dwell time threshold, 89b proceeds to processing related to the separation of the target (target); As an option, the light beam emitted from the handheld device (which was directed at visible objects such as the unicorn in 86a or the cat in 86b) is turned off in 86c, as indicated by the dashed rectangle in 80g, serving as an indicator to the user that a selection has been made; At 80h, extract the most recent image 87b from the camera scene image set 87a, in the direction of light (including the cat in 87c), that was taken before significant movement of the mobile device (i.e., during dwell time); In 80i, objects within the region pointed to by the beam (88a, i.e., in the same direction as the camera) are further separated within the latest camera image; At 80j, it is determined whether the object within the light beam's directing region matches the template and / or classification of the detectable object (in this example, the cat at 88b); In 80k, the attributes of discoverable objects (i.e., objects that the user has successfully aimed at) (e.g., cat meows) are played back on the speaker 88d of the mobile device 88c.

[0153] Figure 9 is a flowchart illustrating how a terminal user performs a light beam-based selection based on the orientation of the mobile device (i.e., vertical orientation in 95a and relative position to Earth's gravity in 95c), following an audio prompt about a detectable object in 91c (e.g., a question about which clothing item to pack next). In the selection process, the user selects a clothing item by directing the light beam at a long-sleeved shirt in 92b. In this case, the detectable object may include actual clothing items rather than printed materials or labels.

[0154] The example steps for this process include: In 90a, one or more voice prompts related to clothing selection are played on the speaker located at 91b of the mobile device 91a; In 90b, the light beam 92c emitted from the mobile device 92d is directed not at the nearby sock 92a, but at the long-sleeved shirt 92b (e.g., the next item to be packed); In step 90c, one or more images 93 of the scene pointed to by the camera (and light rays) are acquired using the mobile device camera 93b; At 90d, IMU data is acquired to determine the orientation of the mobile device. This is typically represented as a vector in three-dimensional space, with X at 94a, Y at 94b, and Z at 94c; In 90e, based on IMU data, the handheld device performs calculations in 95a and calculates the device's attitude relative to the Earth's gravitational field in 95c in 95b; In step 90f, the system determines whether the calculated orientation matches a predetermined target orientation (e.g., vertical) within a specified range. If they match, it indicates the device user's selection. If they match, the system proceeds to step 99b to isolate the selected object. If they do not match (e.g., not vertical), the system proceeds to step 99a, where the camera image is reacquired in step 90c, and the IMU data is reacquired in step 90d. Optionally, in 90g, indicated by the dashed rectangle, turn off the handheld device with the 96c beam (used to point to visible objects in 96a and 96b); At 90h, the latest image (97b) is separated from the previously acquired camera image set (e.g., 97a is an image containing a long-sleeved shirt) (e.g., the image illuminated when the handheld device was held stably and vertically). In 90i, the object (e.g., the long-sleeved shirt in 98a) is separated in the optical beam directional region 98b within the camera image (when the beam and camera are pointing in the same direction); In 90j, determine if the target object matches the template and / or characteristics of the detectable object. If it does not match, return to 99c and rebroadcast an audio prompt and / or a new prompt to assist in finding the detectable object. If it does match, proceed to 99d and perform one or more actions based on the successful detection; and Optionally, in 90k, the success of the discovery can be indicated, for example, by displaying a congratulatory message on the device display in 98c and / or by generating haptic vibrations that the device user feels in 98d.

[0155] According to one embodiment, a system and method are provided for an individual to select (point to, identify, and / or indicate) a location or object (e.g., from among several visible objects) from visual content within the pages of a book, and then interact with the page content based on the selection made using a mobile device. The term “page” as used herein refers to a substantially two-dimensional surface on which visible content can be displayed. A page may be made of one or more materials, such as paper, cardboard, film, cloth, wood, plastic, glass, painted surface, printed surface, textured surface, surface reinforced with three-dimensional elements, flexible surface, or electronic display.

[0156] Similarly, in this specification, the term “book” refers to any collection of one or more pages. Books can include printed materials such as traditional (e.g., bound) books, magazines, pamphlets, newspapers, handwritten notes and drawings, covers, chapters, tattoos, boxes, signs, posters, scrapbooks, and collections of photographs and drawings. Pages may also be displayed on electronic screens such as tablets, e-readers, televisions, projection screens, mobile phones, and other screen displays.

[0157] The content of a page is limited solely by the imagination of the author(s). For example, children's books typically include a combination of text and illustrations. More generally, the content of a page can include any combination of text, symbols (e.g., including a range of symbols available in different languages), logos, specifications, drawings, and / or images. It can also be displayed in color, grayscale, or black and white. The content can depict real or fictional scenarios, or a mixture of both.

[0158] In a further embodiment, there is provided an apparatus, system, and method for which one or more predetermined page layouts, which may include one or more page object templates, are recognized by one or more processors (including a mobile terminal processor). A mobile terminal user can use a light beam generated by the terminal to point to a location within a selected (i.e., user-selected) page. As described later, the beam is generated using a laser diode or a light-emitting diode (LED) with a beamforming optical system.

[0159] The camera in the mobile device is pointed in the same direction as the light beam and can capture one or more images of a location or area on the page pointed to by the mobile device user. As long as the beam is on and the surface is sufficiently reflective, the images captured by the camera may include reflections of the light beam (e.g., incident light reflected from the page).

[0160] Alternatively, the beam can be turned off during image acquisition. Turning off the beam (for example, at least temporarily during image acquisition by the camera) allows the page content to be processed without interference from light beam reflection (e.g., using computer vision techniques). Even when the beam is off, because the beam and camera move simultaneously (i.e., are fixed to or built into the handheld device), the location or region in the camera image where the beam is directed can be identified regardless of the physical position, direction of direction, or overall orientation of the handheld device in (three-dimensional) space.

[0161] The structure of mobile devices is sometimes designed to position the beam reflection at the center of the camera image. However, if there is a small gap between the beam source and the camera sensor (for example, due to physical structural constraints as shown in Figure 15), the beam may not appear at the center of the camera image at all working distances.

[0162] The beam and camera field of view can be adjusted to coincide at a given working distance. In this case, the beam can be designed to appear near the center of the camera image (or at a selected image position) within the working distance. In this configuration, as the distance from the mobile device to the reflective surface changes, the beam's position fluctuates within a limited range (generally, one-dimensional in the axial direction within the image plane along the direction defined by a straight line passing through the center of the camera field of view and the center of the light beam source).

[0163] Given a specific working distance (i.e., the distance from the mobile device to the reflective surface), the beam direction, the camera image acquisition direction, and the physical distance between them (see, for example, Figure 15) can be used to calculate the reflection position using geometry (similar to the geometry that describes parallax). By keeping the physical distance between the beam and the camera small, the beam-illuminated area in the camera image can be kept small within the working distance range used in typical applications.

[0164] Alternatively, the beam and camera can be positioned to project and acquire parallel (i.e., non-converging) rays. In this case, the point of reflection is offset from the center of the camera's field of view by an amount that varies with the working distance. The separation distance between the center of the camera image and the center of the beam decreases as the working distance increases (for example, it approaches zero at infinity). The separation distance can also be kept small by keeping the physical distance between the beam and the camera small.

[0165] Regardless of the alignment configuration, if the camera and beam are positioned in the same location within the mobile device, pointing in the same direction and moving in conjunction, the position of the light beam in the image acquired by the camera (including when the beam is off) can be determined. The beam position in the camera-acquired image is calculated based on the separation distance and directional geometry of the beam and camera, or it is empirically determined through a calibration process using the following procedure (e.g., before deployment): Obtain a baseline image that does not include reflections from the projected light beam (e.g., from a featureless surface); Capture a light beam reflection image that includes one or more reflections generated by the projected light beam; Calculate the difference pixel intensity image by subtracting the baseline image from the light beam reflection image; and The beam direction position is calculated based on pixels that exceed a predetermined light intensity threshold within the subtracted pixel intensity image.

[0166] The beam-direction region can be identified based on the location of pixels exceeding an intensity threshold. Alternatively, a single beam-direction location can be determined from the calculated center position of pixels exceeding the threshold intensity (e.g., the two-dimensional median, mean, or center of the beam diffusion function fitted to the intensity profile). Calibration can be performed at different working distances to map the entire range of the beam-direction region.

[0167] In additional embodiments of this specification, when identifying an object or location to be illuminated in an image acquired by a camera, it may be based on identifying beam reflections as high-luminosity regions (e.g., high intensity within a pixel location region) while the beam is illuminated. In such beam reflection localization, it is also possible to consider the color of the light beam (i.e., identifying high luminosity only within one or more colors related to the light beam spectrum). By understanding the relationship between the directional position and the working distance, an estimate of the distance from the handheld device to the reflective surface can be calculated when a camera-acquired image including beam reflections is available.

[0168] As a further example, based on the spectral sensitivity of the typical human eye, a light beam in the green region of the visible spectrum is likely to be most easily perceived by the vast majority of individuals. Many so-called RGB (red, green, blue) cameras have twice as many green sensor elements as red or blue sensor elements. By utilizing a light beam in the mid-band of the visible light spectrum (e.g., green), it becomes possible to keep the beam intensity low while still making it easily detectable (by both humans and cameras), improving the reliability of reflectance detection in camera-based images and enhancing overall eye safety.

[0169] Depending on further aspects of the apparatus, systems, and methods described herein, a predetermined interactive page layout may include additional attributes relating to object templates, object positions, object orientations, page specifications, text or symbolic formats, and / or visual characteristics related to page content. The position of an image within the page layout can be determined using computer vision techniques (e.g., template matching, neural network classification, machine learning, transformer models) that match one or more images acquired by a camera with a predetermined page layout or object template. Such positioning allows the mobile device processor to identify selected locations, regions, and / or objects within the page (i.e., those pointed to by the light beam).

[0170] In this specification, the term “location” (and related terms such as “location”) is generally used to describe the process of determining the correspondence between a camera-captured image and an interactive page layout. As described below, the location process may involve determining several measurements in the camera-captured image within the page layout, such as horizontal and vertical reference positions (e.g., corner positions (x,y)), orientation (q), and magnification (m). Similarly, the term “location” (and related terms such as “location”) is generally used to describe the process of determining the position of a light beam within a camera image (and therefore within the page layout after the camera image has been located within the layout), and / or the position of objects within the page.

[0171] Within a page layout, object templates and attributes may include, for example, the position of objects and object elements, as well as their shape, size, contours, patterns, colors, and textures. These can be stored in various formats, including bitmaps, vector-based scalable architectures, portable images, object element measurements, 2D or 3D CAD datasets, and text-based descriptions and specifications.

[0172] An object's interactive page layout dataset may further contain or reference a set of actions, properties, and / or functions associated with each page location (or page region). As a result, the referenced object or location (and optionally, other visible objects in the region within the camera image) may reference an interactive dataset, which may include, for example, sounds the object makes, sounds associated with the object's functions, the pronunciation and / or audio elements of the object's name or description, and the spelling of the object's name.

[0173] Examples of additional page location and object-related datasets include audio or visual prompts and guidance related to the beam's pointing location, words and phrases describing objects, questions and cues related to one or more objects within the pointing area, spelling or phoneme displays of selected objects, questions about stories related to objects or locations, additional narrative content, rhythmic features related to selected objects (e.g., those that may form the basis for tactile or vibratory stimuli to the device user's hand), audio feedback of reward or comfort upon making a specific selection, pointers to related additional or sequential content, actions if a selection is not made within a given time, and / or one or more actions performed upon successful selection of a page location or object as part of an interactive sequence (e.g., actions by a mobile device or remote processor). These are described in more detail below.

[0174] In a further example, page and content selection may be performed in two stages, with the "context" (e.g., a specific book, book or magazine type, topic area) being selected first. As a result, only page layouts and / or templates related to the identified context are considered for CV-based matching with camera-acquired images in the second stage (i.e., selection by beampointing).

[0175] Predefined context templates include one or more book covers, magazine covers, book chapters, toys, people, anatomical elements, animals, photographs, drawings, words, phrases, household items, classroom supplies, tools, cars, clothing, etc. For example, pointing to a book cover retrieves an interactive page layout of all pages of the identified book, which is then used to identify a light beam-based selection. Context objects (i.e., those pointed to using the light beam to identify the context) can be "virtual" (e.g., real images printed on a page or displayed on a screen) or "real-world" objects (e.g., a cat or clothing in the device user's environment).

[0176] In some cases, it may not be necessary to direct the light beam to a specific item within the camera-captured image in order to evaluate the context. For example, when directing to a book cover, the cover can be identified without directing to specific parts such as words in the book title or the author's name. The CV method determines the match between the entire image or any part of it acquired by the camera and the entire context template or any part of it. In other cases, such as when referring to a real-world object, the boundary region of the beam-directed area within the camera-captured image is considered when determining the match with a predefined context object template database.

[0177] In addition to the visual properties within object templates (described above), each context template dataset can point to the entire set of available interactive pages (or a subset thereof) associated with the context object, and may provide a method for identifying relevant search terms and interactive page layouts. Once a context is identified, subsequent interactions will only consider the interactive page layouts associated with that context until the end of the context sequence is indicated or another context is identified.

[0178] Context sequence termination signals can arise from indications of a topic change by the device user, such as motion gestures detected by the IMU, prolonged pauses in motion or other predictive responses, speech detected by the microphone, or pressing / releasing device push buttons. Similarly, the process of limiting CV comparison to context page layouts will stop when the end of a narrative or book element is reached.

[0179] If a context selection (e.g., a book cover) is not recognized by cross-validation (CV) or other methods using an image acquired by the mobile device's camera, the device may (e.g., automatically) query a centralized (e.g., remote) repository of predefined context templates. If one or more context matches (e.g., by CV or other methods) are found within the repository dataset, one or more context templates and associated datasets are downloaded to the mobile device (e.g., automatically, wirelessly), enabling the continuation of context-aware operations. If no context matches are found (e.g., an image of an unknown book cover is taken), the context query data is stored in the database for future interactive book content development.

[0180] Regardless of whether a contextual pointing step is used, cross-validation (CV) and AI techniques that match images acquired by a camera to a finite set of page layouts, object templates, and / or contextual templates can significantly reduce the computational resources required (e.g., compared to global classification methods). For example, template matching can consider only layouts related to the pages of the book provided to the child. CV and AI techniques using convolutional neural networks, for instance, can be trained on and consider only the pages of the provided book.

[0181] As yet another example, training networks (and other programming of CV / AI processes on mobile devices) can make extensive use of distributed machine learning resources (e.g., TensorFlow, SageMaker, Watson Studio). The limited nature of comparing camera-based images to a layout or template database can significantly simplify the training (and classification) process of determining whether there is a match with a page layout and / or object template. Training classification networks may be limited to a predetermined layout and / or template database (a children's book library), a subset of data such as a specific collection of books or magazines, or a single book or chapter (where the context is known).

[0182] With such limited datasets, relatively simple classification networks and decision trees can be implemented. Optionally, the classification process can be performed entirely on a mobile device (with limited computing resources) or without sending data to a remote device (for example, to access larger computing resources). Such classification can be performed using neural network (or other computer vision / AI) techniques with hardware commonly found in mobile devices. For example, MobileNet and EfficientNet Lite are platforms designed for mobile devices with sufficient computing power to locate the position of images captured by a camera within a book page.

[0183] Classification based on known layout and / or template datasets (e.g., relatively small compared to global classification methods that identify all objects) can also contribute to improving the accuracy of identifying objects illuminated by a light beam. Such limited classifications may be more robust for the following reasons: 1) images are compared only to a database of discoverable pages (e.g., not all possible objects worldwide), 2) training can be performed using discoverable pages, and 3) the cross-validation matching threshold can be adjusted to help accurately reflect the intent of the device user.

[0184] Thresholds for detecting camera-based images within a page are adjusted based on factors such as hardware (e.g., camera resolution), environment (e.g., lighting, object size), specific application (e.g., presence of multiple objects with similar appearances), user category (e.g., young vs. older users, experienced vs. beginners), or specific users (e.g., considering past object selection success rates).

[0185] In accordance with further aspects of this system and method, a given page layout dataset may include and / or point to datasets describing one or more actions performed by the mobile device and / or connected processor during and / or after selection. Actions performed by the mobile device itself may include playing one or more sounds through the device speaker, displaying one or more lighting patterns on one or more device displays, and / or activating a device haptic unit (the device component may be operationally coupled to the device processor).

[0186] The procedure for initiating an action on one or more external devices includes sending the following to one or more remote processors: interactive page layouts, any context and / or object templates, some or all of a camera image (especially the image used when selecting an object), the time the camera image was acquired, the positioned image (e.g., determined position, orientation, and magnification parameters), the beam-pointing position, the specified object and / or position within the page layout, one or more actions within the page layout dataset, and any feedback elements generated by the handheld device. If a selection is not made within a given time, or if a selection is indicated that does not match an acquired template, this may also be communicated to the external processor.

[0187] When a "dwell" is used to indicate the start of positioning by the device user (described below), the transmitted data may include additional dwell thresholds and measurements. This may include two or more camera images used for motion measurement, measured image movement, IMU-based motion measurements (e.g., gestures or taps), and / or pre-configured dwell amplitude and / or time thresholds. Data related to other signaling mechanisms used to identify that a selection is being made by the device user may also be transmitted. This may include voice or acoustic instructions from the device user (e.g., detected by the device microphone), the timing and identification of activation or release of device switches (or other portable device sensors), etc.

[0188] Interaction facilitated by mobile devices can help “bring to life” printed or displayed content by adding sound, additional visual elements, and / or vibrational stimuli felt in the user’s hand (or other body part). Printed content enhanced by real-time interactive sequences that include content-related feedback can not only provide machine-based guidance while reading, but also become “fun,” potentially helping to maintain emotional engagement, especially during reading by and / or for children. For example, reading a book can be enhanced by adding questions (for parents, guardians, and / or children), additional relevant information, sounds, sound effects, audiovisual presentations of relevant objects, and real-time feedback after discovery.

[0189] Interaction with objects within a book can become a shared experience with parents, friends, guardians, and teachers. Details of continuous content delivery control using mobile devices are described in U.S. Joint Application No. 18 / 091,274, filed December 29, 2022, the disclosures of which are expressly incorporated herein by reference. Control of transitions to new pages / panels and sharing of object selections when viewing books or magazines are described in detail in U.S. Patent No. 11,652,654, filed November 22, 2021, the disclosures of which are all expressly incorporated herein by reference.

[0190] As outlined, parents, peers, guardians, or teachers can support learners as they navigate their Zone of Proximal Development (ZPD). The ZPD is a framework in educational psychology that distinguishes between what learners can achieve independently and what they can achieve with guidance (and even what they cannot achieve with guidance). Making books interactive, especially those that challenge learners, provides readily available means to support and guide ZPD transitions in various areas, even in the absence of human interaction (teachers, family, peers, etc.).

[0191] Such readily available support and instructional tools not only eliminate the need for the presence of individuals with sufficient literacy and skills, but also enable machine-assisted Zones of Practice (ZPD) transitions at the learner's chosen time, place, comfortable environment, and pace. Furthermore, tireless, previously used (e.g., with repetitive reward feedback) personalized devices (e.g., appearance, spoken dialect, knowledge level based on the learner's individual background) can further promote learner acceptance, autonomy, self-motivation, and confidence when using handheld devices. Maintaining a challenging environment by bringing books to life can avoid boredom and loss of interest (and at an interactive pace that is not overwhelming), potentially enhancing learning effectiveness.

[0192] Furthermore, in any form, mobile devices and / or remote processors can simultaneously assess engagement, language ability, reading comprehension, and / or comprehension in a continuous and real-time manner. Assessment metrics may include, for example, the measured time a child spends interacting, the success rate of query-based object discovery (especially within different topic areas, including specific areas of interest such as sports, science, and art), the time required for pointing selections (often related to attention and interest), and the overall progress rate in "discovering" new objects within pages of continuous content such as books and magazines.

[0193] Such evaluations can be compared with past interactions of the same subject (e.g., progress assessment in a specific topic area), the usage of the same or similar interactive sequences by others (e.g., same age group, cultural environment, education level), and / or performance across different groups (e.g., comparisons between geographical, economic, and social clusters).

[0194] Milestone responses that demonstrate various aspects of cognitive processing (e.g., first signs regarding color recognition, phoneme and word distinction, understanding the number of objects, performing simple arithmetic operations, and gesture responses requiring controlled motor skills) can be particularly useful for monitoring early childhood development, assessing the learning speed of older users, determining the feasibility of providing more challenging storylines, and / or improving engagement. Auditory, tactile, and / or visual acuity can also be continuously monitored by mobile devices.

[0195] Mobile devices and / or external processors can record interactions for educational and / or parental monitoring purposes. Automatic recording of interactions, reading engagement, content, progress, vocabulary acquisition, reading fluency, and / or comprehension (e.g., comparisons with other children of the same age) can provide insights into a child's emotional and cognitive development.

[0196] Further examples in this specification include considering a variety of individual factors when shaping interactions on mobile devices, such as: the age and / or educational level of the device user; personal hobbies and / or interests; the educational and / or entertainment value of the page content; and whether the object is expected to be the next discovery in a sequence (e.g., storyline, alphabetical, or numerical). For example, children's books for ages 5 to 9 (i.e., beginner readers) typically contain lighthearted text (generally no more than 2,000 words) and illustrations. Children not only learn how to pronounce words but also recognize sounds typically associated with selected illustrations. Optionally, these considerations can be taken into account when designing page layouts and interactive content.

[0197] Interactive elements (e.g., retrieved from a page layout dataset) are generated by the mobile device and initiate and / or guide interaction to displayable objects (e.g., relationships within a storyline, the next object in a logical sequence, the introduction of new objects or concepts). Prompts may include AI-generated and / or scripted (i.e., pre-configured) sequences and may encompass a combination of visual displays, audible sounds, and / or haptic vibrations.

[0198] Scripted sequences may include additional conditional dependencies (i.e., selection from two or more interaction scenarios) determined based on real-time conditions (e.g., past selection success, time of day, user age). The symbol "[Prompt]" represents one or more prompts or queues associated with an object, and typical scripted prompts presented visually, audibly, and haptically on mobile devices include: Can you find the [prompt]? Please show me the largest [prompt]. What are the characters that follow this [prompt] that is displayed? What animal makes this sound [prompt]? What instrument plays the rhythm of the [prompt] you're feeling? Who is viewing the [prompt] on this page? Find each [prompt] on the page by pointing to each one.

[0199] Furthermore, the processor in a mobile device may include a “personality” driven by AI (Artificial Intelligence Personality, AIP), transformer models, and / or large-scale language models (e.g., ChatGPT, Cohere, GooseAI). The AIP implemented within the mobile device can enhance user interaction by including a friendly appearance, conversational style, physical form, and / or voice, and can also incorporate personal insights about the user (e.g., likes, dislikes, preferences).

[0200] Human-machine interactions enhanced by artificial intelligence platforms (AIPs) are described in detail in U.S. Patent No. 10,915,814, filed June 15, 2020, and U.S. Patent No. 10,963,816, filed October 23, 2020, all of which are expressly incorporated herein by reference. Techniques for determining context from audiovisual content and for virtual agents to generate conversations based on that context are described in detail in U.S. Patent No. 11,366,997, filed April 17, 2021, all of which are expressly incorporated herein by reference.

[0201] Whether used independently or as part of a larger system, a mobile device familiar to an individual (e.g., a child) can be a particularly compelling element, providing auditory, tactile, and / or visual rewards as a result of object selection (or conversely, informing the user that their selection may not be part of the correct storyline). The mobile device may also be colored and / or decorated to make it feel like the child's own possession. Similarly, audio feedback (voice, one or more languages, warning sounds, overall volume) and / or visual feedback (letters, symbols, one or more languages, visual object size) may be pre-selected to suit the individual user's preferences, considerations (e.g., auditory ability), skills, and / or other abilities.

[0202] When used alone (e.g., while reading), operation using a handheld device eliminates the need for accessories and other devices such as computer screens, mice, trackballs, styluses, tablets, and other mobile devices when selecting objects or performing activities. Eliminating these accessories (often designed for older or adult users) further eliminates the need for younger users to understand interactive operation procedures involving these devices and pointing mechanisms. When using a handheld device without a computer screen, interaction with images in books and real-world objects (considering the relative richness comparable to screen-based interaction) can be metaphorically described as "making the world your screen without a screen."

[0203] Feedback from a handheld device identifying selected objects or their locations within the page layout is transmitted through the device speaker, one or more device displays, and / or haptic units. Audible interactions or actions include: sounds or sound effects normally emitted by the selected object, sounds related to a description of an activity using the object, pronunciation of the object's name (including proper nouns), celebratory phrases or sentences, parts or single letters of a name related to the object (e.g., beginning with a letter), descriptions of the object and / or its function, verbal descriptions of one or more object attributes, questions about function and / or object attributes, musical scores related to the object, chime notifications, quotes or sayings related to the object, and verbal quizzes in which the selected object (or the next object to be selected) is the answer.

[0204] Visual interactions and actions may include: displaying the name of a selected object or object category; images or drawings of other objects with a similar appearance; images of objects within an object class or category; outlines of objects; caricatures of objects; words or parts of letters related to the spelling of an object (e.g., initials); one or more colors of an object (e.g., swatches displaying an actual color sample and / or one or more words describing the color); the size of an object (e.g., size in particular relative to other visible objects); descriptions or questions about an object; mathematical problems for which an object (number or symbol) is the solution; the next object in a sequence of objects (e.g., the next letter of the alphabet); phrases for which an object is a missing element, etc.

[0205] Haptic feedback during selection can range from simply using vibration patterns to generate vibrations at frequencies related to the actions or sounds produced by the selected object. Haptic actions may also include vibrations synchronized with (or at least at similar frequencies to) the actions or sounds normally produced by the selected object. For example, pointing to an image of a heart can mimic a cat's purring by generating pulsed vibrations approximately once per second, while pointing to a cat can mimic a cat's purring by generating vibrations and / or sounds lasting 10-15 milliseconds at intervals of approximately 30-40 milliseconds.

[0206] Combinations of visual, auditory, and / or tactile actions are generated by mobile devices and / or via devices operationally coupled to remote processors. These action combinations are generated simultaneously or as tightly timed sequences.

[0207] In accordance with further aspects of this system and method, the light beam emitted from the mobile device is generated using one or more laser diodes, such as those manufactured by OSRAM or ROHM Semiconductor. The laser diodes (and general lasers) generate a coherent, parallel, and monochromatic light source.

[0208] Considering the portability of mobile devices, where a nearby light source can be directed in any direction, using a non-coherent light source (such as a non-oscillating light-emitting diode (LED)) can enhance eye safety, especially when used by children, who are often considered to be in an "uncontrollable" environment from a safety perspective. So-called point light sources using LEDs, such as those from Jenoptik and Marktech Optoelectronics, can produce a non-coherent, parallel, and (optionally) polychromatic light source.

[0209] Optical components associated with an LED point light source control the beam divergence angle, thereby deriving the reflective spot size (see Figure 14) at a typical working distance (e.g., approximately 0.05–1.0 meters when a handheld device points to a book page or a real object). The desired spot size may vary depending on the application environment and usage by different users. For example, toddlers may prefer to point to large, close objects on the pages of children's books, while older children and adults may prefer to point to smaller objects such as individual words or symbols on a page or screen (i.e., using a smaller / more focused beam).

[0210] As an additional example, broad-spectrum (at least compared to lasers) and / or multi-color light sources produced by (non-laser) LEDs may help people with color blindness to perceive beam reflections within the visible spectrum. Multi-color light sources may also be more consistently visible to all users when reflected from various surfaces. For example, a pure green light source may be difficult to see when reflected from a pure red surface (such as a specific area of ​​a page). Multi-color light sources in the red-green region of the visible spectrum may help mitigate this problem. High-energy photons in the deep blue region of the visible spectrum may be avoided for eye safety reasons.

[0211] As yet another example, a device beam source can be operationally synchronized with the device's processor, enabling control of the beam intensity (including turning the beam on / off). For instance, turning the beam on by a handheld device can be used as a prompt to indicate that object selection is expected (e.g., after a voice question).

[0212] As an option, the beam can be turned off when acquiring camera images. This prevents beam reflection and pixel saturation near reflection points (such as "blurring" of pixels due to saturation). By eliminating reflections from objects (which may be considered "noise" during object identification), the computational load during computer vision processing can be reduced and accuracy improved.

[0213] Furthermore, the beam can be turned off when a match is determined as part of a success notification to the user. Conversely, keeping the beam illuminated during operation can suggest to the device user that further searching for page objects or locations is expected. Methods for using a light beam emitted from a mobile device as a pointing indicator in response to questions or cues are further described in joint patent application serial number 18 / 201,094, filed on 23 May 2023, all of which are expressly incorporated herein.

[0214] As described later, device users can indicate (i.e., from the user to the mobile device) that the beam is directed towards the desired target using various methods (e.g., dwell time, voice commands). These methods include turning the beam on and off multiple times. For example, if the waiting time is measured based on not detecting motion in the camera image, the beam can be turned off during each image acquisition while remaining on at other times, allowing the device user to remain involved in the directing process. The beam can be turned off if a predetermined waiting time is exceeded, or if the handheld device detects another selection instruction means.

[0215] Optionally, the beam intensity can also be modulated. For example, this can be based on measuring one or more reflections in an image acquired by a mobile device camera. The reflection intensity can be set so that it is clearly distinguishable to the user against the background (e.g., to take into account ambient lighting conditions, to accommodate the reflectivity of different surfaces, and / or to accommodate visual impairments), but not to be overwhelming (e.g., based on user preference). The beam intensity can be modulated by several methods known in the art (e.g., adjusting the optical beam driving current with transistor-based circuits, using pulse width modulation (PWM)).

[0216] A further aspect of this system and method is the ability to project one or more illumination patterns into the device beam (and / or display them on the device display). Projecting one or more illumination patterns using the beam effectively integrates the role of beampointing with one or more independent displays on the handheld device. The illumination patterns projected into the beam can be formed, for example, using a miniature LED array, LCD filtering, or DLP (i.e., digital light processing with a fine mirror array) techniques.

[0217] These lighting patterns range from simple images (e.g., one or more points, line drawings) to complex, rich, high-resolution images. Lighting patterns may be animations generated to enhance or expand the projected object, image, or text. They may move in proportion to the movement of the projecting mobile device, or they may be dynamically stabilized using images captured by the IMU and / or camera, ensuring a stable appearance even as the user moves the device.

[0218] In yet another example, during an interaction with a mobile device, one or more page objects identified in an image captured by the device's camera may be "enlarged" by the projection of a light beam. For example, if you are looking for a bird in the context of a story (e.g., what you are listening to) and you shine a beam at a squirrel, the beam will project a pair of wings onto the printed image of the squirrel, serving as part of a (funny) interaction that asks if the pointed object is a bird. In yet another example, it is possible to change the apparent color of one or more components of a printed object by illuminating the shape of the entire object (or its individual components) with a selected color in the beam.

[0219] Illumination patterns can be synchronized with other interaction modes during any activity. For example, each word read aloud by the device can be independently and dynamically illuminated, framed, underlined, or have other annotations added. Similarly, projected images can be overlaid on still images (e.g., images of animals) and displayed as if bouncing around on a printed page.

[0220] When images or symbols (e.g., characters that make up words or phrases) are too long or too complex to display all at once, messages or patterns within the beam can be "scrolled." Scrolled text or graphics are displayed segment by segment (e.g., providing a dynamic appearance) in a given direction (e.g., upward, downward, or horizontal). During and after the object selection process using a light beam (i.e., when attention is focused on the beam), messages embedded within the beam (e.g., the name of the object being identified) may be particularly noticeable, effective, and / or meaningful.

[0221] The illumination patterns generated by the beam light source are used to enhance pointing functionality, including control over the size and / or shape of the beam visible to the device user. Within the illumination pattern, the beam size (e.g., related to the number of illuminated pixels) and relative position (the position of the illuminated pixels) are controlled by the mobile device. For example, different beam sizes can be used for different applications (e.g., pointing to characters in text (using a thin beam) and pointing to large cartoon characters (using a larger beam)). By "fine-tuning" the beam position and indicating the pointing direction on the mobile device, it helps direct the user's attention to specific (e.g., nearby) objects within the camera's field of view.

[0222] As a further aspect of this system and method, visible information and / or symbols in a beam projection with optical performance similar to that of a device camera (e.g., common depth of field, minimal distortion even when viewed perpendicular to a reflective surface) tend to prompt (and / or psychologically guide) the mobile device user to adjust the orientation and position of the mobile device so that the information and / or symbols are most visible to both the user and the device camera (e.g., in focus and undistorted). As a result, properly positioned and oriented camera-acquired images (i.e., images captured at a working distance and field of view that are easily visible to both the user and the camera) can facilitate computer vision processing (e.g., improved reliability and accuracy of classification).

[0223] For example, symbols or patterns consisting of a "smiley face" may be easily (perhaps inherently) recognized by infants. With minimal instruction, infants may instinctively position a device in terms of distance and / or angle to obtain a clear, undistorted image of the smiley face, based on their innate recognition of this universally recognizable image. Furthermore, the lens and filter structure of a smiley face light beam can be designed to focus the image only within a predetermined "sweet spot" distance from the reflective surface. Users may not realize that the ability to easily see and identify the projected beam pattern also contributes to improved image quality in camera-based image processing.

[0224] Furthermore, if the projected image has directionality or a typical viewing orientation (e.g., text, upright figure), most users tend to operate their mobile devices so that the projected image faces the object in a typical viewing orientation. Alternatively, the projection pattern and specific image orientation can be maintained by measuring the orientation of the mobile device (e.g., using an IMU) (e.g., using Earth's gravity as a reference). The image is projected in a way that maintains a preferred orientation, such as its relative orientation to other objects in the mobile device user's environment.

[0225] As a further aspect of this system and method, the mobile device user can instruct (the mobile device) that the beam is directed towards the selected object and that selection is being made. Such instruction is performed by various interactive methods, such as: Pressing or releasing a switch (e.g., a push button, other contact sensor, or proximity sensor) that is a component of a mobile device. The mobile device's microphone detects and identifies (e.g., classifies using natural language processing) voice commands (e.g., saying "now") by the device's processor or remote processor. To direct a beam towards an object (i.e., without any actual movement) and hold it for a predetermined "dwell" time (e.g., based on user preference), The handheld device is oriented in a predetermined direction detected by the IMU (e.g., perpendicular to Earth's gravity, tilting the device forward), or Perform a gesture or tap detected by the IMU of the handheld device.

[0226] In the latter exemplary case (where user actions on a mobile device, such as gestures or taps, may cause movement within the camera's field of view), still images can be separated (e.g., from a series of sequentially sampled images) before any motion-based signaling occurs. The camera-acquired images before motion-based signaling are used to identify the indicated visible object or location.

[0227] In yet another embodiment, one method for implementing a dwell time-based approach is to ensure that a certain number of consecutive images (e.g., calculated by dividing a predetermined dwell time by the frame rate) capture substantially stationary visible objects and / or beam reflections. Using CV techniques such as template matching, computer vision, and neural network classification, pairs of consecutively acquired camera images can be compared to calculate one or more spatial offsets. Image movement (e.g., for comparison with a dwell movement threshold) is calculated from one or more spatial offsets, or the sum of offsets at a selected time.

[0228] Measuring dwell time quickly and / or precisely requires a high frame rate for motion measurement based on camera images, resulting in computational power and / or power demands. An alternative method for determining whether sufficient dwell time has elapsed is to use an IMU to evaluate whether the mobile device has remained substantially stationary for a predetermined period.

[0229] To convert analog IMU data into a digital format suitable for processing, analog-to-digital (A / D) conversion techniques well-known in this field can be used. IMU sampling rates generally range from approximately 10 samples per second (lower rates are available if necessary) to approximately 10,000 samples per second, with higher sampling rates resulting in trade-offs with signal noise, cost, power consumption, and circuit complexity. Operating and / or residence time thresholds can be set based on the device user's preferences.

[0230] Gesture-based selection instructions include translational motion, rotation, no motion, tapping the device, and / or device orientation. User intent is indicated, for example, by: Any type of operation (e.g., anything that exceeds the IMU noise level), Movement in a specific direction, Speed ​​exceeding a threshold (e.g., in any direction), Gestures using handheld devices (e.g., known movement patterns), Orienting a handheld device in a predetermined direction, Lightly tap the mobile device with the fingers of your other hand. Smashing a mobile device with an object (e.g., a stylus), Slamming a mobile device against a solid object (e.g., a desk), and / or The act of slamming one mobile device against another.

[0231] As a further example of user intent determination based on IMU data streams, a "tap" on a mobile device may be identified as the result of a user-intentionally moved object (i.e., "object") making contact with a target location on the mobile device surface (i.e., "tap location"). The tap location calculated on the handheld device can be used to convey additional information about the device user's intent (i.e., in addition to object selection). For example, the user's confidence in a selection, indicating the first or last selection among a group of objects, or the intention to "skip ahead" during a continuous interaction sequence can each be signaled based on directional movement and / or tap location on the mobile device.

[0232] A tap is determined when a stationary mobile device collides with a moving object (e.g., fingers of the hand opposite to the one holding the device), when the mobile device itself moves and collides with another object (e.g., a table, another mobile device), or when the colliding object and the mobile device move simultaneously before contact. The IMU data stream before and after the tap helps determine whether an object was used to strike the stationary device, whether the device was forcibly moved to another object, or whether both processes occurred simultaneously.

[0233] Tap locations can be identified using characteristic "signatures" or waveform patterns (e.g., peak force, acceleration direction) within the IMU data stream (particularly accelerometer and gyroscope data), which vary with the tap location. Identification of tap locations on the surface of a mobile device based on inertial (i.e., IMU) measurements and subsequent motion control are described in detail in U.S. Patent No. 11,614,781 filed July 26, 2022, the disclosures of which are expressly incorporated herein by reference.

[0234] As a further aspect of the apparatus and methods described herein, the portable device processor may perform various actions based on selection instructions. If it determines that a camera image does not match any page layout, the portable device can simply acquire subsequent camera images and continue monitoring whether a match is found. Alternatively, or additionally, it may provide prompts and / or queues, which may include repeating a previous interaction, maintaining the beam's illuminated state, and / or presenting new interactions (e.g., retrieval from a template database) to expedite and / or enhance the selection process.

[0235] Identifying the page layout position and / or selected objects may trigger actions performed by the mobile device's processor within the device itself. As mentioned above, these may include, for example, audible, haptic, and / or visual actions that indicate a successful page layout match and / or reinforce the storyline within the page.

[0236] Actions performed by the mobile device processor include sending available information related to the selection process to a remote device. For example, the remote device may perform further actions. The information sent may include camera images, camera image acquisition time, a given camera image light beam direction area, a template for the selected page, any contextual information, and information on one or more selected objects and locations.

[0237] Especially when used in entertainment, educational, and collaborative environments, the ability to transmit object detection results makes mobile devices part of a larger system. For example, when used by children or learners, experiences (e.g., successful / unsuccessful object selection) can be shared, recorded, evaluated, or simply enjoyed with connected parents, relatives, friends, or guardians. By assessing indicators such as literacy, reading comprehension, overall interest in reading, and skill development in real time and substantially continuously, it may be helpful to identify children who could benefit from additional support or resources (at the earliest stage where intervention is most effective). Examples of early intervention support include identifying users who could benefit from adaptation strategies for dyslexia, dyscalculia, autism, or programs for gifted children.

[0238] Optionally, the mobile device may further include one or more photodiodes, optical blood sensors, and / or electrical heart rate sensors. Each of these is operationally connected to the terminal processor. These mobile device components provide auxiliary elements (i.e., inputs) for monitoring and determining user interactions. For example, a data stream from a heart rate monitor may indicate stress or distress during the selection process. Based on the detected level, past interactions, and / or predefined user settings, interactions involving object selection may be limited, delayed, or aborted.

[0239] As an additional example, such portable electronic devices, while not strictly "portable," may be attached to and / or operated from other parts of the human body. For instance, a device that interacts with a user to direct a beam of light at an object may be attached to the arm, leg, foot, or head. Such arrangements may be used to address accessibility issues in individuals with limited upper limb and / or hand movement, individuals lacking sufficient manual dexterity to communicate intentions, individuals with missing hands, and / or situations where hands are required for other activities.

[0240] When using handheld devices, accessibility-related factors are also considered. For example, when an individual with color blindness uses a device, it is possible to avoid certain colors or color patterns in visual operation. The size and brightness of symbols and images displayed on one or more handheld device screens and / or within the beam are adjusted to accommodate individuals with visual impairments. Media containing selectable objects may be braille-enhanced (e.g., including both braille and images) and / or may include patterns and / or textures with raised edges. Pointing using a light beam can be complemented by enhancing beam intensity and / or having a mobile device camera track finger pointing (e.g., within areas containing braille).

[0241] Similarly, if an individual has hearing impairment in one or more speech frequency bands, it is possible to avoid or amplify those frequencies (e.g., depending on the type of hearing impairment) in the voice interactions generated by the mobile device. Haptic interactions can also be adjusted to take into account the individual's heightened or suppressed tactile sensitivity.

[0242] For example, activities involving young children or individuals with cognitive impairments may require considerable "guessing" or guidance from the device user during operation. User assistance during operation and relaxation of pointing accuracy can be considered a form of "interpretive control." Interpretive control includes "guiding" to one or more target responses or reactions (e.g., providing intermediate hints). For example, young children may not fully understand how to operate a mobile device. In such interactions, voice instructions (e.g., "Lift the wand straight up") accompany the operation process and guide the user to a choice.

[0243] Similarly, when the user approaches a selectable object or a predicted response, flashing indicators or repeating sounds (the frequency of which may be related to the proximity of cues or attributes to a particular selection) may be broadcast. On the other hand, responses that do not show an obvious attempt to point the beam may be accompanied by “questioning” indicators (e.g., haptic feedback or buzzer sounds) that serve as prompts to consider alternatives. Further aspects of interpretation control are described in detail in U.S. Patent No. 11,334,178 filed August 6, 2021, and U.S. Patent No. 11,409,359 filed November 19, 2021, all of which are expressly incorporated herein by reference.

[0244] Figure 10 illustrates an exemplary scenario in which a child 101 uses their right hand 104 to manipulate a light beam 100a generated from a mobile device 105 to point to (and select) a picture of a dog 102b. The dog 102b is one of several characters appearing in a cartoon scene 102a depicted on two pages 103a, 103b of a children's book. The child 101 can see a reflection 100b generated by the light beam 100a at the location of the dog 102b on the right-hand page 103b of the printed book. The child can identify and select the dog at the beam position 100b during an interactive sequence using the mobile device 105, resulting in audio feedback (e.g., a barking sound) being played on the device speaker 106 and / or displayed on the device displays 107a, 107b, 107c, and / or within the projected light beam 100a.

[0245] Optionally, the light beam 100a can be illuminated by pressing push button 108 (or other signaling mechanism, such as an audio prompt like "OK" detected by the device microphone, or a motion gesture against Earth's gravity or the orientation of a handheld device sensed by the IMU). Releasing push button 108 (or other signaling mechanism) is used to indicate that a selection has been made (i.e., a dog selection at beam position 100b) and is optionally used to turn off the light beam 100a (e.g., until the next selection is made).

[0246] A mobile device camera (not visible from the viewpoint in Figure 10) pointed at the page in the same direction as beam 100a can capture one or more images of the area indicated by the light beam. By positioning the area indicated by the light beam and the camera's field of view in the same location, objects or locations selected using the light beam can be identified within the interactive book page template.

[0247] Figure 11 continues the illustrative scenario shown in Figure 10, showing the field of view of the mobile device camera 111 within two book pages 113a and 113b. Similar to Figure 10, the mobile device 115 is operated by the child's right hand 114 to point to a selected object (e.g., a dog 112b) or location within the cartoon scene 112a.

[0248] During camera image acquisition, the light beam can remain lit (for example, normally generating visible reflection from page 113b), or it can be selectively turned off momentarily (i.e., during camera acquisition), as shown by the dashed line in Figure 11 (a line crossing the beam path 110a, assuming the lit state). Since the light ray and the camera's optical path are generated inside the mobile terminal 115 and point in the same direction, the position indicated by the light ray (indicated by the crosshair pattern 110b) can be identified in the camera image even when the light ray is off. Turning off the light ray during camera base acquisition avoids image distortion due to light ray reflection and makes it easier to match the camera image with the position in the page layout using CV processing.

[0249] When a predetermined interactive page layout or template (or part thereof) matches the field of view of the mobile device camera 111, the image of the field of view can be superimposed onto the page (see Figure 11). This makes it possible to identify the beam position (i.e., a known position in the camera image) within the page layout (and / or associated database).

[0250] By grasping the position pointed by the device user, access to a predetermined dataset associated with a position (or region) within the interactive page layout is triggered. In the case of the dog shown in FIG. 11, information regarding the specific dog pointed, information regarding dogs in general, and / or the role of the dog in the context of the story may be conveyed.

[0251] For example, the three spherical displays 117a, 117b, 117c on the mobile terminal display the word "DOG" (when pointing to dog 112b). Alternatively, or additionally, the device speakers 116a, 116b can announce the word "dog", play a barking sound, and / or provide proper nouns of the dog and general information regarding the dog. As further examples of rewards and / or feedback by the mobile terminal, a tactile stimulus generated by the mobile terminal may be provided along with the barking sound, or it may be confirmed (using vibration) that a predicted selection (e.g., response to an inquiry by the mobile terminal) has been made.

[0252] Furthermore, the occurrence of the user selection, the selection position within the template, the associated layout dataset, the timing and identification information of the selected object 112b (or the lack of selection) may then control further actions that are directly executed by the mobile terminal 115 and / or trigger further actions to be communicated to one or more remote terminals (not shown).

[0253] Figure 12 shows exemplary parameters that may be calculated using a CV-based method to match an image 121 acquired by a camera with a given page layout or template 123. In this example, the page includes a star symbol 122a, a cartoonish depiction of a unicorn 122b, a cat illustration 122c, text about the cat's name 122d, and a page number 122e. In addition to the page layout 123, page attributes may point to one or more datasets containing additional information (e.g., text, audio clips, queries, sound effects, follow-up prompts) that may be used in subsequent interactions involving the mobile device user.

[0254] Computer vision techniques such as convolutional neural networks, machine learning, deep learning networks, transformer models, and / or template matching may be used to position images acquired by the camera within the layout and / or template dataset. Position parameters may include the reference position of the camera image 121 within the template 123. Horizontal and vertical coordinates (e.g., using a Cartesian coordinate system) are specified, for example, by determining the position of the lower-left corner of the camera's field of view (usually represented as (x,y)) in 125a as its relative position to the page layout's coordinate system (usually with the lower-left corner as the origin (0,0)).

[0255] Because device users may tilt (i.e., change direction) their handheld devices when illuminating the light beam (for example, relative to the page orientation), the camera's field of view may not coincide with the page layout's coordinate system. To compensate for this, the direction angle (usually represented by q) can be calculated based on the image orientation.

[0256] Similarly, users may move their mobile devices at different distances from the page surface. Under these conditions, using a device camera with fixed optics (i.e., no optical zoom), the page area covered by the camera's field of view will change (i.e., a wider area will be covered if the mobile device is held further away from the page). Therefore, the magnification, usually represented by m, is calculated based on the size of the page being covered. Putting these together, (x,y), q, and m can be used to calculate (i.e., overlay) any position in the image acquired by the camera (e.g., the position of the light beam) onto the page layout.

[0257] Figure 13 is an exemplary electronic connection diagram of a mobile terminal 135, showing 132a, 132b, 132c, 132d, 132e, 132f, 132g, 132h, 132i, 132j, 133, and 134, indicating the main direction of information flow during use (i.e., indicated by the direction of the arrows relative to the electronic bus structure 130 that constitutes the backbone of the device circuit). All electronic components can communicate with one or more processors 133 via this electronic bus 130 and / or by a direct path (not shown). In certain applications, some components may be unnecessary or not used.

[0258] The core of a portable handheld device is one or more processors (including microcomputers, microcontrollers, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.) 133, powered by one or more (typically rechargeable or replaceable) batteries 134. As shown in Figure 16, the components of the portable device also include a light beam generation component 132c (e.g., typically a light-emitting diode) and a camera 132d that detects objects in the beam region (which may include reflections from the beam). If both the beam source 132c and the camera 132d are integrated into the handheld terminal body 135, each component may require one or more optical apertures and / or optical transmissions (131b and 131c, respectively) passing through the handheld terminal housing 135 or other structure.

[0259] In applications involving acoustic cues or feedback, a speaker 132f (e.g., electromagnetic coil or piezoelectric) 132f may be used. Similarly, in applications that may involve voice-based user interaction, a microphone 132e captures ambient sounds from the mobile device. When integrated within a handheld device 135, the operation of both the speaker 132f and the microphone 132e is aided by acoustic transmission through the handheld device housing 135 or other structures, for example, by making contact with the device housing and / or by including multiple perforations 131d (further shown, e.g., in 26a of Figure 11).

[0260] For example, in applications that include vibration feedback or that notify the user of an expected selection, a tactile unit 132a (e.g., an eccentric rotating mass or a piezoelectric actuator) is employed. One or more tactile units are mechanically coupled to a position on the device housing (e.g., so that they are felt at a specific location on the device) or fixed to an internal support structure (e.g., designed to be felt more generally across the entire surface of the device).

[0261] Similarly, in applications that include visual feedback or responses after object selection, one or more displays 132b are used to display, for example, characters (as illustrated), words, images, or shapes related to the selected object (or to indicate that the object was pointed to incorrectly). Such one or more displays may be fixed to and / or located outside the main portable device body (as shown in 132b), and / or optical transparency may be utilized within the device housing as shown in 131a.

[0262] During a typical operation, the user may send signals to the mobile device at various times, such as when they are ready to select another object, when they point to a new object, or when they agree on a previously selected object. User signals are indicated by voice feedback detected by the microphone, motion gestures detected by the IMU132g, and the physical orientation of the mobile device. Although shown as a single device in Figure 132g, different implementations may employ distributed subcomponents, for example, individually detecting acceleration, gyroscopic motion, magnetic orientation, and gravitational force. Furthermore, subcomponents can be placed in different areas of the device structure (e.g., distal arms, electrically quiet areas) to improve the signal-to-noise ratio during detected operation.

[0263] User signals can also be indicated using one or more switch devices, including push buttons, toggle switches, contact switches, capacitive switches, and proximity switches. Such switch-based sensors may require structural components on or near the surface of the mobile terminal 131e to transmit force or motion to internally located circuits.

[0264] Communication with the mobile terminal 135 can be implemented using Wi-Fi 132i and / or Bluetooth 132j ​​hardware and protocols (e.g., each using different electromagnetic spectrum regions). In an example scenario using both protocols, Bluetooth 132j ​​for short-range communication may be used, for example, when registering the mobile terminal using a mobile phone or tablet (e.g., identifying the Wi-Fi network and entering the password). Subsequently, the Wi-Fi protocol is adopted, allowing the activated mobile terminal to communicate directly with other remote devices or the World Wide Web.

[0265] Figure 14 shows the beam generation 141a, 141b, 141c, 141d, the beam irradiation 140a and the reflected light 140c optical paths, and the detection of a targeted drawing of a cat 149 on pages 143a and 143b of a book 147 by camera 146a. The drawing of the cat 149 can be selected from page sketches including a cartoon character 142a on the leftmost page 143a, a second cartoon character 142c on the rightmost page 143b, and a unicorn 142b. The electronic circuits 141a, 141b, 141c, 141d, as well as camera 146a and associated optics 144, can be incorporated into the body of a portable device (not shown).

[0266] The components of the beam generation circuit include: 1) a power supply 141a consisting of a rechargeable or replaceable battery, typically built into a portable handheld device; 2) optionally a switch 141b or other electronic control device (e.g., push button, relay, transistor) controlled by the handheld device's processor and / or the device user to switch the pointing beam on and off; 3) a resistor 141c (and / or a transistor that can adjust the beam intensity). The LED is typically configured in the forward bias direction (i.e., low resistance) to limit the current supplied to the beam source. 4) the beam source, typically consisting of a laser oscillator or light-emitting diode 141d.

[0267] A precision optical system (which may consist of multiple optical elements, including those enclosed within a diode-based light source, but not shown) makes the beam 140a nearly parallel and provides a small beam divergence angle as needed. As a result, the dimensions of the beam emitted from the light source 148a may be smaller than those at points further away on the optical path 148b. The divergence angle of the illumination beam 140a and the distance between the light source 141d and the selected object 149 primarily determine the reflected beam spot size 140b.

[0268] Furthermore, the beam 140c reflected by the selected object 140c may continue to diffuse, for example, in Figure 14 the size of the reflected beam 148c is larger than that of the illuminated beam 148b. The size and shape of the illumination spot 140b (and its reflection) may also be affected by the position of the beam relative to the reflective surface (e.g., the angle with respect to the normal of the reflective surface) and / or the shape of the reflective surface. For example, because the horizontal dimension of page 143a on the far left is convexly curved relative to the incident beam 140a, an illumination beam (e.g., Gaussian distribution profile and / or circular) may produce an elliptical (i.e., wider horizontal dimension) reflected spot 140b.

[0269] Light from the field of view of camera 146a is collected by the camera optics 144 and images an image containing a focused reflected beam spot 145b, as indicated by the ray 145a, onto the light-sensing component of camera 146b. Such detected images are digitized using methods known in the prior art and then processed (e.g., using CV techniques) to identify the position indicated within the camera's field of view.

[0270] Figure 15 is an exploded view of the portable device 155, showing the exemplary positions of the light beam source 151a and the camera 156a. These components may be incorporated into the portable device 155 during final assembly. This figure of the portable device 155 also shows the back surfaces of the three spherical displays 157a, 157b, and 157c mounted on the body.

[0271] The light beam generator consists of a laser or non-laser light-emitting diode 151a and may include built-in and / or external optical components (not visible in Figure 15) for forming, structuring, and / or parallelizing the light beam 150. The beam-generating electronics and optics are housed in a subassembly 151b that provides electrical connections to the beam generator and precise control of beam aiming.

[0272] Similarly, the image acquisition process is realized by the focusing optical system 154a incorporated within the threaded housing 154b. This housing allows for the inclusion of additional (optional) optical systems in the optical path for magnification and / or optical filtering (e.g., to remove reflected light generated from the beam). The optical components are attached to the camera assembly 156a (i.e., including the imaging element surface), and this camera assembly is housed within a sub-assembly that provides precise control of the electrical contacts of the camera and the imaging direction.

[0273] As one aspect of the exemplary configuration shown in FIG. 15, the point that the light beam 150 and the image acquisition optical system of the camera 154a are directed in the same direction 152 can be mentioned. As a result, the beam reflection from the visible object occurs within substantially the same region in the acquired camera image, regardless of the overall pointing direction and / or orientation of the handheld device. Depending on the relative alignment and separation (i.e., between the beam light source and the camera 153), the position of the beam reflection may be centered in the acquired camera image (at typical operating distances), or may be somewhat offset from the center. Furthermore, due to the (small-designed) distance 153 between the beam light source 151a and the camera 156a, there may be a slight difference in the beam position when the distance from the portable terminal to the reflecting surface is different.

[0274] Such differences can be estimated using a mathematical approach similar to the method for explaining parallax. As a result, even when the pointing beam is off (i.e., there is no beam reflection in the camera image), the beam position can be determined based on the position where the beam optical system is directed within the camera's field of view. Conversely, the deviation measured at the center position (or other reference point) of the light beam reflection in the camera image can be used, based on geometry, to estimate the distance from the portable terminal (more specifically, the terminal camera) to the visible object.

[0275] Figure 16 is an exemplary flowchart illustrating the procedure for selecting an image of a shoe 162c, which is targeted using a light beam 162b generated by a mobile device 162d, within an interactive page 164a. The shoe 162c is one of the options in a collection of fashion accessories, including a T-shirt 162a. Once its position is identified within the page layout, predetermined acoustic feedback associated with the shoe selection (e.g., type, features) within the interactive page is played through the device speaker 168b. The steps of this selection process are as follows: In 160a, the mobile device's processor obtains a predetermined interactive page layout 161a, which includes a page 161b displaying clothing accessories, including a baseball cap, a T-shirt, and shoes; In 160b, the user directs the light beam 162b generated by the mobile device 162d towards the shoe 162c located near the image of the T-shirt 162a; Optionally (indicated by a dashed box), in 160c, the light emitted from the mobile terminal 163b (directed towards the image of the shoe 163a) may be turned off before the camera-based image acquisition to avoid light reflection; In 160d, using a camera (not shown) and a focusing optical system 164d (shown separately from the mobile device for illustrative purposes), the camera's field of view 164a captures an image that includes the lower half of the T-shirt 164b and the shoes 165c; In step 160e, the page layout and / or template previously obtained in step 160a is compared with the camera-based image 165a (including the lower half of the T-shirt 165b and the shoes 165c) to determine the best match and alignment within the page layout; At 160f, a region including the light beam reflection position or region 166b (regardless of the beam's illumination state) representing the shoe portion 166c by the device user is separated from the camera image 166a; In step 160g, based on the selection location or region within the page layout obtained in step 160a, the selected object (i.e., the shoe) and associated interactive elements are identified from the page layout database; At 160h, based on the shoe selection, interactive operations are performed within the page layout database using audio (using the mobile device speaker 168b) and / or visual (using one or more device displays 168a).

[0276] Figure 17 is a flowchart based on Figure 16, illustrating a two-step process for a user to interact with images and content related to cats. First, the user specifies (i.e., selects) a cat context, and then interacts with one or more interactive pages containing cat-related content. The first step involves using a light beam 172e emitted from the mobile device 172e to select the word "CAT" 172c from a cluster of text options (animal lists 172a, 172b, 172c). In this case, context identification includes step 170e, which uses the beam position to isolate a specific context object in the image acquired by the camera (an optional step not required in other cases, such as identifying a book cover).

[0277] Once the feline context is established by the device user, the device processor identifies and / or retrieves cat-related page layouts (e.g., books, posters, articles 175a about cats). In this second stage, the user can again interact with the page content, including audio and / or visual information and feedback, through operations involving the light beam 176c. The steps of this two-stage process are as follows: In 170a, the mobile device's processor retrieves a page layout and / or object template 171a containing an animal list 171b with images and text related to cats; In 170b, the user uses the light beam 172d generated by the mobile terminal 172e to point to the word "cat" 172c from a list of animals, which includes "cow" 172a and "dog" 172b; Optionally (indicated by a dashed box), in 170c, the light beam emitted from the mobile terminal 172f (directed towards "CAT" 172c) can be turned off, thereby avoiding interference from light beam reflection within the image during text content recognition; In 170d, using a camera (not shown) and a focusing optical system 173d (shown separately from the mobile device for illustrative purposes), an image is captured in which the camera's field of view 173a includes "DOG" 173b and "CAT" 173c; In step 170e (as indicated by the dashed box, step 170e may not be performed in all cases), a region containing the location or region of the light beam reflection 173f that indicates the region specified by the device user is isolated from the camera image 173e (regardless of whether the beam is on or off); In step 170f, in this exemplary case, the template matching method 174a (using the template obtained in step 170a) is used to identify the selection item indicated by the light beam; In 170g, the template matching results identify the word "CAT" 174b as the context for future interaction; At 170h, the mobile device's processor retrieves multiple page layouts 175a related to cats. This includes page 175b, which contains images and text about a cat named Sam; In 170i, the user again uses the light beam 176c generated by the mobile terminal 176d to point to a selected picture of cat 176b located near the picture of unicorn 176a (for example, on the same page); Optionally (indicated by a dashed box), in the 170j, the light beam emitted from the mobile device 176e (directed towards the cat 176c) can be turned off to avoid reflection of the light beam during camera-based image acquisition; At 170k, an image 177a containing a picture of a cat 177c and a nearby unicorn 177b is acquired using a mobile device camera and a focus adjustment optical system 177d (illustrated separately from the mobile device for illustrative purposes); In step 170l, the interactive page layout for the cat, previously acquired in step 170h, is compared with the camera-based image 177d (which includes drawings of the unicorn 177e and the cat 177f) to determine the match and optimal alignment within the layout; At 170m, the region containing the light beam reflection position 178b (regardless of the beam's on / off state) is separated from the camera image 178a. This indicates the area selected by the device user within the cat's region 178c; In step 170n, based on the selection location or region within the page layout obtained in step 170h, the selected object (e.g., a cat named "Sam") and interactive elements (including Sam's purring function 179b) are identified from the page layout dataset; In 170o, based on the selection of a mobile device, auditory (rumbling sound from the mobile device speaker 179f) and visual (displaying the name "SAM" on three device displays 179e, 179d, and 179c) interactive actions are performed.

[0278] Figure 18 is a flowchart illustrating how the on / off switch of the light beam 182c is performed using a push button 181c on a mobile device, with the user making a selection. In this exemplary case, when shoes 182b are selected from the clothing items, an audio explanation (e.g., an explanation of the shoe's performance) is provided via 188c. Furthermore, dynamic distortion of the image caused by button operation is avoided by separating the camera image captured just before the push button is released. The exemplary steps of this process are as follows: In 180a, the mobile device's processor retrieves page layout 181a containing clothing options; In 180b, the user uses their thumb 181d to press the push button 181c on the mobile device 181b, illuminating the pointing light beam (if it is not already lit, and / or to notify the device that a selection will be made soon); In 180c, the user operates the mobile device 182d to direct the light beam 182c towards the shoe 182b (placed on the page adjacent to the baseball cap 182a); In 180d, the mobile device's camera and optics 183c collect an image 183a. The field of view of this image includes the shoe 183b pointed to by the device user; In 180e, the processor in the mobile terminal 184a obtains the state of the push button 184b and determines whether it has been released by the user's thumb 184c; In step 180f, if the push button is not released, step 189a returns to step 180d and waits for the push button to be released. Otherwise, step 189b determines the targeted page area. Optionally (indicated by a dashed box), in 180g, the light beam 185c (directed towards the interactive page area including the baseball cap 185a and shoes 185b) is turned off; In 180h, to avoid the influence of motion in the camera image when the push button is released, the most recently acquired image 186b (i.e., before the push button was released) is separated from the already acquired image sequence 186a; In step 180i, the page layout previously acquired in step 180a is compared with the camera field of view 187b (including part of the hat 187a and the shoes 187c) to determine the match and optimal alignment within the page layout; In 180j, based on the selected beam position within page layout 188a, identify the selected object (i.e., shoe 188b) and retrieve the interactive element from the page layout dataset; and In 180k, audio descriptions related to the selected shoes, obtained from the selected page dataset, are played using the mobile device speaker 188c.

[0279] As described in the detailed explanation, the method may involve prompting, or "mentally inducing," the device user to position the handheld device at a target distance from a visible surface. The visible surface may include text, symbols, images, drawings, or other content that the device user can identify and / or select based on the pointing of the light beam and image acquired by a handheld device camera positioned coaxially with the device beam. Prompting the handheld device to position itself near the target distance may result in the user being able to sense and / or recognize a focused light beam pattern reflected from the visible surface.

[0280] In particular, if the visualized pattern is perceived as desirable (e.g., a smiley face, the outline of a favorite toy) and / or practical (e.g., a pointing arrow, the outline of a finger), mobile device users tend to try to keep the image projected by the beam within focus, and as a result, may hold the mobile device at a distance from the visible surface approximately to the target working distance. As will be discussed later, maintaining the mobile device at approximately the target distance from the visible surface simultaneously controls the field of view of the coaxially positioned device camera.

[0281] When used by young children, certain patterns (e.g., smiles) may elicit positive or desirable responses that arise from or are reinforced by interactions within that environment (e.g., the sight of familiar shapes or faces). As a further example, eye-catching (e.g., including bright colors) and / or unattractive (e.g., angry faces or frightening shapes) beam patterns may attract particular attention from certain device users. Beam image patterns may include human faces projecting various expressions, animal faces (e.g., cute or ferocious appearances), animal forms, recognizable shapes, cartoon characters, toys, circles, rectangles, polygons, arrows, crosses, fingers, hands, letters, numbers, other symbols, emojis, etc. Furthermore, projected beam images can be animated, deformed (e.g., color, intensity, shape), projected as a series of images, and / or flash.

[0282] The display surface on which beam images are visualized is made of paper, cardboard, film, cloth, wood, plastic, or glass. The surface can be painted, printed, textured, reinforced with three-dimensional elements, and / or flexible. The visualization surface may also be any form of electronic display, including e-readers, tablets, signs, and mobile devices. The visualization surface may also be, for example, a component of a book (including cover, pamphlet, box, sign, newspaper, magazine, printable surface, tattoo, and display screen).

[0283] As an additional example, beam images are most easily perceived when the device user positions the handheld device perpendicular to the visible surface (i.e., in focus). At a position nearly perpendicular to the visible surface, the beam image best matches the structured light pattern within the handheld device's beam. If the projected beam is off-center from the normal to the visible surface or at an acute angle, the reflection of the beam image will appear distorted (e.g., stretched in the direction of the handheld device axis away from the surface normal), and this distortion may be amplified by surface curvature and defects.

[0284] In further embodiments of this specification, when a portable device is held near a predetermined working distance and pointed approximately perpendicular to the surface, the results include: 1) maintaining a predetermined target image area with a camera oriented in the same direction as the light beam; 2) avoiding image acquisition of reflective surfaces distorted by viewing from angles not approximately perpendicular to the surface; and / or 3) ensuring that the camera-acquired image of an object on the surface within the field of view is in focus by the camera optics (with depth of field including the target distance, which may be controlled by the device processor, for example, as described below).

[0285] The area captured within the image acquired by a camera is generally considered to be the camera's field of view. In camera systems without zoom capabilities (or other methods of adjusting the focusing optics), the field of view is primarily determined by the distance from the camera's optical sensor array to the visible surface. As a result, if, for example, the camera acquires an image of an area containing a small object located far from the mobile device (more specifically, the camera sensor array), the object may appear as a small dot (i.e., defined by an area with little or no structure and a low number of camera pixels). Conversely, if the device camera is too close to the visible surface, only a small portion of the object may be visible in the image acquired by the camera. Identifying such objects, or even determining the presence of a target object, poses a challenge for CV-based processing. These challenges can potentially be overcome by keeping the area covered by the camera image (i.e., the camera's field of view) at least approximately within a predetermined target image area (i.e., near the target working distance).

[0286] Even when applying CV to less extreme situations, limiting the size range of object profiles to a target range with a sufficient field of view to identify the overall shape of an object while ensuring enough pixels to identify object details can potentially improve object recognition accuracy and robustness within camera-acquired images. Limiting CV-based methods to a size range near the beam's focal length can simplify processes such as neural network training, object template definition, and other CV-based algorithmic approaches.

[0287] In further embodiments of this specification, the ability to change the focal length of the projected beam image allows a mobile device to influence the working distance. Specifically, this is achieved by prompting the device user to follow the change in beam focal length (i.e., making positional adjustments to keep the beam image in focus). As a result, the field of view and object size in the image acquired by the coaxial camera can be controlled (indirectly) through the control of the beam focal length.

[0288] In applications involving interaction with visible objects of varying sizes, adjusting the working distance by controlling the beam focal length can be beneficial. For example, in computer vision processing of images acquired by a camera (such as optical character recognition), shortening the working distance (resulting in a narrower camera field of view and capturing more pixels of the object) is effective when identifying characters printed in small fonts. Conversely, if a book contains large images with little detail (e.g., cartoon illustrations or characters), a longer working distance (i.e., beam focal length) may be beneficial for CV processing of images acquired by a camera.

[0289] Furthermore, when the beam's focal length is changed under the control of the device's processor, the area covered by the camera's field of view (e.g., the target image area) and consequently the size of objects within the camera's sensor array (i.e., the number of pixels affected by each object) can be estimated and / or known in CV-based methods for identifying such objects.

[0290] Within the compact beamforming optical configuration of a portable device, the focal length can be controlled by changing the distance the beam travels to the beamforming optical element, or by changing the beamforming optical element itself (e.g., changing the shape or position of the optical element). One or more of these strategies can be implemented at one or more locations along the beam path. The optical path can be modified by inserting and / or moving one or more reflective surfaces or refractive elements (e.g., to change the beam path).

[0291] The reflective surfaces are components of, for example, microelectromechanical systems (MEMS) devices, and one or more mirrors can be actuated to deflect the entire beam or beam elements (individually). Alternatively, or additionally, one or more polyhedral prisms (usually utilizing internal reflection) can be fixed to an actuator to alter the optical path. As a further example, one or more small deformable lenses can be included in the optical path. The small actuators that change the reflective and refractive elements may employ piezoelectric, electrostatic, and / or electromagnetic mechanisms actuarily coupled to a device processor.

[0292] Multiple reflective surfaces (e.g., 2 to 10 or more) can amplify the effect of minute movements of optically oriented elements that change the focal length. In micro-beamforming elements within handheld devices (e.g., optical paths of less than a few millimeters), changes in the optical path in the sub-millimeter range can change the focal length of the light beam by centimeters.

[0293] The focal length of the target beam may be dynamically changed based on, for example, the interactive visual environment, the type of interaction being performed, individual user preferences, and / or the characteristics of the displayable content (particularly the size and level of detail of selectable objects required for object identification). In a desktop environment where a user points to an object in a handheld book, a shorter beam focal length may be beneficial compared to, for example, pointing to an electronic screen on a large table. Rapid back-and-forth interactive sequences may also benefit from a shorter target beam distance. User settings may take into account an individual's (especially a child's) eyesight, age, and / or motor skills. If a device user selects different books or magazines with different font or image sizes, the beam focal length may be changed to accommodate such content variations (and to ensure comfortable viewing for the user).

[0294] In yet another example, the image size of a light beam may be statically set (e.g., using a light-shielding filter) or dynamically changed. Dynamic control of beam image size can generate light covering a wide range of sizes, from small focal beams (e.g., the size of alphanumeric characters) to areas larger than a book page, using techniques that similarly change the focal length (such as moving or inserting reflective or refractive optical elements). In the latter case, the light source of a portable device may exhibit functionality closer to a flashlight or torchlight compared to so-called "laser pointers." Similar to controlling the focal length, the beam image size can take into account the type of interaction being performed, visual acuity, age, and / or the cognitive abilities of the user (especially children).

[0295] As mentioned above, although not strictly "handheld," this device may be fixed or positioned (temporarily or for extended periods) on body parts other than the hand. For example, it can be attached to a device worn on the user's head. In this case, the emitted light beam not only indicates the direction of the coaxial device camera to the wearer, but can also notify others who can see the visible surface of the area where the center of attention of the visible content may be located. In some applications, a relatively large beam (e.g., comparable to the aforementioned flashlight or strobe) is effective in informing nearby users which page or area of ​​a page is being viewed, along with the information content contained within the beam (and further, prompting the device user to position the device at a target distance from the visible surface).

[0296] In additional examples herein, methods for generating structured light patterns that produce recognizable beam image reflections on a visible surface include generating patterns from multiple addressable (i.e., by a device processor) light sources or reflective surfaces, and / or blocking or deflecting selected light. Exemplary methods include: One or more light-shielding filters (e.g., films or masks) can be inserted into the light beam path to block light of all wavelengths or selected wavelengths. This method generally requires an optical element for projecting the beam and control of the focal plane for shielding and display.

[0297] Digital light processing (DLP) projector technology (e.g., Texas Instruments' DLP Pico system) combines light-emitting elements with movable micromirrors that control the light projection. The device processor controls the micromirror array, enabling dynamic display patterns.

[0298] Liquid crystal (LC) filters or projection arrays (e.g., those used in displays manufactured by Epson or Sony) can be components in the beam's optical path. Addressable LC arrays are controlled by a device processor, enabling dynamic display patterns. Multiple light beam sources (e.g., LEDs, including Mojo Vision's MicroLEDs) are each connected to a device processor, enabling the generation of dynamic beam images. LED arrays can include a relatively small number of light sources (e.g., a 5x7 grid capable of generating alphanumeric characters and symbols) to a very large number of light sources capable of generating image details, including complex animations.

[0299] The device light beam source includes one or more light-emitting diodes (including quantum dots, micro-LEDs, and / or organic LEDs) or laser diodes, and may be monochromatic or polychromatic. Light intensity (including switching the beam on / off) is controlled using multiple methods (e.g., by an operationally coupled device processor), such as controlling the light beam drive current or pulse width modulation of the drive current. Modulation of light intensity and color (e.g., by selecting from different light sources) may be used, for example, to attract the device user's attention, or to provide a "visual reward" (e.g., when a selection is made), if a user selection process is anticipated (e.g., by turning the beam on or on).

[0300] Beam intensity and / or wavelength may be modulated to take into account the environmental conditions of use. For example, when used in a dark environment, the beam intensity can be reduced to conserve power and / or not overwhelm the visual sensitivity of the mobile device user, as detected by the device's photodetector and / or by the ambient light region in the image acquired by the camera (e.g., the area not containing the beam image). Conversely, in bright or visually noisy environments (e.g., to improve visibility for the device user), the beam intensity can be increased.

[0301] Furthermore, it is possible to adjust the beam intensity and / or wavelength by considering the reflectivity (including color) of the visible surface in the direction of illumination. For example, a beam composed mainly of green wavelengths may not reflect well, or may even be difficult to see, when directed towards a drawing area where red pigment is mainly present. The overall color of the reflective surface is estimated from the image acquired by the camera from the neighboring area surrounding the beam image region. The beam color is then adjusted to make the reflection more clearly perceptible for both the device camera and the user. A similar technique can be used to compensate for individual visual acuity and beam intensity preferences.

[0302] As another example, when the visible surface consists of an electronic display (e.g., e-reader, tablet), the reflectivity of these surfaces is generally low, which can reduce the intensity of beam reflection (including reflection from subsurface structures). Dynamically adjusting the beam's color and / or intensity can help ensure that the beam image is strong enough to be visible to the device user.

[0303] Furthermore, the low reflectivity from the surface measured in the beam image region within the camera-acquired image can be used by the mobile device to determine the presence of some form of electronic display screen. Reduction or absence of the beam image can be linked to the determination of whether or not there is a dynamic change in individual objects within the visible surface. The movement or appearance change of a visible object that does not move simultaneously with the rest of the visible surface (e.g., showing movement of a handheld device camera and / or the entire visible surface) may help indicate the presence of an electronic display surface (e.g., one containing dynamic content).

[0304] In beam illumination systems with modifiable beam structure and light patterns (e.g., LC, DLP, LED arrays), the beam pattern can be dynamically changed. For example, it can be changed as a prompting mechanism during the selection process (e.g., facial expression changes, magnification of beam image components) and / or as a reward after selection (e.g., projection of a starburst pattern). Changes in the light pattern may be displayed in conjunction with dynamic changes in beam intensity and / or wavelength.

[0305] Changes in beam projection pattern, intensity, and / or color can also be used by handheld devices as feedback to the device user indicating whether the device is being held near the target distance from the reflective surface. If CV-based analysis of the beam projection area in the camera-acquired image determines that the beam image is out of focus (i.e., the handheld device is not being held at the target distance from the reflective surface), changes in beam pattern, intensity, or color can be used to prompt the user to adjust the device position.

[0306] Such cues could include, for example, removing details within the projected image so that the user can roughly visualize the outline of the beam pattern when not significantly out of focus, and / or altering the beam intensity and / or color to make it more recognizable even in a significantly out-of-focus beam image. Another example is a smile (i.e., one that is well-received by most viewers) being recognizable in the beam image when in focus, which then changes to a frown or a sad face (e.g., still recognizable) as the portable device moves away from the target distance.

[0307] Alternatively, or additionally, the mobile device's position feedback may include other modalities. The use of other feedback modalities is useful when the beam image is excessively out of focus (i.e., the beam image is unrecognizable) and / or when the terminal processor has little or no control over the beam image (e.g., when the beam is formed using a simple light-shielding filter in the optical path).

[0308] Positional feedback can be provided to the device user using audio, other visual sources, or haptic means. For example, if a handheld device has a built-in speaker that is operationally connected to the device processor, it can provide audio prompts or cues when the beam is determined to be in focus (in the image acquired by the camera) or, conversely, out of focus. A display or other light source (i.e., not related to the beam) operationally connected to the device processor can provide similar feedback by changing the display intensity, color, or content. Similarly, a haptic unit can be operationally connected to the device processor and activated when the light beam is out of focus (or, conversely, in focus).

[0309] In any of these warning or notification methods, the degree of focus can be reflected in the feedback. The amplitude, tone, and / or content (including words, phrases, or entire sentences) of sounds played from a speaker can reflect the degree of focus. Similarly, the brightness, color, and / or content of visual cues or prompts can be modulated based on the measured degree of focus. Likewise, the frequency, amplitude, and / or pattern of tactile stimuli can indicate to the device user how close the handheld device is to a target distance from a visible surface.

[0310] As mentioned above, the measurement of the degree of focus in a beam image region within an image acquired by a camera is used as the basis for user feedback (and / or the performance of one or more other actions). Numerous computer vision (CV)-based methods for determining the degree of focus in an entire image or a sub-region of an image are known in the art. On the simpler side of the spectrum of methods (e.g., the easier-to-implement side), an increase in focus is generally associated with an increase in contrast (especially if the image shape does not change). Therefore, the contrast measured within an image region can be used as an indicator of the degree of focus.

[0311] Furthermore, edge sharpness and the associated focal depth can be estimated using multiple kernel-based operators (two-dimensional in the case of images), including the Laplacian calculation. Transforming an image or image region into frequency space using the Fourier transform allows for the identification of in-focus areas within images with large high-frequency component amplitudes.

[0312] Artificial neural networks (ANNs) trained for focusing within images can also provide a measure of focus. When the image region is small and clearly defined (see, for example, Figure 19), and the image profile is based on a predefined small dataset of potential beam projection images, such trained networks can be trained relatively quickly and / or miniaturized (e.g., by reducing the number of nodes and / or layers).

[0313] In further embodiments of this specification, operational questions that may arise when performing one or more actions during an interactive sequence using a light beam directed towards a selectable region on a visible surface include: how to inform the user of the boundaries of the selectable region, when to process the camera-acquired images to identify the content within that region, and when to perform the resulting actions generated by the device. For example, if the device processor applies CV analysis to all camera-acquired images available on the mobile device, the user is likely to be quickly overwhelmed by unintended actions and unable to effectively communicate their intentions.

[0314] One possible solution to these challenges is to inform the user of the boundaries of each area related to selectable actions. When mobile devices "bring" book or magazine pages to life, the size and number of objects contained within a particular area may not be clear. This is because the conditions for areas that generate actions upon selection are unclear. For example, within a page area containing text, individual characters may, in some cases (e.g., during spell learning), become selectable or manipulable objects. In other cases, words, phrases, sentences, paragraphs, comic book bubbles, explanatory text, and associated graphics may each be contained within separate areas that generate actions upon selection (i.e., in a way that is not apparent to the device user).

[0315] One way to provide feedback to the user regarding selected areas or operable boundaries is to provide one or more audible prompts when moving between areas or when temporarily placing the cursor over an operable area. To allow users to easily point to multiple operable areas in a short time or quickly move the beam, audible prompts and prompts (e.g., ping, ding, chime, click, clang) are usually short (e.g., less than 1 second). Furthermore, if multiple operable areas are detected consecutively, continuous prompts may be immediately stopped, and / or some or all of the pending prompts in the queue may be removed. This prevents the device user from becoming overwhelmed with information.

[0316] Depending on the identification information or classification within a "flyover" or "hover" region, it is possible to generate different sounds and prompts. For example, a region that, when selected, triggers an interactive query from a mobile device may generate a more attention-grabbing "question-type" prompt. On the other hand, a region that simply results in an interactive exchange (without a question) may generate a different (e.g., less attention-grabbing) informational prompt. In yet another example, a short sound effect associated with the identified object may be used as a prompt. For example, a prompt for a region containing a cat would consist of a "meow." More generally, prompting a region pointed to by a light beam may depend on predefined page layout characteristics of the region and / or objects (e.g., identified by computer vision techniques) within the camera-acquired image of that region.

[0317] Alternatively, one or more visual cues can be provided using either or both display elements on the mobile device (e.g., indicator LEDs, orbs, display panels) and / or a pointing beam. Similar to the audible prompts described above, the visual prompts displayed when scanning the beam across different operable areas are designed to be concise and not overwhelming to the device user. For example, visual prompts may consist of changes in brightness (including flashing), color, and / or pattern of the displayed content. Similar to the acoustic methods described above, visual prompts may rely on information related to the page layout of a particular area and / or one or more objects determined within the camera-captured image of that area.

[0318] Device users with hearing impairments can utilize visual prompts. Conversely, individuals with visual impairments may prefer auditory cues. Alternatively, haptic prompts and cues can be generated via a haptic unit operationally coupled to the device processor. A haptic prompt is felt in the device user's hand as the beam moves into or lingers in an operable area. The duration and / or intensity of the haptic cue may depend on the page layout information associated with the area and / or one or more objects identified in the camera-acquired image. For example, if the beam is pointing to a specific area, a long-lasting (e.g., more than one second) and strong haptic prompt may be executed, as selecting that area may generate a query to the device user.

[0319] Once the user recognizes that the beam is directed towards an operable area (e.g., by auditory, visual, or tactile means), the user can indicate (i.e., to the mobile device) that a selection is being made. Such indication is made via a switch component (e.g., a push button, toggle, contact sensor, proximity sensor) operationally connected to the device processor and controlled by the device user. Similarly, a selection can be indicated using voice control (e.g., a specific word, phrase, or sound) sensed by a microphone operationally connected to the device processor. Alternatively, the movement (e.g., shaking) and / or orientation (e.g., orientation relative to Earth's gravity) of the mobile device while it is being operated by the device user can be sensed by an IMU operationally connected to the device processor and used to indicate that a selection is being made.

[0320] Furthermore, the method of instruction (e.g., voice or switch use) may be linked to one or more of the aforementioned audible, visual, and / or haptic prompts. During the selection process, prompts are used to inform the user of when instructions are expected and the modality of those instructions. This interactive strategy helps avoid generating unnecessary instructions at unexpected times, generating instructions using the expected modality (e.g., voice versus switch), and providing additional input (i.e., input to the mobile device processor) during the selection process. As an example of the latter, a visual prompt can be provided when a displayed color corresponds to one of several different colored push-button switches (e.g., the contact surfaces may also differ in size and texture). Pressing a push-button of a similar color to the displayed color indicates agreement with the interactive content, while pressing another button may indicate disagreement, an alternative selection, or an unexpected response from the device user.

[0321] Figure 19 is an illustrative diagram (including elements similar to a ray diagram) showing how a structured light pattern is formed using a light-blocking filter 192 and an in-focus reflected beam image 193b generated at a preferred or target working distance between the portable device beam source 195 and the visible surface 196b that reflects the beam. The in-focus reflection encourages the device user (especially if preferred by the user) to operate the handheld device at the target working distance. At distances shorter than the desired working distance 196b (e.g., 196a) or longer (e.g., 196c), a blurred image (i.e., out of focus) may be observed (e.g., 193a and 193c), suppressing the operation and / or movement of the portable device in these areas.

[0322] In this example, a structured light pattern is formed using a light source (e.g., 190 LEDs), passed through one or more optical elements 191a, and then the light pattern is structured by a shielding filter 192. Depending on the imaging characteristics of the light source (e.g., a directional light source array, and / or whether the light source includes a parallel light optical system), some beamforming configurations may not require the optical elements shown in 191a.

[0323] The second group of optical elements 191b projects a structured light pattern as a beam (the ray elements are shown in 194a and 194b). The beam image, a smiley face 193b, appears in focus at the target working distance 196b from the device light source 195. If the handheld device is too close to the reflective surface 196c, the beam image will be out of focus and will appear as 193a. If the beam is configured to diffuse slightly at a typical working distance (e.g., up to 1 meter), the out-of-focus reflected image may appear smaller (compared to the focused beam image or the beam image at a greater distance). Similarly, if the handheld device is too far from the reflective surface at position 196c, the beam image will appear out of focus to the device user (193c) (and will appear larger in the case of a divergent beam).

[0324] Optionally, the optical element may include one or more movable reflective surfaces (e.g., 197) operationally coupled to the device processor. Such surfaces are components of, for example, MEMS devices and / or polyhedral prisms and can change the optical path distance from the light source 190 to the focusing and / or parallelizing optical system 191b. The optical components shown in 191b may include one or more optical elements that introduce and / or deduct light to the reflective surface region 197.

[0325] In the exemplary beam image pattern shown in Figure 19, the bright (i.e., illuminated) elements of the beam image are depicted as dark elements 193b (i.e., on a white background) to illustrate the overall circular beam structure. Similarly, elements that allow light to pass through within the light-shielding filter 192 are shown in white (again, to visualize the circular pattern of the entire beam). In general, light-shielding elements that block structural light create a wide range of image complexity, as long as they block specific wavelengths or all wavelengths (e.g., using films or LCD filters) and do not violate the diffraction limit of light (at specific wavelengths). Systems that directly generate structural light patterns (e.g., LEDs or micro-LED arrays, DLP projections) can also produce beam images with similar complexity and optical constraints.

[0326] Figure 20 is a flowchart illustrating exemplary steps for calculating the beam focal position from beam-directing regions in one or more camera-acquired images 203a, in addition to the procedure for the device user to visually observe the beam reflection and determine the focal position (i.e., near the desired distance). A mechanical evaluation of whether the handheld device is at the desired distance from a visible surface can be performed periodically or continuously to provide positional feedback to the user and to determine whether the user is making a valid beam-directing selection (i.e., within the field of view of the target camera) (206). Feedback to the user indicating whether the handheld device is being held near the desired distance may include visual 207a, auditory 207b, and / or tactile prompts. Exemplary steps of this process are as follows: In 200a, the user directs a structured light beam 201a emitted from a mobile device 201d towards a selected object on the visible surface (in this example, a cat 201c) to generate a beam reflection 201b; In 200b, a camera (not shown) and a focusing optical system 202c (illustrated separately from the mobile terminal body 202d for illustrative purposes) are used to capture an image in which the camera's field of view 202a includes the beam reflection 202b; In 200c, based on recognizing coaxially positioned beams and beam-direction regions in the camera image, and / or optionally using computer vision techniques to compare one or more regions within the camera field of view (e.g., 203b) that may contain beam images (indicated by dashed frames) with a predetermined template (203c) of a focused beam image; In 200d, the device processor measures the focal point in the beam-directing region of the camera-acquired image using kernel operators (shown in 204) and / or additional computer vision techniques to determine the focal point; In 200e, it is determined whether the beam image has sufficient depth of field to fit within the desired distance range. If so, processing of the camera-acquired image continues in 205b; otherwise (i.e., 205a), the user is allowed to continue operating the mobile device in 200a. In 200f, continuous image processing may include the identification of text, symbols, and / or one or more objects in the vicinity of the beam (which is determined to be in focus). Here, the morphology of the feline depicted in 206 is identified. In the 200g, optionally (indicated by a dashed border), the device user is visually notified (e.g., via one or more device displays on the 207a and / or by adjusting the beam intensity (including turning it off) as shown) that the handheld device is properly positioned and selected, in a manner that triggers one or more actions by the device processor, including visual (e.g., audio and / or haptic feedback using the device speaker on the 207a).

[0327] The above-described embodiments are presented for illustrative purposes only. They are not intended to be exhaustive and do not limit the invention to the exact forms disclosed. Many variations and modifications of the embodiments described herein will be apparent to those skilled in the art in light of the above disclosures. It will be understood that various components and features described in a particular embodiment may be added, deleted, and / or replaced in other embodiments depending on the intended use of that embodiment.

[0328] Furthermore, when describing representative examples, the specification may present a method and / or process as a specific sequence of steps. However, unless the method or process depends on the specific sequence of steps described herein, the method or process should not be limited to that specific sequence of steps. As those skilled in the art will understand, other sequences of steps are also possible. Therefore, the specific sequence of steps described in the specification should not be construed as a limitation on the claims.

[0329] While various modifications and alternative forms are possible with respect to the present invention, specific examples are shown in the drawings and described in detail herein. It should be understood that the present invention is not limited to any particular form or method disclosed, but encompasses all modifications, equivalents, and alternatives that fall within the scope of the accompanying claims.

Claims

1. A method for a human to point to a detectable object using a mobile device, the mobile device comprising a terminal processor, a terminal light beam source configured to generate a light beam that produces one or more light beam reflections from one or more visible objects that are visible to a human, a terminal camera arranged such that the camera field of view includes the beam position regions of one or more light beam reflections and operationally connected to the terminal processor, and a terminal speaker operationally connected to the terminal processor, the method being Playing one or more audible cues related to a detectable object via the device's speaker; A person operates a mobile device, and while the emitted light beam is directed from the device's light beam source towards one or more visible objects, a camera image is captured by the device's camera; The device processor separates one or more indicated objects in the beam position region within the camera image; and A method comprising determining, by a device processor, whether one or more indicated objects match a predetermined template of detectable objects.

2. The method according to claim 1, wherein the apparatus light beam source is either a light-emitting diode or a laser diode.

3. The method according to claim 1, wherein the device light beam source is operationally connected to a device processor, and the brightness of the device light beam source is controlled by one or more of the following: adjusting the magnitude of the light beam drive current and pulse width modulation of the light beam drive current.

4. A method according to claim 1, wherein the projected light beam is one or more of the following: parallel, non-coherent, divergent, and patterned.

5. A method according to claim 1, wherein the apparatus light beam source is operably connected to the apparatus processor, and the method further includes turning off the apparatus light beam source when either or both of the following occur: acquisition of a camera image and / or determination of a match between a detectable object and a predetermined template.

6. In the method of claim 1, the beam position region in the camera image is: The device camera acquires a baseline image that does not include reflections from the projected light beam; The device camera acquires a light beam reflection image that includes one or more reflections generated by the projected light beam; The device processor calculates a differential pixel intensity image obtained by at least partially subtracting the baseline image from the light beam reflection image; and A method determined by assigning light beam sensing pixels that exceed a predetermined light intensity threshold in the subtracted pixel intensity image to a beam position region using an instrument processor.

7. A method according to claim 1, wherein one or more audible cues relating to a detectable object include: one or more sounds produced by the detectable object, one or more names of the detectable object, one or more questions relating to the detectable object, one or more descriptions of the detectable object, one or more functional descriptions of the detectable object, a mathematical problem relating to the detectable object, a musical score relating to the detectable object, and one or more related object descriptions relating to the detectable object.

8. The method according to claim 1, wherein the predetermined template of the discoverable object includes one or more of: one or more shapes of the discoverable object, one or more sizes of the discoverable object, one or more colors of the discoverable object, one or more textures of the discoverable object, and one or more patterns within the discoverable object.

9. The method according to claim 1, wherein the detectable object is printed material in any of the following: a book, a book cover, a pamphlet, a box, a sign, a newspaper, or a magazine.

10. A method according to claim 1, further comprising performing an operation by a device processor, at least in part on determining whether a detectable object matches or does not match a predetermined template.

11. In the method according to claim 10, the operation is: Transmitting to one or more remote processors one or more audible cues, camera images, camera image acquisition times, predetermined camera image light beam direction regions, predetermined templates for detectable objects, and one or more of the following: Playing one or more sounds through the device's speaker; Displaying one or more lighting patterns on one or more device displays operationally connected to a device processor; and A method comprising one or more of the following: activating a device haptic unit operationally connected to a device processor.

12. The method according to claim 1, wherein a device switch is operably connected to a device processor, and further includes determining whether the device switch matches a predetermined template of a detectable object when the device processor detects a change in the switch state of the device switch.

13. A method according to claim 1, further comprising a device microphone operationally connected to a device processor, and further comprising determining a match with a predetermined template of a detectable object that occurs when the device processor identifies one or more human-generated identifiable sounds in data acquired by the device microphone.

14. A method according to claim 1, wherein a device inertia measurement unit is operably connected to a device processor, and further includes determining a match with a predetermined template of a detectable object when the device processor identifies either a predetermined handheld device gesture operation or a predetermined handheld device orientation in the data acquired by the device inertia measurement unit.

15. A method for indicating an object detectable by a human using a handheld device, the method comprising a handheld device including a handheld device processor, a handheld device light beam source configured to generate a light beam that generates one or more light beam reflections from one or more visible objects that are visible to a human, a handheld device camera (positioned so that the camera field of view includes the beam position regions of one or more light beam reflections and operationally connected to the handheld device processor), and one or more device displays operationally connected to a device processor, the method being: Displaying one or more visual cues related to a detectable object on one or more device displays; The device camera acquires camera images while the handheld device is operated by a human and a projected light beam is directed from the device's light beam source toward one or more visible objects; The device processor separates one or more indicated objects in the beam position region within the camera image; and A method comprising determining, by a device processor, whether one or more of the objects found match a predetermined template of the objects.

16. A method according to claim 15, further comprising performing an action by a device processor based on a determination of whether a detectable object matches or does not match a predetermined template.

17. In the method according to claim 16, the operation is: Transmitting to one or more remote processors one or more visual cues, camera images, camera image acquisition times, predetermined camera image light beam direction regions, predetermined templates for discoverable objects, and one or more designated objects; Playing one or more sounds through a device speaker operationally connected to a device processor; Displaying one or more lighting patterns on one or more device displays; and A method comprising one or more of the following: activating a device haptic unit operationally connected to a device processor.

18. A method for indicating a human-detectable object using a portable device comprising a device processor, a device light beam source configured to generate a light beam that produces one or more light beam reflections from one or more human-visible objects, a device camera operationally connected to the device processor and positioned such that the camera field of view includes the beam position regions of one or more light beam reflections, and a device haptic unit operationally connected to the device processor, the method being: The device haptic unit generates sensed tactile vibrations in one or more actions associated with detectable objects and one or more tactile frequencies associated with sound; When a handheld device is operated and a projected light beam is directed from the device's light beam source toward one or more visible objects, a camera image is captured by the device's camera; The device processor separates one or more indicated objects in the beam position region within the camera image; and A method comprising determining, by a device processor, whether one or more of a specified object matches a given template of discoverable objects.

19. A method according to claim 18, further comprising performing an action by a device processor based on a determination of whether a detectable object matches or does not match a predetermined template.

20. In the method according to claim 19, the operation is: Transmitting to one or more remote processors any of the following: one or more haptic frequencies, a camera image, the acquisition time of the camera image, a predetermined camera image light beam directional region, a predetermined template of a discoverable object, and one or more specified objects; Playing one or more sounds through a device speaker operationally connected to a device processor; Displaying one or more lighting patterns on one or more device displays operationally connected to a device processor; and A method comprising one or more of the following: activating a device haptic unit operationally connected to a device processor.

21. A method for performing an operation based on a specified page position selected by a human using a handheld device operationally connected to the device processor, comprising: a device processor; a device light source configured to generate a projection light beam that produces one or more light reflections from one or more visible objects visible to a human; and a device camera having a camera field of view arranged to include the reflection positions of one or more light reflections, the method being: The device processor retrieves one or more predetermined interactive page layouts; The device camera captures a camera image while the handheld device is being operated and a projection light beam is directed from the device light source towards a specified page position; The device processor calculates a positioned image based on the matching of the camera image with one or more predetermined interactive page layouts; The device processor identifies a specified page position based on the reflection position in the positioned image; and A method comprising either or both a device processor and a remotely connected processor performing an action based on a specified page position within one or more predetermined interactive page layouts.

22. The method according to claim 21, wherein the device light source is either a light-emitting diode or a laser diode.

23. The method according to claim 21, wherein the device light source is operably connected to a device processor, and further comprises controlling the intensity of the device light source by adjusting the light source drive current and pulse width modulation of the light source drive current, one or more of the above.

24. A method according to claim 21, wherein the projected light beam is any of the following: parallel, non-coherent, divergent, or patterned.

25. A method according to claim 21, wherein one or more predetermined interactive page layouts include one or more of the following: the position of one or more page objects, the shape of one or more page objects, a camera image of one or more page objects, the size of one or more page objects, the color of one or more page objects, the texture of one or more page objects, a pattern within one or more page objects, one or more sounds normally produced by one or more page objects, the name of one or more page objects, a description of one or more page objects, the function of one or more page objects, one or more questions related to one or more page objects, and one or more related objects related to one or more page objects.

26. A method according to claim 21, wherein one or more predetermined interactive page layouts are displayed on one or more of the following: one or more books, book covers, brochures, boxes, signs, newspapers, magazines, posters, tablets, printable surfaces, painted surfaces, textured surfaces, flexible surfaces, and reinforced surfaces with one or more three-dimensional elements, signs, tattoos, mobile devices, tablets, e-readers, televisions, and display screens.

27. A method according to claim 21, wherein calculating a placed image includes calculating the matching position between a camera image and one or more predetermined interactive page layouts using one or more of the following: template matching, computer vision, machine learning, transformer models, and neural network classification.

28. A method according to claim 27, wherein either or both machine learning and neural network classification are trained based on one or more predetermined interactive page layouts.

29. A method according to claim 21, wherein the step of calculating a positioned image includes determining the horizontal position of the camera image, the vertical position of the camera image, the magnification of the camera image, and the orientation of the camera image, which generate a matching position between the camera image and one or more components of a predetermined interactive page layout.

30. In the method of claim 21, a predetermined beam-directing position in the positioned image is: The device camera acquires a baseline image that does not include reflections from the projected light beam; The device camera acquires a light beam reflection image that includes one or more reflections generated by the projected light beam; The device processor calculates a difference pixel intensity image by subtracting the baseline image from the light beam reflection image; and A method determined by assigning the center position of a light beam sensing pixel that exceeds a predetermined light intensity threshold in a subtracted pixel intensity image to a predetermined beam directional position, using an instrument processor.

31. The method of claim 21, wherein the light source is operably connected to the device processor, and further includes one or more of the following: turning on the light source when a predetermined page layout is acquired; turning off the light source before acquiring a camera image; and turning off the light source when acquiring a camera image.

32. In the method of claim 21, the operation is: Transmitting one or more of the following to one or more remote processors: camera image, camera image acquisition time, predetermined beam irradiation position, one or more predetermined interactive page layouts, and specified page positions; Playing one or more sounds through a device speaker operationally connected to a device processor; Displaying one or more lighting patterns on one or more device displays operationally connected to a device processor; and A method comprising one or more actions of activating a device haptic unit operationally connected to a device processor.

33. A method for performing an action based on a specified page position in a specific context selected by a human, using a portable device including a device processor, a device light source configured to generate a projected light beam that produces one or more light reflections from one or more visible objects visible to a human, and a device camera operationally connected to the device processor and aligned so that the camera field of view includes the reflection positions of one or more light reflections, the method being: The device processor obtains one or more predetermined context templates; In a state where a mobile device is operated and a light beam emitted from the device's light source is directed toward a contextual object, a first camera image is acquired by the device's camera; The device processor classifies the identified context based on the match between the first camera image and one or more predetermined context templates; The device processor retrieves one or more predetermined interactive page layouts associated with an identified context; When a mobile device is operated to direct the light beam emitted from the device's light source to a specified page position, a second camera image is captured by the device's camera; The device processor calculates a positioning image based on the matching of the second camera image with one or more predetermined interactive page layouts; Based on the reflection position in the positioning image, the device processor determines the specified page position; and A method comprising one or more actions, in which either or both a device processor and / or a remote connection processor perform actions based at least partially on a specified page position in one or more predetermined interactive page layouts.

34. The method of claim 33, wherein the calculation of the identified context further includes, by the device processor, separating the first camera image into the region of the first camera image pointed to by the projected light beam.

35. The method of claim 33, wherein one or more predetermined context templates include attributes relating to one or more book covers, magazine covers, book chapters, pages, toys, people, anatomical elements, animals, photographs, drawings, words, phrases, household items, classroom supplies, tools, automobiles, and clothing.

36. In the method of claim 33, the operation is: Transmitting one or more predetermined context templates, a first camera image, the acquisition time of the first camera image, a second camera image, the acquisition time of the second camera image, one or more predetermined interactive page layouts, a positioned image, a predetermined beam-pointing position, and a specified page position to one or more remote processors; Playing one or more sounds in a device speaker that is operablely connected to a device processor; Displaying one or more lighting patterns on one or more device displays operationally connected to a device processor; and A method comprising one or more actions of activating a device haptic unit operationally connected to a device processor.

37. A method for performing an action based on a specified page position selected by a human, using a portable device including a device processor, a device light source configured to generate a projected light beam that produces one or more light reflections from one or more visible objects visible to a human, and a device camera operationally connected to the device processor and aligned so that the camera field of view includes the reflection positions of one or more light beam reflections, the method being: The device processor retrieves one or more predetermined interactive page layouts; When a mobile device is operated and a light beam emitted from the device's light source is directed towards a specified page position, the device's camera acquires two or more camera images; The device processor determines that the image movement measured by two or more camera images is less than a predetermined movement threshold, and that the two or more camera images were acquired for an acquisition time exceeding a predetermined dwell time threshold; The device processor calculates a positioning image based on the matching of the camera image with one or more predetermined interactive page layouts; The device processor determines the specified page position based on the reflection position in the positioned image; and A method comprising either or both a device processor and / or a remote connection processor performing an action based on a specified page position in one or more predetermined interactive page layouts, at least in part.

38. A method according to claim 37, wherein one or more spatial offsets are calculated by comparing a pair of consecutively acquired camera images using one or more of template matching, computer vision, and neural network classification, and the image movement is calculated from the sum of the one or more spatial offsets.

39. The method of claim 37, wherein the light source is operably connected to a device processor, and further includes turning on the light source when a predetermined page layout is acquired, turning off the light source before acquiring each of two or more camera images, turning on the light source after acquiring each of two or more camera images, and turning off the light source when it is determined that the amount of image movement is less than a predetermined amount of movement threshold and the acquisition time exceeds a predetermined dwell time threshold.

40. In the method according to claim 37, the operation is: Transmitting one or more camera images from two or more cameras, image movement, a predetermined dwell time threshold, acquisition time for two or more camera images, a predetermined beam pointing position, one or more predetermined interactive page layouts, and a specified page position to one or more remote processors; Playing one or more sounds through a device speaker operationally connected to a device processor; Displaying one or more lighting patterns on one or more device displays operationally connected to a device processor; and A method comprising one or more actions of activating a device haptic unit operationally connected to a device processor.

41. It is a portable device: A device body configured to be operated by the user's hand; Electronic circuits, including the device processor, located within the device body; A device light beam source configured to emit light from the device body and generate a beam image reflection on a visible surface, wherein the beam image reflection is focused at a predetermined distance from the visible surface; and The portable device includes a device camera that is operationally connected to a device processor and is aligned such that the field of view of the device camera includes beam image reflection, wherein the portable device includes: To generate a beam image reflection on a visible surface so that it is in focus at a predetermined distance, prompting the device user to position the handheld device at a predetermined distance from the visible surface; and A portable device configured to perform the following actions to acquire camera images.

42. A portable device according to claim 41, wherein the field of view of the device camera is configured to provide a target image area on a display surface at a predetermined distance.

43. A portable device according to claim 41, wherein the device light beam light source is one or more light-emitting diodes, one or more micro-light-emitting diodes, one or more laser diodes, and one or more digital light processing projector elements, each of which is operably connected to the device processor.

44. A portable device according to claim 41, wherein the beam image reflection is recognized to the device user as one of the following: a smiley face, a human face, an animal face, a toy, a cartoon character, a circle, a rectangle, a polygon, an arrow, a cross, a finger, a hand, one or more letters, one or more numbers, one or more symbols, one or more recognizable shapes, and an animation.

45. A portable device according to claim 41, wherein the structured light pattern that generates beam image reflection is generated by any one of one or more light-shielding filters in the optical path of the device light beam source; liquid crystal filter elements in the optical path of the light beam source (each operationally connected to the device processor); digital light processing projector elements (each operationally connected to the device processor); and a plurality of light beam sources (each operationally connected to the device processor).

46. A portable device according to claim 45, wherein the device processor modifies at least partially the structured light pattern in a camera image based on content determined by the device processor.

47. A portable device according to claim 41, wherein the light beam generator is operably connected to the device processor and is further configured to control the light beam intensity by either or both of the following: adjustment of the light beam drive current and / or pulse width modulation of the light beam drive current.

48. A portable device according to claim 47, wherein the device processor is further configured to control the light beam intensity based on the detected light intensity measured in either or both of the camera images of the beam image reflection region and the ambient light image region that does not include the beam image reflection.

49. A method for prompting a device user to position a portable device at a predetermined distance from a visible surface, wherein the portable device includes a device body configured to be operated by the device user's hand, and comprising electronic circuitry located within the device body, including a device processor; a device light beam source configured to produce a beam image reflection focused at a predetermined distance; and a device camera operably connected to the device processor and aligned such that the device camera's field of view includes the beam image reflection, wherein the method: The device uses a light beam source to generate a beam image reflection on a visible surface that is focused at a predetermined distance, prompting the device user to position the handheld device on the visible surface at a predetermined distance; and A method that includes taking a camera image using a device camera.

50. A method according to claim 49, further comprising: a device processor determining whether the beam state of a beam image reflection in a camera image is in-focus or out-of-focus; and at least partially performing an operation based on the beam state by the device processor.

51. A method according to claim 50, wherein the determination of the beam state comprises the device processor computing one or more of the following for a camera image: a trained neural network, one or more kernel operations, one or more image contrast measurements, and a Fourier transform.

52. A method according to claim 50, wherein the optical beam source is operably connected to an apparatus processor, the operation of which includes changing one or more of the beam intensity, beam color, and beam-structured optical pattern.

53. A method according to claim 50, wherein a haptic unit in a portable device is operably connected to a device processor, and the operation includes activating the haptic unit.

54. A method according to claim 50, wherein a device speaker in a mobile terminal is operationally connected to a device processor, and the operation includes playing an instruction sound on the device speaker.

55. The method according to claim 49, wherein the visible surface comprises one or more of the following: paper, corrugated cardboard, film, cloth, wood, plastic, glass, painted surface, printed surface, textured surface, reinforced surface with three-dimensional elements, flexible surface, and electronic display.

56. The method of claim 49, wherein the visible surface is one or more components of a book, book cover, pamphlet, box, sign, newspaper, magazine, tablet, printable surface, tattoo, mobile device, e-reader, and display screen.

57. A method for prompting a device user to position a portable device at a predetermined distance from a visible surface, wherein the portable device comprises a device body configured to be operated by the device user's hand, and electronic circuitry located within the device body, including a device processor; a device light beam source configured to generate beam image reflections on a displayable surface; one or more movable reflective surfaces located in the device beam path emitted from the device light beam source and each operationally connected to the device processor; and a device camera operationally connected to the device processor and aligned such that the device camera's field of view includes the beam image reflections, wherein the method includes: To generate beam image reflections on a visible surface using a device light beam source; The device processor moves one or more movable reflective surfaces to focus the beam image reflection at a predetermined distance, thereby prompting the device user to position the handheld device at a predetermined distance from the visible plane; and A method that includes acquiring a camera image using a device camera.

58. The method according to claim 57, wherein one or more movable reflective surfaces each include one or more movable mirrors, one or more microelectromechanical systems, and one or more movable prisms.

59. A method according to claim 57, comprising changing either or both the optical beam path distance and / or the optical path direction followed by a device optical beam by moving one or more movable reflective surfaces.

60. A method according to claim 57, further comprising moving one or more movable reflective surfaces by an apparatus processor to generate a beam image reflection focused at the updated distance.