Systems and methods to acquire complete viewable scenes using a portable interactive device
A portable device with a light source and camera provides intuitive interaction with printed content by ensuring comprehensive imaging and feedback, addressing the lack of simple methods for engaging with books and objects, enhancing the reading experience with added content and feedback.
Patent Information
- Application Number
- PCT/US2025/037590
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-15
- Filing Date
- 2025-07-14
- Publication Date
- 2026-01-22
AI Technical Summary
Existing technologies lack intuitive and efficient methods for individuals, particularly children, to interact with printed content in books and other viewable objects using portable devices, requiring precision manual dexterity and understanding of screen-based interactions.
A portable device with a light source and co-aligned camera captures images of viewable surfaces, providing visual, audio, and haptic feedback to ensure comprehensive imaging and interactive engagement with the content, allowing for intuitive interaction without complex manual dexterity.
Enables comprehensive imaging and interactive engagement with printed content, enhancing the reading experience by adding sounds, narratives, and real-time feedback, maintaining emotional engagement and facilitating machine-based guidance.
Smart Images

Figure US2025037590_22012026_PF_FP_ABST
Abstract
Description
[0001] SYSTEMS AND METHODS TO ACQUIRE COMPLETE VIEWABLE SCENES USING A PORTABLE INTERACTIVE DEVICE
[0002] RELATED APPLICATION DATA
[0003] The present application claims benefit of and priority to co-pending U.S. provisional application Serial No. 63 / 671,656, filed July 15, 2024, the entire disclosure of which is expressly incorporated by reference herein.
[0004] TECHNICAL FIELD
[0005] The present application relates generally to devices and methods for an individual to acquire the contents of a book (or other collection of viewable objects) by pointing a light emitted by a portable device toward a book page (or other collection of objects) and using a device camera, co-aligned with the projected light to acquire a complete image of each page region, page, and / or sequential pages. The device utilizes simple interactive signaling that lacks requirements for precision manual dexterity and / or understanding screen-based interactive sequences. Devices and methods herein employ techniques within the fields of mechanical design, electronic design, firmware design, computer programming, optics, computer vision (CV), ergonometric (including child safe) construction, human motor control and human-machine interaction. Devices and methods may provide an individual with means to acquire the contents of a book (or other collection of viewable objects) and to subsequently allow a child or learner, to interact with book contents using the same portable (e.g., handheld) device.
[0006] BACKGROUND
[0007] In recent years, the world has become increasingly reliant on portable electronic devices that have become more powerful, sophisticated and useful to a wide range of users. The devices and methods disclosed herein make use of advances in the fields of optics that include visible light (i.e., frequently referred to as “laser”) pointers and cameras, mobile sound and / or vibration generation (employing miniature coils, piezoelectric elements and / or haptic units), and telecommunications.
[0008] The beam of a visible light pointer (also referred to as a “laser pen”), typically used within business and educational environments, is often generated by a lasing diode with undoped intrinsic (I) semiconductor between p (P) and n (N) type semiconductor regions (i.e., a PEST diode). Within prescribed power levels and when properly operated, such coherent and collimated light sources are generally considered safe. Additionally, if directed at an eye, the corneal reflex (also known as the blink or eyelid reflex) ensures an involuntary aversion to bright light (and foreign bodies).
[0009] However, further eye safety may be attained using a non-coherent, light-emitting diode (LED) source. Such non-coherent sources (so-called “point-source” LEDs) may be collimated using precision (e.g., including so-called “pre-collimating”) optics to produce a light beam with minimal and / or controlled divergence. Point-source LEDs may, if desired, also generate a beam composed of a range of spectral frequencies (i.e., compared with the predominantly monochromatic light produced by a single laser).
[0010] Miniature, light-weight cameras are commonplace in modern mobile devices. The majority of cameras used in mobile settings use either CMOS (complementary metal-oxide- semi conductor) or CCD (charged-coupled device) light-sensing arrays. These different approaches may each be configured to produce advantages related to size, circuit integration complexity, cost, light-sensitivity, dynamic range, temperature stability, noise, and power consumption. Camera modules may additionally include optics, and control and read-out circuitry.
[0011] Speakers associated with televisions, theaters and other stationary venues generally employ one or more electromagnetic moving coils. Within handheld and / or mobile devices, the vibrations of a miniature speaker may be produced using similar electromagnetic coil approaches and / or piezoelectric (sometimes referred to as “buzzer”) designs. Vibrations (e.g., particularly those associated with alerts) may also be generated by a haptic unit (also known as kinesthetic communication). Haptic units generally employ an eccentric (i.e., unbalanced) rotating mass or piezoelectric actuator to produce vibrations (particularly at the low end of the audio spectrum) that can be heard and / or felt.
[0012] Advances in both electronics (i.e., hardware), standardized communications protocols and allocation of dedicated frequencies within the electromagnetic spectrum have led to the development of a wide array of portable devices with abilities to wirelessly communicate with other, nearby devices as well as large-scale communications devices including the World Wide Web and the metaverse. Considerations for which protocols (or combinations of available protocols) to employ within such portable devices include power consumption, communication range (e.g., from a few centimeters to hundreds of meters and beyond), and available bandwidth. Currently, Wi-Fi (e.g., based on the IEEE 802.11 family of standards) and Bluetooth (managed by the Bluetooth Special Interest Group) are used within many portable devices. Less common and / or older communications protocols within portable devices in household settings include Zigbee, Zwave, and cellular- or mobile phone-based networks. In general (i.e., with many exceptions, particularly considering newer standards), compared with Bluetooth, Wi-Fi offers a greater range, greater bandwidth and a more direct pathway to the internet. On the other hand, Bluetooth, including Bluetooth Low Energy (BLE), offers lower power, a shorter operational range (that may be advantageous in some applications), and less complex circuitry to support communications.
[0013] Advances in miniaturization, reduced power consumption and increased sophistication of electronics, including those applied to displays, micro-electromechanical devices (MEMS) including inertial measurement units (IMUs), and telecommunications have revolutionized the mobile device industry. Such portable devices have become increasingly sophisticated, allowing users to concurrently communicate, interact, geolocate, monitor exercise, track health, be warned of hazards, capture videos, perform financial transactions, and so on. Devices and methods that facilitate simple and intuitive methods to scan the contents of a book for subsequent interactions related to the book content may be useful.
[0014] SUMMARY
[0015] In view of the foregoing, devices and methods are provided herein that describe a light-weight, simple-to-use and intuitive portable (e.g., handheld) device that may augment the experience of reading a book and / or exploring objects in the environment of the device user. The device may be particularly well-suited for machine-based interactions by a child or other learner by bringing a book or other objects “to life” by adding sounds, supplemental content, questioning, device-based interaction, incorporating user responses within book-based narratives and, and so on. The portable device may be considered an endpoint platform, connected to and exchanging information with a computer network (e.g., user interaction data).
[0016] Within descriptions herein, references are made to the portable device being “handheld”. Although the device may be readily manipulated by a hand of the device user, it may alternatively, or in addition (e.g., at different times), be manipulated by other parts of the user’ s body including a head, wrist, arm, shoulder, leg, or chest. Attachment of the device to a body part may be aided by one or more supportive components such as a headband, wrist strap, or chest holster. Manipulation by other parts of the body may, for example, free a user’s hands to perform other tasks such as page-tuming (e.g., using either or both hands), or holding a young child while interacting with a book.
[0017] According to one aspect, devices and methods are provided for an individual to acquire a comprehensive (i.e., complete) image of a page by pointing a light emanating from a portable device at, or at least adjacent to, the viewable page. As described more fully below, herein, the term “page” refers to any substantially two-dimensional surface capable of displaying static (e.g., printed) viewable content. A page may, for example, comprise a book page, a book spread (i.e. two adjacent pages), a book cover, a magazine page, a newspaper spread, a brochure, and so on. Viewable pages may, for example, be produced using a printing process on paper, cardboard, cloth, hardboard, plastic, etc.
[0018] A device camera co-aligned with the projected light may capture an image of the viewable page. A camera-acquired image may include the reflection of the light beam (e.g., containing incident light reflected off the object). Alternatively, the beam may be turned off (e.g., momentarily) while one or more images are being acquired. Turning the beam off during camera-based acquisition may allow the contents of the page to be imaged absent interference by light beam reflections.
[0019] The device may then determine whether all edges of the viewable page (or other object of interest, see below) are present within the image. Alternatively, or in addition, the image may be processed (e.g., using neural net-based classification) to determine if a complete page appears to be present (i.e. within the image) and / or if the region of the page includes pixels located at the perimeter of the camera image. If the image of an object (e.g., page) includes (i.e., “touches”) the edge of the acquired image, then some portion of the object is likely beyond the field of view of the camera. In other words, the image of the object may not be comprehensive or complete.
[0020] Optionally, the camera-acquired image may be further processed to determine if the image of the page occupies a sufficient number of pixels (i.e., possess sufficient resolution) to reliably identify, for example, text, drawings and other printed material. A reduced number of pixels within a camera image may arise as a result of 1) a location of the device camera (absent an ability to optically zoom) well away from the page, and / or 2) an orientation of the camera pointing toward the page but well away from a normal to the surface of the page (e.g., see FIG. 9). In the latter case, image “skew” caused by viewing a surface away from the normal to the surface (e.g., at an acute angle) may be further subdivided into: 1) horizontal (e.g., left or right) and / or 2) vertical (e.g., up and down) skew directions.
[0021] If all aspects and / or edges of the page are present within the image and optionally, if the page within the image has been acquired with sufficient resolution, then the camera- acquired image may be classified and / or labeled as containing a complete image of the viewable page. The image may be stored and (optionally) an audio, visual or haptic indication of successful acquisition of a complete image of the page may be provided to the device user.
[0022] On the other hand, if a complete object has not been detected, the device may generate one or more instructions for the user to better position and / or orient the device. The generation of one or more camera (i.e., device) movement instructions may take into account the camera viewing perspective (i.e., camera pose) based on the acquired image. As an example, if the leftmost edge of a page is missing from the camera-acquired image, then the user may be instructed to move the device to the left. More generally, the device may compare the viewing perspective of the acquired image based on movement of the device toward a location and orientation with a computed (e.g., based on geometry) camera pose that ensures all of the page may be seen in the camera-acquired image with adequate resolution.
[0023] Instructions to the device user may direct movement of the device in any, or any combination, of the six degrees of freedom (i.e., left-right, up-down, closer-further, pitch, roll and yaw) that an object may generally be manipulated in a three-dimensional space. The one or more instructions to better position and / or orient the device may be conveyed to the device user by visual, audio and / or tactile means.
[0024] Visual indications may be conveyed by one or more indicator lights (e.g., LEDs) on the device. Visual feedback to the user may also include modulating or structuring (e.g., patterning) the device light used to point toward the page (or other object). A reflected image produced by the device pointing light source may indicate direction (e.g., directional arrow) and / or orientation (e.g., rotational symbol) of a target movement of the device (e.g., see FIG. 10). Movement magnitude may be provided by, for example, the color and / or size of the indicating symbol(s). Providing visual feedback by modulating and / or structuring the light used to point toward the page may help to maintain visual focus by the device user in the region of the page or other object of interest. Audio instructions may be conveyed using a device speaker. Audio instructions may, for example, consist of phrases, words and / or sounds. One or more instructions may indicate a direction and / or orientation to move the device along with indications of magnitudes for such movements. Tactile feedback (e.g., to alert a user that device movement may be required) may alternatively or additionally be enacted using a device haptic unit.
[0025] All forms of user feedback may include modulating a modal frequency (e.g., visual color, acoustic tone, vibrational frequency), repetition frequency and / or intensity of the signal conveyed to the device user. Such feedback may be provided in a continuous manner as the device is moved about. As an example of such feedback, as the device is moved toward a location and / or orientation with an improved view of the page, the flashing of one or more lights, clicking sound and / or tactile pulses may increase in frequency and / or intensity. Conversely, moving or orienting the device in a manner that decreases an ability of the device camera to view the complete page (or other object, see below) results in a reduced frequency and / or intensity of the feedback. Such continuous feedback may continue until an adequate camera viewing location and orientation are achieved
[0026] According to further aspects, devices and methods are provided to apply a similar overall strategy to ensure that an image of any object (i.e., not confined to a viewable page) in the environment of the device user captured by the device camera includes all aspects of the object, including its edges. Similar to the process just described, an individual may point a light emanating from the portable device onto or in the vicinity of (e.g., adjacent to) an object in the user environment. The device may subsequently determine, within an image acquired by the device camera co-aligned with the projected light, whether all edges of the object are present (with adequate resolution) and / or whether any aspect of the object touches the perimeter of the image.
[0027] If a complete object is identified within the camera-acquired image with adequate resolution, then the image and / or any determined identity of the object may be used within subsequent interactive sequences. Such interactions (e.g., enacted by the device) may proceed in a manner that assumes all viewable aspects of the selected object are available within the camera-acquired image.
[0028] If only a portion of the object is present within the camera-acquired image and / or if the object has inadequate resolution within the image, then the user may be instructed to reposition and / or re-orient the device such that the device camera is moved to better acquire all aspects of the object being pointed toward. Such re-positioning and / or re-orienting may be expressed relative to the camera viewing perspective (i.e., camera pose) of the object within the acquired image.
[0029] The one or more instructions to the device user may be conveyed using haptic, visual, and / or audio means via a haptic unit, one or more light indicators (including the projected light), and / or a speaker operatively coupled to the device processor. Instructional feedback to the user may be provided until a complete (i.e., comprehensive) representation of the object can be identified within a camera-acquired image.
[0030] Taken together, the contents of a book or images of other objects in the environment of the device user may be acquired using the same device subsequently used to interact with the book. Augmenting printed content with interactive sequences including added narratives and real-time feedback related to content may not only provide omnipresent, machine-based guidance while reading, but also be “fun” and / or help maintain emotional engagement (without being overwhelming) within a learning environment. Augmenting printed content with interactions involving objects in the real world within a play and / or learning environment, may further help a child or learner relate to, and engage with, the printed material.
[0031] In accordance with an example, a device is provided to acquire a comprehensive image of a page comprising: a device body configured to be manipulated by a device user; electronic circuitry within the device body that includes a device processor; a light source configured to emit a light away from the device body to generate a light reflection on a surface at a light reflection location; and a camera operatively coupled to the device processor and aligned such that a field of view of the camera includes the light reflection location, wherein the device processor is configured to: acquire, using the camera, a camera image when the device is manipulated by the device user to project the light reflection onto or adjacent to the page; identify, within the camera image, one or more identified page edges; and either: a) if the one or more identified page edges comprise all page edges, label the camera image as the comprehensive image of the page, or b) if a) is not met, determine one or more device movement instructions to manipulate the device such that a predicted camera image of the page upon following the one or more device movement instructions contains all of the page edges.
[0032] In accordance with another example, a device is provided to acquire a comprehensive image of a page comprising: a device body configured to be manipulated by a device user; electronic circuitry within the device body that includes a device processor; a light source configured to emit a light away from the device body to generate a light reflection on a surface at a light reflection location; and a camera operatively coupled to the device processor and aligned such that a field of view of the camera includes the light reflection location, wherein the device processor is configured to: acquire, using the camera, a camera image when the device is manipulated by the device user to project the light reflection onto or adjacent to the page; identify, within the camera image, one or more page edge pixels; determine a pixel count of the one or more page edge pixels that are one of camera image edge pixels; and either: a) if the pixel count exceeds a predetermined pixel count threshold, determine one or more device movement instructions to manipulate the device such that a predicted page edge pixel count in a predicted camera image following manipulation that are one of the camera image edge pixels is less than the predetermined pixel count threshold, or b) if a) is not met, label the camera image as the comprehensive image of the page.
[0033] In accordance with a further example, a device is provided to acquire a comprehensive image of an object comprising: a device body configured to be manipulated by a device user; electronic circuitry within the device body that includes a device processor; a light source configured to emit a light away from the device body to generate a light reflection on a surface at a light reflection location; and a camera operatively coupled to the device processor and aligned such that a field of view of the device camera includes the light reflection location, wherein the device processor is configured to: acquire, using the camera, a camera image when the device is manipulated by the device user to project the light reflection onto or adjacent to the object; identify, within the camera image, one or more object edge pixels; determine a pixel count of the one or more object edge pixels that are one of camera image edge pixels; and either: a) if the pixel count exceeds a predetermined pixel count threshold, determine one or more device movement instructions to manipulate the device such that a predicted object edge pixel count in a predicted camera image following manipulation that are one of the camera image edge pixels is less than the predetermined pixel count threshold, or b) if a) is not met, label the camera image as the comprehensive image of the object.
[0034] Other aspects and features including the need for and use of the present invention will become apparent from consideration of the following description taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] A more complete understanding may be derived by referring to the Detailed Description when considered in connection with the following illustrative figures. In the figures, like-reference numbers refer to like-elements or acts throughout the figures. Presented examples are illustrated in the accompanying drawings, in which:
[0036] FIG. 1 illustrates exemplary manipulation of a handheld device to point a light emanating from the device to generate a dumbbell-shaped reflection in the gutter region of a book spread, helping to direct a device camera (co-aligned with the light) to acquire an image of the complete spread.
[0037] FIG. 2 illustrates some of the terminology used in bookmaking including a spread (i.e., two adjacent pages) and the central gutter (i.e., region between the two pages).
[0038] FIG. 3 is an exploded-view drawing of an exemplary handheld device showing locations, pointing directions, and typical relative sizes of a light source and a co-aligned camera.
[0039] FIG. 4 is an exemplary interconnection layout of components within a portable device (in which some components may not be used during some applications) showing predominant directions for the flow of information relative to a bus structure that forms an electronic circuitry backbone.
[0040] FIG. 5 illustrates exemplary reflections produced by the structured light emanating from the handheld device including geometries that include directional orientation, cover larger areas and / or are whimsical.
[0041] FIG. 6 illustrates an exemplary situation in which the field of view of the camera is determined to be pointing to the right of a page spread, resulting in one or more instructions to the device user to move the device translationally to the left.
[0042] FIG. 7 illustrates an exemplary scenario in which the camera field of view is determined to be rotated counterclockwise compared with the spread, resulting in one or more instructions to the device user to rotate the device clockwise.
[0043] FIG. 8 illustrates an exemplary situation in which a book spread is accessible within a camera-acquired image, but is too small for reliable processing, resulting in one or more instructions to the device user to move the device closer to the spread.
[0044] FIG. 9 illustrates an exemplary scenario in which an image of a book spread is skewed as a result of holding the device away from a normal to the spread. FIG. 10 shows examples of visual prompts and directional words that may be used to instruct a device user to move the device to better view a page or other object.
[0045] FIG. 11 is an exemplary flow diagram illustrating steps to acquire a book spread in which audio, haptic and / or visual instructive feedback is provided to a device user to help ensure that a complete spread is acquired.
[0046] FIG. 12 is an exemplary flow diagram illustrating steps to identify an object in the environment of the device user or to provide audio, haptic and / or visual instructive feedback to help ensure the object may be identified with confidence.
[0047] FIG. 13 illustrates the identification of interactive blocks or “chunks” of text that, for example, may be primarily oriented vertically, horizontally, in an arbitrary direction, or using varying directions, and / or displayed with different fonts and / or within a structure such as a text bubble.
[0048] FIG. 14 illustrates an exemplary scenario in which an incomplete image of a book spread includes object edges acquired by pixels located at the periphery of the camera sensor array, suggesting that the camera-acquired image of the object is incomplete.
[0049] FIG. 15 is an exemplary flow diagram illustrating steps to determine whether pages, or regions of a page, are being viewed in a sequential order based on a book’s storyline and, if not, to provide the user with audio, haptic and / or visual instructive feedback.
[0050] DETAILED DESCRIPTION
[0051] Before the examples are described, it is to be understood that the invention is not limited to particular examples described herein, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular examples only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0052] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. It must be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to “a compound” includes a plurality of such compounds and reference to “the polymer” includes reference to one or more polymers and equivalents thereof known to those skilled in the art, and so forth. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither or both limits are included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0053] Certain ranges are presented herein with numerical values being preceded by the term “about.” The term “about” is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.
[0054] Devices and methods are provided herein that describe a light-weight, simple-to-use and intuitive portable (e.g., handheld) device that may bring a printed book “to life” by adding to a printed page, context and content, related sounds and sound effects, visual cues, haptic stimulation, queries, user exchanges, and so on. Effective interactions with the book facilitated by the device require the contents of the book to be known by the device and / or a connected processor.
[0055] In some cases, the contents of a book may be acquired from one or more bookcontent repositories. To download such book contents, the identity of the book may be acquired by the device, for example, by scanning (e.g., using the device camera) an International Standard Book Number (ISBN), by recognizing (e.g., using a device camera) the title and author(s) of the book and / or by using image search to identify the book based on its cover or other identifiable view (e.g., table of contents, introduction). If a book’s contents are not available, an ability to scan contents with the same portable device used during interactions may be useful.
[0056] According to one aspect, devices and methods are provided herein for an individual to acquire (e.g., digitize) the contents of a book, one viewable page (or page spread) at a time, using a portable (e.g., handheld) device. The device may be manipulated by the device user such that a light emanating from the device is pointed toward (i.e., at or at least adjacent to) each viewable page. A camera within the device, co-aligned with the projected light may acquire an image containing a comprehensive (i.e., complete or entire) page with sufficient pixel resolution to be suitable for additional processing (e.g., optical character recognition, image recognition) to generate interactive sequences with the device user based on camera-acquired contents.
[0057] However, if the image does not contain the entire page (e.g., one or more edges of the page cannot be identified within the image) or if the resolution of the image of the page is insufficient to reliably identify contents, the device may generate one or more audible, visual and / or haptic instructions to be followed by the device user to guide the device (i.e., including its camera) toward a location and / or orientation better suited to image the page. Such feedback with the device user may continue until comprehensive images are acquired for each and / or all pages.
[0058] Within descriptions herein, the term “page” refers to any substantially two- dimensional surface capable of displaying viewable content. Pages may be constructed of one or more materials including paper, cardboard, film, cloth, wood, plastic, glass, a painted surface, a printed surface, a textured surface, a surface enhanced with three-dimensional elements, a flexible printed surface, and so on.
[0059] As used in the book-making industry and illustrated in FIG. 2, a “spread” comprises a pair of adjacent (i.e., left and right) pages that may be viewed simultaneously. Acquiring a complete image of a spread may include ensuring that all (four) edges of the spread are visible within a camera-acquired image. Within descriptions herein, acquiring the contents of a “page” includes camera-based imaging of any number of surfaces containing printed content including a single page, collection of viewable pages, book spread, folded surfaces elements that have been unfolded for viewing, front or back book cover, and so on.
[0060] Along similar lines, the term “book” is used herein to refer to any collection of one or more pages. A book may include printed materials such as a traditional (e.g., bound) book, magazine, brochure, newspaper, handwritten notes and / or drawings, book cover, chapter, box, scrapbook, collection of photographs or drawings, and so on.
[0061] Page content is limited only by the imagination of the author(s). Typically, for example, a children’s book may contain a combination of text and drawings. More generally, page content may include any combination of text, symbols (e.g., including the range of symbols available in different languages), logos, specifications, drawings and / or images; and may be displayed in color, shades of gray, or black-and-white. Content may portray real or fictional scenarios, or mixed combinations of both.
[0062] Although it may be possible to image the front and back covers of a book simultaneously (e.g., by turning over an opened book to be face-down) front and back covers may, if desired, be acquired separately. In general, similar criteria for comprehensive images of covers, pages and spreads may be applied; namely that all (four) edges are viewable and the surface of the object occupies a sufficient number of pixels within an image to be processed further, for example, including text and / or object identification.
[0063] Additionally, references herein to a device that is “handheld” may help to visualize typical use of the device. However, the device (including the light beam source and a co-aligned camera) may be attached or connected to, and / or manipulated by, any other part of the human body including a device user’s head, wrist, arm, shoulder, leg, or chest. Manipulation by other parts of the body may, for example, free a user’s hands to perform other tasks such as holding an infant or young child, using fingers or a stylus to point at objects (e.g., text, drawings), drawing on a whiteboard, generating signing or other gestures, and so on.
[0064] Physical attachment of a device to a body part may be aided by one or more supportive structures such as a headband, wrist strap, arm extension, and shoulder or chest holster. Attachment of the portable device to a support structure may be aided by a configuration that allows quick and easy attachments and detachments. For example, one or more attachment points may be held magnetically, using a simple latch mechanism, and / or using a so-called hook and loop fastening system (e.g., manufactured by Velcro). Quick and easy attachments and detachments may facilitate employing (and purchasing) a single portable device manipulated by different body parts, or by different users, at different times.
[0065] Additionally, distinct devices may be specifically designed (e.g., with different device body shapes and / or optical working distances) to be more conveniently manipulated by different body parts. For example, functions of a larger pushbutton switch on a handheld device may be performed by a larger contact (e.g., touch sensitive) switch on a head-mounted device. Additionally, as described further below, the focal and / or working distances of optics associated with both the light source and camera may be greater for a head mounted device compared with a handheld device.
[0066] According to further aspects of the devices and methods herein, the device includes a light source and one or more optical elements to produce an emitted light, projecting from the device. Light may be generated using one or more lasing diodes, such as those manufactured by OSRAM and ROHM Semiconductor. Lasing diodes (and lasers in general) produce coherent, collimated and monochromatic sources of light. Considering the portability of a device in which a close-proximity light source may be pointed in any direction, increased eye safety (especially during use by a child, generally considered an “uncontrolled” environment from a safety perspective) may be attained using a non-coherent source such as a (non-lasing) light-emitting diode (LED). LED so-called point sources, such as those manufactured by Jenoptik and Marktech Optoelectronics may produce non-coherent, collimated and (optionally) polychromatic light sources.
[0067] Optical components associated with LED point sources may control light divergence that, in turn, may guide reflection size at typical working distances (e.g., about five centimeters to one meter when handheld) when the device is used to point toward a page. As described more fully above, the same portable device used to scan book contents may be used during user interactions while reading a book.
[0068] The focal distances (or focal ranges) of the light source and co-aligned camera may be configured to match light-path geometries within different device configurations and / or different applications. For example, when manipulated by a head of a device user to direct the beam at a book page, the focal distance may be extended (e.g., to up to about two meters). Whereas when manipulated by a hand or arm of a device user, focal distances may typically be lessened (e.g., to up to about one meter) as a result of reaching with the device by the hand or arm toward a page.
[0069] Reflection patterns that may be an element of user interactions (e.g., smiley face, directional arrow) at a predetermined working distance are further described in co-pending application Serial No. 18 / 202,456, filed October 20, 2023, the entire disclosure of which is expressly incorporated by reference herein. The use of a light beam to point at an object during user interactions is further described in U.S. Patent No. 11,989,357, filed July 11, 2023, the entire disclosure of which is expressly incorporated by reference herein.
[0070] Within additional examples, wide spectrum (at least compared to a laser) and / or polychromatic light sources produced by (non-lasing) LEDs may help beam reflections to be seen by those who might be color-blind within a region of the visible spectrum. A polychromatic light source may also be viewed more consistently by all users when reflected off various surfaces. As an example, a purely green light source may not be seen easily when reflected off a purely red surface (e.g., region of a page). A polychromatic light source, especially in the red-green portion of the visible spectrum, may help alleviate such issues. More energetic photons within the deep blue end of the visible spectrum may be avoided for reasons related to eye safety. Within further examples, the device beam source may be operatively coupled to the device processor, allowing the intensity of the beam source to be controlled, including turning the beam on and off. When the beam is turned on, and as long as a surface is sufficiently reflective, a camera-acquired image may include a reflection produced by the light beam (e.g., containing incident light reflected off the page).
[0071] Optionally, the beam may be turned off during times when camera images are acquired, for example, to avoid beam reflection(s), and / or pixel saturation at or near a reflection (e.g., as a result of pixels “bleeding” due to saturation). Even while the beam is turned off, because both the beam and camera move at the same time (i.e., affixed to, or embedded within, the portable device body), the location or region a beam is pointed toward within a camera image may be known regardless of the physical position, pointing direction, or overall orientation of the device in (three-dimensional) space. The absence of reflections off objects being pointed at (where reflections may be considered obstructive “noise” when identifying objects) may reduce computational requirements and increase classification accuracy during CV processing of camera-acquired images.
[0072] The beam may also be turned off as an indication to a user, acquisition of a complete page image (or other object) with sufficient resolution. Conversely, leaving the beam turned on during interactions may indicate to the device user that further manipulation of the device is expected to better position the device (or, more specifically, the device camera) to obtain a complete image of the object being pointed at by the projected light.
[0073] Optionally, projected light intensity may also be modulated, for example, based on measurements of one or more reflections within images acquired by the device camera. Reflection intensity may be made clearly discernible to a user over background (e.g., considering ambient lighting conditions, to accommodate for the reflectivity of different surfaces, and / or to accommodate for visual impairment of the user), but not overwhelming (e.g., based on user preference). Beam intensity may be modulated by a number of methods known in the art including regulating the magnitude of the light beam driving current (e.g., with transistor-based circuitry) and / or using pulse width modulation (i.e., PWM).
[0074] As a further aspect of devices and methods herein, structured-light illumination patterns may range from simple imagery (e.g., one or more points, line drawings) to complex, rich, high-resolution images. Additionally, illumination patterns may be animated. Illumination patterns including animations may be generated by, for example, miniature LED arrays, filtering by an addressable liquid crystal (e.g., commonly used within liquid crystal displays, LCDs), or DLP (i.e., digital light processing) employing an array of microscopic mirrors.
[0075] Illumination patterns may: 1) initially encourage the user to manipulate the portable device about a position and orientation suitable for acquiring acceptable camera-based page images, and / or 2) provide visible user feedback that may include one or more instructions when incomplete images of a page are detected. In other words (in addition to pointing), projected imagery may facilitate dual roles of providing initial encouragement for manipulation prior to acquiring a camera-based image, and instructional or acknowledging feedback once an image has been acquired.
[0076] Additionally, emitted light images may enhance or augment the objects, drawings, photographs, and / or words on which they are projected. Patterns may move in concert with portable device movements (e.g., when used to point), or projected patterns may be dynamically stabilized using imagery acquired by the camera, madding the projected image appear stable even when there is movement (over a limited range) of the portable device.
[0077] Projected imagery may (i.e., optionally) contain a directional element (i.e., an image that indicates a unique spatial direction). The directional element of the reflected image may be aligned with an orientation (e.g., horizontal or vertical) of images captured by the coaligned camera, and serve to encourage a user to manipulate the device in a particular orientation. Examples of reflections containing one or more directional patterns (e.g., see FIG. 5) include a directing symbol (e.g., arrow, line, pointing finger), collinear arrangement of objects (e.g., pair of points), or an image elongated in one direction (e.g., ellipse, dumbbell).
[0078] Encouraging a directional orientation of the portable device may, for example, reduce needs to digitally rotate camera-images (e.g., containing text) for processing. Alternatively, or in addition, most video cameras have a greater number of pixels in the (typically assigned) horizontal axis compared with the vertical direction. Common video format aspect ratios (horizontal compared to vertical) are 4:3 and 16:9. Taking advantage of a camera aspect ratio, the long dimension of an acquired image may be aligned with the long dimension of a viewable scene (e.g., book spread). This strategy may facilitate acquiring images with increased resolution (e.g., an object filling image dimensions) and / or relax spatial accuracy when directing the device toward a target object (e.g., making use, if needed, of additional horizontal pixels to capture an image of the object). As an example, a user may be instructed to orient the long axis of a light reflection to align with a (vertical) spread “gutter”, comprising the central region of a book spread between left and right pages (e.g., see FIG. 2). A camera with a horizontal axis aligned perpendicular to the long axis of the light reflection may acquire a greater number of pixels in the horizontal direction. If the horizontal width of a book spread is greater than its height, then an image of the spread may be taken with a greater resolution (e.g., when the device is held closer to a spread) and / or requirements to center the spread within the camera-acquired image may be relaxed (e.g., with additional camera pixels available in the horizontal direction aligned with the long axis of the spread).
[0079] Within further examples, construction of the portable device may strive to place the beam reflection at the center of camera images. However, given a small separation (e.g., as a result of physical construction constraints, for example as shown in FIG. 3) between the beam source and the camera sensor, the beam may not appear at about the center of camera images at all working distances.
[0080] The pointing directions of a beam and camera field-of-view may be aligned to converge at a preferred working distance. In this case, the beam may be made to appear at about the center of camera images (or some other selected camera image location) over a range of working distances. In this configuration, as the distance from the device to a reflective surface varies, the location of the beam may vary over a limited range (generally, in one dimension along an axis in the image plane in a direction defined by a line passing through the center of the camera field-of-view and the center of the light beam source).
[0081] At a particular working distance (i.e., from the device to a reflective surface), a location of a reflection may be computed using geometry (analogous to the geometry describing parallax) given the direction of beam pointing, the direction of camera image acquisition and the physical separation between the two (see, e.g., FIG. 3). By keeping the physical separation between the beam and camera small, a beam pointing region within camera images may be kept small over a range of working distances employed during typical applications.
[0082] Alternatively, the beam and camera may be aligned to project and acquire light rays that are parallel (i.e., non-converging). In this case, a reflection may be offset from the center of the camera’s field-of-view by an amount that varies with working distance. The separation between the center of camera images and the center of the beam decreases as working distances increase (e.g., approaching a zero distance at infinity). By keeping the physical distance separating the beam and camera small, the separation may similarly be kept small.
[0083] Within further examples, based on the spectral sensitivity of a typical human eye, a light beam in the green portion of the visible spectrum may be most readily sensed by most individuals. Many so-called RGB (i.e., red, green, blue) cameras contain twice as many green sensor elements as red or blue elements. Utilizing a light beam in the mid-range of the visible spectrum (e.g., green) may allow beam intensity to be kept low but readily detectable (e.g., by both humans and cameras), increasing overall eye safety.
[0084] As further aspects of devices and methods herein, the device camera, co-aligned with the device beam may continuously acquire camera images, analyzing each image to determine if a partial or complete page is present within the image. In this mode, if no element of a page is present with an image, instructions to the user (described in greater detail below) may be suppressed, at least temporarily so as not to overwhelm with instructions directed at the user (who might be otherwise engaged). Once an image, or partial image, of a page is acquired, the process of acquiring page contents may continue.
[0085] Alternatively, the device user may indicate (i.e., to the device) that the projected light is pointed toward a page and the device should either label the page image as successfully acquired or provide instruction for improved imaging. Such indications may be made via a range of interactive methods, such as:
[0086] 1. pressing or releasing a switch (e.g., pushbutton, or other contact or proximity sensor) that is a component of the device,
[0087] 2. providing a verbal indication (e.g., saying “now”) sensed by a device microphone and identified (e.g., classified using natural language processing) by the device processor or a remote processor,
[0088] 3. point the beam at a location (i.e., absent substantial movement) for a predetermined (e.g., based on user preferences) “dwell” time,
[0089] 4. orienting the device in a predetermined direction (e.g., vertically relative to the gravitational pull of the earth, tipping the device forward) sensed by a device IMU, or
[0090] 5. gesturing or tapping the device, also sensed by a device IMU.
[0091] Within these latter exemplary cases, in which signaling movements of the device by the user (e.g., gesture, tap) may produce motion within the camera’s field-of-view, a stationary image may be isolated (e.g., from a continuously sampled series of images) prior to any process that might produce movement. The camera-acquired image prior to any motionbased signaling may be used to identify a page.
[0092] Within further examples, one method to implement dwell-based methods may involve ensuring a number of consecutive images (e.g., computed from desired dwell time divided by frame rate) to reveal a substantially stationary viewable object and / or beam reflection. CV techniques such as template matching, computer vision, or neural network classification may be used to compute one or more spatial offsets comparing pairs of successively acquired camera images. Image movement (e.g., to compare with a dwell movement threshold) may be computed from the one or more spatial offsets or a sum of offsets over a selected time.
[0093] When determining rapid and / or precise dwell times, movement measurements based on camera images demand high frame rates, and resultant computational and / or power needs. Alternative methods to determine if a sufficient dwell time has elapsed include using an IMU to assess whether the device remains substantially stationary for a predetermined period.
[0094] Conversion of analog IMU data into a digital form, suitable for processing, may use analog-to-digital (A / D) conversion techniques, well-known in the art. IMU sample rates may generally be in a range from about ten (10) samples / second (even lower sample rates may be employed, if desired) to about the thousand (10,000) samples / second where higher IMU sample rates involve trade-offs involving signal noise, cost, power consumption and / or circuit complexity. Such rapid sampling may allow brief and / or precisely measured dwell times. Movement and / or dwell time thresholds may be based on user preferences.
[0095] As a further aspect of devices and methods herein, if the complete the page cannot be determined within the camera-acquired image or the device needs to be moved (e.g., translationally), tilted, and / or held closer to the page to acquire a sufficient number of page- based pixels for processing, then the device may generate one or more instructions for the device user to manipulate the device in a manner to acquire the full-page contents within camera-acquired images.
[0096] At least two general classes of methods may be used to determine if a camera image contains a complete image of an object (including a page) and, if not, to provide user instructions to manipulate the camera positions: 1) traditional computer vision (CV) methods that include object edge detection and / or computing the camera pose (e.g., camera position and / or orientation relative to an object or location), and 2) neural network approaches based on a neural networks trained to classify camera-acquired images.
[0097] Within the field of computer vision (CV), a number of edge detection methods are available to identify the boundaries of an object. One class of edge detection approaches involves computing a first-order spatial derivative and searching for maximum gradients coupled with following local gradient directions. Another class of approaches involves identifying zero-crossings within second order spatial derivatives. Examples of edge detection approaches include a first-order gradient operator, a number of variations of Canny edge detection schemes, a Sobel operator, and a phase stretch transform.
[0098] Once edges within an image are determined, assessing the presence of a complete page may be addressed by: 1) determining where all (e.g., four) edges (including comers) of a page are present within the image, and / or 2) determine if any edges “touch” the periphery (i.e., border pixels) of the camera-acquired image. If a comprehensive page is not present within the image, then device movement and / or orientation instructions to the user may be generated based on: 1) a direction computed from the location(s) of missing edges and / or edges that touch the image border, and / or 2) computing the “pose” or viewing orientation of the camera relative to the page.
[0099] Within the field of computer vision, the task of estimating the pose of a camera given a set of “n” points in space is referred to as a perspective-n-point (PnP) problem. In general, if the locations of three points are known, camera pose may be estimated (e.g., available in software libraries such as OPENCV). As an example, if the positions of at least three corners of the page are known, a camera pose may be compared with a target or “ideal” camera pose located along a line normal to the center of the page at the focal distance of the camera away from the page surface.
[0100] Within neural network (NN) base approaches, a neural network (NN) may be trained to identify (i.e., classify) areas of a page including whether a comprehensive or complete page is present and, if not, a direction and / or magnitude of a movement needed by a camera to best acquire a complete image of the page. For the purposes of descriptions herein, a neural network may refer to any type of network including perceptron, feed-forward, recurrent, modular, and convolution neural networks. Such networks may contain different neural network architectures (number of nodes, number of layers, activation functions). Convolution neural networks (CNNs), including deep CNNs, are commonly used to classify content within images. Alternatively, or in addition, a neural network may be trained to identify the visual characteristics of a page (including a book spread). Approaches similar to those used to identify foreground (e.g., a face) versus a surrounding background may be employed to determine the extent and location of a foreground page versus background (i.e., that generally does not “look like” a page). Within a typical camera view, page characteristics may include being rectangular in shape; and containing text, symbols, drawings and / or images.
[0101] A neural network may be trained to input a camera-acquired image and output either an indication that a complete book page of an appropriate size may be seen within the image, or which movement, or combination of movements, of the device might better position the camera to acquire an image of the page. Training may include labeled datasets of (generally a large number of) images of book pages acquired under a wide range of conditions that might be encountered during use of the portable device. Training datasets may contain images intentionally missing one or more edges (or partial edges) and / or acquired from a non-ideal camera pose.
[0102] Image labels may indicate that an acceptable view of a book page (e.g., containing all four edges and corners) is present within the camera image, or to classify (i.e., based on the camera image) a movement, or combination of movements (e.g., including orientation, direction and magnitude of movement) to better position the camera to view a comprehensive (i.e., complete) page. Such positioning may strive to both 1) position the center of the image of the page in the central region of the camera image, and 2) view the page at (or near) a line emanating from the center of the book page normal to the plane of the page.
[0103] A training library that includes anticipated conditions during device use may improve the robust of classification. Training conditions may include variations in:
[0104] 1) intensities, spectral characteristics (e.g., indoors versus outdoors), positions (i.e., relative to the page) and number of light sources illuminating the page,
[0105] 2) sizes and aspect ratios of pages,
[0106] 3) contents of book pages (e.g., text, drawings) including the degree of detail required to reliably determine content,
[0107] 4) background surrounding the page including variations in focus (e.g., distance from the book spread) and contents.
[0108] Within further examples, network training (and other programming of device CV and Al processes) may take advantage of a wide array of distributed machine learning resources (e.g., TensorFlow, SageMaker, Watson Studio). The restricted nature of comparing camera-based images with the relatively simple geometry of a page may greatly simplify training (and classification) processes to determine the presence, or not, of a complete page within a camera-acquired image, to generate instruction on how to move the device.
[0109] Such confined datasets may also allow relatively simple classification networks and / or decision trees to be implemented. Optionally, classifications may be performed entirely on a device (with confined computing resources) and / or without transmitting to remote devices (e.g., to access more substantial computing resources). Such classifications may be performed using neural network (or other CV and Al) approaches using hardware typically found on mobile devices. As examples, MobileNet and EfficientNet Lite are platforms designed for mobile devices that may have sufficient computational power to determine locations of pages within camera-acquired images.
[0110] As further aspects of devices and methods herein, within three-dimensional space, an object such as a portable (e.g., handheld) device may have six degrees of movement freedom. Thus, instructions to the device user may include movements of any, or any combination, of these six movement elements:
[0111] 1) horizontal translational movement to the left or right,
[0112] 2) vertical translational movement up or down,
[0113] 3) moving closer to or further away from the page (zoom direction),
[0114] 4) rotational movement clockwise or counterclockwise (analogous to airplane roll),
[0115] 5) tip the device forward or backwards (analogous to airplane pitch), and / or
[0116] 6) tip the device to the left or right (analogous to airplane yaw).
[0117] In some cases, combinations of these basic movement elements may be most efficient (e.g., minimum performance time and / or number of steps) to acquire improved positioning of the portable device. For example, if the device is generally pointed toward a page but positioned well below the center of the page, the translational movement in the upward direction along with pitch movement in the forward direction may be required to keep the device pointed toward the center of the page. This exemplary combination of movement elements may help to position the device normal to the page, helping to eliminate skew.
[0118] As further aspects of devices and methods herein, from an imaging perspective, movement of the device closer to, or further from a page may be used to control the size (i.e., number of image pixels) a page occupies within a camera image. As just described, movement in this direction effectively performs a “zoom” function. When too close to a page, the field of view of a held device camera may not include all edges or corners of a page. Conversely, if held too far from the page, an insufficient number of pixels may be available to classify and / or process printed content within the image of the page.
[0119] Alternatively, or in addition to instructing the device user to move the device closer to, or further away from, a page, an optical zoom may be incorporated within the optics associated with the camera within the device. This may be implemented by including one or more moveable optical elements in the light path of the camera image. Moveable elements may include one or more refractive lenses, reflective mirror surfaces, and / or prisms.
[0120] Due to the scale of optical components only small movements of the one or more moveable elements may be necessary to implement a zoom function over a range (e.g., 0.5X to 5x) suitable for a device. Such small movements may be implemented using so-called MEMS (micro-electro-mechanical devices) approaches. Based on the image size of a page or partial page within a camera-acquired image, the device processor may control voltage applied to moveable optical elements to adjust zoom.
[0121] Within further examples, user interface (UI) interactions for audible instructions may consider the relatively long time required to vocalize (and repetitive nature of) an instructional phrase or even a single word compared with, for example, the ability of a device user to quickly manipulate a device for improved viewing of the page. Thus, instructions may be confined to single words, brief phrases and / or sound characteristics. Example of relatively simple single- or double-word instructions include left, move left, right, move right, up, move up, down, move down, in, move in, forward, move forward, toward, move toward, out, move out, move away, move back, clockwise, rotate clockwise, rotate right, counterclockwise, rotate counterclockwise, rotate left, tip forward, tip back, tip backward, turn left, swing left, turn right, and swing right.
[0122] User instructions may also be conveyed by changes in the spectral frequency, repeat frequency, and / or amplitude of one or more sounds. For example, a tone may be heard as higher in pitch and / or louder as an improved camera viewing location and / or orientation are determined during device manipulations. As another example, a brief (e.g., clicking) sound may be repeated at an increased frequency as the device is moved toward an improved viewing location and / or orientation. Audible feedback may (optionally) also include audible feedback when a suitable combination of device location and orientation is achieved. For example, a “ta-da” fanfare sound may be played when a complete page image is acquired.
[0123] Visual instructions to manipulate the device may similarly take into consideration the time required to convey meaning (e.g., movement magnitude and / or direction). Additionally, a user’s eyes might typically be focused he page as device manipulations are being performed, Different approaches may also be enacted depending on whether multiple illumination indicators are available and / or whether the light projected by the device can be patterned under the control of the device processor (e.g., to produce different reflected images or symbols).
[0124] A single illumination source (e.g., a LED on the device or the projected light) may be made to blink at a rate that depends on proximity to a location and orientation of the device to successfully acquire a page. For example, the projected light (e.g., frequently where a user’s eyes might be directed) may blink faster as the device is manipulated to be closer to a location and orientation where the image of a complete page may be acquired.
[0125] Along similar lines, the color and / or intensity of the light may be changed as the device is manipulated. For example, the blue end of the color spectrum might indicate movement away from a location and orientation where a page might be viewed, whereas light toward the red end of the visible spectrum might indicate movement toward a better viewing location. As a further example, a dim visual cue might indicate movement away from an effective viewing location, whereas a bright light may be associated with an improved viewing location.
[0126] According to further aspects, similar methods may be used to acquire comprehensive images of objects in the environment of the device user (i.e., in addition to pages). Methods may ensure that an image of a physical object captured by the device camera includes a complete view of the object (i.e., from the viewing perspective of the camera). The user may point the projected light emanating from the device onto or in the vicinity of the object. An object closest to the location within the image of the beam reflection (whether the device light source is turned on or not) may be recognized. The device may subsequently determine whether the outline of the object is complete, including all edges and / or ensuring that no part of the object encroaches on the image perimeter.
[0127] Similar to the process of acquiring a comprehensive image of a page, if a complete object is identified within a camera-acquired image with adequate resolution, then the acquired image may be processed in a manner that assumes all viewable aspects of the selected object are present (e.g., classify the object using a trained neural network to determine its identity). If only a portion of the object is present within the camera-acquired image and / or if the object has inadequate resolution within the image, then the user may be instructed using haptic, visual, and / or audio methods to re-position and / or re-orient the portable device to better acquire all elements of the object. Instructional feedback to the user may continue until a complete representation of the object is viewable within a camera- acquired image.
[0128] In some cases, it may be advantageous to relax one or more criteria related to locating object edges within camera-acquired images. For example, if one or more small portions of the edges of an object or page cannot be recognized within an image, then object identification and / or determining the contents of a book page may generally proceed with a high level of confidence. A threshold for the portion of the one or more edges that might not be recognized may be predetermined and / or set according to a device user preference.
[0129] An aspect of the device and methods herein is the lack of use of a display on the portable device. A visual indication of a camera field of view reflected directly on the site being imaged is in contrast with devices (e.g., a mobile phone) that use a display to provide visual feedback of camera-acquired imagery.
[0130] As an example of the use of displays on other devices, when making a deposit to a bank, the printed front of a check may be scanned using a mobile device. As the device is manipulated during use, a user generally focuses on looking at a device display to track the field of view of the device camera, ensuring that the field of view includes an image of the front of the check. This process takes the visual focus of the user away from the region of the contents being scanned (i.e. the location of the check). Further, when the back side of the check is to be scanned, the user must generally switch visual focus to the physical location of the check in order to turn it over. Taking an image of the back of the check then requires the user to re-focus on the device display to take the additional image.
[0131] In contrast, devices and methods are described herein in which, in addition to the lack of a need for a display, the user may maintain focus on the contents being scanned. A book page may be observed as it is being scanned. Further, as a page turn is enacted to move on to the next page, there is no need to look elsewhere. The lack of a need to look at a display to monitor the camera field of view may allow a device to be held stationary, without significant movement, during content acquisition as well as when advancing from one page to the next. Thus, the systems and methods herein may intentionally avoid reliance on device displays or external screens to maintain continuous engagement with the target object, enabling novel feedback loops that preserve visual attention by a device user.
[0132] More specifically, when manipulated manually, the device may be held by one hand of a device user at a location and orientation best suited for camera-based imagery. Proprioception (i.e., an ability to perceive the location and movement of body parts) allows the manipulation of this hand while focusing on the projected light reflection and absent a need to visually track the location or orientation of the hand when manipulating the device. The second hand of the device user may then be used to manipulate the book, including turning pages. If the portable device is manipulated by other parts of the body (e.g., head, chest) then both hands may be available to turn book pages or manipulate objects in the environment of the device user.
[0133] Absent a need to look away, (e.g., at a display) during the process of acquiring camera-based imagery, there is no need to reposition the device at a location or orientation best suited for looking at a display or other indicator on the device itself. The lack of a need to shift visual focus, or to adjust device position to initially aim a camera and then look at a display, may simplify and speed the process of acquiring images.
[0134] Within further examples, an action enacted by a device processor may include transmitting available information to a remote device where, for example, further action(s) may be enacted. Transmitted information may include any or all camera images containing complete pages, acquisition times when camera images were acquired, and feedback elements to help direct the user produced by the device. A lack of a device user making any selection (or even device movement) within a prescribed time may also be conveyed to an external processor.
[0135] Interactions facilitated by a device may help to bring printed or displayed content “to life” by adding audio, additional visual elements, and / or vibrational stimulation felt by a hand (or other body part) of a device user. Printed content augmented with real-time interactive sequences including feedback related to content may not only provide machinebased guidance while reading, but also be “fun”, helping to maintain emotional engagement particularly while reading by, and / or to, a child. For example, the reading of a book may be augmented by adding queries, questions (for a parent or guardian, and / or the child), additional related information, sounds, sound effects, audiovisual presentations of related objects, real-time feedback following discoveries, and so on. Interactions with objects in books may be a shared experience with a parent, friend, guardian or teacher. Using a handheld device to control the delivery of serial content is more fully described in co-pending U.S. application Serial No. 18 / 091,274, filed December 29, 2022, the entire disclosure of which is expressly incorporated herein by reference. Sharing the control of advancing to a new page or panel to select objects when viewing a book or magazine is more fully described in U.S. Patent No. 11,652,654, filed November 22, 2021, the entire disclosure of which is expressly incorporated herein by reference.
[0136] Further, the portable device processor may include a “personality” driven by Al (i.e., artificial intelligence personality, AIP), transformer models and / or large language models (e.g., ChatGPT, Cohere, GooseAI). An AIP instantiated within a device may enhance user interactions by including a familiar appearance, interactive format, physical form, and / or voice that may additionally include personal insights (e.g., likes, dislikes, preferences) about the user.
[0137] Human-machine interactions enhanced by an AIP are more fully described in U.S. Patent No. 10,915,814, filed June 15, 2020, and U.S. Patent 10,963,816, filed October 23, 2020, the entire disclosures of which are expressly incorporated herein by reference. Determining context from audiovisual content and subsequently generating conversation by a virtual agent based on such context(s) are more fully described in U.S. Patent No.
[0138] 11,366,997, filed April 17, 2021, the entire disclosure of which is expressly incorporated herein by reference.
[0139] Whether used in isolation or as a part of a larger system, a portable device that is familiar to an individual (e.g., to a child) may be a particularly persuasive element of audible, haptic and / or visual rewards as a result of object selection (or, conversely, notifying a user that a selection may not be a correct storyline component). A handheld device may even be colored and / or decorated to be a child’s unique possession. Along similar lines, audible feedback (voices, one or more languages, alert tones, overall volume), and / or visual feedback (letters, symbols, one or more languages, visual object sizing) may be pre-selected to suit the preferences, accommodations (e.g., hearing abilities), skills and / or other abilities of an individual user.
[0140] When used in isolation (e.g., while reading a book), interactions using a device may eliminate requirements for accessories or other devices such as a computer screen, computer mouse, track ball, stylus, tablet or mobile device while making object selections and performing activities. Eliminating such accessories (often designed for an older or adult user) may additionally eliminate requirements by younger users to understand interactive sequences involving such devices or pointing mechanisms. When using the handheld device without a computer screen, interacting with images in books and / or objects in the real world (given the relative richness of such interactions that approaches that of screen-based interaction) may figuratively be described as using the device to “make the world your screen without a screen”.
[0141] As further aspects of devices and methods herein, viewable information and / or symbols within beam projections may encourage and / or mentally nudge device users to orient and / or position the device such that the information and / or symbols are most viewable (e.g., in focus, not skewed) by both the device user and the device camera. As a consequence, well-positioned and oriented camera-acquired images (i.e., at a working distance and viewing angle readily viewable by the user and camera) may facilitate computer vision processing (e.g., improving classification reliability and accuracy).
[0142] As an example, a symbol or pattern comprising a “smiley face” may be readily recognized by most individuals (perhaps even inherently by a young child). The lens and filter construction of the smiley -face light beam may be designed to present the image in focus, only within a workable (i.e., from an imaging perspective) distance from a reflective surface. A user may be unaware that an ability to readily view and / or identify projected beam patterns also enhances image quality for camera-based image processing. Further, as described above, if the projected image has a directionality or typical viewing orientation (e.g., text, image of a face), then most users may tend to manipulate the device such that the projected image is oriented in a viewing orientation typical for the object.
[0143] As described above, the portable electronic device may be manipulated by a user’s hand and / or other parts of the human body such as an arm, wrist, leg, foot, chest or head. Such positioning may be used to address accessibility issues for individuals with restricted upper limb and / or hand movement, individuals lacking sufficient manual dexterity to convey intent, individuals absent a hand, and / or during situations where a hand may be required for other activities.
[0144] Particular colors and / or color patterns may be avoided within visual interactions when devices are used by individuals with different forms of color blindness. Along similar lines, if an individual has a hearing loss over one or more ranges of audio frequencies, then those frequencies may be avoided or boosted in intensity (e.g., depending on the type of hearing loss) within audio interactions generated by the device. Haptic interactions may also be modulated to account for elevated or suppressed tactile sensitivity of an individual.
[0145] During activities that, for example, involve young children or individuals who are cognitively challenged, the complexity of instructions produced by the device may be restricted. For example, a young child may not fully understand a verbal instruction to “rotate counter clockwise”. Instead, instructions may be broken into two or more steps and / or visual symbols (e.g., arrows, animations) may be used preferentially.
[0146] Similarly, when there is no apparent attempt to point the beam following an instruction, a “questioning” indication (e.g., haptic feedback and / or buzzing sound) may be provided as an alerting prompt. Further aspects of simplified interactive elements are more fully described in U.S. Patent No. 11,334,178, filed August 6, 2021, and U.S. Patent No. 11,409,359 filed November 19, 2021, the entire disclosures of which are expressly incorporated herein by reference.
[0147] FIG. 1 shows an exemplary scenario in which a handheld device 10 is manipulated by the right hand 15 of a device user to direct a light I la, 11b emanating from the device 10 toward a spread 14a (i.e., two adjacent book pages 14b, 14c). The light I la, 1 lb is structured so that the light reflection, viewable by the device user, is in the form of an elongated dumbbell 12. The dumbbell shape 12 provides a linear orientation of the light reflection including circular ends that may be aligned with visual elements of an object. In this exemplary case, the user may be instructed to manipulate the device 10 so that the dumbbell-shaped light reflection 12 aligns with the central (e.g., gutter) region 14d of each spread 14a.
[0148] A device camera (not visible from the viewing perspective in FIG. 1, see for example FIG. 3) within the device 10 moves in conjunction with the projected light source and is aligned such that the projected light reflection 12 is in the central region of the camera’s field of view (FOV) 13. As a result, when the light reflection is manipulated to be in the central gutter region of a book spread, the camera’s FOV 13 is also directed centrally toward the spread 14a.
[0149] The device camera may periodically or continuously acquire images, allowing the device to provide user feedback regarding optimum device position and / or orientation for camera-based acquisition of a comprehensive object (e.g., page or spread). Optionally, the user may signal when a page may be acquired and / or indicate other conditions (e.g., the final pages of book are to be processed or a new book is about to be acquired) by depressing or releasing one or more device switches (e.g., a pushbutton 16). Alternatively, or in addition, other signaling methods may be employed such as shaking or orienting the handheld device (e.g., in a direction relative to the gravitational pull of the earth) sensed within camera-acquired images or an inertial measurement unit (IMU, not shown) and / or using verbal commands sensed by a microphone (not shown).
[0150] In the example shown in FIG. 1, the camera’s FOV 13 is sufficiently large to encircle the entire spread when the dumbbell-shaped light reflection (shown as a white region in FIG. 1) is directed toward the central gutter. This allows the complete contents of the book spread to be extracted from the camera image. The device may then indicate to the device user, a successful acquisition and / or the ability to move on to the next page within the book by aural (e.g., using the device speaker 17), visual (e.g., using one or more of the orb-shaped light sources 18a, 18b, 18c), haptic, or other methods.
[0151] If all edges of a page (including a spread) 14a are not (yet) detectable within a camera-acquired image or the detected edges encircle an area within the image that is too small to accurately extract printed materials, then instruction may be provided by the device to the user to further manipulate the device. As described in greater detail above, this may involve instructing the user to manipulate the device by translation (e.g., move up, down, left or right), rotation, tipping (e.g., left-right, or up-down), and / or moving the device closer to, or further from, the spread. Such continuous or periodic feedback may be provided aurally using a device speaker 17, haptically, and / or visually using one or more sources of light 18a, 18b, 18c including the pointing (i.e., projected) light source.
[0152] FIG. 2 illustrates some of the nomenclature used during the production and distribution of books where, as described more fully above, herein, a “book” may include a traditional (e.g., bound) book, magazine, comic or newspaper-like format. With some exceptions (e.g., pages that involve foldouts), a “spread” 20 comprises a pair of adjacent pages 21a, 21b that may contain drawings (e.g. at 22a) and / or images, and / or text (e.g., at 22b) and / or other symbols. A spread may include a left 21a and right page 21b.
[0153] From an image-processing perspective, the region of a spread 20 may be determined from (generally four) edges that encircle 24a, 24b, 24c, 24d the spread. Edges may be determined based on (generally abrupt) differences in spread content compared with background, expected geometry of a spread including linear edges 24a, 24b, 24c, 24d and right-angled comers, and / or distance measurements (if available) where a spread may be located closer to a device camera compared with background regions of a camera-acquired image.
[0154] A gutter 23a, 23b comprises a region in the middle of a spread where the two pages 21a, 21b of a spread 20 border each other. Within traditional books, book binding elements may be visible between left 21a and right 21b pages. Generally, pages may be curved in the region of the gutter making this region of a spread less easy to read. For this reason, authors and book producers typically avoid placing contents in the gutter.
[0155] FIG. 3 is an exploded-view drawing of a handheld device 35 showing exemplary locations for a projected light source 31a and a camera 36a. Such components may be internalized within the handheld device 35 during final assembly. This view of the device 35 also shows the backsides of three spherical displays 37a, 37b, 37c attached to the main body of the device 35.
[0156] The projected light source may, for example, comprise a lasing or non-lasing lightemitting diode 31a that may also include embedded and / or external optical components (not viewable in FIG. 3) to form, structure and / or collimate the projected light 30. Light generation electronics and optics may be housed in a sub-assembly 31b that provides electrical contacts for the projected light source and precision control over beam aiming.
[0157] Along similar lines, the process of image acquisition is achieved by light gathering optics 34a incorporated within a (threaded) housing 34b that allows further (optional) optics to be included in the light path for magnification and / or optical filtering (e.g., to reject reflected light emanating from the beam and / or radiation in the infrared portion of the electromagnetic spectrum). Optical components are attached to a camera assembly 36a (i.e., including the image-sensing surface) that, in turn, is housed in a sub-assembly providing electrical contacts for the camera and precision control over image detection direction.
[0158] An aspect of the exemplary configuration shown in FIG. 3 includes the light beam 30 and image-acquiring optics of the camera 34a pointing in the same direction 32. As a result, beam reflections off any viewable object occur within about the same region within camera-acquired images, regardless of the overall pointing direction and / or orientation of the portable device when manipulated by the device user.
[0159] Depending on relative alignment and separation (i.e., of the light source 31a and camera 33), the location of the beam reflection may be centered (at a typical working distance) or offset somewhat from the center of camera-acquired images. Additionally, small differences in beam location may occur at different distances from the device to a reflective surface due to the (designed to be small) separation at 33 between the beam source 31a and camera 36a. Such differences may be estimated using mathematical techniques analogous to those describing parallax.
[0160] FIG. 4 is an exemplary electronic interconnection diagram of a portable device 45 illustrating components at 42a, 42b, 42c, 42d, 42e,42f, 42g, 42h, 42i, 42j, 43, 44 and predominant directions for the flow of information during use (i.e., indicated by the directions of arrows relative to an electronic bus structure 40 that forms a backbone for device circuitry). All electronic components may communicate via this electronic bus 40 and / or by direct pathways (not shown) with one or more processors 43. Some components may not be required during different stages of user interactions.
[0161] A core of the portable, device may be one or more processors (including microcomputers, microcontrollers, application-specific integrated circuits (ASICs), field- programmable gate arrays (FPGAs), etc.) 43 powered by one or more (typically rechargeable or replaceable) batteries 44. As shown in FIG. 3, device elements also include a light beam generating component 42c (e.g., typically a light-emitting diode), and camera 42d to detect objects in the region of the beam (that may include a reflection produced by the beam). If embedded within the core of the device 45, both the beam source 42c and camera 42d may require one or more optical apertures and / or optical transparency (41b and 41c, respectively) through any device casing 45 or other structure(s).
[0162] During times when acoustic cues or feedback are generated, a speaker 42f (e.g., electromagnetic coil or piezo-based) 42f may be utilized. Similarly, during applications that might include audio-based user interactions, a microphone 42e may acquire sounds from the environment of the device. If embedded within the device 45, operation of both the speaker 42f and the microphone 42e may be aided by acoustic transparency through the device casing 45 or other structure(s) by, for example, coupling tightly to the device housing and / or including multiple perforations 41d (e.g., as further illustrated at 17 in FIG. 1).
[0163] During applications that, for example, include vibrational feedback and / or to alert a user to further manipulate the device, a haptic unit 42a (e.g., eccentric rotating mass or piezoelectric actuator) may be employed. One or more haptic units may be mechanically coupled to locations on the device housing (e.g., to be felt at specific locations on the device) or may be affixed to internal support structures (e.g., designed to be felt more generally throughout the device surface). Similarly, during times when visual feedback is employed, one or more light sources 42b (e.g., LEDs) may be utilized to indicate, for example, one or more directions to move the device, successful acquisition of a book spread, and so on. Light sources maybe made more visible using diffusive elements (e.g., three diffusing orbs as illustrated by at 42b). They may also include filters or other optical elements to structure light (e.g., producing directional arrows as illustrated by at 42b). The one or more light sources may be affixed and / or exterior to the main device body (as shown at 42b). Additionally, one or more optically transparent windows 41a may be provided when sources are within internal structures (e.g., embedded within a device body).
[0164] During typical interactions, a user may signal to the device at various times such as when ready to acquire another spread, acquire the first spread of a new book, and so on. User signaling may be indicated by verbal feedback sensed by a microphone, as well as movement gestures or physical orientation of the device sensed by an IMU 42g. Although illustrated as a single device at 42g, different implementations may involve distributed subcomponents that, for example, separately sense acceleration, gyroscopic motion, magnetic orientation, and gravitational pull. Additionally, subcomponents may be located in different regions of a device structure (e.g., distal arms, electrically quiet areas) to, for example, enhance signal-to-noise during sensed motions.
[0165] User signaling may also be indicated using one or more switch devices including one or more pushbuttons, toggles, contact switches, capacitive switches, proximity switches, and so on. Such switch-based sensors may require structural components at or near the surface of the device 41e to convey forces and / or movements to more internally located circuitry.
[0166] Telecommunications to and from the device 45 may be implemented using Wi-Fi 42i and / or Bluetooth 42j hardware and protocols (e.g., each using different regions of the electromagnetic spectrum). During exemplary scenarios that employ both protocols, shorter- range Bluetooth 42j may be used, for example, to register a device (e.g., to identify a Wi-Fi network and enter a password) using a mobile phone or tablet. Subsequently, Wi-Fi protocols may be employed to allow the activated device to communicate directly with other, more distant devices and / or the World Wide Web.
[0167] FIG. 5 shows examples of reflection image geometries 50a, 50b, 50c, 50d, 51a, 51b, 51c, 5 Id, 52a, 52b that may be generated when pointing structured light produced by the device toward a page or other object. Due to the use of a white background in FIG. 5, reflected light patterns may appear reversed compared to those reflected off a book spread. In other words, dark regions within reflection geometries depicted in the figure may appear as bright-light regions (e.g., white or another color, or combination of colors) within actual reflections.
[0168] Pointing characteristics that may be incorporated within reflected images (i.e., produced by structured light) include a directional orientation that may be aligned with the directional orientation of a (e.g., linear) object element such as a book spread gutter (e.g., 50a, 50b, 50c, 50d, 51a, 51b, 51c), simple pinhole geometries, for example, generated by one or more masks in the light path producing the reflected image (e.g., 50d, 52a), a large (e.g., easy to find) reflection (e.g.52a, 52b), and / or whimsical designs such as a pointing finger 5 Id or smiley face 52b. As depicted in FIG. 10 below, the same device light source used to point toward an object to be imaged may also be used to provide user feedback.
[0169] FIG. 6 illustrates a scenario in which the device user may be instructed to move the portable device to the left (i.e., a translation movement). In this example, content within the left 62a and right 62b pages of a book spread each includes both text and a drawing 63a, 63b. The user (not shown) manipulated the handheld device (not shown) such that the reflection of the projected light (i.e., structured to generate a smiley face 60) is to the right of the spread gutter region 64.
[0170] As a result, the FOV of the camera (shown as a dashed line outline 61) is sufficiently to the right side of the overall spread that, although all other spread edges are present and oriented appropriately (e.g., horizontally and vertically) within the FOV 61, the left edge 65 and two leftmost corners 66a, 66b are absent within the camera-acquired image. As a result of determining an absence of a left edge (as well as leftmost corners 66a, 66b), instruction(s) to the user may include moving the device to the left.
[0171] FIG. 7 illustrates an exemplary scenario in which the device user may be instructed to move (i.e., rotate) the handheld device clockwise. Content within the left 72a and right 72b pages of a book spread each includes both text and a drawing 73a, 73b. The user has manipulated the handheld device such that the reflection of the projected light (i.e., structured to generate a smiley face 70) in the vicinity of the gutter region 74. However, the co-aligned camera and light source of the device generate a camera FOV 71 and a smiley face that are rotated counterclockwise compared with the spread. (Note that the smiley face within any camera image does not appear rotated because the camera and light source are co-aligned.) As a result, the FOV of the camera (shown as a dashed line outline 71) is sufficiently to the right side of the overall spread that, although all other spread edges are present and oriented approximately (e.g., horizontally and vertically) within the FOV 71, portions of the left edge 75a, 75b (including two leftmost comers 76a, 76b) are absent from the camera- acquired image. As a result, instruction(s) to the user may include rotating the device to the right (e.g., clockwise).
[0172] FIG. 8 illustrates an exemplary scenario in which, although all 4 edges 87a, 87b, 87c, 87d as well as left page 82a and right page 82b of a page spread may be identified within the camera FOV 81, the image of the page spread may occupy an insufficient number of pixels within the overall camera FOV 81 to reliably apply computer vision methods (e.g., OCR, image recognition). Absent an ability to optically zoom, this issue may be resolved by moving the portable device (i.e., or, more specifically, the device camera) closer to the page spread. The device user may be instructed to move the device closer (e.g., an audible command to “move toward”) to enlarge the image of the object within the camera FOV.
[0173] An additional aspect illustrated in FIG. 8 relates to the size (and focus) of the reflected light image (i.e., smiley face at 80) viewable by the device user. For a given size of objects being acquired by the portable device (e.g., a series of pages within a book), a device user may become familiar with the approximate size of a reflected image 80 that routinely allows objects (e.g., pages) to be acquired. Thus, the size of the reflected image along with positioning the reflected image near the center of an object (e.g., directed at a book spread gutter 84) provides visual feedback, encouraging the device user to manipulate the device in a manner suitable for camera-based acquisitions.
[0174] FIG. 9 illustrates an exemplary scenario showing a camera FOV 91 when the device is held by a device user (not shown) below a book spread (i.e., not close to a normal at the center of the page spread) and tipped upward to direct the projected light (i.e., smiley face 90) toward the center of a book spread, 92a, 92b. As a result of viewing the book spread away from a normal to the surface of the spread, the spread appears skewed.
[0175] The degree to which the device is tipped away from the normal to the surface may be estimated by comparing the length of the front edge of the page spread 92b with the length of the back edge 92a. Additionally, edges of the book spread on the left 97a and right 97b sides point inward, becoming progressively closer as the distance away from the camera becomes greater. Compared with an image of the book spread taken from a location normal to the spread surface (e.g., as shown in FIG. 6), the number of pixels covered by the spread decreases as the viewing angle by the device camera becomes more acute. If too severe, image skew may challenge further processing of the image.
[0176] FIG. 10 shows exemplary symbols, and words or phrases the may be used to instruct a device user how to manipulate the device to better acquire an image of an object. Motion symbols 100a, 100b, 101a, 101b, 102a, 102b, 103a, 103b, 104a, 104b may, for example, be structured as a pattern within the projected light emanating from the device (i.e., producing a reflected image). Instructive words or phrases 100c, lOOd, 101c, lOld, 102c, 102d, 103c, 103d, 104c, 104d, 105c, 105d may, for example, also be displayed within the projected light, and / or played on a device speaker.
[0177] Instructional words and / or motion symbols may be divided into those that direct: 1) translational movements, moving the device in any direction from one location to another, and 2) rotational movements, moving the portable device about its center. Examples of translational movement instructions include up 100a, 100c versus down 100b, lOOd; left 101a, 101c versus right 101b, lOld; and toward 102a, 102c versus away 102b, 102d from the object of interest. Examples of rotational movement instructions include rotate the device to the left 103a, 103c in the counterclockwise direction versus right 103b, 103d in the clockwise direction (e.g., analogous to roll on an airplane), tip the device forward 104a, 104c versus backward 104b, 104d (analogous to pitch on an airplane), and turn left 105a, 105c versus right 105b, 105d (analogous to yaw on an airplane).
[0178] FIG. 11 is a flow diagram illustrating exemplary steps employed to acquire a full image of a book spread 111b (or other object) using a portable device 11 le. A user may be instructed to manipulate the device 11 le such that a light 111c projecting from the device is directed toward the gutter region of a book spread 11 lb. An acquired camera image I l la may then be used as input to computer vision methods (e.g., various masks including Canny and Sobel and / or a trained neural network 112 to identify an object and / or its edges (a book spread in this exemplary case).
[0179] If identified edges account for all edges of the spread and / or edges do not touch the periphery of the camera-acquired image 110c, and the area occupied by the spread covers a sufficiently large portion of the overall camera image 1 lOd, then the camera-based image may be labelled as successfully containing a book spread. Optionally, the user may be notified of the successful acquisition by aural 118b, visual 118a and / or haptic methods. Otherwise, based on the locations(s) of image regions containing edges (or, conversely, absent detected edges), one or more instructions to manipulate the device may be formulated 1 lOe and broadcast to the user using one or more light sources 118a (including the projected light source), a device speaker 118b and / or a haptic unit. Steps in this spread acquisition process include:
[0180] 1) at 110a, using a camera (not visible) and focusing optics 11 Id (depicted apart from the body of the device for illustration purposes only), when the device is manipulated by the user such that light 111c emanating from the device points toward a book spread 111b (e.g., ideally directed at the gutter region), acquire a camera-based image that encompasses the camera’s field-of-view at 11 la;
[0181] 2) at 110b, the camera-acquired image may be input to a neural network trained to detect object surfaces and / or edges, and / or additional computer vision methods may be used to identify object edges 112;
[0182] 3) at 110c, if detected edges comprise all edges of an object (e.g., fully encircling a book spread) and / or edges do not touch peripheral pixels of the camera image using threshold values 113c for the completeness of determined edges and / or a number of object edge pixels intersecting with image edge pixels (e.g. accounting for image noise), then proceed to determine if the image of the object is sufficiently large 113b, otherwise proceed to formulate user feedback 113a based on incomplete object components;
[0183] 4) optionally (as indicated by the dashed line outline) at 1 lOd, if the area of the object (e.g., spread) within the camera-based image is sufficiently large to accurately determine content then proceed 114a to classify the image as a successfully acquiring object content, otherwise proceed 114 to formulate user feedback for the undersized image (e.g., to hold the device closer to the object and / or tip the device to reduce skew);
[0184] 5) at 1 lOe, formulate one or more instructions (e.g., involving device translation, rotation, and / or zoom) for the user to manipulate the device based on the presence and / or orientation of the object components including edges;
[0185] 6) at 1 lOf, present the one or more instructions via the handheld device one or more light sources 116a (optionally including the light used for pointing) and / or device speaker 116b;
[0186] 7) at 110g; if the camera-based image of the object 117b is deemed complete and acceptable (e.g., in size), label (e.g., within image metadata) and / or store the image 117a, for example, in a database of successfully acquired images of comprehensive objects;
[0187] 8) optionally (indicated by the dashed line outline) at 1 lOh, indicate to the user that the object has been successfully acquired using the one or more light sources 118a (optionally including the light used for pointing), haptic unit (not shown), and / or device speaker 118b;
[0188] FIG. 12 is a flow diagram illustrating exemplary steps using a portable device 12 If to acquire an identified image (e.g., with confidence) of an object in the environment of a device user. A user may manipulate the device 12 If so that a light 121 emanating from the device is directed toward an object 121b. An image 121a acquired by a camera co-aligned with the projected light 121 d may then be used as input to a neural network 122 (e.g., CNN) trained to identify objects.
[0189] If the neural network identifies the object (e.g., as a shoe) with sufficient confidence (e.g., greater than a predetermined confidence score 123c) then the camera-based image 121a may be tagged as containing a successfully identified object and, optionally, the user may be notified of the successful acquisition by aural 127b, visual 127a and / or haptic (not shown) methods. Otherwise in this case, based on the front portion of the shoe not being fully visible within the camera-acquired image (e.g., one or more edges of the shoe have been acquired by one or more pixels in the periphery of the image), one or more instructions to manipulate the device may be formulated (e.g., to move the device to the left) and broadcast to the user using one or more light sources 125a (e.g., including the projected light source), and / or a device speaker 125b. Steps in this process to acquire an identified object include:
[0190] 1) at 120a, using a camera (not visible) and focusing optics 121e (depicted apart from the body of the handheld device for illustration purposes), when the device is manipulated by the user such that light 12 Id emitted from the device points toward an object 111b such as a shoe (absent an ability to see the toe portion of the shoe 121c), acquire a camera-based image encompassing the camera’s field-of-view I l la;
[0191] 2) at 120b, the camera-acquired image may be input to a neural network 112 trained to identify objects;
[0192] 3) at 120c, if the neural network confidently identifies the object (that may include a predetermined threshold confidence level 123c) then proceed to label the image 123a, otherwise proceed to formulate user feedback 123b based on object components that may not be visible in the camera-acquired image;
[0193] 4) at 120d, formulate one or more instructions that may include directional symbols 124a and / or words 124b to, in this case, move the device to the left;
[0194] 5) at 120e, present the one or more instructions using one or more device light sources 125a and / or device speaker 125b;
[0195] 6) at 120f; if the camera-based image of the object has be identified with a sufficient degree of confidence, label (e.g., within image metadata) the image 126 as containing an identified object;
[0196] 7) optionally (indicated by the dashed line outline) at 120g, indicate to the user that the object in the image has been successfully identified using the one or more light sources 127a (optionally including the light used for pointing) and / or device speaker 127b;
[0197] FIG. 13 illustrates intelligent identification of different interactive blocks or “chunks” of text within book pages. The left page 132a of an illustrated book spread shows distinctive blocks of text that are: 1) printed primarily within a vertical, leftmost column 130a, 2) oriented linearly in a non-horizontal direction 130b, and 3) printed with varying individual character orientations 130c following the circumference of a circle 133. Similarly, the right page 132b shows a selectable block of text 130d that may be identified as a distinct interactive chunk and distinguished from other text based on: 1) its predominantly horizontal block structure, 2) unique font compared with other text, 3) distinctive bold appearance, 4) sentence construct, and / or 4) containment within a text bubble 135 directed toward a drawn interactive character 134.
[0198] During an interactive sequence, the device user (not shown) may direct a light pointer (e.g., smiley face 131) toward a body of text. Even though the pointer 131 may only partially cover a few words of the sentence “How is the circumference of a circle related to its diameter?” 130a the device may identify (e.g., using natural language processing and / or Al) that a complete phrase or sentence 130a is present within a predominantly vertical column of the pointing region. Thus, unless the context of the interactive scenarios suggests a more focused pointing region (e.g., a letter or word), the full sentence (i.e., beyond the area covered by a pointer) may be used as a basis for subsequent interactions.
[0199] In contrast, if the user were to point toward a region at or near the word “diameter” 130b, then the absence of any other words in the region with the same orientation, font size, and / or related meaning would direct one or more subsequent interactions to be based on this single word. Similarly, the phrase, “circle circumference” 130c may be treated as a single linguistic chunk based particularly on its varying radial orientation of individual characters (i.e., following a circumferential path). In each of these cases, the selected word or phrase may be a component of more complete sentences generated by the device (e.g., using natural language processing and / or Al) during interactive sequences.
[0200] As just described, any of one or more of the unique presentation features (e.g., horizontal orientation, font, bold appearance, encased within a bubble 135) of the text 130d on the rightmost page 132b may indicate that the full sentence, “You multiply the diameter by pi, which is approximately 3.1416.” is to be used as a basis for one or more subsequent interactions. Alternatively, or in addition, a user may encircle at least a majority of the text or text bubble 135 to select the entire (i.e., distinct) block of text. Intelligent identification of text blocks may enhance intuitive interactive sequences and / or reduce requirements for precision pointing.
[0201] FIG. 14 illustrates an exemplary scenario in which two edges of a camera-based image of a page at 142 (e.g., a book spread including a gutter at 145) have been acquired by a region of camera sensor pixels that includes one or more pixels (at 143a and 143b) in the periphery of the camera image 140. In this figure, pixels at the periphery of the camerabased image are indicated by an “X” superimposed on the pixel region whereas interior pixels of a camera image are represented by open squares (e.g., at 141a).
[0202] The intersection of the one or more determined edges of the camera-acquired object (i.e., page) with one or more peripheral pixels of the camera sensor array (at 143a and 143b) suggests that at least some portion of the object (e.g., at 144) may be beyond the field of view of the camera 140. In this exemplary case, since edges of the object 142 intersect with peripheral pixels on the upper side of the camera image (i.e., field of view 140), one or more instructions to move the portable device (or, more specifically, the included device camera) translationally in the upward direction (and / or rotate counter-clockwise) may be conveyed to the device user in order to acquire a comprehensive image of the object 142. A projected (i.e., computed) image of the page with the camera moved sufficiently upward and / or rotated may contain all edges of the page (and additionally, a sufficient number of pixels) for further processing.
[0203] In FIG. 14, a camera sensor array comprising thirty -two by sixteen pixels (for a total of five hundred twelve pixels) is shown for illustration purposes. Video cameras (even those that are modest in cost) may typically contain sensor arrays comprised of millions of pixels (i.e., sensor elements). Thus, peripheral pixels located at the edges of the camera sensor array generally occupy a thin sliver of pixel elements (e.g., compared with the illustration of peripheral pixels shown in FIG. 14). Additionally in FIG. 14, the aspect ratio of horizontal 146b versus vertical 146a pixels is sixteen to nine (16:9). This image aspect ratio (as well as four to three, 4:3) is common in modern-day (consumer grade) video cameras.
[0204] FIG. 15 is a flow diagram illustrating exemplary steps using a portable device 15 Id to view and / or acquire the imagery of book pages (e.g., including text and drawings) in sequential order. This process may help to ensure that parts of a storyline are not missed and includes using Al 153 to determine if images appear to be in a logical and / or consecutive sequence. These steps may be particularly useful when a book is not known to the device (e.g., pages have not been previously scanned or reviewed).
[0205] If a user “skips around” to view different pages or portions of a page in an order that does not follow an author’s intended storyline, then the user may be prompted 150f in a manner that encourages following the storyline. Such guidance by the device may encourage a more complete understanding of the book’s contents. Alternatively, or in addition, a user may be advised to point the device 15 Id toward pages, page regions or page spreads in an order that allow the device to be fully aware of book contents, allowing the context of a story to be known to the device. By making a complete storyline and / or context available to an Al, more engaging and / or directed interactive sequences may be generated. Steps in this process to sequentially view and acquire a storyline include:
[0206] 8) at 150a, using a device camera (not visible) and focusing optics 151c (depicted apart from the body of the handheld device for illustration purposes), acquire an initial image 151a as the device 15 Id is manipulated by the user to point toward a page that, in this case, includes a hockey player shooting a puck 151b;
[0207] 9) at 150b, using the device camera (not visible), acquire another image 152a as the device 152c is further manipulated by the user to point toward another page that, in this case, contains a hockey goalie 152b who may be in a position to stop the puck;
[0208] 10) at 150c, the camera-acquired images may be input to an artificial intelligence (e.g. neural network 153) trained to determine if camera-acquired imagery (e.g., containing text and / or images) represent sequential activities or events; 11) at 150d, if images are determined to represent sequential activities, then continue 154a to acquire further imagery, otherwise proceed 154b to steps involved in advising the user that imagery may be incomplete or have been acquired out of sequence;
[0209] 12) optionally (indicated by the dashed line outline) at 150e, formulate one or more instructions (e.g., verbal, directional symbols 155a and / or instructional text) that may direct the user to view content sequentially (e.g., following a storyline);
[0210] 13) at 150f, present one or more general prompts (e.g., sound effect(s), light indication(s)) or formulated instructions using one or more device light sources 125a, a device speaker 125b and / or a device haptic unit;
[0211] 14) optionally (indicated by the dashed line outline) at 150g; indicate to the user, using visible 157a, audible 157b and / or haptic prompts) that camera-acquired imagery appears to represent a sequential storyline; and
[0212] 15) at 150h, transfer (or redirect a pointer) to indicate that the most recently acquired camera image 158a is available for comparison 158b with the next image that is acquired at 150b as the search for sequential content is repeated 159b.
[0213] The foregoing disclosure of the examples has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many variations and modifications of the examples described herein will be apparent to one of ordinary skill in the art in light of the above disclosure. It will be appreciated that the various components and features described with the particular examples may be added, deleted, and / or substituted with the other examples, depending upon the intended use of the examples.
[0214] Further, in describing representative examples, the specification may have presented the method and / or process as a particular sequence of steps. However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular sequence of steps described. As one of ordinary skill in the art would appreciate, other sequences of steps may be possible. Therefore, the particular order of the steps set forth in the specification should not be construed as limitations on the claims.
[0215] While the invention is susceptible to various modifications and alternative forms, specific examples thereof have been shown in the drawings and are herein described in detail. It should be understood that the invention is not to be limited to the particular forms or methods disclosed, but to the contrary, the invention is to cover all modifications, equivalents and alternatives falling within the scope of the appended claims.
Claims
We claim:
1. A device for acquiring a comprehensive image of a page comprising: a device body configured to be manipulated by a device user; electronic circuitry within the device body that includes a device processor; a light source configured to emit a light away from the device body to generate a light reflection on a surface at a light reflection location; and a camera operatively coupled to the device processor and aligned such that a field of view of the camera includes the light reflection location, wherein the device processor is configured to: acquire, using the camera, a camera image when the device is manipulated by the device user to project the light reflection onto or adjacent to the page; identify, within the camera image, one or more identified page edges; and either: a) if the one or more identified page edges comprise all page edges, label the camera image as the comprehensive image of the page, or b) if a) is not met, determine one or more device movement instructions to manipulate the device such that a predicted camera image of the page upon following the one or more device movement instructions contains all of the page edges.
2. A device for acquiring a comprehensive image of a page comprising: a device body configured to be manipulated by a device user; electronic circuitry within the device body that includes a device processor; a light source configured to emit a light away from the device body to generate a light reflection on a surface at a light reflection location; and a camera operatively coupled to the device processor and aligned such that a field of view of the camera includes the light reflection location, wherein the device processor is configured to: acquire, using the camera, a camera image when the device is manipulated by the device user to project the light reflection onto or adjacent to the page; identify, within the camera image, one or more page edge pixels; determine a pixel count of the one or more page edge pixels that are one of camera image edge pixels; and either:a) if the pixel count exceeds a predetermined pixel count threshold, determine one or more device movement instructions to manipulate the device such that a predicted page edge pixel count in a predicted camera image following manipulation that are one of the camera image edge pixels is less than the predetermined pixel count threshold, or b) if a) is not met, label the camera image as the comprehensive image of the page.
3. The device of claim 1 or 2, wherein the page comprises one of a book page, a book spread, a front cover, a back cover, a magazine page, a magazine spread, a newspaper page, a newspaper spread, a scrapbook page, a brochure, a painted surface, a printed surface, a textured surface, a surface enhanced with three-dimensional elements, and a flexible printed surface.
4. A device for acquiring a comprehensive image of an object comprising: a device body configured to be manipulated by a device user; electronic circuitry within the device body that includes a device processor; a light source configured to emit a light away from the device body to generate a light reflection on a surface at a light reflection location; and a camera operatively coupled to the device processor and aligned such that a field of view of the device camera includes the light reflection location, wherein the device processor is configured to: acquire, using the camera, a camera image when the device is manipulated by the device user to project the light reflection onto or adjacent to the object; identify, within the camera image, one or more object edge pixels; determine a pixel count of the one or more object edge pixels that are one of camera image edge pixels; and either: a) if the pixel count exceeds a predetermined pixel count threshold, determine one or more device movement instructions to manipulate the device such that a predicted object edge pixel count in a predicted camera image following manipulation that are one of the camera image edge pixels is less than the predetermined pixel count threshold, or b) if a) is not met, label the camera image as the comprehensive image of the object.
5. The device of any one of claims 1, 2, and 4, wherein the device is configured to be manipulated by one of a hand of the device user, a head of the device user, a wrist of the device user, an arm of the device user, a shoulder of the device user, a leg of the device user, and a chest of the device user.
6. The device of claim 4, wherein if a determined number of camera image pixels detecting the object is determined by the device processor to be less than a predetermined minimum number of the camera image pixels, then determine one or more additional device movement instructions to further manipulate the device such that an additional predicted camera image of the object following the one or more additional device movement instructions contains greater than the predetermined minimum number of the camera image pixels.
7. The device of claim 4, wherein identifying the one or more identified object edges comprises using the camera image as input to one of: a neural network trained to identify object edges, and an edge detection process comprising one or more of a Canny edge detector, a first- order gradient operator, a Sobel operator, and a phase stretch transform.
8. The device of claim 1, wherein identifying the one or more identified page edges comprises using the camera image as input to one of: a neural network trained to identify page edges, and an edge detection process comprising one or more of a Canny edge detector, a first- order gradient operator, a Sobel operator, and a phase stretch transform.
9. The device of claim 2, wherein identifying the one or more identified page edge pixels comprises using the camera image as input to one of: a neural network trained to identify object edges, and an edge detection process comprising one or more of a Canny edge detector, a first- order gradient operator, a Sobel operator, and a phase stretch transform.
10. The device of any one of claims 1, 2, and 4, wherein the light reflection generated by the light source comprises a light reflection image.
11. The device of claim 10, wherein the light reflection image is in focus at a target focal distance from the camera.
12. The device of any one of claims 1, 2, and 4, wherein the light source produces a light reflection image using one of an array of light emitting diodes each operatively coupled to the device processor, a digital light processor operatively coupled to the device processor, an optical filter in a light path of the light, and a liquid crystal filter in the light path of the light operatively coupled to the device processor.
13. The device of claim 12, wherein the light reflection image includes one or more of one or more lines, a dumbbell, one or more rectangles, one or more dots, one or more ellipses, one or more arrows, a pointing finger, a Gaussian profile, and a smiley face.
14. The device of claim 12, wherein the light reflection image is in focus at a target focal distance from the camera.
15. The device of claim 12, wherein the one or more device movement instructions comprise the light reflection image that includes one or both of a translational arrow and a rotational arrow.
16. The device of any one of claims 1, 2, and 4, wherein the one or more device movement instructions comprise one or both of a translational movement of the device and a rotational movement of the device.
17. The device of any one of claims 1, 2, and 4, further comprising an output device on the device body operatively coupled to the device processor for outputting the one or more device movement instructions to the device user.
18. The device of claim 17, wherein the output device comprises a speaker, and wherein the one or more device movement instructions comprise one or more words played on the speaker including left, move left, right, move right, up, move up, down, move down, toward, move toward, move forward, move out, away, move away, clockwise, rotateclockwise, counterclockwise, rotate counterclockwise, tip forward, tip back, turn left, swing left, turn right, and swing right.
19. The device of claim 17, wherein the output device comprises a speaker, and wherein the one or more device movement instructions comprise changing one or more of a sound frequency, a repeat frequency, and an amplitude of one or more sounds played on the speaker.
20. The device of claim 17, wherein the output device comprises one or more indicating light sources, and wherein the one or more device movement instructions comprise changing one or more of a color, a repeat frequency, and an amplitude of one or more visible light indications emitted from the one or more indicating light sources.
21. The device of claim 20, wherein the one or more indicating light sources each comprises a light emitting diode.
22. The device of claim 17, wherein the one or more device movement instructions comprise outputting instructions on the output device comprising one or more of left, move left, right, move right, up, move up, down, move down, in, move in, move forward, out, move out, away, move away, clockwise, rotate clockwise, counterclockwise, rotate counterclockwise, tip forward, tip back, swing left, and swing right.
23. The device of claim 17, wherein the output device comprises a haptic unit, and wherein the one or more device movement instructions comprise changing one or more of a vibrational frequency, a repeat frequency, and an amplitude of one or more vibrations produced by the haptic unit.
18. The device of claim 5, wherein upon determining the one or more device movement instructions, a prompt sound is played on a speaker operatively coupled to the device processor.
23. The device of claim 1 or 2, further comprising an output device on the device body operatively coupled to the device processor and wherein, upon labelling the cameraimage as the comprehensive image of the page, the device processor causes the output device to generate a completion indication to the device user.
24. The device of claim 4, further comprising an output device on the device body operatively coupled to the device processor and wherein, upon labelling the camera image as the comprehensive image of the object, the device processor causes the output device to generate a completion indication to the device user.
25. The device of claim 23 or 24, wherein the output device comprises a speaker, and wherein the completion indication comprises a sound generated by the speaker.
26. The device of any one of claims 1, 2, and 4, further comprising a communications interface operatively coupled to the device processor and wherein the device processor is configured to transmit the comprehensive image to a remote processor via the communications interface.
27. The device of any one of claims 1, 2, and 4, wherein the device processor is configured to control an intensity of the light source by one or both of regulating a light beam driving current, and a pulse width modulation of the light beam driving current.
28. The device of any one of claims 1, 2, and 4, wherein the light source is configured to emit one or more of collimated, non-coherent, diverging, and patterned light.
29. The device of any one of claims 1, 2, and 4, further comprising one or more indicator light sources on the device body operatively coupled to the device processor.
30. The device of claim 29, wherein the one or more device indicator light sources each comprises a light emitting diode.
31. The device of claim 29, wherein the device processor is configured to generate one or more camera movement light indications on the one or more indicator light sources comprising one or more of changing a color of one or more of the one or more device indicator light sources, changing a blink frequency of the one or more deviceindicator light sources, and changing a light intensity of the one or more device indicator light sources.
32. The device of claim 31, wherein the one or more movement instruction light patterns comprise one or more of a directional arrow indicating a target direction, a rotational arrow indicating a target orientation, one or more directional words indicating the target direction, and one or more rotational words indicating the target orientation.
33. The device of any one of claims 1, 2, and 4, wherein the light source comprises an array of light emitting diodes each controlled by the device processor, producing a projected light pattern comprising the light reflection.
34. The device of any one of claims 1, 2, and 4, further comprising an optical filter in the light path of the light source and operatively coupled to the device processor for producing a patterned light reflection comprising the light reflection.
35. The device of claim 34, wherein the optical filter comprises a liquid crystal array controllable by the device processor.
36. A device for acquiring a comprehensive image of a target surface, comprising: a device body configured to be manipulated by a device user; electronic circuitry within the device body including a device processor; a light source configured to emit a structured light pattern away from the device body to produce a light reflection at a reflection location on the target surface; and a camera co-aligned with the light source and operatively coupled to the device processor, the camera having a field of view that includes the reflection location; wherein the device processor is configured to: acquire, using the camera, a camera image when the structured light reflection is projected onto or adjacent to the target surface; determine, using computer vision and / or a trained neural network, whether the image contains a complete representation of the target surface based on detected edges and resolution sufficiency;if the representation is complete, label the image as comprehensive and initiate a completion indication via at least one of an audio, visual, or haptic output; and if the representation is incomplete, generate and output one or more movement instructions to guide the device user to reposition the device to acquire a comprehensive image, wherein the movement instructions are conveyed through at least one of: modulated projected light patterns, audio commands, or haptic signals.
37. The device of claim 36, wherein the structured light comprises a projected image containing one or more directional or symbolic elements selected from the group consisting of: arrows, smiley faces, lines, pointing fingers, and animated icons, wherein the image changes in response to the device’s relative position to the target surface.
38. The device of any one of claims 1, 2, 4, and 36, wherein the device lacks a visual display screen for monitoring camera images.
39. The device of any one of claims 1, 2, 4, and 36, wherein the device body is configured to be mounted on a user's head, chest, wrist, or arm via a support structure, enabling hands-free operation.
40. The device of claim 36, wherein the type and characteristics of feedback are adapted based on stored user accessibility preferences, including compensation for color blindness, hearing impairment, or tactile sensitivity.
41. The device of claim 36, wherein the device processor is further configured to determine whether sequential images of a multi-part media (such as a book or instructional set) are being acquired in narrative or logical order, and to prompt the user if a deviation from expected sequence is detected.
42. The device of claim 36, wherein the structured light pattern comprises one or more symbolic or directional visual elements selected from the groupconsisting of: arrows, ellipses, lines, pointing fingers, smiley faces, and animated icons, the pattern being dynamically modulated by the device processor to convey at least one of direction or orientation for movement of the device to the device user43. The device of claim 36, wherein the device is devoid of a visual display screen configured to render camera images or graphical user interfaces, and wherein all user feedback for device movement is conveyed via at least one of modulated projected light patterns, auditory signals, or haptic outputs.
44. The device of claim 36, wherein the device body is further configured to be supported on a body part of the device user, the body part selected from the group consisting of a head, chest, wrist, arm, leg, and shoulder, and wherein the device body includes one or more attachment structures comprising a strap, band, holster, or harness.
45. The device of claim 36, wherein the device processor is further configured to adjust at least one of a color of the structured light pattern, a pitch or amplitude of an audio signal, or a vibration frequency of a haptic output based on one or more accessibility parameters associated with the device user46. The device of claim 36, wherein the device processor is further configured to analyze a sequence of camera-acquired images and to determine whether the images correspond to a logically sequential or narrative progression, and wherein the device processor provides a prompt to the device user upon detecting a deviation from the sequential progression.
47. A method for acquiring and interacting with a target surface using a portable device comprising a light source and a co-aligned camera, the method comprising:(a) projecting a structured light pattern from the light source onto or adjacent to the target surface;(b) acquiring a camera image including the structured light reflection;(c) determining, via a processor, whether the image includes a complete representation of the target surface based on edge detection and resolution sufficiency;(d) if the image is complete, labeling the image as comprehensive and initiating an output indication to the device user via audio, visual, or haptic feedback; and(e) if the image is incomplete, generating one or more movement instructions to reposition the device to improve the image, and conveying the instructions via at least one of: modulated structured light, auditory cues, or haptic feedback.
48. A method for acquiring and interacting with printed or visual media using a portable device comprising a light source and camera, the method comprising: projecting structured light from the device onto or near a surface; acquiring an image from a camera co-aligned with the projected light; determining, using a processor, whether the image includes a complete and high-resolution view of the surface; if incomplete, providing instructional feedback through modulated projected light, audible instructions, or vibrations; and if complete, storing the image and initiating a real-time interaction sequence based on the acquired image content.
Citation Information
Patent Citations
Systems and methods to specify interactive page locations by pointing a light beam using a handheld device
US11989357B1
Method of determining usability of a document image and an apparatus therefor
US20020150279A1
Method and device for capturing a document
US20150278594A1
Adaptive Enhancement of Scanned Document Pages
US20190089865A1
Edge identification of documents within captured image
US20240112348A1