Real-time Object Detection and Tracking
By using lightweight classifiers and on-device tracking techniques to detect objects on mobile devices and send image data only when the device is stationary, the problem of large consumption of image processing resources and poor user experience on mobile devices is solved, and faster content presentation and more efficient object recognition are achieved.
Patent Information
- Application Number
- CN201880093287.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-05-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2039-01-06
AI Technical Summary
In the prior art, object detection and content presentation in images are inefficient and user experience is poor, especially on mobile devices, image processing resources are consumed largely and bandwidth utilization is insufficient.
By using lightweight classifiers and on-device tracking techniques on mobile devices, image data is sent to the content distribution system only when the device is stationary, content distribution system is used for efficient content identification and cache, and object locations are tracked in real time to present relevant content.
It improves the speed and user experience of content presentation, reduces network communication consumption and computing resources, and achieves faster and more efficient object identification and content selection.
Smart Images

Figure CN112088377B_ABST
Abstract
Description
Technical Field
[0001] This specification describes technologies related to methods and systems for presenting content related to objects identified in an image. Background Art
[0002] Computer vision analysis techniques can be used to detect and recognize objects in an image. For example, optical character recognition (OCR) techniques can be used to recognize text in an image, and edge detection techniques can be used to detect objects in an image (e.g., products, landmarks, animals, etc.). Content related to the detected object can be provided to a user (e.g., the user who captured the image in which the object was detected). Summary of the Invention
[0003] This specification describes technologies related to presenting content related to objects identified in an image.
[0004] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method performed by one or more data processing devices of a mobile device, the method including: detecting the presence of one or more objects depicted in a viewfinder of a camera of the mobile device; in response to detecting the presence of the one or more objects: sending image data representing the one or more objects to a content distribution system that selects content related to the objects depicted in the image; and while waiting to receive content from the content distribution system, tracking the position of each of the one or more objects in the viewfinder of the camera; receiving content related to the one or more objects from the content distribution system; and for each of the one or more objects: determining the current position of the object in the viewfinder based on the tracking; and presenting the received content related to the object at the current position of the object within the viewfinder.
[0005] Other embodiments of this aspect include corresponding devices, methods, systems, and computer programs encoded on a computer storage device configured to perform the actions of the method.
[0006] These and other embodiments may each optionally include one or more of the following features. Detecting the presence of an object in an image may include: capturing a sequence of images using a camera of a mobile device; determining that the camera is substantially stationary based on pixel data of the images in the image sequence; in response to determining that the camera is substantially stationary, capturing a given image after the camera has stopped moving; and analyzing the given image using object detection techniques to detect the presence of an object in the given image. Determining that the camera is substantially stationary based on pixel data of the images in the image sequence may include: identifying the respective positions of each pixel in a first image in the image sequence; for each pixel in the first image: identifying the respective positions of corresponding pixels that match the pixel in one or more subsequent images captured after the first image was captured; and determining the distances between the respective positions of the pixels in the first image and the respective positions of the corresponding pixels in each subsequent image; determining that the camera is substantially stationary based on each determined distance being less than a threshold distance. Presenting content of an object at a current position of the object in a viewfinder of the camera may include presenting the content above or near the object in the viewfinder. Determining a current position of an object in a viewfinder of the camera may include: identifying a first set of pixels corresponding to the object in a given image; and determining the position of a second set of pixels in the viewfinder that matches the first set of pixels. Determining a current position of an object in a viewfinder of the camera may include: receiving a first image representing one or more objects depicted in a viewfinder of a camera of a mobile device; determining a first visual feature of a first set of pixels represented in the first image and associated with the object; receiving a second image representing one or more objects depicted in a viewfinder of a camera of the mobile device; and determining the position of the object in the second image based on the first visual feature. Determining a current position of an object in a viewfinder of the camera may include: determining a second visual feature of a second set of pixels represented in a first image; determining the distance between the first set of pixels and the second set of pixels; and determining the position of the object in the second image based on the first visual feature, the second visual feature, and the determined distance. Determining that the camera is substantially stationary may be based on the position of the object in the first image and the position of the object in the second image. Determining that the camera is substantially stationary may further be based on the time associated with the first image and the time associated with the second image. A content distribution system may analyze image data to identify each of the one or more objects; select content for each of the one or more objects; and pre-cache the content before receiving a request for content related to a given object. Presenting a visual indicator within the viewfinder and for each of the one or more objects, the visual indicator indicating that content related to the object is being identified. Detecting the presence of one or more objects depicted in a viewfinder of a camera of a mobile device may include: processing image data representing the one or more objects depicted in a viewfinder of a camera of a mobile device using a coarse classifier.The rough classifier may include a light-weight model. Classify each of the one or more objects into a corresponding object category; and select a corresponding visual indicator for the object from a plurality of visual indicators based on the corresponding category of each of the one or more objects. Sending image data representing the one or more objects to the content distribution system may further include sending data specifying a location associated with the one or more objects.
[0007] The subject matter described in this specification can be implemented in particular embodiments so as to achieve one or more of the following advantages. Images captured from the viewfinder of a mobile device's camera can be provided (e.g., streamed) to a content distribution system, where the content distribution system provides content related to the objects identified in the images, such that the content is presented more quickly, e.g., in response to a user request to view the content. For example, rather than waiting for the user to select an object in the image (or an interface control for the object) and sending the image to the content distribution system in response to that selection, the image can be automatically sent to the content distribution system to increase the speed of content presentation. The content for the identified objects can be stored in a high-speed memory on the server of the content distribution system (e.g., in a cache or at the top of a memory stack) or on the user's device to further increase the speed of presenting content in response to a user request.
[0008] On-device pixel tracking and / or on-device object detection techniques can be used to ensure that the images sent to the content distribution system have sufficient quality to identify objects and / or that the images include objects that the user may be interested in viewing content for. For example, by tracking the movement of visual content represented by individual pixels or groups of pixels within the viewfinder, the mobile device can determine when the device is stationary or substantially stationary (e.g., moving less than a threshold amount), and when the device is determined to be stationary or substantially stationary, provide the image to the content distribution system. Images captured when the device is substantially stationary can result in higher-quality image processing than when the device is moving, which results in more accurate object identification by the content distribution system. This also avoids using computationally expensive image processing techniques to process low-quality images. The fact that the device is stationary can also indicate the user's interest in one or more objects within the viewfinder field, which can reduce the likelihood of the image being sent and processed unnecessarily.
[0009] By sending the captured images only when the user device is determined to be stationary or substantially stationary, the number of images sent over the network to the content distribution system and processed by the content distribution system can be significantly reduced, resulting in less bandwidth consumption, faster network communication, less demand on the content distribution system, and faster object identification and content selection by the content distribution system. Using object detection technology at the user device to determine whether an image depicts an object and providing only the images that depict an object to the content distribution system can provide similar technical improvements compared to streaming all captured images to the content distribution system.
[0010] By presenting a visual indicator indicating that the content of the object is being recognized, the user experience is improved because the user receives real-time feedback. This makes it clear to the user what the application can detect and for which it can provide content, which helps the user learn to use the application.
[0011] The various features and advantages of the foregoing subject matter are described below with reference to the accompanying drawings. Additional features and advantages are apparent from the subject matter described herein and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram of an environment in which a camera application of a mobile device presents content related to an object identified in the viewfinder of the camera of the mobile device.
[0013] Figure 2 Depicts a sequence of example screen shots of a mobile device presenting content related to an object identified in the viewfinder.
[0014] Figure 3 Depicts a sequence of example screen shots of a mobile device presenting content related to an object identified in the viewfinder.
[0015] Figure 4 is a flowchart of an example process for presenting content related to an object identified in the viewfinder of a camera of a mobile device.
[0016] Like reference numerals and names in the different drawings represent like elements. DETAILED DESCRIPTION
[0017] Generally, the systems and techniques described herein can identify objects depicted in the viewfinder of a mobile device's camera and present information and / or content related to the identified objects, e.g., within the viewfinder. An application can present the viewfinder and detect the presence of one or more objects depicted in the viewfinder. In response, the application can present visual indicators for each detected object and send image data representing the viewfinder's field of view to a content distribution system that provides content based on the objects depicted in the image. The application can track the position of an object in the viewfinder such that when content related to the object is received, the application can present the content at the object (e.g., above, near, or within a threshold distance of the object).
[0018] The application can present information and / or content for an object in response to user interaction with the object's visual indicator (or other user interface control) or automatically, e.g., in response to receiving the content or determining that the object is an object in which the user is interested in viewing content. For example, the application can select an object for which to present content based on the object's position in the viewfinder (e.g., the object closest to the center of the viewfinder) or based on detected user actions (e.g., zooming in on the object or moving the camera such that other objects are no longer in the camera's field of view).
[0019] The application can selectively provide image data to the content distribution system. For example, in response to determining that the mobile device is substantially stationary (e.g., moving no more than a threshold amount), the application can capture image data and send the image data to the content distribution system. Lack of movement can serve as a proxy for image quality and the user's interest in one or more objects in the viewfinder's field of view. For example, if a user moves the mobile device such that the camera is pointed at an object and then stops moving the mobile device, the user may be interested in obtaining content (e.g., information or a user experience) related to that object. As described in more detail below, the application can determine whether the mobile device is substantially stationary by tracking the movement of visual content represented by individual pixels or groups of pixels. By selectively sending image data from the device to the content distribution system, the application can reduce the amount of image data sent to the server and can address issues associated with bandwidth usage while avoiding requiring the user to provide input to select which images to send.
[0020] Figure 1It is a block diagram of environment 100, where the camera application 116 of the mobile device 110 presents content related to an object identified in the viewfinder of the camera 111 of the mobile device 110. The mobile device 110 is an electronic device capable of requesting and receiving resources through a data communication network 140 (such as a local area network (LAN), a wide area network (WAN), the Internet, a mobile network, or a combination thereof). Example mobile devices 110 include smartphones, tablet computers, and wearable devices (e.g., smartwatches). The environment 100 may include a number of mobile devices 110.
[0021] The camera application 116 may be a native application developed for a specific platform. The camera application 116 may control the camera 111 of the mobile device 110. For example, the camera application 116 may be a dedicated application for controlling the camera, a camera - first application that controls the camera 111 for use with other features of the application, or another type of application that can access and control the camera 111. The camera application 116 may present the viewfinder of the camera 111 in the user interface 122 of the camera application 116.
[0022] Generally, the camera application 116 enables a user to view content related to an object depicted in the viewfinder of the camera 111 (e.g., information or a user experience) and / or view content related to an object depicted in an image stored on the mobile device 110 or at another location accessible to the mobile device 110. The viewfinder is a portion of the mobile device display that presents a live image within the field of view of the camera lens. As the user moves the camera 111 (e.g., by moving the mobile device), the viewfinder is updated to present the current field of view of the lens.
[0023] The camera application 116 includes an object detector 117, a user interface generator 118, and an on - device tracker 119. The object detector 117 may use edge detection and / or other object detection techniques to detect an object in the viewfinder. In some embodiments, the object detector 117 includes a rough classifier that determines whether the image includes an object among one or more objects of a particular category (e.g., kind). For example, the rough classifier may detect that the image includes an object of a particular category in the case where an actual object is or is not identified.
[0024] A coarse classifier can detect the presence of a class of objects based on whether an image includes (e.g., depicts) one or more features indicative of the class of objects. The coarse classifier can include a lightweight model to perform a low computational analysis to detect the presence of objects within this (these) class(es) of objects. For example, the coarse classifier can detect a limited set of visual features depicted in an image for each class of objects to determine whether the image contains an object belonging to the class of objects. In a specific example, the coarse classifier can detect whether an image depicts an object classified into one or more of the following categories: text, barcode, landmark, media object (e.g., album cover, movie poster, etc.), or art object (e.g., painting, sculpture, etc.). For a barcode, the coarse classifier can determine whether the image includes parallel lines of different widths.
[0025] In some embodiments, the coarse classifier uses a trained machine learning model (e.g., a convolutional neural network) to determine whether an image includes an object in one or more classes based on the visual features of the image. For example, the machine learning model can be trained using labeled images, where the labeled images are labeled with their corresponding (multiple) object classes. The machine learning model can be trained to classify an image into zero or more specific sets of object classes. The machine learning model can receive data related to the visual features of the image as input and output a classification into zero or more of the object classes in the specific set of object classes.
[0026] The coarse classifier can output data specifying whether a class of objects has been detected in the image. The coarse classifier can also output a confidence value indicating the confidence that the presence of a class of objects has been detected in the image and / or a confidence value indicating the confidence that an actual object (e.g., the Eiffel Tower) is depicted in the image.
[0027] The object detector 117 can receive image data representing the field of view of the camera 111 (e.g., what is presented in the viewfinder) and detect the presence of one or more objects in the image data. If at least one object is detected in the image data, the camera application 116 can provide (e.g., send) the image data to the content distribution system 150 via the network 140. As described below, the content distribution system 150 can identify objects in the image data and provide content related to the objects to the mobile device 110.
[0028] In some embodiments, when camera 111 is determined to be substantially stationary or to have stopped moving, camera application 116 uses object detector 117 to detect the presence of objects in the viewfinder. Camera 111 can be considered substantially stationary when it is determined to be moving less than a threshold amount (e.g., less than 10 millimeters per second). Camera 111 can be considered to have stopped moving when the amount of movement drops below the threshold amount after having moved beyond the threshold amount. By only processing the image data of images captured when camera 111 is determined to be substantially stationary, camera application 116 can ensure that higher quality images are being processed rather than wasting expensive computational resources on low-quality images for which object detection may be impossible or inaccurate.
[0029] In addition, when camera 111 is substantially stationary, there is a greater likelihood than when camera 111 is moving beyond the threshold amount that the user is interested in the content of an object received in the field of view of camera 111. For example, if a user wants to obtain content related to an object, the user may hold mobile device 110 such that the object is stable in the field of view of camera 111. Thus, by only processing the data of images captured when camera 111 is determined to be stationary, expensive computational resources are not wasted detecting and / or discerning objects in images that the user is not interested in the received content.
[0030] Camera application 116 can use on-device tracker 119 (which can be part of camera application 116 or a separate tracker) to determine when camera 111 is substantially stationary. On-device tracker 119 can track the movement of visual content represented by individual pixels or groups of pixels within the viewfinder. For example, on-device tracker 119 can use the pixel data of an image to track the movement of (a) pixel(s) throughout a sequence of images captured by camera 111. To track this movement, on-device tracker 119 can determine the visual characteristics (e.g., color, intensity, etc.) of each pixel (or at least a portion of the pixels) of the image. On-device tracker 119 can then identify the location of the visual content represented by that pixel in subsequent images, e.g., by identifying pixels in subsequent images having the same or similar visual characteristics.
[0031] Since multiple pixels in an image can have the same visual characteristics, the on-device tracker 119 can track the movement of visual content represented by the pixels based on the multiple pixels in the image and their relative positions. For example, if a first group of pixels having the same or similar first visual characteristic is a specific distance away from a second group of pixels having the same or similar second visual characteristic in a first image, and the two groups of pixels are identified as being the same distance apart in a second image, the on-device tracker 119 can determine that the second group of pixels matches the first group of pixels. The on-device tracker 119 can then determine the distance that the visual content represented by the group of pixels has moved in the viewfinder and use that distance to estimate the amount of movement of the camera 111.
[0032] The on-device tracker 119 can estimate the movement of the camera 111 based on the distance that each pixel moves between consecutive images (e.g., in terms of the number of pixels) and the duration between the times at which each image is captured. In this example, the movement can be based on the number of pixels per second. The on-device tracker 119 can convert the measurement to centimeters per second by multiplying the number of pixels per second by a constant value.
[0033] In some embodiments, the mobile device 110 includes a gyroscope and / or an accelerometer capable of detecting the movement of the mobile device 110. The camera application 116 can receive data describing the movement of the mobile device 110 from the gyroscope and / or the accelerometer and use that data to determine whether the mobile device 110 is substantially stationary (e.g., moving less than a threshold amount).
[0034] The on-device tracker 119 can also track the movement of an object detected by the object detector 117 across multiple images. For example, the object detector 117 can output data specifying the position of the object detected in the image (e.g., pixel coordinates). The pixel coordinates can include the coordinates of multiple positions along the boundary of the object, e.g., to outline the object. The on-device tracker 119 can then monitor the movement of the object based on the movement of the visual content depicted by the pixels at that location (e.g., within the pixel coordinates) across subsequent images. As described below, the position of the object in subsequent images can be used to determine where the content of the object is presented in the viewfinder.
[0035] To track the movement of an object, the on-device tracker 119 can analyze each subsequent image to identify the location of visual content presented in the pixel coordinates of the first image in which the object was detected, similar to the way the on-device tracker 119 tracks the movement of individual pixels and groups of pixels. For example, the on-device tracker 119 can determine the visual characteristics of each pixel within the pixel coordinates of the object in the first image and the relative orientation of the pixels (e.g., the distance and direction between pairs of pixels). In each subsequent image, the on-device tracker 119 can attempt to identify pixels with the same (or similar) visual characteristics and the same (or similar) orientation. For example, the orientation of an object within the viewfinder can change based on a change in the orientation of the mobile device 110, a change in the orientation of the object, and / or a change in the distance between the mobile device 110 and the object. Thus, if the distance between each pair of pixels is within a threshold of the distance between pairs of pixels in the previous image, the on-device tracker 119 can determine that the group of pixels in the subsequent image matches the object.
[0036] In another example, if the group of pixels has the same (or similar) shape and visual characteristics (e.g., color and intensity) regardless of orientation or size, the on-device tracker 119 can determine that the group of pixels in the subsequent image matches the object. For example, the user can rotate the mobile device 110 or move closer to the object in an attempt to capture a better image of the object. In this example, the size and orientation of the object may change within the viewfinder, but the shape and color should remain the same or nearly the same.
[0037] In another example, the on-device tracker 119 can identify the edges of an object based on pixel coordinates and track the movement of the edges between images. If the mobile device 110 is substantially stationary, the edges may not move much between consecutive images. Thus, the on-device tracker 119 can locate the position of the edge in the subsequent image by analyzing the pixels near the edge in the previous image (e.g., within a threshold number of pixels).
[0038] The user interface generator 118 can generate and update a user interface 122 that presents the viewfinder of the camera 111 and other content. If the object detector 117 detects the presence of one or more objects in the viewfinder, the user interface generator 118 can present visual indicators in the viewfinder for each detected object. The visual indicators can indicate to the user the objects that have been detected and the objects for which content is being identified for presentation to the user.
[0039] In some embodiments, once the camera application 116 is launched or the camera is activated within the camera application 116, the camera application 116 uses the object detector 117 to detect objects. In this way, visual indicators can be presented in real time to enable the user to quickly request content related to the objects.
[0040] The user interface generator 118 can update the user interface to present a visual indicator for each object at the location of the object within the viewfinder. For example, the user interface generator 118 can present the visual indicator of the object on the object (e.g., using a translucent indicator), near the object, or within a threshold number of pixels from the object, but such that the visual indicator does not block the line of sight of other detected objects. The visual indicator of the object can include a visual highlight of the object (e.g., a visible frame or other shape around the object).
[0041] The user interface generator 118 can present different visual indicators for different categories of objects. For example, the user interface generator 118 can store one or more visual indicators for each category of object that the object detector 117 is configured to detect. When a particular category of object is detected, the user interface generator 118 can select the visual indicator corresponding to that particular category and present the visual indicator at the location of the object detected in the viewfinder. In one example, the visual indicator for text can include a circle with the letter "T" or the word "text" inside the circle (or other shape), the visual indicator for a landmark can include a circle with the letter "L" inside the circle, and the visual indicator for a dog can include a circle with the letter "D" or the word "dog" inside the circle. The user interface generator 118 can present different visual indicators for text, barcodes, media, animals, plants and flowers, cars, faces, landmarks, food, clothing, electronic devices, bottles and jars, and / or other categories of objects that can be detected by the object detector 117. In another example, each visual indicator can include a number corresponding to the object. For example, the visual indicator for the first object detected can include the number 1, the visual indicator for the second object detected can include the number 2, and so on.
[0042] In some embodiments, the visual indicator of the object can be based on the actual object identified in the viewfinder. For example, as described below, the mobile device 110 can include an object identifier, or the content distribution system 150 can perform object detection techniques and visual indicator techniques. In these examples, the actual object can be identified before presenting the visual indicator, and the visual indicator can be selected based on the actual object. For example, the visual indicator for the Eiffel Tower can include a small image of the Eiffel Tower.
[0043] In some embodiments, the user interface generator 118 may select a visual indicator for an object based on the object's category. The user interface generator 118 may present the visual indicator at the location of the object in the viewfinder. If the object is recognized, the user interface generator 118 may replace the visual indicator with a visual indicator selected based on the actual recognized object. For example, the object detector 117 may detect the presence of a landmark in the viewfinder, and the user interface generator 118 may present a visual indicator of the landmark at the location of the detected landmark. If the landmark is determined to be the Eiffel Tower (e.g., by an object identifier at the mobile device 110 or the content distribution system 150), the user interface generator 118 may replace the visual indicator of the landmark with a visual indicator of the Eiffel Tower.
[0044] As described above, the on-device tracker 119 may track the location of an object detected within the viewfinder. The user interface generator 118 may use the location information of each object to move the visual indicator of each detected object such that the visual indicator of the object follows the object within the viewfinder. For example, the user interface generator 118 may continuously (or periodically) update the user interface 122 such that the visual indicator of each object follows the object and is presented at the location of the object in the viewfinder.
[0045] The visual indicator may be interactive. For example, the camera application 116 may detect an interaction with the visual indicator (e.g., selected by the user). In response to detecting the user's interaction with the visual indicator, the camera application 116 may request content related to the object for which the visual indicator is presented from the content distribution system 150.
[0046] The content distribution system 150 includes one or more front-end servers 160 and one or more back-end servers 170. The front-end servers 160 may receive image data from a mobile device (e.g., the mobile device 110). The front-end servers 160 may provide the image data to the back-end servers 170. The back-end servers 170 may identify content related to the objects recognized in the image data and provide the content to the front-end servers 160. In turn, the front-end servers 160 may provide the content to the mobile device from which they received the image data.
[0047] The backend server 170 includes an object recognizer 172, a user interface control selector, and a content selector 174. The object recognizer 172 can process the image data received from the mobile device and recognize an object (if any) in the image data. The object recognizer 172 can use computer vision and / or other object recognition techniques (e.g., edge matching, pattern recognition, greyscale matching, gradient matching, etc.) to recognize an object in the image data.
[0048] In some embodiments, the object recognizer 172 uses a trained machine learning model (e.g., a convolutional neural network) to recognize an object in the image data received from the mobile device. For example, the machine learning model can be trained using labeled images with their corresponding objects. The machine learning model can be trained to recognize and output data identifying the object depicted in the image represented by the image data. The machine learning model can receive data related to the visual features of the image as input and output data identifying the object depicted in the image.
[0049] The object recognizer 172 can also output a confidence value that indicates the confidence that the image depicts the recognized object. For example, the object recognizer 172 can determine the confidence of each object recognized in the image based on the degree of match between the features of the object and the features of the image.
[0050] In some embodiments, the object recognizer 172 includes multiple object recognizer modules, e.g., an object recognizer module for each class of objects that recognizes objects in its corresponding class. For example, the object recognizer 172 can include a text recognizer module that recognizes text in the image data (e.g., recognizes characters, words, etc.), a barcode recognizer module that recognizes (e.g., decodes) barcodes (including QR codes) in the image data, a landmark recognizer module that recognizes landmarks in the image data, and / or other object recognizer modules that recognize objects of a specific class.
[0051] In some embodiments, the camera application 116 provides data specifying the location (e.g., pixel coordinates) within the image where a specific object or a specific class of objects is detected in the image data of the image. This can increase the speed at which the object is recognized by enabling the object recognizer 172 to focus on the image data at that location and / or by enabling the object recognizer 172 to use an appropriate object recognizer module (e.g., an object recognizer module for only the object class specified by the data received from the camera application 16) to recognize the object(s) in the image data. This also reduces the amount of computational resources that will be used by other object recognition modules.
[0052] The content selector 174 can select the content to be provided to the camera application 116 for each object identified in the image data. The content can include information related to the object (e.g., text including the object name and / or facts about the object), visual processing (e.g., other images or videos of the object or related objects), links to resources related to the object (e.g., links to web pages or app pages where the user can purchase the object or view additional information about the object), or experiences related to the object (augmented reality videos, playing music in response to identifying a singer or a poster of a singer), and / or other suitable content. For example, if the object is a barcode, the selected content can include a text-based caption that includes the name of the product corresponding to the barcode and information about the product, a link to a web page or app page where the user can purchase the product, and / or an image of the product.
[0053] The content selector 174 can select visual processing of text related to the identified object. The visual processing can be in the form of a text caption that can be presented at the object in the viewfinder. The text included in the caption can be based on the ranking of facts about the object. For example, more popular facts can be ranked higher. The content selector 174 can select one or more of the captions in the caption for the identified object based on the ranking to provide to the mobile device 110.
[0054] The content selector 174 can select the text of the caption based on the level of confidence output by the object identifier 172. For example, if the confidence level is high (e.g., greater than a threshold), the text can include popular facts about the object or the object name. If the confidence level is low (e.g., less than a threshold), the text can indicate that the object may be what the object identifier 172 detected (e.g., "This may be a golden retriever").
[0055] The content selector 174 can also select interactive controls based on the object(s) identified in the image. For example, if the object identifier 172 detects a phone number in the image, the content selector 174 can select a click-to-call icon that, when interacted with, causes the mobile device 110 to call the identified phone number.
[0056] Content can be stored in content data storage unit 176, which can include a hard disk drive, flash memory, random access memory, or other data storage devices. In some embodiments, content data storage unit 176 includes an index that specifies, for each object and / or each type of object, the content that can be provided for that object or type of object. This index can increase the speed of selecting content for an object or a type of object.
[0057] After the content is selected, the content can be provided to mobile device 110 from which the image data is received, stored in content cache 178 of content distribution system 150, and / or stored at the top of the memory stack of front-end server 160. In this way, in response to a user's request for content, the content can be quickly presented to the user. If the content is provided to mobile device 110, camera application 116 can store the content in content cache 112 or other random access memory. For example, camera application 116 can refer to the object to store the content of the object so that camera application 116 can identify the appropriate content of the object in response to determining to present the content of the object.
[0058] Camera application 116 can present the content of the object in response to a user's interaction with the visual indicator of the object. For example, camera application 116 can detect the user's interaction with the visual indicator of the object and request the content of the object from content distribution system 150. In response, front-end server 160 can obtain the content from content cache 178 or the top of the memory stack and provide the content to mobile device 110 from which the request was received. If the content was provided to mobile device 110 before the user interaction was detected, camera application 116 can obtain the content from content cache 112.
[0059] In some embodiments, the camera application 116 determines (e.g., automatically determines) to present content related to an object. For example, the camera application 116 can determine based on the position of the object in the viewfinder (e.g., the object is closest to the center of the viewfinder) or based on a detected user action (e.g., zooming in on the object or moving the camera such that other objects are no longer within the view of the viewfinder). For example, the object detector 117 can initially detect three objects in the viewfinder, and the user interface generator 118 can present visual indicators for each of the three objects. The user can then interact with the camera application 116 to move the camera 111 (and change its field of view) or zoom in on one of the objects. The camera application 116 can determine that one or more of the objects are no longer depicted in the viewfinder, or determine that one of the objects is now at the center of the viewfinder. In response, the camera application 116 can determine to present the content of the object(s) remaining in the viewfinder and / or now at the center of the viewfinder.
[0060] If the user interacts with the visual indicator to request the content of the object, the user interface generator 118 can freeze the viewfinder user interface 122 (e.g., maintain the presentation of the current image in the viewfinder) and present the content of the object. In this way, the user does not have to keep the mobile device 110 stationary with respect to the object in the field of view to view the content related to the object while the object remains visible in the viewfinder.
[0061] The tasks performed by the mobile device 110 and the content distribution system 150 can be divided between the mobile device 110 and the content distribution system 150 in various ways, e.g., depending on the priorities of the system or user preferences. For example, object detection (e.g., using rough classification), visual indicator selection, object identification, and / or content selection can be distributed between the mobile device 110, the content distribution system 150, and / or other systems.
[0062] In some embodiments, tasks can be distributed between the mobile device 110 and the content distribution system 150 based on the current conditions of the environment 100. For example, if network communication is slow (e.g., due to demand on the network 140), the camera application 116 can perform object identification and content selection tasks instead of sending image data over the network 140. If network communication is fast (e.g., greater than a threshold speed), the mobile device 110 can stream more images to the content distribution system 150 compared to when network communication is slow, e.g., by not considering the movement of the camera 111 or whether an object is detected in the image when determining whether to send image data to the content distribution system 150.
[0063] In some embodiments, the content distribution system 150 includes an object detector 117, such as instead of the camera application 116. In such an example, when the camera application 116 is activated or when the user places the camera application 116 in a requested content mode, the camera application 116 can continuously (e.g., as an image stream) send image data to the content distribution system 150. The requested content mode can allow the camera application 116 to continuously send image data to the content distribution system 116 in order to request content for an object identified in the image data. The content distribution system 150 can detect an object in the image, select a visual indicator for the detected object, and send the visual indicator to the camera application 116 for presentation in the viewfinder. The content distribution system 150 can also continue to process the image data to identify objects, select content for each identified object, and either cache the content or send the content to the camera application 116.
[0064] In some embodiments, the camera application 116 includes an on-device object identifier that identifies objects in the image data. In this example, the camera application 116 can identify an object, present a visual indicator of the identified object, and either request content for the identified object from the content distribution system or identify content from an on-device content datastore. The on-device object identifier can be a lightweight object identifier, where the lightweight object identifier identifies a more limited set of objects or uses object identification techniques that are less computationally expensive than the object identifier 172 of the content distribution system 150. This enables a mobile device with processing power lower than a typical server to execute an object identification program. In some embodiments, the camera application 116 can use the on-device identifier to perform an initial identification of an object and provide the image data to the content distribution system 150 (or another object identification system) for confirmation. The on-device content datastore can also store a more limited set of content or a link to a resource including the content than the content storage unit 176 to conserve the data storage resources of the mobile device 110.
[0065] Figure 2 A sequence of example screenshots 210, 220, and 230 of a mobile device presenting content related to an object identified in the viewfinder is depicted. In the first screenshot 210, the mobile device presents a user interface 212 that includes a viewfinder 213 of a camera. The viewfinder 213 depicts a shoe 214. The user interface 212 can be presented by the camera application, e.g., when the user launches the camera application.
[0066] The camera application (or content distribution system) can process the image data of the viewfinder 213 to detect the presence of any object depicted in the viewfinder 213. In this example, the camera application has detected the presence of the shoe 214 in the viewfinder 213. In response, the camera application can update the user interface 212 to present the user interface 222 of the screen snapshot 220.
[0067] The updated user interface 222 presents a visual indicator 224 of the detected shoe 214 within the viewfinder 213. In this example, the visual indicator 224 is a circle with the letter "s" inside the circle. Other visual indicators with different shapes or other visual features can be presented alternatively. The visual indicator 224 is presented at the shoe 214, for example, on a part of the shoe.
[0068] In the updated user interface, the shoe 214 has moved upward in the viewfinder 213, for example, upward based on the movement of the mobile device and / or its camera. As described above, the tracker on the device can track the position of the detected object within the viewfinder. This allows the camera application to present the visual indicator 224 of the shoe 214 at the position of the shoe 214 within the viewfinder 213.
[0069] The visual indicator 224 indicates that the shoe 214 has been detected and that content related to the shoe is being recognized or has been recognized. As described above, in response to detecting an object in the viewfinder, the camera application can send the image data of the viewfinder to the content distribution system. The visual indicator can also be interactive, such that interacting with the visual indicator causes the camera application to send a request for content related to the shoe 214.
[0070] In response to detecting a user interaction, the camera application can update the user interface 222 to present the user interface 232 of the screen snapshot 230. The updated user interface 232 presents the shoe 214 and content related to the shoe 214. In some embodiments, the user interface 232 presents an image 240 of the shoe 214 (e.g., an image captured when the user interacts with the visual indicator 224), rather than a live image of those within the field of view of the camera's lens. The position of the shoe 214 in the image 240 can be the same as the position of the shoe 214 in the viewfinder 213 when the user interaction with the visual indicator 224 is detected.
[0071] Content related to the shoe 214 includes a caption 234, where the caption 234 designates that the content distribution system has identified the shoe as a "super lightweight shoe". The caption is presented on the portion of the shoe 214 in the image 240. Content related to the shoe also includes: an interactive icon 238 that includes a link to a resource (e.g., an app page or a web page) where a user can purchase the super lightweight shoe; and an interactive icon 237 that, when interacted with, causes the camera app to present content of shoes similar to the super lightweight shoe.
[0072] Figure 3 A sequence of example screen shots 310 and 320 of a mobile device presenting content related to an object identified in a viewfinder is depicted. In the first screen shot 310, the mobile device presents a user interface 312 that includes a viewfinder 313 of a camera. The viewfinder 313 depicts furniture in a room. The user interface 312 can be presented by a camera app, e.g., when the user launches the camera app. When the camera app is launched (or at another time, e.g., when the mobile device becomes stationary), the camera of the mobile device can be pointed at the furniture.
[0073] The user interface 312 includes visual indicators 314 - 317 for each piece of furniture detected in the viewfinder 313. In this example, each visual indicator 314 - 317 presents a corresponding number representing the corresponding piece of furniture. For example, the visual indicator 314 is a circle with the number 1 in it and corresponds to the gray chair in the viewfinder 313.
[0074] Since the viewfinder 313 depicts real - time images of those within the field of view of the camera's lens, moving the camera or zooming out the camera causes different objects to be presented in the viewfinder 313 and / or objects to move out of the viewfinder 313. As shown in the updated user interface 322 in the screen shot 320, the user has zoomed in on the gray chair.
[0075] In response, the camera app can interpret the user's action as an indication that the user wants to receive content related to the gray chair. For example, the camera app can interpret the zooming in on the chair and / or the fact that no other objects are depicted in the viewfinder 313 as an indication that the user is interested in receiving content related to the gray chair. In response, the camera app can present content related to the gray chair, e.g., without detecting an interaction with the visual indicator 314 of the gray chair.
[0076] For example, the user interface 322 presents a caption 326 that includes information about a gray chair, an interactive icon 327 that includes a link to a resource (e.g., an app page or a web page), and an interactive icon 328, where the user can purchase the chair at the link (e.g., a retailer's web page or a web page that includes chairs offered by multiple retailers), and where when interacting with the interactive icon 328, the interactive icon 328 causes the camera application to present content of chairs similar to the gray chair.
[0077] Figure 4 is a flowchart of an example process 400 for presenting content related to an object recognized in the viewfinder of a camera of a mobile device. Operations of process 400 can be performed, for example, by one or more data processing devices (such as Figure 1 the mobile device 110). Operations of process 400 can also be implemented as instructions stored on a non-transitory computer-readable medium. Execution of the instructions causes one or more data processing devices to perform the operations of process 400.
[0078] Detect the presence of one or more objects depicted in the viewfinder of the camera of the mobile device (402). For example, a camera application executing on the mobile device can include an object recognizer that detects the presence of an object depicted in the viewfinder of the camera based on the image data of the viewfinder.
[0079] In some embodiments, in response to determining that the camera is stationary or has stopped moving (e.g., moving no more than a threshold amount), the image data is captured and analyzed. If the camera is moving (e.g., more than a threshold amount), the image data of the viewfinder may not be captured or processed to detect the presence of an object because the image data may not be of good enough quality or the user may not be interested in receiving content related to anything in the camera's field of view.
[0080] For each detected object, present a visual indicator in the viewfinder (404). The visual indicator can indicate to the user of the mobile device that content is being (or has been) identified for the object. The visual indicator of the object can be displayed at the location of the object in the viewfinder.
[0081] The visual indicator presented for each object can be based on the category to which the object has been classified. For example, the visual indicator for text can be different from the visual indicator for a person. The visual indicator for each object can be interactive. For example, interaction by the user with the visual indicator can initiate a request for content related to the object.
[0082] Image data representing the one or more objects is sent to a content distribution system (406). The content distribution system can identify the one or more objects and select content associated with each of the one or more objects. The content distribution system can send the selected content to the mobile device, or store the selected content in a cache or at the top of a memory stack. If stored in a content management system, the content management system can send the content to the mobile device in response to a request for the content received from the mobile device.
[0083] The position of each of the one or more objects in the viewfinder is tracked (408). For example, as described above, pixel tracking techniques can be used to track the position of each object. When an object is visible in the viewfinder, the current position of the object can be continuously tracked while waiting to receive content from the content distribution system or waiting for the user to interact with a visual indicator of the object. In this way, when content is received or the user interacts with a visual indicator of the object, the content can be presented at the position of the object without delay.
[0084] Content is received from the content distribution system (410). The received content can be stored on the mobile device, for example, in a local cache or other high-speed memory.
[0085] The current position of each of the one or more objects is determined based on the tracking (412). For example, as described above, the position of each detected object within the viewfinder can be continuously tracked to maintain the current position of each object. When it is time to present the content, the current position of each object for which the content will be presented is determined based on this tracking.
[0086] The content of each object is presented in the viewfinder (414). The content of the object can be presented at the current position of the object in the viewfinder. In response to the user interacting with a visual indicator of the object, or in response to determining that the object is an object in which the user is interested in receiving content, the content can be presented when the content is received.
[0087] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a backend component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a frontend component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification), or includes any combination of one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”) (e.g., the Internet).
[0088] A computing system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The relationship between the clients and servers is created by computer programs running on the respective computers and having a client-server relationship with each other.
[0089] Although this specification contains many specific implementation details, these should not be construed as limitations on any invention or the scope of the claims, but rather as descriptions of specific features of particular embodiments of a specific invention. Certain features described in the context of separate embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may be directed to a sub-combination or a variation of the sub-combination.
[0090] Similarly, although operations are depicted in the figures in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of the various system modules and components in the above embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0091] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the acts recited in the claims may be performed in a different order and still achieve the desired result. As one example, the processes described in the figures do not necessarily need the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
Claims
1. A method performed by one or more data processing devices of a mobile device, the method comprising: Detecting the presence of a plurality of objects depicted in a viewfinder of a camera of the mobile device; In response to detecting the presence of the plurality of objects: Presenting, within the viewfinder and for each of the plurality of objects, a visual indicator that indicates that content associated with the object is being recognized; Sending image data representing the plurality of objects to a content distribution system that selects content associated with the plurality of objects depicted in the image data, wherein the image data includes data specifying respective positions within the image at which each of the plurality of objects is detected; And While waiting to receive content associated with the plurality of objects depicted in the image data from the content distribution system, tracking the respective positions of each of the plurality of objects within the viewfinder of the camera; Receiving content associated with the plurality of objects from the content distribution system; And For each of the plurality of objects: Based on the tracking, determining the respective current position of the object within the viewfinder; And Presenting, within the viewfinder and at the respective current position of the object, the received content associated with the object.
2. The method according to claim 1, wherein, Detecting the presence of an object in the image includes: Capturing a sequence of images using the camera of the mobile device; Based on pixel data of the images in the image sequence, determining that the camera is substantially stationary; In response to determining that the camera is substantially stationary, capturing a given image after the camera stops moving; and Analyzing the given image using object detection techniques to detect the presence of the object in the given image.
3. The method according to claim 2, wherein Based on pixel data of the images in the image sequence, determining that the camera is substantially stationary includes: Identifying the respective positions of each pixel in a first image of the image sequence; For each pixel in the first image: Identifying the respective positions of corresponding pixels that match the pixel in one or more subsequent images captured after the first image is captured; and Determining the distance between the respective position of the pixel in the first image and the respective positions of the corresponding pixels in each subsequent image; Determining that the camera is substantially stationary based on each determined distance being less than a threshold distance.
4. The method according to claim 1, wherein Presenting the content of the object at the respective current position of the object in the viewfinder of the camera includes presenting the content above or near the object in the viewfinder.
5. The method according to claim 1, wherein, Determining the respective current position of an object in the viewfinder of the camera includes: Identifying a first set of pixels in a given image that corresponds to the object; and Determining the position of a second set of pixels in the viewfinder that matches the first set of pixels.
6. The method according to claim 1, wherein, The content distribution system: Analyzes the image data to identify each of the plurality of objects; Selects content for each of the plurality of objects; and Pre-caches the content before receiving a request for content associated with a given object.
7. The method according to claim 1, further comprising: Classifying each of the plurality of objects into a respective object category; And Select a respective visual indicator from a plurality of visual indicators for each object of the plurality of objects based on the respective category of each object.
8. A system for presenting content, comprising: One or more data processing devices; And A memory storage device in data communication with the data processing device, the memory storage device storing instructions that are executable by the data processing device and that, when executed, cause the data processing device to perform operations, the operations including: Detect the presence of a plurality of objects depicted in the viewfinder of a camera of a mobile device; In response to detecting the presence of the plurality of objects: Present, within the viewfinder and for each object of the plurality of objects, a visual indicator that indicates that content associated with the object is being recognized; Send image data representing the plurality of objects to a content distribution system that selects content associated with the plurality of objects depicted in the image data, wherein the image data includes data specifying the respective locations at which each of the plurality of objects in the image is detected; and While waiting to receive content associated with the plurality of objects depicted in the image data from the content distribution system, track the position of each object of the plurality of objects in the viewfinder of the camera; Receive content associated with the plurality of objects from the content distribution system; and For each object of the plurality of objects: Based on the tracking, determine the respective current position of the object in the viewfinder; and Present, within the viewfinder and at the respective current position of the object, the received content associated with the object.
9. The system according to claim 8, wherein Detecting the presence of an object in the image includes: Capturing a sequence of images using the camera of the mobile device; Based on pixel data of the images in the image sequence, determine that the camera is substantially stationary; In response to determining that the camera is substantially stationary, capture a given image after the camera stops moving; and Analyze the given image using object detection techniques to detect the presence of the object in the given image.
10. The system according to claim 9, wherein, Determining that the camera is substantially stationary based on pixel data of the images in the image sequence includes: Identifying the respective positions of each pixel in a first image of the image sequence; For each pixel in the first image: In one or more subsequent images captured after the first image is captured, identify the respective positions of corresponding pixels that match the pixel; and Determine the distance between the respective position of the pixel in the first image and the respective positions of the corresponding pixels in each subsequent image; Determine that the camera is substantially stationary based on each determined distance being less than a threshold distance.
11. The system according to claim 8, wherein Presenting the content of the object at the respective current position of the object in the viewfinder of the camera includes presenting the content above or near the object in the viewfinder.
12. The system according to claim 8, wherein Determining the respective current position of an object in the viewfinder of the camera includes: Identifying a first set of pixels corresponding to the object in a given image; and Determining the position of a second set of pixels in the viewfinder that matches the first set of pixels.
13. The system according to claim 8, wherein, The content distribution system: Analyze the image data to identify each of the plurality of objects; Select content for each of the plurality of objects; and Pre-cache the content before receiving a request for content related to a given object.
14. The system according to claim 8, wherein The operations include: Classify each of the plurality of objects into a corresponding object category; and Based on the respective category of each of the plurality of objects, select a corresponding visual indicator for the object from a plurality of visual indicators.
15. A non-transitory computer storage medium encoded with a computer program, the program including instructions that, when executed by a data processing device, cause the data processing device to perform operations, the operations including: Detect the presence of a plurality of objects depicted in a viewfinder of a camera of a mobile device; In response to detecting the presence of the plurality of objects: Present within the viewfinder and for each of the plurality of objects a visual indicator that indicates that content related to the object is being recognized; Send image data representing the plurality of objects to a content distribution system that selects content related to the plurality of objects depicted in the image data, wherein the image data includes data specifying the respective locations within the image where each of the plurality of objects is detected; And While waiting to receive from the content distribution system content related to the plurality of objects depicted in the image data, track the position in the viewfinder of the camera of each of the plurality of objects; Receive from the content distribution system content related to the plurality of objects; And For each of the plurality of objects: Based on the tracking, determine the respective current position in the viewfinder of the object; And Present within the viewfinder and at the respective current position of the object the received content related to the object.
16. The non-transitory computer storage medium according to claim 15, wherein, Detecting the presence of an object in the image includes: Capturing a sequence of images using the camera of the mobile device; Based on pixel data of the images in the image sequence, determine that the camera is substantially stationary; In response to determining that the camera is substantially stationary, capture a given image after the camera stops moving; and Use object detection techniques to analyze the given image to detect the presence of the object in the given image.
17. The non-transitory computer storage medium according to claim 16, wherein, Based on pixel data of the images in the image sequence, determining that the camera is substantially stationary includes: Identify the respective position of each pixel in a first image of the image sequence; For each pixel in the first image: In one or more subsequent images captured after the first image is captured, identify the respective position of a corresponding pixel that matches the pixel; and Determine the distance between the respective position of the pixel in the first image and the respective position of the corresponding pixel in each subsequent image; Determine that the camera is substantially stationary based on each determined distance being less than a threshold distance.
18. The non-transitory computer storage medium according to claim 15, wherein, Presenting the content of the object at the respective current position of the object in the viewfinder of the camera includes presenting the content above or near the object in the viewfinder.
19. The non-transitory computer storage medium according to claim 15, wherein, Determining the respective current position in the viewfinder of the camera of the object includes: Identify a first set of pixels corresponding to the object in a given image; and Determine the position of a second set of pixels in the viewfinder that match the first set of pixels.
20. The non-transitory computer storage medium according to claim 15, wherein, The content distribution system:[[]] Analyze the image data to identify each of the plurality of objects; Select content for each of the plurality of objects; and Pre-cache the content before receiving a request for content related to a given object.
21. The non-transitory computer storage medium according to claim 15, wherein, The operations include:[[]] Classify each of the plurality of objects into a corresponding object category; and Based on the respective categories of each of the plurality of objects, select a corresponding visual indicator for the object from a plurality of visual indicators.
Citation Information
Patent Citations
Time scale adaptive motion detection
US20150139484A1