A scheme for retrieving content items and associating them with real-world objects
By identifying real-world objects through the camera of the user's device and matching them with tagged objects based on image recognition, content item notifications are provided, which solves the limitations of the association between content items and real-world objects in existing augmented reality systems and realizes customized social interaction and information sharing.
Patent Information
- Application Number
- CN202110249994.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-12-09
- Filing Date
- 2015-09-18
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2036-03-12
AI Technical Summary
Most existing augmented reality systems have limitations in applications and content, and are unable to effectively provide users with customized and controlled associations between content items and real-world objects.
The system recognizes real-world objects through the camera of the user's device, matches the real-world objects with tagged objects based on image recognition and settings of the tagged objects, and provides notification of content items. The user can select and configure sharing settings to be stored on the server for retrieval by other users.
It enables users to share content items with real-world objects in a customized and controlled manner in an augmented reality system, provides social interaction and information sharing, and enhances the user experience.
Smart Images

Figure CN112906615B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application date of September 18, 2015, application number 201580051147.6, and invention name “Scheme for retrieving content items and associating them with real-world objects using augmented reality and object recognition”.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application is a continuation-in-part of and claims the benefit of U.S. Patent Application No. 14 / 565,236, filed December 9, 2014, entitled “SCHEMES FOR RETRIEVING AND ASSOCIATING CONTENT ITEMS WITH REAL-WORLD OBJECTS USING AUGMENTED REALITY AND OBJECT RECOGNITION,” which claims the benefit of U.S. Provisional Patent Application No. 62 / 057,219, filed September 29, 2014, entitled “SCHEMES FOR RETRIEVING AND ASSOCIATING CONTENT ITEMS WITH REAL-WORLD OBJECTS USING AUGMENTED REALITY AND OBJECT RECOGNITION,” and further claims the benefit of U.S. Provisional Patent Application No. 62 / 057,219, filed September 29, 2014, entitled “METHOD AND APPARATUS FOR RECOGNITION AND MATCHING OF OBJECTS DEPICTED IN IMAGES”, the entire contents and disclosure of which are hereby incorporated by reference herein in their entirety.
[0004] This application is related to U.S. patent application Ser. No. 14 / 565,204, filed Dec. 9, 2014, entitled “METHOD AND APPARATUS FOR RECOGNITION AND MATCHING OF OBJECTS DEPICTED IN IMAGES,” the entire contents and disclosure of which are hereby fully incorporated herein by reference. Background of the Invention
[0005] 1. Field of the Invention
[0006] The present invention relates generally to computer software applications, and more particularly to augmented reality devices and software.
[0007] 2. Related technical discussions
[0008] Augmented reality (AR) is the general concept of modifying the view of reality with computer-generated sensory input. While AR systems already exist, most are limited in terms of applications and content. In most cases, AR applications are single-purpose programs that provide users with content generated and controlled by the developer. Summary of the Invention
[0009] One embodiment provides a method comprising: identifying a real-world object in a scene viewed by a camera of a user device; matching the real-world object to a tagged object based at least in part on image recognition and a shared setting of tagged objects, the tagged object having been tagged with a content item; providing a notification to a user of the user device that the content item is associated with the real-world object; receiving a request for the content item from the user; and providing the content item to the user.
[0010] One embodiment provides a method comprising: identifying a real-world object in an image; matching the real-world object to a tagged object based at least in part on image recognition and a setting that has been assigned to the tagged object, the tagged object having been tagged with a content item, the content item being distinct from the setting; and providing a notification to a user that the content item is associated with the real-world object.
[0011] Another embodiment provides an apparatus comprising: a processor-based device; and a non-transitory storage medium storing a set of computer-readable instructions configured to cause the processor-based device to perform multiple steps, the steps comprising: identifying a real-world object in a scene viewed by a camera of a user device; matching the real-world object with a tagged object based at least in part on image recognition and a shared setting of tagged objects, the tagged object having been tagged with a content item; providing a notification to a user of the user device that the content item is associated with the real-world object; receiving a request for the content item from the user; and providing the content item to the user.
[0012] Another embodiment provides an apparatus comprising: a processor-based device; and a non-transitory storage medium storing a set of computer-readable instructions configured to cause the processor-based device to perform a plurality of steps comprising: identifying a real-world object in an image; matching the real-world object with a tagged object based at least in part on image recognition and a setting that has been assigned to the tagged object, the tagged object having been tagged with a content item, the content item being distinct from the setting; and providing a notification to a user that the content item is associated with the real-world object.
[0013] Another embodiment provides a method comprising: receiving a selection of a real-world object as viewed by a camera of a first user device; receiving a content item to tag to the real-world object; capturing one or more images of the real-world object with the camera; receiving sharing settings from the user, wherein the sharing settings include whether the real-world object will be matched only with images of the real-world object or with images of any object that shares one or more common attributes with the real-world object; and storing the content item, the one or more images of the real-world object, and the sharing settings on a server, wherein the content item is configured to be retrieved by a second user device that views an object that matches the real-world object.
[0014] Another embodiment provides a method comprising: receiving a selection of a real-world object from a first user; receiving a content item to tag to the real-world object; receiving one or more images of the real-world object; receiving settings from the first user, wherein the settings are distinct from the content item and include information about how the real-world object is to be matched with other images; and storing the content item, the one or more images of the real-world object, and the settings on a server, wherein the content item is configured to be retrieved by a second user viewing an object that matches the real-world object.
[0015] Another embodiment provides an apparatus comprising: a processor-based device; and a non-transitory storage medium storing a set of computer-readable instructions configured to cause the processor-based device to perform multiple steps, the steps comprising: receiving a selection of a real-world object as viewed by a camera of a first user device from a user; receiving a content item to tag to the real-world object; capturing one or more images of the real-world object with the camera; receiving sharing settings from the user, wherein the sharing settings include whether the real-world object will be matched only with images of the real-world object or with images of any object that shares one or more common attributes with the real-world object; and storing the content item, the one or more images of the real-world object, and the sharing settings on a server, wherein the content item is configured to be retrieved by a second user device that views an object that matches the real-world object.
[0016] Another embodiment provides an apparatus comprising: a processor-based device; and a non-transitory storage medium storing a set of computer-readable instructions configured to cause the processor-based device to perform a plurality of steps, the steps comprising: receiving a selection of a real-world object from a first user; receiving a content item to be tagged to the real-world object; receiving one or more images of the real-world object; receiving settings from the first user, wherein the settings are distinct from the content item and include information about how the real-world object will be matched with other images; and storing the content item, the one or more images of the real-world object, and the settings on a server, wherein the content item is configured to be retrieved by a second user viewing an object that matches the real-world object.
[0017] A better understanding of the features and advantages of various embodiments of the present invention will be obtained by reference to the following detailed description and accompanying drawings that set forth illustrative embodiments in which the principles of embodiments of the invention are utilized. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other aspects, features and advantages of embodiments of the present invention will become more apparent from the following more particular description of the invention presented in conjunction with the following drawings, in which:
[0019] Figure 1 is a diagram illustrating an apparatus for viewing a real-world scene according to some embodiments of the present invention;
[0020] Figure 2 is a diagram illustrating a method of virtually labeling a real-world object according to some embodiments of the present invention;
[0021] Figure 3 is a diagram illustrating a method of providing content items tagged to real-world objects according to some embodiments of the present invention;
[0022] Figure 4-6 are screenshots illustrating user interfaces according to some embodiments of the present invention;
[0023] Figure 7 is a diagram illustrating an apparatus for viewing a real-world scene according to some embodiments of the present invention;
[0024] Figure 8 is a block diagram illustrating a system that can be used to perform, implement, and / or execute any of the methods and techniques shown and described herein, according to an embodiment of the present invention. Detailed description
[0025] For storytelling purposes, augmented reality (AR) can be used to provide supplemental information about real-world objects. The following discussion will focus on an example embodiment of the present invention that allows users to share stories based on real-world objects through an AR storytelling system. In some embodiments, the AR storytelling system provides users with a way to publish and share memories and stories about real-world objects with a great degree of customization and control. Specifically, in some embodiments, users can tag real-world objects with content items such as text, photos, videos, and audio to share with other users. In some embodiments, the system provides notifications to the user when real-world objects in his / her surroundings have content items tagged to or associated with the real-world objects.
[0026] In some embodiments, using augmented reality, users can "tag" objects with personal content such as photos, videos, or voice recordings. In some embodiments, the AR system can use image recognition and / or a global positioning system (GPS) to detect objects in the real world that have been "tagged" by other users. When a tagged object is detected, the system can issue an alert for the object and allow the wearer or user of the AR device to review the information left in the tag (e.g., watch a video, view a photo, read a note, etc.). Tags can be shared through social networks based on "followed" people. Tags can exist permanently, expire after a user-selected amount of time, or be limited to a specific number of views (e.g., the tag is removed only after one person views it).
[0027] An example application of embodiments of the present invention might include a user tagging their favorite baseball cap with a video of themselves catching a home run. A friend who encounters a similar cap can see the cap and watch the video. As another example, a user tags the menu of their favorite restaurant with a voice note of the user's favorite items. Other users who see the menu can then listen to their friend's recommendations. Furthermore, as another example, a user tags a park bench with a childhood memory of a picnic there. Other users who later pass by can then engage in a nostalgic story.
[0028] Some embodiments of the present invention focus on using image recognition to provide physical reference points for digital data, which allows alerts to be triggered based on the user's interaction with the real world. In some embodiments, to place a tag, the user locates the desired object through an AR device, then selects the type of tag (e.g., photo, video, audio, text, etc.) and writes a "post" (e.g., title, description, hashtag, etc.). When the AR system detects a tagged object in the real world, the system can use any kind of image overlay to alert the user and direct the user to the information. In this way, some embodiments of the present invention can provide social interaction and / or information sharing through augmented reality.
[0029] refer to Figure 1 , showing an apparatus for viewing a real-world scene according to some embodiments of the present invention. Figure 1 , a user device 110 is used to view a real-world scene 120. The real-world scene 120 includes real-world objects such as a car 121, a tree 122, a lamp post 123, and a bench 124. The user device 110 may be a portable electronic device such as a smartphone, a tablet computer, a mobile computer, a tablet-like device, a head-mounted display, and / or a wearable device. The user device 110 includes a display screen 112 and other input / output devices such as a touch screen, a camera, a speaker, a microphone, a network adapter, a GPS sensor, a tilt sensor, and the like. The display screen 112 may display an image of the real-world scene 120 as viewed by an image capture device of the user device 110 as well as computer-generated images to enable storytelling through AR. For example, in Figure 1 In the example, brackets 113 are shown around the image of the car 121 in the display screen 112 to highlight the car 121. The user can select objects in the display screen 112 for virtual interaction. For example, the user can tag one of the objects with a content item or select an object to view the content item tagged to the object. Figure 2-8 To provide additional details of a method for providing AR storytelling.
[0030] Although Figure 1 A user device 110 is shown having a display screen 112 that reproduces an image of a real-world scene, but in some embodiments, the user device 110 may include transparent or semi-transparent lenses, such as those used in eyeglasses and / or head-mounted display devices. In the presence of such a device, the user can view the real-world scene 120 directly through one or more lenses on the device, and the computer-generated image is displayed in a manner so as to overlay the direct view of the real-world scene, thereby producing a visual effect similar to that of the real-world scene. Figure 1 and Figure 4-6 Combined views similar to those shown in .
[0031] refer to Figure 2 , shows an example of a method 300 for virtually tagging a real-world object operating in accordance with some embodiments of the present invention. In some embodiments, the steps of method 300 may be performed by one or more server devices, user devices, or a combination of server and user devices.
[0032] In step 210, a user selection of a real-world object in a scene viewed by a camera of a first user device is received. The scene viewed by the camera of the user device may be an image captured in real time or an image captured with a time delay. In some embodiments, the user selects the object by selecting an area in the image that corresponds to the object. The system may use the image of the selected area for object recognition. In some embodiments, the user may select the object by zooming or framing the image of the real scene so that the object is the primary object in the image. In some embodiments, the system may indicate in the image of the real-world scene the objects that the system automatically recognizes through object recognition. For example, Figure 1 As shown, the system may place brackets around an identified object (in this case, car 121) that the user can select. In some embodiments, the user may first select an area of the image, and the system will then determine the objects associated with the selected area of the image. For example, the user may touch the area above the image of car 121, and the system may identify trees 122 in the area. In some embodiments, the selected real-world object may be previously known to the system, and the system may extract the object from the image by using image processing techniques such as edge detection. For example, the system may not recognize lamppost 123 as a lamppost, but may be able to recognize the lamppost as a separate object based on edge detection. In some embodiments, the user may be prompted to type in attributes such as name, model number, object type, object class, etc. to describe the object so that the system can "learn" about previously unknown objects.
[0033] In step 220, the system receives a content item to be tagged to or associated with a real-world object. The content item can be one or more of the following: a text annotation, an image, an audio clip, a video clip, a hyperlink, etc. The system can provide an interface for a user to enter, upload, and / or capture the content item. For example, the system can provide a comment box and / or a button for selecting a file to upload.
[0034] In step 230, one or more images of the real-world object are captured with a camera of the user device. The captured images may be later processed to determine additional information related to the selected object. The real-world object may be identified based on the shape, color, relative size, and other attributes of the object that can be identified from the image of the real-world object. For example, the system may identify car 121 as a car based on its shape and color. In some embodiments, the system may further identify other attributes of the object. For example, the system may identify the make, model, and model year of car 121 based on a comparison of the captured image with images in a database of known objects. In some embodiments, the captured image may include specific identifying information of the object. For example, the image may include the license plate number of car 121. In other examples, the captured image may include the serial number or model number of an electronic device, a unique identifier for a limited edition collectible item, and the like.
[0035] In some embodiments, step 230 is performed automatically by the system without user input. For example, the system may acquire one or more images of the real-world object when the user selects the object in step 210. In some embodiments, the system may prompt the user to take images of the real-world object. For example, if the object selected in step 210 is not an object identified by the system, the system may prompt the user to take images of the object from different angles so that the object can be better matched with other images of the object. In some embodiments, after step 240, the system may prompt the user to capture additional images of the object based on the sharing settings. For example, if the user wants to tag the content item exactly only to car 121 and not any other car of similar appearance, the system may request an image of unique identifying information, such as the car's license plate (if the license plate is not visible in the current view of the car).
[0036] In general, the system can utilize various known computer object recognition techniques to identify objects in images of real-world scenes. In some embodiments, object recognition can use appearance-based methods that compare an image to reference images of known objects to identify the object. Examples of appearance-based methods include edge matching, grayscale matching, receptive field response histograms, etc. In some embodiments, object recognition can use feature-based methods that rely on matching object features and image features. Examples of feature-based methods include morphological clustering, geometric hashing, scale-invariant feature transforms, interpretation trees, etc. The system can use one or more object recognition methods in combination to improve the accuracy of object recognition.
[0037] In step 240, sharing settings are received from the user. In some embodiments, the sharing settings include whether the real-world object will only be matched with images of the real-world object or with images of any object that shares one or more common attributes with the real-world object. For example, if the user configures the sharing settings so that the real-world object is matched only with images of a specific real-world object, the system may determine and / or request the user to enter a unique identifier associated with the real-world image. The unique identifier may be based on the object's location, serial number, license plate number, or other identifying text or image of the object. For example, car 121 may be uniquely identified based on its license plate number and / or text or graphics painted on its body, bench 124 and lamppost 123 may be uniquely identified based on their location information and / or inscription, baseball caps may be uniquely identified based on a logo or wear on the cap, trophies may be uniquely identified by text on the trophy, and so on. In some embodiments, the system may automatically determine these identifying attributes. For example, if the system detects a license plate, name, or serial number in the image, the system may use that information as an identifying attribute. In another example, if the system identifies the real-world object as a stationary object, such as a lamppost 123, a building, a statue, etc., the system can use the location of the object or the location of the device capturing the image as an identifying attribute. In some embodiments, the system can prompt the user to specify an identifying attribute. For example, the system can prompt the user to take a picture of the serial number of the electronic device or type in the serial number as the identifying attribute.
[0038] If the user chooses to match a real-world object with an image of any object that shares one or more common attributes with the real-world object, the system may prompt the user to select matching attributes. Matching attributes refer to attributes that an object must possess in order to match a tagged object and receive content items tagged to the tagged object. In some embodiments, matching attributes may be based on the object's appearance, attributes of the object determined using image recognition, and / or the object's location. For example, a user may configure car 121 to match any car, cars of the same color, cars of the same make, cars of the same model, or cars of the same make and model year. While the color of a car can be determined based on the image alone, attributes such as the make of the car can be determined using image recognition. For example, the system may match the image of car 121 with reference images of known cars to determine the make, model, and model year of the car. Common attributes may further be location-based attributes. For example, a user may configure bench 124 to match any bench in the same park, any bench in the same city, any bench in a municipal park, any bench above a certain altitude, and so on. The location attribute information of the real-world object may be determined based on the device's GPS information and / or location metadata information of the captured image of the real-world object.
[0039] In some embodiments, the system may perform image recognition and / or position analysis on the selected real-world object and present a list of identified attributes to the user for selection. The user may configure sharing settings by selecting attributes for matching the real-world object with another object from the list. For example, if the user selects car 121 on user device 110, the system may generate a list of attributes such as:
[0040] Object Type: Car
[0041] Color: Gray
[0042] Brand: DeLorean
[0043] Model: DMC-12
[0044] Model Year: 1981
[0045] Location: Hill Valley, California
[0046] The user can then select one or more of the identified attributes to configure matching attributes in the sharing settings. For example, the user can select make and model as matching attributes, and content items tagged to car 121 will only be matched to another image of a DeLorean DMC-12. In some embodiments, the user can manually enter the attributes of the object to configure the sharing settings. For example, the user can configure a location-based sharing setting to any location within 20 miles or 50 miles of Hill Valley, California, etc.
[0047] In some embodiments, the sharing settings also include social network-based sharing settings that control who can view content items tagged to real-world objects based on the user's social network connections. For example, the user can choose to share a content item with: all "followers," only friends, a select group of users, or members of an existing group within a social network, etc.
[0048] Although Figure 2220, 230, and 240 are shown sequentially in FIG, but in some embodiments, these steps may be performed in a different order. For example, in some embodiments, the user may configure sharing settings before typing in the content item to tag to the object. In some embodiments, one or more images of the real-world item may be captured before receiving the content item or after typing in the sharing settings. In some embodiments, at any time during steps 220-240, the user may return to one of the previously performed steps to edit and / or enter information. In some embodiments, an object tagging user interface is provided to allow the user to perform two or more of steps 220, 230, and 240 in any order. In some embodiments, an image of the real-world object is captured when the user selects the real-world object in step 210. The system then determines whether more images are needed based on the sharing settings received in step 240, and prompts the user to capture one or more images as needed.
[0049] In step 250, the content item, the image of the real-world object, and the sharing settings entered in steps 220-240 are stored on a network-accessible database so that the content item is configured to be retrieved by a second user device viewing an object that matches the real-world object based on the sharing settings. Thus, in some embodiments, the content item is said to be tagged to or associated with the real-world object. In some embodiments, the content item, the image of the real-world object, and the sharing settings are stored as entered in steps 220-240. In some embodiments, after step 250, the user can return to edit or delete the information entered in steps 210-240 by logging into the user profile using the user device and / or another user electronic device.
[0050] refer to Figure 3 , which illustrates an example of a method 300 operating in accordance with some embodiments of the present invention. In some embodiments, the steps of method 300 can be used to provide AR storytelling as described herein. In some embodiments, method 300 can be performed by one or more server devices, user devices, or a combination of server and user devices.
[0051] In step 310, real-world objects in the scene viewed by the user device's camera are identified based on image recognition. The system can utilize any known image recognition technique. The scene viewed by the user device's camera can be an image captured in real time or with a time delay. In some embodiments, the user device uploads an image of the real-world scene to a server, and the server identifies real-world objects in the image. In some embodiments, the user device identifies attributes of the real-world objects (such as color, gradient map, edge information) and uploads the attributes to the server to reduce network latency. Real-world objects can be identified based on their shape, color, relative size, and other attributes that can be identified from images of real-world objects. In some embodiments, object recognition can be based on multiple video image frames. In some embodiments, objects can be identified based in part on the GPS location of the user device when the image was captured. For example, the system may be able to distinguish between two visually similar benches based on the device's location. In some embodiments, in step 310, the system will attempt to identify all objects viewed by the user device's camera.
[0052] In step 320, the real world object is matched with the tagged object. The tagged object can be Figure 2 . Thus, in some embodiments, the tagged object has a content item associated with it. The matching in step 320 is based on image recognition of the real-world object and the sharing settings of the tagged object. In some embodiments, the matching can be based on a comparison of the attributes of the real-world object with tagged objects in a tagged object database. For example, if the real-world object in step 310 is a red sports car, then the real-world object may match with tagged objects that specify matching with red cars, sports cars, etc., but will not match with a green sports car whose sharing settings limit sharing to only green cars. In some embodiments, matching is also based on the social network sharing settings of the tagged object. For example, if the tagged object has been specified to be shared only with the author's "friends", and the user using the device to view the real-world object is not connected with the author of the content item in the social networking service, then the real-world object will not match the particular tagged object.
[0053] In some embodiments, the matching in step 320 can also be based on the location of the device viewing the real-world object. For example, if the tagged object has been configured to be shared only with items within a geographic region (GPS coordinates, address, neighborhood, city, etc.), then the tagged object will not be matched with images of the real-world object taken outside of that region. In some embodiments, step 320 is performed for each object in the real-world scene identified in step 310.
[0054] In step 330, once a matching tagged object is found in step 320, a notification is provided to the user. The notification may be provided by one or more of: a sound notification, a pop-up notification, a vibration, and an icon in the display of the scene viewed by the camera of the user device. For example, the system may cause a graphical indicator to appear on the screen to indicate to the user that an item in the real-world scene has been tagged. The graphical indicator may be an icon that overlays the view of the object, brackets that surround the object, or a color / brightness change that highlights the object, etc. In some embodiments, if the object has been tagged with more than one content item, the notification may include an indicator for each content item. For example, if the item has been tagged by two different users or with different types of content items, the graphical indicator may also indicate the author and / or content type of each content item.
[0055] In some embodiments, a notification can be provided based solely on location information, followed by a prompt to the user to search for a tagged object in the surrounding environment. For example, if the system determines that a tagged object is nearby based on a match between the device and the GPS coordinates of the tagged object, the system can cause the device to vibrate or beep. The user can then point the device's camera at surrounding objects to locate the tagged object. The system can cause a graphical indicator to appear on the user's device when the tagged object appears in the camera's view.
[0056] In step 340, a request for a content item is received. The request can be generated by the user device when the user selects a real-world object and / or a graphical indicator in the user interface. For example, the user can tap an icon displayed on the touch screen of the user device to select a tagged object or content item. In some embodiments, if the object has been tagged with multiple content items, the selection can include a selection of content items, and the request can include an indication of a specific content item. In some embodiments, the request can be automatically generated by the user device, and the content item can be temporarily stored or cached on the user device. In some embodiments, requests may only be required for certain content item types. For example, the system can automatically transmit text and image content items after detecting a tagged object, but only stream audio and video type content items to the device after receiving a request from the user.
[0057] In step 350, the content items are provided to the user device. In some embodiments, all content items associated with the selected tagged object are provided to the user device. In some embodiments, only the selected content items are provided. In some embodiments, video and audio type content items are provided in a streaming format.
[0058] In some embodiments, after step 350, the viewing user may choose to annotate the content item and / or tag object by providing his / her own content item.
[0059] refer to Figure 4 , shows an example screenshot of a user interface for providing AR storytelling. The object selection user interface 410 displays an image of a real-world scene viewed by the device's camera and can be used by a user to have real-world objects tagged with virtual content items. In some embodiments, the user can select an area in the real-world scene associated with an object. For example, the user can tap an area, draw a circle around an area, and / or click and drag a rectangle around an area to select an area. In some embodiments, the system captures an image of the selected area and attempts to identify objects in the selected area. In some embodiments, individual objects can be identified based on edge detection. In some embodiments, the system identifies one or more objects identified by the system in the real-world scene presented to the user. For example, the system can overlay brackets 414 or other graphical indicators on the image of the real-world scene. The graphical indicators can be user-selectable.
[0060] After an object has been selected in the real-world scene view, a tagging user interface 420 may be displayed. The tagging user interface 420 may include an image 422 of the object selected to be tagged. The tagging user interface 420 may also include various options for typing and configuring content items. For example, a user may use an "add content" button 424 to type or attach a text note, image, audio file, and / or video file. In some embodiments, the user interface also allows the user to record new images, audio, and / or video to tag an object without leaving the user interface and / or application. The user may use a "share settings" button 426 to configure the various sharing settings discussed herein. For example, a user may select matching attributes, configure social network sharing settings, and / or type location restrictions in the tagging user interface 420.
[0061] refer to Figure 5 , shows additional example screenshots of user interfaces for providing AR storytelling. A tagged object notification user interface 510 can be displayed to a user viewing a real-world scene through a user device. The tagged object notification user interface 510 can include one or more graphical indicators to identify tagged objects in the view of the real-world scene. For example, icon 512 indicates that a lamp post has been tagged with a text annotation, icon 514 indicates that a bench has been tagged with an audio clip, and icon 516 indicates that a car has also been tagged with a text annotation. In some embodiments, other types of graphical indicators are used to indicate that an object has a viewable content item. For example, an icon similar to Figure 4 In some embodiments, the color and / or shading of the object or its surroundings may be changed to distinguish the marked object from its surroundings. For example, Figure 6A car 610 is shown highlighted against a shaded background 620 .
[0062] In some embodiments, the tagged object notification user interface 510 can be the same interface as the tagging user interface 410. For example, when a user opens a program or application (or "app") on a user device, the user is presented with an image of a real-world scene viewed by the user device's camera. If any of the objects in the real-world scene are tagged, a notification can be provided. The user can also select any object in the scene to virtually tag a new content item to the object.
[0063] When a tagged object or content item is selected in the tagged object notification user interface 510, the content item viewing user interface 520 is shown. In the content item viewing user interface 520, the content item selected from the tagged object notification user interface 510 is retrieved and displayed to the user. The content can be full screen or in a variety of formats such as Figure 5 An image of a real-world scene is shown as an overlay 522. The content item may include one or more of the following: a text annotation, an image, an audio clip, and a video clip. In some embodiments, if more than one content item is associated with a tagged object, multiple content items from different authors may be displayed simultaneously.
[0064] The object selection user interface 410, the tagging user interface 420, the tagged object notification user interface 510, and the content item viewing user interface 520 can each be part of a program or application running on a user device and / or a server device. In some embodiments, the user interface can be part of an app on a smartphone that communicates with the AR server to provide images and content that enable AR storytelling on the user interface. In some embodiments, the user interface is generated and provided entirely on the server.
[0065] refer to Figure 7 , shows a diagram illustrating a glasses-type user device for viewing a real-world scene according to some embodiments of the present invention. In some embodiments, the above user interface can also be implemented using glasses-type and / or head-mounted display devices. In the presence of such a device, the user can view the real-world 710 scene directly through the lens 720 on the device, and the computer-generated image 722 can be projected on the lens and / or in the wearer's eyes in a manner so as to overlay the direct view of the real-world scene, thereby producing a scene such as Figure 7The combined view shown. The overlay image 722 may indicate a selected or selectable object for marking and / or identify a marked object as described herein. The user's selection of an object may be detected by gaze detection and / or voice command recognition. In some embodiments, the object may be selected by a touchpad on the device. For example, a user may swipe on the touchpad to "scroll" through the selectable objects viewed through the lens of the device. In some embodiments, the user may add text annotations or audio clip content items by verbal commands received via a microphone of the head-mounted display. In some embodiments, the selection may be made by gesturing in a view of a real-world scene captured and interpreted by a camera on the device. Although Figure 7 Only the frame of one lens is shown in the figure, but it should be understood that the same principles apply to a head-mounted device with two lenses.
[0066] Figure 8 800 is a block diagram illustrating a system that can be used to perform, implement, and / or execute any of the methods and techniques shown and described herein, according to some embodiments of the present invention. System 800 includes user devices 810, 820, and 830, an AR server 840, and a third-party / social network server 870 communicating over a network 805.
[0067] The user device 810 may be any portable user device, such as a smartphone, a tablet computer, a mobile computer or device, a tablet-like device, a head-mounted display, and / or a wearable device. The user device 810 may include a processor 811, a memory 812, a network interface 813, a camera 814, a display 815, one or more other input / output devices 816, and a GPS receiver 817. The processor 811 is configured to execute computer-readable instructions stored in the memory 812 to facilitate reference to the user device. Figure 2-3 One of the multiple steps of the method described. The memory 812 may include RAM and / or hard drive memory devices. The network interface 813 is configured to transmit data to at least the AR server 840 via the network 805 and receive data therefrom. The camera 814 can be any image capture device. The display 815 can be a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a liquid crystal on silicon (LCoS) display, an LED light emitting display, and the like. In some embodiments, the display 815 is a touch display configured to receive input from a user. In some embodiments, the display 815 is a head-mounted display that projects an image into a lens or the wearer's eyes. The display can be used to display the above referenced images in different ways. Figure 1-7The user interface described herein. The GPS receiver 817 can be configured to detect GPS signals to determine coordinates. The determined GPS coordinates can be used by the system 800 to determine the location of the device, image, and / or object. The user device 810 can also include other input / output devices 816, such as a microphone, a speaker, a tilt sensor, a compass, a USB port, an auxiliary camera, a graphics processor, etc. The input / output devices 816 can be used to enter and play back content items. In some embodiments, the tilt sensor and compass are used by the system to better position computer-generated graphics on an image of a real-world scene (except for graphic recognition).
[0068] User devices 820 and 830 may be user devices similar to user device 810 that may be operated by one or more other users to access AR storytelling system 800. It should be understood that although Figure 8 Three user devices are shown in , but any number of users and user devices may access system 800 .
[0069] AR server 840 includes a processor 842, a memory 841, and a network interface 844. Processor 842 is configured to execute computer readable instructions stored in memory 841 to perform the functions referred to herein. Figure 2-3 The AR server 840 may be connected to or include one or more of an object database 860 and a content item database 850. The content item database 850 may store information related to each of the following: tagged real-world objects, content items tagged to each object, and sharing settings associated with the content items and / or tagged objects. In some embodiments, Figure 2 Step 250 in the content item database 850 is performed by storing the information in the content item database 850. The object database 860 may include a database of known objects used by the system's object recognition algorithm to identify one or more real-world objects and their attributes. In some embodiments, an image of a real-world object captured by the camera 814 of the user device 810 is compared with one or more images of the object in the object database 860. When a match occurs, the object database 860 may provide further information about the object, such as the object name, object type, object model, etc.
[0070] In some embodiments, one or more of the object database 860 and the content item database 850 can be part of the AR server 840. In some embodiments, the object database 860 and the content item database 850 can be implemented as a single database. In some embodiments, the AR server 840 also communicates with one or more of the object database 860 and the content item database 850 via the network 805. In some embodiments, the object database 860 can be maintained and controlled by a third party.
[0071] In some embodiments, the object database 860 can "learn" new objects by receiving user-provided images and attributes and adding the user-provided information to its database. For example, when a user takes a picture of a car and enters its make and model, the object database may then be able to recognize another image of the same car and provide information about the make and model of the same car. Although only one AR server 840 is shown, it should be understood that the AR server 840, the object database 860, and the content item database 850 can be implemented using one or more physical devices connected via a network.
[0072] The social network server 870 provides social networking functionality for users to connect with each other and build social networks and social groups. The social network server 870 can be part of the AR server 840 or a third-party service. The connections and built groups in the social network service can be used to configure the sharing settings discussed herein. For example, if a content item is configured to be shared only with the author's "friends", the AR server 840 can query the social network status between the two users from the social network server 870 to determine whether the content item should be provided to the second user. In some embodiments, when a user configures the sharing settings, information can be retrieved from the social network server 870 so that the user can choose among his / her friends and / or social groups to share the content item.
[0073] In some embodiments, one or more of the embodiments, methods, approaches and / or technologies described above can be implemented in one or more computer programs or software applications that can be executed by a processor-based device or system. For example, such a processor-based system can include a processor-based device or system 800, or a computer, entertainment system, game console, graphics workstation, server, client, portable device, a device similar to a tablet computer, etc. This type of computer program can be used to perform the various steps and / or features of the methods and / or technologies described above. In other words, the computer program can be suitable for enabling a processor-based device or system to perform and implement the functions described above or configuring the device or system to perform and implement the functions described above. For example, this type of computer program can be used to implement any embodiment of the methods, steps, technologies or features described above. As another example, this type of computer program can be used to implement any type of tool or similar utility using any one or more of the embodiments, methods, approaches and / or technologies described above. In some embodiments, program code macros, modules, loops, subroutines, calls, etc. within or outside the computer program can be used to perform the various steps and / or features of the methods and / or technologies described above. In some embodiments, the computer program may be stored or included on one or more computer-readable storage or recording media, such as any of the one or more computer-readable storage or recording media described herein.
[0074] Thus, in some embodiments, the present invention provides a computer program product comprising a medium for comprising a computer program for computer input; and a computer program included in the medium, the computer program for causing a computer to perform or execute a plurality of steps, the steps comprising any one or more of the steps involved in any one or more of the embodiments, methods, approaches, and / or techniques described herein. For example, in some embodiments, the present invention provides one or more non-transitory computer-readable storage media storing one or more computer programs adapted to cause a processor-based device or system to perform a plurality of steps, the steps comprising: identifying a real-world object in a scene viewed by a camera of a user device; matching the real-world object with a tagged object based at least in part on image recognition and a shared setting of tagged objects, the tagged object having been tagged with a content item; providing a notification to a user of the user device that the content item is associated with the real-world object; receiving a request for the content item from the user; and providing the content item to the user. In another example, in some embodiments, the present invention provides one or more non-transitory computer-readable storage media storing one or more computer programs adapted to cause a processor-based device or system to perform a plurality of steps comprising: receiving a selection of a real-world object as viewed by a camera of a first user device from a user; receiving a content item to tag to the real-world object; capturing one or more images of the real-world object with the camera; receiving sharing settings from the user, wherein the sharing settings include whether the real-world object will be matched only with images of the real-world object or with images of any object that shares one or more common attributes with the real-world object; and storing the content item, the one or more images of the real-world object, and the sharing settings on a server, wherein the content item is configured to be retrieved by a second user device that views an object that matches the real-world object.
[0075] While the invention disclosed herein has been described with respect to specific embodiments and applications thereof, numerous modifications and variations can be made thereto by those skilled in the art without departing from the scope of the invention as set forth in the claims.
Claims
1. A method for image processing, the method comprising: Recognize real-world objects in images; matching the real-world object to the tagged object based at least in part on image recognition and a shared setting that has been assigned to the tagged object, wherein the tagged object has been tagged with a content item that is distinct from the shared setting and includes one or more of: a text annotation, an image, an audio clip, or a video clip; as well as Notification is provided to a user that the content item is associated with the real-world object. 2 . The method of claim 1 , wherein the sharing setting includes whether the tagged object will only be matched with images of the tagged object or with images of any object that shares one or more common attributes with the tagged object. 3 . The method of claim 1 , wherein the matching of the real-world object to the tagged object is further based on one or more of: a location of the user or a social network connection between the user and an author of the content item.
4. The method of claim 1 or 2, wherein the notification comprises one or more of: a sound notification, a pop-up notification, a vibration, or an icon in a display of a user device.
5. A device for image processing, comprising: processor-based devices; as well as A non-transitory storage medium storing a set of computer-readable instructions configured to cause the processor-based device to perform a plurality of steps comprising: Recognize real-world objects in images; matching the real-world object to the tagged object based at least in part on image recognition and a shared setting that has been assigned to the tagged object, wherein the tagged object has been tagged with a content item that is distinct from the shared setting and includes one or more of: a text annotation, an image, an audio clip, or a video clip; and Notification is provided to a user that the content item is associated with the real-world object. 6 . The apparatus of claim 5 , wherein the sharing setting includes whether the tagged object will be matched only with images of the tagged object or with images of any object that shares one or more common attributes with the tagged object. 7 . The apparatus of claim 5 , wherein the matching of the identified real-world object to the tagged object is further based on one or more of: a location of the user or a social network connection between the user and an author of the content item.
8. The apparatus of claim 5, wherein the notification comprises one or more of: a sound notification, a pop-up notification, a vibration, or an icon in a display of the user device.
9. A method for image processing, the method comprising: receiving a selection of a real-world object from a first user; receiving a content item to be tagged to the real-world object; receiving one or more images of the real-world object; receiving sharing settings from the first user, wherein the sharing settings are distinct from the content item and include information regarding how the real-world object is to be matched with other images; as well as storing the content item, the one or more images of the real-world object, and the sharing settings on a server, wherein the content item is configured to be retrieved by a second user viewing an object that matches the real-world object, The content item includes one or more of the following: a text annotation, an image, an audio clip, or a video clip.
10. The method of claim 9, wherein the sharing settings further include restrictions on who can view the content item.
11. The method of any one of claims 9 and 10, wherein the sharing setting is entered by the first user by selecting from a list of properties corresponding to the real-world object.
12. The method of claim 11, wherein the attribute list is generated by object recognition performed on the one or more images of the real-world object.
13. A device for image processing, the device comprising: processor-based devices; as well as A non-transitory storage medium storing a set of computer-readable instructions configured to cause the processor-based device to perform a plurality of steps comprising: receiving a selection of a real-world object from a first user; receiving a content item to be tagged to the real-world object; receiving one or more images of the real-world object; receiving sharing settings from the first user, wherein the sharing settings are distinct from the content item and include information regarding how the real-world object is to be matched with other images; and storing the content item, the one or more images of the real-world object, and the sharing settings on a server, wherein the content item is configured to be retrieved by a second user viewing an object that matches the real-world object, The content item includes one or more of the following: a text annotation, an image, an audio clip, or a video clip.
14. The apparatus of claim 13, wherein the sharing settings further include restrictions on who can view the content item.
15. The apparatus of any one of claims 13 and 14, wherein the sharing setting is entered by the first user by selecting from a list of properties corresponding to the real-world object.
16. The apparatus of claim 15, wherein the attribute list is generated by automatic object recognition performed on the one or more images of the real-world object.