Platform for registering and processing visual coding

A multi-system approach with contextual validation improves visual encoding accuracy and security by leveraging multiple recognition systems and contextual data, addressing inefficiencies and security risks in visual encoding systems.

JP7832147B2Active Publication Date: 2026-03-17GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing visual encoding systems face challenges in accurately interpreting visual patterns due to insufficient image detail and vulnerability to malicious activities, leading to inefficiencies and security risks.

Method used

A computing system utilizing multiple recognition systems operating on different data signals and processing techniques collaboratively to recognize visual codes across varying distances and environments, incorporating contextual data for validation and fraud prevention.

Benefits of technology

Enhances accuracy and security in visual code processing by leveraging multiple recognition systems and contextual information, reducing fraud and improving object recognition, even in suboptimal imaging conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007832147000001
    Figure 0007832147000001
  • Figure 0007832147000002
    Figure 0007832147000002
  • Figure 0007832147000003
    Figure 0007832147000003
Patent Text Reader

Abstract

To provide a platform for registering and / or processing visual encodings.SOLUTION: The present disclosure relates generally to the processing of machine-readable visual encodings in view of contextual information. One embodiment of aspects of the present disclosure comprises obtaining image data descriptive of a scene that includes a machine-readable visual encoding; processing the image data with a first recognition system configured to recognize the machine-readable visual encoding; processing the image data with a second, different recognition system configured to recognize a surrounding portion of the scene that surrounds the machine-readable visual encoding; identifying a stored reference associated with the machine-readable visual encoding based at least in part on one or more first outputs generated by the first recognition system based on the image data and based at least in part on one or more second outputs generated by the second recognition system based on the image data; and performing one or more actions responsive to identification of the stored reference.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to visual encoding. More particularly, the present disclosure relates to a platform for registration and / or processing of visual encoding.

Background Art

[0002] Visual patterns are used to encode information in the form of visual encoding. For example, barcodes are commonly used to communicate information about products within a store. Searching for information encoded in such a pattern generally requires imaging with sufficient acuity to distinguish certain features of the pattern (e.g., to detect and process the bars of a barcode). In some practical applications, obtaining an image of the pattern with sufficient detail to interpret the pattern presents many challenges. Additionally, since visual patterns are not always interpretable by a user, the user may be vulnerable to malicious activities associated with improper visual patterns.

Summary of the Invention

Means for Solving the Problems

[0003] Aspects and advantages of embodiments of the present disclosure are described in part in the following description, or can be learned from the description, or can be known through practice of the embodiments.

[0004] One exemplary aspect of the present disclosure relates to a computer implementation method for processing machine-readable visual coding. The method includes the step of acquiring image data describing a scene containing machine-readable visual coding by a computing system comprising one or more computing devices. The method includes the step of processing the image data by the computing system using a first recognition system configured to recognize machine-readable visual coding. The method includes the step of processing the image data by the computing system using a second distinct recognition system configured to recognize the surrounding portion of the scene enclosing the machine-readable visual coding. The method includes the step of the computing system identifying stored criteria related to machine-readable visual coding, at least in part on one or more first outputs generated by the first recognition system based on the image data, and at least in part on one or more second outputs generated by the second recognition system based on the image data. The method includes the step of the computing system performing one or more actions in response to the identification of the stored criteria.

[0005] Another exemplary aspect of the present disclosure relates to a computing system for performing actions in response to the recognition of machine-readable visual coding. The computing system includes one or more processors and one or more non-temporary computer-readable media for storing instructions together, which, when executed by one or more processors, cause the computing system to perform actions. These actions include acquiring image data describing a scene containing machine-readable visual coding. These actions include transmitting the image data to a recognition server system for processing by the recognition server system, which comprises a first recognition system configured to recognize machine-readable visual coding and a second distinct recognition system configured to recognize the surrounding portion of a scene containing machine-readable visual coding. These actions include performing one or more actions based at least in part on an association between stored criteria and image data, wherein the association is determined by the recognition server system using one or more first outputs generated by the first recognition system and one or more second outputs generated by the second recognition system.

[0006] Another exemplary aspect of the present disclosure relates to one or more non-temporary computer-readable media for storing instructions together, which, when executed by one or more processors, cause one or more processors to perform an operation. These operations include receiving image data describing a scene including machine-readable visual coding, wherein the image data further describes the context of the machine-readable visual coding. These operations include associating the image data with a stored criterion. Associating includes determining a first similarity between the machine-readable visual coding and the stored criterion using a first recognition system, and determining a second similarity between the context and the stored criterion using a second recognition system. These operations include initiating one or more operations based on the association.

[0007] Other aspects of this disclosure cover a variety of systems, apparatus, non-temporary computer-readable media, user interfaces, and electronic devices.

[0008] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood by referring to the following description and the appended claims. The appended drawings are incorporated herein and form part thereof, illustrating exemplary embodiments of this disclosure and, together with the description, are useful in illustrating the relevant principles.

[0009] A detailed discussion of embodiments intended for those skilled in the art is provided herein, and this specification refers to the accompanying drawings. [Brief explanation of the drawing]

[0010] [Figure 1] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 2] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 3] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 4] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 5] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 6] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 7] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 8] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 9A] This figure shows machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 9B] This figure shows the progress of rendering on machine-readable visual coding according to exemplary embodiments of the present disclosure. [Figure 10A] This figure shows an image of a scene including machine-readable visual coding according to an exemplary embodiment of the present disclosure. [Figure 10B] This figure shows an image of a scene including machine-readable visual coding according to an exemplary embodiment of the present disclosure. [Figure 10C] This figure shows an image of a scene including machine-readable visual coding according to an exemplary embodiment of the present disclosure. [Figure 10D] This figure shows the mapping of image features in a scene including machine-readable visual coding according to an exemplary embodiment of the present disclosure. [Figure 11A] This figure shows an image of a scene including machine-readable visual coding according to an exemplary embodiment of the present disclosure. [Figure 11B] This figure shows the mapping of image features in a scene including machine-readable visual coding according to an exemplary embodiment of the present disclosure. [Figure 11C] This figure shows the generation of reference data by a user according to an exemplary embodiment of the present disclosure. [Figure 11D] This figure shows the generation of reference data by a user according to an exemplary embodiment of the present disclosure. [Figure 12] This figure shows a package delivery system using machine-readable visual coding according to an exemplary embodiment of the present disclosure. [Figure 13A] This is a block diagram of an exemplary computing system according to an exemplary embodiment of the present disclosure. [Figure 13B] This is a block diagram of an exemplary computing device according to an exemplary embodiment of the present disclosure. [Figure 13C] This is a block diagram of an exemplary computing device according to an exemplary embodiment of the present disclosure. [Figure 13D] This is a block diagram of an exemplary recognition system according to an exemplary embodiment of the present disclosure. [Figure 13E] This is a block diagram of an exemplary recognition system according to an exemplary embodiment of the present disclosure. [Figure 14] This is a flowchart illustrating an exemplary method according to an exemplary embodiment of the present disclosure.

Best Mode for Carrying Out the Invention

[0011] Reference numerals repeated throughout multiple drawings are intended to identify the same features in various implementations.

[0012] Summary Generally, the present disclosure is directed to the processing of visual encoding in view of context information. The context information can be used in some embodiments to assist in the recognition, identification, and / or processing of visual encoding. The use of context information according to aspects of the present disclosure can advantageously provide a robust platform for recognizing and processing visual encoding.

[0013] One exemplary aspect of this disclosure relates to a computing system and method that uses multiple data signals and / or multiple processing systems to process or recognize a visual code (which may, for example, correspond to an action point in the real world). The use of multiple systems, each operating on a different data signal and / or with a different processing technique, may enable the system to effectively recognize the visual code over a wide range of different distances. For example, in the case of a visual code placed on a wide surface or otherwise in a particular environment or scene, a user may capture an image from a visual code at various distances from the visual code, including situations where the visual code is outside the field of view or is completely or partially obscured, for example, from a close-up view to several meters or even meters away. Conventional visual code processing systems cannot handle this variability in image distance. To address this challenge, some exemplary implementations of this disclosure leverage multiple different recognizers that operate collaboratively (e.g., in parallel and / or in series) to process the visual code (e.g., to produce or unlock an augmented reality experience). As an example, some exemplary systems may employ three different recognition systems: a near-field visual coding reader that directly processes the visual coding; an image recognition-based system that recognizes semantic entities, objects, and / or other content known to be contained within the surrounding area of ​​the coding; and a visual positioning system capable of recognizing location based on visual or spatial features. Other exemplary systems may operate based on various other forms of contextual data, such as time and location (e.g., provided by a GPS system). Some exemplary systems may also feature a low-power digital signal processor that triggers the above systems to operate when the visual coding is first detected. Each of these systems may offer best performance at different distances.For example, a near-range reader may provide relatively best performance at close range (e.g., 10 centimeters), an image recognition-based system may provide relatively best performance at medium range (e.g., 10 meters), and a visual positioning system may provide relatively best performance at longer ranges. Various logical or triaging algorithms can be used to provide collaborative action from multiple systems. For example, the most efficient system may operate first, and if such a system is unable to recognize the coding, the next most efficient system may be triggered. As another example, multiple systems may operate in parallel, and the results may be combined (e.g., by a voting / trust or first-to-recognize paradigm). Using multiple different recognition systems in this way can enable visual coding processing with improved accuracy and efficiency. For example, the use of a visual positioning system allows a computing system to eliminate ambiguity between the same visual coding placed in multiple different locations (e.g., visual coding present on widely distributed movie posters).

[0014] Another exemplary aspect of this disclosure relates to location-based payment fraud protection using contextual data. Specifically, one exemplary application of visual coding is to enable payments. However, some visual coding (e.g., QR Code®) adheres to public protocols and allows anyone to generate codes that look identical to a human viewer. Thus, a malicious party could print similar visual coding and place it against existing commercial coding. Malicious coding could reroute users scanning the coding to a fraudulent payment portal. To address this issue, exemplary implementations of this disclosure may use contextual information to validate visual coding and / or delay or prevent users from making coding-based payments against fraudulent coding. Specifically, as described above, the computing system may include and use multiple different processing systems, each operating on different signals and / or using different processing techniques. Accordingly, exemplary implementations of this disclosure enable associating a visual code with some or all of such different signals or information, including, for example, visual feature points surrounding the code, surrounding semantic entity data, location data such as GPS data, ambient noise (e.g., high-speed noise, payment terminal noise), and / or other contextual information. Whenever a visual code is scanned, these additional data points can then be evaluated and associated with the visual code. This adds validation to each visual code in a texture that can later be used to detect incorrect coding. Specifically, as an example, if a visual code is scanned for the first time and the contextual data cannot be confirmed for such a visual code, the system may warn the user before proceeding. As another example, if a new visual code is scanned and has contextual information that matches an existing code, the new visual code may be blocked or subject to an additional validation routine.Therefore, continuing with the example given above, a contextual "fingerprint" may be generated over time for authenticated and registered visual coding in a particular location or business. In that case, if a malicious party seeks to cover or replace existing coding, the contextual "fingerprint" for the malcode will match the existing coding, triggering one or more anti-fraud techniques.

[0015] Another exemplary aspect of this disclosure relates to a technique that enables improved object (e.g., product) recognition using contextual data. Specifically, the use of barcodes is a technique that has been proven over many years. However, the indexes that the technique uses to find results are often flawed or incomplete. Even with standardized barcodes, coverage is often only around 80%. To address this challenge, an exemplary implementation of this disclosure may use visual information (e.g., as well as other contextual signals) surrounding a visual code as a supplementary path for recognizing an object (e.g., a return result) when a user attempts to scan the visual code. For example, consider an image of product packaging containing a visual code (e.g., a barcode). If the visual code is not effectively registered in the relevant index, a typical system may return incorrect results or no results at all. However, the exemplary systems described herein may use other features to identify an object. For example, it may recognize text or an image of a product present on the packaging, which can then be searched as a product or a general web query. In another example, multiple images of visual codes or the surrounding scene may be joined together in a "session" to form a more complete understanding of the product and update the index. Specifically, if a user scans for a visual code that is not present in the index, but contextual signals are used to recognize the relevant object / product, the relevant object / product and visual code may be added to the index. Thus, mapping codes to objects in the index may evolve or be supplemented over time based on anonymized aggregate user involvement. The proposed system may also work to eliminate ambiguity between multiple potential solutions for visual codes. For example, sometimes mismatched product codes may contain the same ID for different products.Therefore, some exemplary implementations may capture the visual features surrounding the visual coding and use those visual features as secondary identifying information, for example, to distinguish between competing ID maps.

[0016] The proposed system and method can be used in several different applications or use cases. For example, the proposed visual coding platform could be used to authenticate and / or enable package delivery to secure or other access-restricted locations. For example, an image (e.g., captured by a camera-enabled doorbell) could be used to authenticate expensive deliveries and coordinate a very short unlock / relock protocol with the door, including a smart lock. During this time, the package delivery person could place the product inside the secure location rather than manually receiving the package. Thus, in an example, each delivery person could have a temporarily generated code that attaches their account to any delivery unlock. Each package may have a unique code that can associate the visual coding platform with the product purchase. If value criteria set by the delivery person / recipient are met, the door can be unlocked when the delivery person presents the package code to the doorbell camera. The door may remain unlocked for a short period of time, until the door is closed, or until the delivery person scans its badge. Thus, visual coding can represent data beyond mere identification of the object itself. For example, the processing of visual coding could trigger a detailed process (performed, for example, by multiple cloud-based or IoT devices) for security verification of fully automated (i.e., no user intervention) package delivery.

[0017] More specifically, systems and methods according to embodiments of the present disclosure can identify one or more machine-readable visual codes in a scene. In some embodiments, machine-readable visual codes may include, and combinations thereof, one-dimensional (1D) patterns, two-dimensional (2D) patterns, three-dimensional (3D) patterns, four-dimensional (4D) patterns, and combinations of 1D and 2D patterns with other visual information, such as photographs, sketches, drawings, logos, etc. One exemplary embodiment of a machine-readable visual code is a QR code®. A scene may include a figure of one or more machine-readable visual codes that provides a visual and / or spatial context for the machine-readable visual codes within it, and into which they are unfolded. For example, in some embodiments, a scene may include only a figure of a machine-readable visual code, but in some embodiments, a scene may include the machine-readable visual codes, and any objects, people, or structures to which the machine-readable visual codes are attached, as well as other objects, people, or structures near and / or otherwise visible in a captured image of the scene.

[0018] The systems and methods of this disclosure can identify one or more machine-readable visual codes in a scene based on processing image data describing the scene. In some examples, the image data may include photographs or other spectral imaging data, metadata, and / or their encodings or other representations. For example, photographs or other spectral imaging data may be captured by one or more sensors of a device (e.g., a user device). Photographs or other spectral imaging data may be acquired from one or more imaging sensors on a device (e.g., one or more cameras on a mobile phone) and, in some exemplary embodiments, may be stored as bitmap image data. In some embodiments, photographs or other spectral imaging data may be acquired from one or more exposures from each of a plurality of imaging sensors on a device. The imaging sensors may capture spectral information including wavelengths visible to the human eye and / or invisible to the human eye (e.g., infrared). Metadata can provide context for image data and may include, in some examples, geographical data (e.g., location relative to nearby mapped elements such as GPS location, roads and / or businesses, device pose relative to imaging features such as machine-readable visual coding), temporal data, telemetry data (e.g., device orientation, velocity, acceleration, altitude, etc., related to image data), account information related to user accounts for one or more services (e.g., accounts corresponding to devices related to image data), and / or preprocessing data (e.g., data generated by image data preprocessing, such as by the device generating the image data). In one embodiment, the preprocessing data may include depth mapping of features of imaging data captured by one or more imaging sensors.

[0019] Examples of image data may include encoding or other representations of imaging data captured by one or more imaging sensors. For example, in some embodiments, the image data may include hash values ​​generated based on bitmap imaging data. In this way, a representation of bitmap imaging data may be used in place of or in addition to the bitmap imaging data itself (or one or more parts thereof) within the image data.

[0020] Image data can be processed by one or more recognition systems. For example, one or more parts (and / or the entirety) of image data can be processed by multiple recognition systems. The recognition systems may comprise a general-purpose image recognition model and / or one or more models configured for specific recognition tasks. For example, in some embodiments, the recognition systems may include a face recognition system, an object recognition system, a landmark recognition system, a deep mapping system, a machine-readable visual coding recognition system, an optical character recognition system, a semantic analysis system, and the like.

[0021] In some embodiments, image data describing a scene containing machine-readable visual coding may be processed by a coding recognition system configured to recognize machine-readable visual coding. In some embodiments, only a portion of the image data describing the machine-readable visual coding (e.g., a segment of the image) may be processed by the coding recognition system, but the coding recognition system may be configured to process all or part of the image data. For example, in some embodiments, a preprocessing system recognizes the presence of machine-readable visual coding and extracts a portion of the image data relating to the machine-readable visual coding. The extracted portion may then be processed by the coding recognition system for identification by the coding recognition system. Alternatively or in addition, the entire image data may be processed directly by the coding recognition system.

[0022] In some embodiments, image data may be further processed by one or more additional recognition systems. For example, image data may be processed by a coding recognition system configured to recognize machine-readable visual codes, and by another recognition system distinct from the coding recognition system. In one embodiment, the other recognition system may be configured to recognize various aspects of the context related to the machine-readable visual codes. For example, the other recognition system may include a recognition system configured to recognize one or more parts of a scene described by the image data (e.g., the surrounding part of the scene that encloses one or more machine-readable visual codes contained within the scene). In some embodiments, the other recognition system may include a recognition system configured to recognize spectral information conveyed within the image data (e.g., infrared exposure). In some embodiments, the other recognition system may include a recognition system configured to process metadata contained within the image data.

[0023] In some embodiments, recognizing one or more features of image data may involve associating those features with one or more stored criteria. In some embodiments, stored criteria may be a set of data registered in correspondence with one or more machine-readable visual codes (and / or one or more entities associated therewith). For example, stored criteria may be created when one or more machine-readable visual codes are generated to register data associated with one or more machine-readable visual codes (e.g., metadata, nearby objects, people, or structures associated with the implementation of one or more machine-readable visual codes). In some embodiments, data describing stored criteria may be received by one or more recognition systems simultaneously with image data describing a scene containing one or more machine-readable visual codes. For example, stored criteria may include data describing the context in which the machine-readable visual codes are located (e.g., location, time, environmental information, nearby devices, networks, etc.) which is updated at or near (e.g., simultaneously) the time when the image data describing the scene containing the machine-readable visual codes is captured and / or received.

[0024] In some embodiments, a stored set of criteria may be associated with the same predetermined algorithm and / or standard for generating machine-readable visual codes. For example, a machine-readable visual code may be generated by encoding one or more data items (e.g., instructions and / or information) into a visual pattern according to a predetermined algorithm and / or standard. In this way, a coding recognition system may recognize that the machine-readable visual code is associated with at least one of the stored criteria of the stored set of criteria, and may process the machine-readable visual code according to a predetermined algorithm and / or standard to decode the visual pattern and retrieve one or more data items.

[0025] A stored criterion may include information describing one or more machine-readable visual codes and / or information describing the context associated with one or more of the machine-readable visual codes. For example, a stored criterion may include data describing contextual features of a machine-readable visual code, such as visual features that one or more data items do not encode (e.g., design and / or aesthetic features, shape, size, orientation, color, etc.). In this way, one or more machine-readable visual codes may each be associated with one or more stored criteria (or, optionally, one or more sub-criteria of a single stored criterion) at least in part on the unencoded contextual features of one or more machine-readable visual codes. In one example, two machine-readable visual codes may encode the same data item (e.g., an instruction to perform an action). However, image data describing machine-readable visual codes may show that one of the machine-readable visual codes is depicted in a different shape than, for example, another of those machine-readable visual codes, and a coding recognition system can thereby distinguish the machine-readable visual codes and associate each of them with different stored criteria and / or subcriteria.

[0026] For example, a stored criterion may include data describing other contexts related to one or more machine-readable visual codes, such as a scene in which one or more machine-readable visual codes can be found and / or metadata related to one or more machine-readable visual codes. In this way, portions of image data describing contextual information can also be recognized as being related to one or more stored criterions. For example, one or more image recognition systems may process image data describing contextual information for recognizing one or more people, objects, and / or structures represented thereby, and may optionally determine whether there is any relationship (e.g., relative placement, semantic association, etc.) between such contextual information and one or more machine-readable visual codes in the scene. In some examples, one or more recognition systems may associate image data with one or more stored criterions in terms of metadata contained by the image data (e.g., by comparing location or other metadata with locations or other features related to stored criterions).

[0027] In some embodiments, stored criteria describing one or more scenes including machine-readable visual coding can be authenticated and / or registered. For example, an entity associated with a particular implementation of machine-readable visual coding may choose to register and / or authenticate stored criteria associated with it. In this way, other entities (e.g., malicious entities, false entities, conflicting entities, etc.) may be prohibited from generating, storing, or otherwise documenting stored criteria for the same scenes. In some embodiments, such prohibitions can ensure that certain entities cannot fraudulently falsify or otherwise spoof the authenticated context data of registered, stored criteria. Registration may involve performing a “sweep” (e.g., 180 or 360 degrees) using a camera to capture various visual characteristics of the scene.

[0028] In some embodiments, portions of image data describing one or more machine-readable visual codes and portions of image data describing contextual information may each be associated with one or more stored criteria by one or more recognition systems. In some embodiments, portions of image data describing contextual information may be processed by one or more recognition systems different from those processing portions of image data describing one or more machine-readable visual codes (although in some embodiments, the same recognition system may process both portions). Each association may involve determining similarity with different confidence levels. In some embodiments, a confidence level for a portion of image data describing one or more machine-readable visual codes that is similar to a feature of a stored criterion may be compared to a confidence level for a portion of image data describing contextual information that is similar to a feature of the same stored criterion (or, for example, a sub-criterion).

[0029] In some embodiments, the processing of machine-readable visual coding is determined to be successful (e.g., recognized, identified, etc.) if at least one of the confidence levels is greater than or equal to a predetermined confidence threshold and / or target value. In some embodiments, a composite confidence score may be required for the completion of processing (e.g., for verification). In some embodiments, a portion of the image data describing contextual information may be processed by a different recognition system as a fail-safe or alternative option in response to a determination that another portion of the image data describing one or more machine-readable visual codings has not been recognized with a sufficiently high confidence score.

[0030] Upon successful recognition, the recognition system may initiate actions. For example, actions may include verifying that image data corresponds to machine-readable visual coding, and verifying that image data corresponds to approved and / or machine-readable visual coding within a given context (e.g., within an appropriate location). In some embodiments, such verification may be received by the device that produced the image data (e.g., a captured image). In some embodiments, such verification may be received by one or more provider systems (e.g., a third-party system) related to machine-readable visual coding. Verification may include verification indicators that may include security certificates required to process data encoded in machine-readable visual coding. Other actions may include initiating a secure connection between the device and another system (e.g., for secure data exchange).

[0031] The systems and methods according to embodiments of this disclosure convey several technical effects and advantages. For example, processing machine-readable visual coding in light of contextual information, as disclosed herein, can provide improved robustness of the recognition process against noise, data loss, measurement errors, and / or defects. In some embodiments, the systems and methods according to embodiments of this disclosure can provide recognition and processing of machine-readable visual coding even when such coding cannot be clearly resolved by the imaging sensor of the imaging device, and can improve recognition capabilities beyond the limitations of the imaging device (e.g., insufficient sensor and / or optical resolution). Thus, in some cases, visual coding can be recognized / processed more efficiently because data of multiple types or modalities is used to recognize the visual coding, thereby reducing the number of images that need to be processed to recognize the coding. This more efficient processing can result in savings of computing resources, such as processor usage, memory usage, and bandwidth usage.

[0032] Additional technical advantages arising from the improved recognition techniques according to the embodiments of this disclosure include enabling the creation of encodings in a smaller size and in a visual configuration that facilitates integration into various implementations. In this way, less material and effort will be consumed in the implementation of machine-readable visual encoding. Lower barriers to implementation will also enable broader adoption and lead to increased efficiency in data communication, such as by compactly encoding data within visual patterns, in order to reduce data transmission costs.

[0033] Additional technical advantages include the ability to communicate large amounts of data using a given machine-readable visual code. For example, some machine-readable visual codes may be generated to correspond to standards that provide a given number of visual "bits" to encode data of a given size (e.g., a print area, a display area, etc.). In some examples, at least some of the visual "bits" may be used for error correction and / or alignment for processing the code. Advantageously, systems and methods according to embodiments of the present disclosure may provide improved error correction, alignment, and / or data communication without expending additional visual "bits," and in some embodiments, may maintain compatibility with code recognition systems that do not process contextual data. Some embodiments may, in some cases, provide improved analysis by using contextual information to distinguish equivalent machine-readable visual codes, allowing a single machine-readable visual code to be deployed in multiple contexts toward lower manufacturing costs (e.g., through economies of scale in display, printing, distribution, etc.) while retaining the ability for fine-grained record keeping.

[0034] Additional technical benefits include improved security for processing machine-readable visual codes. In some embodiments, a malicious party might attempt to modify and / or replace one or more machine-readable visual codes in order to assert control over any device processing machine-readable visual codes. Exemplary embodiments may prevent the success of such an attack (or misuse or misplacement of legitimate machine-readable visual codes) by comparing one or more features of a machine-readable visual code and / or one or more contextual features thereof with a stored standard and exposing inconsistencies in the attacker's machine-readable visual code. In this way, embodiments can also reduce attempts to commit fraud via machine-readable visual codes. In some embodiments, attempts to deceive a user by using a device to process machine-readable visual codes by modifying the codes can be reduced. Similarly, attempts to deceive a service provider (e.g., an entity associated with generating machine-learned visual codes) can be reduced by ensuring that only accurate and authentic machine-readable visual codes are processed by the user device.

[0035] Additional technical advantages include improved control over data communicated by or in accordance with the processing of machine-readable visual coding. For example, the processing of machine-readable visual coding may be restricted in light of its context, such that certain contextual conditions are required to perform actions related to machine-readable visual coding. Such control may be exercised retrospectively, after the machine-readable visual coding has been generated, displayed (e.g., printed), and / or distributed, enabling fine-grained control while reducing customization costs (e.g., printing consumables, individualized distribution costs, coding dedicated to unique identification, and / or the size of visual "bits").

[0036] The exemplary embodiments of this disclosure will now be discussed in more detail with reference to the drawings.

[0037] Exemplary devices and systems Figures 1 to 9B illustrate exemplary embodiments of machine-readable visual coding according to aspects of this disclosure. Exemplary figures conforming to specific geometries, layouts, and configurations are provided herein, but it should be understood that the examples provided herein are for illustrative purposes only. Adaptations and modifications of the exemplary examples provided herein are limited to the scope of this disclosure.

[0038] Figure 1 shows an exemplary embodiment of machine-readable visual coding including the shape of a glyph. The machine-readable visual coding 100 in Figure 1 includes the glyph 102 of the letter "G" filled with a visual pattern 104. In some embodiments, the visual pattern 104 includes a shape 106 against a background, and the visual contrast between the shape 106 and the background may represent a "bit" value of information. In some embodiments, the geometric relationships of the shapes 106 within the visual pattern (e.g., the relative size of the shapes 106, the distance between the shapes 106, the contours of the filled and / or unfilled areas, etc.) may be used to encode information. For example, a recognition system may be trained to recognize distinct features of the visual pattern 104 and associate the image data describing the machine-readable visual coding 100 with corresponding stored criteria. While Figure 1 shows a visual pattern 104 consisting of individual shapes, it should be understood that a visual pattern may include a continuous set of features that encode information. For example, an alternative machine-readable visual coding 200 may comprise a different glyph and / or another visual pattern. For example, alternative visual patterns may include waveforms that individually possess information-coding features such as amplitude, frequency, phase, and line thickness. Figure 1 shows individual glyphs, but it should be understood that two or more glyphs may be combined (for example, to form a word, logo, etc.). In this way, machine-readable visual coding can be integrated into, or form part of, a desired aesthetic value using individual and / or continuous visual features.

[0039] Figures 2 and 3 illustrate exemplary embodiments of machine-readable visual coding, which comprises a visual pattern combined with other visual information. In Figure 2, which shows exemplary coding 300, the visual pattern 304 surrounds a central space 308 containing text and / or graphical information 310 (e.g., a brand name, descriptor, etc.). In some embodiments, the central space 308 may not contain coded data, but the content of the central space 308 (e.g., text information 310) may provide context for processing the visual pattern 304. For example, image data describing machine-readable visual coding 300 may include information describing the central space 308 as contextual data, and a semantic recognition system may parse the text information 310 to assist and / or complement the recognition of the visual pattern 304. In Figure 3, exemplary coding 400 comprises a visual pattern 404 surrounding a central space 408 containing a photograph of a person or other avatar 410. As will be discussed with reference to text information 310, photograph 410 may provide context to assist and / or complement the recognition or subsequent processing of the visual pattern 404. For example, the visual pattern 404 may communicate instructions to perform an action, such as communicating encrypted data to a user account, and photograph 410 may correspond to a user associated with the user account. In some embodiments, for example, a different destination would no longer be associated with the account associated with photograph 410, and the mismatch would be detectable, so that a malicious party would not be able to redirect the encrypted data stream by simply changing the visual pattern 404 to specify a different destination.

[0040] While visual patterns 304 and 404 include several composed circular shapes, it should be understood that any number of visual patterns can be used to encode data, leveraging multiple shapes, geometric configurations, hues, and / or color values. In addition, while some variants can be encoded as visual "bits" (e.g., light / dark as 0 / 1), each variant can optionally encode more than one bit, and some variants (e.g., different shapes and / or sizes) can correspond to a given value or data object.

[0041] Figure 4 shows one exemplary embodiment of an encoding 500 comprising a visual pattern 504 of shapes 506a, 506b. Shapes 506a, 506b may encode information based on value (e.g., a range from white to black in a grayscale embodiment, and optionally, a range binned to one or more predetermined values ​​such as white, light blue, dark blue, etc.), orientation (e.g., positive radial direction, such as shape 506a, or negative radial direction, such as shape 506b), and / or location (e.g., based on the quadrants of the encoding 500). Another exemplary embodiment in Figure 5 shows an exemplary encoding 600 comprising a visual pattern 604 of multiple shapes 606 of different sizes, values, and locations. As demonstrated in Figure 6, an exemplary encoding 700 may comprise a visual pattern 704 containing multiple shapes 706, each of which may be associated with a value, location, orientation, and size; therefore, the visual pattern is not limited to a variation of any single shape. Another exemplary encoding may employ a visual pattern having a continuous shape configured in the form of concentric rings. Groups of positions may also correspond to certain data classes, so that the shape, value, position, orientation, and size may correspond to encoded data. For example, one or more rings of the encoding may be assigned to encode a particular class of data (e.g., an identification number), and the length and / or value of each arc segment in the ring may represent that value.

[0042] The aforementioned features can be combined and / or reconfigured to produce configurable machine-readable visual codes, as shown in Figures 7 and 8. For example, the exemplary code 900 in Figure 7 comprises both concentric rings and radial directional line segments, as well as dispersed circular shapes (for example, for registering orientation in an image). The exemplary code 1000 in Figure 8 comprises concentric rings 1006a, shapes 1006b, and radially dispersed trajectories 1006c. It should also be understood that the central areas 308 and 408 in Figures 2 and 3 can similarly be implemented in any of the codes from Figures 1 to 8.

[0043] As shown in Figures 1 to 8, the machine-readable visual coding provided in this disclosure can be integrated into and form part of substantially any desired aesthetic value by using individual and / or consecutive features to encode data. In addition, the configurability of the machine-readable visual coding provided in this disclosure may provide an easily recognizable class of machine-readable visual coding (by human perception and / or machine perception).

[0044] For example, differences in contours, layouts, and shapes used in patterns, colors, etc., such as the differences in each of Figures 1 through 8, may correspond to different algorithms and / or standards used to generate machine-readable visual codes. Thus, the aforementioned differences may indicate to one or more recognition systems the type of recognition model required to read the machine-readable visual codes according to different algorithms and / or standards used to generate the machine-readable visual codes. In this way, systems and methods according to embodiments of the present disclosure may provide efficient processing by deploying specially trained recognition models. However, in some embodiments, the aforementioned easily perceptible structural differences may also be advantageously employed to help a general coding recognition model immediately and accurately recognize different classes of coding and minimize errors due to classification confusion.

[0045] Figure 9A shows another exemplary embodiment of the encoding 1100, which includes a visual pattern 1104 that employs circular shapes 1106a, 1106b to encode information, with several shapes 1106c additionally used for orientation registration. In some embodiments, a user device may capture an image of the encoding 1100 and display a rendering 1150 of the image that is manipulated to indicate an indicator of the processing progress of the encoding 1100. For example, as shown in Figure 9B, the rendering 1150 may proportionally display portions of the full-color encoding 1100 and "grayed-out" portions as processing progresses, such that the angle 1152 sweeps from 0 to 360 degrees during processing. In some embodiments, the progress rendering 1150 may be notified by actual status updates or estimations thereof. In some embodiments, the rendering is an augmented reality rendering such that a progress overlay is rendered to appear in real time on top of the image of the encoding 1100 (for example, within an open app that displays the field of view of a camera on a user device). It should be understood that any other encoding of any of the exemplary encodings in Figures 1 to 8 may be used to display the progress rendering. Such progress renderings can sweep across the coding in an angular manner, a linear manner, or by tracing one or more contours of the coding (e.g., the glyphs in Figure 1).

[0046] Figures 10A–10C demonstrate one embodiment of the present disclosure for recognizing machine-readable visual codes in light of their context. Figure 10A shows an image 1200 of a map 1202 on which machine-readable visual codes 1204 are printed. Adjacent to the machine-readable visual codes 1204 is text material 1206. While Figure 10A is discussed herein in terms of a “printed” “map,” it should be understood that the machine-readable visual codes 1204 can be displayed on a screen or other device as part of a larger display 1202 in substantially any form suitable for capture within the image 1200.

[0047] In some embodiments, the imaging device may be able to decompose the machine-readable visual code 1204 sufficiently to directly decode its content. In some embodiments, prior to decoding the content and / or executing the code communicated by that content, additional contextual information related to the machine-readable visual code 1204 (e.g., within image 1200) may be processed by one or more recognition systems. For example, map 1202 may be processed to recognize its identification as a map and / or one or more locations mapped thereon. In some embodiments, text information 1206 may be processed (e.g., via OCR and / or semantic analysis) to recognize an association between the machine-readable visual code 1204 and “station”. Stored criteria related to the machine-readable visual code 1204 may include information that associates the machine-readable visual code 1204 with the map, “station”, and / or one or more locations on the map. By comparing the contextual information within image 1200 with the stored criteria, the confidence level related to the recognition and / or verification of the machine-readable visual code 1204 may be increased. In this way, if the imaging device is unable or does not sufficiently decompose the machine-readable visual code 1204 to directly decode its content (for example, due to insufficient lighting, a slow shutter speed, etc.), the increased level of confidence obtained by recognizing contextual information may be "satisfied" from a criterion in which any missing information is stored.

[0048] Figure 10B shows image 1210 of the same map 1202 captured from a greater distance. In some embodiments, the machine-readable visual code 1204 cannot be decomposed to a sufficient level of detail to be directly recognized. However, contextual information contained within image 1210 can be recognized, such as the relative positioning of the machine-readable visual code 1204 in map 1202, the presence of text labels 1212a, 1212b, additional machine-readable visual codes 1214, and large mapping features 1216a, 1216b. Each, some, or all of these contextual features can be recognized and compared with contextual data in a stored reference that relates to the machine-readable visual code 1204. In this way, even if the imaging device cannot or does not decompose the machine-readable visual code 1204 sufficiently to directly decode its content, the recognition of contextual information and the association of the contextual information with a stored reference can enable retrieval of the data encoded by the code 1204 from the stored reference.

[0049] Figure 10C shows an image 1220 of a scene containing a map 1202 on which a machine-readable visual code 1204 is printed. Examples of recognizable contextual features in the scene include nearby objects such as a large logo 1222, a frame 1224 in which the map 1202 is mounted, benches 1226a and 1226b, contrasting architectural features such as the joint 1228 of the wall panel behind the map 1202, the boundary of the room such as the boundary between the wall and the ceiling 1230, and lighting elements such as the light fixture 1232. Each, some, or all of these contextual features can be recognized and compared with contextual data in a stored reference that relates to the machine-readable visual code 1204. In this way, even if the imaging device cannot or does not resolve the machine-readable visual code 1204 at all (or is far from its contours and / or edges), the recognition of contextual information and the association of contextual information with a stored reference can enable retrieval of the data encoded by the code 1204 from the stored reference.

[0050] The recognition of contextual features in 3D space, such as the room boundaries, objects, and architectural features described above, may involve processing image 1220 using a visual positioning system (VPS). For example, in some embodiments, image 1220 may be processed using a VPS model to detect surfaces, edges, corners, and / or other features and generate a feature mapping (e.g., a point cloud as shown in Figure 10D). Such a mapping may be packaged into image data describing the scene and / or generated by one or more recognition systems, and the mapping may be compared to data describing the mapping in a stored reference. In some embodiments, the mapping (e.g., a point cloud and / or a set of anchors) may be generated from image 1220 and / or associated image data (e.g., collected sensor data describing the scene, including spatial measurements from LiDAR). As an addition or alternative, spectral data (e.g., from the invisible portion of the electromagnetic spectrum) may be collected and used to generate the mapping. In some embodiments, multiple exposures can be combined (e.g., from multiple points in time, from multiple sensors on the device, etc.) to determine the deep mapping (e.g., by triangulation). In this way, the VPS can be used to associate the image 1220 (and / or image data describing that image) with a stored criterion for recognition, verification, and / or processing of the machine-readable visual code 1204.

[0051] While the exemplary embodiments shown in Figures 10A-10D relate to information display (maps), it should be understood that the systems and methods of this disclosure process image data describing a wide range of objects with machine-readable visual coding, such as posters, advertisements, billboards, publications, flyers, vehicles, personal mobility vehicles (e.g., scooters, motorcycles, etc.), rooms, walls, furniture, buildings, streets, signs, tags, pet collars, printed clothing, medical bracelets, etc. For example, some embodiments of the systems and methods according to the aspects of this disclosure may include object recognition (e.g., product recognition) for information retrieval from a web server. For example, image data describing a scene containing objects with machine-readable visual coding may be captured. The image data may include contextual information surrounding the machine-readable visual coding (e.g., other features of the object, such as the image and / or text on the image), and contextual information surrounding the object (e.g., where the object is located, nearby objects, etc.). In some embodiments, one or more recognition systems may process the machine-readable visual coding, and one or more other recognition systems may process the contextual data. For example, an image recognition system may recognize and / or identify the type of an object and / or the type of objects surrounding it. In some examples, a semantic recognition system may process one or more semantic entities recognized in a scene (e.g., text and / or images on an object) to assist in the processing of machine-readable visual coding. Depending on the identification of an object, information describing that object may be retrieved from a web server (e.g., a web server associated with the recognition system) and rendered on the display device of a user device (e.g., a user device that captured image data).

[0052] Figure 11A shows another exemplary embodiment of image 1300 of a scene including machine-readable visual coding 1304. In the shown scene, the storefront displays machine-readable visual coding 1304 on its exterior wall. As discussed above, various features shown in the scene may be recognized and / or processed for comparison with stored criteria in order to assist and / or complement the recognition and / or verification of machine-readable visual coding 1304. For example, external architectural features such as door frames 1310, window frames 1312, windowsills 1314, and awnings 1316 may be recognized. In addition, internal architectural features such as lighting features 1318 may also be recognized if visible.

[0053] One or more of the above features can be recognized using VPS. For example, a feature map 1320 may be generated as shown in Figure 11B. In some examples, anchor points 1322 may be located and optionally connected to segments 1324 to mark and / or trace the features of interest in the scene. In some embodiments, the feature map 1320 may be compared to a stored reference feature map to determine its association (e.g., similarity).

[0054] In some embodiments, a feature map 1320 is generated and can be stored in a stored reference based on one or more captured images of the scene. For example, in one embodiment shown in Figure 11C, a user 1324 associated with a user account corresponding to a stored reference can use an imaging device (e.g., a smartphone 1326) to capture multiple images of the scene from multiple bandage points along a path 1328. In this way, the image data captured from multiple angles can be processed to generate a robust feature map that describes the scene to be stored in a stored reference. Once documented, the feature map can then be used to compare with other feature maps from one or more recognition systems (e.g., a feature map generated from only one bandage point, in some cases).

[0055] The stored criteria may optionally include additional contextual information about objects and structures surrounding the scene. For example, as shown in Figure 11D, user 1324 may generate panoramic image data by using the imaging device 1326 to collect images from multiple bandage points along a path 1330 (for example, by user 1324 rotating). In this way, feature maps may be generated not only for the scene shown in image 1300 in Figure 11A, but also (if desired) for the surrounding objects and structures. The stored criteria associated with machine-readable visual coding 1304 may include all feature maps associated with that machine-readable visual coding 1304 and may provide a robust set of contextual data for use by one or more recognition systems according to embodiments of the present disclosure.

[0056] In some embodiments, the machine-readable visual coding 1304 may be desired to be used in one of several locations associated with the user account of user 1324. In some embodiments, the stored criteria may include contextual data associated with each of several locations, each associated with a subcrib.

[0057] In some embodiments, the context associated with one machine-readable visual code may include another machine-readable visual code. For example, in some embodiments, the processing of data communicated in one machine-readable visual code (e.g., the execution of an instruction communicated therein) may be configured to be conditional on the recognition of another machine-readable visual code within the image data.

[0058] For example, in one embodiment shown in Figure 12, a system 1400 for secure package delivery may be implemented according to an aspect of this disclosure. For example, a delivery person 1402 may be provided with a first machine-readable visual code 1403 (e.g., on a badge, on a device), and a package 1404 may be provided with a second machine-readable visual code 1405 (e.g., printed thereon). A “smart” home system 1420 may include a camera 1422 (e.g., a doorbell camera, a security camera) capable of capturing a scene that includes both the first machine-readable visual code 1403 and the second machine-readable visual code 1405 within its field of view. In some embodiments, both the first machine-readable visual code 1403 and the second machine-readable visual code 1405 may be required to match a given level of trust before a “smart” lock 1424 may allow access to a secure area for storing the package 1404 (e.g., a package locker, a garage area, a front door, a doorway, etc.). In some embodiments, an additional “smart” component 1426 may provide additional contextual information (e.g., delivery time, estimated delivery time, estimated package dimensions, estimated number of packages, any overrides that may prevent access permissions, weather conditions that may affect image distortion, 3D imaging data to prevent spoofing using printed images of packages). The home system 1420 may optionally be connected to the network 1430 to communicate image data (e.g., to process recognition tasks on the server 1440) and / or to receive reference information for local processing of one or more recognition tasks (e.g., to receive expected parameters of machine-readable visual coding 1403 and / or facial recognition data to verify the identification of the delivery person 1402). In some embodiments, the first machine-readable visual coding 1403 may be generated to provide restricted access only across a time window. For example, data describing the machine-readable visual coding 1403 (including any associated time window) may be stored in a stored reference on the server 1440 and optionally transmitted to the home system 1420.In this way, in some embodiments, access may be granted to verified individuals only when they possess the verified package and when they access it at a specified time.

[0059] Figure 13A shows a block diagram of an exemplary computing system 1500 for recognizing and / or processing machine-readable visual coding according to an exemplary embodiment of the present disclosure. The system 1500 includes a user computing device 1502, a server computing system 1530, and a provider computing system 1560, which are communicably coupled over a network 1580.

[0060] The user computing device 1502 may be any type of computing device, such as a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device (e.g., computer-enabled glasses, a watch, etc.), an embedded computing device, or any other type of computing device.

[0061] The user computing device 1502 includes one or more processors 1512 and memory 1514. The one or more processors 1512 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors operably connected. The memory 1514 may include one or more non-temporary computer-readable storage media, such as RAM, SRAM, DRAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 1514 can store data 1516 and instructions 1518 executed by the processors 1512 to cause the user computing device 1502 to perform operations.

[0062] The user computing device 1502 may include one or more sensors 1520. For example, the user computing device 1502 may include a user input component 1521 that receives user input. For example, the user input component 1521 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may be useful for implementing a virtual keyboard. Other exemplary user input components include a microphone, a conventional keyboard, or other means by which the user can provide user input. The user computing device 1502 may include one or more imaging sensors 1522 (e.g., CCD, CMOS, RADAR, LIDAR, etc.). Each of the imaging sensors 1522 may be the same or different, and each may consist of one or more different lens configurations. One or more imaging sensors may be located on one side of the user computing device 1502, and one or more imaging sensors may be located on the opposite side of the user computing device 1502. The imaging sensor may include a sensor that captures a wide range of visible and / or invisible light spectrum. The user computing device 1502 may include one or more geospatial sensors 1523 (e.g., GPS) for measuring, recording, and / or interpolating location data. The user computing device 1502 may also include one or more transform sensors 1524 (e.g., accelerometers, etc.) and one or more rotation sensors 1525 (e.g., inclinometers, gyroscopes, etc.). In some embodiments, one or more geospatial sensors 1523, one or more transform sensors 1524, and one or more rotation sensors 1525 may cooperate to determine and record the position and / or orientation of the device 1502 and, in combination with one or more imaging sensors 1522, determine the pose of the device 1502 relative to the imaged scene.

[0063] The user device 1502 may include image data 1527 collected and / or generated using any or all of the sensors 1520. The user device 1502 may include one or more recognition models 1528 for performing a recognition task on the collected image data 1527. For example, the recognition models 1528 may be a variety of machine-learned models, such as neural networks (e.g., deep neural networks) or other types of machine-learned models including nonlinear and / or linear models, or otherwise may include such machine-learned models. The neural network may include a feedforward neural network, a recurrent neural network (e.g., a long-short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks.

[0064] In some implementations, one or more recognition models 1528 may be received from a server computing system 1530 via a network 1580, stored in user computing device memory 1514, and then used or otherwise implemented by one or more processors 1512. In some implementations, a user computing device 1502 may implement multiple parallel instances of a single recognition model 1528 (for example, to perform parallel recognition tasks, such as performing a recognition task on a portion of image data 1527 describing machine-readable visual coding and a recognition task on a portion of image data 1527 describing context data).

[0065] As an addition or alternative, one or more recognition models 1540 may be included within a server computing system 1530 that communicates with a user computing device 1502 according to a client-server relationship, or otherwise stored and implemented by the server computing system 1530. The server computing system 1530 includes one or more processors 1532 and memory 1534. One or more processors 1532 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors operably connected. Memory 1534 may include one or more non-temporary computer-readable storage media, such as RAM, SRAM, DRAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. Memory 1534 may store data 1536 and instructions 1538 executed by the processors 1532 to cause the server computing system 1530 to perform operations. In some implementations, the server computing system 1530 includes one or more server computing devices, or is otherwise implemented by server computing devices. In cases where the server computing system 1530 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or any combination thereof.

[0066] In some examples, the recognition model 1540 may be implemented by a server computing system 1530 as part of a web service (e.g., a machine-readable visual coding recognition and / or verification service). Thus, one or more recognition models 1528 may be stored and implemented on a user computing device 1502, and / or one or more recognition models 1540 may be stored and implemented on a server computing system 1530. For example, one or more recognition tasks may be shared and / or distributed among one or more recognition models 1528 and one or more recognition models 1540.

[0067] In some examples, a user computing device 1502 transmits image data 1527 to a server computing system 1530 via a network 1580 for recognition and / or processing of any machine-readable visual coding described within the image data 1527. The server computing system 1530 may process the image data 1527 using one or more recognition models 1540 to associate the image data 1527 with a stored reference 1542.

[0068] In some embodiments, the stored criterion 1542 may correspond to encoded data 1544 that can assist in decoding the machine-readable visual coding in the image data 1527 (for example, by indicating one or more specific encoding recognition models to be used, by detailing the algorithms used to generate / interpret the machine-readable visual coding, such as one or more of the encoding schemes demonstrated in Figures 1 to 11A). In some embodiments, the stored criterion may include general contextual data 1546 relating to the entire stored criterion 1542. For example, the stored criterion 1542 may relate to multiple machine-readable visual codings. In one exemplary embodiment, each coding may be installed on an instance of a particular type of object or structure (e.g., a device, map, advertisement, scooter, car, wall, building, etc.). Each instance may share some amount of contextual information such that the scene containing each encoding shares at least some overlapping context (for example, the image of each encoding may show at least a portion of a scooter, a map of a particular location, a particular advertisement, a logo or photograph incorporated within a machine-readable visual encoding as described herein, etc.). General contextual data 1546 may include information describing the context shared between encodings.

[0069] In some embodiments, the stored criterion 1542 may include context data 1550 associated with a first subcriterion 1548, and optionally, context data 1554 associated with a second subcriterion 1552. Subcriteria 1548, 1552 may, in some implementations, be used to classify context information associated with a subset of multiple codings associated with the stored criterion 1542. For example, to continue using the language of the embodiments described above, a subset of multiple codings corresponding to different subsets of instances of a particular type of object or structure may be associated with one or more subcriteria. For example, each instance of an object or structure (e.g., each scooter, each restaurant location, etc.) may be associated with its own context (e.g., location information, appearance, etc.), which may be stored in context data 1550, 1554, respectively. In this way, machine-readable visual coding may be associated with the stored criterion 1542 (e.g., corresponding to entities, user accounts, projects, categories, etc.) and subcriteria 1548, 1552 (e.g., corresponding to specific implementations, application examples, etc.).

[0070] Once associated, the server computing system 1530 may begin operating in accordance with the operation instruction 1556. In some embodiments, each stored reference 1542 and / or subreference 1548, 1552 may correspond to the same or different operation instruction 1556. In some embodiments, the operation instruction 1556 includes verifying a machine-readable visual code described by image data 1527. Verification may include, for example, sending a verification indicator to a user computing device 1502, in which case the user computing device 1502 may perform additional operations (e.g., processing any data items encoded in the machine-readable visual code). For example, the verification indicator may include a security proof required to process the data encoded in the machine-readable visual code. The security proof may, in some embodiments, be required by the user computing device 1502 (e.g., by an application stored on and / or running on the user computing device 1502, by a web server via a browser interface running on the user computing device 1502, etc.). For example, in the system 1400 for secure delivery of packages shown in Figure 12, the “smart” home system 1420 may, in some embodiments, require such security proof before the “smart” lock 1424 allows access to the secure area to store the package 1404. Referring again to Figure 13A, in some embodiments, security proof is required by the provider computing system 1560 to access services provided by the provider computing system 1560 (for example, to initiate a secure connection with the provider computing system 1560).

[0071] In some embodiments, the user device 1502 may be configured to process only machine-readable visual codes that are validated by the server computing system. In some embodiments, when the user device 1502 captures one or more machine-readable visual codes that are not associated with any stored criteria, or are partially associated with stored criteria, but whose context conflicts with context data 1546, 1550, 1554 stored within the stored criteria, the user device 1502 may be configured to display a warning message indicating this. For example, the user of the user device 1502 may be prompted to accept and / or reject a message indicating the lack of validation of the machine-readable visual code (e.g., a message rendered on the user device's display) before proceeding to process the data encoded by the machine-readable visual code.

[0072] In some embodiments, verification may include transmitting a verification indicator to a provider computing system 1560 via the network 1580. The provider computing system 1560 may be associated with one or more machine-readable visual codes (which may be described, for example, within image data 1527). For example, one or more machine-readable visual codes may be generated to communicate data (e.g., information, instructions) to interact with services of entities associated with the provider computing system 1560.

[0073] The provider computing system 1560 includes one or more processors 1562 and memory 1564. The one or more processors 1562 may be any suitable processing device (e.g., a processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and may be one processor or multiple processors operably connected. The memory 1564 may include one or more non-temporary computer-readable storage media, such as RAM, SRAM, DRAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 1564 can store data 1566 and instructions 1568 executed by the processors 1562 to cause the provider computing system 1560 to perform operations. In some implementations, the provider computing system 1560 includes one or more server computing devices, or is otherwise implemented by server computing devices. In some embodiments, the provider computing system 1560 comprises a server computing system 1530.

[0074] The provider computing system 1560 may include service instructions 1570 for providing services, such as providing services to a user of a user computing device 1502. In some embodiments, the provider computing system 1560 initiates provisioning a service in response to an indicator that the user computing device 1502 has captured and / or otherwise processed verified machine-readable visual coding. For example, the provider computing system 1560 may provide a service related to a particular instance of image data 1527 based on an indicator from the server computing system 1530 that the image data includes contextual data related to a stored criterion 1542 or sub-criteria 1548, 1552 related to a particular instance of machine-readable visual coding. In this way, the provider computing system 1560 may, for example, adjust or otherwise improve the services provided.

[0075] In some embodiments, the service instruction 1570 may include executable instructions for running an application on the user device 1502. In some embodiments, the user of the user device 1502 may scan machine-readable visual coding according to the embodiments of this disclosure to download or otherwise obtain access to executable code for running the application (or processing within the application).

[0076] In some embodiments, context data 1572 may be generated on and / or stored in the provider computing system 1560 and / or sent to the server computing system 1530. In some embodiments, context data 1572 may be used to update context data 1546, 1550, and 1554. In some embodiments, context data 1572 may be used to verify the authenticity of image data 1527. For example, the provider computing system 1560 may be configured to provide services related to a particular instance of machine-readable visual coding, and a particular instance of machine-readable visual coding may correspond to context data 1572. In some embodiments, context data 1572 may be generated to ensure that image data 1527 can be processed to initiate a unique transaction. For example, in one embodiment, a set of context data 1572 may be generated that must be described by image data 1527 for verification by the server computing system 1530. For example, the provider computing system 1560 may generate a visual pattern or other context perceptible by one or more sensors 1520 that will be displayed near or with the target machine-readable visual code (e.g., complementary machine-readable visual code). In such an example, verification by the server computing system 1530 may be conditional on both the target machine-readable visual code and the complementary machine-readable visual code described by the image data 1527. In one example, the complementary machine-readable visual code may be used in one embodiment of a system 1400 for secure package delivery, as shown in Figure 12.

[0077] While the examples described above refer to the recognition and / or processing of arbitrary machine-readable visual coding described in image data 1527 by the server computing system 1530, it should be understood that the recognition task can be distributed between the user computing device 1502 and the server computing system 1530. For example, the user computing device 1502 may use one or more recognition models 1528 to process one or more portions of image data 1527 and send the same and / or different portions of image data 1527 to the server computing system 1530 for processing by the server computing system 1530 using one or more recognition models 1540. In some embodiments, one or more recognition models 1528 may include coding recognition models for recognizing arbitrary machine-readable visual coding described by image data 1527. In some embodiments, based on the processing of image data 1527 by the user computing device 1502 using coding recognition models, the user computing device 1502 may communicate the image data 1527 to the server computing system 1530 via the network 1580 for additional processing.

[0078] For example, in one instance, user computing device 1502 may determine that one or more machine-readable visual codes are described by image data 1527. User computing device 1502 may process image data 1527 using a coding recognition model to decode the machine-readable visual codes. User computing device 1502 may send data describing a machine-readable visual code (e.g., data decoded from image data 1527 and / or the code itself) to server computing device 1530 to determine whether there is a stored criterion associated with (or potentially associated with) the machine-readable visual code. This association may be made by server computing system 1530 by processing the data sent to server computing system 1530, which describes a machine-readable visual code (e.g., data decoded from image data 1527 and / or the code itself). If stored criteria 1542 related to machine-readable visual coding exist, the server computing system may transmit the stored criteria 1542 and / or related context data to the user computing device 1502 for comparison with image data 1527 (e.g., general context data 1546 and / or context data 1550, 1554 specific to one or more sub-criteria 1548, 1552). In such embodiments, the context data collected by the user computing device 1502 may remain on the device (e.g., to achieve additional privacy, low latency, etc.).

[0079] In some examples, one or more recognition models 1528 and one or more recognition models 1540 may be, or otherwise include, various machine learning models. Exemplary machine learning models include neural networks or other multilayer nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.

[0080] A machine learning model can be trained using various training or learning techniques, such as backpropagation. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (for example, based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update parameters over several training iterations. In some implementations, performing backpropagation may include performing abbreviated temporal backpropagation. A model trainer can perform several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained. In particular, a model trainer can train a model based on a set of training data. Training data may include, for example, machine-readable visual coding, image data describing it, and image data describing it.

[0081] In some implementations, if the user gives consent, training examples may be provided by the user computing device 1502. Thus, in such implementations, the model 1528 provided to the user computing device 1502 can be trained by the training computing system on user-specific data received from the user computing device 1502. In some cases, this process is referred to as model personalization.

[0082] A model trainer may contain computer logic used to provide desired functionality. A model trainer can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, a model trainer includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, a model trainer includes one or more sets of computer executable instructions stored in a tangible computer-readable storage medium such as a RAM hard disk or optical or magnetic media.

[0083] In some implementations, the input to the machine learning-trained recognition model of this disclosure (e.g., input, data, and / or training examples) (included in any one of the following, e.g., user computing device 1502, server computing system 1530, etc.) may be image data. The machine learning-trained model may process the image data to produce outputs. For example, the machine learning-trained model may process the image data to produce an image recognition output (e.g., recognition of image data, latent embedding of image data, encoded representation of image data, hash of image data, etc.). For another example, the machine learning-trained model may process the image data to produce an image segmentation output. For yet another example, the machine learning-trained model may process the image data to produce an image classification output. For yet another example, the machine learning-trained model may process the image data to produce an image data modification output (e.g., modification of image data, etc.). For yet another example, the machine learning-trained model may process the image data to produce an encoded image data output (e.g., encoded and / or compressed representation of image data, etc.). For yet another example, the machine learning-trained model may process the image data to produce an upscaled image data output. As another example, a machine learning model can process image data and generate predictive outputs.

[0084] In some embodiments, the input to the machine learning-trained recognition model of this disclosure (e.g., input, data, and / or training examples), which is included in any one of the following: a user computing device 1502, a server computing system 1530, etc., may be latent coded data (e.g., a latent spatial representation of the input). The machine learning-trained model may process the latent coded data to produce an output. As an example, the machine learning-trained model may process the latent coded data to produce a recognition output. As another example, the machine learning-trained model may process the latent coded data to produce a reconstruction output. As yet another example, the machine learning-trained model may process the latent coded data to produce a search output. As yet another example, the machine learning-trained model may process the latent coded data to produce a reclustering output. As yet another example, the machine learning-trained model may process the latent coded data to produce a prediction output.

[0085] In some implementations, the inputs (e.g., inputs, data, and / or training examples) to the machine learning-trained recognition model of this disclosure (e.g., contained in any one of the following: user computing device 1502, server computing system 1530, etc.) may be statistical data. The machine learning-trained model may process the statistical data to produce outputs. As an example, the machine learning-trained model may process the statistical data to produce a recognition output. As another example, the machine learning-trained model may process the statistical data to produce a prediction output. As yet another example, the machine learning-trained model may process the statistical data to produce a classification output. As yet another example, the machine learning-trained model may process the statistical data to produce a segmentation output. As yet another example, the machine learning-trained model may process the statistical data to produce a segmentation output. As yet another example, the machine learning-trained model may process the statistical data to produce a visualization output. As yet another example, the machine learning-trained model may process the statistical data to produce a diagnostic output.

[0086] In some embodiments, the input to the machine learning-trained recognition model of this disclosure (e.g., input, data, and / or training example) (included in any one of the following: user computing device 1502, server computing system 1530, etc.) may be sensor data. The machine learning-trained model may process the sensor data to generate outputs. As an example, the machine learning-trained model may process the sensor data to generate a recognition output. As another example, the machine learning-trained model may process the sensor data to generate a prediction output. As yet another example, the machine learning-trained model may process the sensor data to generate a classification output. As yet another example, the machine learning-trained model may process the sensor data to generate a segmentation output. As yet another example, the machine learning-trained model may process the sensor data to generate a segmentation output. As yet another example, the machine learning-trained model may process the sensor data to generate a visualization output. As yet another example, the machine learning-trained model may process the sensor data to generate a diagnostic output. As yet another example, the machine learning-trained model may process the sensor data to generate a detection output.

[0087] Network 1580 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or any combination thereof, and may include any number of wired or wireless links. Generally, communication over Network 1580 can be carried over any type of wired and / or wireless connection using a wide range of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formatting (e.g., HTML, XML), and / or protection methods (e.g., VPN, Secure HTTP, SSL).

[0088] Figure 13A shows one exemplary computing system that may be used to implement the present disclosure. Other computing systems may be used similarly. For example, in some implementations, the user computing device 1502 may include a model trainer and a training dataset. In such implementations, models 1528 and 1540 can be both trained and used locally on the user computing device 1502. In some such implementations, the user computing device 1502 may implement a model trainer based on user-specific data to personalize model 1528.

[0089] Figure 13B shows a block diagram of an exemplary computing device 1582 according to an exemplary embodiment of the present disclosure. The computing device 1582 may be a user computing device or a server computing device.

[0090] The computing device 1582 includes several applications (for example, applications 1 to N). Each application includes its own machine learning library and pre-trained models. For example, each application may include a pre-trained model. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and so on.

[0091] As shown in Figure 13B, each application can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0092] Figure 13C shows a block diagram of an exemplary computing device 1584 according to an exemplary embodiment of the present disclosure. The computing device 1584 may be a user computing device or a server computing device.

[0093] Computing device 1584 contains several applications (e.g., applications 1 through N). Each application communicates with a central intelligence layer. Exemplary applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, and a browser application. In some implementations, each application can communicate with the central intelligence layer (and the models stored within it) using an API (e.g., a common API across all applications).

[0094] The central intelligence layer contains several machine learning models. For example, as shown in Figure 13C, each machine learning model (e.g., Model) may be provided to each application and managed by the central intelligence layer. In other implementations, two or more applications may share a single machine learning model. For example, in some implementations, the central intelligence layer may provide a single model (e.g., SingleModel) to all applications. In some implementations, the central intelligence layer is contained within the operating system of the computing device 1584, or otherwise implemented by the operating system.

[0095] The central intelligence layer can communicate with the central device data layer. The central device data layer may be a centralized repository of data for the computing device 1584. As shown in Figure 13C, the central device data layer can communicate with several other components of the computing device, such as one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0096] Exemplary Recognition System Configuration Figure 13D shows a block diagram of an exemplary recognition system 1590 according to an exemplary embodiment of the present disclosure. In some implementations, the recognition system 1590 receives a set of image data 1527 describing a scene containing one or more machine-readable visual codes, and is trained to provide recognition data 1592 (e.g., confidence levels of the associations between them) describing associations between one or more of the machine-readable visual codes and one or more stored criteria as a result of receiving the image data 1527. In some implementations, the recognition system 1590 may include an encoding recognition model 1540a that is operable to process (e.g., decode) one or more machine-readable visual codes described within the image data 1527. The recognition system 1590 may also include an image recognition model 1540b for recognizing and / or otherwise processing the contextual image data 1527 (e.g., for performing deep mapping or other VPS techniques to recognize objects, people, and / or structures described within the image data 1527). Image data 1527 can flow directly into each of models 1540a and 1540b, and / or sequentially through them.

[0097] Figure 13E shows a block diagram of an exemplary recognition system 1594 according to an exemplary embodiment of the present disclosure. Recognition system 1594 is similar to recognition system 1590 in Figure 13D, except that it further includes a context component 1596 that can directly process context data stored in image data 1527 (e.g., sensor data such as metadata associated with one or more images of the image data).

[0098] In some embodiments, the recognition data 1592 may include a composite score describing the reliability associated with each of the recognition models 1540a and 1540b. In some implementations, the recognition data 1592 may include a sum (weighted or unweighted). In some embodiments, the recognition data 1592 may be determined based on the higher reliability associated with each of the recognition models 1540a and 1540b.

[0099] Estimative method Figure 14 shows a flowchart of exemplary method 1600 according to exemplary embodiments of the present disclosure. Figure 14 shows the steps performed in a specific order for illustrative and discussion purposes, but the methods of the present disclosure are not limited to the order or arrangement shown. Various steps of method 1600 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0100] In 1610, the computing system obtains image data describing a scene, including machine-readable visual coding.

[0101] In 1620, the computing system processes image data using a first recognition system configured to recognize machine-readable visual codes.

[0102] In 1630, the computing system processes image data using a second distinct recognition system configured to recognize the surrounding portion of the scene that encloses the machine-readable visual coding. In some embodiments, the image data includes an image related to the machine-readable visual coding and contained within the surrounding portion of the scene, including information displays and / or advertisements. In some embodiments, the second distinct recognition system comprises a visual positioning system configured to extract visual features of the surrounding portion of the scene. In some embodiments, the second distinct recognition system comprises a semantic recognition system configured to recognize semantic entities related to the machine-readable visual coding and referenced within the surrounding portion of the scene. In some embodiments, the second recognition system processes metadata (e.g., location data) related to the image data.

[0103] In 1640, the computing system identifies a stored criterion related to machine-readable visual coding, at least in part on one or more first outputs generated by a first recognition system based on image data, and at least in part on one or more second outputs generated by a second recognition system based on image data. In some embodiments, the identification is based on a composite output based on one or more first outputs and one or more second outputs. In some embodiments, at least one of the one or more first outputs may fail to satisfy a target value, and in response to the determination that at least one of the one or more first outputs fails to satisfy a target value, the computing system may generate one or more second outputs by the second recognition system.

[0104] In 1650, the computing system performs one or more actions in response to the identification of a stored criterion. In some embodiments, the computing system generates a verification indicator. In some embodiments, the verification indicator is configured to provide a security proof required to process data encoded by machine-readable visual coding. In some embodiments, the data encoded in machine-readable visual coding is associated with a request to obtain access to a secure area, the scene includes two or more machine-readable visual codes, and the security proof is required to initiate service for the request to obtain access to the secure area. In some embodiments, the request to obtain access to a secure area is associated with a package delivery entity for the delivery of a package to the secure area, the package includes at least one of two or more machine-readable visual codes.

[0105] Additional disclosure The technologies discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions performed and the information transmitted to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionalities among their components. For example, the processes discussed herein can be implemented using a single device or component, or multiple devices or components, operating in combination. Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0106] This subject matter has been described in detail with respect to various specific exemplary embodiments, each example given for illustrative purposes only and not as an limitation of this disclosure. Those skilled in the art, upon understanding the above, will readily be able to create modifications, variations, and equivalents to such embodiments. Therefore, this disclosure does not preclude the inclusion of such modifications, variations, and / or additions to the subject matter, as will be readily apparent to those skilled in the art. For example, features shown or described as part of one embodiment may also be used in conjunction with another embodiment to result in further embodiments. Thus, this disclosure is intended to cover such modifications, variations, and equivalents. [Explanation of Symbols]

[0107] 100 Machine-readable visual coding 102 glyphs 104 Visual Patterns 106 Shape 200 Machine-Readable Visual Coding 300 coding, machine-readable visual coding 304 Visual Patterns 308 Central Space 310 Text and / or graphical information 400 encoding 404 Visual Patterns 408 Central Space 410 A person's photograph or another person's avatar 500 encoding 504 Visual Patterns 506a shape 506b shape 600 encoding 604 Visual Patterns 606 Shape 700 encoding 704 Visual Patterns 706 Shape 900 encoding 1000 encoding 1006a Concentric ring 1006b shape 1006c Radially dispersed orbits 1100 encoding 1104 Visual Patterns 1106a Circular shape 1106b Circular shape 1106c shape 1150 rendering, progress rendering 1152 angle 1200 images 1202 Map, Display 1204 Machine-readable visual coding, coding 1206 Text materials, text information 1202 Map 1204 Machine-readable visual coding 1210 images 1212a Text Label 1212b Text Label 1214 Machine-readable visual coding 1216a Mapping Features 1216b Mapping Features 1220 images 1222 Large Logo 1224 frames 1226a Bench 1226b Bench 1228 Joint 1230 Boundary 1232 Lighting 1300 images 1304 Machine-readable visual coding 1310 Door frame 1312 Window frame 1314 Window sill 1316 Awning 1318 Lighting Features 1320 Feature Map 1322 Anchor point 1324 segments, users 1326 Smartphones, imaging devices 1328 routes 1330 Route 1400 System 1402 Delivery driver, facial recognition data 1403 First machine-readable visual coding 1404 Package 1405 Second machine-readable visual coding 1420 "Smart" Home System, Home System 1422 Camera 1424 "Smart" Lock 1426 "Smart" Components 1430 Network 1440 Servers 1500 computing systems, systems 1502 User computing device, device, user device 1512 processors 1514 Memory, User Computing Device Memory 1516 data 1518 command 1520 Sensor 1521 User Input Components 1522 Imaging Sensor 1523 Geospatial Sensor 1524 Conversion Sensor 1525 Rotation Sensor 1527 Image data 1528 Recognition Model, Model 1530 Server Computing System 1532 processors 1534 memory 1536 data 1538 command 1540 Recognition Model, Model 1540a Encoding recognition model, model, recognition model 1540b Image recognition model, model, recognition model 1542 Memorized standards 1544 Encoded Data 1546 Context Data 1548 First substandard, substandard 1550 Context Data 1552 Second substandard, substandard 1554 Context data 1556 Instructions, operation instructions 1560 Provider Computing Systems 1562 processors 1564 memory 1566 data 1568 command 1570 Instructions, Service Instructions 1572 Context Data 1580 Network 1582 Computing Devices 1584 Computing Devices 1590 Recognition System 1592 Recognition data 1594 Recognition System 1596 Context Components 1600 methods

Claims

1. A computing system, One or more processors, The system comprises one or more non-temporary computer-readable media for storing instructions, wherein the instructions are executable by one or more processors to cause the computing system to perform an operation, Decoding the data using a first recognition system for processing first image data that describes a machine-readable visual code that encodes the data to be decoded, Transmitting the decoded data to a server computing system for identifying the stored criteria, wherein the stored criteria include one or more stored feature maps related to a scene. Generating a feature map using a second recognition system for processing second image data describing a scene, wherein the second recognition system includes a visual positioning system. Transmitting the aforementioned feature map to the server computing system, The server computing system receives a verification indicator from which a match is shown between the generated feature map and one or more stored feature maps. Based on the aforementioned verification indicator, the augmented reality rendering is rendered. A computing system that includes this.

2. A computing system, One or more processors, The system comprises one or more non-temporary computer-readable media for storing instructions, wherein the instructions are executable by one or more processors to cause the computing system to perform an operation, Decoding data using a first recognition system for processing first image data describing a machine-readable visual code that encodes the data to be decoded, wherein the first image data represents an image of product packaging including the machine-readable visual code, Transmitting the decoded data to a server computing system for identifying stored criteria, wherein the stored criteria include one or more stored feature maps related to a scene and correspond to a product related to the product packaging. This involves generating a feature map using a second recognition system to process a second image data that describes the scene, Transmitting the aforementioned feature map to the server computing system, The server computing system receives a verification indicator from which a match is shown between the generated feature map and one or more stored feature maps. Based on the aforementioned verification indicator, the augmented reality rendering is rendered. A computing system that includes this.

3. The computing system according to claim 1 or 2, wherein the feature map describes features in three-dimensional space.

4. The computing system according to claim 2, wherein the second recognition system comprises a visual positioning system.

5. The aforementioned operation, A computing system according to claim 1 or 4, comprising locating anchor points with respect to features within the aforementioned scene.

6. The computing system according to claim 1 or 2, comprising an image sensor, wherein the image captured by the image sensor comprises the first image data and the second image data.

7. The aforementioned operation, The computing system according to claim 1 or 2, comprising determining the pose of the computing system relative to the scene using a transformation sensor or a rotation sensor.

8. One or more non-temporary computer-readable media for storing instructions, wherein the instructions are executable by one or more processors to cause a computing system to perform an operation, Decoding the data using a first recognition system for processing first image data that describes a machine-readable visual code that encodes the data to be decoded, Transmitting the decoded data to a server computing system for identifying the stored criteria, wherein the stored criteria include one or more stored feature maps related to a scene. Generating a feature map using a second recognition system for processing second image data describing a scene, wherein the second recognition system includes a visual positioning system. Transmitting the aforementioned feature map to the server computing system, The server computing system receives a verification indicator from which a match is shown between the generated feature map and one or more stored feature maps. Based on the aforementioned verification indicator, the augmented reality rendering is rendered. One or more non-temporary computer-readable media, including [the specified text].

9. One or more non-temporary computer-readable media for storing instructions, wherein the instructions are executable by one or more processors to cause a computing system to perform an operation, Decoding data using a first recognition system for processing first image data describing a machine-readable visual code that encodes the data to be decoded, wherein the first image data represents an image of product packaging including the machine-readable visual code, Transmitting the decoded data to a server computing system for identifying stored criteria, wherein the stored criteria include one or more stored feature maps related to a scene and correspond to a product related to the product packaging. This involves generating a feature map using a second recognition system to process a second image data that describes the scene, Transmitting the aforementioned feature map to the server computing system, The server computing system receives a verification indicator from which a match is shown between the generated feature map and one or more stored feature maps. Based on the aforementioned verification indicator, the augmented reality rendering is rendered. One or more non-temporary computer-readable media, including [the specified text].

10. One or more non-temporary computer-readable media according to claim 8 or 9, wherein the feature map describes features in three-dimensional space.

11. The one or more non-temporary computer-readable media according to claim 9, wherein the second recognition system comprises a visual positioning system.

12. The aforementioned operation, One or more non-temporary computer-readable media according to claim 8 or 11, comprising locating anchor points with respect to features within the aforementioned scene.

13. The operation comprises capturing an image using an image sensor, wherein the image comprises the first image data and the second image data, one or more non-temporary computer-readable media according to claim 8 or 9.

14. The aforementioned operation, One or more non-temporary computer-readable media according to claim 8 or 9, comprising determining the pose of the computing system relative to the scene using a conversion sensor or a rotation sensor.

15. A computing system, One or more processors, The system comprises one or more non-temporary computer-readable media for storing instructions, wherein the instructions are executable by one or more processors to cause the computing system to perform an operation, Receiving a first feature map generated by a first client computing device using a scene recognition system for processing scene image data describing a scene, wherein the scene recognition system includes a visual positioning system. Associating the feature map described above with a stored criterion, Receiving data decoded from a second computing device, wherein the decoded data is acquired by the second computing device using an encoding recognition system for processing first image data describing a machine-readable visual encoding that encodes the decoded data, Using the decoded data, identify the stored criteria, Receiving a second feature map from the aforementioned second computing device, Comparing the second feature map and the first feature map, Based on the comparison, augmented reality rendering is initiated for the second computing device. A computing system that includes this.

16. A computing system, One or more processors, The system comprises one or more non-temporary computer-readable media for storing instructions, wherein the instructions are executable by one or more processors to cause the computing system to perform an operation, Receiving a first feature map generated by a first client computing device using a scene recognition system for processing scene image data that describes the scene, Associating the feature map described above with a stored criterion, Receiving data decoded from a second computing device, wherein the decoded data is acquired by the second computing device using an encoding recognition system for processing first image data describing a machine-readable visual code encoding the decoded data, the first image data representing an image of product packaging including the machine-readable visual code, and the stored criteria corresponding to a product associated with the product packaging. Using the decoded data, identify the stored criteria, Receiving a second feature map from the aforementioned second computing device, Comparing the second feature map and the first feature map, Based on the comparison, augmented reality rendering is initiated for the second computing device. A computing system that includes this.

17. The aforementioned operation, The computing system according to claim 15 or 16, comprising transmitting a verification indicator to the second computing device.

18. The computing system according to claim 16, wherein the scene recognition system includes a visual positioning system.

19. The aforementioned comparison, The computing system according to claim 15 or 16, comprising determining the similarity between the first feature map and the second feature map.

20. The computing system according to claim 15 or 16, wherein the first feature map is generated from a plurality of vantage points of the scene, and the second feature map is generated from a single vantage point of the scene.

Citation Information

Patent Citations

  • Own-vehicle position detection system using scenic image recognition

    JP2011215052A

  • Terminal device, information processing apparatus, display control method, and display control program

    JP2015148957A