Augmented reality system
By processing sensor data and obfuscating redundant information in the AR system, only the necessary positioning data is transmitted to the remote computing device, thus solving the problems of insufficient computing resources and leakage of sensitive information, and achieving efficient location orientation determination and information security.
Patent Information
- Application Number
- CN202110627513.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-10
- Filing Date
- 2021-06-04
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2041-06-04
AI Technical Summary
Existing AR systems suffer from insufficient computing resources and power consumption when determining location and orientation, especially when running AR applications on mobile computing devices, and the transmitted data may leak sensitive information.
By processing sensed data in the AR system to identify and obscure redundant information, only the necessary data for determining location and orientation is transmitted, and then transmitted to a remote computing device for further processing. Neural networks and object recognition algorithms are used to identify and remove sensitive information.
This improves the computational efficiency and accuracy of location determination in AR systems, while protecting sensitive information from leakage and reducing data transmission volume and computational load.
Smart Images

Figure CN113778221B_ABST
Abstract
Description
BACKGROUND TECHNICAL FIELD
[0002] The present disclosure relates to augmented reality (AR) systems. The invention has particular, but non-exclusive, relevance to the security of data used to determine the position and orientation of an AR system.
[0003] Description of Related Art
[0004] AR devices provide experiences to users in which a representation of a real-world environment is augmented with computer-generated sensory information. In order to accurately provide these experiences to users, the position and orientation of the AR device is determined so that the computer-generated sensory information can be seamlessly integrated into the representation of the real world. An alternative term for AR is "mixed reality," which refers to the merging of real and virtual worlds.
[0005] Augmenting a real-world environment with computer-generated sensory information can include using sensory information that encompasses one or more sensory modalities, including, for example, visual information (in the form of images, which in some cases can be text or simple icons), auditory information (in the form of audio), tactile information (in the form of touch), somatosensory information (related to the nervous system), and olfactory information (related to smell).
[0006] Superimposing sensory information onto a real-world (or "physical") environment can be done constructively (by making additions to the natural environment) or destructively (by making subtractions from the natural environment or masking the natural environment). AR changes the user's perception of the real-world environment in which they are located, whereas virtual reality (VR) replaces the real-world environment in which the user is located with a completely simulated (i.e., computer-generated) environment.
[0007] AR devices include, for example, AR-enabled smartphones, AR-enabled mobile computers such as tablets, and AR head-mounted devices including AR glasses. The position and orientation of an AR device relative to the environment in which it is located is typically determined based on sensor data collected by the AR device or associated with the AR device through a localization process. SUMMARY
[0008] According to a first aspect of the disclosure, there is provided an augmented reality, AR, system comprising: one or more sensors arranged to generate sensing data representing at least a portion of an environment in which the AR system is located; a storage device for storing the sensing data generated by the one or more sensors; one or more communication modules for transmitting positioning data to be used to determine a position and orientation of the AR system; and one or more processors arranged to: obtain the sensing data representing the environment in which the AR system is located; process the sensing data to identify a first portion of the sensing data representing redundant information; derive the positioning data for determining the position and orientation of the AR system, wherein the positioning data is derived from the sensing data and the first portion is obfuscated during the derivation of the positioning data; and transmit at least a portion of the positioning data using the one or more communication modules.
[0009] According to a second aspect of the disclosure, there is provided a computer- implemented data processing method for an augmented reality, AR, system, the method comprising: obtaining sensing data representing an environment in which the AR system is located; processing the sensing data to identify a first portion of the sensing data representing redundant information; deriving positioning data for determining a position and orientation of the AR system, wherein the positioning data is derived from the sensing data and the first portion is obfuscated during the derivation of the positioning data; and transmitting at least a portion of the positioning data.
[0010] According to a third aspect of the disclosure, there is provided a non-transitory computer-readable storage medium comprising computer-readable instructions which, when executed by at least one processor, cause the at least one processor to: obtain sensing data representing an environment in which an augmented reality, AR, system is located; process the sensing data to identify a first portion of the sensing data representing redundant information; derive positioning data for determining a position and orientation of the AR system, wherein the positioning data is derived from the sensing data and the first portion is obfuscated during the derivation of the positioning data; and transmit at least a portion of the positioning data. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is a schematic diagram of an AR system according to an example.
[0012] Figure 2 is shown a flow diagram illustrating a computer-implemented method for data processing of an augmented reality system according to an example.
[0013] Figure 3 is schematically illustrated an example of an augmented reality system implementing a method according to an example.
[0014] Figure 4An example of an augmented reality system implementing a method according to an example is schematically illustrated.
[0015] Figure 5 An example of an augmented reality system implementing a method according to an example is schematically illustrated.
[0016] Figure 6 A non-transitory computer-readable storage medium comprising computer-readable instructions according to an example is schematically illustrated. DETAILED DESCRIPTION
[0017] The details of systems and methods according to examples will become apparent from the following description, with reference to the drawings. In this description, for the purposes of explanation, numerous specific details of certain examples are set forth. Reference in this specification to "an example," or similar language, indicates that a particular feature, structure, or characteristic described in connection with the example is included in at least that one example, but not necessarily in other examples. It should also be noted that some examples are described schematically, with certain features omitted and / or necessarily simplified in order to facilitate explanation and understanding of the concepts upon which the examples are based.
[0018] Described herein are systems and methods related to data processing in the context of an augmented reality (AR) system. An AR system provides an augmented reality experience to a user in which virtual objects that can include perceptual information are used to augment a representation or perception of a real-world environment. The representation of the real-world environment can include data originating from sensors, which can also be referred to as sensory data, corresponding to one or more sensory modalities, such as vision (in the form of image data), hearing (in the form of audio data), touch (in the form of tactile data), neurology (in the form of somatosensory data), and smell (in the form of olfactory data).
[0019] Sensing data can represent a physical quantity that can be measured by a sensor. A sensor can be a device configured to measure a physical quantity, such as light, depth, motion, sound, etc., and convert it into a signal, e.g., an electrical signal. Examples of sensors include image sensors, haptic sensors, motion sensors, depth sensors, microphones, sound navigation and ranging (Sonar) devices, light detection and ranging (LiDAR) devices, radio detection and ranging (RADAR), global positioning system GPS sensors, and sensors included in an inertial measurement unit (IMU), such as accelerometers, gyroscopes, and in some cases magnetometers. For example, an image sensor can convert light into a digital signal. Image sensors include image sensors that operate in the visible light spectrum, but additionally or alternatively can include image sensors that operate outside the visible light spectrum, e.g., in the infrared spectrum. Thus, sensing data associated with an image captured by a sensor can include image data representing the image captured by the sensor. However, in other examples, sensing data can additionally or alternatively include audio data representing sound (e.g., that can be measured by a microphone), or another type of sensor-originated data representing a different physical quantity that can be measured by a corresponding type of sensor (e.g., haptic, somatosensory, or olfactory data). In some cases, sensing data can be source data or “raw data” that is output directly from a sensor (e.g., sensor data). In such cases, the sensing data can be acquired from the sensor, e.g., through direct transmission of the data or by reading the data from an intermediate storage device where the data is stored. In other cases, the sensing data can be pre-processed: e.g., further processing can be applied to the sensing data after it has been acquired by the sensor and before it is processed by a processor. In some examples, the sensing data includes a processed version of the sensor data that is output by the sensor. For example, raw sensory input can be processed to convert low-level information into higher-level information (e.g., shapes are extracted from images for object recognition).
[0020] To provide an AR experience, the position and orientation of the AR system within the real-world environment is determined by a localization process. Determining the position and orientation of the AR system allows for accurate integration of virtual objects into the representation of the real world such that the user of the AR system experiences an immersive fusion of the real world and virtual augmentation. The position and orientation of the AR system can be collectively referred to as a “geopose” or “geographically anchored pose,” which represents the spatial position of the AR system and the orientation or “pose” of the AR system, which specifies the pitch, roll, and yaw according to a coordinate system.
[0021] To determine the position and orientation of an AR system, localization data can be processed to determine the relative position of the AR system within an environment. The localization data can be derived from sensing data that provides information representative of the environment in which the AR system is located and / or information related to the orientation and / or motion of the AR system. For example, a portion of image data generated by an image sensor included by the AR system can be selected for inclusion in the localization data. Alternatively or additionally, the image data can be processed to identify a set of feature points and construct feature descriptors that encode information related to the feature points, enabling the feature points to be distinguished. These feature points and descriptors are used to identify objects and structures within the environment, enabling the relative position of the AR system to be determined. The localization data can be derived from multiple types of sensing data generated by different sensors. For example, image data or data representative of a set of feature points and descriptors can be used in conjunction with motion data generated from an inertial measurement unit during localization to accurately identify and track the position and orientation of the AR system as it moves through a real-world environment. Alternatively or additionally, the image data can be supplemented by depth information generated by a depth sensor or derived from LiDAR, RADAR, and other output devices to identify the relative positions of objects in the images represented by the image data.
[0022] As AR services and systems become more prevalent, it is desirable for AR experiences to persist across multiple device types, operating systems, and over time. To this end, AR experiences can be stored, and certain AR functionality can be implemented by remote systems, such as an AR cloud implemented by one or more remote computing devices. The AR cloud can include or implement a real-time spatial (i.e., "three-dimensional" or "3D") map of the real world, for example, in the form of a point cloud. One such AR functionality that can be performed by the AR cloud is localization. In this case, an AR system arranged to provide an AR experience to a user provides localization data to the AR cloud, and the AR cloud determines the position and orientation of the AR system based on the localization data and the real-time spatial map of the real world. In other examples, the AR cloud can include or implement a real-time spatial map of a particular portion of a real-world environment. Position data (or "geo-pose data") representative of the position and orientation of an AR system relative to the environment can then be provided by the AR cloud to the AR system.
[0023] Although some AR systems are capable of performing localization without the use of remote systems, for example, AR systems arranged to perform simultaneous localization and mapping (SLAM) can be capable of providing AR experiences without the need to transmit data to an AR cloud, performing localization of an AR device can require significant computational power. Localization is a challenge for AR applications executing on AR systems that are mobile computing devices, such as general-purpose smartphones and general-purpose tablet computing devices, which have a relatively small amount of available computing resources and / or power.
[0024] In this way, performing certain AR functions away from the AR system can allow data storage and processing performed by the AR system to be kept to a minimum necessary, allowing the AR system to have a size, weight, and form factor that is practical and advantageous for long-term use and / or daily use of the AR system.
[0025] Performing localization of the AR system in a remote computing device (e.g., one or more servers implementing an AR cloud) can also allow for determining the relative positions of multiple AR systems in an environment and correlating them when providing an AR experience using multiple AR systems.
[0026] Certain examples described herein relate to an AR system arranged to provide localization data to one or more external computing devices to determine a position and orientation of the AR system. The localization data is derived from sensing data generated by sensors of the AR system, and when the localization data is derived, a first portion of the sensing data that is representative of redundant information is obfuscated. Redundant information includes information that is not used in determining the position and orientation of the AR system, such as sensitive information. The first portion of data is obfuscated in such a way that the localization data transmitted to the one or more remote computing devices can not be used to determine the redundant information captured by the one or more sensors. For example, obfuscating the first portion can include modifying the first portion, or in some cases, excluding or "removing" the first portion when the localization data is derived. In either case, obfuscating the first portion is done in such a way that it is not possible to determine the redundant information from the localization data, e.g., by reverse engineering the localization data.
[0027] Figure 1 An example of an AR system 100 is shown, which can be embodied as a single device, such as an AR-enabled smartphone or an AR headset. Alternatively, the AR system 100 can be implemented by multiple devices, which can be communicatively coupled via wired or wireless means. For example, an AR device, such as a mobile computer or an AR-enabled smartphone, is in communication with one or more AR accessories, such as an AR headset.
[0028] The AR system 100 comprises one or more sensors 102 arranged to generate sensing data representative of at least a portion of the environment in which the AR system 100 is located. The one or more sensors 102 comprise one or more cameras for generating image data representative of a portion of the environment falling within the field of view of the one or more cameras. The field of view can be delimited in vertical and / or horizontal directions, depending on the number and position of the cameras. For example, the cameras can be arranged to face substantially the same direction as the direction in which the user’s head is facing, e.g. in the case that the user is wearing an AR headset, the field of view of the one or more cameras can comprise all or part of the user’s field of view. Alternatively, the field of view can comprise a wider area, e.g. completely surrounding the user. The cameras can comprise stereo cameras, from which the AR system can derive depth information indicating the distance to objects in the environment using stereo matching. Alternatively or in addition, the sensors 102 can comprise depth sensors, infrared cameras, sonar transceivers, LiDAR systems, RADAR systems, etc. for generating depth information.
[0029] The sensors 102 can also comprise a position sensor for determining the position and / or orientation (collectively, position or pose) of the user of the AR system 100. The position sensor can comprise a global positioning system (GPS) module, as well as one or more accelerometers, one or more gyroscopes and / or a Hall effect magnetometer (electronic compass) for determining orientation, e.g. comprised in an IMU.
[0030] The AR system 100 comprises a storage 104 for storing the sensing data 106 generated by the one or more sensors 102. The storage 104 can be embodied as any suitable combination of non-volatile storage and / or volatile storage. For example, the storage 104 can comprise one or more solid state drives (SSDs), as well as non-volatile random access memory (NVRAM) and / or volatile random access memory (RAM), e.g. static random access memory (SRAM) and dynamic random access memory (DRAM). Other types of memory can be included, such as removable storage, synchronous DRAM, etc.
[0031] The AR system 100 further comprises one or more communication modules 108 for transmitting positioning data 110 to be used for positioning of the AR system 100. For example, the communication modules 108 can transmit the positioning data 110 to one or more remote computing devices implementing an AR cloud that provides positioning functionality to identify the position and orientation of the AR system 100 within the environment. Alternatively, the remote computing devices can forward the positioning data 110 to a further one or more remote computing devices implementing the AR cloud.
[0032] The communication module 108 can be arranged to transmit the positioning data 110 over any suitable wireless communication type. For example, the communication module 108 can use Wi-Fi ® , Bluetooth ® , infrared, cellular frequency radio waves, or any other suitable wireless communication type. Alternatively or additionally, the communication module 108 can be arranged to transmit data over a wired connection.
[0033] The AR system 100 comprises one or more processors 112. The processors 112 can comprise various processing units, including central processing units (CPUs), graphics processing units (GPUs), and / or specialized neural processing units (NPUs) for efficiently performing neural network operations. In accordance with the present disclosure, neural networks can be used for certain tasks, including object detection, as will be described in greater detail below. The one or more processors 112 can comprise other specialized processing units, such as application-specific integrated circuits (ASICs), digital signal processors (DSPs), or field-programmable gate arrays (FPGAs).
[0034] The storage 104 holds machine-readable instructions in the form of program code 114 that, when executed by the one or more processors 112, cause the AR system 100 to perform the methods described below. The storage 104 is also arranged to store further data for performing the described methods. The further data in this example comprises the sensing data 106 generated by the one or more sensors 102.
[0035] It will be appreciated that the AR system 100 can comprise other components not shown in Figure 1 , for example a user interface for providing an AR experience to a user of the AR system 100. The user interface can comprise any suitable combination of input devices and output devices. Input devices include, for example, a touchscreen interface for receiving user input, actuatable buttons for receiving user input, sensors such as motion sensors, or microphones suitable for sensing user input. Output devices can include a display such as a touchscreen display, a speaker, a haptic feedback device, and the like.
[0036] Figure 2 An example of a method 200 performed by the AR system 100 in accordance with the present disclosure is shown. It will be appreciated that while the method 200 is described with respect to the AR system 100, the same method 200 can be performed by any suitable AR system arranged to provide an AR experience to a user and to perform positioning for determining a position and orientation of the AR system by transmitting positioning data 110 to one or more remote computing devices, for example by associating the positioning data generated by the AR system 100 with data comprised in an AR cloud.
[0037] At a first block 202, the AR system 100 acquires sensing data 106 representing an environment in which the AR system 100 is located. Acquiring the sensing data 106 can include accessing one or more portions of the storage 104 that store the sensing data 106. In some examples, the sensing data 106 is generated by the one or more sensors 102 and stored directly on the storage 104. Alternatively or additionally, at least a portion of the sensing data 106 can be generated by the one or more sensors 102 and processed prior to being stored on the storage 104. For example, data generated by an image sensor can represent light intensity values generated based on light captured at a plurality of pixel sensors included in the image sensor. This data can be processed to generate image data representing an image of the environment. Alternatively, the sensing data 106 can be streamed directly from the one or more sensors 102.
[0038] At a block 204, the AR system 100 processes the sensing data 106 to identify a first portion of the sensing data 106 that represents redundant information. When the sensing data 106 is generated using the one or more sensors 102, redundant information related to the environment can be collected. Redundant information includes information that is not used to determine a location and orientation of the AR system 100. The redundant information can include information related to one or more dynamic objects, e.g., objects that are located in the environment and move through the environment. Alternatively or additionally, for example, where the sensing data 106 includes image data representing an image of the environment, the redundant information can relate to sensitive information, which can include representations of people, specific parts of people (such as their faces, their clothing, identification cards they are wearing), objects in the environment (e.g., credit cards), sensitive information displayed on digital displays (such as computer monitors or television units), information such as text printed on paper, etc. In these cases, processing the sensing data 106 to identify the first portion of the sensing data 106 that represents redundant information includes identifying portions of the image data that represent one or more objects in the image of the environment. Where the sensing data 106 includes audio data generated by a microphone in the AR system 100, the redundant information can include speech of one or more people in the environment, e.g., when reading out credit card details, etc. Depth or location information represented by sensing data 106 generated by a sensor (such as a depth sensor, sonar, RADAR, LiDAR, etc.) can be redundant in nature if it relates to objects that are dynamic in the environment or the arrangement and location of the objects within the environment is confidential.
[0039] What is considered sensitive information can vary depending on the environment in which the AR system 100 is being used and the type of sensing data 106 being processed. For example, where the AR system 100 is being used in a manufacturing context to aid in the design and / or construction of a product, objects in the environment in which the AR system 100 is operating can have confidential properties, e.g., related to trade secrets, unreleased products, and confidential intellectual property. In high-security environments, such as military or government buildings or facilities, a higher degree of sensitivity can be assigned to objects that would not otherwise be considered sensitive.
[0040] Processing the sensing data 106 to identify a first portion of the sensing data 106 that represents redundant information can include using one or more object recognition (or“object detection,”“object identification,”“object classifier,”“image segmentation”) algorithms. These algorithms are configured to detect instances of objects of a particular class in a real-world environment, e.g., from an image / audio representation of the real-world environment represented by the sensing data 106, and the location of the redundant information within the sensing data, e.g., the location of instances of objects in an image. The one or more object recognition algorithms used can be implemented for other purposes in the AR system 100. For example, the one or more object recognition algorithms can also be used in other processes to understand the environment in which the AR system 100 is located. In this case, the output of these processes can be used in the method 200 described herein, which allows the method 200 to be performed without substantially increasing the amount of processing performed by the AR system 100.
[0041] Where the predetermined class is a human face, the object recognition algorithm can be used to detect the presence of a human face in an image represented by image data included in the sensing data. In some cases, a particular instance of an object can be identified using the one or more object recognition algorithms. For example, the instance can be a particular human face. Other examples of such object recognition include identifying or detecting instances of expressions (e.g., facial expressions), gestures (e.g., hand gestures), audio (e.g., identifying one or more particular sounds in an audio environment), thermal energy signatures (e.g., identifying objects such as faces in infrared representations or“heat maps”). Thus, in examples, the type of“object” that is detected can correspond to the type of representation of the real-world environment. For example, for a visual or image representation of a real-world environment, object recognition can involve identifying particular articles, expressions, gestures, etc., while for an audio representation of a real-world environment, object recognition can involve identifying particular sounds or sound sources. In some examples, object recognition can involve detecting motion of the identified object. For example, in addition to identifying an instance of a particular type of object (e.g., a car in an audio / visual representation of a real-world environment), object recognition can also detect or determine the motion of the instance of the object (e.g., the identified car).
[0042] In an example, processing the sensing data 106 to identify a first portion of the sensing data that represents redundant information can include implementing a support vector machine (SVM) or a neural network to perform object recognition, although other types of object recognition methods can also be present.
[0043] A neural network generally includes a number of interconnected neurons that form a directed weighted graph, where the vertices (corresponding to neurons) or edges (corresponding to connections) of the graph are respectively associated with weights. The weights can be adjusted for particular purposes throughout the training of the neural network, thereby changing the output of individual neurons and thus the output of the neural network as a whole. In a convolutional neural network (CNN), a fully connected layer generally connects every neuron in one layer to every neuron in another layer. Thus, as part of an object classification process, a fully connected layer can be used to identify overall features of an input, such as whether a particular class of object or a particular instance belonging to a particular class is present in the input (e.g., an image, a video, a sound).
[0044] A neural network can be trained to perform object detection, image segmentation, sound / speech recognition, etc. by processing sensing data, e.g., to determine whether an object of a predetermined class of objects is present in a real-world environment represented by the sensing data. Training a neural network in this manner can generate one or more kernels associated with at least some layers, such as layers other than the input and output layers of the neural network. Thus, an output of the training can be a plurality of kernels associated with a predetermined neural network architecture (e.g., different kernels are associated with respective different layers of the multi-layer neural network architecture). The kernel data can be considered to correspond to weight data representing weights to be applied to image data, as each element of a kernel can be considered to respectively correspond to a weight. Each of these weights can be multiplied by a corresponding pixel value of an image patch to convolve the kernel with the image patch as described above.
[0045] A kernel can allow identification of features to an input of a neural network. For example, in the case of image data, some kernels can be used to identify edges in an image represented by the image data, and other kernels can be used to identify horizontal or vertical features in the image (although this is not limiting, and other kernels are possible). The precise features that a kernel is trained to identify can depend on the image characteristics that the neural network is trained to detect, such as the class of objects. A kernel can have any size. A kernel can sometimes be referred to as a “filter kernel” or “filter.” Convolution generally involves multiplication operations and addition operations, sometimes referred to as multiply-accumulate (or “MAC”) operations. Thus, a neural network accelerator configured to implement a neural network can include multiply-accumulator (MAC) units configured to perform these operations.
[0046] Following the training phase, the neural network (which may be referred to as a trained neural network) can be used, for example, to detect the presence of objects of a predetermined object category in an input image. This process may be called "classification" or "inference." Classification typically involves convolving the kernel acquired during the training phase with portions of the input originating from the sensor (e.g., image patches of the image input to the neural network) to generate feature maps. These feature maps can then be processed using at least one fully connected layer, for example, to classify objects; however, other types of processing can also be performed.
[0047] In the case of using a Region Convolutional Neural Network (R-CNN) to generate bounding boxes (e.g., to identify the location of detected objects in an image), the processing may first involve using CNN layers, followed by a Region Scheme Network (RPN). In examples where CNNs are used to perform image segmentation, such as using a Fully Convolutional Neural Network (FCN), the processing may include using a CNN, followed by deconvolutional layers (i.e., transposed convolutions), and then upsampling.
[0048] return Figure 2 In method 200, at third block 206, AR system 100 derives positioning data 110 to determine the position and orientation of AR system 100. Positioning data 110 is derived from sensing data 106, and a first portion is blurred during the export of the positioning data. The first portion of sensing data 106 can be blurred in various ways; for example, the value represented by the first portion of sensing data 106 can be modified during the export of positioning data 110. In other examples, positioning data 110 can be derived from a second portion of sensing data 106, which is different from the first portion of sensing data 106, and blurring the first portion can include excluding the first portion from the export of positioning data 110. Blurring the first portion of the data is performed in a manner that prevents it from containing redundant information and makes it impossible to determine redundant information.
[0049] In one example, the localization data 110 includes image data representing at least a portion of the environment derived from the sensory data 106. In this case, deriving the localization data 110 can include modifying the image data included in the sensory data 106. For example, one or more segmentation masks associated with one or more objects in the image can be used to identify a first portion of the sensory data 106, which can then be modified such that it no longer represents redundant information. The modification in this context can involve modifying pixel intensity values represented by the first portion. Alternatively or additionally, the modification can include removing or deleting the first portion when deriving the localization data 110. The one or more segmentation masks can be generated based on the object identification methods described above. The one or more segmentation masks define a representation of the one or more objects in the image and identify portions of the image data representing these objects. Alternatively, in the case where the localization data 110 includes image data derived from the image data included in the sensory data, deriving the localization data can include selecting a second portion of the sensory data as the localization data 110 that does not include the first portion.
[0050] At the fourth block 208, the AR system 100 transmits at least a portion of the localization data 110 to one or more remote computing devices including or implementing an AR cloud using, for example, the one or more communication modules 108 for performing localization of the AR system 100. Since the first portion of the sensory data 106 is obfuscated during the derivation of the localization data 110, the redundant information, such as sensitive information, can be prevented from being transmitted to the one or more remote computing devices. In this way, the method 200 can prevent the localization data 110 from being used to determine sensitive information about the environment in which the AR system 100 is located in the case where the redundant information includes sensitive information. In some cases, certain entities can intercept communications from the AR system 100 to the one or more remote computing devices, in which case any communications that can be intercepted do not include data that can be used to determine sensitive information. In some cases, the one or more computing devices including or implementing the AR cloud can be provided and / or managed by multiple service providers, and thus limiting access to data that can be used to determine sensitive information is desirable.
[0051] Figure 3The AR system 100 is shown with acquired sensing data 302, which includes image data representing an image of an environment. The AR system 100 processes the image data to identify two objects 304 and 306 in the image represented by a first portion of the sensing data 302. The AR system 100 generates segmentation masks 308 and 310 associated with the two objects 304 and 306. In this case, the segmentation masks 308 and 310 represent portions of the image that include the objects 304 and 306. The AR system 100 derives localization data 312 from the sensing data 302 that includes the image data. The first portion of the sensing data 302, which is the image data representing the one or more objects 304, is excluded from the localization data 312. However, it will be appreciated that in some examples, supplemental data can be provided in the localization data 312 that represents the portions of the image in which the objects 304 and 306 are located. For example, a label can be provided indicating that these portions of the image are restricted and therefore not shown.
[0052] In some cases, further localization data can be generated and transmitted after the transmission of the localization data 312. In this case, when the one or more objects 304 and 306 are dynamic objects that are moving between frames of the captured image data, the subsequent localization data can include image data representing these portions of the image so that accurate localization can still be performed based on these areas of the environment. By excluding data representing redundant information in the localization data 312, the amount of data transmitted can be reduced, allowing for faster communication between the AR system 100 and one or more remote computing devices 314, and also improving the efficiency of determining the position and orientation of the AR system 100, as less data is used for localization using AR cloud processing.
[0053] While in this example the segmentation masks 308 and 310 represent bounding boxes that include each of the detected objects 304 and 306, in other examples the segmentation masks 308 and 310 represent portions of the image that are the same size and shape as the detected objects 304 and 306. In some cases, it can be sufficient to identify a portion of the sensing data 302 that represents portions of the image that are the same size and shape as the detected objects. In other cases, the shape and size of the objects themselves can be inherently sensitive, so by identifying a portion of the data that includes and obscures the shape of the objects 304 and 306, information related to the size and shape of the objects 304 and 306 can not be included in the localization data 312. While bounding boxes have been used to represent the segmentation masks 308 and 310 in this example, it should be understood that other shapes of the segmentation masks 308 and 310 can be used, including regular and irregular polygons, curved shapes, and any other suitable shape. In some cases, the size and shape of the segmentation mask or masks 308 and 310 used can depend on the class of the objects 304 and 306 that have been detected in the image. Dynamic objects captured in the image can also affect the accuracy of the determination of the position and orientation (or “geo- pose determination”), so by excluding data representing these objects, the accuracy of the geo- pose determination can be improved.
[0054] The AR system 100 transmits at least a portion of the localization data 312 for receipt by one or more remote computing devices 314 that implement the AR cloud 314. The one or more remote computing devices 314 include a point cloud 316 that represents a real-time spatial map of the real-world environment, including the environment in which the AR system 100 is located. This point cloud 316 is used with the localization data 312 to determine the position and orientation of the AR system 100. The AR system can then receive position data 318 from the one or more remote computing devices 314 that represents the position and orientation of the AR system 100.
[0055] In some examples, the localization data 312 includes metadata that identifies portions of the image that represent one or more objects 304 and 306 in the image. By identifying portions of the image that represent one or more objects 304 and 306 in the image within the localization data, the one or more remote computing devices 314 can be informed of portions of the image data included in the localization data that are not to be processed. Thus, the remote computing devices 314 can avoid wasting computational resources in attempting to process these portions of data, and / or can use this information to ensure that these portions of data do not impact the results of the determination of the position and orientation of the AR system 100.
[0056] As noted above, the localization data 110 can alternatively or in addition to include other data in addition to image data. Figure 4An example is shown in which the AR system 100 acquires sensing data 402 including image data representing an image of an environment. The AR system 100 processes the image data to identify one or more objects 404 and 406 represented by a first portion of the sensing data 402. The AR system 100 derives positioning data 408 including data representing a set of one or more feature points generated using the image data. At least a portion of the positioning data 408 representing at least a portion of the one or more feature points is then transmitted. Feature points can include, for example, edges, corners, blobs, ridges, and other relevant features in the image. Generating feature points from an image can include using one or more feature detection algorithms, such as scale-invariant feature transform (SIFT), fast accelerated segment test feature (FAST), local binary pattern (LBP), and other known feature detection algorithms. In some examples, feature points are associated with corresponding feature descriptors, and the feature descriptors can be included in the positioning data 408.
[0057] Although the feature points and descriptors do not include image data representing an image of the environment, in some cases data representing the feature points and descriptors can be processed, e.g., through a feature inversion process, to identify redundant information. In such cases, providing positioning data 408 that does not include data representing feature points associated with the redundant information can inhibit or prevent reconstruction of the original image including the redundant information.
[0058] In some examples, deriving the positioning data 408 can include processing a second portion of the sensing data 402 that does not include the first portion of the sensing data 402 to generate the set of one or more feature points. Alternatively, the positioning data 408 can be derived by processing the sensing data 402 to generate the set of one or more feature points including using the first portion of the sensing data 402, and subsequently removing data representing certain feature points from the set of one or more feature points corresponding to the first portion of the sensing data 402. In either case, one or more segmentation masks can be generated to identify the first portion of the sensing data 402 as described above.
[0059] Figure 5An example is shown in which the AR system 100 acquires sensing data 502 that includes image data representing an image of an environment. The AR system 100 processes the image data to identify one or more objects 504 and 506 represented by a first portion of the sensing data 502. The AR system derives positioning data 508 that includes data representing a point cloud 510. The point cloud 510 is a 3D representation of the environment in the form of a plurality of points. At least a portion of the positioning data 508 including the point cloud 510 is then transmitted to one or more remote computing devices 314 to determine a position and orientation of the AR system 100. In this case, deriving the positioning data 508 can include generating data representing the point cloud 510 from a second portion of the sensing data 502 that does not include the first portion. Alternatively, deriving the positioning data 508 can include generating the point cloud 510 from the sensing data 502 using both the first portion and the second portion of the sensing data 502, and subsequently removing data representing points in the point cloud 510 that are generated based on the first portion.
[0060] In some examples, the transmitted positioning data 312, 408, and 508 can include a combination of the above-described data. For example, the positioning data can include a combination of image data, feature points and descriptors, and / or a point cloud 510.
[0061] The sensing data 106 can include other combinations of different types of data, such as image data and depth data generated by a depth sensor (such as a sonar transceiver), a RADAR system, or a LiDAR system and representing relative positions of one or more objects in the environment. In this case, processing the sensing data 106 can include processing the image data to identify one or more objects in the image and identifying a first portion of the sensing data 106 that includes image data and depth data associated with the one or more objects. The AR system 100 can derive positioning data that includes any of the depth data and the image data, data representing feature points and descriptors, and / or data representing a point cloud, the first portion of the sensing data being obfuscated when the positioning data 110 is derived. For example, the depth data representing depths of the detected one or more objects can not be included in the positioning data 110.
[0062] As noted above, the sensory data 106 can include audio data representing sound captured from the environment. In this case, the method can include processing the sensory data 106 to identify a first portion of the sensory data 106 that represents speech or sound emitted by one or more artifacts in the environment, such as a machine. The localization data 110 including the audio data can then be derived with the first portion of the sensory data obfuscated, such as by deriving the localization data 110 from a second portion of the sensory data 106 that is different from the first portion of the sensory data 106, or by first deriving the localization data from all of the sensory data 106 and subsequently removing the portion of the localization data 110 derived from the first portion.
[0063] Figure 6 A non-transitory computer-readable storage medium 602 is shown that includes computer-readable instructions 606-612 that, when executed by one or more processors 604, cause the one or more processors 604 to perform the method shown in blocks 606-612 as described above and in Figure 6 FIG. 2. The examples and variations described above with respect to the method 200 of Figures 1 to 5 FIG. 2 also apply to the computer-readable instructions 606-612 included on the computer-readable storage medium 602.
[0064] Other examples are also contemplated in which an AR system that locally performs geo- pose determination can occasionally or periodically transmit localization data 110 to one or more remote computing devices in order to verify the accuracy of the geo-pose determination and, in some cases, subsequently correct and / or resynchronize the geo-pose determination of the AR system. In these cases, the method 200 can be applied such that the localization data 110 transmitted does not include data representing redundant information or that can be used to determine redundant information.
[0065] It should be understood that any of the features described with respect to any one example can be used alone or in combination with other features described, and any feature can be used in combination with one or more features of any other example or any combination of any other examples. Further, equivalents and modifications not described above are also within the scope of the following claims.
Claims
1. An augmented reality, AR, system comprising: one or more sensors arranged to generate sensing data representative of at least a portion of an environment in which the AR system is located; storage means for storing sensing data generated by the one or more sensors; one or more communication modules for transmitting positioning data to be used in determining a position and orientation of the AR system; and one or more processors arranged to: obtain sensing data representative of the environment in which the AR system is located; process the sensing data to identify a first portion of the sensing data representative of sensitive information relating to an object in the environment, wherein the object comprises an object assigned a sensitivity that varies according to the environment, and whether information relating to the object is considered to be in the first portion depends on the sensitivity assigned to the object in the environment in which the object is located; derive positioning data for determining a position and orientation of the AR system, wherein the positioning data is derived from the sensing data, and the first portion is obfuscated during derivation of the positioning data; and transmit at least a portion of the positioning data using the one or more communication modules.
2. The AR system of claim 1, wherein the positioning data is derived from a second portion of the sensing data, the second portion being a different portion than the first portion, and obfuscating the first portion of the sensing data comprises excluding the first portion of sensing data during derivation of the positioning data.
3. The AR system of claim 1, wherein the one or more processors are arranged to receive position data representative of a position and orientation of the AR system relative to the environment.
4. The AR system of claim 1, wherein the one or more sensors comprise any of: an image sensor; a microphone; a depth sensor; a sonar transceiver; a light detection and ranging, LiDAR, sensor; and a radio angle direction and ranging, RADAR, sensor.
5. The AR system of claim 1, wherein the sensing data comprises image data representative of an image of the environment, and processing the sensing data to identify the first portion of sensing data comprises identifying a portion of the image data representative of one or more objects in the image, and optionally, wherein the at least a portion of the positioning data transmitted comprises metadata identifying the portion of the image representative of the one or more objects, and optionally, wherein the at least a portion of the positioning data transmitted using the one or more communication modules comprises any of: data representative of a set of one or more feature points generated using the image data; and data representative of a point cloud generated using the image data. 6. The AR system of claim 5, wherein processing the sensory data to identify the first portion of sensory data comprises processing the image data to generate one or more segmentation masks associated with the one or more objects in the image, the one or more segmentation masks identifying the first portion of sensory data, and Optionally, wherein the one or more segmentation masks are generated by processing the image data using a neural network to identify the one or more objects in the image of the environment, and Optionally, wherein the segmentation masks represent portions of the image that include the one or more objects in the image.
7. A computer-implemented data processing method for an augmented reality (AR) system, the method comprising: obtaining sensory data representing an environment in which the AR system is located; processing the sensory data to identify a first portion of the sensory data representing sensitive information relating to an object in the environment, wherein the object comprises an object assigned a sensitivity that varies according to the environment, and whether information relating to the object is considered to be in the first portion depends on the sensitivity assigned to the object in the environment in which the object is located; deriving positioning data for determining a position and orientation of the AR system, wherein the positioning data is derived from the sensory data, and the first portion is obfuscated during derivation of the positioning data; and transmitting at least a portion of the positioning data.
8. The computer-implemented data processing method of claim 7, wherein the positioning data is derived from a second portion of the sensory data, the second portion being a different portion to the first portion, and obfuscating the first portion of sensory data comprises excluding the first portion of sensory data during derivation of the positioning data.
9. The computer-implemented data processing method of claim 7, comprising receiving position data representing a position and orientation of the AR system relative to the environment.
10. The computer-implemented data processing method of claim 7, wherein the sensory data comprises image data representing an image of the environment, and processing the sensory data to identify the first portion of sensory data comprises identifying a portion of the image data representing one or more objects in the image, and wherein Optionally, the at least a portion of the positioning data transmitted comprises metadata identifying the portion of the image representing the one or more objects, Optionally, wherein the at least a portion of the positioning data transmitted comprises any of: data representing feature points generated using the image data; and data representing a point cloud generated using the image data.
11. The computer-implemented data processing method of claim 10, wherein processing the sensory data to identify the first portion of sensory data comprises processing the image data to generate one or more segmentation masks associated with the one or more objects in the image, the one or more segmentation masks identifying the first portion of sensory data, and wherein Optionally, the one or more segmentation masks are generated by processing the image data using a neural network to identify the one or more objects in the image of the environment, Optionally, the one or more segmentation masks represent portions of the image that include the one or more objects.
12. A non-transitory computer-readable storage medium comprising computer-readable instructions that, when executed by at least one processor, cause the at least one processor to: obtain sensory data representing an environment in which an augmented reality (AR) system is located; processing the sensing data to identify a first portion of the sensing data that represents sensitive information related to an object in the environment, wherein the objects include objects assigned a sensitivity that varies according to the environment, and whether information related to the objects is considered in the first portion depends on the sensitivity assigned to the objects in the environment in which the objects are located; derive positioning data for determining a position and orientation of the AR system, wherein the positioning data is derived from the sensory data, and the first portion is obfuscated during derivation of the positioning data; and transmit at least a portion of the positioning data.
Citation Information
Patent Citations
Data processing
CN110633576A