Predictive Field of View (FOV) and Queues for Enforcing Compliance of Data Capture and Transmission with Real-Time and Near-Real-Time Video
By predicting the FOV of a video camera and controlling it to exclude unauthorized objects, the system enforces compliance in real-time data capture and transmission, addressing the limitations of existing technologies in secure environments.
Patent Information
- Application Number
- JP2024523577
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-21
- Filing Date
- 2022-10-19
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-10-19
AI Technical Summary
Existing video capture and processing systems fail to enforce compliance in real-time or near-real-time data capture and transmission in secure environments, as they either capture, store, and process unauthorized data or are incompatible with real-time applications due to processing requirements.
A method that predicts the field of view (FOV) of a video camera, recognizes unauthorized objects, and controls the camera to prevent their capture, ensuring that only authorized objects are included in the video signal, thereby enforcing compliance with data capture and transmission regulations.
This solution effectively prevents the capture and transmission of unauthorized data in real-time or near-real-time, ensuring compliance with security regulations and maintaining the integrity of sensitive information.
Smart Images

Figure 0007695029000001 
Figure 0007695029000002 
Figure 0007695029000003
Abstract
Description
Technical Field
[0001] Claim of Priority This patent application claims the benefit of priority to U.S. Patent Application No. 17 / 507,111, filed on October 21, 2021, which is hereby incorporated by reference in its entirety.
Background Art
[0002] The present invention relates to video capture and processing for enforcing compliance in the capture and transmission of data in a private, restricted, or secure environment in real-time or near real-time.
[0003] Description of Related Art Video camera technology is becoming even more widespread around the world today. For example, devices such as head-mounted cameras, robot control cameras, semi-autonomous or autonomous robots, mobile phones, desktop or table computers, near-eye displays, and handheld game systems may include cameras and related software that enable video capture, display, and transmission. These devices are used to provide one-way or two-way video communication in real-time or near real-time. Privacy and security concerns arise when video that should not be captured, saved, displayed, or transmitted is so done, either intentionally or unintentionally. The privacy of individuals, companies, or countries may be violated, sometimes illegally. In certain restricted environments, such as military or corporate proprietary or secure environments, there are strict controls governing what visual information can be captured, saved, displayed, or transmitted.
[0004] To limit the capture or transmission of unwanted videos, some existing systems monitor the video when it is captured. These systems use human processing, artificial intelligence (AI), computational algorithms, or combinations thereof to identify visual information of concern (such as a person's face or a company's proprietary information), and then remove or obscure that information from the video file data. In these systems, the recording device may even be shut off to prevent further capture of the information of concern. However, all of the existing systems described capture, store, and process the information of concern. The data of concern is stored (even if only temporarily) and processed, so the risk of data leakage still exists, and therefore these systems cannot meet the requirements of certain secure or restrictive environments. The processing required to remove or obscure information from a video file makes these systems incompatible with applications that require real-time or near-real-time video capture and transmission.
[0005] Video capture that enforces compliance with data capture and transmission in real-time or near-real-time may be required for various applications for individual users, companies, or countries. Such applications may include, but are not limited to, inspection / process review, supplier quality management, internal audits, equipment or system troubleshooting, factory operations, factory collaboration, verification and validation, equipment repair and upgrade, and equipment or system training. In these applications, it may be necessary to capture and transmit videos of local scenes that include information of concern or real-time or near-real-time, either unidirectionally or bidirectionally, in order to achieve efficient and effective communication. As a special case, compliance with data capture and transmission may be implemented in an extended reality environment.
[0006] Augmented Reality (AR) refers to the generation of two-dimensional or three-dimensional (3D) videographic or other media that are registered overlaid on objects in the surrounding environment. Artificial “markers,” also known as “sources,” with unique and easily distinguishable features are placed on users, objects, or scenes and can be used for various purposes. These markers are used to identify and locate specific objects, trigger the display of computer-generated media, or determine the user's position and orientation.
[0007] In certain video or AR environments, such as remote repair or inspection, concerns that are mainly held by customers and increased by the advancement of the video camera industry to maximize the Field of View (FOV) are that the user of the video being captured, transmitted, or displayed locally (either a field technician or an expert, but mainly a field technician) may inadvertently or intentionally look away from the object of interest and capture video of another part of the scene that should not be captured or transmitted. To avoid unintentional or intentional transmission of a wide FOV, compliance with some degree of data capture and transmission may be required by customer requirements, industry regulations, national security, or country-specific laws. Current technologies include physically covering the area around the object of interest with cloth or a waterproof sheet so that it is not captured in the video signal, mechanically narrowing the FOV, or isolating the video before transmission and having a security-certified domain expert review and edit the captured video signal. These are time-consuming activities. More generally and more costly is to move the equipment in question to a special secure location such as an empty garage or storage facility so that no site-unrelated items remain. In many cases, equipment removal, physical draping, or post-capture editing may not be sufficient to meet compliance requirements or may be unrealistic and costly to implement in a quasi-real-time interactive situation. Depending on the situation, there are country laws that prevent any kind of post-capture editing for reasons related to national security or ITAR - International Traffic in Arms Regulations.
[0008] U.S. Patents Nos. 10,403,046 and 10,679,425, "Field of View (FOV) and Key Code Limited Augmented Reality to Enforce Data Capture and Transmission Compliance," disclose enforcing alignment conditions between the orientation of a video camera and markers within a scene to avoid capturing data that is excluded in a real-time interactive situation. This can be done, for example, by determining whether the alignment conditions for markers within the local scene are met such that the orientation of the video camera is such that the FOV of the video camera is within a user-defined acceptable FOV around the marker. To meet the alignment conditions, another sensor can be used to detect the presence of the marker within the FOV of the sensor. The FOV of the camera or sensor can be reduced to create a buffer zone and further ensure that the FOV of the camera does not go outside the acceptable FOV. If the alignment conditions are not met, the video camera is controlled to exclude at least a portion of the camera FOV that is outside the user-defined acceptable FOV from capture within the video signal. For example, this can be done by turning off the power of the video camera or narrowing the FOV. The markers can also be used as a failsafe to ensure that images of particularly sensitive areas of the scene are not captured or transmitted. When another sensor detects these markers, the video camera is shut down. The system may signal the user, for example, where "green" means the alignment conditions are met, "yellow" means the engineer's eyes are starting to wander, and "red" means the alignment conditions are violated and the camera is disabled.In this system, by enforcing alignment conditions and using another sensor to detect other markers in a delicate area, especially in more stringent environments where compliance requires that part of the scene or tagged objects cannot be captured (detected) by the video camera itself and further output to the video signal, or in environments where real-time or near-real-time interactions are required, it is designed.
Summary of the Invention
[0009] The following is a summary of the present invention to provide a basic understanding of some aspects of the present invention. This summary is not intended to identify important or critical elements of the present invention or to describe the scope of the present invention in detail. Its sole purpose is to present some concepts of the present invention in a simplified form as a preamble to the more detailed description and definition of the claims that will be presented later.
[0010] The present invention provides a method for predicting the FOV of a video camera and realizes the enforcement of compliance for real-time and near-real-time video data capture and transmission.
[0011] To provide the enforcement of real-time or near-real-time video data capture and transmission compliance, the present invention predicts the future FOV of the video camera, recognizes unauthorized objects, and queues or controls the video camera to prevent their capture in the video signal. The predicted FOV can also be used to enforce the alignment conditions of the video camera for authorized objects.
[0012] In an embodiment, a three-dimensional ground truth map of a local scene including one or more objects is generated using a 3D sensor such as a 3D camera, LIDAR, or sonar. One or more of the objects in the ground truth map are identified as not permitted or potentially permitted. A 2D or 3D model may be used to represent the identified permitted and non-permitted objects. A video camera captures a video signal at a frame rate within the camera's field of view (CFOV) in the direction of the local scene. The system determines a pose including the position and orientation of the video camera within the local scene of the current frame, and calculates a predicted pose and a predicted FOV (PFOV) for one or more future frames from the pose of the current frame and measurements of the speed and acceleration of the video camera. The one or more predicted FOVs are compared with the ground truth map to recognize and identify non-permitted objects and objects that may be permitted. If a non-permitted object is recognized, the video camera is controlled to prevent the capture of the non-permitted object and its inclusion in one or more future frames of the video signal.
[0013] In certain embodiments, a queue can be generated to change the orientation of the video camera to prevent the capture of non-permitted objects and their inclusion in one or more future frames of the video signal. After generating the queue, the predicted FOV is updated and it is determined whether the updated predicted FOV contains non-permitted objects. If the queue cannot prevent the capture of non-permitted objects in the updated predicted FOV, the video camera is controlled to prevent the capture of non-permitted objects and their inclusion in the video signal. As a result, non-permitted objects are not included in the video signal and thus do not reach downstream circuits, processes, or networks to which the video camera is connected.
[0014] In various embodiments, an object may be defined and prohibited based on other attributes rather than just on content. For example, any object identified as being too close or too far from a video camera may be designated as not permitted. The movement speed (such as velocity and acceleration) of an object entering a local scene is defined as an object, and if its movement speed exceeds a maximum value, it may be defined as not permitted. Further, the movement speed (such as velocity and acceleration) of a video camera may be defined as an object and not permitted if its movement speed exceeds a maximum value.
[0015] In different embodiments, a video camera is trained on permitted objects and steered away from non-permitted objects. The system determines whether the current and predicted pointing directions of the camera satisfy an alignment condition for one of the permitted objects. If not, the system generates a queue to change the pointing direction of the video camera to enforce the alignment condition. If the queue cannot enforce the alignment condition, the video camera becomes inactive. A ground truth map can be used to verify recognized permitted objects.
[0016] In various embodiments, the pointing direction of the video camera is linked or controlled by the movement of the user (e.g., a head-mounted video camera), a manually manipulated user-controlled manipulator (e.g., a robotic arm), or a fully automated manually manipulated manipulator (e.g., an AI-controlled robotic arm, or a semi-autonomous or autonomous robot). For example, before capturing a non-permitted object, an audio, video, or vibratory cue may be presented to the user to deflect the user's head away from the non-permitted object or to return to a permitted object before violating an alignment condition. Similarly, a cue can be presented to the user of the robotic arm to deflect the camera or return to a permitted object before capturing a non-permitted object. In a fully automated system, the cue may disable the system and cause the camera to point in a direction away from the non-permitted object or in the direction of the permitted object.
[0017] In different embodiments, the video camera can be controlled to prevent the capture of non-permitted objects in various ways. The video camera can be mechanically controlled to change the pointing direction, optically controlled to narrow the CFOV, blurred before the non-permitted object reaches the detector array (e.g., by changing the f / #), changing the illumination of the local scene to blind the sensor, electrically controlled to turn off the power, disabling the electrochemical layer, amplifier, or A / D converter of the ROIC of the detector array, or selectively turning off the pixels of the non-permitted object.
[0018] These and other features and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the preferred embodiments, taken in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
DETAILED DESCRIPTION OF THE INVENTION
[0020] Video capture that enforces compliance with data capture and transmission in real time or near real time may be required for a variety of applications for individual users, enterprises, or countries. Such applications may include, but are not limited to, inspection / process review, supplier quality management, internal auditing, equipment or system troubleshooting, factory operations, factory collaboration, verification and validation, equipment repair and upgrade, and equipment or system training. In these applications, in order to achieve efficient and effective communication, it may be necessary to capture problem information or video of the local scene including real time or near real time and transmit it unidirectionally or bidirectionally. As a special case, compliance with data capture and transmission may be implemented in an extended reality environment. The pointing direction of the video camera is linked or controlled by the movement of the user (e.g., head-mounted video camera or handheld video camera), a manually controlled manipulator by the user (e.g., robotic arm), or a fully automated manually controlled manipulator (e.g., AI-controlled robotic arm, or semi-autonomous or autonomous robot).
[0021] The present invention is directed to these and other similar applications where compliance with respect to the capture and transmission of some data may be required by customer requests, industry regulations, national security, or country-specific laws. In some cases, for compliance, it may be necessary to prevent a video camera from including in the video signal output for display or transmission a portion of a scene or an object tagged with a particular tag. In more stringent environments, for compliance, it may be required that a portion of a scene or a tagged object cannot be stored in the camera's memory chip and further cannot be output to the video signal. The memory chip may be the memory chip alone or may be a video display or video transmission chip that includes persistent memory. The required level of compliance is determined by various factors and may change during the interval between the capture and display or transmission of the video signal, or during that interval.
[0022] The present invention predicts the FOV of a video camera, recognizes unauthorized objects, controls the video camera to prevent the capture of unauthorized objects and their inclusion in one or more future frames of the video signal. The predicted FOV can also be used to enforce the alignment conditions of the video camera for authorized objects. A queue can be used to prompt a correction of the pointing direction of the video camera to prevent the capture of unauthorized objects before it occurs or to enforce it before the alignment condition is lost. If no corrective action is taken by the queue, the video camera is controlled to prevent the capture of unauthorized objects or to penalize the loss of the alignment condition. As a result, unauthorized objects are not included in the video signal and thus do not reach downstream circuits, processing, or networks to which the video camera may be connected. This can be implemented in real-time or near real-time, or at a slower rate if the application does not require such performance, by placing a delay line or a temporary memory chip between the ROIC of the video camera and the memory chip. For example, the slowest video frame rate acceptable to most users is about 24 frames per second (fps), i.e., about 42 milliseconds (ms). A time delay of less than 42 ms is generally acceptable to most users. High-speed video cameras are 120 fps, i.e., about 8 ms. At these frame rates, a one-frame delay is certainly real-time. Utilizing the predicted FOV can enforce compliance in the capture and transmission of data for a single image or image sequence within the video signal.
[0023] Referring now to FIG. 1, in one embodiment, a video capture, display, and transmission device 10, such as a video goggle or a handheld unit (such as a tablet or a mobile phone), has a pointing direction 12 that is linked to the movement of the technician (e.g., whether the on-site technician 14 is looking at the unit or pointing at the unit). The device 10 includes a sensor 16 (e.g., a 3D camera, LIDAR, sonar) configured to capture signals sensed within the FOV 18 of sensors around the pointing direction 12, and a video camera 20 (e.g., a 2D or 3D CMOS, CCD, or SWIR camera) configured to capture light within the camera FOV 22 around the pointing direction 12 and within the sensor FOV 18 to form a video signal. The sensor (16) is used to generate a 3D ground truth map 24 of the local scene, in which certain objects 26 and 28 are shown as permitted objects and objects 30 and 32 are shown as non-permitted objects. The objects may be represented in the ground truth map as actual sensor data or as computer-generated models.
[0024] In this example, the on-site technician 16 may move within the manufacturing facility to confirm the presence and location of certain objects, repair or maintain certain objects, or use certain objects. These objects may be considered "permitted" objects. The technician may even be prompted or instructed to maintain the pointing direction 12 with respect to the designated objects to perform certain tasks (verification, repair, use, etc.). The on-site technician 16 can capture, display, and transmit "permitted" objects. The on-site technician 16 cannot capture "non-permitted" objects in memory, let alone display or transmit them.
[0025] Allowed objects or non-allowed objects may depend on the use, the type of objects within the local scene, the authority of on-site technicians, supervisors, and any remote staff, and any specific restrictions applicable to the compliance of excluded data and transmission. This information may be stored in the device. The device processes the sensed signals to detect objects, identify their positions, and determine whether they are allowed objects or non-allowed objects. The background of the local scene and the ground truth map are permitted or set as not permitted by default. When an object that does not exist in the ground truth map moves into the scene, the object is not properly permitted until it is identified and marked as permitted. Similarly, an unrecognized object is not properly permitted until it is identified and marked as permitted. Instead of being based on exactly that content, objects may be defined and not permitted based on other attributes. For example, any object identified as being too close (< minimum distance) or too far (> maximum distance) from a video camera may be designated as not permitted. The moving speed (such as speed and acceleration) of an object entering the video camera or the local scene is defined as an object, and if its moving speed exceeds the maximum value, it may be defined as not permitted. This device compares the recognized objects with the ground truth map to verify whether they are the same object, whether they are permitted or not permitted, and their locations, greatly improving the accuracy and reliability of object recognition.
[0026] To prevent the capture and transmission of excluded data, one or more predictive camera FOVs 35 for a video camera are utilized to recognize unauthorized objects and control the video camera to interrupt and stop the transfer of the image to the memory chip 34 where the video signal is formed. For example, if the current camera FOV 22 and the pointing direction 12 for one or more predictive camera FOVs 35 meet the alignment condition (e.g., the pointing direction is within a few degrees from the recommended line of sight (LOS)) for an authorized object 26 to perform some task and do not include any unauthorized objects 30 or 32, the image captured by the video camera is transferred to the memory chip 34, where it is formed into a video signal and can be displayed to a field technician or transmitted remotely (e.g., for storage or display to another remote user).
[0027] If both conditions are met, the device can generate a positive cue (e.g., a green "Good") to enhance the user's focus on the authorized object. If the user's pointing direction starts to move away from the authorized object or starts to move towards an unauthorized object but has not yet violated either condition, the device can generate a prompt cue (e.g., a yellow "Move left") for the user to take corrective action. If the user's pointing direction changes to a point where it violates the alignment condition or the capture of an unauthorized object is imminent, the device can control the video camera to prevent the capture of the unauthorized object and its inclusion in the video signal, or deactivate the camera and issue an interrupt cue (e.g., a red "Deactivate video camera").
[0028] If any condition is violated, the device controls the camera to prevent the capture of video signals containing unauthorized objects, or issues an "interrupt" 36 when the alignment conditions are not met. When the alignment conditions are violated, the video camera is typically turned off by cutting the power to the video camera, deactivating the electrochemical top layer of the detector array or ROIC, or pointing the video camera in a completely different direction. In the case of a violation that moves to capture an unauthorized object, in addition to these options, the video camera can be controlled to optically narrow the camera FOV, selectively blur a portion of the camera FOV (e.g., by changing the f / #), change the illumination of the local scene to blind the sensor, or selectively turn off or blur the pixels of the detector array corresponding to the unauthorized object.
[0029] The same method can also be applied to a robotic arm for remotely controlling the direction of the video camera and a fully autonomous robot that uses the video camera as part of a vision system. In the case of a robotic arm, the "time delay" can ensure that protected data is not captured and sent to a remote site or other location where the user is present. In the case of a fully autonomous robot, the "time delay" can ensure that protected data is not captured or used by the robot or sent to other locations.
[0030] This method can be applied to applications where only authorized objects are present or to local scenes where only unauthorized objects are present.
[0031] Referring now to FIGS. 2, 3, and 4, in one embodiment, the video capture and display device 100 includes a video camera 102, a sensor 104, and an interrupt processor 106 that controls the video camera to prevent unauthorized objects 132 from being captured in the video signal and to enforce alignment conditions on authorized objects 130. The device 100 is coupled to a "platform" 107 such as a user, robotic arm, robot, etc. that controls the orientation of the device.
[0032] The sensor 104 includes a 3D sensor 124 such as a 3D camera, LIDAR, sonar, etc., and a sensor processor 126 that captures sensed signals within the sensor FOV (SFOV) 128 along the orientation direction 110 and forms a three-dimensional ground truth map 127 in which one or more authorized or unauthorized objects are identified. The identified objects may be represented by the sensed signals or a computer-generated model. As described below, the sensor can also be used to determine the pose of the camera. The sensor 104 is separated from a video processor 144 or any other circuit or network downstream of the video capture and display device 100 to which the video camera may be connected to distribute the video signal.
[0033] The video camera 102 captures light within the camera's field of view (CFOV) 108 along the orientation direction 110 in the local scene. The video camera appropriately includes a power supply 111, an optical system 112 that collects light within the CFOV, a detector array 114 that senses and integrates light and converts photons to electrons to form an image, a readout integrated circuit (ROIC) 116 that includes an amplifier and an A / D converter for reading out an image sequence at a frame rate, and a memory chip 118 that stores the image sequence and forms a video signal 120.
[0034] The interrupt processor 106 determines a pose 158 that includes the position and orientation of the video camera within the local scene of the current frame 159. This can be done by using a gyroscope 150 that measures the 6DOF pose of the video camera (e.g., x, y, z, and rotations about the x, y, z axes), or by correlating the signals sensed in the sensor FOV 128 of the current frame with the ground truth map 127. The interrupt processor 106 calculates the predicted pose 160 and predicted FOV (PFOV) 162 of one or more future frames 163 from the pose of the current frame and the measurements of the velocity and acceleration of the video camera obtained by the motion sensor 152, suitably from 6DOF. The interrupt processor 106 compares one or more predicted FOVs 162 with the ground truth map to recognize and identify unauthorized objects 132 and potentially authorized objects 130.
[0035] Due to the predictive nature of the system, the interrupt processor 106 can change the pointing direction of the video camera without controlling or turning off the video camera, to prevent the capture of unauthorized objects and the generation of a queue 140 for preventing it from being included in one or more future frames of the video signal. The queue 140 is configured to prevent the video camera from moving towards unauthorized objects before it occurs. After generating the queue, the interrupt processor updates one or more predicted FOVs and determines whether the updated predicted FOV contains unauthorized objects. If the queue cannot prevent the capture of unauthorized objects in the updated predicted FOV, the interrupt processor 106 issues an interrupt 136 to control the video and prevent the capture of unauthorized objects and their inclusion in the video signal.
[0036] The video camera is trained on the permitted object 130 and, when away from the non-permitted object 132, the interrupt processor 106 determines whether the current and predicted pointing directions of the camera satisfy the alignment conditions for one of the permitted objects. If not, the system generates a queue 140 that changes the pointing direction of the video camera to enforce the alignment conditions. If the queue cannot enforce the alignment conditions, the video camera becomes inactive. Losing the alignment conditions does not necessarily mean that the camera captures a non-permitted object. However, if the video camera deviates from a permitted object and the queue cannot correct the problem, turning off the video camera at least temporarily is effective for the platform to maintain proper alignment with the permitted object and to be trained to perform the tasks at hand. The length of time the video camera is turned off can be changed to more effectively train local or remote users or robots.
[0037] In various embodiments, the objects may be defined and prohibited based on other attributes rather than just on content. For example, any object identified as being too close or too far from the video camera may be designated as non-permitted. The movement speed of the video camera (such as speed and acceleration) may be defined as an object and non-permitted if its movement speed exceeds a maximum value.
[0038] The video processor 144 processes the video signal 120 to form a video output signal 146 for display 148 or for transmission. If no unauthorized object is detected, the video output signal 146 is a normal video signal. When an interruption is issued, the video output signal 146 may not receive a video signal if the power is interrupted, may receive a blank or noisy video signal if the ROIC is deactivated, or may receive a video signal in which pixels corresponding to unauthorized objects are deleted or hidden.
[0039] Referring now to FIG. 5, one way to determine the 6DOF pose of a video camera is to use a sensor 200 that shares the pointing direction 202 with the video camera 204. As shown in this example, a pair of 3D sensors 202 are placed on both sides of the video camera 204 to surround and align with the camera FOV 208, providing a sensor FOV 206. During the operation and control of the video camera 204, the sensors 202 capture 3D sensing signals and form a self-generated map 210 in the sensor FOV 206. An interrupt processor collates the 3D sensing signals with a 3D ground truth map 212 (generated by the same sensor) to determine the pose 214 of the current frame of the video camera. It may be possible to localize itself using a Simultaneous Localization and Mapping (SLAM) algorithm (or similar), and then align the ground truth map 212 with the self-generated map 210 with reference to the self-generated map 210 to determine the exact position and orientation.
[0040] In an AR environment, the pointing direction of the video camera is linked to the movement of the on-site technician (the head of the technician in the case of goggles, the hand in the case of a handheld unit, etc.). The video signal is captured within the FOV of the objects in the local scene at a distance of the length of the technician's arm from the technician. The technician receives a hand gesture for operating the object from an expert at a remote location, registers it, and superimposes it on the video signal to create an augmented reality that instructs the user to operate the object. In various implementation forms, when the expert views the video signal captured by the technician's video camera and transmitted in real time to the remote site and responds, when the expert interacts with the replica of the object in real time, or when the expert generates an "off-the-shelf" instruction offline by responding to the video or interacting with the replica of the object, hand gestures are provided.
[0041] What is a concern is that there is a possibility that the technician may intentionally or unconsciously look away from the target object and capture the video of another part of the scene or object that should not be captured or transmitted. The present invention automatically controls the video camera in such a restricted AR environment under the control of the technician / customer / master ("user"), and provides a system and method for excluding a part of the scene for compliance with data capture and transmission without interfering with the AR overlay that instructs the technician to operate the object.
[0042] Referring to FIGS. 6 and 7, an embodiment of the AR system 310 includes a video capture, display, and transmission device 312 such as video goggles or a handheld unit (such as a tablet or mobile phone), and its pointing direction 314 is linked to the movement of the technician (for example, whether the field technician 316 is looking at the unit or pointing at the unit). The field technician 316 operates an object 318 within the local scene 320, for example, to perform maintenance on the object or to receive instructions regarding how to operate the object. In this example, the object 318 is an access panel, and the local scene includes a tractor. The access panel is a "permitted" object. For the sake of explanation, the cabin 319 is a "not permitted" object.
[0043] The device 312 includes a 3D sensor 313 (such as a 3D camera, LIDAR, or sonar) that captures sensed signals, and a video camera 315 that captures an image of the object 318 in the local on-site scene 320 and transmits a video signal 326 via a communication link 328 or network to a remote site, possibly in another country. The sensor 313 and the sensed signals are separated from the communication link 328 or network. The video signal 326 is presented to an expert 330 on a computer workstation 332 equipped with a device 334 for capturing hand gestures of the expert. The expert 330 operates the object in the video (either by hand or via a tool) to perform a task. The device 334 captures the hand gesture, converts it into an animated hand gesture 336, and transmits it to the user 316 via the communication link 328 for registration and overlay on the display. The expert 330 may provide instructions in the form of voice, text, or other media, in addition to or instead of hand gestures, to support the AR environment. The AR environment itself is implemented using application software 340 and configuration files 342 and 344 on remote and local computer systems 332, 312, and a server 346.
[0044] According to the present invention, an AR system or method determines whether the orientation direction 314 of the local video capture and display device 312 satisfies the alignment condition 350 for the permitted object 318. The alignment condition 350 associates the orientation direction 314 of the camera with the line of sight (LOS) 355 to the permitted object. Both need to maintain a specified deviation 357 given as an angle or distance with respect to the permitted object 318. The deviation 357 is fixed, for example, at plus or minus 5 degrees and set by the user.
[0045] The key code 360 of one or more technicians / customer, master, or expert can be used to identify the technician / master / expert, define the pairing of permitted objects, and specify the tolerance values that define the alignment condition 350. With the key code, the on-site technician / customer / master or expert can control the video camera and prevent the on-site technician from capturing and / or transmitting data that violates the customer's or country's policies or legal requirements. With the key code, the owner of the code can, for example, via the GUI, specify permitted or non-permitted objects or pre-load their selections for specific local scenes or tasks. To establish the alignment condition, the user can specify a distance of 24 inches, creating a circle with a radius of 24 inches centered on the LOS to the permitted object. Alternatively, the user can also specify a deviation of plus / minus 5 degrees with respect to the pointing direction, which may create a circle with a radius of, for example, 12 inches with respect to the working distance of the nominal arm length. Instead of or in addition to specifying a distance or angle, the user can specify and place additional markers (non-permitted objects) within the scene around the permitted object that define the outer boundary or tolerance values of the alignment condition. To activate the video capture and display device 312, at least a successful pairing with the identified permitted object is required. Another key code managed by the expert may be provided to the on-site technician at a remote location to prevent non-compliant data from being transmitted and received remotely. This key code is enforced in the on-site technician's AR system. In addition to the distance, the key code may also have a known set of permitted or non-permitted shapes pre-entered.
[0046] The 3D sensor 313 captures the signals sensed within the sensor FOV 354 and generates a 3D ground truth map in which the permitted objects 318 and non-permitted objects 319 are shown and identified. The video camera 315 captures the light within the camera FOV 356, and the light is read out and stored as an image in a memory chip, where the image is formed into a video signal. The interrupt processor updates the pose of the video camera for the current frame, calculates the predicted poses 359 of one or more predicted camera FOVs 361 for future frames, and compares those predicted camera FOVs with the ground truth map to recognize and identify permitted 318 or non-permitted 319 objects. The interrupt processor uses the detected sensed signals or gyroscope measurements, etc., to determine whether the pointing direction 314 meets the alignment condition 350 for the permitted object 318, and whether the FOV 356 of the video camera or the predicted camera FOV 361 contains or will soon contain a non-permitted object 319.
[0047] Regarding the satisfaction of the alignment conditions for the permitted object 318, as long as the technician is looking at the direction of the permitted object 318, the permitted object is captured in the camera's FOV 356, the alignment conditions are met (green "Good"), enabling the camera to capture an image and form a video signal. When the technician's eye (the camera) begins to wander, the permitted object approaches the edge of the alignment condition 350 (yellow "Correct left" or "Correct right"), and a cue (audio, video, vibration) is issued to prompt the technician to continue focusing on the permitted object 318 and the task at hand. The camera remains active, and the video is captured for transmission and display. If the technician's LOS wanders too far, the predicted camera FOV 361 loses pairing with the permitted object 318, thereby violating the alignment condition 350, deactivating the video camera, and a cue (red, "Deactivate camera") is issued. The length of time the video camera is deactivated can be controlled to "train" the technician to continue focusing on the permitted object.
[0048] Regarding preventing the capture of the non-permitted object 319 in the video signal, if the non-permitted object 319 is captured in the predicted camera FOV 361, an interrupt can be issued to control the video camera by issuing a similar cue, and any camera image containing the non-permitted object 319 is not transferred to the memory chip and not included in the video signal. This includes cutting off the power to the video camera, disabling the top electrochemical layer of the detector array, or deactivating the ROIC, dumping or editing the time-delayed image before it is transferred to the memory chip. A cue such as red "Please deactivate the camera" can be issued to the technician to notify the technician of the violation regarding the non-permitted object and prompt the technician to change the LOS.
[0049] Although some exemplary embodiments of the present invention have been shown and described, those skilled in the art will conceive of numerous variations and alternative embodiments. Such variations and alternative embodiments are contemplated and may be made without departing from the spirit and scope of the present invention as defined in the appended claims.
Claims
1. A method for preventing the capture and transmission of data excluded from a local scene from a video signal, the method comprising: Generating a ground truth map of the local scene including one or more objects; Identifying one or more of the objects in the ground truth map as not permitted; Capturing a video signal at a frame rate within the camera's field of view (CFOV) in the direction of the camera within the local scene using a video camera; Determining a pose including the position and orientation of the video camera within the local scene of the current frame; Receiving measurements of the velocity and acceleration of the direction of the video camera; Calculating one or more predicted fields of view (PFOVs) for one or more future frames from the pose of the current frame and the measurements of the velocity and acceleration; Comparing the one or more predicted FOVs with the ground truth map to recognize and identify non-permitted objects, and Controlling the video camera to prevent the capture of non-permitted objects and their inclusion in the one or more future frames of the video signal if a non-permitted object is recognized.
2. The ground truth map is three-dimensional, and The method according to claim 1, further comprising generating the three-dimensional ground truth map of the local scene using a sensor.
3. The method according to claim 2, wherein the sensor is one of a 3D video camera, LIDAR, or sonar.
4. The method according to claim 2, wherein the ground truth map includes at least one computer-generated 2D or 3D model of the non-permitted objects.
5. Using the sensor to capture signals sensed within a three-dimensional sensor FOV (SFOV) along the pointing direction, The method according to claim 3, further comprising comparing the signals sensed within the three-dimensional sensor FOV with the three-dimensional ground truth map to determine the pose of the current frame.
6. Delivering the video signal to a network, and Separating the sensor from the network, the method according to claim 2.
7. Measuring the moving speed of an object entering the local scene or the moving speed of the video camera relative to the local scene, and The method according to claim 1, further comprising deactivating the video camera when the measured moving speed exceeds a maximum value.
8. The method according to claim 1, wherein a gyroscope of the video camera determines the pose of the current frame.
9. When a non-permitted object is recognized, Generating a queue to change the pointing direction of the video camera to prevent capture of the non-permitted object and its inclusion in one or more future frames of the video signal, After generating the queue, updating the predicted FOV and determining whether the updated predicted FOV contains the non-permitted object, If the queue cannot prevent capture of the non-permitted object within the updated predicted FOV, controlling the video camera to prevent capture of the non-permitted object and its inclusion in the video signal, the method according to claim 1.
10. The method according to claim 9, wherein one of a user, a user-controlled robot, or an autonomous robot responds to the queue for changing the pointing direction of the video camera.
11. One or more of the objects of the ground truth map are permitted, Determining whether the current and predicted pointing directions satisfy alignment conditions for the permitted objects, If the current and predicted pointing directions do not satisfy the alignment conditions, generating a queue for changing the pointing direction of the video camera to enforce the alignment conditions, After generating the queue, updating the current and predicted pointing directions to determine whether they satisfy the alignment conditions for the permitted objects, and If the queue cannot enforce the alignment conditions, deactivating the video camera, the method according to claim 1.
12. The local scene is an augmented reality (AR) environment, Receiving instructions in the form of audio, text, or video media related to the permitted objects from a remote location, and Registering the instructions in the video signal and displaying the augmented reality to the user to instruct the user, further comprising The pointing direction of the video camera is linked to the movement of the user, The method according to claim 11, wherein the queue prompts the user to change the user's pointing direction to enforce the alignment conditions.
13. The video camera includes a detector array having an electrochemical layer that senses and integrates light to form an image, a readout integrated circuit (ROIC) having an amplifier and an A / D converter that reads out an image sequence at a frame rate, and a memory chip that stores the image sequence to form the video signal. The step of controlling the video camera to prevent the capture of unauthorized objects includes mechanically controlling the video camera to change the pointing direction, optically controlling the video camera to narrow the CFOV or blur the unauthorized object, and electrically controlling the video camera to cut off the power, disable the electrochemical layer, amplifier, or A / D converter, or selectively turn off the pixels of the unauthorized object, the method according to claim 1 comprising one or more of.
14. The video camera includes a detector array having an electrochemical layer that senses and integrates light to form an image. The step of controlling the video camera to prevent the capture of unauthorized objects includes disabling the electrochemical layer, the method according to claim 1.
15. The step of controlling the video camera to prevent the capture of unauthorized objects includes selectively turning off the camera pixels of the unauthorized object, the method according to claim 1.
16. A method for preventing the capture and transmission of data excluded from a local scene from a video signal, the method comprising: generating a ground truth map of the local scene including one or more objects; identifying one or more of the objects in the ground truth map as unauthorized; using a video camera to capture a video signal at a frame rate within the field of view (CFOV) of the camera in the pointing direction within the local scene; Determining a pose including the position and orientation of the video camera within the local scene of the current frame, Receiving measurements of the velocity and acceleration of the pointing direction of the video camera, Calculating one or more predicted pointing directions and a predicted FOV (PFOV) for one or more future frames from the pose of the current frame and the measurements of the velocity and acceleration, Comparing the predicted FOV with the ground truth map to recognize and identify non-permitted objects, and Generating a first queue for changing the pointing direction of the video camera to prevent the capture of non-permitted objects and their inclusion in the one or more future frames of the video signal if a non-permitted object is recognized. A method comprising.
17. The ground truth map includes one or more permitted objects, After generating the queue, updating the predicted FOV to determine whether the updated predicted FOV includes the non-permitted object, The method according to claim 16, further comprising controlling the video camera to prevent the capture of the non-permitted object and its inclusion in the video signal if the first queue cannot prevent the capture of the non-permitted object within the updated predicted FOV.
18. The ground truth map includes one or more permitted objects, Determining whether the current and predicted pointing directions satisfy alignment conditions for permitted objects, Generating a second queue for changing the pointing direction of the video camera to enforce the alignment conditions if the current and predicted pointing directions do not satisfy the alignment conditions, After generating the first or the second queue, updating the current and predicted pointing directions and the predicted FOV to determine whether they meet the alignment conditions for the permitted objects and whether the updated predicted FOV contains the non-permitted objects, and The method according to claim 16, further comprising deactivating the video camera if either the first queue or the second queue fails to enforce the alignment condition, or if capture of the non-permitted object cannot be prevented. **Claim 19** A method for preventing capture and transmission of data excluded from a local scene in a video signal, the method comprising: Generating a ground truth map of the local scene including one or more objects; Identifying one or more of the objects in the ground truth map as permitted; Capturing a video signal at a frame rate within the field of view (CFOV) of a camera in the pointing direction within the local scene using a video camera; Determining a pose including the position and orientation of the video camera within the local scene of the current frame; Receiving measurements of the velocity and acceleration of the pointing direction of the video camera; Calculating one or more predicted pointing directions for one or more future frames from the pose of the current frame and the measurements of the velocity and acceleration; Determining whether the current and predicted pointing directions meet the alignment conditions for the permitted objects, and Generating a first queue for changing the pointing direction of the video camera to enforce the alignment condition if the current and predicted pointing directions do not meet the alignment conditions. **Claim 20** After generating the first queue, updating the current and predicted pointing directions to determine whether they meet the alignment conditions to the permitted object, and The method according to claim 19, further comprising deactivating the video camera when the first queue cannot enforce the alignment condition.
Citation Information
Patent Citations
Monitor camera
JP2004179971A
Object tracking method and object tracking apparatus
JP2005057743A
Automatic imaging apparatus
JP2009147647A
Video management device, video management method and program
JP2017118168A
System and method used for playback of panoramic video content
JP2017528947A