Method for adding annotations and interfaces to control panels and screens in augmented reality applications

The AR system addresses remote troubleshooting challenges by generating interactive overlays and recorded tutorials, improving guidance clarity and reducing errors and costs in device installation and maintenance.

JP7838612B2Active Publication Date: 2026-04-01FUJIFILM BUSINESS INNOVATION CORP
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing AR technologies for remote assistance in troubleshooting electronic devices face challenges such as voice-only communication errors, time-consuming on-site visits, and difficulties in real-time video streaming due to user movement, especially when both hands are needed for device operation.

Method used

An AR system that generates interactive overlays on device screens using anchor images for stable video streaming, allowing remote experts to guide users through steps asynchronously, with features like automatic detection of occlusions, perspective correction, and recorded tutorials.

Benefits of technology

Enhances remote troubleshooting by providing clear, hands-free guidance, reducing errors and costs, and enabling users to complete tasks independently with recorded procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838612000001
    Figure 0007838612000001
  • Figure 0007838612000002
    Figure 0007838612000002
  • Figure 0007838612000003
    Figure 0007838612000003
Patent Text Reader

Abstract

To provide methods for generating an overlay display to which annotations can be added on a screen using an augmented reality technique.SOLUTION: A method of the present invention includes: pausing a video received for display on a second device from a first device on the second device; while pausing the video received from the first device on the second device, accepting input including annotations for the paused video on the second device; and generating an augmented reality (AR) overlay corresponding to a portion of the paused video on a display of the first device in response to input made to the portion of the paused video on the second device if a pause duration of the input exceeds a threshold.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to augmented reality (AR) systems, and more specifically, to systems and methods for generating control panels and screen interfaces that can be used with AR.

Background Art

[0002] In related technologies, AR applications have been implemented that provide an interface so that a user can operate a vehicle dashboard or a stereo system. In other applications, AR can be utilized during an Internet browsing session to add an overlay for assisting in navigating the Internet to a web page.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 1

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

[0005] The present invention aims to provide a method for generating an overlay display on a screen that can be annotated, using augmented reality technology. [Means for solving the problem]

[0006] The embodiments described herein concern AR implementations that can stream a tuned display of a display (e.g., a computer screen, a touch liquid crystal display (LCD), a digital control panel, or a control panel for an electronic device) with live (real-time) or automated agent-added overlays that guide the user through steps (e.g., which buttons to click or tap on the screen, where to enter text, etc.). Embodiments include registration, i.e., detecting the boundaries of target objects so that the AR overlay is properly displayed on the screen even if the user moves the camera. In another embodiment, marks can be created based on underlying content and automatically removed when an action is performed. In yet another embodiment, automatic detection of occlusions is performed to display the instruction overlay in a realistic manner. Finally, an automated process can take existing video material (e.g., how-to videos for the LCD display of an electronic device such as a multi-function device (MFD)) and extract anchor images to be used to initiate the registration process.

[0007] In this embodiment, the AR interface can be extended to live remote assistance tasks, allowing remote experts to connect with people sharing live streams from mobile or head-mounted devices to diagnose and resolve on-site problems. With the advent of live streaming services, live remote assistance is becoming a way for users to troubleshoot increasingly difficult problems. While background technology tools focus on allowing remote users to annotate or demonstrate solutions, they do not take into account the time and effort users need to expend following instructions. Often, instructions need to be repeated until the user fully understands, and in some cases, both hands are needed to operate physical equipment, making live video streaming from mobile devices difficult. To address these issues, this embodiment makes available an AR-based tool for a remote assistance interface that can automatically record steps during a live stream that users can view asynchronously.

[0008] Aspects of the present disclosure may include a method that stabilizes video received from a first device for display on a second device, and generates an augmented reality (AR) overlay on the display of the first device corresponding to the stabilized portion of video in response to an input made to a portion of the video stabilized on the second device.

[0009] Fixing the video received from the first device for display on the second device may include identifying one or more anchor images in the video, determining a target object which is a two-dimensional surface based on the identified one or more anchor images, and performing perspective correction of the video based on the target object which is a two-dimensional surface for display on the second device.

[0010] Identifying the one or more anchor images in the video may include searching a database for one or more anchor images that match one or more images in the video.

[0011] Identifying the one or more anchor images in the video may include detecting a quadrilateral on the video received from the first device, matching a three-dimensional plane to the two-dimensional points of the detected quadrilateral, tracking the three-dimensional plane that matches the two-dimensional points of the detected quadrilateral, and receiving a selection of the one or more anchor images in the video via the second device.

[0012] The method may further include cropping the target object in the video for display on the second device.

[0013] The target object may be a display screen.

[0014] Generating the AR overlay corresponding to a portion of the fixedly displayed video on the display of the first device may be performed live in response to an input made to a portion of the fixedly displayed video on the second device.

[0015] Generating the AR overlay corresponding to a portion of the fixedly displayed video on the display of the first device may include generating the AR overlay on the display of the second device in response to the input, receiving an instruction to provide the AR overlay to the first device, and transmitting and displaying the AR overlay on the first device when the instruction is received.

[0016] The method may further include tracking one or more hands and fingers in the video and occluding portions of the AR overlay that overlap the one or more hands and fingers on the display of the first device.

[0017] Fixing the video received from the first device for display on the second device includes pausing the video on the display of the second device, and the input may include an annotation.

[0018] Generating an augmented reality (AR) overlay corresponding to a portion of the fixed video on the display of the first device may include generating a video clip including the annotation when the pause of the input exceeds a timeout threshold, and providing the video clip to the display of the first device.

[0019] Further, when it is determined that the first device is placed, playing the video clip on the display of the first device, and when it is determined that the first device is in the user's hand, providing the video for display on the second device.

[0020] The AR overlay may include an instruction for moving the cursor of the first device from a first location to a second location.

[0021] The method may further include saving the AR overlay for playback by the first device.

[0022] Aspects of the present disclosure include a non - transient computer - readable medium storing instructions for performing a process, the instructions including fixing a video received from a first device for display on a second device, and generating an augmented reality (AR) overlay corresponding to a portion of the video fixed on the second device on the display of the first device in response to an input made to the portion of the video fixed on the second device.

[0023] Aspects of the present disclosure include a system comprising means for fixing a video received from a first device for display on a second device, and means for generating an augmented reality (AR) overlay on the display of the first device corresponding to a portion of the fixed video in response to an input made to a portion of the fixed video on the second device. [Brief explanation of the drawing]

[0024] [Figure 1] An exemplary flow for overlaying at least one of an AR interface and / or annotations on a screen or control panel is shown according to the embodiment. [Figure 2] An example of an overlay on the device panel captured from a user device is shown in the embodiment. [Figure 3] An example of a three-dimensional overlay node with transitions is shown using an embodiment. [Figure 4] The panel shown here has been perspective-corrected according to the example. [Figure 5] This shows an example of an overlay where a hand and finger mask is implemented so that the overlay is positioned below the hand or fingers. [Figure 6] Examples of recording and playback of an AR interface are shown using the embodiments. [Figure 7] A flowchart of the annotation and recording process is shown according to the example. [Figure 8] An example of a computer device according to the embodiment is shown. [Modes for carrying out the invention]

[0025] The following detailed description provides details of the drawings and embodiments of this application. Descriptions of elements that overlap between reference numerals and figures have been omitted for clarity. Terms used throughout this description are provided as examples and are not intended to be limiting. For example, the use of the term “automatic” may include fully automatic or semi-automatic implementations, depending on the desired implementation of those skilled in the art, including user or administrator control over specific aspects of the implementation. Selections may be made by the user via a user interface or other input means, or through a desired algorithm. The embodiments described herein may be used individually or in combination, and the functions of the embodiments may be implemented through any means according to the desired implementation.

[0026] There are several potential challenges in remotely assisting customers in troubleshooting advanced electronic devices such as MFDs. For example, voice-only communication is prone to errors, and sending a service engineer to the customer's location can be time-consuming and costly.

[0027] To address this situation, many electronics manufacturers are creating how-to videos. When videos are insufficient, customers still require live assistance from service engineers. In one embodiment, an AR system is provided that is configured to provide AR overlays on screens and control panels, such as the computer / smartphone screen when a customer installs a new MFD driver, or the MFD's LCD screen when a worker touches buttons to configure the MFD. In particular, the embodiment utilizes the surfaces of screens and control panels, which are inherently two-dimensional surfaces, to provide overlays that are superior to annotations and related technology implementations.

[0028] In the implementation of the underlying technology, customers install screen-sharing software that allows remote engineers to view, control, and remotely move a cursor to guide them. Furthermore, in many cases, users are forced to rely solely on video, such as taking pictures of the LCD or control panel with a smartphone to show the remote engineer what they are seeing.

[0029] Implementing such background technologies presents challenges in installing screen-sharing software on a PC. For example, a customer may already be requesting assistance with installing other software, a company may not readily allow the installation of new software, the computer may not be connected to the internet, or there may be no screen-sharing application available for mobile devices.

[0030] Furthermore, with video streams, remote engineers may become confused if the user moves their phone, and communication can be significantly impaired as instructions are limited to verbal commands (for example, "Yes, click this red button in the bottom left, no, not this button, but that one, then press all these buttons at the same time and hold them down for three seconds").

[0031] To address these issues, the embodiments support an AR interface and overlay system corresponding to a control panel and screen (e.g., a computer screen, touchscreen, MFD, or a typical digital control panel found in electrical appliances such as microwave ovens and car stereo systems). With only a mobile device utilizing the AR interface of the embodiments described herein, the user can point the mobile device's camera at the screen / LCD / panel, and a remote engineer can guide the user by adding interactive overlay instructions.

[0032] Figure 1 shows an exemplary flow for overlaying at least one of an AR interface and / or annotations on a screen or control panel according to an embodiment. The flow begins when a local user connects to the remote assistance system via a user device.

[0033] In the embodiment, the system performs image tracking as the basis for detecting and tracking a screen or control panel. In 101, the system searches a database for anchor images that match the streaming content. Depending on the desired implementation, data can be added automatically or manually to the database of anchor images that indicate the objects to be detected. An anchor image is an image that has been processed to extract key points.

[0034] If the screen or LCD display is static and belongs to a known device (such as the LCD panel of a known MFD), the reference image is either pre-added to the application or retrieved from an online database and downloaded to the application. For example, in the case of an MFD, there may be a set of images showing the LCD control panel of a particular MFD device, and as soon as these types of control panels appear in the camera's field of view, the application can be automatically detected and tracked. Similarly, a set of images can be created for a common standard laptop model. Thus, if an anchor image is found within the application or can be retrieved from an online database, such an anchor image is used in 103.

[0035] If an anchor image is not found (102), the application also supports dynamic registration of unseen objects or LCD displays, in which case the rectangle detector can be used together with the AR plane detector. Specifically, when a service engineer or local user taps the screen, the application can be configured to run the rectangle or rectangle detector in the current frame and test the projection of the four corners of three-dimensional space intersecting with a known AR plane. A three-dimensional plane is then created that coincides with the two-dimensional points of the rectangle and is tracked in three-dimensional space by the AR framework, which then selects the anchor image in 104.

[0036] Once these reference images are confirmed, the video frames captured by the application are perspective-corrected so that the remote engineer can view a stable version of the area, and an Augmented Reality Overlay (ARO) can be created in 105. The remote assistant can then annotate the stream in 106, and the application system then determines in 107 whether there are any objects obstructing the screen. If there are, the annotations are hidden in 109; otherwise, the annotations are displayed in 108.

[0037] Once the application detects and tracks an anchor, the remote engineer can click on the screen to create an overlay. The mark is sent to the application and displayed at the corresponding location in AR. In this example, the tracked three-dimensional rectangle uses a WebView as a texture, and the mark created by the remote engineer is recreated in Hypertext Markup Language (HTML), so that what both users are seeing matches.

[0038] Depending on the desired implementation, overlaid marks can be masked and displayed on top of the display surface to improve the AR user experience. Such embodiments are useful when the device is a touch panel (digital touchscreen or physical buttons) and the customer may have difficulty seeing part of the display surface during interaction.

[0039] In one embodiment, the application facilitates dynamic overlays, allowing service engineers to create overlays containing multiple steps (e.g., "Enter something in this text box here, then click the 'OK' button"). In this case, the service engineer clicks / taps the text box, then moves to the OK button and clicks / taps it. Subsequently, the overlay is sent to the customer as an animation of what needs to be done, showing the movement from the customer's current position to the text box (e.g., an arc following the outline of the text box is highlighted), followed by another arc moving from the text box to the OK button. The steps can be numbered to make the order of actions to be performed clearer and to allow the customer to reproduce the steps they perform (which would be impossible if multiple overlays and mouse positions were transmitted in real time).

[0040] Unlike traditional screen sharing, dynamic overlays can be convenient for end users because they cannot always follow an entire sequence while looking at the display. Users may want to first view the sequence in AR and then perform the steps on the actual display. Furthermore, some steps require holding down multiple buttons, which may not be easily communicated using real-time overlays. Using the dynamic overlays described herein, service engineers can comfortably create a series of steps and then send them to remote customers after they have been correctly created. This asynchronous video collaboration, while synchronous video collaboration in other respects, is similar to what users can do in a text-based chat system: compose and edit text messages before committing and pressing "send."

[0041] In the implementation, various types of overlays can be used. For example, some actions require dragging a finger or mouse pointer along a path, while others simply require moving the finger / mouse to a different location. Several types of overlays can represent these differences, for example, a light arrow and a thick arrow. Depending on the desired implementation, overlays can be extended with text tooltips.

[0042] In the embodiment, detecting the current mouse / cursor position can also be easily done. Like teaching a child by the hand, the AR overlay can take into account the current finder / cursor position and show the user where to go next. For example, during a software installation process, it may not be clear where the user's cursor should be placed. Some UI elements require clicking inside a text box first. If a service engineer specifies clicking within an area, but the user's cursor is outside that area, the application can automatically display an arc from the current user mouse position to the text box position, making it clear that the cursor needs to move there first.

[0043] In some embodiments, automatic overlays can be facilitated. In some embodiments, procedures received during a live session can be recorded and played back later. For example, if the application detects that the same anchor image is included in the object being videotaped, it can automatically suggest playing back a previously recorded overlay instead of repeatedly calling a service engineer. This feature allows customers to operate the device themselves without needing live communication with a service engineer.

[0044] In the examples, it is also possible to verify whether an action has been performed. In some scenarios, a button may need to be held down for several seconds. When engineers create the overlay, they don't need to hold down the area for the required time (e.g., 10 seconds); they can specify the time. However, the user must hold down the button for the specified time. In addition to displaying the duration in a tooltip, the examples can also easily count how long the cursor / fingert was held at the specified location.

[0045] Figure 2 shows an exemplary overlay on a device panel captured from a user device, according to the embodiment. As shown in Figure 2, real-time rectangle detection is used to track the control panel captured by the user device. Three-dimensional overlay nodes can be generated and applied using the framework according to the embodiment, and texturing plane nodes are available in any view. Figure 3 shows an example of a three-dimensional overlay node with transitions, according to the embodiment.

[0046] For network communication, the user device can function as a web server and WebSocket server using appropriate libraries. Frames captured by the application are sent to the remote engineer as images, and the created marks are sent back to the application and recreated on the web view, which is used as a texture. For bidirectional audio communication, a WebRTC (Web Real-Time Communication) based solution can be used between the web browser and the application. Once a three-dimensional plane is fitted and then tracked by an AR framework, the frame is perspective-corrected and sent to the remote engineer. Figure 4 shows a perspective-corrected panel in an embodiment. Perspective correction allows the remote engineer to see a cropped and adjusted live camera view of the display captured by the application's end user. The remote engineer can create any overlay.

[0047] Through the embodiments, we can provide an AR system that overlays an AR interface on a two-dimensional surface, particularly in a live scenario, and specifically to create an overlay that is useful for guiding the user by obscuring the hand and detecting the mouse / finger position. Figure 5 shows an example of an overlay in which a hand and finger mask is implemented so that the overlay is positioned under the hand or fingers. Depending on the desired implementation, the hand and finger mask can also be implemented to track the hand and position the overlay under the hand or fingers. Such a mask can be obtained via a segmentation network or using a hand tracking model that tracks the hand or fingers in real time. Thus, if there is an object blocking the screen in 107 in Figure 1, the annotation added in 109 can be hidden.

[0048] In another embodiment, the AR remote assistance system can also generate browsing instructions for the system. Accumulating shared visual representations of the work environment can be useful in addressing many problems on site. Step-by-step instructions from an expert allow the user to complete a task, which can sometimes be difficult. During this time, the user may have to either put down the equipment or ignore visual input. Furthermore, the user may forget the exact details of how to perform certain steps and need the remote expert to repeat the instructions.

[0049] To address these issues, the embodiment extends the AR interface to support the creation of asynchronous tutorial procedures in a live remote assistance system. Figure 6 shows an example of recording and playback of the AR interface according to the embodiment. In the embodiment, instructions from a remote expert are automatically or manually saved as a video clip during a live video call. Then, when a local user needs to complete the procedure, they can view the saved video clip in another video player to complete the task. While the task is being completed, the remote expert can see a live view of the recording that the local user is viewing in a subwindow. The local user can switch to a live camera view at any time.

[0050] In the embodiment, video clip instructions are automatically generated whenever the remote expert is actively using a keyboard, mouse, or other peripheral device. The remote expert can also create instructions manually.

[0051] Figure 7 shows a flowchart of the annotation and recording process according to an embodiment. At 700, the local user connects to the remote assistant. Multiple functions can be supported during the connection. The local user can share a stream with the remote assistant at 701. In such an embodiment, the local user streams content to the remote user from a mobile device or heads-up device. In this case, when the remote expert begins providing annotations to the user stream, recording of the new clip automatically starts in the background while the live video session continues. The system records until the remote expert stops annotating the stream and a timeout occurs. The remote expert can optionally pause the user video and add more expressive annotations, depending on the desired implementation. As shown in Figure 7, the remote assistant can view the stream at 704 and add annotations as needed. While the stream is running, the remote assistant can pause annotations at 707. If annotations are paused for a threshold period (e.g., a few seconds), a timeout occurs at 709. At that point, the system saves the video clip as a procedure at 711.

[0052] In another example, the remote assistant shares the stream with the local user at 702. Sometimes, for example, if the local user is trying to resolve a software system issue, the remote expert may share their screen to demonstrate how to resolve the specific issue using their software tools. In this case, the remote expert actively uses their mouse and keyboard to demonstrate the "steps" that the system can record. Again, a timeout is used to determine when the steps are complete. At 705, the remote assistant begins interacting with the stream by providing annotations or controlling an interface or panel on the screen. The flow can continue with saving the video clip, as shown in 707 and beyond. In another embodiment, the remote expert can also manually create the video clip by clicking a button on the interface. This is useful if the remote user wants to create a clip using their own video camera or load an external clip.

[0053] In another embodiment, a local user may place the user device to perform a function indicated by the remote assistant in 703. The placement of the user device can be detected based on the accelerometer, gyroscope, or through other hardware of the device, according to the desired implementation. Even if the user is attempting to hold the device, the background process system may detect slight anomalies in the accelerometer and gyroscope data to determine that the device is being held. However, once the user places the device, the accelerometer and gyroscope data become static, and the background process can determine that the device is no longer in the user's hands. In this way, the system can automatically switch between displaying recorded procedures (when the device is placed) and live streams (when the device is in the user's hands). In 706, when it is detected that the device has been placed, the application switches to the procedure view. The procedure view is maintained until the local user picks up the device in 708. Then, the application returns to the live view in 710.

[0054] These approaches can be combined to help local users complete challenging tasks. For example, when interacting with a complex interface, a remote expert can annotate the user's live stream, automatically creating one clip. Then, while the user pauses to complete the task, the remote expert can annotate the same or similar interface in their own stream, automatically creating another clip. Alternatively, another clip can be manually loaded from a recorded stream of another user who has dealt with the same problem.

[0055] Similarly, local users can switch between live video streaming and clip review using automatic or manual methods, depending on the desired implementation.

[0056] By default, the system turns off the local user microphone when a local user is reviewing (viewing) a clip. Also, by default, the most recently recorded clip is displayed first. Furthermore, users can move between different media clips using standard vertical swipe gestures and navigate within clips using horizontal swipe gestures. In this way, local users can seamlessly switch the device from a live streaming tool to a lightweight tutorial review tool.

[0057] If the user is streaming from a head-up display, they can switch between live streaming and the review interface by issuing a verbal command. On mobile devices, the user can switch interfaces by issuing a verbal command or by pressing a button.

[0058] Through the embodiments described herein, a remote assistance system can be facilitated that automatically records procedures in a live stream that can be viewed asynchronously by the user.

[0059] Figure 8 shows an example of a computer device according to an embodiment. The computer device may take the form of a laptop, personal computer, mobile device, tablet, or other device, depending on the desired implementation. The computer device 800 may include a camera 801, a microphone 802, a processor 803, memory 804, a display 805, an interface (I / F) 806, and a compass sensor 807. The camera 801 may include any type of camera configured to record any type of video, depending on the desired implementation. The microphone 802 may include any type of microphone configured to record any type of audio, depending on the desired implementation. The display 805 may include a touchscreen display configured to receive touch input to facilitate instructions for performing the functions described herein, or, depending on the desired implementation, a conventional display such as a liquid crystal display (LCD) or any other display. The I / F 806 may include a network interface to facilitate connection of the computer device 800 to external elements such as servers and other arbitrary devices, depending on the desired implementation. The processor 803 may be in the form of a hardware processor such as a central processing unit (CPU), or a combination of hardware and software units depending on the desired implementation. The orientation sensor 807 may include at least one of any form of gyroscope and accelerometer configured to measure all kinds of orientation measurements, such as tilt angle, orientation relative to x, y, and z, axis, and acceleration (e.g., gravity), depending on the desired implementation. The orientation sensor measurement may also include gravity vector measurement to indicate the gravity vector of the device, depending on the desired implementation. The computer device 800 may be used as a device for a local user or as a device for a remote assistant, depending on the desired implementation.

[0060] In this embodiment, the processor 803 is configured to fix and continuously display video received from the first device (e.g., a local user device) for display on a second device (e.g., a remote assistant device). Then, in response to input made to a portion of the fixedly displayed video on the second device, it generates an augmented reality (AR) overlay on the display of the first device corresponding to the fixedly displayed portion of the video, for example, as shown in Figures 3 to 5.

[0061] Depending on the desired implementation, the processor 803 may be configured to fix and display the video received from the first device for display on the second device by identifying one or more anchor images in the video, determining the object to be displayed in the two-dimensional display based on the identified anchor images, and performing perspective correction of the video based on the object to be displayed on the second device, which is a two-dimensional surface for display on the second device, as described in Figure 1. As described herein, the object to be displayed may include a two-dimensional panel surface such as a panel display (e.g., displayed on an MFD), a keypad, a touchscreen, a display screen (e.g., on a computer, mobile device, or other device), and other physical or displayed interfaces depending on the desired implementation. Depending on the desired implementation, the anchor images may include buttons, dials, icons, or other objects that are presumed to be on the panel surface.

[0062] Depending on the desired implementation, the processor 803 may be configured to crop the video to a target object for display on a second device, as shown in Figure 4. In this way, the video can be cropped so that only a display screen, panel display, or other target object is provided to the second device.

[0063] Depending on the desired implementation, the processor 803 is configured to identify one or more anchor images in a video by searching a database for one or more anchor images that match one or more images in the video, as described in Figure 101. The database may be stored and accessed remotely by a storage system, server, or other means, depending on the desired implementation. In the embodiment, AR overlays may also be stored in the database for retrieval or future playback by the first device.

[0064] The processor 803 is configured to identify one or more anchor images in a video by detecting a rectangle in the video received from the first device, matching a three-dimensional plane to the two-dimensional points of the detected rectangle, tracking the three-dimensional plane that matches the two-dimensional points of the detected rectangle, and receiving a selection of one or more anchor images in the video via the second device, as described in Figure 1. Since most panels and displayed interfaces tend to take the form of a rectangle or square, the embodiments utilize a rectangle or square detector known in the art, although the detector can be modified according to the desired implementation. For example, in embodiments including a circular interface, a circular surface detector can be used instead. Furthermore, after a rectangle or square is detected, a three-dimensional plane is mapped to the two-dimensional points of the detected rectangle / square (e.g., mapping to the corners of the rectangle), which can then be tracked according to a desired implementation known in the prior art. Once a panel is detected, the user can select anchor images (e.g., buttons, dials on the panel) that can be incorporated into the AR system in real time.

[0065] As shown in Figures 1 to 6, to facilitate real-time interaction between a remote assistant and a local user, an AR overlay corresponding to a portion of a fixed-display video can be displayed live on the display of the first device in response to input made on a portion of the fixed-display video of the second device. In another embodiment, depending on the desired implementation of the remote assistant, the generation of the AR overlay can be delayed and deployed asynchronously. The remote assistant can check the AR overlay on its own device and then give instructions (such as touching a confirmation button) to the device to send the AR overlay to the local user device for display. In this way, the remote assistant can create AR annotations or provide other AR overlays to preview before deploying them to the local user. In the embodiment, input made on a portion of a fixed-display video can include free-form annotations. Furthermore, if the AR overlay involves selecting a specific panel button or moving a cursor to click a specific portion, the AR overlay can include markings for moving the cursor of the first device from a first location to a second location. Such markings can be implemented in any way (e.g., by arrows or lines that trace a path) depending on the desired implementation.

[0066] As shown in Figure 5, the processor 803 can be configured to track one or more hands and fingers in the video and to occlude the portion of the AR overlay that overlaps with one or more hands and fingers on the display of the first device. Tracking of at least one of the hands and fingers can be carried out through any desired implementation. Such an embodiment allows the AR overlay to be displayed in a realistic manner on the remote user's device.

[0067] As shown in Figure 7, the processor 803 can be configured to pause the video on the display of the second device, thereby fixing the video received from the first device for display on the second device. In such an embodiment, depending on the desired implementation, the remote assistant can pause the video stream to create annotations or provide other AR overlays. Furthermore, if the input pause exceeds a timeout threshold, the processor 803 can be configured to generate a video clip containing annotations and provide the video clip on the display of the first device, as shown in 707, 709, and 711 of Figure 7, thereby generating an AR overlay on the display of the first device corresponding to the fixed portion of the video. The timeout can be set depending on the desired implementation.

[0068] The processor 803 may be configured to play a video clip on the display of the first device if it is determined that the first device is present, as shown in 703, 706, 708, and 710 of Figure 7, and to provide the video for display on the second device if it is determined that the first device is in the user's hands.

[0069] Some parts of the detailed description are presented with respect to algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are means used by those skilled in the field of data processing technology to convey the essence of its innovations to others skilled in the field. An algorithm is a set of defined steps leading to a desired final state or result. In the examples, the steps performed require a specific amount of physical operation to achieve a specific result.

[0070] Unless otherwise specified, as will be clear from the preceding explanation, throughout this specification, any description using terms such as “processing,” “computing,” “calculating,” “determining,” and “displaying” may include the operation and processing of a computer system or other information processing device that manipulates and converts data represented as physical (electron) quantities in the registers or memory of a computer system into other data similarly represented as physical quantities in the memory or registers of a computer system or other information storage, transmission, or display device.

[0071] The embodiments may also relate to apparatus for performing the operations described herein. Such apparatus may include one or more general-purpose computers that may be specifically constructed for a required purpose or selectively operated or reconfigured by one or more computer programs. Such computer programs may be stored on computer-readable media such as computer-readable storage media or computer-readable signal media. Computer-readable storage media may include, but are not limited to, specific media such as optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices and drives, or other suitable tangible or non-temporary media for storing electronic information. Computer-readable signal media may include media such as carrier waves. The algorithms and representations presented herein are not inherently related to any particular computer or other apparatus. Computer programs may include pure software implementations containing instructions for performing the operations of a desired implementation.

[0072] Various general-purpose systems can be used with the programs and modules provided in the embodiments herein, or it may be advantageous to construct more specialized equipment to perform the desired method steps. Furthermore, the embodiments are not described in terms of a specific programming language. It will be understood that various programming languages ​​may be used to implement the teachings of the embodiments described herein. Instructions in a programming language may be executed by one or more processing units, such as a central processing unit (CPU), a processor, or a control unit.

[0073] As is known in the art, the operations described above can be performed by hardware, software, or any combination of software and hardware. Various embodiments of the embodiment can be implemented using circuits and logic devices (hardware), while other embodiments can be implemented using instructions stored in a machine-readable medium (software), which, when executed by a processor, perform the method for carrying out the implementation of the present application. Furthermore, some embodiments of the present application may be performed by hardware alone, and other embodiments may be performed by software alone. Moreover, the various functions described may be performed by a single unit or distributed among numerous components in any various way. When performed by software, the method may be executed by a processor such as a general-purpose computer based on instructions stored in a computer-readable medium. If necessary, the instructions may be stored in the medium in a compressed and encrypted form.

[0074] Other implementations of this application will be apparent to those skilled in the art, given the practice of the specification and teachings of this application. The various aspects and components of the described embodiments may be used individually or in any combination. The specification and embodiments are intended to be considered illustrative only, and the true scope and spirit of this application are set forth by the following claims.

Claims

1. The video received from the first device is paused on the second device in order to be displayed on the second device. The second device accepts input including annotations for a video that has been paused, If the pause time of the input exceeds a threshold, the second device generates an augmented reality (AR) overlay on the display of the first device corresponding to the paused portion of the video for the input made to the paused portion of the video. Includes, Augmented reality (AR) overlay corresponding to the paused portion of the video is generated on the display of the first device. If the pause time of the input exceeds the threshold, a video clip including the annotation is generated. The video clip is provided to the display of the first device, If it is determined that the first device is present, the video clip is played on the display of the first device. If it is determined that the first device is in the user's possession, the video is provided for display on the second device. Methods that include...

2. The video received from the first device is paused on the second device in order to be displayed on the second device. The second device accepts input including annotations for a video that has been paused, The second device generates a dynamic overlay corresponding to the input made on the paused portion of the video when the pause time of the input exceeds a threshold, The aforementioned dynamic overlay is an animation, The second device generates a video clip including the animation, The second device transmits the video clip to the display of the first device. Methods that include...

3. The aforementioned dynamic overlay includes multiple operation steps, The method according to claim 2.

4. While sharing a live video stream between a first device used by a first user and a second device used by a second user, the live stream is paused on the second device. The second device receives a first input, including annotations, for the live stream that has been paused on the second device. If the pause time of the first input exceeds a threshold, a first augmented reality (AR) overlay corresponding to the first input is generated. A first video clip including the first AR overlay is generated, The first video clip is provided to the display of the first device. While the live stream is stopped in the first device, On the second device, generate a second video clip that is different from the first video clip. Methods that include...

5. Generating the second video clip is The system receives a second input from the second user, including annotations, to the same or similar interface as the live stream. A second AR overlay corresponding to the second input is generated, To generate a second video clip including the second AR overlay. The method according to claim 4, including the method described in claim 4.

6. The method according to claim 4 or 5, wherein the first device is capable of switching between displaying the video clip and the live stream.

Citation Information

Patent Citations

  • Electronic still camera

    JP2001136418A

  • Collaborative dynamic video commenting device and method

    JP2002507027A

  • Remote support method and system, and program

    JP2011034315A

  • Video instruction display method, system, terminal, and program for synchronously superposing instruction image on imaged moving image

    JP2015179947A

  • Display system, display device, terminal device, and program

    JP2019079314A