Apparatus, display apparatus, method of defining boundary, and computer-program product
The apparatus and method allow users to define VR safety boundaries using interactive inputs without full rotation, ensuring safety and comfort by integrating 360-degree image capture and display technologies.
Patent Information
- Application Number
- PCT/CN2024/107427
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-01-29
AI Technical Summary
Existing VR systems require users to rotate 360 degrees to define safety boundaries, which is inconvenient and may not align with user needs, potentially leading to collisions and accidents.
An apparatus and method that uses 360-degree environmental image capture, obstacle detection, and image stitching, allowing users to define safety boundaries through interactive inputs such as head movements, eye movements, controllers, or gestures without rotating 360 degrees, and displays the results for review and adjustment.
Enables users to comfortably and accurately set virtual reality boundaries within a 360-degree environment, enhancing safety and user experience by integrating multiple input methods and displaying the boundary for confirmation or adjustment.
Smart Images

Figure CN2024107427_29012026_PF_FP_ABST
Abstract
Description
APPARATUS, DISPLAY APPARATUS, METHOD OF DEFINING BOUNDARY, AND COMPUTER-PROGRAM PRODUCTTECHNICAL FIELD
[0001] The present invention relates to display technology, more particularly, to an apparatus, a display apparatus, a method of defining a boundary, and a computer-program product.BACKGROUND
[0002] Virtual Reality (VR) is a technology that uses computer systems to create and experience virtual worlds. It simulates human sensory experiences such as vision, hearing, and touch, allowing users to immerse themselves in a virtual environment and interact with virtual objects. VR technology typically requires specialized equipment, such as VR devices, gloves, and controllers, to provide an immersive experience. When users wear VR devices, they see virtual scenes displayed on screens and interact with the virtual environment using head tracking and controllers.SUMMARY
[0003] In one aspect, the present disclosure provides an apparatus, comprising a memory; and one or more processors; wherein the memory and the one or more processors are connected with each other; and the memory stores computer-executable instructions for controlling the one or more processors to obtain a plurality of original images captured by the plurality of cameras; perform an image stitching process on the plurality of original images to obtain a stitched image; perform object detection and localization based on the stitched image; define a boundary based on a user input, and a result of the object detection and localization; and have the boundary displayed for review and / or adjustment.
[0004] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to switch views within the stitched image based on a first user input; and define the boundary based on a second user input, and the stitched image comprising one or more detected objects; wherein the first user input and the second user input are different from each other.
[0005] Optionally, the user input includes a combination of the head movement and one or both of the hand gesture and the controller input; the memory stores computer-executable instructions for controlling the one or more processors to switch views within the stitched image based on a first user input from the head movement; and define the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.
[0006] Optionally, the user input includes a combination of the eye movement and one or both of the hand gesture and the controller input; the memory stores computer-executable instructions for controlling the one or more processors to switch views within the stitched image based on a first user input from the eye movement; and define the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.
[0007] Optionally, the user input includes a combination of the head movement and the eye movement; the memory stores computer-executable instructions for controlling the one or more processors to switch views within the stitched image based on a first user input from the head movement; and define the boundary based on a second user input from the eye movement, and the stitched image comprising one or more detected objects.
[0008] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to determine whether an input from a hand gesture or a controller is received; and determine whether an input from a head movement is received.
[0009] Optionally, based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is received, the memory stores computer-executable instructions for controlling the one or more processors to switch views within the stitched image based on a first user input from the head movement; and define the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.
[0010] Optionally, based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is not received, the memory stores computer-executable instructions for controlling the one or more processors to switch views within the stitched image based on a first user input from an eye movement; and define the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.
[0011] Optionally, based on determination that the input from the hand gesture or the controller is not received, the memory stores computer-executable instructions for controlling the one or more processors to switch views within the stitched image based on a first user input from the head movement; and define the boundary based on a second user input from the eye movement, and the stitched image comprising one or more detected objects.
[0012] Optionally, upon completion of defining an entire 360-degree boundary, the memory stores computer-executable instructions for controlling the one or more processors to have the boundary displayed for review and / or adjustment.
[0013] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to detect obstacles within the boundary; upon detecting an obstacle, highlight the obstacle, and prompt a user to address the obstacle; and receive a user input to adjust the boundary.
[0014] Optionally, upon determination that the obstacle is not removed, the memory stores computer-executable instructions for controlling the one or more processors to exclude the obstacle from the boundary to maximize safety.
[0015] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to receive a user input for select a viewpoint for adjustment; have the specific viewpoint displayed in a front viewing area; and adjust the boundary based on an additional user input.
[0016] Optionally, to perform the image stitching process on the plurality of original images to obtain the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to perform an image preprocessing process on the plurality of original images to obtain a plurality of pre-processed images; perform a feature point detection and matching process, extracting one or more feature points from each image, and determining correspondence between images based on matched feature points; perform perspective transformation, determining a mapping relationship between images based on image coordinates of matched feature points; and perform an image postprocessing process to obtain the stitched image.
[0017] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to perform object detection and localization based on the stitched image.
[0018] Optionally, to perform object detection and localization based on the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to perform an object detection process; perform an object segmentation process; and perform a semantic segmentation process and / or a panoptic segmentation process.
[0019] Optionally, the stitched image is a stitched image of a panoramic view.
[0020] In another aspect, the present disclosure provides a display apparatus, comprising the apparatus described herein, a plurality of cameras, and a display panel connected to the apparatus.
[0021] In another aspect, the present disclosure provides a method of defining a boundary, comprising obtaining a plurality of original images captured by a plurality of cameras; performing an image stitching process on the plurality of original images to obtain a stitched image; performing object detection and localization based on the stitched image; defining a boundary based on a user input, and a result of the object detection and localization; and having the boundary displayed for review and / or adjustment.
[0022] In another aspect, the present disclosure provides a computer-program product, comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform obtaining a plurality of original images captured by a plurality of cameras; performing an image stitching process on the plurality of original images to obtain a stitched image; performing object detection and localization based on the stitched image; defining a boundary based on a user input, and a result of the object detection and localization; and having the boundary displayed for review and / or adjustment.
[0023] BRIEF DESCRIPTION OF THE FIGURES
[0024] The following drawings are merely examples for illustrative purposes according to various disclosed embodiments and are not intended to limit the scope of the present invention.
[0025] FIG. 1 illustrates a method of defining a boundary.
[0026] FIG. 2 illustrates a method of defining a boundary.
[0027] FIG. 3 illustrates an image preprocessing process in some embodiments according to the present disclosure.
[0028] FIG. 4 illustrates a feature point detection and matching process in some embodiments according to the present disclosure.
[0029] FIG. 5 illustrates perspective transformation in some embodiments according to the present disclosure.
[0030] FIG. 6 illustrates an image postprocessing process in some embodiments according to the present disclosure.
[0031] FIG. 7 illustrates an object detection process in some embodiments according to the present disclosure.
[0032] FIG. 8 illustrates an object segmentation process in some embodiments according to the present disclosure.
[0033] FIG. 9 illustrates a semantic segmentation process in some embodiments according to the present disclosure.
[0034] FIG. 10 illustrates a panoptic segmentation process in some embodiments according to the present disclosure.
[0035] FIG. 11 illustrates various modes of boundary definition in some embodiments according to the present disclosure.
[0036] FIG. 12 illustrates a continuous angle switching mode in some embodiments according to the present disclosure.
[0037] FIG. 13 illustrates a discrete angle switching mode in some embodiments according to the present disclosure.
[0038] FIG. 14 illustrates a mirror switching mode in some embodiments according to the present disclosure.
[0039] FIG. 15 illustrates a process where an obstacle within an initial boundary is identified and either highlighted for removal or automatically excluded from a safety area.
[0040] FIG. 16 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure.
[0041] FIG. 17 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure.
[0042] FIG. 18 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure.
[0043] FIG. 19 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure.DETAILED DESCRIPTION
[0044] The disclosure will now be described more specifically with reference to the following embodiments. It is to be noted that the following descriptions of some embodiments are presented herein for purpose of illustration and description only. It is not intended to be exhaustive or to be limited to the precise form disclosed.
[0045] When a user wears a virtual reality device, their eyes are unable to see the real world, potentially leading to collisions and accidents during the virtual reality experience. To prevent such incidents, it is necessary to define a safe boundary to ensure user safety. Related solutions require users to rotate 360 degrees to observe and define the boundary, which may not align with user needs.
[0046] In some embodiments, a stationary mode is used for defining the boundary. FIG. 1 illustrates a method of defining a boundary. Referring to FIG. 1, in the stationary mode, a circular safety boundary is automatically defined around the user's current position. For example, the safety boundaries may have radii of 1.7m, 2.5m, and 3.5m. During the boundary setup, the system prompts the user to ensure that the surrounding area is clear of obstacles to prevent collisions with limbs or the head.
[0047] If the user reaches the safety boundary, the system switches to a pass-through view, displaying the real environment to ensure the user's safety.
[0048] In some embodiments, a custom mode is used for defining the boundary. FIG. 2 illustrates a method of defining a boundary. Referring to FIG. 2, in the custom mode, users define the safety boundary by manually rotating in a circle with the controller. In an automatic boundary defining method, users wear the virtual reality device and rotate in a circle. The system collects real-time depth information to detect obstacles and automatically defines the custom safety boundary. This boundary is then presented to the user for confirmation, allowing further adjustments with the controller before finalizing.
[0049] Once the safety boundary is defined, if the user's gesture crosses the boundary, an additional circle appears around the gesture as a warning. If the user reaches the boundary, the system switches to a pass-through view to display the real environment, ensuring user safety.
[0050] The custom mode requires users to rotate in a circle with the controller to manually define the safety boundary. The automatic boundary defining method, equipped with a depth camera, also needs the user to rotate in a circle to scan the surrounding environment and automatically generate a custom boundary. This requirement for the user to rotate 360 degrees does not fully meet user needs.
[0051] Accordingly, the present disclosure provides, inter alia, an apparatus, a display apparatus, a method of defining a boundary, and a computer-program product that substantially obviate one or more of the problems due to limitations and disadvantages of the related art. In one aspect, the present disclosure provides an apparatus. In some embodiments, the apparatus includes a plurality of cameras; a memory; and one or more processors. Optionally, the memory and the one or more processors are connected with each other. Optionally, the memory stores computer-executable instructions for controlling the one or more processors to obtain a plurality of original images captured by the plurality of cameras; perform an image stitching process on the plurality of original images to obtain a stitched image; perform object detection and localization based on the stitched image; define a boundary based on an interactive user input, and the stitched image comprising one or more detected objects; and have the boundary displayed for review and / or adjustment.
[0052] The present disclosure provides a method and apparatus that allows users to define safety boundaries without the need to rotate 360 degrees. The method and apparatus according to the present disclosure leverages 360-degree environmental image capture, obstacle detection and localization, and image stitching. The method and apparatus according to the present disclosure integrates head movements, eye movements, controllers, or gestures as interactive inputs to control the environment scene. Once the boundary is defined, the system displays the results through an overhead map view, allowing users to make further adjustments and / or confirmations. The apparatus according to the present disclosure in some embodiments includes a plurality of cameras to capture 360° images of the user's surroundings without the need for the user to turn their head. The system can automatically detect and locate obstacles in the surrounding environment and then generate the safety boundary automatically. It also supports manual boundary definition, allowing users to control the environment scene using head movements, eye movements, controllers, or gestures without needing to rotate in a circle. Once defined, the system displays the safety boundary results through an overhead map view, allowing the user to confirm or make further adjustments.
[0053] In one aspect, the present disclosure provides an apparatus for defining a boundary. In some embodiments, the apparatus includes a plurality of cameras. In some embodiments, the plurality of cameras are integrated into a virtual reality device. In one example, the plurality of cameras are distributed around the virtual reality device to allow capturing 360-degree images of the surrounding environment without requiring any user action.
[0054] In alternative embodiments, the plurality of cameras are integrated into a neck-mounted device. This is particularly suitable for a glasses-type virtual reality product. In one example, the plurality of cameras are distributed around the neck-mounted device to allow capturing 360-degree images of the surrounding environment without requiring any user action.
[0055] Various appropriate implementations may be practiced according to the present disclosure. In some embodiments, the plurality of cameras are configured to actively obtain depth information. In some embodiments, the plurality of cameras are structured light cameras. In one example, a structured light camera includes a projector and a receiver. The projector is configured to cast a light pattern onto the object's surface, and the receiver is configured to capture the deformation of this pattern. The apparatus for defining the boundary further includes one or more processors configured to analyze the deformation and determine the object's depth. As used herein, in the context of obtaining depth information, the term “deformation” refers to changes or distortions in a projected light pattern when it is cast onto the surface of an object. These deformations occur because the surface of the object is not flat, causing the light pattern to bend, stretch, compress, or otherwise change shape in response to the object's contours and features. By capturing and analyzing these deformations, the system can infer the depth and three-dimensional shape of the object.
[0056] In some embodiments, the plurality of cameras are time of flight (TOF) cameras. In one example, a time of flight camera includes a projector and a receiver. When the light source emits a signal, it reflects off the object and is captured by the receiver. By measuring the travel time of the light signal, the distance between the light source and the object is calculated. Time of flight cameras provide a complete depth map of the scene in one shot, have no scanning components, offer fast imaging speeds, and have a low computational load, making them widely used.
[0057] In one example, the apparatus for defining the boundary includes four cameras configured to actively obtain depth information, wherein each camera has a 90° field of view.
[0058] In some embodiments, the plurality of cameras are configured to passively obtain depth information. In some embodiments, the plurality of cameras are stereo cameras. In one example, a stereo camera includes two lenses to capture depth data by comparing the disparity between images taken from two different angles. The disparity is then used to calculate the distance to the object.
[0059] In one example, the apparatus for defining the boundary includes eight cameras configured to passively obtain depth information, wherein each camera has a 45° field of view.
[0060] In some embodiments, the apparatus for defining the boundary further includes one or more processors. In some embodiments, the one or more processors are configured to obtain a plurality of original images captured by the plurality of cameras. In some embodiments, the one or more processors are further configured to perform an image stitching process on the plurality of original images to obtain a stitched image. Optionally, the stitched image is a stitched image of panoramic view (e.g., 360-degree panoramic view) . Image stitching is a technique used to combine two or more images with overlapping areas to create a single image with a wider or 360-degree panoramic view, providing more information to support various subsequent processes. It is widely used in machine vision fields such as motion detection and tracking, augmented reality, image stabilization, resolution enhancement, and video compression. As used herein, the term “panoramic view” refers to field of view that equals or exceeds the field of view of the human eye (generally considered 70° by 160°) . A panoramic view may encompass 360° along a given plane, for example along the horizontal plane. In one embodiment, the panoramic view is a spherical field of view.
[0061] In some embodiments, to perform the image stitching process on the plurality of original images to obtain the stitched image, the one or more processors are configured to perform an image preprocessing process on the plurality of original images to obtain a plurality of pre-processed images. Optionally, the image preprocessing process includes denoising to remove noise. Optionally, the image preprocessing process includes image enhancement to enhance the detailed texture features of the image, improving the accuracy of feature detection. FIG. 3 illustrates an image preprocessing process in some embodiments according to the present disclosure.
[0062] The inventors of the present disclosure discover that the image preprocessing process is a crucial step in ensuring the accuracy and quality of the subsequent image stitching. By performing denoising, the system removes unwanted noise that can obscure important details and introduce errors in feature detection. Image enhancement further refines the pre-processed images by emphasizing texture features, edges, and other critical details, thereby facilitating more precise matching of corresponding points across images. As illustrated in FIG. 3, the left side shows an original image with potential noise and less defined features, while the right side demonstrates the enhanced image with clearer textures and reduced noise. This preprocessing not only improves the robustness of the feature detection and matching algorithms but also contributes to the overall fidelity of the stitched image, ensuring a seamless and visually coherent final output.
[0063] In some embodiments, to perform the image stitching process on the plurality of original images to obtain the stitched image, the one or more processors are further configured to perform a feature point detection and matching process. Optionally, the feature point detection and matching process includes extracting one or more feature points (e.g., key feature points) from each image, such as corners, edges, and textures, and determines the correspondence between images based on matched feature points. FIG. 4 illustrates a feature point detection and matching process in some embodiments according to the present disclosure.
[0064] The inventors of the present disclosure discover that the feature point detection and matching process is essential for ensuring the accuracy of image alignment in the stitching process. "Key feature points" refer to distinct and identifiable points within an image that are selected based on their unique and invariant properties, which enable reliable detection, description, and matching across multiple images. These points are characterized by their ability to remain consistent under various transformations such as scale, rotation, and affine changes. Key feature points are utilized to establish correspondences between images, facilitating processes such as image stitching, alignment, and recognition. By extracting key feature points such as corners, edges, and textures from each image, the system can identify unique and stable points that can be reliably matched across multiple images. This matching process establishes the correspondence between different images, enabling the system to accurately align them. As illustrated in FIG. 4, the feature points detected in the original images are used to create a network of corresponding points, ensuring that the images can be seamlessly stitched together. This process not only facilitates the creation of a coherent and visually accurate stitched image but also helps in maintaining the geometric and photometric consistency across the combined images. The accurate matching of feature points is crucial for achieving high-quality results in applications such as virtual reality, panoramic photography, and various machine vision tasks.
[0065] In some embodiments, to perform the image stitching process on the plurality of original images to obtain the stitched image, the one or more processors are further configured to perform perspective transformation. Optionally, the perspective transformation includes solving a homography matrix. Optionally, the perspective transformation includes determining a mapping relationship between images based on image coordinates of matched feature points. The goal is to transform the perspective of the image to be registered to match that of the reference image, ensuring correct stitching and generating a stitched image with spatial consistency and natural appearance. FIG. 5 illustrates perspective transformation in some embodiments according to the present disclosure.
[0066] The inventors of the present disclosure discover that the perspective transformation is essential for aligning multiple images accurately in the stitching process. By solving a homography matrix, the system establishes a mapping relationship between the image coordinates of matched feature points from different images. This relationship allows the transformation of the perspective of each image to align with the reference image. As shown in FIG. 5, the transformation adjusts the grid lines from an unaligned state on the left to a correctly aligned state on the right. This ensures that the overlapping regions of the images are correctly stitched together, maintaining spatial consistency and producing a stitched image with a natural and cohesive appearance. This step is crucial for correcting geometric distortions and achieving seamless integration of multiple images into a single, panoramic view.
[0067] In some embodiments, to perform the image stitching process on the plurality of original images to obtain the stitched image, the one or more processors are further configured to perform an image postprocessing process to obtain a stitched image. Optionally, the stitched image is an image of panoramic view (e.g., 360-degree panoramic view) . Optionally, the image postprocessing process includes performing an image fusion process to eliminate seams that may exist between stitched images, improving the quality of the stitched image. FIG. 6 illustrates an image postprocessing process in some embodiments according to the present disclosure.
[0068] The inventors of the present disclosure discover that the image postprocessing process is a critical step in the image stitching workflow, aimed at refining the stitched image. After aligning and stitching the images, the one or more processors perform an image fusion process to eliminate any visible seams between the stitched images. This step ensures that transitions between images are smooth and indistinguishable, resulting in a high-quality stitched image. As shown in FIG. 6, the left side depicts the initial stitched image with noticeable seams and misalignments, while the right side illustrates the final stitched image after postprocessing, which appears seamless and visually coherent. The image fusion process enhances the overall aesthetic and spatial consistency of the stitched image, making it suitable for applications requiring high visual fidelity, such as panoramic photography and virtual reality environments.
[0069] In some embodiments, the one or more processors are further configured to perform object detection (e.g., obstacle detection) and localization based on the stitched image.
[0070] In some embodiments, to perform object detection and localization based on the stitched image, the one or more processors are configured to perform an object detection process. Optionally, the object detection process includes identifying and locating an object of interest (e.g., people, vehicles, animals) in the stitched image. Optionally, the object detection process further includes determining a bounding box and a class label of the object. FIG. 7 illustrates an object detection process in some embodiments according to the present disclosure.
[0071] The inventors of the present disclosure discover that the object detection process is integral to the system's ability to identify and interact with the environment accurately. In one example, by leveraging advanced machine learning algorithms, the system can analyze the stitched image to detect and classify objects with high precision. The bounding box, as shown in FIG. 7, serves not only to highlight the object's location but also to provide spatial context, which is crucial for subsequent processes such as obstacle avoidance and navigation. Furthermore, the class label offers semantic information that can be used to tailor interactions based on the object's type, enhancing the user's experience and safety. For instance, distinguishing between static obstacles like furniture and dynamic obstacles like people allows the system to implement more sophisticated and responsive safety measures. This dual approach of localization and classification ensures that the virtual reality environment remains both immersive and secure.
[0072] In some embodiments, to perform object detection and localization based on the stitched image, the one or more processors are configured to perform an object segmentation process. Optionally, the object segmentation process includes segmenting the object from a background in the stitched image, separating the object from the background. Optionally, the object segmentation process further includes determining a mask or a contour of the object. FIG. 8 illustrates an object segmentation process in some embodiments according to the present disclosure.
[0073] The inventors of the present disclosure discover that the object segmentation process plays a pivotal role in distinguishing objects from their backgrounds in the stitched image. By accurately segmenting objects, the system can isolate them from the surrounding environment, which is crucial for detailed analysis and interaction. As depicted in FIG. 8, the segmentation process involves creating masks or contours that precisely outline each object, allowing for clear separation from the background. This capability enhances the system's ability to perform subsequent tasks, such as obstacle avoidance and path planning, by providing more granular information about the objects's hapes and boundaries. Additionally, object segmentation improves the system's overall understanding of the scene, facilitating more intelligent and context-aware responses to dynamic changes in the environment.
[0074] In some embodiments, to perform object detection and localization based on the stitched image, the one or more processors are configured to perform a semantic segmentation process. Optionally, the semantic segmentation process includes assigning each subpixel in the composite to a predefined class (e.g., road, building, trees) . Optionally, semantic segmentation does not distinguish between different instances of objects but divides the entire image into different semantic regions. FIG. 9 illustrates a semantic segmentation process in some embodiments according to the present disclosure.
[0075] The inventors of the present disclosure discover that the semantic segmentation process further enhances the system's ability to understand and interpret the environment by categorizing every subpixel in the stitched image into predefined classes such as roads, buildings, and trees. Unlike object segmentation, which isolates individual objects, semantic segmentation focuses on dividing the entire image into meaningful regions based on semantic content. As illustrated in FIG. 9, each region of the image is labeled with a specific class, providing a comprehensive understanding of the scene's layout. This process does not differentiate between instances of the same class but instead groups all pixels belonging to a particular category together. This semantic information is vital for applications such as autonomous navigation and environmental mapping, where understanding the context and relationships between different regions is crucial for making informed decisions and interactions.
[0076] In some embodiments, to perform object detection and localization based on the stitched image, the one or more processors are configured to perform a panoptic segmentation process. The panoptic segmentation process is a more advanced form of semantic segmentation that not only segments the image into different semantic regions but also distinguishes between different instances of the same class. Optionally, the panoptic segmentation process includes assigning each subpixel in the composite to a predefined class (e.g., road, building, trees) and an instance. Optionally, the panoptic segmentation process further includes determining a mask or a contour of each object instance. FIG. 10 illustrates a panoptic segmentation process in some embodiments according to the present disclosure.
[0077] The inventors of the present disclosure discover that the panoptic segmentation process combines the strengths of both semantic and instance segmentation, providing a comprehensive understanding of the scene by not only classifying each subpixel into predefined categories but also distinguishing between different instances of the same class. As depicted in FIG. 10, this process involves assigning each subpixel in the stitched image to a specific class (e.g., road, building, trees) and to an individual instance within that class. This dual classification enables the system to generate detailed masks or contours for each object instance, allowing for precise obstacle detection and localization. The panoptic segmentation process enhances the system's ability to interact with complex environments by accurately identifying and differentiating between multiple objects, leading to improved navigation, obstacle avoidance, and scene understanding. This level of detail is crucial for applications that require a high degree of situational awareness and interaction with dynamic environments.
[0078] In some embodiments, the one or more processors are further configured to define a boundary based on an interactive user input, and the stitched image comprising one or more detected objects. Optionally, the stitched image is an image of panoramic view (e.g., 360-degree panoramic view) . In the present disclosure, the apparatus is configured to define the boundary using an interactive user input without needing the user to rotate their entire body. Optionally, the user input is an interactive input including at least one of a head movement, an eye movement, a hand gesture, or a controller input.
[0079] The functionality of gestures and controllers is similar. For controllers, the system detects the controller's position and projects a ray forward from the controller. The intersection of this ray with the ground marks the selected point. By pressing the controller button, the user can confirm the point. By holding down the button and moving the controller, the user can draw the safety boundary on the ground.
[0080] Similarly, for gestures, the system detects the hand's position and projects a ray forward from the hand. The intersection of this ray with the ground also marks the selected point. A common gesture for confirmation is the "pinch" gesture using the thumb and forefinger. By sustaining the pinch gesture, the user can draw the safety boundary on the ground.
[0081] In some embodiments, the interactive user input includes a combination of two or more of the head movement, the eye movement, the hand gesture, or the controller input. In some embodiments, the interactive user input includes a combination of the head movement and one or both of the hand gesture and the controller input. In alternative embodiments, the interactive user input includes a combination of the eye movement and one or both of the hand gesture and the controller input. In alternative embodiments, the interactive user input includes a combination of the head movement and the eye movement. By employing these methods, users can conveniently and accurately define custom safety boundaries, ensuring a safe VR experience without unnecessary physical movement.
[0082] In some embodiments, the interactive user input includes a combination of the head movement and one or both of the hand gesture and the controller input. Optionally, the one or more processors are configured to switch views within a panoramic view (e.g., a 360-degree panoramic image) based on a first user input from the head movement, and are configured to define the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0083] In some embodiments, the user can control the view within the panoramic image (e.g., the 360-degree panoramic image) by moving their head. From an ergonomics perspective, the comfortable range for head movement is ±45 degrees. In one example, when the user's head turns left or right beyond 45 degrees, the one or more processors are configured to perform a view switch. The new view direction within the 360-degree panorama is displayed within the user's comfortable head movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without straining their neck.
[0084] In some embodiments, the boundary is defined based on the second user input from the controller input. In some embodiments, the one or more processors are configured to detect the controller's position and project a ray forward in a virtual reality image. An intersection of the ray with the ground in the virtual reality image marks the selected point. The user can confirm this point by pressing the controller button and draw the boundary by holding down the button and moving the controller.
[0085] In some embodiments, the boundary is defined based on the second user input from the hand gesture. In some embodiments, the one or more processors are configured to detects the hand's position and projects a ray forward in a virtual reality image. An intersection of the ray with the ground marks the selected point. Optionally, the one or more processors are configured to detect a hand gesture. In one example, the one or more processors are configured to detect a pinch gesture wherein the pinch gesture is a gesture having the thumb and forefinger pinching each other. By sustaining the pinch gesture, the user can draw the safety boundary on the ground.
[0086] The user continues this process, switching views with head movements and defining boundary points with gestures or the controller, until the entire 360-degree boundary is defined. By combining head movement for view switching and gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.
[0087] In alternative embodiments, the interactive user input includes a combination of the eye movement and one or both of the hand gesture and the controller input. Optionally, the one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the eye movement, and are configured to define the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0088] In some embodiments, the user can control the view within the panoramic image (e.g., the 360-degree panoramic image) by moving their eye. From an ergonomics perspective, the comfortable range for eye movement is ±30 degrees. In one example, when the user's eye turns left or right beyond 30 degrees, the one or more processors are configured to perform a view switch. The new view direction within the 360-degree panorama is displayed within the user's comfortable eye movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without straining their eyes.
[0089] In some embodiments, the boundary is defined based on the second user input from the controller input. In some embodiments, the one or more processors are configured to detect the controller's position and project a ray forward in a virtual reality image. An intersection of the ray with the ground marks the selected point. The user can confirm this point by pressing the controller button and draw the boundary by holding down the button and moving the controller.
[0090] In some embodiments, the boundary is defined based on the second user input from the hand gesture. In some embodiments, the one or more processors are configured to detect the hand's position and projects a ray forward in a virtual reality image. An intersection of the ray with the ground in the virtual reality image marks the selected point. Optionally, the one or more processors are configured to detect a hand gesture. In one example, the one or more processors are configured to detect a pinch gesture wherein the pinch gesture is a gesture having the thumb and forefinger pinching each other. By sustaining the pinch gesture, the user can draw the safety boundary on the ground.
[0091] The user continues this process, switching views with eye movements and defining boundary points with gestures or the controller, until the entire 360-degree boundary is defined. By combining eye movement for view switching and gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.
[0092] In alternative embodiments, the interactive user input includes a combination of the head movement and the eye movement. When users are unable to input using gestures or controllers, this allows for head movement to control the panoramic view (e.g., the 360-degree panoramic view) and eye movement to define the boundary.
[0093] In some embodiments, the one or more processors are configured to have an initial view of a panoramic image (e.g., a 360-degree panoramic image) displayed, configured to calculate a gaze coordinate of a user's pupils in real time, identifying a specific location in the panorama image, configured to determine whether a gaze point of the user’s gaze is on the ground. In one example, the one or more processors are configured to define a boundary point based on determination that the gaze point of the user’s gaze is on the ground, and the gaze point is maintained for a duration longer than a threshold (gaze time) . In another example, the one or more processors are configured to define a boundary point based on determination that the gaze point of the user’s gaze is on the ground and a blink of the user’s eye is detected.
[0094] In some embodiments, the one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a user input from the head movement. In some embodiments, the user can control the view within the panoramic image by moving their head. From an ergonomics perspective, the comfortable range for head movement is ±45 degrees. In one example, when the user's head turns left or right beyond 45 degrees, the one or more processors are configured to perform a view switch. The new view direction within the 360-degree panorama is displayed within the user's comfortable head movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without straining their neck.
[0095] The user continuously defines the boundary by selecting points on the ground with their gaze. The selection is confirmed through sustained gaze or a blink. The head movement enables the user to switch views seamlessly, ensuring that the entire 360° environment can be covered. The user continues this process, switching views with head movements and defining boundary points with their gaze until the entire 360-degree boundary is defined. By combining head movement for view switching and eye movement for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment, even without using hand gestures or controllers.
[0096] FIG. 11 illustrates various modes of boundary definition in some embodiments according to the present disclosure. Referring to FIG. 11, the workflow prioritizes the use of gestures or controllers when received and head movements when received, to minimize the need for fine control of the eyes and head. The process begins with the system being ready to define the safety boundary. The one or more processors are configured to determine whether an input from a hand gesture or a controller is received; and determine whether an input from a head movement is received. Based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is received, the one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and configured to define the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0097] Based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is not received, the one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from an eye movement; and configured to define the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0098] Based on determination that the input from the hand gesture or the controller is not received, the one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and configured to define the boundary based on a second user input from the eye movement.
[0099] This workflow ensures that the user can effectively and comfortably set their virtual reality boundaries using the most suitable input methods available, thereby enhancing the overall user experience and safety.
[0100] The view switching functionality mentioned above aims to allow users to define custom safety boundaries without needing to physically rotate in a circle. This requires displaying every angle of the 360° panoramic image in front of the user to complete the boundary definition. Various appropriate view switching modes may be implemented in the present disclosure. Examples of view switching modes include a continuous angle switching mode, a discrete angle switching mode, and a mirror switching mode. In the continuous angle switching mode, the view switches continuously by small increments, giving the user a smooth transition through the panoramic image. In the discrete angle switching mode, the view switches in larger, discrete increments, jumping from one specific angle to another. In the mirror switching mode, the view switches by mirroring the current view to cover the opposite angle. Various view switching modes may be explained using the head movement combined with gestures or controllers.
[0101] FIG. 12 illustrates a continuous angle switching mode in some embodiments according to the present disclosure. Referring to FIG. 12, in the continuous angle switching mode, the user continuously switches the view angle within the 360° panoramic image while defining the custom boundary. In some embodiments, the one or more processors are configured to have a front perspective of the 360° panoramic image displayed, and configured to confirm a starting point for the boundary based on an input from the hand gesture or the controller. The one or more processors are configured to define the boundary based on an input from the hand gesture or the controller.
[0102] In some embodiments, the one or more processors are configured to rotate the view continuously as the user’s head moves. The 360-degree panoramic image smoothly transitions, allowing the user to maintain a consistent flow in defining the boundary. In some embodiments, the one or more processors are configured to rotate the view upon determination that the head movement reaches an angle of 45° from the center.
[0103] During the continuous view rotation, the one or more processors are configured to define boundary points based on the input from the hand gesture or the controller. The process continues until the user returns to the initial view, completing the loop of the 360-degree panoramic image. The entire 360° view is covered with continuous angle switching, ensuring a seamless and comprehensive definition of the safety boundary.
[0104] FIG. 13 illustrates a discrete angle switching mode in some embodiments according to the present disclosure. Referring to FIG. 13, in the discrete angle switching mode, the 360-degree panoramic image is divided into discrete view segments, and the one or more processors are configured to switch between these segments to define the boundary. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees.
[0105] In some embodiments, the one or more processors are configured to have a first segment of the 360° panoramic image displayed, and configured to confirm a starting point for the boundary based on an input from the hand gesture or the controller. The one or more processors are configured to define boundary points in a present segment (e.g., the first segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0106] In some embodiments, upon determination that the boundary points in the present segment are defined, the one or more processors are configured to automatically rotate the 360-degree panoramic image by 90 degrees to a next segment. The one or more processors are configured to define boundary points in the next segment (e.g., a second segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0107] The process is reiterated for each discrete segment until all segments have been covered and the boundary is fully defined. The process ends when the user returns to the initial view segment, completing the full 360-degree boundary definition. Figure 13 illustrates this process, showing how the 360-degree panoramic image is divided into discrete segments. The arrows indicate the rotation from one segment to the next, with the user defining the boundary within each discrete view. This mode allows for efficient and structured boundary definition without requiring continuous head movement.
[0108] FIG. 14 illustrates a mirror switching mode in some embodiments according to the present disclosure. Referring to FIG. 14, in the mirror switching mode, the 360-degree panoramic image is divided into discrete view segments with a mirrored approach to minimize user input and simplify the boundary definition process. The one or more processors are configured to switch between these segments to define the boundary. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees. Segments 1 and 3 remain unchanged, while segments 2 and 4 are mirrored horizontally.
[0109] In some embodiments, the one or more processors are configured to have a present segment (e.g., the first segment) of the 360° panoramic image displayed, and configured to confirm a starting point for the boundary based on an input from the hand gesture or the controller. The one or more processors are configured to define boundary points in a present segment (e.g., the first segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0110] In some embodiments, upon determination that the boundary points in the present segment are defined, the one or more processors are configured to automatically mirror the view of the second segment, displaying it as if the user is looking in the opposite direction. The one or more processors are configured to define boundary points in the next segment (e.g., a second segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0111] The process is reiterated for each discrete segment until all segments have been covered and the boundary is fully defined. For example, upon determination that the boundary points in the second segment are defined, the one or more processors are configured to automatically mirror the view of the third segment, displaying it as if the user is looking in the opposite direction. The one or more processors are configured to define boundary points in the third segment of the 360-degree panoramic image, based on an input from the hand gesture or the controller. Upon determination that the boundary points in the third segment are defined, the one or more processors are configured to automatically mirror the view of the fourth segment, displaying it as if the user is looking in the opposite direction. The one or more processors are configured to define boundary points in the fourth segment of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0112] The user completes the boundary definition by moving through the mirrored segments and the unchanged segments in a sequence, minimizing the need for repeated confirmations and cancellations. The process ends when the user has covered all segments and returns to the initial segment, completing the full 360-degree boundary definition. Figure 14 illustrates this process, showing how the 360-degree panoramic image is divided into mirrored and unchanged segments. The arrows indicate the transition between segments, with the user defining the boundary in each mirrored and unchanged view. This method reduces the number of input confirmations required and streamlines the boundary definition process.
[0113] Upon completion of defining boundary points, the one or more processors are configured to have the boundary displayed (e.g. as a whole) for review and / or adjustment. This can be shown in an overhead view, as illustrated in FIG. 15, where the user can make necessary adjustments and confirm the boundary.
[0114] In some embodiments, the one or more processors are configured to have the boundary displayed in an overhead view, and are configured to automatically detect any obstacles within the boundary.
[0115] In some embodiments, upon detecting an obstacle, the one or more processors are configured to highlight the obstacle, and are configured to prompt a user to address the obstacle. In one example, the one or more processors are configured to provide a recommendation to remove the obstacle in the boundary to ensure safety. In some embodiment, upon determination that the obstacle is not removed, the one or more processors are configured to automatically exclude the obstacle from the boundary to maximize safety.
[0116] In some embodiments, the one or more processors are configured to receive a user input to adjust the boundary, e.g., to expand or contract the boundary as needed. In some embodiments, the one or more processors are configured to receive a user input for select a specific viewpoint in the overhead view for adjustment, and are configured to have the specific viewpoint displayed in a front comfortable viewing area. In some embodiments, the one or more processors are configured to adjust the boundary based on an input of the hand gesture, the controller, or the eye movement.
[0117] After making necessary adjustments, the user confirms the final custom safety boundary. The confirmed boundary is then used to ensure the user's safety during the virtual reality experience. FIG. 15 illustrates a process where an obstacle within an initial boundary is identified and either highlighted for removal or automatically excluded from a safety area. The overhead view allows the user to see the entire boundary and make precise adjustments before final confirmation. This ensures the safety boundary is accurately defined and free from obstacles, providing a secure VR environment for the user.
[0118] In another aspect, the present disclosure provides a method of defining a boundary. FIG. 16 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure. Referring to FIG. 16, the method in some embodiments includes obtaining a plurality of original images captured by the plurality of cameras; performing an image stitching process on the plurality of original images to obtain a stitched image; performing object detection and localization based on the stitched image; defining a boundary based on an interactive user input, and the stitched image comprising one or more detected objects; and having the boundary displayed for review and / or adjustment.
[0119] In some embodiments, the method includes performing an image stitching process on the plurality of original images to obtain a stitched image. Optionally, the stitched image is a stitched image of 360-degree panoramic view. FIG. 17 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure. Referring to FIG. 17, in some embodiments, performing the image stitching process on the plurality of original images to obtain the stitched image includes at least one of performing an image preprocessing process on the plurality of original images to obtain a plurality of pre-processed images; performing a feature point detection and matching process; performing perspective transformation; or performing an image postprocessing process to obtain the stitched image.
[0120] In some embodiments, the method includes performing an image preprocessing process on the plurality of original images to obtain a plurality of pre-processed images. Optionally, the image preprocessing process includes denoising to remove noise. Optionally, the image preprocessing process includes image enhancement to enhance the detailed texture features of the image, improving the accuracy of feature detection.
[0121] The inventors of the present disclosure discover that the image preprocessing process is a crucial step in ensuring the accuracy and quality of the subsequent image stitching. By performing denoising, the system removes unwanted noise that can obscure important details and introduce errors in feature detection. Image enhancement further refines the pre-processed images by emphasizing texture features, edges, and other critical details, thereby facilitating more precise matching of corresponding points across images. As illustrated in FIG. 3, the left side shows an original image with potential noise and less defined features, while the right side demonstrates the enhanced image with clearer textures and reduced noise. This preprocessing not only improves the robustness of the feature detection and matching algorithms but also contributes to the overall fidelity of the stitched image, ensuring a seamless and visually coherent final output.
[0122] In some embodiments, the method includes performing a feature point detection and matching process. Optionally, the feature point detection and matching process includes extracting key feature points from each image, such as corners, edges, and textures, and determines the correspondence between images based on matched feature points.
[0123] The inventors of the present disclosure discover that the feature point detection and matching process is essential for ensuring the accuracy of image alignment in the stitching process. By extracting key feature points such as corners, edges, and textures from each image, the system can identify unique and stable points that can be reliably matched across multiple images. This matching process establishes the correspondence between different images, enabling the system to accurately align them. As illustrated in FIG. 4, the feature points detected in the original images are used to create a network of corresponding points, ensuring that the images can be seamlessly stitched together. This process not only facilitates the creation of a coherent and visually accurate stitched image but also helps in maintaining the geometric and photometric consistency across the combined images. The accurate matching of feature points is crucial for achieving high-quality results in applications such as virtual reality, panoramic photography, and various machine vision tasks.
[0124] In some embodiments, the method includes performing perspective transformation. Optionally, the perspective transformation includes solving a homography matrix. Optionally, the perspective transformation includes determining a mapping relationship between images based on image coordinates of matched feature points. The goal is to transform the perspective of the image to be registered to match that of the reference image, ensuring correct stitching and generating the stitched image with spatial consistency and natural appearance.
[0125] The inventors of the present disclosure discover that the perspective transformation is essential for aligning multiple images accurately in the stitching process. By solving a homography matrix, the system establishes a mapping relationship between the image coordinates of matched feature points from different images. This relationship allows the transformation of the perspective of each image to align with the reference image. As shown in FIG. 5, the transformation adjusts the grid lines from an unaligned state on the left to a correctly aligned state on the right. This ensures that the overlapping regions of the images are correctly stitched together, maintaining spatial consistency and producing the stitched image with a natural and cohesive appearance. This step is crucial for correcting geometric distortions and achieving seamless integration of multiple images into a single, panoramic view.
[0126] In some embodiments, the method includes performing an image postprocessing process to obtain a stitched image. Optionally, the stitched image is an image of panoramic view (e.g., the 360-degree panoramic view) . Optionally, the image postprocessing process includes performing an image fusion process to eliminate seams that may exist between stitched images, improving the quality of the stitched image.
[0127] The inventors of the present disclosure discover that the image postprocessing process is a critical step in the image stitching workflow, aimed at refining the stitched image. After aligning and stitching the images, the one or more processors perform an image fusion process to eliminate any visible seams between the stitched images. This step ensures that transitions between images are smooth and indistinguishable, resulting in a high-quality stitched image. As shown in FIG. 6, the left side depicts the initial stitched image with noticeable seams and misalignments, while the right side illustrates the final stitched image after postprocessing, which appears seamless and visually coherent. The image fusion process enhances the overall aesthetic and spatial consistency of the stitched image, making it suitable for applications requiring high visual fidelity, such as panoramic photography and virtual reality environments.
[0128] In some embodiments, the method further includes performing object detection and localization based on the stitched image. FIG. 18 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure. Referring to FIG. 18, in some embodiments, the method includes performing an object detection process. Optionally, the object detection process includes identifying and locating an object of interest (e.g., people, vehicles, animals) in the stitched image. Optionally, the object detection process further includes determining a bounding box and a class label of the object. FIG. 7 illustrates an object detection process in some embodiments according to the present disclosure.
[0129] The inventors of the present disclosure discover that the object detection process is integral to the system's ability to identify and interact with the environment accurately. In one example, by leveraging advanced machine learning algorithms, the system can analyze the stitched image to detect and classify objects with high precision. The bounding box, as shown in FIG. 7, serves not only to highlight the object's location but also to provide spatial context, which is crucial for subsequent processes such as obstacle avoidance and navigation. Furthermore, the class label offers semantic information that can be used to tailor interactions based on the object's type, enhancing the user's experience and safety. For instance, distinguishing between static obstacles like furniture and dynamic obstacles like people allows the system to implement more sophisticated and responsive safety measures. This dual approach of localization and classification ensures that the virtual reality environment remains both immersive and secure.
[0130] In some embodiments, the method includes performing an object segmentation process. Optionally, the object segmentation process includes segmenting the object from a background in the stitched image, separating the object from the background. Optionally, the object segmentation process further includes determining a mask or a contour of the object. FIG. 8 illustrates an object segmentation process in some embodiments according to the present disclosure.
[0131] The inventors of the present disclosure discover that the object segmentation process plays a pivotal role in distinguishing objects from their backgrounds in the stitched image. By accurately segmenting objects, the system can isolate them from the surrounding environment, which is crucial for detailed analysis and interaction. As depicted in FIG. 8, the segmentation process involves creating masks or contours that precisely outline each object, allowing for clear separation from the background. This capability enhances the system's ability to perform subsequent tasks, such as obstacle avoidance and path planning, by providing more granular information about the objects's hapes and boundaries. Additionally, object segmentation improves the system's overall understanding of the scene, facilitating more intelligent and context-aware responses to dynamic changes in the environment.
[0132] In some embodiments, the method includes performing a semantic segmentation process. Optionally, the semantic segmentation process includes assigning each subpixel in the composite to a predefined class (e.g., road, building, trees) . Optionally, semantic segmentation does not distinguish between different instances of objects but divides the entire image into different semantic regions. FIG. 9 illustrates a semantic segmentation process in some embodiments according to the present disclosure.
[0133] The inventors of the present disclosure discover that the semantic segmentation process further enhances the system's ability to understand and interpret the environment by categorizing every subpixel in the stitched image into predefined classes such as roads, buildings, and trees. Unlike object segmentation, which isolates individual objects, semantic segmentation focuses on dividing the entire image into meaningful regions based on semantic content. As illustrated in FIG. 9, each region of the image is labeled with a specific class, providing a comprehensive understanding of the scene's layout. This process does not differentiate between instances of the same class but instead groups all pixels belonging to a particular category together. This semantic information is vital for applications such as autonomous navigation and environmental mapping, where understanding the context and relationships between different regions is crucial for making informed decisions and interactions.
[0134] In some embodiments, the method includes performing a panoptic segmentation process. The panoptic segmentation process is a more advanced form of semantic segmentation that not only segments the image into different semantic regions but also distinguishes between different instances of the same class. Optionally, the panoptic segmentation process includes assigning each subpixel in the composite to a predefined class (e.g., road, building, trees) and an instance. Optionally, the panoptic segmentation process further includes determining a mask or a contour of each object instance. FIG. 10 illustrates a panoptic segmentation process in some embodiments according to the present disclosure.
[0135] The inventors of the present disclosure discover that the panoptic segmentation process combines the strengths of both semantic and instance segmentation, providing a comprehensive understanding of the scene by not only classifying each subpixel into predefined categories but also distinguishing between different instances of the same class. As depicted in FIG. 10, this process involves assigning each subpixel in the stitched image to a specific class (e.g., road, building, trees) and to an individual instance within that class. This dual classification enables the system to generate detailed masks or contours for each object instance, allowing for precise obstacle detection and localization. The panoptic segmentation process enhances the system's ability to interact with complex environments by accurately identifying and differentiating between multiple objects, leading to improved navigation, obstacle avoidance, and scene understanding. This level of detail is crucial for applications that require a high degree of situational awareness and interaction with dynamic environments.
[0136] In some embodiments, the method further includes defining a boundary based on an interactive user input, and the stitched image comprising one or more detected objects. Optionally, the stitched image is an image of panoramic view (e.g., the 360-degree panoramic view) . In the present disclosure, the method defines the boundary utilizing an interactive user input without needing the user to rotate their entire body. Optionally, the user input is an interactive input including at least one of a head movement, an eye movement, a hand gesture, or a controller input.
[0137] The functionality of gestures and controllers is similar. For controllers, the system detects the controller's position and projects a ray forward from the controller. The intersection of this ray with the ground marks the selected point. By pressing the controller button, the user can confirm the point. By holding down the button and moving the controller, the user can draw the safety boundary on the ground.
[0138] Similarly, for gestures, the system detects the hand's position and projects a ray forward from the hand. The intersection of this ray with the ground also marks the selected point. A common gesture for confirmation is the "pinch" gesture using the thumb and forefinger. By sustaining the pinch gesture, the user can draw the safety boundary on the ground.
[0139] In some embodiments, the interactive user input includes a combination of two or more of the head movement, the eye movement, the hand gesture, or the controller input. In some embodiments, the interactive user input includes a combination of the head movement and one or both of the hand gesture and the controller input. In alternative embodiments, the interactive user input includes a combination of the eye movement and one or both of the hand gesture and the controller input. In alternative embodiments, the interactive user input includes a combination of the head movement and the eye movement. By employing these methods, users can conveniently and accurately define custom safety boundaries, ensuring a safe VR experience without unnecessary physical movement.
[0140] FIG. 19 is a flow chart illustrating a method of defining a boundary in some embodiments according to the present disclosure. In some embodiments, the interactive user input includes a combination of the head movement and one or both of the hand gesture and the controller input. In some embodiments, the method includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement, and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0141] In some embodiments, the user can control the view within the panoramic image (e.g., the 360-degree panoramic image) by moving their head. From an ergonomics perspective, the comfortable range for head movement is ±45 degrees. In one example, when the user's head turns left or right beyond 45 degrees, the one or more processors are configured to perform a view switch. The new view direction within the 360-degree panorama is displayed within the user's comfortable head movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without straining their neck.
[0142] In some embodiments, the boundary is defined based on the second user input from the controller input. In some embodiments, the method includes detecting the controller's position and project a ray forward in a virtual reality image. An intersection of the ray with the ground in the virtual reality image marks the selected point. The user can confirm this point by pressing the controller button and draw the boundary by holding down the button and moving the controller.
[0143] In some embodiments, the boundary is defined based on the second user input from the hand gesture. In some embodiments, the method includes detecting the hand's position and projects a ray forward in a virtual reality image. An intersection of the ray with the ground marks the selected point. Optionally, the method includes detecting a hand gesture. In one example, the method includes detecting a pinch gesture wherein the pinch gesture is a gesture having the thumb and forefinger pinching each other. By sustaining the pinch gesture, the user can draw the safety boundary on the ground.
[0144] The user continues this process, switching views with head movements and defining boundary points with gestures or the controller, until the entire 360-degree boundary is defined. By combining head movement for view switching and gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.
[0145] In alternative embodiments, the interactive user input includes a combination of the eye movement and one or both of the hand gesture and the controller input. In some embodiments, the method includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the eye movement, and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0146] In some embodiments, the user can control the view within the panoramic image (e.g., the 360-degree panoramic image) by moving their eye. From an ergonomics perspective, the comfortable range for eye movement is ±30 degrees. In one example, when the user's eye turns left or right beyond 30 degrees, the one or more processors are configured to perform a view switch. The new view direction within the 360-degree panorama is displayed within the user's comfortable eye movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without straining their eyes.
[0147] In some embodiments, the boundary is defined based on the second user input from the controller input. In some embodiments, the method includes detecting the controller's position and project a ray forward in a virtual reality image. An intersection of the ray with the ground marks the selected point. The user can confirm this point by pressing the controller button and draw the boundary by holding down the button and moving the controller.
[0148] In some embodiments, the boundary is defined based on the second user input from the hand gesture. In some embodiments, the method includes detecting the hand's position and projects a ray forward in a virtual reality image. An intersection of the ray with the ground in the virtual reality image marks the selected point. Optionally, the method includes detecting a hand gesture. In one example, the method includes detecting a pinch gesture wherein the pinch gesture is a gesture having the thumb and forefinger pinching each other. By sustaining the pinch gesture, the user can draw the safety boundary on the ground.
[0149] The user continues this process, switching views with eye movements and defining boundary points with gestures or the controller, until the entire 360-degree boundary is defined. By combining eye movement for view switching and gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.
[0150] In alternative embodiments, the interactive user input includes a combination of the head movement and the eye movement. When users are unable to input using gestures or controllers, this allows for head movement to control the 360-degree panoramic view and eye movement to define the boundary. In some embodiments, the method includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement, and defining the boundary based on a second user input from the eye movement.
[0151] In some embodiments, the method includes having an initial view of a panoramic image (e.g., a 360-degree panoramic image) displayed, calculating a gaze coordinate of a user's pupils in real time, identifying a specific location in the panorama image, determining whether a gaze point of the user’s gaze is on the ground. In one example, the method includes defining a boundary point based on determination that the gaze point of the user’s gaze is on the ground, and the gaze point is maintained for a duration longer than a threshold (gaze time) . In another example, the method includes defining a boundary point based on determination that the gaze point of the user’s gaze is on the ground and a blink of the user’s eye is detected.
[0152] In some embodiments, the method includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a user input from the head movement. In some embodiments, the user can control the view within the panoramic image by moving their head. From an ergonomics perspective, the comfortable range for head movement is ±45 degrees. In one example, when the user's head turns left or right beyond 45 degrees, the one or more processors are configured to perform a view switch. The new view direction within the 360-degree panorama is displayed within the user's comfortable head movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without straining their neck.
[0153] The user continuously defines the boundary by selecting points on the ground with their gaze. The selection is confirmed through sustained gaze or a blink. The head movement enables the user to switch views seamlessly, ensuring that the entire 360° environment can be covered. The user continues this process, switching views with head movements and defining boundary points with their gaze until the entire 360-degree boundary is defined. By combining head movement for view switching and eye movement for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment, even without using hand gestures or controllers.
[0154] Referring to FIG. 11, the method in some embodiments includes determining whether an input from a hand gesture or a controller is received; and determining whether an input from a head movement is received. Based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is received, the method in some embodiments further includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0155] Based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is not received, the method in some embodiments further includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from an eye movement; and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0156] Based on determination that the input from the hand gesture or the controller is not received, the method in some embodiments further includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and defining the boundary based on a second user input from the eye movement.
[0157] This workflow ensures that the user can effectively and comfortably set their virtual reality boundaries using the most suitable input methods available, thereby enhancing the overall user experience and safety.
[0158] The view switching functionality mentioned above aims to allow users to define custom safety boundaries without needing to physically rotate in a circle. This requires displaying every angle of the 360° panoramic image in front of the user to complete the boundary definition. Various appropriate view switching modes may be implemented in the present disclosure. Examples of view switching modes include a continuous angle switching mode, a discrete angle switching mode, and a mirror switching mode. In the continuous angle switching mode, the view switches continuously by small increments, giving the user a smooth transition through the panoramic image. In the discrete angle switching mode, the view switches in larger, discrete increments, jumping from one specific angle to another. In the mirror switching mode, the view switches by mirroring the current view to cover the opposite angle. Various view switching modes may be explained using the head movement combined with gestures or controllers.
[0159] Referring to FIG. 12, in the continuous angle switching mode, the user continuously switches the view angle within the 360° panoramic image while defining the custom boundary. In some embodiments, the method further includes having a front perspective of the 360° panoramic image displayed, and confirming a starting point for the boundary based on an input from the hand gesture or the controller. The method further includes defining the boundary based on an input from the hand gesture or the controller.
[0160] In some embodiments, the method further includes rotating the view continuously as the user’s head moves. The 360-degree panoramic image smoothly transitions, allowing the user to maintain a consistent flow in defining the boundary. In some embodiments, the method further includes rotating the view upon determination that the head movement reaches an angle of 45° from the center.
[0161] During the continuous view rotation, the method includes defining boundary points based on the input from the hand gesture or the controller. The process continues until the user returns to the initial view, completing the loop of the 360-degree panoramic image. The entire 360° view is covered with continuous angle switching, ensuring a seamless and comprehensive definition of the safety boundary.
[0162] Referring to FIG. 13, in the discrete angle switching mode, the 360-degree panoramic image is divided into discrete view segments, and the method includes switching between these segments to define the boundary. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees.
[0163] In some embodiments, the method includes having a first segment of the 360°panoramic image displayed, and confirming a starting point for the boundary based on an input from the hand gesture or the controller. The method includes defining boundary points in a present segment (e.g., the first segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0164] In some embodiments, upon determination that the boundary points in the present segment are defined, the method includes automatically rotating the 360-degree panoramic image by 90 degrees to a next segment. The method includes defining boundary points in the next segment (e.g., a second segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0165] The process is reiterated for each discrete segment until all segments have been covered and the boundary is fully defined. The process ends when the user returns to the initial view segment, completing the full 360-degree boundary definition. Figure 13 illustrates this process, showing how the 360-degree panoramic image is divided into discrete segments. The arrows indicate the rotation from one segment to the next, with the user defining the boundary within each discrete view. This mode allows for efficient and structured boundary definition without requiring continuous head movement.
[0166] Referring to FIG. 14, in the mirror switching mode, the 360-degree panoramic image is divided into discrete view segments with a mirrored approach to minimize user input and simplify the boundary definition process. The method includes switching between these segments to define the boundary. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees. Segments 1 and 3 remain unchanged, while segments 2 and 4 are mirrored horizontally.
[0167] In some embodiments, the method includes having a present segment (e.g., the first segment) of the 360° panoramic image displayed, and confirming a starting point for the boundary based on an input from the hand gesture or the controller. The method includes defining boundary points in a present segment (e.g., the first segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0168] In some embodiments, upon determination that the boundary points in the present segment are defined, the method includes automatically mirroring the view of the second segment, displaying it as if the user is looking in the opposite direction. The method includes defining boundary points in the next segment (e.g., a second segment) of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0169] The process is reiterated for each discrete segment until all segments have been covered and the boundary is fully defined. For example, upon determination that the boundary points in the second segment are defined, the method includes automatically mirroring the view of the third segment, displaying it as if the user is looking in the opposite direction. The method includes defining boundary points in the third segment of the 360-degree panoramic image, based on an input from the hand gesture or the controller. Upon determination that the boundary points in the third segment are defined, the method includes automatically mirroring the view of the fourth segment, displaying it as if the user is looking in the opposite direction. The method includes defining boundary points in the fourth segment of the 360-degree panoramic image, based on an input from the hand gesture or the controller.
[0170] The user completes the boundary definition by moving through the mirrored segments and the unchanged segments in a sequence, minimizing the need for repeated confirmations and cancellations. The process ends when the user has covered all segments and returns to the initial segment, completing the full 360-degree boundary definition. Figure 14 illustrates this process, showing how the 360-degree panoramic image is divided into mirrored and unchanged segments. The arrows indicate the transition between segments, with the user defining the boundary in each mirrored and unchanged view. This method reduces the number of input confirmations required and streamlines the boundary definition process.
[0171] In some embodiments, the method further includes having the boundary displayed for review and / or adjustment, upon completion of defining boundary points. This can be shown in an overhead view, as illustrated in FIG. 15, where the user can make necessary adjustments and confirm the boundary.
[0172] In some embodiments, the method includes having the boundary displayed in an overhead view, and automatically detecting any obstacles within the boundary.
[0173] In some embodiments, upon detecting an obstacle, the method includes highlighting the obstacle, and prompting a user to address the obstacle. In one example, the method includes providing a recommendation to remove the obstacle in the boundary to ensure safety. In some embodiment, upon determination that the obstacle is not removed, the method includes automatically excluding the obstacle from the boundary to maximize safety.
[0174] In some embodiments, the method includes receiving a user input to adjust the boundary, e.g., to expand or contract the boundary as needed. In some embodiments, the method includes receiving a user input for select a specific viewpoint in the overhead view for adjustment, and having the specific viewpoint displayed in a front comfortable viewing area. In some embodiments, the method includes adjusting the boundary based on an input of the hand gesture, the controller, or the eye movement.
[0175] After making necessary adjustments, the user confirms the final custom safety boundary. The confirmed boundary is then used to ensure the user's safety during the virtual reality experience. FIG. 15 illustrates a process where an obstacle within an initial boundary is identified and either highlighted for removal or automatically excluded from a safety area. The overhead view allows the user to see the entire boundary and make precise adjustments before final confirmation. This ensures the safety boundary is accurately defined and free from obstacles, providing a secure VR environment for the user.
[0176] In another aspect, the present disclosure provides a a computer-program product comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon. In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to perform obtaining a plurality of original images captured by a plurality of cameras; performing an image stitching process on the plurality of original images to obtain a stitched image; performing object detection and localization based on the stitched image; defining a boundary based on an interactive user input, and the stitched image comprising one or more detected objects; and having the boundary displayed for review and / or adjustment. Optionally, the stitched image is a stitched image of panoramic view (e.g., 360-degree panoramic view) .
[0177] In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform defining a boundary based on an interactive user input, and the stitched image comprising one or more detected objects, wherein the stitched image is an image of panoramic view (e.g., 360-degree panoramic view) .
[0178] In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first interactive user input; and defining the boundary based on a second interactive user input. Optionally, the first interactive user input and the second interactive user input are different from each other.
[0179] In some embodiments, the interactive user input includes a combination of the head movement and one or both of the hand gesture and the controller input. In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0180] In some embodiments, the interactive user input includes a combination of the eye movement and one or both of the hand gesture and the controller input. In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the eye movement; and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0181] In some embodiments, the interactive user input includes a combination of the head movement and the eye movement. In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and defining the boundary based on a second user input from the eye movement.
[0182] In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform determining whether an input from a hand gesture or a controller is received; and determining whether an input from a head movement is received.
[0183] In some embodiments, based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is received, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0184] In some embodiments, based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is not received, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from an eye movement; and defining the boundary based on a second user input from one or both of the hand gesture and the controller input.
[0185] In some embodiments, based on determination that the input from the hand gesture or the controller is not received, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and defining the boundary based on a second user input from the eye movement.
[0186] In some embodiments, upon completion of defining boundary points, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform having the boundary displayed for review and / or adjustment.
[0187] In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform automatically detecting any obstacles within the boundary; upon detecting an obstacle, highlighting the obstacle, and prompting a user to address the obstacle; and receiving a user input to adjust the boundary.
[0188] In some embodiments, upon determination that the obstacle is not removed by the user, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform automatically excluding the obstacle from the boundary to maximize safety.
[0189] In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform receiving a user input for select a specific viewpoint in the overhead view for adjustment; having the specific viewpoint displayed in a front comfortable viewing area; and adjusting the boundary based on an interactive user input.
[0190] In some embodiments, to perform the image stitching process on the plurality of original images to obtain the stitched image, the computer-readable instructions are executable by one or more processors to cause the one or more processors to perform an image preprocessing process on the plurality of original images to obtain a plurality of pre-processed images; a feature point detection and matching process, extracting key feature points from each image, and determining correspondence between images based on matched feature points; perspective transformation, determining a mapping relationship between images based on image coordinates of matched feature points; and an image postprocessing process to obtain the stitched image.
[0191] In some embodiments, the computer-readable instructions are executable by one or more processors to cause the one or more processors to further perform object detection and localization based on the stitched image.
[0192] In some embodiments, to perform object detection and localization based on the stitched image, the computer-readable instructions are executable by one or more processors to cause the one or more processors to perform an object detection process; an object segmentation process; and a semantic segmentation process and / or a panoptic segmentation process.
[0193] All or some of steps of the method, functional modules / units in the system and the device disclosed above may be implemented as software, firmware, hardware, or suitable combinations thereof. In a hardware implementation, a division among functional modules / units mentioned in the above description does not necessarily correspond to the division among physical components. For example, one physical component may have a plurality of functions, or one function or step may be performed by several physical components in cooperation. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium) . The term computer storage medium includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules or other data, as is well known to one of ordinary skill in the art. A computer storage medium includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, Digital Versatile Disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other medium which may be used to store desired information, and which may be accessed by a computer. In addition, a communication medium typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery medium, as is well known to one of ordinary skill in the art.
[0194] The foregoing description of the embodiments of the invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form or to exemplary embodiments disclosed. Accordingly, the foregoing description should be regarded as illustrative rather than restrictive. Obviously, many modifications and variations will be apparent to practitioners skilled in this art. The embodiments are chosen and described in order to explain the principles of the invention and its best mode practical application, thereby to enable persons skilled in the art to understand the invention for various embodiments and with various modifications as are suited to the particular use or implementation contemplated. It is intended that the scope of the invention be defined by the claims appended hereto and their equivalents in which all terms are meant in their broadest reasonable sense unless otherwise indicated. Therefore, the term “the invention” , “the present invention” or the like does not necessarily limit the claim scope to a specific embodiment, and the reference to exemplary embodiments of the invention does not imply a limitation on the invention, and no such limitation is to be inferred. The invention is limited only by the spirit and scope of the appended claims. Moreover, these claims may refer to use “first” , “second” , etc. following with noun or element. Such terms should be understood as a nomenclature and should not be construed as giving the limitation on the number of the elements modified by such nomenclature unless specific number has been given. Any advantages and benefits described may not apply to all embodiments of the invention. It should be appreciated that variations may be made in the embodiments described by persons skilled in the art without departing from the scope of the present invention as defined by the following claims. Moreover, no element and component in the present disclosure is intended to be dedicated to the public regardless of whether the element or component is explicitly recited in the following claims.
Claims
1.An apparatus, comprising:a memory; andone or more processors;wherein the memory and the one or more processors are connected with each other; andthe memory stores computer-executable instructions for controlling the one or more processors to:obtain a plurality of original images captured by the plurality of cameras;perform an image stitching process on the plurality of original images to obtain a stitched image;perform object detection and localization based on the stitched image;define a boundary based on a user input, and a result of the object detection and localization; andhave the boundary displayed for review and / or adjustment.2.The apparatus of claim 1, wherein the memory stores computer-executable instructions for controlling the one or more processors to:switch views within the stitched image based on a first user input; anddefine the boundary based on a second user input, and the stitched image comprising one or more detected objects;wherein the first user input and the second user input are different from each other.3.The apparatus of claim 1, wherein the user input includes a combination of the head movement and one or both of the hand gesture and the controller input;the memory stores computer-executable instructions for controlling the one or more processors to:switch views within the stitched image based on a first user input from the head movement; anddefine the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.4.The apparatus of claim 1, wherein the user input includes a combination of the eye movement and one or both of the hand gesture and the controller input;the memory stores computer-executable instructions for controlling the one or more processors to:switch views within the stitched image based on a first user input from the eye movement; anddefine the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.5.The apparatus of claim 1, wherein the user input includes a combination of the head movement and the eye movement;the memory stores computer-executable instructions for controlling the one or more processors to:switch views within the stitched image based on a first user input from the head movement; anddefine the boundary based on a second user input from the eye movement, and the stitched image comprising one or more detected objects.6.The apparatus of claim 1, wherein the memory stores computer-executable instructions for controlling the one or more processors to:determine whether an input from a hand gesture or a controller is received; anddetermine whether an input from a head movement is received.7.The apparatus of claim 6, wherein, based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is received, the memory stores computer-executable instructions for controlling the one or more processors to:switch views within the stitched image based on a first user input from the head movement; anddefine the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.8.The apparatus of claim 6, wherein, based on determination that the input from the hand gesture or the controller is received, and the input from the head movement is not received, the memory stores computer-executable instructions for controlling the one or more processors to:switch views within the stitched image based on a first user input from an eye movement; anddefine the boundary based on a second user input from one or both of the hand gesture and the controller input, and the stitched image comprising one or more detected objects.9.The apparatus of claim 6, wherein, based on determination that the input from the hand gesture or the controller is not received, the memory stores computer-executable instructions for controlling the one or more processors to:switch views within the stitched image based on a first user input from the head movement; anddefine the boundary based on a second user input from the eye movement, and the stitched image comprising one or more detected objects.10.The apparatus of any one of claims 1 to 9, wherein, upon completion of defining an entire 360-degree boundary, the memory stores computer-executable instructions for controlling the one or more processors to have the boundary displayed for review and / or adjustment.11.The apparatus of claim 10, wherein the memory stores computer-executable instructions for controlling the one or more processors to:detect obstacles within the boundary;upon detecting an obstacle, highlight the obstacle, and prompt a user to address the obstacle; andreceive a user input to adjust the boundary.12.The apparatus of claim 11, wherein, upon determination that the obstacle is not removed, the memory stores computer-executable instructions for controlling the one or more processors to exclude the obstacle from the boundary to maximize safety.13.The apparatus of claim 11, wherein the memory stores computer-executable instructions for controlling the one or more processors to:receive a user input for select a viewpoint for adjustment;have the specific viewpoint displayed in a front viewing area; andadjust the boundary based on an additional user input.14.The apparatus of any one of claims 1 to 13, wherein, to perform the image stitching process on the plurality of original images to obtain the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to:perform an image preprocessing process on the plurality of original images to obtain a plurality of pre-processed images;perform a feature point detection and matching process, extracting one or more feature points from each image, and determining correspondence between images based on matched feature points;perform perspective transformation, determining a mapping relationship between images based on image coordinates of matched feature points; andperform an image postprocessing process to obtain the stitched image.15.The apparatus of any one of claims 1 to 14, wherein the memory stores computer-executable instructions for controlling the one or more processors to perform object detection and localization based on the stitched image.16.The apparatus of claim 15, wherein, to perform object detection and localization based on the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to:perform an object detection process;perform an object segmentation process; andperform a semantic segmentation process and / or a panoptic segmentation process.17.The apparatus of any one of claims 1 to 16, wherein the stitched image is a stitched image of a panoramic view.18.A display apparatus, comprising the apparatus of any one of claims 1 to 17, a plurality of cameras, and a display panel connected to the apparatus.19.A method of defining a boundary, comprising:obtaining a plurality of original images captured by a plurality of cameras;performing an image stitching process on the plurality of original images to obtain a stitched image;performing object detection and localization based on the stitched image;defining a boundary based on a user input, and a result of the object detection and localization; andhaving the boundary displayed for review and / or adjustment.20.A computer-program product, comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform:obtaining a plurality of original images captured by a plurality of cameras;performing an image stitching process on the plurality of original images to obtain a stitched image;performing object detection and localization based on the stitched image;defining a boundary based on a user input, and a result of the object detection and localization; andhaving the boundary displayed for review and / or adjustment.
Citation Information
Patent Citations
Image processing method and device and video processing method and device
CN107018336A
Multi-group-articulated-vehicle perimeter video panoramic-display system and method
CN109429039A
DISTORTION CORRECTION FOR VEHICLE panoramic video CAMERA PROJECTIONS
CN110475107A
Method and system for obtaining panoramic image of ship based on machine vision
CN113191974A
Method, apparatus and computer program product for indicating a seam of an image in a corresponding area of a scene
US20180063426A1