Device, display device, method of demarcating and computer program product

By using multi-camera image stitching and interactive input, virtual reality boundaries can be automatically or manually defined, solving the problem in existing technologies where users need to rotate 360 ​​degrees to define boundaries, thus achieving a safe and comfortable virtual reality experience.

CN121752976APending Publication Date: 2026-03-27BOE TECHNOLOGY GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing virtual reality technology, users cannot safely define boundaries without rotating 360 degrees when wearing VR devices, leading to potential collision and accident risks, and existing boundary definition methods do not meet user needs.

Method used

It uses multiple cameras to capture 360-degree environmental images, and through image stitching, obstacle detection and localization, combined with interactive inputs such as head movement, eye movement, controller or posture, it automatically or manually delineates safety boundaries, and displays the results from a top-down perspective to allow users to adjust them.

Benefits of technology

It enables the safe delineation of virtual reality boundaries without rotating 360 degrees, improving the safety and comfort of the user experience, and supports the rapid setting and adjustment of custom boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121752976A_ABST
    Figure CN121752976A_ABST
Patent Text Reader

Abstract

An apparatus is provided. The apparatus includes a memory, and one or more processors. The memory and the one or more processors are connected to each other. The memory stores computer executable instructions for controlling the one or more processors to: obtain a plurality of raw images captured by the plurality of cameras; performing an image stitching process on the plurality of original images to obtain a stitched image; performing target detection and positioning based on the spliced image; delimiting boundaries based on the interactive user input and the results of target detection and localization; and causing the boundary to be displayed for viewing and / or adjustment. The stitched image is a stitched image of a panoramic field of view.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to display technology, and more particularly to an apparatus, a display device, a method for defining boundaries, and a computer program product. Background Technology

[0002] Virtual reality (VR) is a technology that uses computer systems to create and experience virtual worlds. It simulates human sensory experiences such as sight, hearing, and touch, allowing users to immerse themselves in virtual environments and interact with virtual objects. VR technology typically requires specialized equipment, such as VR headsets, gloves, and controllers, to provide an immersive experience. When users wear VR headsets, they see virtual scenes displayed on a screen and interact with the virtual environment using head tracking and controllers. Summary of the Invention

[0003] In one aspect, this disclosure provides an apparatus comprising: a memory; and one or more processors; wherein the memory and the one or more processors are interconnected; and the memory stores computer-executable instructions for controlling the one or more processors to: acquire a plurality of raw images captured by a plurality of cameras; perform an image stitching process on the plurality of raw images to obtain a stitched image; perform target detection and localization based on the stitched image; delineate boundaries based on user input and the results of the target detection and localization; and display the boundaries for viewing and / or adjustment.

[0004] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to: switch views within the stitched image based on a first user input; and delineate the boundary based on a second user input and the stitched image including one or more detected targets; wherein the first user input and the second user input are different from each other.

[0005] Optionally, the user input includes a combination of head movements and gestures, a combination of head movements and controller input, or a combination of head movements, gestures, and controller input; the memory stores computer-executable instructions for controlling the one or more processors to: switch views within the stitched image based on a first user input from the head movements; and delineate the boundaries based on a second user input from one or both of the gestures and the controller input, and the stitched image including one or more detected targets.

[0006] Optionally, the user input includes a combination of eye movement and gesture, a combination of eye movement and controller input, or a combination of eye movement, gesture, and controller input; the memory stores computer-executable instructions for controlling the one or more processors to: switch views within the stitched image based on a first user input from the eye movement; and delineate the boundary based on a second user input from one or both of the gesture and the controller input, and the stitched image including one or more detected targets.

[0007] Optionally, the user input includes a combination of head movements and eye movements; the memory stores computer-executable instructions for controlling the one or more processors to: switch views within the stitched image based on a first user input from the head movements; and delineate the boundaries based on a second user input from the eye movements and the stitched image including one or more detected targets.

[0008] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to: determine whether input is received from a gesture or controller; and determine whether input is received from head movement.

[0009] Optionally, based on determining that input has been received from the gesture or the controller and input has been received from the head movement, the memory stores computer-executable instructions for controlling the one or more processors to: switch views within the stitched image based on a first user input from the head movement; and delineate the boundaries based on a second user input from one or both of the gesture and the controller input and the stitched image including one or more detected targets.

[0010] Optionally, based on determining that the input from the gesture or the controller has been received and no input from the head movement has been received, the memory stores computer-executable instructions for controlling the one or more processors to: switch views within the stitched image based on a first user input from eye movement; and delineate the boundary based on a second user input from one or both of the gesture and the controller input, and the stitched image including one or more detected targets.

[0011] Optionally, based on the determination that no input from the gesture or the controller is received, the memory stores computer-executable instructions for controlling the one or more processors to: switch views within the stitched image based on a first user input from the head movement; and delineate the boundary based on a second user input from the eye movement and the stitched image including one or more detected targets.

[0012] Optionally, upon completion of defining the entire 360-degree boundary, the memory stores computer-executable instructions for controlling the one or more processors to: display the boundary for viewing and / or adjustment.

[0013] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to: detect obstacles within the boundary; when an obstacle is detected, highlight the obstacle and prompt the user to handle the obstacle; and receive user input to adjust the boundary.

[0014] Optionally, when it is determined that the obstacle has not been removed, the memory stores computer-executable instructions for controlling the one or more processors to: remove the obstacle from the boundary to maximize security.

[0015] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to: receive user input for selecting a viewpoint for adjustment; display a specific viewpoint in the forward viewing area; and adjust the boundary based on further user input.

[0016] Optionally, in order to perform the image stitching process on the plurality of original images to obtain the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to: perform an image preprocessing process on the plurality of original images to obtain a plurality of preprocessed images; perform a feature point detection and matching process to extract one or more feature points from each image and determine the correspondence between images based on the matched feature points; perform a perspective transformation to determine the mapping relationship between images based on the image coordinates of the matched feature points; and perform an image postprocessing process to obtain the stitched image.

[0017] Optionally, the memory stores computer-executable instructions for controlling the one or more processors to perform target detection and localization based on the stitched image.

[0018] Optionally, in order to perform target detection and localization based on the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to: perform a target detection process; perform a target segmentation process; and perform a semantic segmentation process and / or a panoramic segmentation process.

[0019] Optionally, the stitched image is a panoramic stitched image.

[0020] On the other hand, this disclosure provides a display device, including the device described herein, a plurality of cameras, and a display panel connected to the device.

[0021] On the other hand, this disclosure provides a method for delineating boundaries, comprising: acquiring multiple original images captured by multiple cameras; performing an image stitching process on the multiple original images to obtain a stitched image; performing target detection and localization based on the stitched image; delineating boundaries based on user input and the results of the target detection and localization; and displaying the boundaries for viewing and / or adjustment.

[0022] On the other hand, this disclosure provides a computer program product including a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to: acquire a plurality of original images captured by a plurality of cameras; perform an image stitching process on the plurality of original images to obtain a stitched image; perform target detection and localization based on the stitched image; delineate boundaries based on user input and the results of the target detection and localization; and display the boundaries for viewing and / or adjustment. Attached Figure Description

[0023] The following figures are merely illustrative examples based on various disclosed embodiments and are not intended to limit the scope of the invention.

[0024] Figure 1 The method for defining boundaries is shown.

[0025] Figure 2 The method for defining boundaries is shown.

[0026] Figure 3 An image preprocessing procedure according to some embodiments of the present disclosure is illustrated.

[0027] Figure 4 The feature point detection and matching process according to some embodiments of this disclosure is illustrated.

[0028] Figure 5 Perspective transformations according to some embodiments of this disclosure are shown.

[0029] Figure 6Image post-processing procedures according to some embodiments of this disclosure are illustrated.

[0030] Figure 7 The target detection process according to some embodiments of this disclosure is illustrated.

[0031] Figure 8 The target segmentation process according to some embodiments of this disclosure is illustrated.

[0032] Figure 9 The semantic segmentation process according to some embodiments of the present disclosure is illustrated.

[0033] Figure 10 A panoramic segmentation process according to some embodiments of the present disclosure is illustrated.

[0034] Figure 11 Various patterns of boundary delineation according to some embodiments of this disclosure are shown.

[0035] Figure 12 The continuous angle switching mode according to some embodiments of the present disclosure is shown.

[0036] Figure 13 The discrete angle switching mode according to some embodiments of the present disclosure is shown.

[0037] Figure 14 The mirror switching modes according to some embodiments of the present disclosure are shown.

[0038] Figure 15 The process is shown to identify obstacles within the initial boundary and highlight them for removal or to automatically exclude them from the safe area.

[0039] Figure 16 This is a flowchart illustrating a method for delineating boundaries according to some embodiments of the present disclosure.

[0040] Figure 17 This is a flowchart illustrating a method for delineating boundaries according to some embodiments of the present disclosure.

[0041] Figure 18 This is a flowchart illustrating a method for delineating boundaries according to some embodiments of the present disclosure.

[0042] Figure 19 This is a flowchart illustrating a method for delineating boundaries according to some embodiments of the present disclosure. Detailed Implementation

[0043] This disclosure will now be described in more detail with reference to the following embodiments. It should be noted that the following description of some embodiments presented herein is for illustrative and descriptive purposes only. It is not exhaustive or limited to the precise forms disclosed.

[0044] When users wear virtual reality (VR) devices, their eyes cannot see the real world, potentially leading to collisions and accidents during the VR experience. To prevent such incidents, safety boundaries must be defined to ensure user safety. Related solutions require users to rotate 360 ​​degrees to observe and define these boundaries, which may not meet user needs.

[0045] In some embodiments, the stationary mode is used to define the boundaries. Figure 1 The method for defining boundaries is shown. (Refer to...) Figure 1 In stationary mode, a circular safety boundary is automatically drawn around the user's current location. For example, the radius of the safety boundary can be 1.7m, 2.5m, and 3.5m. During boundary setting, the system prompts the user to ensure that there are no obstacles in the surrounding area to prevent collisions with limbs or the head.

[0046] If a user reaches the safety boundary, the system switches to a perspective view to display the real environment to ensure the user's safety.

[0047] In some embodiments, a custom pattern is used to define the boundaries. Figure 2 The method for defining boundaries is shown. (Refer to...) Figure 2 In custom mode, the user manually rotates a full circle using the controllers to define a safety boundary. In the automatic boundary definition method, the user wears the virtual reality device and rotates a full circle. The system collects real-time depth information to detect obstacles and automatically defines a custom safety boundary. This boundary is then presented to the user for confirmation, allowing for further adjustments using the controllers before completion.

[0048] Once a safety boundary is defined, if a user's gesture crosses the boundary, an additional circle appears around the gesture as a warning. If the user reaches the boundary, the system switches to a perspective view to display the real environment, ensuring user safety.

[0049] The custom mode requires the user to manually define the safety boundary by rotating the controller 360 degrees. The automatic boundary definition method, when equipped with a depth camera, also requires the user to rotate 360 ​​degrees to scan the surrounding environment and automatically generate the custom boundary. This requirement for the user to rotate 360 ​​degrees does not fully meet the user's needs.

[0050] Therefore, this disclosure provides, in particular, an apparatus, a display device, a method for delineating boundaries, and a computer program product that substantially eliminates one or more problems caused by the limitations and disadvantages of the prior art. In one aspect, this disclosure provides an apparatus. In some embodiments, the apparatus includes a plurality of cameras, a memory, and one or more processors. Optionally, the memory is connected to one or more processors. Optionally, the memory stores computer-executable instructions for controlling one or more processors to: acquire a plurality of raw images captured by the plurality of cameras; perform an image stitching process on the plurality of raw images to obtain a stitched image; perform target detection and localization based on the stitched image; delineate boundaries based on interactive user input and the stitched image including one or more detected targets; and display the boundaries for viewing and / or adjustment.

[0051] This disclosure provides a method and apparatus that allows a user to define safety boundaries without rotating 360 degrees. The method and apparatus according to this disclosure utilize 360-degree environmental image capture, obstacle detection and localization, and image stitching. The method and apparatus according to this disclosure integrate head movements, eye movements, controllers, or gestures as interactive inputs to control the environmental scene. Once the boundaries are defined, the system displays the results via an overhead map view, allowing the user to make further adjustments and / or confirmation. In some embodiments, the apparatus according to this disclosure includes multiple cameras to capture 360° images of the user's surroundings without requiring the user to turn their head. The system is capable of automatically detecting and locating obstacles in the surrounding environment and then automatically generating safety boundaries. It also supports manual boundary definition, allowing the user to control the environmental scene using head movements, eye movements, controllers, or gestures without rotating a full circle. Once defined, the system displays the safety boundary results via an overhead map, allowing the user to confirm or make further adjustments.

[0052] On one hand, this disclosure provides an apparatus for delineating boundaries. In some embodiments, the apparatus includes a plurality of cameras. In some embodiments, the plurality of cameras are integrated into a virtual reality device. In one example, the plurality of cameras are distributed around the virtual reality device to allow capturing 360-degree images of the surrounding environment without requiring any user action.

[0053] In an alternative embodiment, multiple cameras are integrated into the neckwear device. This is particularly suitable for glasses-based virtual reality products. In one example, multiple cameras are distributed around the neckwear device to allow for the capture of 360-degree images of the surrounding environment without requiring any user movement.

[0054] Various suitable implementations can be practiced according to this disclosure. In some embodiments, multiple cameras are configured to actively acquire depth information. In some embodiments, the multiple cameras are structured light cameras. In one example, the structured light camera includes a projector and a receiver. The projector is configured to project a light pattern onto the surface of an object, and the receiver is configured to capture deformations of the pattern. The means for delineating boundaries also includes one or more processors configured to analyze the deformations and determine the depth of the target. As used herein, in the context of acquiring depth information, the term “deformation” refers to a change or distortion in the projected light pattern when it is projected onto the surface of an object. These deformations occur because the surface of the object is uneven, causing the light pattern to bend, stretch, compress, or otherwise change shape in response to the contours and features of the object. By capturing and analyzing these deformations, the system can infer the depth and three-dimensional shape of the object.

[0055] In some embodiments, the multiple cameras are time-of-flight (TOF) cameras. In one example, a TOF camera includes a projector and a receiver. When a light source emits a signal, it is reflected from an object and captured by the receiver. The distance between the light source and the object is calculated by measuring the propagation time of the light signal. TOF cameras provide a complete depth map of a scene in a single shot, have no scanning components, offer fast imaging speeds, and have low computational loads, making them widely used.

[0056] In one example, the device for delineating the boundary includes four cameras configured to actively acquire depth information, each with a 90° field of view.

[0057] In some embodiments, multiple cameras are configured to passively acquire depth information. In some embodiments, the multiple cameras are stereo cameras. In one example, the stereo camera includes two lenses to capture depth data by comparing the parallax between images taken from two different angles. The parallax is then used to calculate the distance to an object.

[0058] In one example, the device for delineating the boundary includes eight cameras configured to passively acquire depth information, each with a 45° field of view.

[0059] In some embodiments, the means for defining boundaries further includes one or more processors. In some embodiments, the one or more processors are configured to acquire multiple raw images captured by multiple cameras. In some embodiments, the one or more processors are also configured to perform an image stitching process on the multiple raw images to obtain a stitched image. Optionally, the stitched image is a stitched image of a panoramic field of view (e.g., a 360-degree panoramic field of view). Image stitching is a technique for combining two or more images with overlapping areas to create a single image with a wider field of view or a 360-degree panoramic field of view, thereby providing more information to support various subsequent processes. It is widely used in machine vision fields such as motion detection and tracking, augmented reality, image stabilization, resolution enhancement, and video compression. As used herein, the term "panoramic field of view" refers to a field of view equal to or greater than the human eye's field of view (typically considered to be 70° by 160°). A panoramic field of view may include 360° along a given plane, such as a horizontal plane. In one embodiment, the panoramic field of view is a spherical field of view.

[0060] In some embodiments, to perform an image stitching process on multiple original images to obtain a stitched image, one or more processors are configured to perform an image preprocessing process on the multiple original images to obtain multiple preprocessed images. Optionally, the image preprocessing process includes denoising to remove noise. Optionally, the image preprocessing process includes image enhancement to enhance the detail and texture features of the image and improve the accuracy of feature detection. Figure 3 An image preprocessing procedure according to some embodiments of the present disclosure is illustrated.

[0061] The inventors of this disclosure have discovered that image preprocessing is a crucial step in ensuring the accuracy and quality of subsequent image stitching. By performing denoising, the system removes unwanted noise that can blur important details and introduce errors in feature detection. Image enhancement further refines the preprocessed image by emphasizing texture features, edges, and other key details, thereby facilitating more accurate matching of corresponding points across images. Figure 3 As shown, the left side displays the original image with potential noise and fewer defined features, while the right side displays the enhanced image with clearer texture and reduced noise. This preprocessing not only improves the robustness of feature detection and matching algorithms but also contributes to the overall fidelity of the stitched images, ensuring a seamless and visually consistent final output.

[0062] In some embodiments, in order to perform an image stitching process on multiple original images to obtain a stitched image, one or more processors are further configured to perform a feature point detection and matching process. Optionally, the feature point detection and matching process includes extracting one or more feature points (e.g., key feature points), such as corner points, edges, and textures, from each image, and determining the correspondence between images based on the matched feature points. Figure 4The feature point detection and matching process according to some embodiments of this disclosure is illustrated.

[0063] The inventors of this disclosure have discovered that the feature point detection and matching process is essential to ensure the accuracy of image alignment during the stitching process. "Key feature points" refer to distinct and identifiable points within an image selected based on their unique and invariant properties, enabling reliable detection, description, and matching across multiple images. These points are characterized by their ability to remain consistent under various transformations such as scaling, rotation, and affine changes. Key feature points are used to establish correspondences between images, thereby facilitating processes such as image stitching, alignment, and recognition. By extracting key feature points, such as corner points, edges, and textures, from each image, the system can identify unique and stable points that can be reliably matched across multiple images. This matching process establishes correspondences between different images, enabling the system to accurately align them. Figure 4 As shown, feature points detected in the original image are used to create a network of corresponding points, ensuring that the images can be seamlessly stitched together. This process not only facilitates the generation of coherent and visually accurate stitched images but also helps maintain geometric and photometric consistency across the combined images. Precise matching of feature points is crucial for achieving high-quality results in applications such as virtual reality, panoramic photography, and various machine vision tasks.

[0064] In some embodiments, to perform an image stitching process on multiple original images to obtain a stitched image, one or more processors are further configured to perform a perspective transformation. Optionally, the perspective transformation includes solving for a homography matrix. Optionally, the perspective transformation includes determining a mapping relationship between the images based on the image coordinates of matched feature points. The aim is to transform the viewpoint of the images to be registered to match the viewpoint of the reference image, ensuring correct stitching and generating a stitched image with spatial consistency and a natural appearance. Figure 5 Perspective transformations according to some embodiments of this disclosure are shown.

[0065] The inventors of this disclosure have discovered that perspective transformation is essential for accurately aligning multiple images during the stitching process. By solving the homography matrix, the system establishes a mapping relationship between the image coordinates of matching feature points from different images. This relationship allows for the transformation of the viewpoint of each image to be aligned with a reference image. Figure 5 As shown, this transformation adjusts the grid lines from a misaligned state on the left to a correctly aligned state on the right. This ensures that overlapping areas of the image are correctly stitched together, maintaining spatial consistency and producing a visually natural stitched image. This step is crucial for correcting geometric distortion and achieving seamless integration of multiple images into a single panoramic view.

[0066] In some embodiments, in order to perform an image stitching process on multiple original images to obtain a stitched image, one or more processors are further configured to perform an image post-processing process to obtain the stitched image. Optionally, the stitched image is a panoramic view (e.g., a 360-degree panoramic view). Optionally, the image post-processing process includes performing an image fusion process to eliminate seams that may exist between the stitched images, thereby improving the quality of the stitched image. Figure 6 Image post-processing procedures according to some embodiments of this disclosure are illustrated.

[0067] The inventors of this disclosure have discovered that image post-processing is a crucial step in the image stitching workflow, aimed at refining the stitched images. After image alignment and stitching, one or more processors perform an image fusion process to eliminate any visible seams between the stitched images. This step ensures that the transitions between images are smooth and indistinguishable, resulting in high-quality stitched images. Figure 6 As shown, the left side depicts the initial stitched image with obvious seams and misalignment, while the right side shows the final stitched image after post-processing, which appears seamless and visually coherent. The image fusion process enhances the overall aesthetics and spatial consistency of the stitched image, making it suitable for applications requiring high visual fidelity, such as panoramic photography and virtual reality environments.

[0068] In some embodiments, one or more processors are also configured to perform object detection (e.g., obstacle detection) and localization based on the stitched image.

[0069] In some embodiments, for performing object detection and localization based on a stitched image, one or more processors are configured to perform an object detection process. Optionally, the object detection process includes identifying and localizing a target of interest (e.g., a person, vehicle, or animal) in the stitched image. Optionally, the object detection process also includes determining the bounding box and category label of the target. Figure 7 The target detection process according to some embodiments of this disclosure is illustrated.

[0070] The inventors of this disclosure have discovered that object detection is indispensable for a system's ability to accurately identify and interact with its environment. In one example, utilizing advanced machine learning algorithms, the system can analyze stitched images to detect and classify objects with high precision. Figure 7As shown, bounding boxes are used not only to highlight the location of targets but also to provide spatial context, which is crucial for subsequent processes such as obstacle avoidance and navigation. Furthermore, category labels provide semantic information that can be used to customize interactions based on the type of target, thereby enhancing the user experience and safety. For example, distinguishing between static obstacles like furniture and dynamic obstacles like people allows the system to implement more sophisticated and responsive safety measures. This dual approach of localization and classification ensures that the virtual reality environment remains immersive and safe.

[0071] In some embodiments, for performing target detection and localization based on a stitched image, one or more processors are configured to perform a target segmentation process. Optionally, the target segmentation process includes segmenting a target from the background in the stitched image, separating the target from the background. Optionally, the target segmentation process also includes determining a mask or contour of the target. Figure 8 The target segmentation process according to some embodiments of this disclosure is illustrated.

[0072] The inventors of this disclosure have discovered that the target segmentation process plays a crucial role in distinguishing targets from their background in stitched images. By accurately segmenting targets, the system can isolate them from their surroundings, which is essential for detailed analysis and interaction. Figure 8 As described, the segmentation process involves creating a mask or outline that precisely delineates each target, allowing for clear separation from the background. This capability enhances the system's ability to perform subsequent tasks such as obstacle avoidance and path planning by providing more refined information about the object's shape and boundaries. Furthermore, target segmentation improves the system's overall understanding of the scene, thereby facilitating a more intelligent and context-aware response to dynamic changes in the environment.

[0073] In some embodiments, to perform object detection and localization based on the stitched image, one or more processors are configured to perform a semantic segmentation process. Optionally, the semantic segmentation process includes assigning each subpixel in the synthetic image to a predefined category (e.g., road, building, tree). Optionally, the semantic segmentation does not distinguish between different instances of the object, but rather divides the entire image into different semantic regions. Figure 9 The semantic segmentation process according to some embodiments of the present disclosure is illustrated.

[0074] The inventors of this disclosure have discovered that semantic segmentation further enhances the system's ability to understand and interpret its environment by classifying each sub-pixel in a stitched image into predefined categories such as roads, buildings, and trees. Unlike target segmentation, which isolates individual targets, semantic segmentation focuses on dividing the entire image into meaningful regions based on semantic content. Figure 9As shown, each region of the image is labeled with a specific category, providing a comprehensive understanding of the scene layout. This process does not distinguish between instances of the same category, but rather groups all pixels belonging to a specific category together. This semantic information is crucial for applications such as autonomous navigation and environment mapping, where understanding the context and relationships between different regions is essential for making reliable decisions and interactions.

[0075] In some embodiments, to perform object detection and localization based on the stitched image, one or more processors are configured to perform a panoramic segmentation process. Panoramic segmentation is a more advanced form of semantic segmentation that not only segments the image into different semantic regions but also distinguishes different instances of the same category. Optionally, the panoramic segmentation process includes assigning each sub-pixel in the synthesized image to a predefined category (e.g., road, building, tree) and instance. Optionally, the panoramic segmentation process also includes determining a mask or contour for each target instance. Figure 10 The panoramic segmentation process is illustrated in some embodiments of the present disclosure.

[0076] The inventors of this disclosure have discovered that the panoramic segmentation process combines the advantages of semantic and instance segmentation, providing a comprehensive understanding of the scene by not only classifying each subpixel into a predefined category but also distinguishing different instances of the same category. For example... Figure 10 As described, this process involves assigning each subpixel in the stitched image to a specific category (e.g., road, building, tree) and an individual instance within that category. This dual classification enables the system to generate detailed masks or contours for each target instance, allowing for accurate obstacle detection and localization. The panoramic segmentation process enhances the system's ability to interact with complex environments by accurately identifying and distinguishing multiple targets, resulting in improved navigation, obstacle avoidance, and scene understanding. This level of detail is crucial for applications requiring high context awareness and interaction with dynamic environments.

[0077] In some embodiments, one or more processors are further configured to delineate boundaries based on interactive user input and a stitched image including one or more detected targets. Optionally, the stitched image is a panoramic view (e.g., a 360-degree panoramic view). In this disclosure, the device is configured to delineate boundaries using interactive user input without requiring the user to rotate their entire body. Optionally, the user input is interactive input, including at least one of head movement, eye movement, gesture, or controller input.

[0078] The gesture and controller functions similarly. For the controller, the system detects the controller's position and emits a ray forward from the controller. The intersection of this ray and the ground marks the selected point. The user can confirm this point by pressing a button on the controller. By holding down the button and moving the controller, the user can draw safety boundaries on the ground.

[0079] Similarly, for gestures, the system detects the hand's position and emits a ray forward from the hand. The intersection of this ray and the ground is also marked as the selected point. A common gesture used for confirmation is a "pinch" gesture using the thumb and forefinger. By maintaining the pinch gesture, the user can draw safety boundaries on the ground.

[0080] In some embodiments, interactive user input includes a combination of two or more of head movements, eye movements, gestures, or controller input. In some embodiments, interactive user input includes a combination of head movements and gestures, a combination of head movements and controller input, or a combination of head movements, gestures, and controller input. In alternative embodiments, interactive user input includes a combination of eye movements and gestures, a combination of eye movements and controller input, or a combination of eye movements, gestures, and controller input. In alternative embodiments, interactive user input includes a combination of head movements and eye movements. By employing these methods, users can easily and accurately define custom safety boundaries, thereby ensuring a safe VR experience without unnecessary body movement.

[0081] In some embodiments, interactive user input includes a combination of head movements and gestures, a combination of head movements and controller input, or a combination of head movements, gestures, and controller input. Optionally, one or more processors are configured to switch views within a panoramic field of view (e.g., a 360-degree panoramic image) based on a first user input from head movements, and are configured to define boundaries based on a second user input from one or both of gestures and controller input.

[0082] In some embodiments, users can control the view within a panoramic image (e.g., a 360-degree panoramic image) by moving their head. From an ergonomic perspective, a comfortable range of head movement is ±45 degrees. In one example, when the user's head turns more than 45 degrees to the left or right, one or more processors are configured to perform a view switching. The new viewing direction within the 360-degree panorama is displayed within the user's comfortable range of head movement. This ensures that users can comfortably view different directions within the 360-degree panorama without causing neck fatigue.

[0083] In some embodiments, boundaries are defined based on second user input from the controller input. In some embodiments, one or more processors are configured to detect the position of the controller and emit a ray forward in the virtual reality image. The intersection of the ray and the ground in the virtual reality image marks the selected point. The user can confirm the point by pressing a controller button and draw the boundary by holding down the button and moving the controller.

[0084] In some embodiments, boundaries are defined based on second user input from gestures. In some embodiments, one or more processors are configured to detect the position of a hand and emit a ray forward in the virtual reality image. The intersection of the ray with the ground marks the selected point. Optionally, one or more processors are configured to detect gestures. In one example, one or more processors are configured to detect a pinch gesture, where a pinch gesture is a gesture in which the thumb and forefinger are pinched together. By maintaining the pinch gesture, the user can draw a safe boundary on the ground.

[0085] The user continues the process, using head movements to switch views and using gestures or controllers to define boundary points until the entire 360-degree boundary is defined. By combining head movements for view switching with gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.

[0086] In alternative embodiments, interactive user input includes a combination of eye movement and gestures, a combination of eye movement and controller input, or a combination of eye movement, gestures, and controller input. Optionally, one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from eye movement, and are configured to define boundaries based on a second user input from one or both of gestures and controller input.

[0087] In some embodiments, users can control the view within a panoramic image (e.g., a 360-degree panoramic image) by moving their eyes. From an ergonomic perspective, the comfortable range of eye movement is ±30 degrees. In one example, when the user's eyes turn more than 30 degrees to the left or right, one or more processors are configured to perform a view switch. The new viewing direction within the 360-degree panorama is displayed within the user's comfortable eye movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without causing eye strain.

[0088] In some embodiments, boundaries are defined based on second user input from the controller input. In some embodiments, one or more processors are configured to detect the position of the controller and emit a ray forward in the virtual reality image. The intersection of the ray and the ground marks the selected point. The user can confirm the point by pressing a controller button and draw the boundary by holding down the button and moving the controller.

[0089] In some embodiments, boundaries are defined based on second user input from gestures. In some embodiments, one or more processors are configured to detect the position of a hand and emit a ray forward in the virtual reality image. The intersection of the ray and the ground in the virtual reality image marks the selected point. Optionally, one or more processors are configured to detect gestures. In one example, one or more processors are configured to detect a pinch gesture, where a pinch gesture is a gesture in which the thumb and forefinger are pinched together. By maintaining the pinch gesture, the user can draw a safe boundary on the ground.

[0090] The user continues the process, using eye tracking to switch views and using gestures or controllers to define boundary points until the entire 360-degree boundary is defined. By combining eye tracking for view switching with gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.

[0091] In an alternative embodiment, interactive user input includes a combination of head movements and eye movements. This allows head movements to control the panoramic field of view (e.g., a 360-degree panoramic field of view) and eye movements to define boundaries when the user cannot input using gestures or controllers.

[0092] In some embodiments, one or more processors are configured to display an initial view of a panoramic image (e.g., a 360-degree panoramic image), to calculate the gaze coordinates of a user's pupils in real time, to identify specific locations in the panoramic image, and to determine whether the user's gaze point is on the ground. In one example, one or more processors are configured to delineate boundary points based on determining that the user's gaze point is on the ground and that the gaze point is maintained for a duration longer than a threshold (gaze duration). In another example, one or more processors are configured to delineate boundary points based on determining that the user's gaze point is on the ground and detecting a blink of the user's eyes.

[0093] In some embodiments, one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on user input from head movements. In some embodiments, the user can control the view within the panoramic image by moving their head. From an ergonomic perspective, the comfortable range of head movements is ±45 degrees. In one example, when the user's head turns more than 45 degrees to the left or right, one or more processors are configured to perform a view switch. The new viewing direction within the 360-degree panorama is displayed within the user's comfortable range of head movements. This ensures that the user can comfortably view different directions within the 360-degree panorama without causing neck fatigue.

[0094] Users continuously define boundaries by selecting points on the ground using their gaze. Selections are confirmed through sustained gaze or blinking. Head movements allow users to seamlessly switch views, ensuring the entire 360° environment is covered. Users continue this process, using head movements to switch views and their gaze to define boundary points until the entire 360° boundary is defined. By combining head movements for view switching with eye movements for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360° environment, even without the use of gestures or controllers.

[0095] Figure 11 Various patterns of boundary delineation according to some embodiments of this disclosure are illustrated. (Refer to...) Figure 11 The workflow prioritizes using gestures or controllers when they are received, and head movements when they are received, to minimize the need for fine-grained control of the eyes and head. The process begins with the system preparing to define safety boundaries. One or more processors are configured to determine whether input from a gesture or controller is received; and to determine whether input from a head movement is received. Based on the determination that input from a gesture or controller and input from a head movement are received, one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from the head movement; and are configured to define boundaries based on a second user input from one or both of the gesture and controller inputs.

[0096] Based on the determination that input from gestures or controllers has been received, and no input from head movements has been received, one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from eye movements; and are configured to delineate boundaries based on a second user input from one or both of gesture and controller inputs.

[0097] Based on the determination that no input from gestures or controllers is received, one or more processors are configured to switch views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movements; and are configured to delineate boundaries based on a second user input from eye movements.

[0098] This workflow ensures that users can effectively and comfortably set their virtual reality boundaries using the most suitable input methods available, thereby enhancing the overall user experience and security.

[0099] The aforementioned view-switching functionality is designed to allow users to define custom safety boundaries without physically rotating a full circle. This requires displaying each angle of a 360° panoramic image in front of the user to complete the boundary demarcation. Various suitable view-switching modes can be implemented in this disclosure. Examples of view-switching modes include a continuous angle-by-angle switching mode, a discrete angle-by-angle switching mode, and a mirror switching mode. In the continuous angle-by-angle switching mode, the view switches continuously in small increments, giving the user a smooth transition through the panoramic image. In the discrete angle-by-angle switching mode, the view switches in larger discrete increments, jumping from one specific angle to another. In the mirror switching mode, the view switches by mirroring the current view to cover the opposite angle. Head movements combined with gestures or controllers can be used to interpret the various view-switching modes.

[0100] Figure 12 An angle-by-angle continuous switching mode is illustrated in some embodiments according to this disclosure. (Refer to...) Figure 12 In the angle-by-angle continuous switching mode, the user continuously switches perspectives within the 360° panoramic image while defining custom boundaries. In some embodiments, one or more processors are configured to display a frontal view of the 360° panoramic image and are configured to determine the starting point of the boundary based on input from gestures or a controller. One or more processors are configured to define the boundary based on input from gestures or a controller.

[0101] In some embodiments, one or more processors are configured to continuously rotate the view as the user's head moves. The 360-degree panoramic image transitions smoothly, allowing the user to maintain consistent coherence when defining boundaries. In some embodiments, one or more processors are configured to rotate the view when a head movement reaches an angle of 45° from the center.

[0102] During continuous view rotation, one or more processors are configured to delineate boundary points based on input from gestures or controllers. This process continues until the user returns to the initial view, completing the loop of the 360-degree panoramic image. The entire 360° view is covered with continuous angle switching, ensuring seamless and comprehensive delineation of safety boundaries.

[0103] Figure 13 Angle-by-angle discrete switching modes according to some embodiments of this disclosure are illustrated. (Refer to...) Figure 13 In the angle-by-angle discrete switching mode, the 360-degree panoramic image is divided into discrete view segments, and one or more processors are configured to switch between these segments to define boundaries. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees.

[0104] In some embodiments, one or more processors are configured to display a first segment of the 360° panoramic image and are configured to identify the starting point of the boundary based on input from gestures or a controller. One or more processors are configured to delineate boundary points in the current segment (e.g., the first segment) of the 360° panoramic image based on input from gestures or a controller.

[0105] In some embodiments, when boundary points in the current segment are determined to be delimited, one or more processors are configured to automatically rotate the 360-degree panoramic image by 90 degrees to the next segment. One or more processors are configured to delimit boundary points in the next segment (e.g., a second segment) of the 360-degree panoramic image based on input from gestures or a controller.

[0106] This process is repeated for each discrete segment until all segments are covered and the boundaries are fully defined. The process ends when the user returns to the initial view segment, thus completing the full 360-degree boundary definition. Figure 13 The process is illustrated, showing how a 360-degree panoramic image is divided into discrete segments. Arrows indicate rotations from one segment to the next, where the user defines boundaries within each discrete view. This mode allows for efficient and structured boundary delineation without requiring continuous head movement.

[0107] Figure 14 Mirroring switching modes according to some embodiments of this disclosure are shown. (Refer to...) Figure 14 In the mirror switching mode, a mirroring method is used to divide the 360-degree panoramic image into discrete view segments to minimize user input and simplify the boundary delineation process. One or more processors are configured to switch between these segments to delineate boundaries. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees. Segments 1 and 3 remain unchanged, while segments 2 and 4 are horizontally mirrored.

[0108] In some embodiments, one or more processors are configured to display a current segment (e.g., a first segment) of the 360° panoramic image and are configured to identify the starting point of the boundary based on input from a gesture or controller. One or more processors are configured to delineate boundary points within the current segment (e.g., the first segment) of the 360° panoramic image based on input from a gesture or controller.

[0109] In some embodiments, when boundary points in the current segment are determined to be delimited, one or more processors are configured to automatically mirror the view of the second segment, displaying it as if the user were viewing it from the opposite direction. One or more processors are configured to delimit boundary points in the next segment (e.g., the second segment) of the 360-degree panoramic image based on input from gestures or a controller.

[0110] This process is repeated for each discrete segment until all segments are covered and boundaries are fully defined. For example, when boundary points in the second segment are determined, one or more processors are configured to automatically mirror the view of the third segment, displaying it as if the user were viewing it from the opposite direction. One or more processors are configured to define boundary points in the third segment of the 360-degree panoramic image based on input from gestures or a controller. Upon determining that boundary points in the third segment have been defined, one or more processors are configured to automatically mirror the view of the fourth segment, displaying it as if the user were viewing it from the opposite direction. One or more processors are configured to define boundary points in the fourth segment of the 360-degree panoramic image based on input from gestures or a controller.

[0111] Users complete boundary delineation by sequentially mirroring and reverting to unchanged segments, minimizing the need for repeated confirmations and cancellations. The process ends when the user has covered all segments and returned to the initial segment, thus completing the full 360-degree boundary delineation. Figure 14 The process is illustrated, showing how a 360-degree panoramic image is divided into mirrored and unaltered segments. Arrows indicate the transitions between segments, and the user defines the boundaries in each mirrored and unaltered view. This method reduces the amount of input confirmation required and simplifies the boundary definition process.

[0112] Once the boundary points have been defined, one or more processors are configured (e.g., as a whole) to display the boundaries for viewing and / or adjustment. This can be shown in a top view, such as... Figure 15 As shown, users can make necessary adjustments and confirm the boundaries.

[0113] In some embodiments, one or more processors are configured to display the boundary in a top view and are configured to automatically detect any obstacles within the boundary.

[0114] In some embodiments, when an obstacle is detected, one or more processors are configured to highlight the obstacle and prompt the user to handle it. In one example, one or more processors are configured to provide a suggestion to remove the obstacle from the boundary to ensure safety. In some embodiments, when it is determined that the obstacle has not been removed, one or more processors are configured to automatically exclude the obstacle from the boundary to maximize safety.

[0115] In some embodiments, one or more processors are configured to receive user input to adjust the boundaries, for example, to expand or shrink the boundaries as needed. In some embodiments, one or more processors are configured to receive user input for selecting a specific viewpoint in a top-down view for adjustment, and are configured to display the specific viewpoint in a comfortable viewing area in front. In some embodiments, one or more processors are configured to adjust the boundaries based on gesture, controller, or eye-tracking input.

[0116] After making the necessary adjustments, the user confirms the final custom security boundary. This confirmed boundary is then used to ensure the user's safety during the virtual reality experience. Figure 15 The process demonstrates identifying obstacles within the initial boundary and highlighting them for removal or automatically excluding them from the safe zone. The top-down view allows users to see the entire boundary and make precise adjustments before final confirmation. This ensures that the safe boundary is accurately defined and free of obstacles, providing users with a safe VR environment.

[0117] On the other hand, this disclosure provides a method for defining boundaries. Figure 16 This is a flowchart illustrating a method for delineating boundaries according to some embodiments of the present disclosure. (Refer to...) Figure 16 In some embodiments, the method includes: acquiring a plurality of raw images captured by a plurality of cameras; performing an image stitching process on the plurality of raw images to obtain a stitched image; performing target detection and localization based on the stitched image; defining a boundary based on interactive user input and the stitched image including one or more detected targets; and displaying the boundary for viewing and / or adjustment.

[0118] In some embodiments, the method includes performing an image stitching process on multiple original images to obtain a stitched image. Optionally, the stitched image is a stitched image with a 360-degree panoramic view. Figure 17 This is a flowchart illustrating a method for delineating boundaries according to some embodiments of the present disclosure. (Refer to...) Figure 17In some embodiments, an image stitching process is performed on multiple original images to obtain a stitched image, including at least one of the following: performing an image preprocessing process on multiple original images to obtain multiple preprocessed images; performing a feature point detection and matching process; performing a perspective transformation; or performing an image postprocessing process to obtain a stitched image.

[0119] In some embodiments, the method includes performing an image preprocessing procedure on a plurality of original images to obtain a plurality of preprocessed images. Optionally, the image preprocessing procedure includes denoising to remove noise. Optionally, the image preprocessing procedure includes image enhancement to enhance the detail and texture features of the images and improve the accuracy of feature detection.

[0120] The inventors of this disclosure have discovered that image preprocessing is a crucial step in ensuring the accuracy and quality of subsequent image stitching. By performing denoising, the system removes unwanted noise that can blur important details and introduce errors in feature detection. Image enhancement further refines the preprocessed image by emphasizing texture features, edges, and other key details, thereby facilitating more accurate matching of corresponding points across images. Figure 3 As shown, the left side displays the original image with potential noise and fewer defined features, while the right side displays the enhanced image with clearer texture and reduced noise. This preprocessing not only improves the robustness of feature detection and matching algorithms but also contributes to the overall fidelity of the stitched images, ensuring a seamless and visually consistent final output.

[0121] In some embodiments, the method includes performing a feature point detection and matching process. Optionally, the feature point detection and matching process includes extracting key feature points, such as corner points, edges, and textures, from each image, and determining the correspondence between images based on the matched feature points.

[0122] The inventors of this disclosure have discovered that feature point detection and matching processes are essential for ensuring the accuracy of image alignment during the stitching process. By extracting key feature points, such as corner points, edges, and textures, from each image, the system can identify unique and stable points that can be reliably matched across multiple images. This matching process establishes correspondences between different images, enabling the system to accurately align them. Figure 4 As shown, feature points detected in the original image are used to create a network of corresponding points, ensuring that the images can be seamlessly stitched together. This process not only facilitates the generation of coherent and visually accurate stitched images but also helps maintain geometric and photometric consistency across the combined images. Precise matching of feature points is crucial for achieving high-quality results in applications such as virtual reality, panoramic photography, and various machine vision tasks.

[0123] In some embodiments, the method includes performing a perspective transformation. Optionally, the perspective transformation includes solving for a homography matrix. Optionally, the perspective transformation includes determining a mapping relationship between images based on the image coordinates of matched feature points. The aim is to transform the viewpoint of the image to be registered to match the viewpoint of the reference image, ensuring correct stitching and generating a stitched image with spatial consistency and a natural appearance.

[0124] The inventors of this disclosure have discovered that perspective transformation is essential for accurately aligning multiple images during the stitching process. By solving the homography matrix, the system establishes a mapping relationship between the image coordinates of matching feature points in different images. This relationship allows for the transformation of the viewpoint of each image to be aligned with a reference image. Figure 5 As shown, this transformation adjusts the grid lines from a misaligned state on the left to a correctly aligned state on the right. This ensures that overlapping areas of the image are correctly stitched together, maintaining spatial consistency and producing a visually natural stitched image. This step is crucial for correcting geometric distortion and achieving seamless integration of multiple images into a single panoramic view.

[0125] In some embodiments, the method includes performing an image post-processing procedure to obtain a stitched image. Optionally, the stitched image is a panoramic view (e.g., a 360-degree panoramic view). Optionally, the image post-processing procedure includes performing an image fusion procedure to eliminate seams that may exist between the stitched images, thereby improving the quality of the stitched image.

[0126] The inventors of this disclosure have discovered that image post-processing is a crucial step in the image stitching workflow, aimed at refining the stitched images. After image alignment and stitching, one or more processors perform an image fusion process to eliminate any visible seams between the stitched images. This step ensures that the transitions between images are smooth and indistinguishable, resulting in high-quality stitched images. Figure 6 As shown, the left side depicts the initial stitched image with obvious seams and misalignment, while the right side shows the final stitched image after post-processing, which appears seamless and visually coherent. The image fusion process enhances the overall aesthetics and spatial consistency of the stitched image, making it suitable for applications requiring high visual fidelity, such as panoramic photography and virtual reality environments.

[0127] In some embodiments, the method further includes: performing target detection and localization based on the stitched image. Figure 18 This is a flowchart illustrating a method for delineating boundaries according to some embodiments of the present disclosure. (Refer to...) Figure 18In some embodiments, the method includes performing an object detection process. Optionally, the object detection process includes identifying and locating objects of interest (e.g., people, vehicles, animals) in the stitched image. Optionally, the object detection process further includes determining bounding boxes and category labels for the objects. Figure 7 The target detection process according to some embodiments of this disclosure is illustrated.

[0128] The inventors of this disclosure have discovered that object detection is indispensable for a system's ability to accurately identify and interact with its environment. In one example, utilizing advanced machine learning algorithms, the system can analyze stitched images to detect and classify objects with high precision. Figure 7 As shown, bounding boxes are used not only to highlight the location of a target but also to provide spatial context, which is crucial for subsequent processes such as obstacle avoidance and navigation. Furthermore, category labels provide semantic information that can be used to customize interactions based on the type of target, thereby enhancing the user experience and safety. For example, distinguishing between static obstacles like furniture and dynamic obstacles like people allows the system to implement more sophisticated and responsive safety measures. This dual approach of localization and classification ensures that the virtual reality environment remains immersive and safe.

[0129] In some embodiments, the method includes performing a target segmentation process. Optionally, the target segmentation process includes segmenting a target from the background in a stitched image, separating the target from the background. Optionally, the target segmentation process further includes determining a mask or contour of the target. Figure 8 The target segmentation process according to some embodiments of this disclosure is illustrated.

[0130] The inventors of this disclosure have discovered that the target segmentation process plays a crucial role in distinguishing targets from their background in stitched images. By accurately segmenting targets, the system can isolate them from their surroundings, which is essential for detailed analysis and interaction. Figure 8 As described, the segmentation process involves creating a mask or outline that precisely delineates each target, allowing for clear separation from the background. This capability enhances the system's ability to perform subsequent tasks such as obstacle avoidance and path planning by providing more refined information about the object's shape and boundaries. Furthermore, target segmentation improves the system's overall understanding of the scene, thereby facilitating a more intelligent and context-aware response to dynamic changes in the environment.

[0131] In some embodiments, the method includes performing a semantic segmentation process. Optionally, the semantic segmentation process includes assigning each subpixel in the synthesized image to a predefined category (e.g., road, building, tree). Optionally, the semantic segmentation does not distinguish between different instances of the target, but rather divides the entire image into different semantic regions. Figure 9The semantic segmentation process according to some embodiments of the present disclosure is illustrated.

[0132] The inventors of this disclosure have discovered that semantic segmentation further enhances the system's ability to understand and interpret its environment by classifying each sub-pixel in a stitched image into predefined categories such as roads, buildings, and trees. Unlike target segmentation, which isolates individual targets, semantic segmentation focuses on dividing the entire image into meaningful regions based on semantic content. Figure 9 As shown, each region of the image is labeled with a specific category, providing a comprehensive understanding of the scene layout. This process does not distinguish between instances of the same category, but rather groups all pixels belonging to a specific category together. This semantic information is crucial for applications such as autonomous navigation and environment mapping, where understanding the context and relationships between different regions is essential for making reliable decisions and interactions.

[0133] In some embodiments, the method includes performing a panoramic segmentation process. The panoramic segmentation process is a more advanced form of semantic segmentation that not only segments an image into different semantic regions but also distinguishes different instances of the same category. Optionally, the panoramic segmentation process includes assigning each sub-pixel in the synthesized image to a predefined category (e.g., road, building, tree) and instance. Optionally, the panoramic segmentation process also includes determining a mask or contour for each target instance. Figure 10 The panoramic segmentation process is illustrated in some embodiments of the present disclosure.

[0134] The inventors of this disclosure have discovered that the panoramic segmentation process combines the advantages of semantic and instance segmentation, providing a comprehensive understanding of the scene by not only classifying each subpixel into a predefined category but also distinguishing different instances of the same category. For example... Figure 10 As described, this process involves assigning each subpixel in the stitched image to a specific category (e.g., road, building, tree) and an individual instance within that category. This dual classification enables the system to generate detailed masks or contours for each target instance, allowing for accurate obstacle detection and localization. The panoramic segmentation process enhances the system's ability to interact with complex environments by accurately identifying and distinguishing multiple targets, resulting in improved navigation, obstacle avoidance, and scene understanding. This level of detail is crucial for applications requiring high context awareness and interaction with dynamic environments.

[0135] In some embodiments, the method further includes defining boundaries based on interactive user input and a stitched image including one or more detected targets. Optionally, the stitched image is a panoramic view (e.g., a 360-degree panoramic view). In this disclosure, the method utilizes interactive user input to define boundaries without requiring the user to rotate their entire body. Optionally, the user input is interactive input, including at least one of head movement, eye movement, gesture, or controller input.

[0136] The gesture and controller functions similarly. For the controller, the system detects the controller's position and emits a ray forward from the controller. The intersection of this ray and the ground marks the selected point. The user can confirm this point by pressing a button on the controller. By holding down the button and moving the controller, the user can draw safety boundaries on the ground.

[0137] Similarly, for gestures, the system detects the hand's position and emits a ray forward from the hand. The intersection of this ray and the ground is also marked as the selected point. A common gesture used for confirmation is a "pinch" gesture using the thumb and forefinger. By maintaining the pinch gesture, the user can draw safety boundaries on the ground.

[0138] In some embodiments, interactive user input includes a combination of two or more of head movements, eye movements, gestures, or controller input. In some embodiments, interactive user input includes a combination of head movements and gestures, a combination of head movements and controller input, or a combination of head movements, gestures, and controller input. In alternative embodiments, interactive user input includes a combination of eye movements and gestures, a combination of eye movements and controller input, or a combination of eye movements, gestures, and controller input. In alternative embodiments, interactive user input includes a combination of head movements and eye movements. By employing these methods, users can easily and accurately define custom safety boundaries, thereby ensuring a safe VR experience without unnecessary body movement.

[0139] Figure 19 This is a flowchart illustrating a method for defining boundaries according to some embodiments of the present disclosure. In some embodiments, interactive user input includes a combination of head movements and gestures, a combination of head movements and controller input, or a combination of head movements, gestures, and controller input. In some embodiments, the method includes: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movements, and defining boundaries based on a second user input from one or both of gestures and controller input.

[0140] In some embodiments, users can control the view within a panoramic image (e.g., a 360-degree panoramic image) by moving their head. From an ergonomic perspective, a comfortable range of head movement is ±45 degrees. In one example, when the user's head turns more than 45 degrees to the left or right, one or more processors are configured to perform a view switching. The new viewing direction within the 360-degree panorama is displayed within the user's comfortable range of head movement. This ensures that users can comfortably view different directions within the 360-degree panorama without causing neck fatigue.

[0141] In some embodiments, the boundary is defined based on a second user input from the controller input. In some embodiments, the method includes: detecting the position of the controller and emitting a ray forward in the virtual reality image. The intersection of the ray and the ground in the virtual reality image marks the selected point. The user can confirm the point by pressing a controller button and draw the boundary by holding down the button and moving the controller.

[0142] In some embodiments, the boundary is defined based on second user input from a gesture. In some embodiments, the method includes: detecting the position of a hand and emitting a ray forward in a virtual reality image. The intersection of the ray with the ground marks the selected point. Optionally, the method includes: detecting a gesture. In one example, the method includes: detecting a pinch gesture, wherein the pinch gesture is a gesture in which the thumb and forefinger pinch each other. By maintaining the pinch gesture, the user can draw a safe boundary on the ground.

[0143] The user continues the process, using head movements to switch views and using gestures or controllers to define boundary points until the entire 360-degree boundary is defined. By combining head movements for view switching with gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.

[0144] In alternative embodiments, interactive user input includes a combination of eye movement and gestures, a combination of eye movement and controller input, or a combination of eye movement, gestures, and controller input. In some embodiments, the method includes: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from eye movement, and defining boundaries based on a second user input from one or both of gestures and controller input.

[0145] In some embodiments, users can control the view within a panoramic image (e.g., a 360-degree panoramic image) by moving their eyes. From an ergonomic perspective, the comfortable range of eye movement is ±30 degrees. In one example, when the user's eyes turn more than 30 degrees to the left or right, one or more processors are configured to perform a view switch. The new viewing direction within the 360-degree panorama is displayed within the user's comfortable eye movement range. This ensures that users can comfortably view different directions within the 360-degree panorama without causing eye strain.

[0146] In some embodiments, the boundary is defined based on a second user input from the controller input. In some embodiments, the method includes: detecting the position of the controller and emitting a ray forward in the virtual reality image. The intersection of the ray and the ground marks the selected point. The user can confirm the point by pressing a controller button and draw the boundary by holding down the button and moving the controller.

[0147] In some embodiments, the boundary is defined based on second user input from a gesture. In some embodiments, the method includes: detecting the position of a hand and emitting a ray forward in a virtual reality image. The intersection of the ray and the ground in the virtual reality image marks the selected point. Optionally, the method includes: detecting a gesture. In one example, the method includes: detecting a pinch gesture, wherein the pinch gesture is a gesture in which the thumb and forefinger pinch each other. By maintaining the pinch gesture, the user can draw a safe boundary on the ground.

[0148] The user continues the process, using eye tracking to switch views and using gestures or controllers to define boundary points until the entire 360-degree boundary is defined. By combining eye tracking for view switching with gestures or controllers for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360-degree environment.

[0149] In an alternative embodiment, the interactive user input includes a combination of head movements and eye movements. This allows head movements to control the 360-degree panoramic view and eye movements to define boundaries when the user cannot input using gestures or controllers. In some embodiments, the method includes: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movements, and defining boundaries based on a second user input from eye movements.

[0150] In some embodiments, the method includes: displaying an initial view of a panoramic image (e.g., a 360-degree panoramic image), calculating the gaze coordinates of a user's pupils in real time, identifying a specific location in the panoramic image, and determining whether the user's gaze point is on the ground. In one example, the method includes: delineating boundary points based on determining that the user's gaze point is on the ground and that the gaze point is maintained for a duration longer than a threshold (gaze duration). In another example, the method includes: delineating boundary points based on determining that the user's gaze point is on the ground and detecting a blink of the user's eyes.

[0151] In some embodiments, the method includes switching views within a panoramic image (e.g., a 360-degree panoramic image) based on user input from head movements. In some embodiments, the user can control the view within the panoramic image by moving their head. From an ergonomic perspective, a comfortable range of head movements is ±45 degrees. In one example, one or more processors are configured to perform a view switching when the user's head turns more than 45 degrees to the left or right. The new viewing direction within the 360-degree panorama is displayed within the user's comfortable range of head movements. This ensures that the user can comfortably view different directions within the 360-degree panorama without causing neck fatigue.

[0152] Users continuously define boundaries by selecting points on the ground using their gaze. Selections are confirmed through sustained gaze or blinking. Head movements allow users to seamlessly switch views, ensuring the entire 360° environment is covered. Users continue this process, using head movements to switch views and their gaze to define boundary points until the entire 360° boundary is defined. By combining head movements for view switching with eye movements for boundary definition, users can comfortably and accurately set their virtual reality boundaries within a 360° environment, even without the use of gestures or controllers.

[0153] Reference Figure 11 In some embodiments, the method includes: determining whether input from a gesture or controller is received; and determining whether input from head movement is received. Based on determining that input from a gesture or controller and input from head movement are received, in some embodiments, the method further includes: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movement; and defining boundaries based on a second user input from one or both of the gesture and controller inputs.

[0154] Based on the determination that input from a gesture or controller has been received, and no input from head movement has been received, in some embodiments the method further includes: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from eye movement; and defining boundaries based on a second user input from one or both of the gesture and controller inputs.

[0155] Based on the determination that no input from a gesture or controller has been received, in some embodiments, the method further includes: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movement; and defining boundaries based on a second user input from eye movement.

[0156] This workflow ensures that users can effectively and comfortably set their virtual reality boundaries using the most suitable input methods available, thereby enhancing the overall user experience and security.

[0157] The aforementioned view-switching functionality is designed to allow users to define custom safety boundaries without physically rotating a full circle. This requires displaying each angle of a 360° panoramic image in front of the user to complete the boundary delineation. Various suitable view-switching modes can be implemented in this disclosure. Examples of view-switching modes include a continuous angle-by-angle switching mode, a discrete angle-by-angle switching mode, and a mirror switching mode. In the continuous angle-by-angle switching mode, the view switches continuously in small increments, giving the user a smooth transition through the panoramic image. In the discrete angle-by-angle switching mode, the view switches in larger discrete increments, jumping from one specific angle to another. In the mirror switching mode, the view switches by mirroring the current view to cover the opposite angle. Head movements combined with gestures or controllers can be used to interpret the various view-switching modes.

[0158] Reference Figure 12 In the angle-by-angle continuous switching mode, the user continuously switches viewing angles within a 360° panoramic image while defining custom boundaries. In some embodiments, the method further includes: displaying a frontal view of the 360° panoramic image, and confirming the starting point of the boundary based on input from gestures or a controller. The method also includes: defining the boundary based on input from gestures or a controller.

[0159] In some embodiments, the method further includes continuously rotating the view as the user's head moves. The 360-degree panoramic image transitions smoothly, allowing the user to maintain consistent coherence when defining boundaries. In some embodiments, the method further includes rotating the view when the head movement reaches an angle of 45° from the center.

[0160] During continuous view rotation, the method includes defining boundary points based on input from gestures or a controller. This process continues until the user returns to the initial view, completing the loop of the 360-degree panoramic image. The entire 360° view is covered with continuous angle switching, ensuring seamless and comprehensive delineation of safety boundaries.

[0161] Reference Figure 13 In the angle-by-angle discrete switching mode, the 360-degree panoramic image is divided into discrete view segments, and the method includes switching between these segments to define boundaries. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees.

[0162] In some embodiments, the method includes: displaying a first segment of a 360° panoramic image and identifying the starting point of a boundary based on input from a gesture or controller. The method also includes: defining boundary points within the current segment (e.g., the first segment) of the 360° panoramic image based on input from a gesture or controller.

[0163] In some embodiments, when determining that boundary points are delimited in the current segment, the method includes: automatically rotating the 360-degree panoramic image by 90 degrees to the next segment. The method also includes: delimiting boundary points in the next segment (e.g., a second segment) of the 360-degree panoramic image based on input from a gesture or controller.

[0164] This process is repeated for each discrete segment until all segments are covered and the boundaries are fully defined. The process ends when the user returns to the initial view segment, thus completing the full 360-degree boundary definition. Figure 13 The process is illustrated, showing how a 360-degree panoramic image is divided into discrete segments. Arrows indicate rotations from one segment to the next, where the user defines boundaries within each discrete view. This mode allows for efficient and structured boundary delineation without requiring continuous head movement.

[0165] Reference Figure 14 In the mirror switching mode, a mirroring method is used to divide the 360-degree panoramic image into discrete view segments to minimize user input and simplify the boundary delineation process. The method includes switching between these segments to delineate boundaries. In some embodiments, the 360-degree panoramic image is divided into N equal discrete view segments. In one example, the 360-degree panoramic image is divided into four segments, each covering 90 degrees. Segments 1 and 3 remain unchanged, while segments 2 and 4 are horizontally mirrored.

[0166] In some embodiments, the method includes: displaying a current segment (e.g., a first segment) of the 360° panoramic image, and confirming the starting point of a boundary based on input from a gesture or controller. The method also includes: delineating boundary points within the current segment (e.g., the first segment) of the 360° panoramic image based on input from a gesture or controller.

[0167] In some embodiments, when determining that boundary points are delimited in the current segment, the method includes: automatically mirroring the view of the second segment, displaying it as if the user were viewing it from the opposite direction. The method also includes: delimiting boundary points in the next segment (e.g., the second segment) of the 360-degree panoramic image based on input from a gesture or controller.

[0168] The process is repeated for each discrete segment until all segments are covered and boundaries are fully defined. For example, when determining that boundary points in the second segment have been defined, the method includes: automatically mirroring the view of the third segment, displaying it as if the user were viewing it from the opposite direction. The method includes: defining boundary points in the third segment of the 360-degree panoramic image based on input from gestures or a controller. When boundary points in the third segment have been defined, the method includes: automatically mirroring the view of the fourth segment, displaying it as if the user were viewing it from the opposite direction. The method includes: defining boundary points in the fourth segment of the 360-degree panoramic image based on input from gestures or a controller.

[0169] Users complete boundary delineation by sequentially mirroring and reverting to unchanged segments, minimizing the need for repeated confirmations and cancellations. The process ends when the user has covered all segments and returned to the initial segment, thus completing the full 360-degree boundary delineation. Figure 14 The process is illustrated, showing how a 360-degree panoramic image is divided into mirrored and unaltered segments. Arrows indicate the transitions between segments, and the user defines the boundaries in each mirrored and unaltered view. This method reduces the amount of input confirmation required and simplifies the boundary definition process.

[0170] In some embodiments, the method further includes: upon completion of delineating the boundary points, displaying the boundary for viewing and / or adjustment. This can be shown in a top view, such as... Figure 15 As shown, users can make necessary adjustments and confirm the boundaries.

[0171] In some embodiments, the method includes: displaying the boundary in a top view and automatically detecting any obstacles within the boundary.

[0172] In some embodiments, when an obstacle is detected, the method includes: highlighting the obstacle and prompting a user to handle the obstacle. In one example, the method includes: providing a suggestion to remove the obstacle in the boundary to ensure safety. In some embodiments, when it is determined that the obstacle has not been removed, the method includes: automatically excluding the obstacle from the boundary to maximize safety.

[0173] In some embodiments, the method includes receiving user input to adjust a boundary, for example, expanding or shrinking the boundary as needed. In some embodiments, the method includes receiving user input for selecting a specific viewpoint in a top-down view for adjustment, and displaying the specific viewpoint in a comfortable viewing area in front. In some embodiments, the method includes adjusting the boundary based on input from gestures, controllers, or eye movements.

[0174] After making the necessary adjustments, the user confirms the final custom security boundary. This confirmed boundary is then used to ensure the user's safety during the virtual reality experience. Figure 15 The process demonstrates identifying obstacles within the initial boundary and highlighting them for removal or automatically excluding them from the safe zone. The top-down view allows users to see the entire boundary and make precise adjustments before final confirmation. This ensures that the safe boundary is accurately defined and free of obstacles, providing users with a safe VR environment.

[0175] On the other hand, this disclosure provides a computer program product comprising a non-transient tangible computer-readable medium having computer-readable instructions thereon. In some embodiments, the computer-readable instructions may be executed by one or more processors to cause the one or more processors to: acquire a plurality of raw images captured by a plurality of cameras; perform an image stitching process on the plurality of raw images to obtain a stitched image; perform target detection and localization based on the stitched image; delineate boundaries based on interactive user input and the stitched image including one or more detected targets; and cause the boundaries to be displayed for viewing and / or adjustment. Optionally, the stitched image is a stitched image of a panoramic view (e.g., a 360-degree panoramic view).

[0176] In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: defining boundaries based on interactive user input and a stitched image including one or more detected targets, wherein the stitched image is a panoramic view (e.g., a 360-degree panoramic view).

[0177] In some embodiments, computer-readable instructions may be executed by one or more processors to cause the one or more processors to further perform: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first interactive user input; and defining boundaries based on a second interactive user input. Optionally, the first interactive user input and the second interactive user input are different from each other.

[0178] In some embodiments, interactive user input includes a combination of head movements and gestures, a combination of head movements and controller input, or a combination of head movements, gestures, and controller input. In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movements; and defining boundaries based on a second user input from one or both of gestures and controller input.

[0179] In some embodiments, interactive user input includes a combination of eye movement and gestures, a combination of eye movement and controller input, or a combination of eye movement, gestures, and controller input. In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from eye movement; and defining boundaries based on a second user input from one or both of gestures and controller input.

[0180] In some embodiments, interactive user input includes a combination of head movements and eye movements. In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on first user input from head movements; and defining boundaries based on second user input from eye movements.

[0181] In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: determining whether input is received from a gesture or controller; and determining whether input is received from head movement.

[0182] In some embodiments, based on determining that input has been received from a gesture or controller and input has been received from head movement, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movement; and defining boundaries based on a second user input based on one or both of the gesture and controller inputs.

[0183] In some embodiments, based on determining that input from a gesture or controller has been received, and no input from head movement has been received, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from eye movement; and defining boundaries based on a second user input from one or both of the gesture and controller inputs.

[0184] In some embodiments, based on the determination that no input from a gesture or controller has been received, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: switching views within a panoramic image (e.g., a 360-degree panoramic image) based on a first user input from head movement; and defining boundaries based on a second user input from eye movement.

[0185] In some embodiments, upon completion of the delineation of boundary points, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform actions such as displaying the boundary for viewing and / or adjustment.

[0186] In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: automatically detect any obstacles within the boundary; when an obstacle is detected, highlight the obstacle and prompt the user to handle the obstacle; and receive user input to adjust the boundary.

[0187] In some embodiments, when it is determined that the obstacle has not been removed by the user, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: automatically removing the obstacle from the boundary to maximize security.

[0188] In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further perform: receiving user input for selecting a specific viewpoint in a top view for adjustment; displaying the specific viewpoint in a comfortable viewing area in front; and adjusting the boundaries based on interactive user input.

[0189] In some embodiments, in order to perform an image stitching process on multiple original images to obtain a stitched image, computer-readable instructions may be executed by one or more processors to cause the one or more processors to: perform an image preprocessing process on the multiple original images to obtain multiple preprocessed images; perform a feature point detection and matching process to extract key feature points from each image and determine the correspondence between the images based on the matched feature points; perform a perspective transformation to determine the mapping relationship between the images based on the image coordinates of the matched feature points; and perform an image postprocessing process to obtain the stitched image.

[0190] In some embodiments, computer-readable instructions may be executed by one or more processors to cause one or more processors to further: perform object detection and localization based on the stitched image.

[0191] In some embodiments, in order to perform object detection and localization based on stitched images, computer-readable instructions may be executed by one or more processors to cause one or more processors to perform: an object detection process; an object segmentation process; and a semantic segmentation process and / or a panoramic segmentation process.

[0192] All or some steps of the methods disclosed above, the functional modules / units in the system, and the devices can be implemented as software, firmware, hardware, or a suitable combination thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division between physical components. For example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable storage medium, which may include computer storage media (or non-transient media) and communication media (or transient media). The term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data, as is known to those skilled in the art. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible by a computer. Additionally, communication media are typically embodied as computer-readable instructions, data structures, program modules, or other data in modulated data signals, such as carrier waves or other transmission mechanisms, and include any information transmission medium as is known to those skilled in the art.

[0193] For illustrative and descriptive purposes, the foregoing description of embodiments of the invention has been provided. It is not exhaustive, nor is it intended to limit the invention to the precise forms or exemplary embodiments disclosed. Therefore, the foregoing description should be considered illustrative rather than restrictive. Clearly, many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described to explain the principles of the invention and its best mode of practical application, thereby enabling those skilled in the art to understand the various embodiments of the invention and the various modifications suitable for the particular use or implementation contemplated. The scope of the invention is intended to be defined by the appended claims and their equivalents, wherein, unless otherwise stated, all terms are to be interpreted in their broadest reasonable sense. Therefore, the terms “the invention,” “the present invention,” etc., do not necessarily limit the scope of the claims to the specific embodiments, and references to exemplary embodiments of the invention do not imply limitation of the invention, nor should such limitation be inferred. The invention is defined only by the spirit and scope of the appended claims. Furthermore, these claims may involve the use of “first,” “second,” etc., followed by nouns or elements. These terms should be understood as nomenclature and should not be construed as limiting the number of elements modified by these nomenclatures unless a specific number has been given. Any advantages and benefits described may not apply to all embodiments of the invention. It should be understood that changes to the described embodiments can be made by those skilled in the art without departing from the scope of the invention as defined by the appended claims. Furthermore, the elements and components in this disclosure are not intended for public distribution, whether or not they are expressly recited in the appended claims.

Claims

1. An apparatus comprising: Memory; as well as One or more processors; Wherein, the memory and the one or more processors are connected to each other; and The memory stores computer-executable instructions for controlling the one or more processors to: Obtain multiple raw images captured by the multiple cameras; An image stitching process is performed on the multiple original images to obtain a stitched image; Target detection and localization are performed based on the stitched image; The boundary is defined based on user input and the results of target detection and localization; and The boundary is displayed for viewing and / or adjustment.

2. The apparatus according to claim 1, wherein, The memory stores computer-executable instructions for controlling the one or more processors to: Based on the first user input, switch views within the stitched image; as well as The boundary is defined based on the second user input and the stitched image including one or more detected targets; The first user input and the second user input are different from each other.

3. The apparatus according to claim 1, wherein, The user input includes a combination of head movements and gestures, a combination of head movements and controller input, or a combination of head movements, gestures, and controller input. The memory stores computer-executable instructions for controlling the one or more processors to: Based on the first user input from the head movement, switch views within the stitched image; as well as The boundary is defined based on a second user input from one or both of the gesture and the controller input, and the stitched image including one or more detected targets.

4. The apparatus according to claim 1, wherein, The user input includes a combination of eye movement and gesture, a combination of eye movement and controller input, or a combination of eye movement, gesture, and controller input. The memory stores computer-executable instructions for controlling the one or more processors to: Based on first user input from the eye movement, switch views within the stitched image; as well as The boundary is defined based on a second user input from one or both of the gesture and the controller input, and the stitched image including one or more detected targets.

5. The apparatus according to claim 1, wherein, The user input includes a combination of head movements and eye movements; The memory stores computer-executable instructions for controlling the one or more processors to: Based on the first user input from the head movement, switch views within the stitched image; as well as The boundary is defined based on a second user input from the eye movement and the stitched image including one or more detected targets.

6. The apparatus according to claim 1, wherein, The memory stores computer-executable instructions for controlling the one or more processors to: Determine if input from a gesture or controller has been received; as well as Determine whether input from head movement has been received.

7. The apparatus according to claim 6, wherein, Based on the determination that input has been received from the gesture or the controller, and input has been received from the head movement, the memory stores computer-executable instructions for controlling the one or more processors to: Based on the first user input from the head movement, switch views within the stitched image; as well as The boundary is defined based on a second user input from one or both of the gesture and the controller input, and the stitched image including one or more detected targets.

8. The apparatus according to claim 6, wherein, Based on the determination that input from the gesture or the controller has been received and no input from the head movement has been received, the memory stores computer-executable instructions for controlling the one or more processors to: Switch views within the stitched image based on first user input from eye tracking; as well as The boundary is defined based on a second user input from one or both of the gesture and the controller input, and the stitched image including one or more detected targets.

9. The apparatus according to claim 6, wherein, Based on the determination that no input has been received from the gesture or the controller, the memory stores computer-executable instructions for controlling the one or more processors to: Based on the first user input from the head movement, switch views within the stitched image; as well as The boundary is defined based on a second user input from the eye movement and the stitched image including one or more detected targets.

10. The apparatus according to any one of claims 1 to 9, wherein, Upon completion of defining the entire 360-degree boundary, the memory stores computer-executable instructions for controlling the one or more processors to: display the boundary for viewing and / or adjustment.

11. The apparatus according to claim 10, wherein, The memory stores computer-executable instructions for controlling the one or more processors to: Detect obstacles within the boundary; When an obstacle is detected, the obstacle is highlighted and the user is prompted to deal with it. as well as Receive user input to adjust the boundary.

12. The apparatus according to claim 11, wherein, When it is determined that the obstacle has not been removed, the memory stores computer-executable instructions for controlling the one or more processors to: remove the obstacle from the boundary to maximize safety.

13. The apparatus according to claim 11, wherein, The memory stores computer-executable instructions for controlling the one or more processors to: Receive user input for selecting and adjusting the viewpoint; Display a specific viewpoint in the forward field of view area; and The boundaries are adjusted based on additional user input.

14. The apparatus according to any one of claims 1 to 13, wherein, In order to perform the image stitching process on the plurality of original images to obtain the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to: The multiple original images are subjected to an image preprocessing process to obtain multiple preprocessed images; Perform a feature point detection and matching process to extract one or more feature points from each image and determine the correspondence between images based on the matched feature points; Perform perspective transformation to determine the mapping relationship between images based on the image coordinates of matched feature points; as well as A post-processing procedure is performed to obtain the stitched image.

15. The apparatus according to any one of claims 1 to 14, wherein, The memory stores computer-executable instructions for controlling the one or more processors to perform target detection and localization based on the stitched image.

16. The apparatus according to claim 15, wherein, In order to perform target detection and localization based on the stitched image, the memory stores computer-executable instructions for controlling the one or more processors to: Perform the target detection process; Perform the target segmentation process; as well as Perform semantic segmentation and / or panoptic segmentation.

17. The apparatus according to any one of claims 1 to 16, wherein, The stitched image is a panoramic image.

18. A display device comprising a device according to any one of claims 1 to 17, a plurality of cameras, and a display panel connected to said device.

19. A method for delineating boundaries, comprising: Obtain multiple raw images captured by multiple cameras; An image stitching process is performed on the multiple original images to obtain a stitched image; Target detection and localization are performed based on the stitched image; The boundary is defined based on user input and the results of target detection and localization; as well as The boundary is displayed for viewing and / or adjustment.

20. A computer program product comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform: Obtain multiple raw images captured by multiple cameras; An image stitching process is performed on the multiple original images to obtain a stitched image; Target detection and localization are performed based on the stitched image; The boundary is defined based on user input and the results of target detection and localization; as well as The boundary is displayed for viewing and / or adjustment.