A gaze object three-dimensional expression automatic collection system based on head-mounted display bidirectional monitoring
Patent Information
- Application Number
- CN202410338718.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-25
AI Technical Summary
然而,在用户在佩戴头戴式显示器移动的过程中,现有技术并不能自动检测用户的注视行为,并生成被注视物体的三维模型表达,用于新的真实和虚拟背景中的渲染和交互
[0030]1、本发明提供一种基于头显双向监测的注视物体三维表达自动收集系统,通过头戴式显示器的用户注视检测,自动识别并采集环境中被注视的物体,生成并存储该物体的三维高斯神经表达,用于融合到其他真实或者虚拟的场景中;也就是说,本发明基于用户行为,智能判断并识别环境中的被关注物体,然后收集图像信息用于三维物体生成和存储,实现了用户注视(感兴趣)物体的无感收集;由此可见,本发明提高了生成的虚实融合场景的丰富度和代入感,提高了用户体验和虚拟现实场景的交互性;在叙事融合场景中,用户熟悉的或者感兴趣的物体被自然的融合其中,增加了用户的亲切感,代入感和沉浸感,更有进行进一步交互的欲望。
Smart Images

Figure CN118279774B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of wearable devices, virtual reality technology, and artificial intelligence technology, and particularly relates to an automatic collection system for three-dimensional representation of gazed objects based on bidirectional monitoring of a head-mounted display. Background Technology
[0002] For specific data acquisition devices, such as Figure 1 The traditional methods shown typically involve directly acquiring multi-view image sequences of a scene or object, combining them with corresponding pose information to train a 3D Gaussian neural representation of the scene or object, and then rendering it. However, existing technologies cannot automatically detect the user's gaze behavior and generate a 3D model representation of the gazed object for rendering and interaction in new real and virtual backgrounds during the user's movement while wearing a head-mounted display. Summary of the Invention
[0003] To address the aforementioned issues, this invention provides an automatic collection system for 3D representations of gazed objects based on bidirectional monitoring of a head-mounted display. This system automatically detects the user's gaze behavior while the user is moving while wearing the head-mounted display and automatically generates a 3D model representation of the gazed object for rendering and interaction in new real and virtual backgrounds.
[0004] An automatic collection system for 3D representation of gazed objects based on bidirectional monitoring of a head-mounted display includes a user attention detection module, a gazed object image acquisition module, an image ROI generation module, an image semantic segmentation module, and a representation module.
[0005] The user attention detection module is used to detect whether a user wearing a head-mounted display is looking at a specific object in the surrounding environment;
[0006] The object-gazing image acquisition module is used to continuously acquire images of the object being gazed at when the user gazes at a specific object;
[0007] The image ROI generation module is used to select the corresponding object bounding box as the ROI region from continuously acquired images according to the direction of eye gaze;
[0008] The image semantic segmentation module is used to segment the ROI region from the image based on a segmentation algorithm;
[0009] The expression module is used to obtain the three-dimensional Gaussian neural expression of the object corresponding to the segmented ROI region based on the trained three-dimensional Gaussian neural expression model, and to perform virtual-real fusion of the generated object three-dimensional neural expression and the real environment collected by the head-mounted display or the stored three-dimensional virtual environment.
[0010] Furthermore, the user attention detection module determines whether the user is looking at a specific object in the surrounding environment by judging whether the user's gaze time on the same target object exceeds a set threshold.
[0011] Furthermore, the head-mounted display has an internal camera installed inside for collecting information about the user's eye movements, and an external camera installed outside for detecting images of objects in the real environment; when the user looks at a specific object, the user moves their position so that the external camera of the head-mounted display can collect images of different sides of the object being looked at.
[0012] Furthermore, the head-mounted display has an internal camera installed inside for collecting information about the user's eye movements, and an external camera installed outside for detecting images of objects in the real environment.
[0013] The image ROI generation module uses an object detection algorithm to perform object detection on the acquired image, and obtains the bounding boxes and corresponding object labels of all objects in the image.
[0014] The internal camera monitors the direction in which the user's eyes are looking, and based on this direction, it makes a location determination in the image and selects the bounding box of the object in that direction as the ROI region.
[0015] Furthermore, when there are more than two object bounding boxes in the direction the user's eyes are looking at and the representation module only produces a 3D Gaussian neural representation of one object, all object bounding boxes are listed as potential target groups to be focused on.
[0016] The location of the mobile user or potential target group changes the relative positional relationship between the user and the potential target group. Objects that move out of the user's eye gaze direction are removed from the potential target group.
[0017] Determine if there are more than two object bounding boxes in the current potential target group. If so, select one of the following filtering mechanisms to select the final object bounding boxes:
[0018] Based on the user's historical follow list, historical interest trajectory, and personal profile, potential targets for attention are filtered in a personalized way.
[0019] Filtering is based on the distance or occlusion relationship between potential targets and users;
[0020] Filtering based on the size of potential targets;
[0021] Assist in screening potential targets using voice input;
[0022] Increase the fixation time threshold to assist in the screening of potential targets of interest.
[0023] Furthermore, the expression module includes an image pose calculation submodule, a 3D Gaussian model training submodule, a 3D Gaussian expression storage submodule, a 3D Gaussian model preview submodule, and a 3D Gaussian model virtual-real fusion submodule;
[0024] The image pose calculation submodule is used to calculate the global position and pose information of the ROI region;
[0025] The three-dimensional Gaussian model training submodule is used to obtain the three-dimensional Gaussian neural representation of the object corresponding to the current ROI region based on the ROI region and the global position and pose information of the ROI region.
[0026] The three-dimensional Gaussian expression storage submodule is used to store the three-dimensional Gaussian neural expression of the object corresponding to the current ROI region;
[0027] The 3D Gaussian model preview submodule is used to render the 3D Gaussian neural representation of the object corresponding to the current ROI region in an independent window;
[0028] The 3D Gaussian model virtual-real fusion submodule is used to perform virtual-real fusion of the 3D neural representation of the object corresponding to the current ROI region and the real environment or the stored 3D virtual environment collected by the head-mounted display.
[0029] Beneficial effects:
[0030] 1. This invention provides an automatic collection system for 3D representations of gazed objects based on bidirectional monitoring of a head-mounted display. Through user gaze detection using a head-mounted display, it automatically identifies and collects gazed objects in the environment, generates and stores the 3D Gaussian neural representation of the object, and integrates it into other real or virtual scenes. In other words, this invention intelligently judges and identifies objects of interest in the environment based on user behavior, then collects image information for 3D object generation and storage, achieving seamless collection of objects viewed (of interest). Therefore, this invention improves the richness and immersion of the generated virtual-real fusion scenes, enhances user experience and interactivity of virtual reality scenes. In narrative fusion scenes, familiar or interesting objects are naturally integrated, increasing user familiarity, immersion, and engagement, and fostering a greater desire for further interaction.
[0031] 2. This invention provides an automatic collection system for three-dimensional representation of gazed objects based on bidirectional monitoring of a head-mounted display. Utilizing the special design of the head-mounted display device, an inward-facing camera is used to collect the user's state, and an outward-facing camera is used to collect the environment. Then, an algorithm module links the two together to obtain data that is truly valuable to the user.
[0032] 3. This invention provides an automatic collection system for 3D representations of gazed objects based on bidirectional monitoring of a head-mounted display. It employs a 3D Gaussian model training submodule to obtain the 3D Gaussian neural representation of the object corresponding to the current Region of Interest (ROI). The 3D Gaussian model representation has higher quality and can store geometric and textural details of objects with high fidelity. Simultaneously, this invention also employs a 3D Gaussian representation storage submodule to store the 3D Gaussian neural representation of the object corresponding to the current ROI. The storage space for the 3D Gaussian neural representation is smaller than that of the explicit model, while maintaining higher information density. In other words, this invention improves the storage efficiency and quality of virtual objects.
[0033] 4. This invention provides an automatic collection system for three-dimensional representations of gazed objects based on bidirectional monitoring of a head-mounted display. The stored three-dimensional neural representations can be smoothly and naturally rotated and scaled 360°, making them suitable for three-dimensional virtual spaces and improving the interactivity of virtual reality scenes. Attached Figure Description
[0034] Figure 1 A schematic diagram of the existing 3D Gaussian neural representation;
[0035] Figure 2 This invention provides a schematic diagram illustrating the wearing and use of a head-mounted display.
[0036] Figure 3 A diagram illustrating the structure of the three-dimensional representation system for gaze objects provided by this invention;
[0037] Figure 4 Flowchart of the user attention detection module provided by this invention;
[0038] Figure 5 This is a schematic diagram of the virtual-real fusion module for the three-dimensional Gaussian model provided by the present invention;
[0039] Figure 6 A device diagram of the three-dimensional representation system for gaze objects provided by the present invention. Detailed Implementation
[0040] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0041] This invention enables the automatic generation of a 3D model representation of an object being viewed by a user while they are moving around wearing a head-mounted display, and this model can be used for rendering in new real or virtual backgrounds. For example... Figure 2 As illustrated in the example, a user is walking while wearing a head-mounted display device. During this walk, they notice two objects in their path: a flower arrangement and a pine cone. Based on this patented technology, the user's interest in the pine cone is detected, and a 3D model of the pine cone is generated for use in future application rendering.
[0042] System configuration diagram as follows Figure 3 As shown, the generation system includes a user attention detection module, a gaze object image acquisition module, an image ROI generation module, an image semantic segmentation module, and an expression module. The expression module includes an image pose calculation submodule, a 3D Gaussian model training submodule, a 3D Gaussian model storage submodule, a 3D Gaussian model preview submodule, and a 3D Gaussian model virtual-real fusion submodule.
[0043] The user attention detection module is used to detect whether a user wearing a head-mounted display is looking at a specific object in the surrounding environment. It should be noted that the user's eye movement information can be collected by the camera inside the head-mounted display, and the images of objects within the field of view can be detected by the camera outside the head-mounted display. This allows the system to analyze whether the user is interested in a specific object in the surrounding environment. If it is confirmed that the user is looking at an object, the system returns the location information of the object being looked at. At the same time, the condition for determining whether the user is looking at a specific target object is based on whether the gaze duration exceeds a judgment threshold. If the gaze duration on the same target object exceeds 3 seconds, the target object is judged as the target object of attention. The specific time threshold can be adjusted based on the actual scenario.
[0044] The object-gazing image acquisition module is used to continuously acquire images of the object being gazed at when the user gazes at a specific object. It should be noted that, based on the external camera of the head-mounted display, images are continuously acquired when the user gazes at an object. As the user moves, different viewing angles can be changed to acquire images of different sides of the object.
[0045] The image ROI (Region of Interest) generation module is used to select the corresponding object bounding boxes as ROI regions from continuously acquired images based on the direction of eye gaze. In other words, the image ROI generation module is used to predict the bounding boxes of the objects the user is looking at in real-time acquired images after monitoring the user's eye movement information. Specifically, at a certain moment, the object detection algorithm model is used to perform object detection on the acquired images to obtain the bounding boxes and corresponding labels of all objects in the image. Then, the camera in the head-mounted display will monitor the direction of the user's eye gaze, and based on this direction, it will make orientation judgments in the image and select the corresponding object bounding boxes as ROI regions.
[0046] It should be noted that when two or more relatively close target objects appear in the user's gaze direction, the system first determines the number of 3D Gaussian neural representations of the objects that the current expression module needs or can generate. If the number of 3D Gaussian neural representations of the objects that the expression module needs or can generate is greater than one, further filtering is unnecessary, and multiple objects are included together as the target objects of interest. Subsequent data recording, image segmentation, and 3D reconstruction operations are then performed individually on each individual object in the object cluster. Otherwise, all multiple objects are listed as potential targets of interest, awaiting further filtering. The filtering method is as follows:
[0047] The location of the mobile user or potential target group changes the relative positional relationship between the user and the potential target group. Objects that move out of the user's eye gaze direction are removed from the potential target group.
[0048] Determine if there are more than two object bounding boxes in the current potential target group. If so, select one of the following filtering mechanisms to further filter out the final object bounding boxes:
[0049] Based on the user's historical follow list, historical interest trajectory, and personal profile, potential targets for attention are filtered in a personalized way.
[0050] The filtering is based on the distance or occlusion relationship between potential targets and users. Targets that are too far away or have too large an occlusion area are removed. The specific values can be adjusted based on the actual scenario.
[0051] The selection is based on the size of the potential target; targets that are too large or too small are eliminated. The specific values can be adjusted based on the actual scenario.
[0052] Assist in screening potential targets using voice input;
[0053] Increase the threshold for judging gaze time. For example, only when a user gazes at the same target object for more than 5 seconds will the target object be judged as a target object of interest, so as to assist in the screening of potential targets of interest.
[0054] The image semantic segmentation module is used to segment the ROI region from the image based on a segmentation algorithm; such as Figure 4 As shown, firstly, the external camera captures images ( Figure 4 ①) For each frame of image captured by the camera, an object detection algorithm is used to detect all objects in it, and the bounding box of each object is marked. Figure 4 ②). Then, eye movement information is detected to obtain the direction of the user's gaze to achieve focus ( Figure 4③). And map this to the bounding boxes of detected objects in the environment to obtain objects whose viewer's attention lingers for a longer period (e.g., ...). Figure 4 ④); Use the image semantic segmentation module to segment the object from the background. Figure 4 ⑤); Finally, segmented images of the same object at different times are obtained. Figure 4 ⑥).
[0055] The expression module is used to obtain the three-dimensional Gaussian neural expression of the object corresponding to the segmented ROI region based on the trained three-dimensional Gaussian neural expression model, and to perform virtual-real fusion of the generated object three-dimensional neural expression and the real environment collected by the head-mounted display or the stored three-dimensional virtual environment.
[0056] The expression module includes an image pose calculation submodule, a 3D Gaussian model training submodule, a 3D Gaussian expression storage submodule, a 3D Gaussian model preview submodule, and a 3D Gaussian model virtual-real fusion submodule.
[0057] The image pose calculation submodule is used to calculate the global position and pose information of the ROI region. Specifically, the image pose calculation submodule extracts feature points in the ROI region image and calculates the same feature points in different ROI region images. Based on visual algorithms, it calculates the pose information corresponding to each frame of the ROI region image.
[0058] The three-dimensional Gaussian model training submodule is used to obtain the three-dimensional Gaussian neural representation of the object corresponding to the current ROI region based on the ROI region and the global position and pose information of the ROI region.
[0059] The three-dimensional Gaussian expression storage submodule is used to store the three-dimensional Gaussian neural expression of the object corresponding to the current ROI region;
[0060] The 3D Gaussian model preview submodule is used to render the 3D Gaussian neural representation of the object corresponding to the current ROI region in an independent window to achieve a preview effect;
[0061] The 3D Gaussian model virtual-real fusion submodule is used to fuse the 3D neural representation of the object corresponding to the current ROI region with the real environment captured by the head-mounted display or the stored 3D virtual environment, and the fused result can be viewed in 360 degrees or interacted with by the user, such as... Figure 5 As shown.
[0062] It should be noted that the equipment involved in this system, such as... Figure 6 As shown, it mainly includes a gaze detection device, an image acquisition device, a human-computer interaction device, a storage device, a computing device, and a display device.
[0063] In summary, this invention provides an automatic collection system for 3D representations of gazed objects based on bidirectional monitoring of a head-mounted display. By detecting user gaze through a head-mounted display, it automatically identifies and collects gazed objects in the environment, generates and stores the 3D Gaussian neural representation of the object, and integrates it into other real or virtual scenes. In other words, this invention intelligently judges and identifies objects of interest in the environment based on user behavior, and then collects image information for 3D object generation and storage, achieving seamless collection of objects gazed (of interest) by the user.
[0064] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. An automatic collection system for three-dimensional representations of gazed objects based on bidirectional monitoring using a head-mounted display, characterized in that, It includes a user attention detection module, a gaze object image acquisition module, an image ROI generation module, an image semantic segmentation module, and an expression module; The user attention detection module is used to detect whether a user wearing a head-mounted display is looking at a specific object in the surrounding environment; the head-mounted display has an internal camera installed inside to collect the user's eye movement information, and an external camera installed outside to detect images of objects in the real environment. The object-gazing image acquisition module is used to continuously acquire images of the object being gazed at when the user gazes at a specific object; The image ROI generation module is used to select the corresponding object bounding boxes as ROI regions from continuously acquired images according to the direction of eye gaze; wherein, the image ROI generation module uses an object detection algorithm to perform object detection on the acquired images to obtain the bounding boxes and corresponding object labels of all objects in the image; The internal camera monitors the direction in which the user's eyes are looking, and based on this direction, it makes orientation determination in the image and selects the bounding box of the object in that orientation as the ROI region. When there are more than two object bounding boxes in the direction the user's eyes are looking at, and the representation module only produces a 3D Gaussian neural representation of one object, all object bounding boxes are listed as potential target groups to be focused on. The location of the mobile user or potential target group changes the relative positional relationship between the user and the potential target group. Objects that move out of the user's eye gaze direction are removed from the potential target group. Determine if there are more than two object bounding boxes in the current potential target group. If so, select one of the following filtering mechanisms to select the final object bounding boxes: Based on the user's historical follow list, historical interest trajectory, and personal profile, potential targets for attention are filtered in a personalized way. Filtering is based on the distance or occlusion relationship between potential targets and users; Filtering based on the size of potential targets; Assist in screening potential targets using voice input; Increase the fixation time threshold to assist in the screening of potential targets of interest; The image semantic segmentation module is used to segment the ROI region from the image based on a segmentation algorithm; The expression module is used to obtain the three-dimensional Gaussian neural expression of the object corresponding to the segmented ROI region based on the trained three-dimensional Gaussian neural expression model, and to perform virtual-real fusion of the generated object three-dimensional neural expression and the real environment collected by the head-mounted display or the stored three-dimensional virtual environment.
2. The automatic collection system for three-dimensional representation of gazed objects based on bidirectional monitoring of a head-mounted display as described in claim 1, characterized in that, The user attention detection module determines whether a user is looking at a specific object in the surrounding environment by judging whether the user's gaze time on the same target object exceeds a set threshold.
3. The automatic collection system for three-dimensional representation of gazed objects based on bidirectional monitoring of a head-mounted display as described in claim 1, characterized in that, The head-mounted display has an internal camera for collecting information about the user's eye movements and an external camera for detecting images of objects in the real environment. When the user looks at a specific object, the user moves their position so that the external camera of the head-mounted display can capture images of different sides of the object being looked at.
4. The automatic collection system for three-dimensional representation of gazed objects based on bidirectional monitoring of a head-mounted display as described in claim 1, characterized in that, The expression module includes an image pose calculation submodule, a 3D Gaussian model training submodule, a 3D Gaussian expression storage submodule, a 3D Gaussian model preview submodule, and a 3D Gaussian model virtual-real fusion submodule. The image pose calculation submodule is used to calculate the global position and pose information of the ROI region; The three-dimensional Gaussian model training submodule is used to obtain the three-dimensional Gaussian neural representation of the object corresponding to the current ROI region based on the ROI region and the global position and pose information of the ROI region. The three-dimensional Gaussian expression storage submodule is used to store the three-dimensional Gaussian neural expression of the object corresponding to the current ROI region; The 3D Gaussian model preview submodule is used to render the 3D Gaussian neural representation of the object corresponding to the current ROI region in an independent window; The 3D Gaussian model virtual-real fusion submodule is used to perform virtual-real fusion of the 3D neural representation of the object corresponding to the current ROI region and the real environment or the stored 3D virtual environment collected by the head-mounted display.
Citation Information
Patent Citations
Character input device and method based on eye-gaze tracking and speech recognition
CN103076876A
Dynamic mixed reality content in virtual reality
CN117425870A