Multi-camera imaging method and system for head-mounted device
By combining user attention and head movement data, image fusion weight and gaze weight are calculated, dynamic adjustment and real-time feedback of images in multi-camera imaging technology of head-mounted devices are realized, improving the user's immersion and visual experience.
Patent Information
- Application Number
- CN202510606898.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The multi-camera imaging technology of existing head-mounted devices fails to effectively combine user attention and image acquisition, resulting in inconsistent with the user's current gaze area, affecting the immersion and interactive experience, and lacking dynamic camera parameter adjustment and real-time feedback mechanisms.
By obtaining user gaze direction and head movement data, the image fusion weight and gaze weight are calculated, and combined with the pre-trained fusion network and motion prediction model, image weighted fusion and imaging parameter feedback control are carried out to achieve stable generation and real-time adjustment of main viewing angle images.
It enhances the immersive visual experience of the head-mounted device, reduces image delay and flickering, and improves the local image quality balance of the image and the real-time system response.
Smart Images

Figure CN120128692B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-camera imaging technology, and in particular to a multi-camera imaging method and system for a head-mounted device. Background Art
[0002] With the rapid development of wearable technologies such as augmented reality, virtual reality, and mixed reality, head-mounted devices (HMDs) have been widely used in various fields, including consumer electronics, industrial manufacturing, medical training, and distance education. To enhance the immersive visual experience users receive while wearing the devices, more and more HMDs are equipped with multiple cameras to achieve wide-field-of-view, multi-angle, and high-resolution environmental perception and imaging capabilities.
[0003] Existing head-mounted devices often use multi-camera image stitching or fusion to generate the final view, aiming to present a wider visual range or enhance the image's spatial depth information. These methods typically rely on fixed image quality metrics (brightness, contrast, and overlap) for image fusion processing, but fail to fully consider the user's real-time attention distribution and gaze changes. As a result, the generated image may not align with the user's current gaze area, affecting immersion and interactive experience.
[0004] In terms of image quality control, some solutions attempt to enhance overall image clarity through algorithms, but lack local evaluation and feedback mechanisms for different areas of the fused image, and are unable to effectively achieve dynamic adjustment of specific camera imaging parameters, resulting in image quality degradation in some areas that cannot be corrected in a timely manner.
[0005] At the same time, with the deep integration of head-mounted devices with head and eye movement sensors, user behavior data has become highly accessible and valuable. However, existing solutions lack technical solutions that closely integrate user attention behavior with image acquisition and preprocessing, failing to fully tap its potential to improve response speed and visual experience. Summary of the Invention
[0006] The present invention proposes a multi-camera imaging method and system for head-mounted devices, which combines user attention, intelligent image fusion, timing consistency optimization, image quality feedback control and predictive loading mechanism to achieve more natural, stable, clear and low-latency main perspective image generation, providing users with a better visual interaction experience.
[0007] A multi-camera imaging method for a head-mounted device, comprising:
[0008] Acquire images and corresponding image information from multiple cameras of the head-mounted device and integrate them into an image data matrix; obtain user attention information, including gaze direction, eye tracking data, and head movement data;
[0009] Obtaining the gaze weight of the image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtaining the fusion weight based on the image information;
[0010] The images are fused using gaze weights and fusion weights, and a primary-view image is generated according to a preset imaging mode. The fusion weights are only used for weighted calculations in the overlapping area of the images, and the sum of the fusion weights of all images in the overlapping area is one. The image fusion process introduces inter-frame temporal consistency loss.
[0011] The main view image is segmented into regions according to the image region labels, the clarity of the segmented regions is evaluated, and the imaging parameters of the camera are feedback controlled based on the evaluation results;
[0012] Evaluate the stability of the user's head movement data and determine the viewpoint movement method based on the evaluation results to adapt to stable imaging under severe head movement;
[0013] The user's attention information is input into the pre-trained action prediction model to obtain the user's line of sight change trend and preload relevant perspective images to reduce delays.
[0014] As a preferred technical solution of the present invention, obtaining the gaze weight includes:
[0015] The angle between the user's gaze direction vector and the normal vector of each camera image acquisition plane is calculated; a decreasing weight function is constructed based on the angle value, and the initial gaze weight values of all camera images are normalized by the Softmax function to obtain the final gaze weight distribution.
[0016] As a preferred technical solution of the present invention, the acquisition of the fusion weight includes:
[0017] The image information corresponding to each camera image is input into a pre-trained fusion weight estimation network model to output the fusion weight of each image; the image information includes the image resolution, sharpness index, scene brightness value, signal-to-noise ratio and edge clarity score; the fusion weight estimation network model is a lightweight convolutional neural network trained based on historical imaging data.
[0018] As a preferred technical solution of the present invention, generating the main perspective image includes:
[0019] Obtaining a field of view according to a preset imaging mode, wherein the imaging modes include panoramic mode, human eye mode, and macro mode; the fusion weight is only used for weighted calculation of a predetermined overlapping area of each camera image, wherein the overlapping area is calculated based on the camera installation parameters of the head-mounted device;
[0020] Perform weighted fusion of the image data based on the field of view, gaze weight, and fusion weight to obtain a primary perspective image; the field of view is used to determine the size of the fused primary perspective image, the gaze weight is used to perform a single weighting on all image data, and the fusion weight is used to perform a weighted summation on all image data in the overlapping area;
[0021] While generating the main-view image, the image area labeling information is constructed, the image is divided into multiple single-camera imaging areas and multiple multi-camera fusion areas, and the corresponding original image source is marked in each fusion area.
[0022] As a preferred technical solution of the present invention, the inter-frame temporal consistency loss introduced in the image fusion process includes:
[0023] According to the number of overlapping images in the area, the corresponding fusion image detection threshold is obtained, and the fusion weight of each image is tested. If the fusion weight of the image is lower than the detection threshold, the inter-frame temporal consistency loss is used for optimization. The optimization using the inter-frame temporal consistency loss includes fusing adjacent frame images, calculating the pixel difference of the fusion area on the time axis, and constraining the continuity of image content changes through a temporal smoothing loss function; the temporal smoothing loss function is the weighted square difference of the pixel value difference between the current frame and the previous frame of the fusion image in the corresponding fusion area.
[0024] As a preferred technical solution of the present invention, the feedback control of the imaging parameters of the camera based on the evaluation results includes:
[0025] The clarity of each segmented area in the primary view image is evaluated. The clarity evaluation indicators include regional contrast, edge density, frequency domain energy distribution, and image gradient statistics. Based on the evaluation results, they are compared with a preset clarity threshold. When the regional clarity is lower than the threshold, the camera corresponding to the area is identified, and a control instruction is sent back to the camera to adjust its imaging parameters. The imaging parameters include focal length, exposure time, and sensitivity. Feedback control is performed in real time or quasi-real time.
[0026] As a preferred technical solution of the present invention, determining the viewing angle movement mode according to the evaluation result includes:
[0027] Perform dynamic stability analysis on the user's head movement data, calculate the speed, acceleration, and angular velocity of the head movement, and classify the stability assessment results into three states: stable, shaky, and violent;
[0028] When the evaluation result is stable, the perspective image is transitioned by smooth interpolation to maintain a natural and gradual change;
[0029] When the evaluation result is shaking, an inertial buffer mechanism is introduced based on the inter-frame timing information of image fusion to delay the response to view angle switching to suppress shaking;
[0030] When the assessment result is severe, the fast view lock mechanism is enabled to directly switch to the camera view corresponding to the user's current gaze direction to ensure image stability.
[0031] As a preferred technical solution of the present invention, the action prediction model includes:
[0032] Based on a recurrent neural network structure, it uses eye tracking data and head movement data from historical time series as input, combined with the gaze direction change trend within the time window, to predict the gaze point or gaze area within a predetermined time period in the future;
[0033] The viewing angle area corresponding to the prediction result is mapped to the image data matrix, and the corresponding image is acquired, cached or pre-fused in advance to reduce image delay and loading jams when the viewing angle changes suddenly.
[0034] A multi-camera imaging system for a head-mounted device, comprising:
[0035] Image data acquisition module: obtains images and corresponding image information, and integrates them into an image data matrix; obtains user attention information;
[0036] Weight acquisition module: obtains the gaze weight of image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtains the fusion weight based on the image information;
[0037] Image fusion module: fuses images through gaze weight and fusion weight, and generates the main view image according to the preset imaging mode;
[0038] Feedback control module: divides the main view image into regions according to the image area labels, evaluates the clarity of the segmented regions, and performs feedback control on the camera's imaging parameters based on the evaluation results;
[0039] Viewpoint movement determination module: This module evaluates the stability of the user's head movement data and determines the viewpoint movement method based on the evaluation results;
[0040] Preloading module: Input the user's attention information into the pre-trained action prediction model, obtain the user's line of sight change trend, and preload relevant perspective images.
[0041] The present invention has the following advantages:
[0042] The present invention obtains the gaze weight of the image by obtaining the user's gaze direction, realizes the image weighted fusion method with the user's current visual attention area as the core, makes the generated main perspective image more consistent with the user's actual field of view, and enhances the immersive visual experience; introduces the inter-frame temporal consistency loss in the fusion process, and constrains the pixel changes of the fused image between consecutive frames through temporal smoothing, significantly reduces image flickering and jumping phenomena, and improves visual stability in dynamic scenes.
[0043] The present invention divides the main perspective image into different imaging areas, evaluates the clarity of each area, and feeds back the evaluation results to the corresponding camera, adjusting its imaging parameters in real time, thereby effectively improving the overall image's detail expression and local image quality balance. Based on head movement data, the present invention classifies the user's head stability into multiple states, and uses smooth transition, inertial delay, fast locking and other methods to control the perspective switching, providing a natural, smooth and stable image transition effect in different motion states.
[0044] By building a pre-trained motion prediction model and combining it with user attention data, we can predict the future gaze area in advance and preload and cache images, effectively reducing the imaging delay when the perspective changes suddenly, and improving the real-time and continuity of the system response. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only schematic diagrams of the present invention. Those skilled in the art can also derive other drawings based on the provided drawings without inventive effort.
[0046] Figure 1 This is a schematic structural diagram of a multi-camera imaging system for a head-mounted device adopted in an embodiment of the present invention. DETAILED DESCRIPTION
[0047] Embodiment 1, a multi-camera imaging method for a head-mounted device, comprising the following steps:
[0048] Step S1: Acquire images and corresponding image information from multiple cameras of the head-mounted device and integrate them into an image data matrix; obtain user attention information, including gaze direction, eye tracking data, and head movement data;
[0049] Furthermore, the head-mounted device integrates multiple camera modules with different orientations. Each camera performs simultaneous exposure and image acquisition operations through synchronous control logic, generating multiple image signals corresponding to the current frame. After acquisition, each image channel simultaneously records the corresponding image information, including at least the acquisition timestamp, camera number, camera position and attitude parameters (pitch angle, yaw angle), image resolution, exposure information, and ambient brightness value. These images and image information are structured and encapsulated to form a unified image data matrix.
[0050] At the same time, the user's attention information is acquired in real time through the eye tracking module and inertial measurement unit (IMU) integrated into the head-mounted device. This attention information includes: a three-dimensional unit vector of the user's current gaze direction; eye tracking data, including at least the gaze point location, pupil center location, and gaze stability index; and head movement data, including head posture (pitch, yaw, roll), angular velocity, and linear velocity.
[0051] The above multi-source information data are aligned on the time axis and synchronously input into the subsequent weight acquisition module to achieve joint modeling of perspective correlation and image quality.
[0052] Step S2: Obtaining a gaze weight of the image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtaining a fusion weight based on the image information;
[0053] The acquisition of the gaze weight includes:
[0054] Calculate the angle between the user's gaze direction vector and the normal vector of each camera's image acquisition plane. Specifically, the user's gaze direction vector is obtained through the eye tracking system within the headset. This vector represents the position of the user's current gaze point relative to the headset's coordinate system. Simultaneously, the normal vectors of the image acquisition planes of the multiple cameras are extracted. The angle between this normal vector and the user's gaze direction vector is then calculated to obtain a degree value representing the user's gaze angle.
[0055] Based on this angle, a decreasing weighting function is constructed. As the angle between the user's gaze direction and the camera image plane increases, the weight decreases, indicating that the user's attention to the image data decreases. This decreasing function uses an exponential decay function or a Gaussian function to effectively focus on the user's primary viewing area.
[0056] The Softmax function is used to normalize the initial gaze weight values of all camera images. Specifically, the Softmax function normalizes the gaze weight values of all camera images to ensure that the sum of the gaze weight values of each camera image is 1, thus obtaining the final gaze weight distribution.
[0057] The acquisition of the fusion weight includes:
[0058] The image information corresponding to each camera image is input into a pre-trained fusion weight estimation network model. This network model uses a lightweight convolutional neural network (CNN) to output the fusion weight of each camera image by learning image features from historical imaging data.
[0059] The image information includes:
[0060] Image resolution: The resolution of each image directly affects the image quality and subsequent processing effects;
[0061] Sharpness index: Evaluate the clarity of an image by calculating the edge clarity of the image or using methods such as the Laplace operator;
[0062] Scene brightness value: Standardize the brightness of each image to facilitate illumination compensation for subsequent images;
[0063] Signal-to-noise ratio (SNR): Calculates the signal-to-noise ratio of each image to characterize the image quality and noise level;
[0064] Edge clarity score: Analyze the edge clarity of the image through methods such as image gradient information or Canny operator.
[0065] During the training phase, the fusion weight estimation network model learns based on historical image data and adjusts network parameters through error backpropagation to achieve optimal image fusion results. The training data includes images from a variety of different environments and calibrated weight labels to ensure that the network can accurately predict fusion weights for images from various scenarios.
[0066] Step S3: The images are fused using the gaze weight and the fusion weight, and a primary view image is generated according to a preset imaging mode; the fusion weight is only used for weighted calculation of the image overlap area, and the sum of the fusion weights of all images in the image overlap area is one; the image fusion process introduces inter-frame temporal consistency loss;
[0067] Based on the camera position parameters and field of view settings of the head-mounted device, the overlapping area of each camera image is predetermined. The overlapping area is determined by the intersection of the viewing angles of multiple cameras, taking into account parameters such as the camera installation angle, focal length, and imaging range.
[0068] During image fusion, gaze weights and fusion weights are used to weight the image data. Gaze weights primarily weight the overall image, determining the weight of the image in the final composite image based on the user's gaze direction, ensuring that the gazed area receives more attention. Fusion weights are only used to weight overlapping image regions. This means that overlapping image regions between multiple cameras are weighted and fused based on their image quality (such as resolution and clarity), ensuring that optimal image detail is preserved in the fused region.
[0069] Obtain a suitable field of view according to a preset imaging mode, which includes:
[0070] Panoramic mode: Combines multiple images to create a wide field of view, suitable for displaying a complete image of the surrounding environment;
[0071] Human eye mode: performs image processing based on the human visual system, simulating the natural observation method of the human eye;
[0072] Macro mode: Capture close-up objects or scenes with high precision, highlighting details.
[0073] In the process of generating the main perspective image, the gaze weight is used to perform a single weighting on the image to ensure that the image in the gaze direction has a higher proportion; in the overlapping area, the fusion weight is used to perform weighted summation on the multi-camera images to obtain the final main perspective image.
[0074] When generating the primary view image, region labels are constructed on the image to clearly identify the source of each image region. This information divides the image into multiple single-camera imaging regions and multiple multi-camera fusion regions. Each fusion region is labeled with its corresponding original image source, ensuring that users can visually identify the specific camera data corresponding to each image region.
[0075] In the image fusion process, in order to optimize the temporal consistency of image content, inter-frame temporal consistency loss is introduced. Specifically, it includes:
[0076] Based on the number of overlapping images within a region, the fusion image detection threshold for each overlapping region is calculated, and it is determined which regions require temporal optimization for image fusion. If the image fusion weight is lower than the detection threshold, adjacent frame images are used for fusion. By calculating the pixel differences in the fusion region on the time axis, a temporal smoothing loss function is used to constrain the continuity of image content changes, ensuring a natural and smooth transition over time. The temporal smoothing loss function calculates the weighted square difference of the pixel values of the current frame and the previous frame within the corresponding fusion region to constrain the temporal consistency of the image and avoid unnatural jumps or breaks in the picture.
[0077] Step S4: Segmenting the primary view image into regions according to the region labels, performing clarity evaluation on the segmented regions, and performing feedback control on the imaging parameters of the camera based on the evaluation results;
[0078] After the primary view image is generated, it is divided into multiple independent regions based on the region label information. Each region represents a different imaging source, including single-camera imaging regions and multi-camera fusion regions.
[0079] For each segmented region, a clarity evaluation is performed. The clarity evaluation indicators include:
[0080] Regional contrast: This function calculates the contrast value of an image region and analyzes the difference in pixel brightness within the region. Regions with high contrast usually indicate rich details.
[0081] Edge Density: Calculates the number and density of edges within a region using an edge detection algorithm (Canny edge detection). Edges are important indicators of image detail, and a higher density of edges generally indicates a clearer image.
[0082] Frequency Domain Energy Distribution: This function performs frequency domain analysis on an image region to calculate the energy distribution of its frequency components. Regions where frequency domain energy is concentrated in high frequencies generally indicate richer details and clearer images.
[0083] Image gradient statistics: This is done by analyzing image gradients to assess the rate of change in pixel values within a region. Larger gradient values indicate significant structural or detail changes within the image, generally reflecting image clarity.
[0084] The evaluation results are compared with the preset clarity threshold. If the clarity of a segmented area is lower than the threshold, the image quality of that area is considered to be poor and needs to be optimized.
[0085] If the clarity of an area is determined to be lower than a preset threshold, the camera corresponding to the area will be automatically identified and feedback instructions will be sent back to the camera to adjust its imaging parameters. Specific feedback control includes:
[0086] Focus adjustment: For areas with blurred image details, the camera needs to adjust the focus to make the image in that area clearer.
[0087] Exposure time adjustment: If the image in a certain area is too dim, you need to increase the exposure time to improve image brightness and clarity.
[0088] Sensitivity adjustment: In low-light conditions, increasing sensitivity can help improve image quality and reduce noise.
[0089] Step S5: performing a stability assessment on the user's head motion data, and determining a viewing angle movement method based on the assessment results to adapt to stable imaging under conditions of intense head motion;
[0090] Perform dynamic stability analysis on the user's head movement data. This analysis determines the state of the user's head movement based on the head's velocity, acceleration, and angular velocity changes, and categorizes it into the following states:
[0091] Steady state: The head moves slowly with small changes, and the user maintains a relatively stable posture.
[0092] Shaking state: Slight movement or slight tremor of the head causes slight shaking or blurring of the image.
[0093] Severe state: The user's head turns rapidly or performs violent movements, resulting in obvious dynamic instability in the image.
[0094] Using sensors or head tracking data, the system calculates head speed, acceleration, and angular velocity in real time. Based on these values, the system categorizes stability into three states: stable, shaky, and violent. Each state has a different threshold range, and the type of head movement is determined based on the threshold.
[0095] Based on the evaluation results of head movement data, the camera adjusts the viewing angle to adapt to different head movements and ensure image stability. Specific adaptation methods include:
[0096] Steady State: In this state, head movement is relatively stable, with no noticeable change in image quality. Smooth interpolation technology is used to smoothly transition between perspectives, maintaining a natural, gradual change. This smooth interpolation method uses cubic spline interpolation to ensure that the image transitions without jumps or abrupt changes.
[0097] Shake: When the user's head shakes slightly, an inertial buffering mechanism is introduced, combining inter-frame timing information from the image fusion process to delay the response time of perspective switching and mitigate image fluctuations through buffering. This ensures a relatively stable image and reduces the jitter caused by slight head movement. The inertial buffering mechanism performs a weighted average of adjacent frames to ensure smooth image changes, avoiding jumps or discontinuities caused by brief head movements.
[0098] In situations of intense head movement, traditional smooth interpolation and inertial buffering are ineffective in stabilizing the image. In these situations, the fast view lock mechanism is activated to quickly lock the view to the user's current gaze direction. This ensures that the camera adjusts the view in a timely manner, ensuring a stable image that matches the user's actual gaze.
[0099] Step S6: Input the user's attention information into the pre-trained action prediction model, obtain the user's line of sight change trend, and preload relevant perspective images to reduce delays.
[0100] The action prediction model is a lightweight time series prediction model built based on a recurrent neural network (RNN) structure, specifically including a long short-term memory network (LSTM) or a gated recurrent unit (GRU), to balance prediction accuracy and model operation efficiency.
[0101] The motion prediction model uses the user's eye tracking data and head movement data within a time window as input, and uses the following features for modeling: the changing trajectory of the user's gaze vector in the temporal dimension; the head posture angle (pitch, yaw, roll) and its derivatives (speed, acceleration); the spatial projection trajectory of the gaze point on the image plane; and the gaze area stability index between historical frames (gaze point dwell time, jitter frequency).
[0102] During model training, a multi-task loss function is introduced to jointly optimize the gaze point regression error and the viewpoint change classification error, enhancing the model's generalization ability in different scenarios. During the inference phase, the model outputs the probability distribution of the gaze point or gaze area within a certain time period in the future (e.g., 200ms to 500ms). Based on this distribution, the following operations are performed:
[0103] Map the viewpoint index corresponding to the high-probability fixation area to the image data matrix; start the fast image acquisition process and extract the relevant image data of the area from multiple cameras in advance;
[0104] If the image area is in the camera overlap area, perform image pre-fusion processing (optional), including color space alignment, preliminary weighted fusion and caching;
[0105] If the predicted area changes significantly, the image progressive preloading mechanism is enabled, that is, the central area is loaded first and the surrounding areas are loaded in layers;
[0106] The processing results are cached in a high-speed image buffer module, and the current main perspective image is seamlessly accessed when the actual perspective switch is triggered, significantly reducing image loading delays and screen freezes.
[0107] Example 2, a multi-camera imaging system for a head-mounted device, see Figure 1 As shown, it includes the following modules:
[0108] Image data acquisition module: obtains images and corresponding image information, and integrates them into an image data matrix; obtains user attention information;
[0109] Weight acquisition module: obtains the gaze weight of image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtains the fusion weight based on the image information;
[0110] Image fusion module: fuses images through gaze weight and fusion weight, and generates the main view image according to the preset imaging mode;
[0111] Feedback control module: divides the main view image into regions according to the image area labels, evaluates the clarity of the segmented regions, and performs feedback control on the camera's imaging parameters based on the evaluation results;
[0112] Viewpoint movement determination module: This module evaluates the stability of the user's head movement data and determines the viewpoint movement method based on the evaluation results;
[0113] Preloading module: Input the user's attention information into the pre-trained action prediction model, obtain the user's line of sight change trend, and preload relevant perspective images.
[0114] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-camera imaging method for a head-mounted device, characterized in that: include: Acquire images and corresponding image information from multiple cameras of the head-mounted device and integrate them into an image data matrix; Obtain user attention information, including gaze direction, eye tracking data, and head movement data; Obtaining the gaze weight of the image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtaining the fusion weight based on the image information; The images are fused by gaze weight and fusion weight, and the primary view image is generated according to the preset imaging mode; The fusion weight is only used for weighted calculation of the image overlap area, and the sum of the fusion weights of all images in the image overlap area is one; the image fusion process introduces inter-frame temporal consistency loss; The main view image is segmented into regions according to the image region labels, the clarity of the segmented regions is evaluated, and the imaging parameters of the camera are feedback controlled based on the evaluation results; Evaluate the stability of the user's head movement data and determine the viewpoint movement method based on the evaluation results to adapt to stable imaging under severe head movement; Input the user's attention information into the pre-trained action prediction model to obtain the user's gaze change trend and preload relevant perspective images to reduce latency; In the above steps, generating the main perspective image includes: Obtaining a field of view according to a preset imaging mode, wherein the imaging mode includes a panoramic mode, a human eye mode, and a macro mode; the fusion weight is only used for weighted calculation of a predetermined overlapping area of each camera image; the overlapping area is obtained by calculating the camera installation parameters of the head-mounted device; Perform weighted fusion of the image data based on the field of view, gaze weight, and fusion weight to obtain a primary perspective image; the field of view is used to determine the size of the fused primary perspective image, the gaze weight is used to perform a single weighting on all image data, and the fusion weight is used to perform a weighted summation on all image data in the overlapping area; While generating the main-view image, the image area labeling information is constructed, the image is divided into multiple single-camera imaging areas and multiple multi-camera fusion areas, and the corresponding original image source is marked in each fusion area.
2. The multi-camera imaging method for a head-mounted device according to claim 1, characterized in that: The acquisition of the gaze weight includes: The angle between the user's gaze direction vector and the normal vector of each camera image acquisition plane is calculated; a decreasing weight function is constructed based on the angle value, and the initial gaze weight values of all camera images are normalized by the Softmax function to obtain the final gaze weight distribution.
3. The multi-camera imaging method for a head-mounted device according to claim 1, characterized in that: The acquisition of the fusion weight includes: The image information corresponding to each camera image is input into a pre-trained fusion weight estimation network model to output the fusion weight of each image; the image information includes the image resolution, sharpness index, scene brightness value, signal-to-noise ratio and edge clarity score; the fusion weight estimation network model is a lightweight convolutional neural network trained based on historical imaging data.
4. The multi-camera imaging method for a head-mounted device according to claim 1, characterized in that: The inter-frame temporal consistency loss introduced in the image fusion process includes: According to the number of overlapping images in the area, the corresponding fusion image detection threshold is obtained, and the fusion weight of each image is tested. If the fusion weight of the image is lower than the detection threshold, the inter-frame temporal consistency loss is used for optimization. The optimization using the inter-frame temporal consistency loss includes fusing adjacent frame images, calculating the pixel difference of the fusion area on the time axis, and constraining the continuity of image content changes through a temporal smoothing loss function; the temporal smoothing loss function is the weighted square difference of the pixel value difference between the current frame and the previous frame of the fusion image in the corresponding fusion area.
5. The multi-camera imaging method for a head-mounted device according to claim 1, characterized in that: The feedback control of the imaging parameters of the camera based on the evaluation result includes: The clarity of each segmented area in the primary view image is evaluated. The clarity evaluation indicators include regional contrast, edge density, frequency domain energy distribution, and image gradient statistics. Based on the evaluation results, they are compared with a preset clarity threshold. When the regional clarity is lower than the threshold, the camera corresponding to the area is identified, and a control instruction is sent back to the camera to adjust its imaging parameters. The imaging parameters include focal length, exposure time, and sensitivity. Feedback control is performed in real time or quasi-real time.
6. The multi-camera imaging method for a head-mounted device according to claim 1, characterized in that: Determining the viewing angle movement mode according to the evaluation result includes: Perform dynamic stability analysis on the user's head movement data, calculate the speed, acceleration, and angular velocity of the head movement, and classify the stability assessment results into three states: stable, shaky, and violent; When the evaluation result is stable, the perspective image is transitioned by smooth interpolation to maintain a natural and gradual change; When the evaluation result is shaking, an inertial buffer mechanism is introduced based on the inter-frame timing information of image fusion to delay the response to view angle switching to suppress shaking; When the assessment result is severe, the fast view lock mechanism is enabled to directly switch to the camera view corresponding to the user's current gaze direction to ensure image stability.
7. The multi-camera imaging method for a head-mounted device according to claim 1, characterized in that: The action prediction model includes: Based on a recurrent neural network structure, it uses eye tracking data and head movement data from historical time series as input, combined with the gaze direction change trend within the time window, to predict the gaze point or gaze area within a predetermined time period in the future; The viewing angle area corresponding to the prediction result is mapped to the image data matrix, and the corresponding image is acquired, cached or pre-fused in advance to reduce image delay and loading jams when the viewing angle changes suddenly.
8. A multi-camera imaging system for a head-mounted device, characterized in that: The system applies the multi-camera imaging method for a head-mounted device according to any one of claims 1 to 7, comprising: Image data acquisition module: obtains images and corresponding image information, and integrates them into an image data matrix; obtains user attention information; Weight acquisition module: obtains the gaze weight of image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtains the fusion weight based on the image information; Image fusion module: fuses images through gaze weight and fusion weight, and generates the main view image according to the preset imaging mode; Feedback control module: divides the main view image into regions according to the image area labels, evaluates the clarity of the segmented regions, and performs feedback control on the camera's imaging parameters based on the evaluation results; Viewpoint movement determination module: This module evaluates the stability of the user's head movement data and determines the viewpoint movement method based on the evaluation results; Preloading module: Input the user's attention information into the pre-trained action prediction model, obtain the user's line of sight change trend, and preload relevant perspective images.
Citation Information
Patent Citations
Virtual reality interaction method for constructing multi-modal fusion based on application scene features
CN119473014A
Intelligent control method and system for AR equipment
CN119847330A