Multi-camera imaging method and system for head-mounted device
By acquiring user attention data in a head-mounted device, performing image weighted fusion, and combining timing consistency loss and action prediction model, the problem of inconsistent images and user gaze areas in the prior art is solved, achieving a higher quality and stable visual experience.
Patent Information
- Application Number
- CN202510606898.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing multi-camera imaging technology for head-mounted devices fails to fully consider the user's real-time attention distribution and line of sight changes, resulting in the generated image being inconsistent with the user's current gaze area, affecting the immersion and interactive experience.
By obtaining the user's gaze direction, eye tracking data and head movement data, calculating the gaze weight and fusion weight of the image data, performing image weighting and fusion, and introducing inter-frame timing consistency loss to optimize the image generation process. At the same time, the camera imaging parameters are adjusted based on the sharpness evaluation of the image area, and the viewing angle image is preloaded through the action prediction model.
It realizes a more natural, stable, clear and low-latency main viewing image generation, enhances the user's visual interaction experience, and improves the visual stability of the image in dynamic scenes.
Smart Images

Figure CN120128692A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-camera imaging technology, and particularly to a multi-camera imaging method and system for a head-mounted device. Background Art
[0002] With the rapid development of wearable technologies such as augmented reality, virtual reality, and mixed reality, head-mounted devices have been widely used in multiple fields such as consumer electronics, industrial manufacturing, medical training, and distance education. To enhance the immersive visual experience obtained by users during device wearing, more and more head-mounted devices are equipped with multiple cameras to achieve environmental perception and imaging capabilities with a large field of view, multiple angles, and high resolution.
[0003] In the prior art, multi-camera image stitching or fusion methods are mostly used in head-mounted devices to generate a final view to present a wider visible range or enhance the spatial depth information of the image. Such methods usually rely on fixed image quality metrics (brightness, contrast, overlap) for image fusion processing, but do not fully consider the real-time attention distribution and line-of-sight change trend of users, resulting in the generated image may not be consistent with the current gaze area of the user, thus affecting the immersion and interaction experience.
[0004] In terms of image quality control, some solutions attempt to enhance the overall image clarity through algorithms, but lack a local evaluation and feedback mechanism for different regions of the fused image, and cannot effectively achieve dynamic adjustment of specific camera imaging parameters, resulting in the decline of the image quality in some regions that cannot be corrected in time.
[0005] At the same time, with the deep integration of head-mounted devices with user head movement and eye movement sensors, user behavior data has high availability and utilization value. However, in the existing solutions, there is still a lack of technical solutions that closely combine user attention behavior with image acquisition and image preprocessing processes, and the potential in improving response speed and visual experience has not been fully explored. Summary of the Invention
[0006] The present invention proposes a multi-camera imaging method and system for a head-mounted device, which combines user attention, image intelligent fusion, temporal consistency optimization, image quality feedback control, and prediction loading mechanism to achieve the generation of a more natural, stable, clear, and low-latency main perspective image, and provide a better visual interaction experience for users.
[0007] A multi-camera imaging method for a head-mounted device includes: Obtaining images from multiple cameras of the head-mounted device and corresponding image information, and integrating them into an image data matrix; obtaining user attention information, including gaze direction, eye movement tracking data, and head movement data; Obtain the fixation weight of the image data based on the angle between the user's fixation direction and the image acquisition plane of each camera, and obtain the fusion weight based on the image information; Fuse the images through the fixation weight and the fusion weight, and generate a main perspective image according to the preset imaging mode; the fusion weight is only used for the weighted calculation of the image overlapping area, and the sum of the fusion weights of all images within the image overlapping area is one; the frame - to - frame temporal consistency loss is introduced during the image fusion process; Segment the main perspective image according to the figure area label, evaluate the clarity of the segmented areas, and perform feedback control on the imaging parameters of the camera based on the evaluation results; Evaluate the stability of the user's head movement data, and determine the viewing angle movement mode according to the evaluation results to adapt to stable imaging in the case of violent head movement; Input the user's attention information into the pre - trained action prediction model, obtain the trend of the user's gaze change, and pre - load the relevant perspective images to reduce latency.
[0008] As a preferred technical solution of the present invention, the obtaining of the fixation weight includes: Calculate the included angle value between the user's fixation direction vector and the normal vector of the image acquisition plane of each camera; construct a decreasing weight function based on this included angle value, and normalize the initial fixation weight values of all camera images through the Softmax function to obtain the final fixation weight distribution.
[0009] As a preferred technical solution of the present invention, the obtaining of the fusion weight includes: Input the image information corresponding to each camera image into the pre - trained fusion weight estimation network model to output the fusion weight of each image; the image information includes the resolution, sharpness index, scene brightness value, signal - to - noise ratio, and edge clarity score of the image; the fusion weight estimation network model is a lightweight convolutional neural network and is trained based on historical imaging data.
[0010] As a preferred technical solution of the present invention, the generating of the main perspective image includes: Obtain the field of view range according to the preset imaging mode, and the imaging mode includes a panoramic mode, a human - eye mode, and a macro mode; the image fusion is only performed within the pre - determined overlapping area of each camera image, and the overlapping area is calculated from the camera installation parameters of the head - mounted device; Perform weighted fusion on the image data according to the field of view range, fixation weight, and fusion weight to obtain the main perspective image; the field of view range is used to determine the size of the fused main perspective image, the fixation weight is used for single - time weighting of all image data, and the fusion weight is used for weighted summation of all image data within the overlapping area; While generating the main perspective image, construct the figure area label information, divide the image into multiple single-camera imaging areas and multiple multi-camera fusion areas, and label the corresponding original image sources in each fusion area.
[0011] As a preferred technical solution of the present invention, the inter-frame temporal consistency loss introduced in the image fusion process includes: According to the number of overlapping images in the area, obtain the corresponding fusion image detection threshold, and detect the fusion weights of each image. If it is lower than the detection threshold, use the inter-frame temporal consistency loss for optimization, including fusing adjacent frame images, calculating the pixel difference of the fusion area on the time axis, and constraining the continuity of image content changes through the temporal smoothing loss function; the temporal smoothing loss function is the weighted square difference of the pixel values of the current frame and the previous frame of the fusion image in the corresponding fusion area.
[0012] As a preferred technical solution of the present invention, the feedback control of the imaging parameters of the camera based on the evaluation result includes: Evaluate the clarity of each figure segmentation area in the main perspective image. The evaluation indicators include regional contrast, edge density, frequency domain energy distribution, and image gradient statistical values; based on the comparison between the evaluation result and the preset clarity threshold, when the regional clarity is lower than the threshold, identify the camera corresponding to the area, and send a control command back to the camera to adjust its imaging parameters; the imaging parameters include focal length, exposure time, and sensitivity, and the feedback control is performed in real-time or quasi-real-time.
[0013] As a preferred technical solution of the present invention, the determination of the viewing angle movement mode according to the evaluation result includes: Conduct dynamic stability analysis on the user's head movement data, calculate the change amplitudes of the speed, acceleration, and angular velocity of the head movement, and divide the stability evaluation result into three states: stable, shaking, and violent; When the evaluation result is stable, use the method of smooth interpolation to transition the perspective image and maintain natural gradual change; When the evaluation result is shaking, introduce an inertial buffer mechanism in combination with the inter-frame temporal information of image fusion to delay the response to viewing angle switching to suppress jitter; When the evaluation result is violent, enable the fast viewing angle locking mechanism to directly switch to the camera viewing angle corresponding to the user's current gaze direction to ensure image stability.
[0014] As a preferred technical solution of the present invention, the action prediction model includes: Based on the structure of a recurrent neural network, use the eye movement tracking data and head movement data in the historical time series as inputs, and combine the change trend of the gaze direction within the time window to predict the gaze landing point or the fixation area within a predetermined future time period; Map the perspective area corresponding to the prediction result to the image data matrix, and perform acquisition, caching, or pre-fusion processing of the corresponding image in advance to reduce image latency and loading jitter during perspective mutation.
[0015] A multi-camera imaging system for a head-mounted device, comprising: Image data acquisition module: acquire an image and corresponding image information, and integrate them into an image data matrix; acquire the attention information of the user; Weight acquisition module: acquire the fixation weight of the image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and acquire the fusion weight based on the image information; Image fusion module: fuse the images through the fixation weight and the fusion weight, and generate a main perspective image according to the preset imaging mode; Feedback control module: divide the main perspective image into regions according to the map area label, evaluate the sharpness of the divided regions, and perform feedback control on the imaging parameters of the camera based on the evaluation result; Perspective movement determination module: evaluate the stability of the user's head movement data, and determine the perspective movement mode according to the evaluation result; Preloading module: input the attention information of the user into a pre-trained action prediction model, acquire the changing trend of the user's line of sight, and perform preloading of relevant perspective images.
[0016] The present invention has the following advantages: The present invention acquires the fixation weight of an image by obtaining the user's gaze direction, realizes an image weighted fusion method centered on the user's current visual attention area, makes the generated main perspective image more conform to the user's actual field of view, and enhances the immersive visual experience; introduces an inter-frame temporal consistency loss during the fusion process, and constrains the pixel change between consecutive frames of the fused image through temporal smoothing, significantly reducing the image flickering and jumping phenomena, and improving the visual stability in dynamic scenes.
[0017] The present invention divides the main perspective image into different imaging regions, evaluates the sharpness region by region, and feeds back the evaluation result to the corresponding camera to adjust its imaging parameters in real time, effectively improving the detail expressiveness and local image quality balance of the overall image; classifies the user's head stability into multiple states based on the head movement data, and respectively controls the perspective switching by means of smooth transition, inertial delay, fast locking, etc., and can provide a natural, smooth and stable image transition effect in different motion states.
[0018] By constructing a pre-trained action prediction model, combining the user's attention data, predicting the future fixation area in advance and performing preloading and caching processing of the image, the imaging latency during perspective mutation is effectively reduced, and the real-time performance and continuity of the system response are improved. Brief Description of the Drawings
[0019] Figure 1 FIG. 1 is a schematic structural diagram of a multi-camera imaging system for a head-mounted device according to an embodiment of the present invention. Detailed Description of the Embodiments
[0020] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0021] Embodiment 1. A multi-camera imaging method for a head-mounted device includes the following steps: Step S1: Obtain images from multiple cameras of the head-mounted device and corresponding image information, and integrate them into an image data matrix; obtain the attention information of the user, including the gaze direction, eye movement tracking data, and head movement data; Further, a plurality of camera modules with different directions are integrated on the head-mounted device. Each camera performs simultaneous exposure and image acquisition operations through a synchronous control logic to generate a multiplexed image signal corresponding to the current frame. After each image is acquired, the corresponding image information is synchronously recorded, including at least the acquisition timestamp, camera number, camera position and attitude parameters (pitch angle, yaw angle), image resolution, exposure information, and ambient brightness value. The images and image information are structurally encapsulated and uniformly form an image data matrix.
[0022] Meanwhile, the attention information of the user is obtained in real time through an eye movement tracking module and an inertial measurement unit (IMU) integrated in the head-mounted device. The attention information includes: a three-dimensional unit vector of the current gaze direction of the user; eye movement tracking data, including at least the gaze point position, pupil center position, and gaze stability index; head movement data, including head attitude (pitch angle, yaw angle, roll angle), angular velocity, and linear velocity.
[0023] The above multi-source information data is aligned on the time axis and synchronously input into the subsequent weight acquisition module to realize the joint modeling of the perspective correlation degree and the image quality.
[0024] Step S2: Obtain the gaze weight of the image data based on the angle between the gaze direction of the user and the image acquisition plane of each camera, and obtain the fusion weight based on the image information; The acquisition of the gaze weight includes: Calculate the angular value between the user's gaze direction vector and the normal vectors of the image acquisition planes of each camera. Specifically, obtain the user's gaze direction vector through the eye tracking system inside the head-mounted device, and this vector represents the position of the user's current fixation point relative to the coordinate system of the head-mounted device; at the same time, extract the normal vectors of the image acquisition planes from multiple cameras. Then, calculate the angle between this normal vector and the user's gaze direction vector to obtain a degree value representing the user's gaze angle.
[0025] Construct a decreasing weight function based on this angular value, that is, as the angle between the user's gaze direction and the camera image plane increases, the weight value decreases, indicating that the user's attention to this image data decreases. This decreasing function uses an exponential decay function or a Gaussian function to effectively focus on the user's main viewing area.
[0026] Use the Softmax function to normalize the initial fixation weight values of all camera images. Specifically, the Softmax function normalizes the fixation weight values of all camera images to ensure that the sum of the fixation weight values of each camera image is 1, thereby obtaining the final fixation weight distribution.
[0027] The acquisition of the fusion weights includes: Input the image information corresponding to each camera image into a pre-trained fusion weight estimation network model. This network model uses a lightweight convolutional neural network (CNN) to output the fusion weights of each camera image by learning the image features in the historical imaging data.
[0028] The image information includes: Image resolution: The resolution of each image directly affects the image quality and subsequent processing effects; Sharpness index: Evaluate the sharpness of the image by calculating methods such as the edge sharpness of the image or the Laplacian operator; Scene brightness value: Standardize the brightness of each image for subsequent image light compensation; Signal-to-noise ratio (SNR): Calculate the signal-to-noise ratio of each image to characterize the quality and noise level of the image; Edge sharpness score: Analyze the edge sharpness of the image through methods such as the gradient information of the image or the Canny operator.
[0029] The fusion weight estimation network model learns based on historical image data during the training phase, and adjusts the network parameters through error backpropagation to achieve the optimal image fusion effect. The training data includes images in various different environments and calibrated weight labels to ensure that the network can accurately predict the fusion weights for images in various scenarios.
[0030] Step S3: The images are fused using the fixation weight and the fusion weight, and a main perspective image is generated according to a preset imaging mode; the fusion weight is only used for weighted calculation of the image overlapping area, and the sum of the fusion weights of all images within the image overlapping area is one; a frame - to - frame temporal consistency loss is introduced during the image fusion process. According to the camera position parameters and the set field of view range of the head - mounted device, the overlapping area of each camera image is determined in advance. The overlapping area is determined by the perspective intersection area between multiple cameras, taking into account parameters such as the installation angle, focal length, and imaging range of the cameras.
[0031] During image fusion, the fixation weight and the fusion weight are used to perform weighted processing on the image data. The fixation weight mainly acts on the overall image weighting, determining the weight distribution of the image in the final composite image based on the user's fixation direction, ensuring that the fixation area receives more attention. The fusion weight is only used for weighted calculation of the image overlapping area, that is, the overlapping image areas between multiple cameras will be weighted and fused according to their image quality (such as resolution, clarity, etc.), ensuring that the optimal image details in the fusion area are retained.
[0032] According to the preset imaging mode, an appropriate field of view range is obtained. The imaging modes include: Panoramic mode: Multiple images are synthesized to generate a wide field of view, suitable for displaying a complete image of the surrounding environment; Human - eye mode: Image processing is performed according to the human visual system, simulating the natural viewing method of the human eye; Macro mode: High - precision imaging is performed for close - range objects or scenes, highlighting details.
[0033] During the generation of the main perspective image, the fixation weight is used to perform single - weighted processing on the image, ensuring that the image in the fixation direction has a higher proportion; while within the overlapping area, the fusion weight is used to perform weighted summation on the multi - camera images to obtain the final main perspective image.
[0034] When generating the main perspective image, in order to clarify the source of each image area, graphic area label information is constructed on the image. This information divides the image into multiple single - camera imaging areas and multiple multi - camera fusion areas. Each fusion area will be marked with its corresponding original image source, ensuring that users can visually identify the specific camera data corresponding to each image area.
[0035] During the image fusion process, in order to optimize the temporal consistency of the image content, a frame - to - frame temporal consistency loss is introduced. Specifically, it includes: Based on the number of overlapping images within the region, calculate the fusion image detection threshold for each overlapping region, and determine which regions require temporal optimization for image fusion. If the fusion weight of the image is lower than the detection threshold, use adjacent frame images for fusion. By calculating the pixel differences of the fusion region on the time axis, use the temporal smoothing loss function to constrain the continuity of the change in image content, ensuring a natural and smooth transition of the image over time. The temporal smoothing loss function calculates the weighted squared difference of the pixel values between the current frame and the previous frame within the corresponding fusion region to constrain the temporal consistency of the image and avoid unnatural jumps or breaks in the picture.
[0036] Step S4: Segment the main perspective image according to the figure area label, evaluate the clarity of the segmented regions, and perform feedback control on the imaging parameters of the camera based on the evaluation results; After the main perspective image is generated, divide the image into multiple independent regions according to the figure area label information. Each region represents a different imaging source, including single camera imaging regions and multiple camera fusion regions.
[0037] For each segmented region, perform clarity evaluation. The evaluation metrics include: Region contrast: By calculating the contrast value of the image region, analyze the difference in pixel brightness within the region. Regions with high contrast usually indicate rich details.
[0038] Edge density: Through an edge detection algorithm (Canny edge detection), calculate the number of edges and their density within the region. Edges are an important sign of image details, and regions with high-density edges usually correspond to clearer images.
[0039] Frequency domain energy distribution: Perform frequency domain analysis on the image region and calculate the energy distribution of its frequency components. Regions where the frequency domain energy is concentrated in the high-frequency part usually indicate rich details and clearer images.
[0040] Image gradient statistic value: Through image gradient analysis, evaluate the rate of change of pixel values within the region. Larger gradient values indicate obvious structural or detail changes in the image, usually reflecting the clarity of the image.
[0041] Compare the evaluation results with a preset clarity threshold. If the clarity of a segmented region is lower than the threshold, it is considered that the image quality of this region is poor and needs to be optimized.
[0042] When it is determined that the clarity of the region is lower than the preset threshold, automatically identify the camera corresponding to this region and send a feedback command to this camera to adjust its imaging parameters. Specific feedback controls include: Focal length adjustment: For regions with relatively blurred image details, the camera needs to adjust the focal length to make the imaging of this region clearer.
[0043] Exposure time adjustment: If the image in the area is too dim, the exposure time needs to be increased to improve the image brightness and clarity.
[0044] ISO adjustment: In low light conditions, increasing the ISO helps improve image quality and reduce noise.
[0045] Step S5: Evaluate the stability of the user's head movement data, and determine the viewing angle movement method according to the evaluation result to adapt to stable imaging in the case of violent head movement; For the user's head movement data, perform dynamic stability analysis. The analysis is based on the change amplitudes of the head's speed, acceleration, and angular velocity to judge the state of the user's head movement, and classify it into the following states: Steady state: The head movement speed is slow, the change amplitude is small, and the user maintains a relatively stable posture.
[0046] Shaking state: The head has slight movement or tiny tremors, resulting in slight jitter or blurring of the image.
[0047] Violent state: The user's head rotates rapidly or makes violent movements, resulting in obvious dynamic instability of the image.
[0048] Through sensors or head movement tracking data, calculate the head's movement speed, acceleration, and angular velocity change amplitude in real time. Based on the calculated speed, acceleration, and angular velocity, divide the stability evaluation result into three states: steady, shaking, and violent. Each state has different threshold intervals, and the type of head movement is determined according to the threshold.
[0049] According to the evaluation result of the head movement data, adjust the viewing angle movement method to adapt to different head movement states and ensure image stability. The specific adaptation methods include: Steady state: In this state, the head movement is relatively steady and will not cause significant changes in image quality. Smoothly transition the viewing angle image through smooth interpolation technology to maintain a natural slow change effect. The smooth interpolation method is implemented through cubic spline interpolation to ensure that the image does not jump or change suddenly during the transition.
[0050] Shaking state: When the user's head has slight shaking, combine the inter-frame timing information in the image fusion process and introduce an inertial buffer mechanism to delay the response time of the viewing angle switch and slow down the image change through buffer processing. At this time, the image will remain relatively stable and reduce the jitter effect caused by the slight head shaking. The inertial buffer mechanism performs weighted averaging on adjacent frame images to ensure smooth change of the image content and avoid image jumping or incoherence caused by short-term head movement.
[0051] Intense state: In the case of intense head movement, traditional smooth interpolation and inertial buffering cannot effectively stabilize the image. At this time, the fast view locking mechanism is enabled to quickly lock the view in the current gaze direction of the user. Ensure that the camera adjusts the view in a timely manner so that the image is stable and meets the actual line-of-sight requirements of the user.
[0052] Step S6: Input the user's attention information into the pre-trained action prediction model to obtain the trend of the user's gaze change, and pre-load relevant perspective images to reduce latency.
[0053] The action prediction model is a lightweight time series prediction model constructed based on the structure of a recurrent neural network (RNN), specifically including a long short-term memory network (LSTM) or a gated recurrent unit (GRU), in order to balance prediction accuracy and model operation efficiency.
[0054] The action prediction model takes the eye movement tracking data and head movement data of the user within a time window as input, and uses the following features for modeling: the change trajectory of the user's gaze vector in the time series dimension; the head pose angles (Pitch, Yaw, Roll) and their derivatives (velocity, acceleration); the spatial projection trajectory of the fixation point on the image plane; the stability index of the fixation area between historical frames (fixation point residence time, jitter frequency).
[0055] During the model training process, a multi-task loss function is introduced to jointly optimize the gaze landing point regression error and the view change classification error, and enhance the generalization ability of the model in different scenarios. In the inference stage, the model outputs the probability distribution of the gaze landing point or the fixation area within a certain future time (such as 200ms - 500ms), and performs the following operations according to the distribution result: Map the view index corresponding to the high-probability fixation area to the image data matrix; start the fast image acquisition process, and pre-extract the relevant image data of this area from multiple cameras; If the image area is in the camera overlapping area, perform image pre-fusion processing (optional), including color space alignment, preliminary weighted fusion and caching; If the predicted area changes greatly, enable the image progressive pre-loading mechanism, that is, preferentially load the central area and load the surrounding area layer by layer; Cache the processing result in the high-speed image buffer module, and seamlessly access the current main view image when the actual view switch is triggered, significantly reducing the image loading latency and frame stuttering.
[0056] Embodiment 2, A multi-camera imaging system for a head-mounted device, see Figure 1 as shown, including the following modules: Image data acquisition module: Obtain images and corresponding image information, and integrate them into an image data matrix; obtain the user's attention information; Weight acquisition module: Obtain the fixation weight of image data based on the angle between the user's fixation direction and the image acquisition plane of each camera, and obtain the fusion weight based on the image information; Image fusion module: Fusion the images through the fixation weight and the fusion weight, and generate the main perspective image according to the preset imaging mode; Feedback control module: Segment the main perspective image according to the figure area label, evaluate the clarity of the segmented area, and perform feedback control on the imaging parameters of the camera based on the evaluation result; Viewpoint movement determination module: Evaluate the stability of the user's head movement data, and determine the viewpoint movement method according to the evaluation result; Preloading module: Input the user's attention information into the pre-trained action prediction model, obtain the trend of the user's gaze change, and preload the relevant perspective images.
[0057] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-camera imaging method for a head-mounted device, characterized in that: include: Acquire images and corresponding image information from multiple cameras of the head-mounted device and integrate them into an image data matrix; Obtain user attention information, including gaze direction, eye tracking data, and head movement data; Obtaining the gaze weight of the image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtaining the fusion weight based on the image information; The images are fused by using the gaze weight and the fusion weight, and a main-view image is generated according to a preset imaging mode; The fusion weight is only used for weighted calculation of the image overlap area, and the sum of the fusion weights of all images in the image overlap area is one; the image fusion process introduces inter-frame temporal consistency loss; The main view image is segmented into regions according to the image region labels, the clarity of the segmented regions is evaluated, and the imaging parameters of the camera are feedback-controlled based on the evaluation results; Evaluate the stability of the user's head movement data and determine the viewpoint movement method based on the evaluation results to adapt to stable imaging under severe head movement; The user's attention information is input into the pre-trained action prediction model to obtain the user's line of sight change trend and preload relevant perspective images to reduce delays.
2. A multi-camera imaging method for a head mounted device according to claim 1, characterized in that: The acquisition of the gaze weight includes: The angle between the user's gaze direction vector and the normal vector of each camera image acquisition plane is calculated; a decreasing weight function is constructed based on the angle value, and the initial gaze weight values of all camera images are normalized by the Softmax function to obtain the final gaze weight distribution.
3. The multi-camera imaging method for a head mounted device according to claim 1, characterized in that: The acquisition of the fusion weight includes: The image information corresponding to each camera image is input into a pre-trained fusion weight estimation network model to output the fusion weight of each image; the image information includes the image resolution, sharpness index, scene brightness value, signal-to-noise ratio and edge clarity score; the fusion weight estimation network model is a lightweight convolutional neural network trained based on historical imaging data.
4. The multi-camera imaging method for a head mounted device according to claim 1, characterized in that: Generating the main viewing angle image comprises: The field of view range is obtained according to a preset imaging mode, wherein the imaging mode includes a panoramic mode, a human eye mode, and a macro mode; the image fusion is performed only within a predetermined overlapping area of each camera image, wherein the overlapping area is obtained by calculating the camera installation parameters of the head mounted device; The image data is weightedly fused according to the field of view, the gaze weight and the fusion weight to obtain a main-view image; the field of view is used to determine the size of the fused main-view image, the gaze weight is used to perform a single weighting on all image data, and the fusion weight is used to perform a weighted summation on all image data in the overlapping area; While generating the main-view image, the image area labeling information is constructed, the image is divided into multiple single-camera imaging areas and multiple multi-camera fusion areas, and the corresponding original image source is marked in each fusion area.
5. The multi-camera imaging method for a head mounted device according to claim 1, characterized in that: The inter-frame temporal consistency loss introduced in the image fusion process includes: According to the number of overlapping images in the area, the corresponding fusion image detection threshold is obtained, and the fusion weight of each image is detected. If it is lower than the detection threshold, the inter-frame temporal consistency loss is used for optimization, including fusing adjacent frame images, calculating the pixel difference in the fusion area on the time axis, and constraining the continuity of image content changes through a temporal smoothing loss function; the temporal smoothing loss function is the weighted square difference of the pixel value difference between the current frame and the previous frame of the fusion image in the corresponding fusion area.
6. The multi-camera imaging method for a head mounted device according to claim 1, characterized in that: The feedback control of the imaging parameters of the camera based on the evaluation result includes: The clarity of each segmented area in the main view image is evaluated, and the evaluation indicators include regional contrast, edge density, frequency domain energy distribution and image gradient statistics; based on the evaluation result, it is compared with a preset clarity threshold. When the regional clarity is lower than the threshold, the camera corresponding to the area is identified, and a control instruction is sent back to the camera to adjust its imaging parameters; the imaging parameters include focal length, exposure time and sensitivity, and feedback control is performed in real time or quasi real time.
7. The multi-camera imaging method for a head mounted device according to claim 1, characterized in that: Determining the viewing angle movement mode according to the evaluation result includes: Perform dynamic stability analysis on the user's head movement data, calculate the speed, acceleration and angular velocity of the head movement, and divide the stability assessment results into three states: stable, shaky and violent; When the evaluation result is stable, the perspective image is transitioned by smooth interpolation to maintain a natural and slow change; When the evaluation result is shaking, an inertial buffer mechanism is introduced in combination with the inter-frame timing information of image fusion to delay the response to the view switching to suppress the shaking; When the evaluation result is severe, the fast view lock mechanism is enabled to directly switch to the camera view corresponding to the user's current gaze direction to ensure image stability.
8. The multi-camera imaging method for a head mounted device according to claim 1, characterized in that: The action prediction model includes: Based on the recurrent neural network structure, the eye tracking data and head movement data in the historical time series are used as input, combined with the gaze direction change trend in the time window, to predict the gaze point or gaze area in the future predetermined time period; The viewing area corresponding to the prediction result is mapped to the image data matrix, and the corresponding image is acquired, cached or pre-fused in advance to reduce image delay and loading jams when the viewing angle changes suddenly.
9. A multi-camera imaging system for a head-mounted device, characterized in that: The system applies a multi-camera imaging method for a head mounted device as described in any one of claims 1 to 8, including: Image data acquisition module: obtain images and corresponding image information, and integrate them into image data matrix; obtain user attention information; Weight acquisition module: obtains the gaze weight of image data based on the angle between the user's gaze direction and the image acquisition plane of each camera, and obtains the fusion weight based on the image information; Image fusion module: fuses images through gaze weight and fusion weight, and generates the main view image according to the preset imaging mode; Feedback control module: divides the main view image into regions according to the image region labels, evaluates the clarity of the segmented regions, and performs feedback control on the imaging parameters of the camera based on the evaluation results; Viewpoint movement determination module: evaluates the stability of the user's head movement data and determines the viewpoint movement mode based on the evaluation results; Preloading module: Input the user's attention information into the pre-trained action prediction model, obtain the user's line of sight change trend, and preload related perspective images.
Citation Information
Patent Citations
Method and system for fusing multiple images
CA2822150A1
Eye movement gaze image prediction method based on hierarchical gaze image and conditional random field
CN108596243A
Image processing method and device, augmented reality system, computer device and medium
CN112887646A
Image processing method and device, equipment and storage medium
CN115767252A
Image processing method and device
CN116114260A
Cited By
Micro-display system and head-mounted display device
CN120812235A
Image data processing method and system based on heterogeneous computing platform
CN122244160A