Image processing method and apparatus, and electronic device, storage medium and program product
By identifying and replacing flicker-free areas with flicker-free areas in XR devices, the problem of image flickering in indoor lighting environments has been solved, improving the user experience and the smoothness of image display.
Patent Information
- Application Number
- PCT/CN2025/087394
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-27
- Filing Date
- 2025-04-07
- Publication Date
- 2025-12-04
AI Technical Summary
In existing XR devices, the frequency of ambient light emission is mismatched with the camera's frame rate in indoor lighting conditions, resulting in noticeable flickering in the ambient image and affecting the user's experience.
By acquiring the Nth frame environmental image and the previous reference image, flickering stripes are identified and replaced with non-flickering stripe areas in the reference image to eliminate flickering. Image synthesis is then performed in conjunction with head movement direction to optimize the image processing workflow.
It effectively eliminates environmental image flicker, improves the user experience of XR devices, reduces image display latency and object misalignment, and enhances the smoothness of image display.
Smart Images

Figure CN2025087394_04122025_PF_FP_ABST
Abstract
Description
Image processing methods, apparatuses, electronic devices, storage media, and software products
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410666524.8, filed in China on May 27, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present invention relates to the field of image processing technology, and in particular to an image processing method, apparatus, electronic device, storage medium and program product. Background Technology
[0004] The extended reality (XR) industry is developing rapidly. XR includes augmented reality (AR), virtual reality (VR), and mixed reality (MR). Some XR devices have see-through cameras to capture environmental images. However, the frequency of ambient light emission can interfere with the camera's frame rate, especially when indoor lights are on. This severe interference causes noticeable flickering in the captured environmental images, resulting in discomfort for the XR device wearer and negatively impacting their experience. Summary of the Invention
[0005] This invention provides an image processing method, apparatus, electronic device, storage medium, and program product to solve the problem of flickering in environmental images acquired by existing XR devices, which affects the user experience of XR device wearers.
[0006] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0007] In a first aspect, embodiments of the present invention provide an image processing method, including a step of performing image synthesis, wherein the step of performing image synthesis includes:
[0008] Obtain the environment image of the Nth frame;
[0009] At least one environmental image preceding the Nth frame of the environmental image is obtained as a reference image;
[0010] Obtain the flicker stripes of the Nth frame environmental image;
[0011] For each flicker stripe in the Nth frame environmental image, a target reference image is selected from the reference image. The position in the target reference image corresponding to the flicker stripe is a non-flicker stripe region. The flicker stripe in the Nth frame environmental image is replaced with the non-flicker stripe region to obtain the processed environmental image as the Nth frame environmental image to be displayed or merged.
[0012] Optionally, the conditions for performing the image compositing step include one of the following:
[0013] When the extended reality device is in perspective mode, the image compositing step is performed;
[0014] When the extended reality device is in perspective mode and the emission frequency of ambient light does not meet the preset relationship with the camera's frame rate, the image compositing step is performed.
[0015] When the extended reality device is in perspective mode, the emission frequency of ambient light and the camera's frame rate do not meet the preset relationship, and light is detected in the external environment, the image synthesis step is performed.
[0016] The preset relationship is as follows: the camera's frame rate is an integer multiple of the ambient light emission frequency.
[0017] Optionally, the method further includes one of the following:
[0018] If the perspective mode is not enabled on the extended reality device, the image compositing step is not performed;
[0019] When the extended reality device is in perspective mode and the emission frequency of ambient light and the camera's frame rate meet a preset relationship, the acquired Nth frame environmental image is used as the Nth frame environmental image to be displayed or to be fused.
[0020] When the extended reality device is in perspective mode, the emission frequency of ambient light and the camera's frame rate do not meet the preset relationship, and no light is detected in the external environment, the acquired Nth frame environmental image is used as the Nth frame environmental image to be displayed or to be merged.
[0021] The preset relationship is as follows: the camera's frame rate is an integer multiple of the ambient light emission frequency.
[0022] Optionally, acquiring the flicker stripes of the Nth frame environmental image includes:
[0023] The target image is converted to grayscale to obtain a grayscale image, wherein the target image includes the Nth frame environment image and the reference image;
[0024] For each pixel row of the grayscale image, if all pixel values in the pixel row are the same, or if the pixel row has different pixel values, determine the number m of pixels with the same pixel value in the pixel row. If the difference between m and the total number of pixels M in the pixel row is less than or equal to a preset threshold, determine that the pixel row is a blinking row.
[0025] The strobe stripes are determined based on all the flickering lines of the grayscale image.
[0026] Optionally, obtaining at least one environmental image preceding the Nth frame environmental image as a reference image includes:
[0027] Obtain the n environmental images preceding the Nth frame environmental image, where n is an integer greater than or equal to 1;
[0028] The reference image is an image whose similarity to the Nth environmental image is greater than or equal to a similarity threshold among the nth environmental images.
[0029] Optionally, for each flicker stripe in the Nth frame environmental image, a target reference image is selected from the reference image, wherein the position corresponding to the flicker stripe in the target reference image is a non-flicker stripe region, and the flicker stripe in the Nth frame environmental image is replaced with the non-flicker stripe region, including:
[0030] Search steps: For each flicker stripe in the Nth frame of the environmental image, search in the ith frame of the reference image whether the position corresponding to the flicker stripe is a flicker stripe, where the initial value of i is N-1;
[0031] Replacement step: If the position in the i-th frame of the environmental image corresponding to the flicker stripe is a non-flicker stripe area, replace the flicker stripe with the non-flicker stripe area;
[0032] Update step: If the position in the i-th frame of the environmental image corresponding to the strobe stripe is a strobe stripe, execute i = i-1 and return to the search step.
[0033] Optionally, the conditions for performing the image synthesis step include: performing the image synthesis step if the detected head movement speed does not exceed a speed threshold;
[0034] The method further includes: when the head movement speed is detected to exceed the speed threshold, the acquired Nth frame environmental image is used as the Nth frame environmental image to be displayed or to be fused.
[0035] Optionally, for each flicker stripe in the Nth frame environmental image, a target reference image is selected from the reference image, wherein the position corresponding to the flicker stripe in the target reference image is a non-flicker stripe region, and the flicker stripe in the Nth frame environmental image is replaced with the non-flicker stripe region, including:
[0036] For each flicker stripe in the Nth frame environmental image, select a target reference image that is temporally closest to the Nth frame environmental image from the reference images;
[0037] The non-flicker stripe region in the selected target reference image is horizontally translated by a first distance in the opposite direction of the head movement direction to obtain the translated non-flicker stripe region;
[0038] The non-flicker stripe region after translation is used to replace the flicker stripe.
[0039] Optionally, the first distance is sinα*D, where α is the angle of head rotation, α=β*Δt, β is the angular velocity of the head, Δt is the time interval for the camera to acquire environmental images, and D is the distance from the human eye to the image plane.
[0040] Optional, also includes:
[0041] The Nth frame environmental image to be fused is fused with the virtual image to be displayed to obtain the fused image to be displayed.
[0042] Optional, also includes:
[0043] Acquire the user's eye image and determine the user's gaze point based on the eye image;
[0044] Based on the gaze point, the image of the gaze region in the image to be processed is determined; the image to be processed is the Nth frame environment image to be displayed or the fused image to be displayed, and the fused image to be displayed is obtained by fusing the Nth frame environment image to be fused with the virtual image to be displayed;
[0045] Determine the image processing complexity of the image of the gaze region and the image processing complexity of the image to be processed;
[0046] The target resolution and target refresh rate of the image to be processed are determined based on the image processing complexity of the image to be processed and the image rendering data capacity of the image processor.
[0047] Based on the image processing complexity of the gaze region, the target resolution, and the target refresh rate, determine the first resolution and the first refresh rate corresponding to the image of the gaze region;
[0048] Determine the second resolution and second refresh rate of the non-focused region of the image to be processed;
[0049] The image of the gaze region is rendered according to the first resolution and the first refresh rate, and the image of the non-gaze region is rendered according to the second resolution and the second refresh rate.
[0050] The rendered image of the gaze region is stitched together with the rendered image of the non-gaze region to obtain the rendered image.
[0051] Display the rendered image.
[0052] Optionally, at least one of the first refresh rate and the second refresh rate is the same as the target refresh rate;
[0053] and / or
[0054] The first resolution is equal to the product of a first value and the target resolution, wherein the first value is greater than 0 and less than 1;
[0055] and / or
[0056] The second resolution is equal to the difference between the target resolution and the first resolution;
[0057] and / or
[0058] The first value is related to the image processing complexity of the image of the gaze region;
[0059] and / or
[0060] The first value is the sum of the ratio of the image processing complexity of the image of the gaze region to the image processing complexity of the image to be processed and a preset adjustment coefficient.
[0061] Optionally, the target resolution and target refresh rate of the image to be processed are determined based on the image processing complexity of the image to be processed and the image rendering data capacity of the image processor, including:
[0062] Based on the image processing complexity of the image to be processed, estimate the amount of data to be rendered in the image to be processed;
[0063] Based on the amount of data to be rendered, determine whether the image processor's image rendering data capacity exceeds the capacity threshold;
[0064] If the amount of data to be rendered does not exceed the image rendering data capacity of the image processor, a preset resolution is used as the target resolution, and a preset refresh rate is used as the target refresh rate.
[0065] If the amount of data to be rendered exceeds the image rendering data capacity of the image processor, a preset resolution is used as the target resolution, the preset refresh rate is reduced, and the reduced refresh rate is used as the target refresh rate.
[0066] Optionally, displaying the rendered image includes:
[0067] If the rendering of the image to be processed in the Nth frame has been completed before the vertical synchronization signal of the Nth frame arrives, the rendered image to be processed is displayed.
[0068] If the rendering of the image to be processed in the Nth frame is not completed before the vertical synchronization signal of the Nth frame arrives, the image generated by time warping of the previous frame of the image to be processed is displayed.
[0069] Optional, also includes:
[0070] When the perspective mode is not enabled on the extended reality device, after receiving the vertical synchronization signal of the Nth frame, the rendering thread renders the virtual image to be displayed in the left eye based on the predicted head motion data at the intermediate moment between the vertical synchronization signals of the N+1th frame and the N+2th frame.
[0071] At the midpoint between the vertical synchronization signal of frame N and the vertical synchronization signal of frame N+1, the rendered virtual image to be displayed for the left eye is passed to the asynchronous time warp thread for correction, and the rendering thread renders the virtual image to be displayed for the right eye.
[0072] After the vertical synchronization signal arrives in the N+1th frame, the rendered virtual image to be displayed for the right eye is passed to the asynchronous time warp thread for correction, and the corrected virtual image to be displayed for the left eye is then displayed.
[0073] The right eye virtual image to be displayed is rendered at the midpoint between the vertical synchronization signal of frame N+1 and the vertical synchronization signal of frame N+2.
[0074] Optionally, the acquired Nth frame environment image is used as the Nth frame environment image to be displayed or to be merged, including:
[0075] At the midpoint between the vertical synchronization signal of frame N and the vertical synchronization signal of frame N+1, the binocular environmental images of the camera are acquired, including the left-eye environmental image and the right-eye environmental image of frame N.
[0076] The binocular environment image is rendered based on the predicted head motion data at the intermediate moment between the vertical synchronization signal of frame N+1 and the vertical synchronization signal of frame N+2.
[0077] After the vertical synchronization signal arrives in the N+1th frame, the rendered binocular environment image is displayed.
[0078] Optionally, acquiring the Nth frame environmental image and acquiring at least one frame environmental image preceding the Nth frame environmental image as a reference image includes: acquiring the Nth frame environmental image and at least one frame environmental image preceding the Nth frame environmental image at an intermediate time between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame.
[0079] The processed environment image is obtained as the Nth frame environment image to be displayed or merged, and then includes:
[0080] Based on the predicted head motion data at the moment of the vertical synchronization signal of the (N+1)th frame, an asynchronous time distortion correction operation is performed on the environmental image of the Nth frame to be displayed or fused.
[0081] After the vertical synchronization signal of the N+1th frame arrives, the environmental image of the Nth frame to be displayed or merged is displayed after asynchronous time warp correction.
[0082] In a second aspect, embodiments of the present invention provide an image processing apparatus, including an image compositing module, the image compositing module comprising:
[0083] The first acquisition submodule is used to acquire the Nth frame of the environmental image;
[0084] The second acquisition submodule is used to acquire at least one frame of environmental image before the Nth frame of environmental image as a reference image;
[0085] The third acquisition submodule is used to acquire the flicker stripes of the Nth frame environmental image;
[0086] The image synthesis submodule is used to select a target reference image from the reference image for each flicker stripe in the Nth frame environmental image, wherein the position in the target reference image corresponding to the target flicker stripe is a non-flicker stripe region, and replace the flicker stripe in the Nth frame environmental image with the non-flicker stripe region to obtain a processed environmental image as the Nth frame environmental image to be displayed or to be merged.
[0087] Thirdly, embodiments of the present invention provide an electronic device, including a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the image processing method described in the first aspect above.
[0088] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the image processing method described in the first aspect above.
[0089] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the image processing method as described in the first aspect.
[0090] In this embodiment of the invention, the Nth frame environmental image captured by the camera of the XR device and at least one frame of environmental image before the Nth frame external environment image are image synthesized to eliminate flicker stripes in the Nth frame environmental image and improve the user experience of the XR device. Attached Figure Description
[0091] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0092] Figure 1 is a schematic flowchart of one embodiment of the image processing method of the present invention;
[0093] Figure 2 is a second schematic flowchart of the image processing method according to an embodiment of the present invention;
[0094] Figure 3 is a schematic diagram of the flicker stripes in an image according to an embodiment of the present invention;
[0095] Figure 4 is a schematic diagram of the Nth frame environmental image and reference image according to an embodiment of the present invention;
[0096] Figure 5 is a schematic diagram of the method of replacing the target strobe stripe with a non-strobe stripe region according to an embodiment of the present invention;
[0097] Figure 6 is a schematic flowchart of the image processing method according to an embodiment of the present invention (third one).
[0098] Figure 7 is one of the schematic diagrams showing the positional relationship between the head and the image in an embodiment of the present invention;
[0099] Figure 8 is a schematic diagram of the change in the position of an object in an image when the head moves, according to an embodiment of the present invention;
[0100] Figure 9 is a schematic flowchart of the image processing method according to an embodiment of the present invention (third one).
[0101] Figure 10 is a timing diagram of frame interpolation according to an embodiment of the present invention;
[0102] Figure 11 is a second schematic diagram of the positional relationship between the head and the image in an embodiment of the present invention;
[0103] Figure 12 is a timing diagram of image display in an extended reality device without the perspective mode enabled, according to an embodiment of the present invention.
[0104] Figure 13 is a timing diagram of direct image display in an extended reality device with perspective mode enabled, according to an embodiment of the present invention.
[0105] Figure 14 is a timing diagram of the image synthesis and display after the extended reality device is in perspective mode when the extended reality device is in perspective mode according to an embodiment of the present invention.
[0106] Figure 15 is a schematic diagram of the image processing device according to an embodiment of the present invention;
[0107] Figure 16 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0108] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0109] Please refer to Figure 1. An embodiment of the present invention provides an image processing method applied to an XR device. The XR device includes a camera, which is a see-through camera. The image processing method includes:
[0110] Step 10: Performing the image compositing step, which includes:
[0111] Step 11: Acquire the Nth frame of the environmental image; N is an integer greater than 1; the environmental image is acquired by the camera.
[0112] Step 12: Obtain at least one environmental image frame preceding the Nth frame environmental image as a reference image;
[0113] Step 13: Obtain the flicker stripes of the Nth frame environmental image;
[0114] The flickering stripes are caused by the flickering phenomenon and can also be called flickering stripes, flickering areas, forehead wrinkles, water ripples, dark stripes, etc.
[0115] Step 14: For each flicker stripe in the Nth frame environmental image, select a target reference image from the reference image. The position in the target reference image corresponding to the target flicker stripe is a non-flicker stripe region. Replace the flicker stripe in the Nth frame environmental image with the non-flicker stripe region to obtain the processed environmental image as the Nth frame environmental image to be displayed or merged.
[0116] The non-flicker stripe region refers to a region without flicker stripes. In this embodiment of the invention, the Nth frame environmental image captured by the XR device's camera and at least one frame of environmental image preceding the Nth frame external environment image are combined to eliminate flicker stripes in the Nth frame environmental image and improve the user experience of the XR device.
[0117] In some embodiments, optionally, the conditions for performing the image compositing step include performing the image compositing step when the extended reality device has the perspective mode (also known as perspective function) enabled.
[0118] In some embodiments, the method optionally further includes: not performing the image synthesis step when the extended reality device is not in perspective mode. In this case, the camera can also capture environmental images for purposes such as pose recognition.
[0119] Referring to Figure 2, in some embodiments, the XR device may first detect whether perspective mode is enabled. If perspective mode is not enabled, the image compositing step is not performed. At this time, the camera can also capture environmental images for purposes such as pose recognition. If perspective mode is enabled, the image compositing step is then performed.
[0120] In some embodiments, optionally, referring to Figure 2, the conditions for performing the image compositing step include: performing the image compositing step when the extended reality device is in perspective mode and the emission frequency of ambient light and the camera's frame rate do not satisfy a preset relationship. The preset relationship is: the camera's frame rate is an integer multiple of the emission frequency of ambient light, i.e., the camera's exposure time is an integer multiple of 1 / f, where f is the emission frequency of ambient light. When the preset relationship is satisfied, no flickering occurs. When the preset relationship is not satisfied, flickering occurs.
[0121] In some embodiments, optionally, the conditions for performing the image compositing step include: when the extended reality device is in perspective mode, the emission frequency of ambient light and the camera's frame rate do not satisfy a preset relationship, and light is detected in the external environment, the image compositing step is performed. The preset relationship is that the camera's frame rate is an integer multiple of the emission frequency of ambient light. In some embodiments, the presence or absence of light in the external environment can be detected by analyzing the content of the environmental image.
[0122] In some embodiments, optionally, the conditions for performing the image synthesis step include: when the extended reality device is in perspective mode, the emission frequency of ambient light does not meet a preset relationship with the camera's frame rate, and no light is detected in the external environment, the acquired Nth frame environmental image is used as the Nth frame environmental image to be displayed or to be fused.
[0123] In some embodiments, optionally, referring to Figure 2, the method further includes: when the extended reality device is in perspective mode and the emission frequency of ambient light and the camera's frame rate satisfy a preset relationship, the acquired Nth frame environmental image is used as the Nth frame environmental image to be displayed or to be merged; wherein, the preset relationship is: the camera's frame rate is an integer multiple of the emission frequency of ambient light.
[0124] Please refer to Figure 3. Typically, the flickering stripes F in an image are elongated regions composed of multiple consecutive flashing lines l. The flickering stripes F move vertically from top to bottom along the image's longitudinal direction, and their horizontal color is generally consistent. The areas in the image of Figure 3 without flickering stripes are called non-flickering stripe areas.
[0125] In some embodiments, optionally, acquiring the flicker stripes of the Nth frame environmental image includes:
[0126] Step 131: Convert the target image to grayscale to obtain a grayscale image, wherein the target image includes the Nth frame environment image and the reference image;
[0127] Alternatively, the formula Y = 0.229R + 0.587G + 0.114B can be used to convert the RGB image to a grayscale image.
[0128] Step 132: For each pixel row of the grayscale image, if all pixel values in the pixel row are the same, or if the pixel row has different pixel values, determine the number m of pixels with the same pixel value in the pixel row. If the difference between m and the total number of pixels M in the pixel row is less than or equal to a preset threshold, determine that the pixel row is a blinking row.
[0129] For example, the difference between the number of pixels with the same pixel value *m* in a pixel row and the total number of pixels *M* in the pixel row is *d*, i.e., *d* = *M*m. Assuming a preset threshold *T* = 5, if *d* <= 5, then the pixel row is a flickering row; otherwise, it is not. Of course, the preset threshold is not limited to 5; it can be other values, set according to specific circumstances.
[0130] Step 133: Determine the strobe stripes based on all the flicker lines of the grayscale image.
[0131] For the target image, identify all flickering lines. Multiple consecutive flickering lines form strobe stripes, which can be labeled as F. i (x i ,y i ,h i ), where (x i ,y i ) represents the coordinates of the upper left corner of the flickering stripe, h i This indicates the height (i.e., number of rows) of the flicker stripes. For the target image, identify and mark all flicker stripes.
[0132] In some embodiments, the n frames of environmental images preceding the Nth frame can be directly obtained as reference images. For example, the 3 frames of environmental images preceding the Nth frame can be directly obtained as reference images.
[0133] In some embodiments, reference images may be selected from the n environmental images preceding the Nth environmental image, where a portion of the images with a high similarity to the Nth true environmental image are chosen. For example, two of the three environmental images preceding the Nth environmental image may be used as reference images.
[0134] Optionally, referring to Figures 2 and 4, at least one frame of environmental image preceding the Nth frame of environmental image can be obtained as a reference image, including:
[0135] Step 121: Obtain the n environmental images preceding the Nth frame environmental image, where n is an integer greater than or equal to 1;
[0136] Step 122: Use the image in the n-frame environment image that has a similarity greater than or equal to the similarity threshold with the N-frame environment image as the reference image.
[0137] Referring to Figure 4, it can be seen that the Nth frame environmental image has two reference images, and the similarity between the two reference images and the Nth frame environmental image is relatively high.
[0138] In some embodiments, optionally, for each flicker stripe in the Nth frame environmental image, a target reference image is selected from the reference image, wherein the position corresponding to the flicker stripe in the target reference image is a non-flicker stripe region, and the flicker stripe in the Nth frame environmental image is replaced with the non-flicker stripe region, including:
[0139] Search steps: For each flicker stripe in the Nth frame of the environmental image, search in the ith frame of the reference image whether the position corresponding to the flicker stripe is a flicker stripe, where the initial value of i is N-1;
[0140] Replacement step: If the position in the i-th frame of the environmental image corresponding to the flicker stripe is a non-flicker stripe area, replace the flicker stripe in the N-th frame of the environmental image with the non-flicker stripe area;
[0141] Update step: If the position in the i-th frame of the environmental image corresponding to the strobe stripe is a strobe stripe, execute i = i-1 and return to the search step.
[0142] For example, as shown in Figure 5, using the three environmental images preceding the Nth frame as reference images, the flickering stripes in each environmental image (N, N-1, N-2, N-3) are first identified and marked. For flickering stripe F1 in the Nth frame, a search is performed at the corresponding position in the (N-1)th frame. If this position is not a flickering stripe in the (N-1)th frame, the pixel value at the corresponding position in the (N-1)th frame replaces the pixel value at the corresponding position in the Nth frame. For example, this can be done by copying the pixel value at the corresponding position in the (N-1)th frame and overwriting the pixel value at the corresponding position in the Nth frame. For flickering stripe F2 in the Nth frame, if the corresponding position in the (N-1)th frame is also a flickering stripe, the search continues at the corresponding position in the (N-2)th frame, and so on, until a non-flickering stripe area is found. Based on the above method, a full-color image with flicker stripes eliminated is synthesized until all flicker stripes in the Nth frame of the environmental image are replaced with the pixel values of the non-flicker stripe areas in the reference image.
[0143] In some embodiments, for multiple flicker stripes in the Nth frame environmental image, the target reference image with the closest time sequence may have all the corresponding non-flicker stripe regions. In this case, the number of target reference images is 1. It is only necessary to obtain the non-flicker stripe region corresponding to each flicker stripe in the target reference image and replace the corresponding flicker stripe with the reference to the non-flicker stripe region.
[0144] In some embodiments, for multiple flicker stripes in the Nth frame environmental image, the most recent target reference image only contains non-flicker stripe regions corresponding to some of the flicker stripe locations. In this case, it is necessary to obtain the non-flicker stripe regions corresponding to the other flicker stripe locations from other target reference images. In this embodiment, the pixel values of the flicker stripes in the Nth frame environmental image are replaced with the pixel values of the non-flicker stripe regions in the reference image at the most recent moment of the Nth frame environmental image, because images that are relatively close usually have a high degree of similarity.
[0145] Please refer to Figure 6. An embodiment of the present invention provides an image processing method applied to an XR device. In this embodiment, the XR device may be, for example, a helmet-mounted display (HMD) or VR glasses, an XR device worn on the head or face. The XR device includes a camera, which is a see-through camera. The image processing method includes:
[0146] Step 20: If the head movement speed is detected to be below the speed threshold, perform the image synthesis step, which includes:
[0147] Step 21: Acquire the Nth frame of the environmental image; N is an integer greater than 1; the environmental image is acquired by the camera.
[0148] Step 22: Obtain at least one environmental image frame preceding the Nth frame environmental image as a reference image;
[0149] Step 23: Obtain the flicker stripes of the Nth frame environmental image;
[0150] Step 24: For each flicker stripe in the Nth frame environmental image, select a target reference image from the reference image. The position in the target reference image corresponding to the flicker stripe is a non-flicker stripe region. Replace the flicker stripe in the Nth frame environmental image with the non-flicker stripe region to obtain the processed environmental image as the Nth frame environmental image to be displayed or merged.
[0151] In some embodiments, optionally, the image processing method further includes: if the head movement speed is detected to exceed a speed threshold, using the acquired Nth frame environmental image as the Nth frame environmental image to be displayed or to be merged. If the head movement speed is detected to exceed the speed threshold, it indicates that the user wearing the XR device is moving their head too fast. If the image compositing step is performed at this time, image compositing takes time, which may cause a delay (lag) in image display. To avoid the delay (lag) caused by image compositing, the Nth frame environmental image is directly obtained as the Nth frame environmental image to be displayed or to be merged for subsequent processing, such as image fusion processing, rendering processing, etc.
[0152] In this embodiment of the invention, referring to Figure 7, the XR device can move with the user's head movement. As shown in Figure 8, the position of object A is different in the Nth frame, the (N-1)th frame, and the (N-2)th frame. If the non-flicker stripe region in the (N-1)th or (N-2)th frame is directly used to replace the corresponding flicker stripe in the Nth frame, object A will obviously be misaligned. Therefore, in this embodiment, it is first necessary to find the non-flicker stripe region corresponding to the flicker stripe in the Nth frame in the reference image, and then horizontally translate this non-flicker stripe region so that the part of object A in the Nth frame coincides with the part of object A in the non-flicker stripe region. However, due to the translation, some areas of F1 in the Nth frame are not completely replaced (this part is not included in the reference image). In this case, the unreplaced areas can be replaced by the corresponding area of the Nth frame itself, or the content of the unreplaced area can be predicted based on the pixels around the unreplaced area.
[0153] In some embodiments, optionally, for each flicker stripe in the Nth frame environmental image, a target reference image is selected from the reference image, wherein the position corresponding to the flicker stripe in the target reference image is a non-flicker stripe region, and the flicker stripe in the Nth frame environmental image is replaced with the non-flicker stripe region, including:
[0154] Step 241: For each flicker stripe in the Nth frame environmental image, select a target reference image that is temporally closest to the Nth frame environmental image from the reference images; the position corresponding to the flicker stripe in the target reference image is a non-flicker stripe region;
[0155] Assuming that the flicker stripe F1 in the Nth frame environmental image has corresponding non-flicker stripe regions in both the N-1th and N-2th frame environmental images, and the N-1th frame environmental image is closer to the Nth frame environmental image in terms of time sequence, then the N-1th frame environmental image is selected as the target reference image corresponding to the flicker stripe F1.
[0156] For example, regarding a flickering stripe F1 in the Nth frame of the environmental image, it can first be determined whether the position corresponding to the flickering stripe F1 in the (N-1)th frame of the environmental image is a non-flickering stripe area. If so, the (N-1)th frame of the environmental image is used as the target reference image. If not, it is then determined whether the position corresponding to the flickering stripe F1 in the (N-2)th frame of the environmental image is a non-flickering stripe area, and so on, until a target reference image is found where the position corresponding to the flickering stripe F1 is a non-flickering stripe area.
[0157] Step 242: Shift the selected non-flicker stripe region in the target reference image horizontally by a first distance in the opposite direction of head movement to obtain the shifted non-flicker stripe region;
[0158] Please refer to Figure 7. Since the human head and the XR device are relatively fixed, the translation distance of the human eye (i.e., the first distance) during the time interval of the camera acquiring environmental images can be represented by sinα*D.
[0159] Optionally, the first distance is sinα*D, where α is the angle of head rotation, α=β*Δt, β is the angular velocity of the head, Δt is the time interval for the camera to acquire environmental images, and D is the distance from the human eye to the image plane.
[0160] Step 243: Replace the flicker stripes with the translated non-flicker stripe region.
[0161] In this embodiment of the invention, when performing image synthesis, the head movement direction and head movement speed (angular velocity of the head) are combined to perform image synthesis, thereby avoiding stuttering problems under head movement, ensuring smooth display, and preventing misalignment of objects in the synthesized image.
[0162] In some embodiments, the method further includes: fusing the Nth frame environmental image to be fused with the virtual image to be displayed to obtain a fused image to be displayed. Subsequent operations such as rendering are then performed on the fused image to be displayed for display.
[0163] In some other embodiments, the Nth frame of the environment image to be displayed may also be rendered for display.
[0164] In existing extended reality processing technologies, the extended reality system renders the entire field of view based on data reported by sensors, using pose prediction. After post-rendering processing, the image is output, and the extended reality display terminal directly scans and displays it according to the time sequence. However, the full-view clarity is not the same as the real visual situation of the human eye, and it also causes a great waste of rendering resources. This is mainly because the human eye has an ultra-high-definition area within 3° of the fovea, where it can perceive fine details. As the viewpoint expands outward, the sensitivity of the human eye to fine details gradually decreases. For example, in an area of about 20°, the human eye can perceive text and symbols, but at a larger viewpoint, the human eye can only perceive colors and outlines.
[0165] To address the above problems, embodiments of the present invention provide an image processing method applied to an XR device, wherein the XR device includes a camera, the camera being a see-through camera, and the image processing method includes:
[0166] Step 31: Acquire the user's eye image and determine the user's gaze point based on the eye image;
[0167] Step 32: Determine the image of the gaze region in the image to be processed based on the gaze point; the image to be processed is the Nth frame environment image to be displayed or the fused image to be displayed, and the fused image to be displayed is obtained by fusing the Nth frame environment image to be fused with the virtual image to be displayed;
[0168] Step 33: Determine the image processing complexity of the image of the gaze region and the image processing complexity of the image to be processed;
[0169] In some embodiments, optionally, the image processing complexity of the fused image to be displayed can be determined based on the image complexity of the Nth frame environmental image and the image complexity of the virtual image to be displayed. The image complexity of the Nth frame environmental image can be determined based on the image content of the environmental image. The image complexity of the virtual image to be displayed can be determined based on the image content of the virtual image to be displayed.
[0170] Step 34: Determine the target resolution and target refresh rate of the image to be processed based on the image processing complexity of the image to be processed and the image rendering data capacity of the image processor;
[0171] Refresh rate, also known as frame rate, refers to how many times per second a display device redraws the image, or the number of times the display is refreshed per second, measured in Hertz (Hz). A higher refresh rate results in a more stable, natural, and clearer image, with less strain on the user's eyes. Conversely, a lower refresh rate leads to more flickering and jittering, causing faster eye fatigue.
[0172] Step 35: Determine the first resolution and first refresh rate corresponding to the image of the gaze region based on the image processing complexity of the gaze region, the target resolution, and the target refresh rate;
[0173] Step 36: Determine the second resolution and second refresh rate of the non-focused region of the image to be processed;
[0174] Step 37: Render the image of the gaze region according to the first resolution and the first refresh rate, and render the image of the non-gaze region according to the second resolution and the second refresh rate;
[0175] Step 38: Stitch the rendered image of the gaze region with the rendered image of the non-gaze region to obtain the rendered image;
[0176] Step 39: Display the rendered image.
[0177] In some embodiments, optionally, at least one of the first refresh rate and the second refresh rate is the same as the target refresh rate.
[0178] In some embodiments, optionally, the first resolution is equal to the product of a first value and the target resolution, wherein the first value is greater than 0 and less than 1.
[0179] In some embodiments, optionally, the second resolution is equal to the difference between the target resolution and the first resolution.
[0180] In some embodiments, optionally, the first value is related to the image processing complexity of the image of the gaze region.
[0181] In some embodiments, optionally, the first value is related to the image processing complexity of the image of the gaze region. Optionally, the first value is the sum of the ratio of the image processing complexity of the image of the gaze region to the image processing complexity of the image to be processed and a preset adjustment coefficient.
[0182] In some embodiments, optionally, the target resolution and target refresh rate of the image to be processed are determined based on the image processing complexity of the image to be processed and the image rendering data capacity of the image processor, including:
[0183] Step 341: Estimate the amount of data to be rendered in the image to be processed based on the image processing complexity of the image to be processed;
[0184] Step 342: Based on the amount of data to be rendered, determine whether the image processor's image rendering data capacity exceeds the capacity threshold;
[0185] Step 343: If the amount of data to be rendered does not exceed the image rendering data capacity of the image processor, a preset resolution is used as the target resolution, and a preset refresh rate is used as the target refresh rate.
[0186] Step 344: If the amount of data to be rendered exceeds the image rendering data capacity of the image processor, a preset resolution is used as the target resolution, the preset refresh rate is reduced, and the reduced refresh rate is used as the target refresh rate.
[0187] In this embodiment of the invention, the image rendering data capacity of the image processor (GPU) is monitored to see if it exceeds the capacity threshold. If it does not exceed the capacity threshold, a preset resolution can be used as the target resolution and a preset refresh rate can be used as the target refresh rate. If it exceeds the capacity threshold, the refresh rate of the image to be processed can be reduced to avoid delay (lag) in image display.
[0188] Please refer to Figure 9, which is a flowchart illustrating an image processing method according to an embodiment of the present invention. The image processing method includes:
[0189] Step 1': Determine the data information of the image to be processed;
[0190] Step 1: Capture an image of the user's eyes and determine the user's gaze point based on the image;
[0191] Step 2: Determine the location information of the gaze area based on the user's gaze point;
[0192] Step 3: Calculate the image processing complexity of the gaze region in the image to be processed and the image processing complexity of the image to be processed;
[0193] Step 3': Monitor whether the amount of data to be rendered by the image processor (GPU) exceeds the image rendering data capacity;
[0194] Step 3: Determine the target resolution and target refresh rate of the image to be processed;
[0195] Optionally, if the amount of data to be rendered does not exceed the image rendering data capacity of the image processor, a preset resolution is used as the target resolution and a preset refresh rate is used as the target refresh rate.
[0196] If the amount of data to be rendered exceeds the image rendering data capacity of the image processor, a preset resolution is used as the target resolution, the preset refresh rate is reduced, and the reduced refresh rate is used as the target refresh rate.
[0197] Step 4: Calculate the first resolution and first refresh rate of the image corresponding to the gaze region based on the image processing complexity of the gaze region and the target resolution and target refresh rate;
[0198] Step 5: Calculate the second resolution and second refresh rate of the non-focal regions in the image to be processed;
[0199] Step 6: Render the image of the gaze area according to the first resolution and the first refresh rate, and render the image of the non-gaze area according to the second resolution and the second refresh rate;
[0200] Step 7: Stitch the rendered image of the gaze region with the rendered image of the non-gaze region to obtain the rendered image to be processed;
[0201] Step 8: Display the rendered image to be processed.
[0202] To reduce rendering latency in display scenes, some extended reality devices employ time-warp (TW) technology. Time-warp is a technique that corrects image frames based on changes in user actions after rendering. It addresses rendering latency by warping (or correcting) the rendered scene data. Because the time-warp process occurs closer to the display time, the resulting image is closer to what the user expects. Furthermore, since time-warp only processes two-dimensional images, it's similar to affine transformations in image processing, minimizing system overhead.
[0203] Asynchronous Timewarp (ATW) further optimizes the aforementioned timewarp technique by separating rendering and timewarp into two different threads. This allows the rendering and timewarp steps to be executed asynchronously, reducing the overall runtime of both processes. For example, when extended reality applications cannot maintain a sufficient frame rate, the asynchronous timewarp thread reprocesses the previously rendered scene data based on the current user posture to generate a frame (intermediate frame) that matches the current user posture, thus reducing image jitter and improving latency. This technique is crucial for reducing latency and alleviating motion sickness. Without ATW, if no new rendered frame is output when the Vsync signal arrives (meaning the GPU (Graphics Processing Unit) has not finished rendering the Nth frame), the data from the previous N-1 frame will be displayed twice. This does not cause any issues when the user's head remains stationary. However, when the user rotates from their old head position to a new one, the image corresponding to the old head position is still displayed on the screen. At this point, the brain has switched to the new head position, but the eyes are receiving content from the old head position. This mismatch in information can cause dizziness, and the greater the angle of rotation, the stronger this dizziness becomes. Asynchronous time warp addresses this by generating a new frame based on the latest completed frame from the rendering process before each vertical synchronization. If no new rendering frame is output when the Vsync signal arrives (i.e., the GPU has not finished rendering the Nth frame), the image from the (N-1)th frame is used to correct the data and generate the new image frame corresponding to the new head position (scanning and outputting the corresponding Nth frame). This method ensures that a new frame is generated based on the latest head position before each Vsync signal arrives, significantly reducing dizziness and improving VR comfort.
[0204] In some embodiments, after reducing the preset refresh rate, frame interpolation is required to ensure image quality. Optionally, displaying the rendered image includes:
[0205] Step 391: If the rendering of the image to be processed in the Nth frame has been completed before the vertical synchronization (Vsync) signal of the Nth frame arrives, display the rendered image to be processed.
[0206] Vsync (Vertical Synchronization) refers to a technique used in computer graphics processing to synchronize screen refresh rate and graphics rendering. It can also be called a frame synchronization signal.
[0207] Vsync is a synchronization signal applied between two frames, indicating the end of the previous frame and the beginning of the next. This signal is activated once before each frame scan and determines the display device's field frequency, i.e., the number of times the screen refreshes per second, also known as the display device's refresh rate. For example, the refresh rate of a display device can be 60Hz, 120Hz, etc., meaning 60 or 120 refreshes per second, and a display frame time is 1 / 60th of a second, 1 / 120th of a second, etc. This vertical synchronization signal is generated, for example, by a display driver (e.g., a graphics card) and is used to synchronize the gate drive signals, data signals, etc., required to display one frame of an image during the display process.
[0208] Step 392: If the rendering of the image to be processed in the Nth frame has not been completed before the vertical synchronization signal of the Nth frame arrives, display the image generated by time warping the previous frame of the image to be processed.
[0209] In some embodiments, the alternative method for performing temporal warp (TW) on an image is as follows:
[0210] Please refer to the frame interpolation timing diagram in Figure 10. After displaying the (N+1)th frame, the camera refresh rate decreases. Therefore, when rendering of the (N+3)th frame begins (after the Vsync signal of the (N+3)th frame arrives), the camera has not yet acquired the image of the (N+3)th frame, nor has it performed the image compositing step (eliminating flicker stripes). At this point, a new image is generated for display after TW (transfer-over-time) of the (N+2)th frame image. The specific steps are as follows:
[0211] As shown in Figure 11, the head rotates from pose 1 to pose 2. The angle sensor can provide the quaternion Quant for the current pose. Given the quaternions Quant1 and Quant2 at these two moments (pose 1 and pose 2), the rotation quaternion QuantR between them is equal to Quant1. -1 *Quant2. Quant1 -1 =Quant1* / ||Quant1||,Quant1 * = (w, -x, -y, -z) and ||Quant1|| = 1. Then, generate the reprojection matrix R based on QuantantR(x,y,z,w), and multiply the pixel at each position on the image by this matrix to obtain a new position, and add the corresponding pixel value to generate the image after TW. When the camera frame rate decreases, render the N+3rd frame image. If the rendering can be completed before the vertical sync signal of the N+4th frame arrives, the display of the N+3rd frame uses the image rendered in the N+3rd frame; otherwise, it uses the image generated by TW from the N+2th frame.
[0212] The reprojection matrix R can be represented as follows:
[0213] Where w is the real part of the quaternion, and x, y, z are the imaginary parts of the quaternion.
[0214] In some embodiments, the extended reality device does not enable perspective mode and will directly display the virtual image without displaying the environment image. Please refer to Figure 12. Optionally, the image processing method of this embodiment of the invention further includes:
[0215] When the extended reality device is not in perspective mode (only virtual images are displayed, not ambient images), after receiving the vertical synchronization (Vsync) signal of the Nth frame, the rendering thread (MRT) renders the virtual image to be displayed for the left eye based on the predicted head motion data at the midpoint between the vertical synchronization signals of the N+1th frame and the N+2th frame (i.e., 1.5 frames later).
[0216] At the midpoint between the vertical sync signal of frame N and the vertical sync signal of frame N+1, the rendered virtual image to be displayed for the left eye is passed to the asynchronous time warp (ATW) thread for correction, and the rendering thread renders the virtual image to be displayed for the right eye.
[0217] After the vertical synchronization signal arrives in the N+1th frame, the rendered virtual image to be displayed for the right eye is passed to the asynchronous time warp thread for correction, and the corrected virtual image to be displayed for the left eye is then displayed.
[0218] The right eye virtual image to be displayed is rendered at the midpoint between the vertical synchronization signal of frame N+1 and the vertical synchronization signal of frame N+2.
[0219] In some embodiments, when the extended reality device is in perspective mode, the environmental image can be displayed directly without displaying the virtual image. Referring to Figure 13, optionally, the acquired Nth frame environmental image is used as the Nth frame environmental image to be displayed or merged, including:
[0220] At the midpoint between the vertical synchronization signal of frame N and the vertical synchronization signal of frame N+1, the binocular environmental images of the camera are acquired, including the left-eye environmental image and the right-eye environmental image of frame N.
[0221] The binocular environment image is rendered based on the predicted head motion data from the midpoint between the vertical synchronization signal of frame N+1 and the vertical synchronization signal of frame N+2 (i.e., a time one frame later than the midpoint between the vertical synchronization signal of frame N and the vertical synchronization signal of frame N+1).
[0222] After the vertical synchronization signal arrives in the N+1th frame, the rendered binocular environment image is displayed.
[0223] In some embodiments, when the extended reality device is in perspective mode, it can display the Nth frame of the environment image obtained after the image compositing step. Optionally, as shown in Figure 14, optional...
[0224] Acquiring the Nth frame environmental image, and acquiring at least one frame environmental image preceding the Nth frame environmental image as a reference image, includes: at the midpoint between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame, acquiring the Nth frame environmental image and at least one frame environmental image preceding the Nth frame environmental image (such as the previous n frames environmental images).
[0225] After obtaining the Nth frame environmental image and at least one frame of environmental image preceding the Nth frame environmental image, the above-mentioned step of removing flicker stripes is performed, namely: obtaining the flicker stripes of the Nth frame environmental image; for each flicker stripe in the Nth frame environmental image, selecting a target reference image from the reference image, wherein the position corresponding to the flicker stripe in the target reference image is a non-flicker stripe region, and replacing the flicker stripe in the Nth frame environmental image with the non-flicker stripe region, thereby obtaining a processed environmental image as the Nth frame environmental image to be displayed or to be merged.
[0226] The processed environment image is obtained as the Nth frame environment image to be displayed or merged, and then includes:
[0227] Based on the predicted head motion data at the moment of the vertical synchronization signal of the (N+1)th frame, the Nth frame environmental image to be displayed or fused is subjected to asynchronous time warp (ATW) correction.
[0228] After the vertical synchronization signal of the N+1th frame arrives, the environmental image of the Nth frame to be displayed or fused (including the left-eye image and the right-eye image) after asynchronous time warp correction is performed is displayed.
[0229] Please refer to Figure 15. This embodiment of the invention also provides an image processing apparatus 20, including an image compositing module 21, the image compositing module 21 comprising:
[0230] The first acquisition submodule 211 is used to acquire the Nth frame of the environment image;
[0231] The second acquisition submodule 212 is used to acquire at least one frame of environmental image before the Nth frame of environmental image as a reference image;
[0232] The third acquisition submodule 213 is used to acquire the flicker stripes of the Nth frame environmental image;
[0233] The image synthesis submodule 214 is used to select a target reference image from the reference image for each flicker stripe in the Nth frame environmental image, wherein the position corresponding to the flicker stripe in the target reference image is a non-flicker stripe region, and the non-flicker stripe region is used to replace the flicker stripe in the Nth frame environmental image to obtain a processed environmental image as the Nth frame environmental image to be displayed or to be merged.
[0234] Optionally, the conditions for performing the image compositing step include one of the following:
[0235] When the extended reality device is in perspective mode, the image compositing step is performed;
[0236] When the extended reality device is in perspective mode and the emission frequency of ambient light does not meet the preset relationship with the camera's frame rate, the image compositing step is performed.
[0237] When the extended reality device is in perspective mode, the emission frequency of ambient light and the camera's frame rate do not meet the preset relationship, and light is detected in the external environment, the image synthesis step is performed.
[0238] The preset relationship is as follows: the camera's frame rate is an integer multiple of the ambient light emission frequency.
[0239] Optionally, the image processing device 20 further includes one of the following:
[0240] The first processing module is used to prevent the image compositing step from being performed when the perspective mode of the extended reality device is not enabled;
[0241] The second processing module is used to take the acquired Nth frame environmental image as the Nth frame environmental image to be displayed or to be fused when the extended reality device is in perspective mode and the emission frequency of ambient light and the frame rate of the camera meet a preset relationship.
[0242] The third processing module is used to take the acquired Nth frame environmental image as the Nth frame environmental image to be displayed or to be merged when the extended reality device is in perspective mode, the emission frequency of ambient light and the camera's photosensitive frame rate do not meet the preset relationship, and no light is detected in the external environment.
[0243] The preset relationship is as follows: the camera's frame rate is an integer multiple of the ambient light emission frequency.
[0244] Optionally, the third acquisition submodule 213 is used to perform grayscale processing on the target image to obtain a grayscale image, wherein the target image includes the Nth frame environment image and the reference image; for each pixel row of the grayscale image, if the pixel values in the pixel row are all the same, or if the pixel row has different pixel values, the number m of pixels with the same pixel value in the pixel row is determined; if the difference between m and the total number of pixels M in the pixel row is less than or equal to a preset threshold, the pixel row is determined to be a flickering row; and the strobe stripe is determined based on all the flickering rows of the grayscale image.
[0245] Optionally, the second acquisition submodule 212 is used to acquire n environmental images preceding the Nth frame environmental image, where n is an integer greater than or equal to 1; and to use the image in the n frames environmental images whose similarity to the Nth frame environmental image is greater than or equal to a similarity threshold as the reference image.
[0246] Optionally, the image synthesis submodule 214 is configured to perform:
[0247] Search steps: For each flicker stripe in the Nth frame of the environmental image, search in the ith frame of the reference image whether the position corresponding to the flicker stripe is a flicker stripe, where the initial value of i is N-1;
[0248] Replacement step: If the position in the i-th frame of the environmental image corresponding to the target flicker stripe is a non-flicker stripe region, replace the flicker stripe in the N-th frame of the environmental image with the non-flicker stripe region;
[0249] Update step: If the position in the i-th frame of the environmental image corresponding to the target strobe stripe is a strobe stripe, execute i = i-1 and return to the search step.
[0250] Optionally, the conditions for performing the image synthesis step include: performing the image synthesis step if the detected head movement speed does not exceed a speed threshold;
[0251] The image processing device 20 further includes a fourth processing module, used to use the acquired Nth frame environmental image as the Nth frame environmental image to be displayed or to be fused when the head movement speed is detected to exceed the speed threshold.
[0252] Optionally, the image synthesis submodule 214 is configured to, for each flicker stripe in the Nth frame environmental image, select a target reference image that is temporally closest to the Nth frame environmental image from the reference image; horizontally shift the non-flicker stripe region in the selected target reference image by a first distance in the opposite direction of the head movement direction to obtain a shifted non-flicker stripe region; and replace the target flicker stripe with the shifted non-flicker stripe region.
[0253] Optionally, the first distance is sinα*D, where α is the angle of head rotation, α=β*Δt, β is the angular velocity of the head, Δt is the time interval for the camera to acquire environmental images, and D is the distance from the human eye to the image plane.
[0254] Optionally, the image processing device 20 further includes:
[0255] The first fusion module is used to fuse the Nth frame environmental image to be fused with the virtual image to be displayed to obtain the fused image to be displayed.
[0256] Optionally, the image processing device 20 further includes:
[0257] The first acquisition module is used to acquire the user's eye image and determine the user's gaze point based on the eye image;
[0258] The first determining module is used to determine the image of the gaze region in the image to be processed based on the gaze point; the image to be processed is the Nth frame environment image to be displayed or the fused image to be displayed, and the fused image to be displayed is obtained by fusing the Nth frame environment image to be fused with the virtual image to be displayed;
[0259] The second determining module is used to determine the image processing complexity of the image of the gaze region and the image processing complexity of the image to be processed;
[0260] The third determining module is used to determine the target resolution and target refresh rate of the image to be processed based on the image processing complexity of the image to be processed and the image rendering data capacity of the image processor.
[0261] The fourth determining module is used to determine the first resolution and the first refresh rate corresponding to the image of the gaze region based on the image processing complexity of the gaze region, the target resolution, and the target refresh rate;
[0262] The fifth determining module is used to determine the second resolution and second refresh rate of the non-focused region of the image to be processed;
[0263] A first rendering module is used to render the image of the gaze region according to the first resolution and the first refresh rate, and to render the image of the non-gaze region according to the second resolution and the second refresh rate.
[0264] The stitching module is used to stitch together the rendered image of the gaze region with the rendered image of the non-gaze region to obtain the rendered image.
[0265] The first display module is used to display the rendered image.
[0266] Optionally, at least one of the first refresh rate and the second refresh rate is the same as the target refresh rate;
[0267] and / or
[0268] The first resolution is equal to the product of a first value and the target resolution, wherein the first value is greater than 0 and less than 1;
[0269] and / or
[0270] The second resolution is equal to the difference between the target resolution and the first resolution;
[0271] and / or
[0272] The first value is related to the image processing complexity of the image of the gaze region;
[0273] and / or
[0274] The first value is the sum of the ratio of the image processing complexity of the image of the gaze region to the image processing complexity of the image to be processed and a preset adjustment coefficient.
[0275] Optionally, the third determining module is configured to: estimate the amount of data to be rendered in the image to be processed based on the image processing complexity of the image to be processed; determine whether the image rendering data volume capacity of the image processor exceeds a capacity threshold based on the amount of data to be rendered; if the amount of data to be rendered does not exceed the image rendering data volume capacity of the image processor, use a preset resolution as the target resolution and a preset refresh rate as the target refresh rate; if the amount of data to be rendered exceeds the image rendering data volume capacity of the image processor, use the preset resolution as the target resolution, reduce the preset refresh rate, and use the reduced refresh rate as the target refresh rate.
[0276] Optionally, the first display module is configured to, before the arrival of the vertical synchronization signal of the Nth frame, if the rendering of the image to be processed in the Nth frame has been completed, display the rendered image of the image to be processed; and before the arrival of the vertical synchronization signal of the Nth frame, if the rendering of the image to be processed in the Nth frame has not been completed, display an image generated by time warping of the previous frame image of the image to be processed.
[0277] Optionally, the image processing device 20 further includes:
[0278] The second rendering module is used to render the virtual image to be displayed in the left eye based on the predicted head motion data at the intermediate moment between the vertical synchronization signals of the N+1 and N+2 frames after receiving the vertical synchronization signal of the Nth frame, when the perspective mode of the extended reality device is not enabled.
[0279] The first correction module is used to transmit the rendered virtual image to be displayed in the left eye to the asynchronous time warp thread for correction at the intermediate moment between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame.
[0280] The third rendering module is used to render the virtual image to be displayed to the right eye at the intermediate moment between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame.
[0281] The second correction module is used to transmit the rendered virtual image to be displayed in the right eye to the asynchronous time warp thread for correction after the vertical synchronization signal of the N+1th frame arrives.
[0282] The second display module is used to display the corrected left-eye virtual image after the arrival of the vertical synchronization signal of the N+1th frame, and to display the rendered right-eye virtual image at the intermediate moment between the vertical synchronization signal of the N+1th frame and the vertical synchronization signal of the N+2th frame.
[0283] Optionally, the first processing module is used to acquire binocular environmental images of the camera at the intermediate time between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame, wherein the binocular environmental images include the left-eye environmental image and the right-eye environmental image of the Nth frame.
[0284] The image processing device 20 further includes:
[0285] The fourth rendering module is used to render the binocular environment image based on the predicted head motion data at the intermediate moment between the vertical synchronization signal of frame N+1 and the vertical synchronization signal of frame N+2.
[0286] The third display module is used to display the rendered binocular environment image after the vertical synchronization signal of the N+1th frame arrives.
[0287] Optionally, the second acquisition module 212 is used to acquire the Nth frame environmental image and at least one frame of environmental image preceding the Nth frame environmental image at the intermediate time between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame.
[0288] The third correction module is used to perform asynchronous time warp correction on the Nth frame environmental image to be displayed or fused based on the predicted head motion data at the time of the vertical synchronization signal of the N+1th frame.
[0289] The fourth display module is used to display the Nth frame environmental image to be displayed or merged after asynchronous time warp correction operation is performed after the vertical synchronization signal of the N+1th frame arrives.
[0290] Please refer to Figure 16. This embodiment of the invention also provides an electronic device 30, including a processor 31, a memory 32, and a computer program stored in the memory 32 and executable on the processor 31. When the computer program is executed by the processor 31, it implements the various processes of the above-described embodiment of the video ringback tone playback method applied to the ringback tone platform and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0291] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described image processing method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0292] This application also provides a computer program product, including computer instructions. When executed by a processor, the computer instructions implement the various processes of the method embodiments shown in Figures 1, 2, 6, or 9 above, and achieve the same technical effects. To avoid repetition, these will not be described again here.
[0293] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0294] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0295] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. An image processing method, characterized by, The method comprises a step of performing image synthesis, which comprises: acquiring an Nth frame of environment image; acquiring at least one frame of environment image before the Nth frame of environment image as a reference image; acquiring stroboscopic stripes of the Nth frame of environment image; for each stroboscopic stripe in the Nth frame of environment image, selecting a target reference image from the reference image, the corresponding position of the stroboscopic stripe in the target reference image being a non-stroboscopic stripe area, and replacing the stroboscopic stripe in the Nth frame of environment image with the non-stroboscopic stripe area to obtain a processed environment image as the Nth frame of environment image to be displayed or fused.
2. The image processing method of claim 1, wherein, The conditions for the step of performing image synthesis include one of the following: in the case that the extended reality device is in a see-through mode, the step of performing image synthesis is executed; in the case that the extended reality device is in the see-through mode and the light-emitting frequency of the ambient light and the light-sensing frame rate of the camera do not satisfy a preset relationship, the step of performing image synthesis is executed; in the case that the extended reality device is in the see-through mode, the light-emitting frequency of the ambient light and the light-sensing frame rate of the camera do not satisfy the preset relationship, and it is detected that there is light in the external environment, the step of performing image synthesis is executed; wherein the preset relationship is that the light-sensing frame rate of the camera is an integer multiple of the light-emitting frequency of the ambient light.
3. The image processing method of claim 1 or 2, characterized in that, The method further comprises one of the following: in the case that the extended reality device is not in the see-through mode, the step of performing image synthesis is not executed; in the case that the extended reality device is in the see-through mode and the light-emitting frequency of the ambient light and the light-sensing frame rate of the camera satisfy the preset relationship, the acquired Nth frame of environment image is taken as the Nth frame of environment image to be displayed or fused; in the case that the extended reality device is in the see-through mode, the light-emitting frequency of the ambient light and the light-sensing frame rate of the camera do not satisfy the preset relationship, and it is detected that there is no light in the external environment, the acquired Nth frame of environment image is taken as the Nth frame of environment image to be displayed or fused; wherein the preset relationship is that the light-sensing frame rate of the camera is an integer multiple of the light-emitting frequency of the ambient light.
4. The image processing method of claim 1, wherein, The acquisition of the stroboscopic stripes of the Nth frame of environment image comprises: gray processing a target image to obtain a gray image, the target image comprising the Nth frame of environment image and the reference image; for each pixel row of the gray image, if the pixel values in the pixel row are all the same, or if the pixel row has pixel values that are not the same, determining the number m of pixels with the same pixel value in the pixel row, and if the difference between m and the total number M of pixels in the pixel row is less than or equal to a preset threshold, determining the pixel row as a stroboscopic row; determining the stroboscopic stripes according to all the stroboscopic rows of the gray image.
5. The image processing method of claim 1, wherein, The acquisition of the reference image before the Nth frame of environment image comprises: acquiring n frames of environment image before the Nth frame of environment image, n being an integer greater than or equal to 1; taking the images in the n frames of environment image that have a similarity greater than or equal to a similarity threshold with the Nth frame of environment image as the reference image.
6. The image processing method of claim 1, wherein, The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of:
8. The image processing method of claim 1 or 7, characterized by, The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of:
9. The image processing method of claim 8, wherein, The method comprises the following steps of:
10. The image processing method of claim 1, wherein, The method comprises the following steps of: The method comprises the following steps of:
11. The image processing method of claim 1 or 10, characterized by, The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: The method comprises the following steps of: determining a first resolution and a first refresh rate corresponding to the gaze region image according to image processing complexity of the gaze region, the target resolution and the target refresh rate; determining a second resolution and a second refresh rate of the image of the non-gaze region in the image to be processed; rendering the image of the gaze region according to the first resolution and the first refresh rate, and rendering the image of the non-gaze region according to the second resolution and the second refresh rate; splicing the rendered image of the gaze region and the rendered image of the non-gaze region to obtain a rendered image; displaying the rendered image.
12. The image processing method of claim 11, wherein at least one of the first refresh rate and the second refresh rate is the same as the target refresh rate; and / or the first resolution is equal to a product of a first value and the target resolution, the first value being greater than 0 and less than 1; and / or the second resolution is equal to a difference between the target resolution and the first resolution; and / or the first value is related to the image processing complexity of the gaze region image; and / or the first value is a sum of a ratio of the image processing complexity of the gaze region image to the image processing complexity of the image to be processed and a preset adjustment coefficient.
13. The image processing method of claim 11, wherein, determining a target resolution and a target refresh rate of the image to be processed according to image processing complexity of the image to be processed and image rendering data volume capacity of an image processor, comprising: estimating a to-be-rendered data volume of the image to be processed according to the image processing complexity of the image to be processed; judging whether the image rendering data volume capacity of the image processor exceeds a capacity threshold according to the to-be-rendered data volume; in a case where the to-be-rendered data volume does not exceed the image rendering data volume capacity of the image processor, taking a preset resolution as the target resolution and taking a preset refresh rate as the target refresh rate; in a case where the to-be-rendered data volume exceeds the image rendering data volume capacity of the image processor, taking a preset resolution as the target resolution and reducing the preset refresh rate to take the reduced refresh rate as the target refresh rate.
14. The image processing method of claim 11, wherein, the displaying of the rendered image comprises: before a vertical synchronization signal of an Nth frame arrives, if rendering of the image to be processed of the Nth frame is completed, displaying the rendered image to be processed; before the vertical synchronization signal of the Nth frame arrives, if rendering of the image to be processed of the Nth frame is not completed, displaying an image generated by time warping of a last frame image of the image to be processed.
15. The image processing method of claim 1, wherein, further comprising: in a case where the extended reality device does not start a see-through mode, after receiving the vertical synchronization signal of the Nth frame, a rendering thread performs rendering of a left-eye to-be-displayed virtual image based on predicted head movement data at a middle time between a vertical synchronization signal of an N+1th frame and a vertical synchronization signal of an N+2th frame. the left eye virtual image to be displayed is rendered in the middle time between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame, and the right eye virtual image to be displayed is rendered by the rendering thread; after the vertical synchronization signal of the N+1th frame arrives, the right eye virtual image to be displayed is delivered to the asynchronous time warp thread for correction, and the left eye virtual image to be displayed after correction is displayed; the right eye virtual image to be displayed after rendering is displayed in the middle time between the vertical synchronization signal of the N+1th frame and the vertical synchronization signal of the N+2th frame.
16. The image processing method of claim 3, wherein, the acquired Nth frame environment image is taken as the Nth frame environment image to be displayed or to be fused, comprising: in the middle time between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame, the binocular environment image of the camera is acquired, and the binocular environment image comprises a left eye environment image of the Nth frame and a right eye environment image; the binocular environment image is rendered based on the prediction head motion data in the middle time between the vertical synchronization signal of the N+1th frame and the vertical synchronization signal of the N+2th frame; after the vertical synchronization signal of the N+1th frame arrives, the rendered binocular environment image is displayed.
17. The image processing method of claim 1, wherein the Nth frame environment image is acquired, and at least one frame environment image before the Nth frame environment image is acquired as a reference image, comprising: in the middle time between the vertical synchronization signal of the Nth frame and the vertical synchronization signal of the N+1th frame, the Nth frame environment image and at least one frame environment image before the Nth frame environment image are acquired; the processed environment image is taken as the Nth frame environment image to be displayed or to be fused, and then further comprising: the Nth frame environment image to be displayed or to be fused is subjected to an asynchronous time warp correction operation based on the prediction head motion data at the time of the vertical synchronization signal of the N+1th frame; after the vertical synchronization signal of the N+1th frame arrives, the Nth frame environment image to be displayed or to be fused after the asynchronous time warp correction operation is displayed.
18. An image processing apparatus characterized by comprising: comprising an image synthesis module, the image synthesis module comprising: a first acquisition submodule for acquiring an Nth frame environment image; a second acquisition submodule for acquiring at least one frame environment image before the Nth frame environment image as a reference image; a third acquisition submodule for acquiring a stroboscopic fringe of the Nth frame environment image; an image synthesis submodule for selecting a target reference image from the reference image for each stroboscopic fringe in the Nth frame environment image, the corresponding position of the stroboscopic fringe in the target reference image being a non-stroboscopic fringe area, and replacing the stroboscopic fringe in the Nth frame environment image with the non-stroboscopic fringe area to obtain a processed environment image as the Nth frame environment image to be displayed or to be fused.
19. An electronic device, comprising: comprising: a processor, a memory, and a program stored on the memory and executable on the processor, the program being executed by the processor to implement the steps of the image processing method of any one of claims 1 to 17.
20. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps of the image processing method according to any one of claims 1 to 17.
21. A computer program product, characterised in that, The computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps of the image processing method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Image processing device, signal processing device, and program
CN102356631A
Video processing method, storage medium, electronic equipment and video live broadcast system
CN112492375A
Image processing method and device, electronic equipment and computer readable storage medium
CN113674189A
Augmented reality image processing method, system and device and storage medium
CN116665004A
Image processing method and device, electronic equipment, storage medium and program product
CN118521492A