Image display method and device, equipment and storage medium

By judging the overlap between the hand object and the virtual screen in the head-mounted device, selecting three-dimensional grid information and combining the hand image, the problem of virtual objects blocking the hand is solved and the image display effect is improved.

CN120390073APending Publication Date: 2025-07-29BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410124010.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In head-mounted devices, virtual objects block the user's hands, resulting in poor image display effect.

Method used

By judging the overlap between the hand object and the virtual screen, select three-dimensional grid information, combine the target image of the hand object, and display the virtual screen and the three-dimensional hand image to form a correct occlusion relationship.

Benefits of technology

This improves the image display effect, forms a correct occlusion relationship between the three-dimensional hand image and the virtual object, and enhances the user's perception of the hand position.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390073A_ABST
    Figure CN120390073A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image display method and device, equipment and a storage medium, and the method comprises the steps: responding to a received real scene image which is collected by a camera module and comprises a hand object, and judging whether the hand object is overlapped with a virtual screen at a preset spatial position or not; if the hand object is overlapped with the virtual screen, target grid information corresponding to the hand object is selected from three-dimensional grid information corresponding to the real scene image, and the target grid information is used for representing three-dimensional space information corresponding to the hand object; intercepting a target image corresponding to the hand object from the real scene image, and combining the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object; and displaying the virtual screen and the three-dimensional hand image in a display interface of the head-mounted device according to the preset spatial position overlapped with the virtual screen and the three-dimensional spatial information corresponding to the hand object. The display effect of the image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of head-mounted devices, and in particular, to an image display method, apparatus, device, and storage medium. Background Art

[0002] With the development of head-mounted device technology, in order to combine elements of the real scene and improve the authenticity of the display of head-mounted devices, head-mounted devices can display images that combine the real scene and virtual objects. Among them, the virtual object can be a two-dimensional virtual object; for example, a virtual display screen. The virtual object can also be a three-dimensional virtual object; for example, a virtual aircraft model, a virtual cartoon character, etc.

[0003] In the prior art, when displaying an image that combines the real scene and virtual objects, generally the virtual object is superimposed in front of the real scene, so that the real scene does not block the virtual object, ensuring the integrity of the virtual object display.

[0004] However, the inventors found that the prior art has at least the following technical problems: when the real scene includes the user's hand, the virtual object will block the user's hand (for example, the virtual display screen will block the user's hand), resulting in the user being unable to perceive the position of their hand, so the display effect of the image is poor. Summary of the Invention

[0005] Embodiments of the present disclosure provide an image display method, apparatus, device, and storage medium, which can improve the display effect of the image.

[0006] In a first aspect, embodiments of the present disclosure provide an image display method, which is applied to a head-mounted device equipped with a camera module. The method includes:

[0007] In response to receiving a real scene image including a hand object collected by the camera module, determining whether the hand object overlaps with a virtual screen at a preset spatial position;

[0008] If the hand object overlaps with the virtual screen, selecting target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object;

[0009] Cropping a target image corresponding to the hand object from the real scene image, and combining the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object;

[0010] Displaying the virtual screen and the three-dimensional hand image on the display interface of the head-mounted device according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object.

[0011] In a second aspect, embodiments of the present disclosure provide an image display device, which is applied to a head-mounted device equipped with a camera module. The device includes:

[0012] A receiving module, configured to, in response to receiving a real-scene image including a hand object collected by the camera module, determine whether the hand object overlaps with a virtual screen at a preset spatial position;

[0013] A selection module, configured to, if the hand object overlaps with the virtual screen, select target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object;

[0014] A cropping module, configured to crop a target image corresponding to the hand object from the real-scene image, and combine the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object;

[0015] A display module, configured to display the virtual screen and the three-dimensional hand image in a display interface of the head-mounted device according to the preset spatial position where the virtual screen overlaps and the three-dimensional spatial information corresponding to the hand object.

[0016] In a third aspect, embodiments of the present disclosure provide an electronic device, including:

[0017] A processor, and a memory communicatively connected to the processor;

[0018] The memory stores computer-executable instructions;

[0019] The processor executes the computer-executable instructions stored in the memory to implement the image display method as described in the first aspect above.

[0020] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the image display method as described in the first aspect above is implemented.

[0021] In a fifth aspect, embodiments of the present disclosure provide a computer program product, including a computer program, which implements the image display method as described in the first aspect above when executed by a processor.

[0022] The image display method, device, equipment and storage medium provided in this embodiment, the method includes: in response to receiving a real-scene image collected by a camera module and including a hand object, determining whether the hand object overlaps with a virtual screen at a preset spatial position; if the hand object overlaps with the virtual screen, selecting target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object; intercepting a target image corresponding to the hand object from the real-scene image, and combining the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object; and displaying the virtual screen and the three-dimensional hand image on the display interface of the head-mounted device according to the preset spatial position where the virtual screen overlaps and the three-dimensional spatial information corresponding to the hand object. In the embodiment of the present application, the target image corresponding to the hand object is combined with the target grid information corresponding to the hand object to obtain a three-dimensional hand image that can be separated from the real-scene image, so that a correct occlusion relationship can be formed between the three-dimensional hand image and the virtual object, thereby improving the display effect of the image. Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic diagram of an application scenario of an image display method provided in an embodiment of the present disclosure;

[0025] Figure 2 It is a flowchart of an image display method provided in an embodiment of the present disclosure;

[0026] Figure 3 It is a flowchart of another image display method provided in an embodiment of the present disclosure;

[0027] Figure 4 It is a structural block diagram of an image display device provided in an embodiment of the present disclosure;

[0028] Figure 5 It is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. Detailed Embodiments

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0031] With the development of head-mounted device technology, in order to combine elements of the real scene and improve the authenticity of the display of the head-mounted device, the head-mounted device can display an image combining the real scene and virtual objects. Among them, the virtual object can be a two-dimensional virtual object; for example, a virtual display screen. The virtual object can also be a three-dimensional virtual object; for example, a virtual airplane model, a virtual cartoon character, etc.

[0032] In a virtual reality device based on an MR (Mixed Reality) characteristic scene, there are a large number of VST (Video See Through) pictures; among them, gesture interaction is the default interaction method in the VST pictures. At the same time, the VST pictures are taken by the device and necessarily include images of the user's hands. In virtual reality interaction, a virtual display screen or a virtual object will be superimposed on the VST pictures. At this time, it is necessary to handle the occlusion relationship between the VST user hand image and the real hand image.

[0033] In the prior art, when displaying an image combining the real scene and virtual objects, generally the virtual object is superimposed in front of the real scene, so that the real scene will not occlude the virtual object, ensuring the integrity of the virtual object display. However, when the real scene includes the user's hand, the virtual object will occlude the user's hand (for example, the virtual display screen will occlude the user's hand), resulting in the user being unable to perceive the position of their hand, so the display effect of the image is poor.

[0034] Therefore, it can be seen that how to handle the occlusion relationship between the virtual object and the user's hand to improve the display effect of the image is an urgent problem to be solved at present.

[0035] To solve the above problems, the present embodiment provides the following technical concept: First, in response to receiving a real - scene image including a hand object collected by a camera module, it is determined whether the hand object overlaps with a virtual screen at a preset spatial position. Second, if the hand object overlaps with the virtual screen, target grid information corresponding to the hand object is selected from the three - dimensional grid information corresponding to the real - scene image, where the target grid information is used to represent the three - dimensional spatial information corresponding to the hand object. Then, a target image corresponding to the hand object is intercepted from the real - scene image, and the target image is combined with the target grid information to obtain a three - dimensional hand image corresponding to the hand object. Finally, according to the preset spatial position corresponding to the virtual screen and the three - dimensional hand image corresponding to the hand object, the virtual screen and the three - dimensional hand image are displayed on the display interface of the head - mounted device.

[0036] In the present application, since the target image corresponding to the hand object is combined with the target grid information corresponding to the hand object to obtain a three - dimensional hand image that can be separated from the real - scene image, a correct occlusion relationship can be formed between the three - dimensional hand image and the virtual object, thus improving the display effect of the image.

[0037] The application scenarios of the embodiments of the present disclosure are explained below:

[0038] The image display method provided by the embodiments of the present disclosure can be applied to the display scenarios of virtual display screens and hand objects. Figure 1 It is a schematic diagram of the application scenario of an image display method provided by the embodiments of the present disclosure. As Figure 1 shown, the display image includes a real - scene image 101, a virtual display screen 102, and a three - dimensional hand image 103. Among them, through this image display method, the three - dimensional hand image 103 of the hand object can be determined, and the three - dimensional hand image 103 is separated from the real - scene image 101, so that a correct occlusion relationship can be formed between the three - dimensional hand image 103 and the virtual display screen 102, thus improving the display effect of the image. The image display method provided by the embodiments of the present disclosure is described in detail below with specific embodiments.

[0039] Figure 2 It is a flowchart of an image display method provided by the embodiments of the present disclosure. This image display method can be applied to a head - mounted device equipped with a camera module. As Figure 2 shown, the method includes:

[0040] S201. In response to receiving a real - scene image including a hand object collected by a camera module, determine whether the hand object overlaps with a virtual screen at a preset spatial position.

[0041] In an embodiment of the present disclosure, the camera module may be a binocular camera. By using the binocular camera, real-scene images from two different angles can be obtained. In this way, three-dimensional information corresponding to the real-scene images can be obtained from the real-scene images from the two different angles.

[0042] Optionally, the data collected by the camera module is a video stream including multiple frames of real-scene images. Among them, the three-dimensional picture corresponding to the current frame image is displayed on the display interface of the head-mounted device.

[0043] In an embodiment of the present disclosure, the virtual screen may be located at any preset spatial position in the display interface of the head-mounted device. Wherein, the preset spatial position is used to define the shape, size, and depth information of the virtual screen.

[0044] In some embodiments, the coordinate area corresponding to the virtual screen may be stored in advance. Correspondingly, determining whether the hand object overlaps with the virtual screen at the preset spatial position may include: obtaining the three-dimensional coordinates of multiple vertices included in the hand object, and if the three-dimensional coordinates of any vertex are within the coordinate area corresponding to the virtual screen, it is determined that the hand object overlaps with the virtual screen at the preset spatial position; if the three-dimensional coordinates of the multiple vertices are not within the coordinate area corresponding to the virtual screen, it is determined that the hand object does not overlap with the virtual screen at the preset spatial position.

[0045] S202: If the hand object overlaps with the virtual screen, select the target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object.

[0046] In an embodiment of the present disclosure, the real-scene image includes pixel regions corresponding to multiple objects respectively; correspondingly, selecting the target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image includes: determining the three-dimensional grid information corresponding to the real-scene image, where the three-dimensional grid information includes grid information corresponding to multiple pixel regions respectively; according to the target pixel region corresponding to the hand object, select the target grid information corresponding to the target pixel region from the grid information corresponding to multiple pixel regions respectively.

[0047] Exemplarily, the three-dimensional grid information may be a mesh (grid) map corresponding to the real-scene image. The real-scene image may include: a pixel region corresponding to the hand, a pixel region corresponding to the table, and a pixel region corresponding to the ground. The grid map corresponding to each pixel region may be stored in advance. For example, the hand grid map corresponding to the hand pixel region, the table grid map corresponding to the table pixel region, and the ground grid map corresponding to the ground pixel region may be stored in advance.

[0048] Optionally, determining the three-dimensional grid information corresponding to the real scene image includes: for the pixel area corresponding to each object, extracting multiple target pixels from the pixel area according to a preset sampling ratio; determining the three-dimensional grid vertices corresponding to each of the multiple target pixels, and constructing multiple adjacent and non-overlapping triangular facets through the three-dimensional grid vertices corresponding to the multiple target pixels to obtain grid information of the pixel area corresponding to the object; combining the grid information of the pixel areas corresponding to the multiple objects to obtain the three-dimensional grid information corresponding to the real scene image.

[0049] In the embodiments of the present disclosure, the numerical value of the preset sampling ratio is not specifically limited and can be set and modified as needed. For example, the preset sampling ratio is 1:10, then one target pixel is extracted for every 10 pixels, and the target pixel is determined to be a three-dimensional grid vertex. It should be noted that the larger the preset sampling ratio, the more sampling points there are, and the higher the accuracy of the corresponding grid information, but the data that needs to be collected will also increase; conversely, the smaller the preset sampling ratio, the smaller the sampling points, and the lower the accuracy of the corresponding grid information, but the data that needs to be collected will also decrease.

[0050] Optionally, to improve the accuracy of the grid information corresponding to the hand object, the sampling ratio for the hand object can be increased. Accordingly, for each pixel region corresponding to the object, multiple target pixels are extracted from the pixel region at a preset sampling ratio, including: for each pixel region corresponding to the object, if the object is a hand object, the multiple target pixels are extracted from the pixel region at a first preset sampling ratio; if the object is not a hand object, the multiple target pixels are extracted from the pixel region at a second preset sampling ratio; wherein the first preset sampling ratio is greater than the second preset sampling ratio.

[0051] In the embodiment of the present disclosure, the values of the first preset sampling ratio and the second preset sampling ratio are not specifically limited and can be set and modified as needed.

[0052] S203. Capture a target image corresponding to the hand object from the real scene image, and combine the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object.

[0053] In an embodiment of the present disclosure, an AI (Artificial Intelligence) model can be used to identify and extract a hand object captured from a real-world scene image. Accordingly, capturing a target image corresponding to the hand object from the real-world scene image includes: identifying the hand object in the real-world scene image using an artificial neural network model to obtain a recognition result corresponding to the hand object; and capturing a target image corresponding to the hand object from the real-world scene image based on the recognition result corresponding to the hand object.

[0054] In some embodiments, in order to improve the accuracy of the target image corresponding to the captured hand object, a small hand area can be first divided, and then the target image corresponding to the hand object can be captured from the hand area. Accordingly, capturing the target image corresponding to the hand object from the real-scene image includes: selecting a partial scene image including the hand object from the real-scene image, and capturing the target image corresponding to the hand object from the partial scene image; wherein, the display ratio of the partial scene image is greater than the display ratio of the real-scene image.

[0055] S204. Display the virtual screen and the three-dimensional hand image on the display interface of the head-mounted device according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object.

[0056] In the embodiments of the present disclosure, the three-dimensional hand image and the virtual screen can be rendered through the occlusion relationship between the virtual screen and the three-dimensional hand image. Accordingly, this step is: determining the occlusion relationship between the virtual screen and the three-dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object; rendering the virtual object and the three-dimensional hand image according to the occlusion relationship between the virtual screen and the three-dimensional hand image to obtain a composite image, and displaying the composite image on the display interface of the head-mounted device.

[0057] Optionally, the three-dimensional hand image includes the three-dimensional coordinates of multiple first vertices corresponding to the hand object, and the preset spatial position includes the three-dimensional coordinates of multiple second vertices corresponding to the virtual screen. In this case, the occlusion relationship between the virtual screen and the three-dimensional hand image can be determined through the depth information in the three-dimensional coordinates. Accordingly, determining the occlusion relationship between the virtual screen and the three-dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object includes: for each first vertex in the three-dimensional hand image, selecting a target vertex at the same position from the multiple second vertices corresponding to the virtual screen, where the same position means that the coordinate values in the X-axis direction and the Y-axis direction are both the same; if the coordinate value of the first vertex in the Z-axis direction is less than the coordinate value of the target vertex in the Z-axis direction, the first vertex in the three-dimensional hand image is occluded by the virtual screen; if the coordinate value of the first vertex in the Z-axis direction is greater than the coordinate value of the target vertex in the Z-axis direction, the virtual screen is occluded by the first vertex in the three-dimensional hand image.

[0058] In some embodiments, if the first vertex in the three-dimensional hand image is occluded by the virtual screen, the rendering priority of the first vertex is lower than the rendering priority of the virtual screen, and in this case, the first vertex will be occluded by the virtual screen in the obtained composite image.

[0059] In some other embodiments, if the virtual screen is occluded by a first vertex in the three-dimensional hand image, the rendering priority of the first vertex is higher than that of the virtual screen. In this case, the virtual screen will be occluded by the first vertex in the obtained composite image.

[0060] Embodiments of the present disclosure provide an image display method, which includes: in response to receiving a real-scene image collected by a camera module and including a hand object, determining whether the hand object overlaps with a virtual screen at a preset spatial position; if the hand object overlaps with the virtual screen, selecting target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object; intercepting a target image corresponding to the hand object from the real-scene image, and combining the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object; and displaying the virtual screen and the three-dimensional hand image on the display interface of the head-mounted device according to the preset spatial position where the virtual screen overlaps and the three-dimensional spatial information corresponding to the hand object. In the embodiments of the present application, the target image corresponding to the hand object is combined with the target grid information corresponding to the hand object to obtain a three-dimensional hand image that can be separated from the real-scene image. In this way, a correct occlusion relationship can be formed between the three-dimensional hand image and the virtual object, thereby improving the display effect of the image.

[0061] It should be noted that the present application can track the hand object and then implement interactive operations with the virtual object. Correspondingly, as Figure 3 shown, the method includes:

[0062] S301. Determine multiple tracking points corresponding to the hand object, perform trajectory tracking on the multiple tracking points, and obtain an operation instruction of the hand object for the virtual object.

[0063] Optionally, different moving gestures of the hand object correspond to different operation instructions. Multiple moving gestures and their corresponding operation instructions can be preset and stored respectively.

[0064] Correspondingly, performing trajectory tracking on the multiple tracking points to obtain an operation instruction of the hand object for the virtual object includes: performing trajectory tracking on the multiple tracking points to obtain position information corresponding to each of the multiple tracking points within a preset duration, where the position information includes first position information before the tracking point moves and second position information after the tracking point moves; obtaining the moving gesture of the hand object according to the position information corresponding to each of the multiple tracking points; and determining the operation instruction corresponding to the moving gesture as the operation instruction of the hand object for the virtual object.

[0065] In the embodiments of the present disclosure, the value of the preset duration is not specifically limited and can be set and modified as needed. Exemplarily, the preset duration can be 1 s, 5 s, 10 s, etc.

[0066] S302. Control the virtual object to perform the operation corresponding to the operation instruction.

[0067] Exemplarily, the virtual object to be controlled is a virtual display screen; the operation instruction is: the zoom-in operation of the virtual display screen. Correspondingly, this step is to zoom in on the virtual display screen.

[0068] It should be noted that, in order to increase the fusion degree between the three-dimensional hand image and the virtual screen and improve the display effect of the synthesized image, feathering processing or transparency processing can be performed on some areas of the three-dimensional hand image.

[0069] In some embodiments, feathering processing is performed on the edge region of the three-dimensional hand image, where the distance between the edge region and the edge of the three-dimensional hand image is less than a first preset distance.

[0070] In this embodiment, the value of the first preset distance is not specifically limited and can be set and modified as needed. Exemplarily, the first preset distance can be 0.5 cm. Optionally, the ratio between the first preset distance and the width of the three-dimensional hand image is a preset ratio. In this embodiment, the value of the preset ratio is not specifically limited and can be set and modified as needed. Exemplarily, the ratio between the first preset distance and the width of the three-dimensional hand image is 1 / 10.

[0071] Here, since feathering processing is performed on the edge region of the three-dimensional hand image, increasing the fusion degree between the three-dimensional hand image and the virtual screen, the display effect of the synthesized image is improved.

[0072] In some other embodiments, the three-dimensional hand image includes a hand and an arm, and transparency processing is performed on the target arm region of the arm, where the distance between the target arm region and the hand is greater than a second preset distance.

[0073] In this embodiment, the value of the second preset distance is not specifically limited and can be set and modified as needed. Optionally, the second preset distance can be the length of the forearm. At this time, transparency processing can be performed on the upper arm.

[0074] Here, since transparency processing is performed on the target arm region of the arm, and the target arm region is far from the hand, in this way, the interference of the target arm region can be avoided through the transparency processing, further increasing the fusion degree between the three-dimensional hand image and the virtual screen, and thus improving the display effect of the synthesized image.

[0075] Figure 4 It is a structural block diagram of an image display device provided by an embodiment of the present disclosure. This image display device is applied to a head-mounted device equipped with a camera module. Refer to Figure 4, the device includes: a receiving module 401, a selection module 402, a cropping module 403, and a display module 404.

[0076] Among them, the receiving module 401 is configured to, in response to receiving a real - world scene image including a hand object collected by a camera module, determine whether the hand object overlaps with a virtual screen at a preset spatial position;

[0077] The selection module 402 is configured to, if the hand object overlaps with the virtual screen, select target grid information corresponding to the hand object from the three - dimensional grid information corresponding to the real - world scene image, where the target grid information is used to represent the three - dimensional spatial information corresponding to the hand object;

[0078] The cropping module 403 is configured to crop a target image corresponding to the hand object from the real - world scene image, and combine the target image with the target grid information to obtain a three - dimensional hand image corresponding to the hand object;

[0079] The display module 404 is configured to display the virtual screen and the three - dimensional hand image on a display interface of the head - mounted device according to the preset spatial position where the virtual screen overlaps and the three - dimensional spatial information corresponding to the hand object.

[0080] According to one or more embodiments of the present disclosure, the display module 404 is configured to determine an occlusion relationship between the virtual screen and the three - dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three - dimensional hand image corresponding to the hand object; render the virtual object and the three - dimensional hand image according to the occlusion relationship between the virtual screen and the three - dimensional hand image to obtain a composite image, and display the composite image on the display interface of the head - mounted device.

[0081] According to one or more embodiments of the present disclosure, the three - dimensional hand image includes three - dimensional coordinates of a plurality of first vertices corresponding to the hand object, and the preset spatial position includes three - dimensional coordinates of a plurality of second vertices corresponding to the virtual screen; the display module 404 is configured to, for each first vertex in the three - dimensional hand image, select a target vertex at the same position from the plurality of second vertices corresponding to the virtual screen, where the same position means that the coordinate values in the X - axis direction and the Y - axis direction are the same; if the coordinate value of the first vertex in the Z - axis direction is less than the coordinate value of the target vertex in the Z - axis direction, the first vertex in the three - dimensional hand image is occluded by the virtual screen; if the coordinate value of the first vertex in the Z - axis direction is greater than the coordinate value of the target vertex in the Z - axis direction, the virtual screen is occluded by the first vertex in the three - dimensional hand image.

[0082] According to one or more embodiments of the present disclosure, the real - world scene image includes pixel regions corresponding to multiple objects respectively; correspondingly, the selection module 402 is configured to determine three - dimensional grid information corresponding to the real - world scene image, where the three - dimensional grid information includes grid information corresponding to each of the multiple pixel regions respectively; and select target grid information corresponding to the target pixel region from the grid information corresponding to each of the multiple pixel regions according to the target pixel region corresponding to the hand object.

[0083] According to one or more embodiments of the present disclosure, the selection module 402 is configured to extract multiple target pixels from each pixel region corresponding to an object according to a preset sampling ratio; determine three - dimensional grid vertices corresponding to each of the multiple target pixels respectively, and construct multiple non - overlapping adjacent triangular patches through the three - dimensional grid vertices corresponding to each of the multiple target pixels respectively to obtain grid information of the pixel region corresponding to the object; and combine the grid information of the pixel regions corresponding to multiple objects respectively to obtain three - dimensional grid information corresponding to the real - world scene image.

[0084] According to one or more embodiments of the present disclosure, the selection module 402 is configured to, for each pixel region corresponding to an object, if the object is a hand object, extract multiple target pixels from the pixel region according to a first preset sampling ratio; if the object is not a hand object, extract multiple target pixels from the pixel region according to a second preset sampling ratio; where the first preset sampling ratio is greater than the second preset sampling ratio.

[0085] According to one or more embodiments of the present disclosure, the cropping module 403 is configured to identify the hand object in the real - world scene image through an artificial neural network model to obtain an identification result corresponding to the hand object; and crop a target image corresponding to the hand object from the real - world scene image according to the identification result corresponding to the hand object.

[0086] According to one or more embodiments of the present disclosure, the cropping module 403 is configured to select a local scene image including the hand object from the real - world scene image, and crop a target image corresponding to the hand object from the local scene image; where the display ratio of the local scene image is greater than the display ratio of the real - world scene image.

[0087] According to one or more embodiments of the present disclosure, the device further includes: an operation module; the operation module is configured to determine multiple tracking points corresponding to the hand object, perform trajectory tracking on the multiple tracking points to obtain an operation instruction of the hand object for the virtual object; and control the virtual object to execute the operation corresponding to the operation instruction.

[0088] According to one or more embodiments of the present disclosure, the operation module is configured to perform trajectory tracking on the multiple tracking points to obtain position information respectively corresponding to each of the multiple tracking points within a preset time period, where the position information includes first position information before the tracking point moves and second position information after the tracking point moves; obtain a moving gesture of the hand object according to the position information respectively corresponding to each of the multiple tracking points; and determine an operation instruction corresponding to the moving gesture as an operation instruction of the hand object for the virtual object.

[0089] According to one or more embodiments of the present disclosure, the device further includes: a processing module; the processing module is configured to perform feathering processing on an edge region of the three-dimensional hand image, where the distance between the edge region and the edge of the three-dimensional hand image is less than a first preset distance; and / or, the three-dimensional hand image includes a hand and an arm, and the processing module is configured to perform transparency processing on a target arm region of the arm, where the distance between the target arm region and the hand is greater than a second preset distance.

[0090] Among them, the receiving module 401, the selection module 402, the interception module 403, and the display module 404 are connected in sequence. The image display device provided in this embodiment can execute the technical solutions of the above method embodiments, and the implementation principles and technical effects are similar, which will not be elaborated here in this embodiment.

[0091] Figure 5 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present disclosure. Refer to Figure 5 , the electronic device 500 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present disclosure.

[0092] Such as Figure 5As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0093] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0094] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0095] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0096] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.

[0097] The above-mentioned computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.

[0098] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0100] The units involved in the embodiments described in this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0101] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0103] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an image display method, which is applied to a head-mounted device equipped with a camera module, the method comprising:

[0104] In response to receiving an image of a real scene including a hand object captured by a camera module, determining whether the hand object overlaps with a virtual screen at a preset spatial position;

[0105] If the hand object overlaps with the virtual screen, selecting target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real scene image, wherein the target grid information is used to represent the three-dimensional space information corresponding to the hand object;

[0106] intercepting a target image corresponding to the hand object from the real scene image, and combining the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object;

[0107] According to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object, the virtual screen and the three-dimensional hand image are displayed in the display interface of the head-mounted device.

[0108] According to one or more embodiments of the present disclosure, the virtual screen and the three-dimensional hand image are displayed in the display interface of the head-mounted device according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object, including: determining the occlusion relationship between the virtual screen and the three-dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object; rendering the virtual object and the three-dimensional hand image according to the occlusion relationship between the virtual screen and the three-dimensional hand image to obtain a composite image, and displaying the composite image in the display interface of the head-mounted device.

[0109] According to one or more embodiments of the present disclosure, the three-dimensional hand image includes three-dimensional coordinates of a plurality of first vertices corresponding to the hand object, and the preset spatial position includes three-dimensional coordinates of a plurality of second vertices corresponding to the virtual screen; determining an occlusion relationship between the virtual screen and the three-dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object includes: for each first vertex in the three-dimensional hand image, selecting a target vertex at the same position from the plurality of second vertices corresponding to the virtual screen, where the same position means that the coordinate values in the X-axis direction and the Y-axis direction are both the same; if the coordinate value of the first vertex in the Z-axis direction is less than the coordinate value of the target vertex in the Z-axis direction, the first vertex in the three-dimensional hand image is occluded by the virtual screen; if the coordinate value of the first vertex in the Z-axis direction is greater than the coordinate value of the target vertex in the Z-axis direction, the virtual screen is occluded by the first vertex in the three-dimensional hand image.

[0110] According to one or more embodiments of the present disclosure, the real-scene image includes pixel regions respectively corresponding to a plurality of objects; correspondingly, selecting target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image includes: determining the three-dimensional grid information corresponding to the real-scene image, where the three-dimensional grid information includes grid information respectively corresponding to the plurality of pixel regions; according to the target pixel region corresponding to the hand object, selecting the target grid information corresponding to the target pixel region from the grid information respectively corresponding to the plurality of pixel regions.

[0111] According to one or more embodiments of the present disclosure, determining the three-dimensional grid information corresponding to the real-scene image includes: for the pixel region corresponding to each object, extracting a plurality of target pixels from the pixel region according to a preset sampling ratio; determining three-dimensional grid vertices respectively corresponding to the plurality of target pixels, and constructing a plurality of adjacent and non-overlapping triangular patches through the three-dimensional grid vertices respectively corresponding to the plurality of target pixels to obtain the grid information of the pixel region corresponding to the object; combining the grid information of the pixel regions respectively corresponding to the plurality of objects to obtain the three-dimensional grid information corresponding to the real-scene image.

[0112] According to one or more embodiments of the present disclosure, for the pixel region corresponding to each object, extracting a plurality of target pixels from the pixel region according to a preset sampling ratio includes: for the pixel region corresponding to each object, if the object is a hand object, extracting a plurality of target pixels from the pixel region according to a first preset sampling ratio; if the object is not a hand object, extracting a plurality of target pixels from the pixel region according to a second preset sampling ratio; wherein, the first preset sampling ratio is greater than the second preset sampling ratio.

[0113] According to one or more embodiments of the present disclosure, wherein intercepting the target image corresponding to the hand object from the real-scene image includes: identifying the hand object in the real-scene image through an artificial neural network model to obtain the identification result corresponding to the hand object; intercepting the target image corresponding to the hand object from the real-scene image according to the identification result corresponding to the hand object.

[0114] According to one or more embodiments of the present disclosure, intercepting the target image corresponding to the hand object from the real-scene image includes: selecting a local scene image including the hand object from the real-scene image, and intercepting the target image corresponding to the hand object from the local scene image; wherein, the display ratio of the local scene image is greater than the display ratio of the real-scene image.

[0115] According to one or more embodiments of the present disclosure, it further includes: determining a plurality of tracking points corresponding to the hand object, performing trajectory tracking on the plurality of tracking points to obtain an operation instruction of the hand object for the virtual object; controlling the virtual object to execute the operation corresponding to the operation instruction.

[0116] According to one or more embodiments of the present disclosure, performing trajectory tracking on the plurality of tracking points to obtain an operation instruction of the hand object for the virtual object includes: performing trajectory tracking on the plurality of tracking points to obtain the position information respectively corresponding to each of the plurality of tracking points within a preset time period, wherein the position information includes the first position information before the tracking point moves and the second position information after the tracking point moves; obtaining the moving gesture of the hand object according to the position information respectively corresponding to each of the plurality of tracking points; determining the operation instruction corresponding to the moving gesture as the operation instruction of the hand object for the virtual object.

[0117] According to one or more embodiments of the present disclosure, it further includes: feathering an edge region of the three-dimensional hand image, where the distance between the edge region and the edge of the three-dimensional hand image is less than a first preset distance; and / or, the three-dimensional hand image includes a hand and an arm, and making a target arm region of the arm transparent, where the distance between the target arm region and the hand is greater than a second preset distance.

[0118] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an image display device, which is applied to a head-mounted device equipped with a camera module. The device includes:

[0119] a receiving module, configured to, in response to receiving a real-scene image including a hand object collected by the camera module, determine whether the hand object overlaps with a virtual screen at a preset spatial position;

[0120] a selection module, configured to, if the hand object overlaps with the virtual screen, select target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object;

[0121] a cropping module, configured to crop a target image corresponding to the hand object from the real-scene image, and combine the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object;

[0122] a display module, configured to display the virtual screen and the three-dimensional hand image on a display interface of the head-mounted device according to the preset spatial position where the virtual screen overlaps and the three-dimensional spatial information corresponding to the hand object.

[0123] According to one or more embodiments of the present disclosure, the display module is configured to determine an occlusion relationship between the virtual screen and the three-dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object; perform rendering on the virtual object and the three-dimensional hand image according to the occlusion relationship between the virtual screen and the three-dimensional hand image to obtain a composite image, and display the composite image on the display interface of the head-mounted device.

[0124] According to one or more embodiments of the present disclosure, the three-dimensional hand image includes three-dimensional coordinates of a plurality of first vertices corresponding to the hand object, and the preset spatial position includes three-dimensional coordinates of a plurality of second vertices corresponding to the virtual screen; the display module is configured to, for each first vertex in the three-dimensional hand image, select a target vertex at the same position from the plurality of second vertices corresponding to the virtual screen, where the same position means that the coordinate values in the X-axis direction and the Y-axis direction are both the same; if the coordinate value of the first vertex in the Z-axis direction is less than the coordinate value of the target vertex in the Z-axis direction, the first vertex in the three-dimensional hand image is blocked by the virtual screen; if the coordinate value of the first vertex in the Z-axis direction is greater than the coordinate value of the target vertex in the Z-axis direction, the virtual screen is blocked by the first vertex in the three-dimensional hand image.

[0125] According to one or more embodiments of the present disclosure, the real-scene image includes pixel regions respectively corresponding to a plurality of objects; correspondingly, the selection module is configured to determine three-dimensional grid information corresponding to the real-scene image, where the three-dimensional grid information includes grid information respectively corresponding to the plurality of pixel regions; and select target grid information corresponding to the target pixel region from the grid information respectively corresponding to the plurality of pixel regions according to the target pixel region corresponding to the hand object.

[0126] According to one or more embodiments of the present disclosure, the selection module is configured to, for the pixel region corresponding to each object, extract a plurality of target pixels from the pixel region according to a preset sampling ratio; determine three-dimensional grid vertices respectively corresponding to the plurality of target pixels, and construct a plurality of adjacent and non-overlapping triangular patches through the three-dimensional grid vertices respectively corresponding to the plurality of target pixels to obtain grid information of the pixel region corresponding to the object; and combine the grid information of the pixel regions respectively corresponding to the plurality of objects to obtain three-dimensional grid information corresponding to the real-scene image.

[0127] According to one or more embodiments of the present disclosure, the selection module is configured to, for the pixel region corresponding to each object, if the object is a hand object, extract a plurality of target pixels from the pixel region according to a first preset sampling ratio; if the object is not a hand object, extract a plurality of target pixels from the pixel region according to a second preset sampling ratio; where the first preset sampling ratio is greater than the second preset sampling ratio.

[0128] According to one or more embodiments of the present disclosure, the interception module is configured to identify the hand object in the real-scene image through an artificial neural network model to obtain an identification result corresponding to the hand object; and intercept a target image corresponding to the hand object from the real-scene image according to the identification result corresponding to the hand object.

[0129] According to one or more embodiments of the present disclosure, the cropping module is configured to select a partial scene image including the hand object from the real scene image, and crop a target image corresponding to the hand object from the partial scene image; wherein, a display ratio of the partial scene image is greater than a display ratio of the real scene image.

[0130] According to one or more embodiments of the present disclosure, the apparatus further includes: an operation module; the operation module is configured to determine a plurality of tracking points corresponding to the hand object, perform trajectory tracking on the plurality of tracking points to obtain an operation instruction of the hand object for the virtual object; and control the virtual object to execute an operation corresponding to the operation instruction.

[0131] According to one or more embodiments of the present disclosure, the operation module is configured to perform trajectory tracking on the plurality of tracking points to obtain position information respectively corresponding to each of the plurality of tracking points within a preset time period, where the position information includes first position information before the tracking point moves and second position information after the tracking point moves; obtain a moving gesture of the hand object according to the position information respectively corresponding to each of the plurality of tracking points; and determine an operation instruction corresponding to the moving gesture as an operation instruction of the hand object for the virtual object.

[0132] According to one or more embodiments of the present disclosure, the apparatus further includes: a processing module; the processing module is configured to perform feathering processing on an edge region of the three-dimensional hand image, where a distance between the edge region and an edge of the three-dimensional hand image is less than a first preset distance; and / or, the three-dimensional hand image includes a hand and an arm, and the processing module is configured to perform transparency processing on a target arm region of the arm, where a distance between the target arm region and the hand is greater than a second preset distance.

[0133] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0134] The memory stores computer-executable instructions;

[0135] The processor executes the computer-executable instructions stored in the memory to implement the image display method as described in the first aspect and various possible designs of the first aspect above.

[0136] Fourthly, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image display method described in the first aspect above and various possible designs of the first aspect.

[0137] Fifthly, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the image display method described in the first aspect above and various possible designs of the first aspect.

[0138] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0139] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0140] Although the subject matter has been described in language specific to structural features and / or methodological act logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. An image display method, characterized in that, Applied to a head-mounted device equipped with a camera module, the method includes: In response to receiving a real-scene image captured by the camera module and including a hand object, determining whether the hand object overlaps with a virtual screen at a preset spatial position; If the hand object overlaps with the virtual screen, selecting target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object; Cropping a target image corresponding to the hand object from the real-scene image and combining the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object; According to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object, displaying the virtual screen and the three-dimensional hand image on the display interface of the head-mounted device.

2. The method according to claim 1, wherein The step of displaying the virtual screen and the three-dimensional hand image on the display interface of the head-mounted device according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object includes: Determining an occlusion relationship between the virtual screen and the three-dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object; Rendering the virtual object and the three-dimensional hand image according to the occlusion relationship between the virtual screen and the three-dimensional hand image to obtain a composite image, and displaying the composite image on the display interface of the head-mounted device.

3. The method according to claim 2, wherein The three-dimensional hand image includes the three-dimensional coordinates of multiple first vertices corresponding to the hand object, and the preset spatial position includes the three-dimensional coordinates of multiple second vertices corresponding to the virtual screen; The step of determining an occlusion relationship between the virtual screen and the three-dimensional hand image according to the preset spatial position corresponding to the virtual screen and the three-dimensional hand image corresponding to the hand object includes: For each first vertex in the three-dimensional hand image, selecting a target vertex at the same position from the multiple second vertices corresponding to the virtual screen, where the same position means that the coordinate values in the X-axis direction and the Y-axis direction are both the same; If the coordinate value of the first vertex in the Z-axis direction is less than the coordinate value of the target vertex in the Z-axis direction, the first vertex in the three-dimensional hand image is occluded by the virtual screen; if the coordinate value of the first vertex in the Z-axis direction is greater than the coordinate value of the target vertex in the Z-axis direction, the virtual screen is occluded by the first vertex in the three-dimensional hand image.

4. The method according to claim 1, characterized in that, The real-scene image includes pixel regions corresponding to multiple objects respectively; Correspondingly, the step of selecting target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image includes: Determining the three-dimensional grid information corresponding to the real-scene image, where the three-dimensional grid information includes grid information corresponding to the multiple pixel regions respectively; Selecting the target grid information corresponding to the target pixel region from the grid information corresponding to the multiple pixel regions respectively according to the target pixel region corresponding to the hand object.

5. The method according to claim 4, wherein Determining the three-dimensional grid information corresponding to the real-scene image includes: For each pixel region corresponding to an object, extracting a plurality of target pixels from the pixel region according to a preset sampling ratio; Determining the three-dimensional grid vertices respectively corresponding to the plurality of target pixels, and constructing a plurality of adjacent and non-overlapping triangular patches through the three-dimensional grid vertices respectively corresponding to the plurality of target pixels, to obtain the grid information of the pixel region corresponding to the object; Combining the grid information of the pixel regions respectively corresponding to the plurality of objects to obtain the three-dimensional grid information corresponding to the real-scene image.

6. The method according to claim 5, wherein The extracting a plurality of target pixels from the pixel region according to a preset sampling ratio for each pixel region corresponding to an object includes: For each pixel region corresponding to an object, if the object is a hand object, extracting a plurality of target pixels from the pixel region according to a first preset sampling ratio; if the object is not a hand object, extracting a plurality of target pixels from the pixel region according to a second preset sampling ratio; Wherein, the first preset sampling ratio is greater than the second preset sampling ratio.

7. The method according to claim 1, wherein Wherein, intercepting the target image corresponding to the hand object from the real-scene image includes: Identifying the hand object in the real-scene image through an artificial neural network model to obtain the identification result corresponding to the hand object; According to the identification result corresponding to the hand object, intercepting the target image corresponding to the hand object from the real-scene image.

8. The method according to claim 7, wherein Intercepting the target image corresponding to the hand object from the real-scene image includes: Selecting a local scene image including the hand object from the real-scene image, and intercepting the target image corresponding to the hand object from the local scene image; wherein, the display ratio of the local scene image is greater than the display ratio of the real-scene image.

9. The method according to claim 1, characterized in that It further includes: Determining a plurality of tracking points corresponding to the hand object, and performing trajectory tracking on the plurality of tracking points to obtain an operation instruction of the hand object for the virtual object; Controlling the virtual object to execute the operation corresponding to the operation instruction.

10. The method according to claim 9, wherein Performing trajectory tracking on the plurality of tracking points to obtain an operation instruction of the hand object for the virtual object includes: Performing trajectory tracking on the plurality of tracking points to obtain the position information respectively corresponding to the plurality of tracking points within a preset time period, wherein the position information includes the first position information before the tracking point moves and the second position information after the tracking point moves; Obtaining the moving gesture of the hand object according to the position information respectively corresponding to the plurality of tracking points; Determining the operation instruction corresponding to the moving gesture as the operation instruction of the hand object for the virtual object.

11. The method according to any one of claims 1 to 10, characterized in that, It further includes: Performing feathering processing on the edge region of the three-dimensional hand image, wherein the distance between the edge region and the edge of the three-dimensional hand image is less than a first preset distance; And / or, The three-dimensional hand image includes a hand and an arm, and performing transparency processing on the target arm region of the arm, wherein the distance between the target arm region and the hand is greater than a second preset distance.

12. An image display device, characterized in that, Applied to a head-mounted device equipped with a camera module, the device includes: A receiving module, configured to determine whether the hand object overlaps with a virtual screen at a preset spatial position in response to receiving a real-scene image collected by the camera module and including the hand object; A selection module, configured to, if the hand object overlaps with the virtual screen, select target grid information corresponding to the hand object from the three-dimensional grid information corresponding to the real-scene image, where the target grid information is used to represent the three-dimensional spatial information corresponding to the hand object; A cropping module, configured to crop a target image corresponding to the hand object from the real-scene image and combine the target image with the target grid information to obtain a three-dimensional hand image corresponding to the hand object; A display module, configured to display the virtual screen and the three-dimensional hand image on the display interface of the head-mounted device according to the preset spatial position where the virtual screen overlaps and the three-dimensional spatial information corresponding to the hand object.

13. An electronic device, characterized in that, Includes: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the image display method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the image display method according to any one of claims 1 to 11 is implemented.

15. A computer program product, characterized in that, Includes a computer program, which when executed by a processor, implements the image display method according to any one of claims 1 to 11.