Method and device for displaying holographic image, and computer-readable storage medium and electronic device

By updating the local user-side image processing model in the holographic image display method to personalize the configuration, the problems of low accuracy and large memory consumption in the prior art are solved, achieving higher accuracy holographic image display and reducing the burden on the user-side.

WO2026001370A1PCT designated stage Publication Date: 2026-01-02BOE TECHNOLOGY GROUP CO LTD +1

Patent Information

Application Number
PCT/CN2025/094336
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2025-05-12
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing holographic image display methods cannot configure personalized image processing models for users, resulting in low accuracy and large models consuming a lot of memory, increasing the burden on the user end.

Method used

By sending a video call request to the remote user terminal, the parameters of the target remote model are determined. Based on these parameters, the original image processing model of the local user terminal is updated to generate a personalized target image processing model. Then, the viewpoint map is determined and image interlacing processing is performed to display the target holographic image.

Benefits of technology

It improves the accuracy of holographic images, reduces the system burden on the user end, and reduces the dependence on large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025094336_02012026_PF_FP_ABST
    Figure CN2025094336_02012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of naked-eye 3D technology, and relates to a method and device for displaying a holographic image, and a computer-readable storage medium and an electronic device. The method comprises: sending a video call request to a second client where a remote user corresponding to a local user is located, and in response to a call answering message corresponding to the video call request that is fed back by the second client, determining target remote model parameters corresponding to the remote user; on the basis of the target remote model parameters, performing parameter updating on an original image processing model in a first client, so as to obtain a target image processing model; on the basis of the target image processing model and remote image data of the remote user, determining a viewpoint image of the remote user; performing image interleaving processing on the viewpoint image, so as to obtain a target holographic image of the remote user, and displaying the target holographic image. The present disclosure improves the accuracy of a target holographic image.
Need to check novelty before this filing date? Find Prior Art

Description

Holographic image display method and device, computer readable storage medium and electronic device

[0001] Cross-reference to Related Applications

[0002] The present disclosure takes the application file with the application number: 202410832423.3, the application date: June 25, 2024, and the invention name: holographic image display method and device, computer readable storage medium, and electronic device as the priority, and the entire content of this Chinese patent application is incorporated by reference herein. TECHNICAL FIELD

[0003] Embodiments of the present disclosure relate to the field of naked eye 3D technology, in particular, to a holographic image display method, a holographic image display device, a computer readable storage medium, and an electronic device. BACKGROUND

[0004] In the existing holographic image display method, the holographic effect is realized based on a large language model. However, this method has the following defects: on the one hand, it cannot configure a personalized image processing model for the user, thereby reducing the accuracy of the obtained holographic image; on the other hand, the large model occupies a large amount of memory, thereby increasing the burden on the user end.

[0005] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0006] The purpose of the present disclosure is to provide a holographic image display method, a holographic image display device, a computer readable storage medium, and an electronic device, thereby at least partially overcoming the problem of low accuracy of holographic images and heavy burden on the user end due to the limitations and defects of related technologies.

[0007] According to one aspect of the present disclosure, a holographic image display method is provided, configured in a first user end where a local user is located, the holographic image display method comprising:

[0008] sending a video call request to a second user end where a remote user corresponding to the local user is located, determining a target remote model parameter corresponding to the remote user in response to a call answering message corresponding to the video call request fed back by the second user end;

[0009] updating the original image processing model in the first user end based on the target remote model parameter to obtain a target image processing model;

[0010] determine a view point map of the remote user according to the target image processing model and remote image data of the remote user;

[0011] perform image interweaving processing on the view point map to obtain a target holographic image of the remote user, and display the target holographic image.

[0012] In an exemplary embodiment of the present disclosure, determining a target remote model parameter corresponding to the remote user comprises:

[0013] querying a first remote model parameter corresponding to the remote user in the first user terminal according to a remote user identifier of the remote user, and querying a second remote model parameter corresponding to the remote user in the second user terminal;

[0014] comparing a first parameter version number of the first remote model parameter and a second parameter version number of the second remote model parameter, and selecting a target remote model parameter from the first remote model parameter and the second remote model parameter based on a version number comparison result.

[0015] In an exemplary embodiment of the present disclosure, performing parameter updating on an original image processing model in the first user terminal based on the target remote model parameter to obtain a target image processing model comprises:

[0016] calling the original image processing model in the first user terminal, and extracting a local model parameter in the original image processing model;

[0017] replacing the local model parameter based on the target remote model parameter to obtain a target image processing model associated with the remote user.

[0018] In an exemplary embodiment of the present disclosure, the target image processing model comprises a depth map calculation model, a depth fusion module, a projection transformation module, and an image rectification model;

[0019] The depth map calculation model is configured to determine depth image data of the remote image data, and the depth fusion module is configured to fuse the depth image data to obtain fused depth image data.

[0020] The depth map calculation model is configured to determine depth image data of the remote image data, and the depth fusion module is configured to fuse the depth image data to obtain fused depth image data.

[0021] projecting the fused depth image to a new view point pixel plane corresponding to the local user based on the projection transformation module to obtain a depth map projection result;

[0022] projecting the depth map image data to obtain a viewpoint image of the remote user.

[0023] In an example embodiment of the present disclosure, the depth map determination model comprises a feature encoder, a context encoder, a correlation pyramid, and a disparity update module.

[0024] The depth map determination model is obtained by the following manners:

[0025] The feature encoder is shared with the context encoder to obtain the depth map determination model.

[0026] The feature encoder is shared with the context encoder to obtain the depth map determination model.

[0027] The feature encoder is shared with the context encoder to obtain the depth map determination model.

[0028] The feature encoder is shared with the context encoder to obtain the depth map determination model.

[0029] In an example embodiment of the present disclosure, the depth map determination model is obtained by the following manners:

[0030] The feature encoder is shared with the context encoder to obtain the depth map determination model.

[0031] The feature encoder is shared with the context encoder to obtain the depth map determination model.

[0032] The feature encoder is shared with the context encoder to obtain the depth map determination model.

[0033] In an example embodiment of the present disclosure, the image rectification model comprises a plurality of first convolution layers, a plurality of first down-sampling layers, and a plurality of first up-sampling layers, and the number of layers of the plurality of first down-sampling layers is the same as that of the plurality of first up-sampling layers.

[0034] The image correction model is used to correct the depth map projection result, to obtain a viewpoint image of the remote user, including:

[0035] The first multi-layer convolution layer is used to perform convolution processing on the depth map projection result, to obtain local features of different scales;

[0036] The first multi-layer down-sampling layer is used to perform down-sampling on the local features of different scales, to obtain region features of local regions;

[0037] The first multi-layer up-sampling layer and the region features of the local regions are used to correct the depth map projection result, to obtain the viewpoint image of the remote user.

[0038] In an exemplary embodiment of the present disclosure, the viewpoint image is subjected to image interleaving processing, to obtain a target holographic image of the remote user, including:

[0039] The interpolation range of the viewpoint image is calculated based on a vertex shader, and the interpolation content of the viewpoint image is calculated based on a fragment shader;

[0040] The viewpoint image is subjected to image interleaving processing based on the interpolation range and the interpolation content, to obtain the target holographic image of the remote user.

[0041] In an exemplary embodiment of the present disclosure, the target holographic image is displayed, including:

[0042] The first user eye height of the local user and the second user eye height of the remote user are determined;

[0043] The display image height of the target holographic image at the first user end is adjusted according to the first user eye height and the second user eye height, and the target holographic image after height adjustment is displayed.

[0044] In an exemplary embodiment of the present disclosure, the first user eye height of the local user and the second user eye height of the remote user are determined, including:

[0045] The first user image of the local user is captured based on a first camera group associated with the first user end, and the first left and right eye key point coordinates of the local user are determined based on the first user image;

[0046] The first user eye height of the local user on the first user end is determined according to the first left and right eye key point coordinates;

[0047] determining second pixel coordinates of the second left and right eyes of the remote user on the target holographic image according to the target holographic image, and determining a second user eye height of the remote user on the first user terminal according to the second pixel coordinates.

[0048] In an example embodiment of the present disclosure, the display image height of the target holographic image on the first user terminal is adjusted according to the first user eye height and the second user eye height, including:

[0049] According to the first user eye height and the second user eye height, an image height adjustment value of the target holographic image is determined, and the display image height of the target holographic image on the first user terminal is adjusted according to the image height adjustment value.

[0050] In an example embodiment of the present disclosure, the original image processing model is obtained by the following way:

[0051] A first historical user image of the local user is obtained, and the first historical user image is preprocessed to obtain a first target user image;

[0052] The first target user image is input into the network model to be trained to obtain a first prediction result, and a loss function is constructed according to a viewpoint label image corresponding to the first target user image and the first prediction result;

[0053] Based on the loss function, the parameters in the network model to be trained are adjusted to obtain the original image processing model.

[0054] In an example embodiment of the present disclosure, the first historical user image is preprocessed to obtain a first target user image, including:

[0055] Based on a preset face recognition algorithm, the first historical user image is face-recognized to obtain a first historical face image corresponding to the local user;

[0056] The first historical face image is preprocessed to obtain a first target user image; wherein the preprocessing manner includes at least one of adding noise, color conversion, clipping and lifting.

[0057] In an example embodiment of the present disclosure, the original image model is updated by the following way:

[0058] In response to the feedback result of the picture display quality corresponding to the target holographic image fed back by the second user terminal in response to the video call end instruction, it is determined whether the original image processing model needs to be optimized.

[0059] In a case where it is determined that model optimization processing is required for the original image processing model, local image data of a local user is collected based on the first camera group, and current local model parameters of the original image model are optimized based on the local image data, to obtain an original image model after optimization processing.

[0060] In an exemplary embodiment of the present disclosure, the original image model is also updated in parameters in the following manner:

[0061] A self-learning update frequency of the original image model is determined, and current local model parameters of the original image model are optimized based on the self-learning update frequency and local image data of a local user collected by the first camera group, to obtain an original image model after optimization processing.

[0062] According to an aspect of the present disclosure, there is provided a holographic image display device configured in a first user terminal where a local user is located, the holographic image display device comprising:

[0063] A remote model parameter determination module configured to send a video call request to a second user terminal where a remote user corresponding to the local user is located, and determine target remote model parameters corresponding to the remote user in response to a call answering message corresponding to the video call request fed back by the second user terminal.

[0064] A model parameter update module configured to update parameters of an original image processing model in the first user terminal based on the target remote model parameters, to obtain a target image processing model.

[0065] A view point map determination module configured to determine a view point map of the remote user according to the target image processing model and remote image data of the remote user.

[0066] A holographic image display module configured to perform image interleaving processing on the view point map, to obtain a target holographic image of the remote user, and display the target holographic image.

[0067] According to an aspect of the present disclosure, there is provided a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the holographic image display method of any one of the above aspects.

[0068] According to an aspect of the present disclosure, there is provided an electronic device comprising:

[0069] a processor; and

[0070] a memory configured to store executable instructions of the processor.

[0071] The processor is configured to execute the display method of the holographic image via executing the executable instructions.

[0072] The display method of the holographic image provided by the embodiments of the present disclosure, on one hand, by sending a video call request to a second user terminal where a remote user corresponding to a local user is located, responding to a call answering message corresponding to the video call request fed back by the second user terminal, determining a target remote model parameter corresponding to the remote user; then updating a parameter of an original image processing model in the first user terminal based on the target remote model parameter to obtain a target image processing model; further determining a viewpoint graph of the remote user according to the target image processing model and remote image data of the remote user; finally performing image interleaving processing on the viewpoint graph to obtain a target holographic image of the remote user, and displaying the target holographic image, realizing that the target image processing model personalized for the remote user is configured based on the target remote model parameter, and the accuracy of the obtained target holographic image is improved; on the other hand, since the target image processing model personalized for the remote user can be configured based on the target remote model parameter, a large model does not need to be configured in the first user terminal, thereby solving the problem that in the prior art, the memory occupied by the large model is large, and thereby the burden on the user terminal is large, and the system burden of the first user is reduced.

[0073] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0074] The drawings incorporated into the specification and forming a part of the specification, show embodiments consistent with the present disclosure, and together with the specification, serve to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.

[0075] FIG. 1 schematically shows a flowchart of a display method of a holographic image according to an example embodiment of the present disclosure.

[0076] FIG. 2 schematically shows an example diagram of a holographic remote video communication system according to an example embodiment of the present disclosure.

[0077] FIG. 3 schematically shows a structural example diagram of a user terminal according to an example embodiment of the present disclosure.

[0078] FIG. 4 schematically shows a structural example diagram of an original image processing model according to an example embodiment of the present disclosure.

[0079] FIG. 5 schematically shows a structural example diagram of a depth map calculation model according to an example embodiment of the present disclosure.

[0080] FIG. 6 schematically shows a structural example diagram of an image correction model according to an example embodiment of the present disclosure.

[0081] FIG. 7 schematically shows a flow example diagram of a training method of an original image processing model according to an example embodiment of the present disclosure.

[0082] FIG. 8 schematically shows a scene example diagram of image processing by an original image processing model according to an example embodiment of the present disclosure.

[0083] FIG. 9 schematically shows a method flow diagram of a specific determination process of a viewpoint map according to an example embodiment of the present disclosure.

[0084] FIG. 10 schematically shows a scene example diagram of a specific calculation process of depth image data according to an example embodiment of the present disclosure.

[0085] FIG. 11 schematically shows an example diagram of a specific distribution of a machine group according to an example embodiment of the present disclosure.

[0086] FIG. 12 schematically shows a scene example diagram of an image correction process according to an example embodiment of the present disclosure.

[0087] FIG. 13 schematically shows a block diagram of a display device of a holographic image according to an example embodiment of the present disclosure.

[0088] FIG. 14 schematically shows an example diagram of an electronic device for implementing a display method of a holographic image according to an example embodiment of the present disclosure. DETAILED DESCRIPTION

[0089] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the disclosure. One skilled in the relevant art will recognize, however, that the techniques of the disclosure can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures have not been described in detail so as not to obscure the aspects of the disclosure.

[0090] Furthermore, the accompanying drawings are only schematic and are non-limiting. Identical, corresponding or similar elements present in several figures are designated with the same reference numeral, and will not be repeatedly presented in the description of the figures. Some of the blocks in the drawings are functional blocks, which can be implemented in software, hardware or a combination thereof. The software can be stored in a memory and executed on a processor.

[0091] With the continuous development of 3D display technology and the continuous pursuit of a better life, naked eye 3D display has become a popular product in the display field; among them, the basic principle of naked eye 3D display is: real-time rendering of the pixel position (viewpoint) of the left and right eyes according to the three-dimensional coordinates of the local user's eyes in the screen coordinate system, and displaying the left and right eye viewpoint images on the naked eye 3D screen at the same time through the mapping algorithm; under this premise, the left and right eyes of the user will receive these two images, and then the human brain synthesizes a 3D scene to give the user a sense of being there.

[0092] Relying on 3D display technology and naked eye 3D display screen, holographic remote video communication system emerges as the times require; traditional video communication system can realize real-time video call between two people in different geographical locations, while holographic video communication system expands 2D video communication to 3D with the help of naked eye 3D display, so that users can see a stereoscopic display screen, thus having a more immersive and realistic experience. At present, in order to ensure the robustness and stability of the algorithm, the core algorithm in the system is generally implemented based on deep learning algorithm; usually, excellent and general deep learning algorithm is a large model trained on a large amount of data; however, due to the characteristics of complex network structure and numerous parameters of the large model, the large model has a high time consumption in actual use, which cannot meet the real-time requirement of holographic remote video communication; however, although network compression operations such as pruning, distillation and reducing resolution can improve the real-time performance of the algorithm, the performance such as image quality and algorithm generality is also reduced; therefore, how to make the holographic remote video communication system balance the display effect and time consumption under the condition of fixed hardware (GPU, CPU, etc.) has become a problem to be solved; that is, how to balance the algorithm delay and algorithm performance has become a difficulty of holographic remote video communication system.

[0093] Based on this, the example embodiments of the present disclosure provide a holographic image display method, which is a holographic remote video communication system scheme with self-learning capability; in the actual application process, the example embodiments of the present disclosure add a self-learning module to the holographic remote video communication system, train and evaluate the core algorithm module in the system based on the image data of the local end user accumulated in the video communication process, optimize the model parameters of the specified local end user, and through the introduction of the self-learning module, the generality requirement of the holographic remote video communication system for the algorithm model is reduced (from being applicable to all users to being applicable to specified users, that is, the problem complexity is reduced), so that a network model with fewer parameters and lower operation amount can be used to realize, thereby improving the system display frame rate and improving the user experience effect on the premise of ensuring the display effect.

[0094] In an example embodiment, the holographic image display method described in the example embodiments of the present disclosure can be run on a first user end where a local user is located. The first user end can be a display with naked-eye 3D display effect or an all-in-one machine, etc. The present example does not make special limitations on this. Of course, a person skilled in the art can also run the method of the present disclosure on other platforms according to needs, and the present example embodiment does not make special limitations on this.

[0095] In a possible example embodiment, the holographic image display method described in the example embodiments of the present disclosure can also be run on a server or a cloud server, etc. Wherein, the holographic image display method running on the server or the cloud server means that the determination process of the target image processing model, the determination process of the viewpoint graph and the generation process of the target holographic image are implemented in the server or the cloud server. Further, when the holographic image display method is run on the server or the cloud server, the image processing model can be replaced with a large language model for implementation, or the image processing model described in the example embodiments of the present disclosure can also be directly used, and the present example does not make special limitations on this.

[0096] Specifically, referring to FIG. 1, the holographic image display method can include the following steps:

[0097] Step S110. Sending a video call request to a second user end where a remote user corresponding to the local user is located, determining a target remote model parameter corresponding to the remote user in response to a call answering message corresponding to the video call request fed back by the second user end;

[0098] Step S120. Parameter updating the original image processing model in the first user end based on the target remote model parameter to obtain a target image processing model;

[0099] Step S130. Determine a viewpoint map of the remote user according to the target image processing model and remote image data of the remote user.

[0100] Step S140. Perform image interleaving processing on the viewpoint map to obtain a target holographic image of the remote user, and display the target holographic image.

[0101] In the holographic image display method described above, on the one hand, by sending a video call request to the second user terminal where the remote user corresponding to the local user is located, responding to the call answering message corresponding to the video call request fed back by the second user terminal, the target remote model parameter corresponding to the remote user is determined; then the original image processing model in the first user terminal is updated based on the target remote model parameter to obtain a target image processing model; then the viewpoint map of the remote user is determined according to the target image processing model and the remote image data of the remote user; finally, the target holographic image of the remote user is obtained by performing image interleaving processing on the viewpoint map, and the target holographic image is displayed, which realizes the configuration of the personalized target image processing model for the remote user based on the target remote model parameter, and improves the accuracy of the obtained target holographic image; on the other hand, since the personalized target image processing model can be configured for the remote user based on the target remote model parameter, there is no need to configure a large model in the first user terminal, thereby solving the problem that the memory occupied by the large model is large in the prior art, and the burden on the user terminal is large, thereby reducing the system burden of the first user.

[0102] In the following, the holographic image display method described in the example embodiments of the present disclosure will be explained and described in detail in combination with the accompanying drawings.

[0103] Firstly, the holographic remote video communication system involved in the example embodiments of the present disclosure is explained and described. Specifically, referring to FIG. 2, the holographic remote video communication system can include a first user terminal 210, a server 220 and a second user terminal 230; wherein the first user terminal and the second user terminal can be in communication connection with the server through wired network or wireless network; in the actual application process, the first user terminal can be referred to as a local user terminal, and the second user terminal can be referred to as a remote user terminal.

[0104] Secondly, referring to FIG. 3, each user terminal can include a naked-eye 3D display 301, a plurality of color cameras or depth cameras 302 and a host computer and other hardware devices; at the same time, the software modules mainly include a network transmission module 303, a core algorithm module 304, an image acquisition module 305, a stereoscopic display module 306 and a self-learning module 307, etc.

[0105] In practical application, the holographic remote video communication system can include the following four states:

[0106] The first state is a static state, that is, the system is currently in a stop running state and does not perform any processing and operation; the second state is an initialization state; the third state is a communication state; and the fourth state is a self-learning state.

[0107] The initialization state can include two parts, namely system initialization and user initialization. For system initialization, when the system is started for the first time, the general parameters of the current algorithm model (the general parameter version label is 0) are obtained in the cloud. The algorithm model parameters are obtained by pre-training on general data sets, and can be used by any user to meet the real-time requirements of the system, but the quality of the generated new view image is poor. For user initialization, when a user uses the system for the first time, the system generates a unique ID for the user, and uses the general model parameters as the initial model parameters of the user. In addition, the system prompts the user to upload local user image data or the user to perform a specified action (hand waving, clapping, or fist clenching, etc. common hand movements in daily communication, and head movements such as turning left, right, up, and down, blinking, smiling, and other facial expression movements) to collect a video data using the system camera for preliminary optimization training of the general model parameters. At the same time, the user can choose whether to start the self-learning function of the system (starting means that the user allows the system to collect user pictures for model parameter updating during video communication) and the selection of the algorithm model parameter update frequency.

[0108] The communication state is the video call state corresponding to the example embodiments of the present disclosure. The self-learning state is the state of self-learning of the original image processing model. In the communication state, the local system collects camera data (color image and depth image) in real time and transmits it to the remote end during the video call process of the first user end and the second user end, and receives the camera data transmitted from the remote end. Further, the local core algorithm module generates the left and right eye image data of the current user according to the received camera data and the position of the user's two eyes, and transmits the left and right eye image data to the stereoscopic display module for image arrangement processing, and finally displays the left and right eye image data on the naked eye 3D display. The user's two eyes receive the left and right eye image data displayed on the screen, and the brain synthesizes the left and right eye image data to form a holographic stereoscopic display effect. At the same time, according to the user's selection of whether to start the system self-learning, data collection (saving the local collected camera data at a certain frequency, and a too high frequency can increase the algorithm delay) can be performed.

[0109] It needs to be added here that, generally, within a certain range, the network model complexity of the deep learning algorithm is proportional to its expression ability (problem solving ability); the example embodiments of the present disclosure utilize the self-learning module to optimize the model parameters for a specified user, so the performance requirements of the system for the network model are reduced from needing to adapt to most users (possibly in the order of ten thousand) to achieving optimality on the data set of the specified user (a single person), the difficulty of the problem that the network model needs to solve is greatly reduced, and the complexity and parameter amount of the network model can be reduced; and for all users in the communication system, the algorithm model used by the system is the same (the deep learning network model structure is the same), but the parameters in the model are personalized, and the system has the ability to continuously optimize and improve the effect, bringing better experience effect to the user.

[0110] In the following, the original image processing model involved in the example embodiments of the present disclosure will be explained and described. Specifically, referring to FIG. 4, the original image processing model can include an input layer 401, a depth map calculation model 402, a depth fusion module 403, a projection transformation module 404, an image rectification model 405, and an output layer 406; where the specific application of each model or module in the image processing process is described below, and will not be repeated here.

[0111] In an example embodiment, referring to FIG. 5, the depth map calculation model described herein can include a feature encoder 501, a context encoder 502, a correlation pyramid 503, and a disparity update module 504; where the specific application of each module in the image processing process is described below, and will not be repeated here.

[0112] In an example embodiment, referring to FIG. 6, the image rectification model described herein can include a plurality of first convolution layers 601, a plurality of first down-sampling layers 602, and a plurality of first up-sampling layers 603, the number of layers of the plurality of first down-sampling layers and the plurality of first up-sampling layers is the same; where the specific application of each module in the image processing process is described below, and will not be repeated here.

[0113] In the following, the specific training process of the original image processing model will be explained and described. Specifically, referring to FIG. 7, the specific training process of the original image processing model can include the following steps:

[0114] Step S710, obtaining a first historical user image of a local user, and pre-processing the first historical user image to obtain a first target user image;

[0115] Step S720, input the first target user image into the network model to be trained to obtain a first prediction result, and construct a loss function according to the viewpoint label image corresponding to the first target user image and the first prediction result.

[0116] Step S730, adjust the parameters in the network model to be trained based on the loss function to obtain the original image processing model.

[0117] In an exemplary embodiment, the first historical user image is preprocessed to obtain the first target user image, which can be achieved by the following manner: performing face recognition on the first historical user image based on a preset face recognition algorithm to obtain a first historical face image corresponding to the local user; preprocessing the first historical face image to obtain the first target user image; wherein the preprocessing manner includes at least one of adding noise, color conversion, cropping, and lifting.

[0118] In the following, the specific training process of the original image processing model will be further explained and described. Specifically, in the actual application process, the training process of the original image processing model can be divided into the following stages: data preparation stage, training strategy determination stage, model evaluation stage, learning pause and recovery stage, and training stop stage. Among them:

[0119] For the data preparation stage, first, the data needs to be preprocessed, that is, to filter out effective training data; wherein the filtering of effective training data can be assisted by the existing face detection and face recognition algorithm to recognize the stored user data (i.e. the first historical user image), eliminate image data not containing the local user and image data containing multiple people, and only keep single-person data containing the specified user (i.e. the local user) for training; at the same time, the specific judgment process of whether it is the local user can be based on the comparison result of the image and the image data submitted at the system initialization; if it is the same person, it is considered that the image is the image data containing the user; secondly, if the image data uploaded by the user is less at the initialization, the system needs to expand the data set by means of increasing noise, color conversion, cropping, lifting, etc. to avoid the phenomenon of model parameter overfitting due to too small data set in the model training and optimization process;

[0120] For the training strategy determination stage, the new view generation algorithm model can be iteratively optimized in an overall training manner; specifically, in the training process, each time iteration, a group of camera data (i.e. the first target user image) is randomly extracted as the input data of the model, and a color image (i.e. the view label image corresponding to the first target user image) is randomly extracted from it as the supervision of the output image (the view of the image is the new view position in the network model input), the MSE (mean square error) between the output image of the network model and the supervision image is calculated as the loss function Loss of the network model for back propagation, so that the parameters in the network model are updated;

[0121] For the model evaluation stage, the self-learning module takes PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) as the evaluation indicators of the algorithm model output by the self-learning module, to compare the similarity between the color image of the expected view and the new view image generated by the trained algorithm. Only when both indicators are improved, the current self-learning process is considered effective, and the version number is incremented by 1 as the version number of the new model parameters, and the PSNR and SSIM values are marked. For example, assuming that the expected new view image (view label image) is I, and the algorithm-generated new view image (first prediction result) is K; the specific calculation formulas of the mean square error MSE, PSNR, and SSIM between the two images can be shown in the following formulas (1), (2), and (3):

[0122] where W and H are the height and width of the image, I(i,j) is the pixel value of the pixel point at position i,j in the view label image; K(i,j) is the pixel value of the pixel point at position i,j in the first prediction result; μ and σ are the mean and variance of the image, σ IK represent the covariance of the two images; MAX is the maximum pixel value of the pixel points in the view label image; c1 and c2 are constant terms;

[0123] For the learning pause and resume stage, deep learning training usually requires a long training time. If the user starts the communication function during the self-learning process, the self-learning module pauses the training and saves the training information (gradient, learning rate, etc.) and model parameters at this time. After the user's communication is completed, the training information and model parameters are read to resume the training.

[0124] For the training stop phase, when the evaluation indicators PSNR and SSIM no longer improve or the network model's Loss no longer decreases or reaches the maximum iteration number, the training of the self-learning module is stopped; if the evaluation indicators are better than the old version stored locally at this time, the algorithm model parameters stored locally are updated, the version number is incremented by 1, and the PSNR value and SSIM value at this time are marked.

[0125] It needs to be supplemented here that, for the consideration of the hardware computing power of the local end (i.e., the first user end), the local end system only collects and trains the data of the local end user, and the local end only stores the model parameters of the remote end user and does not save the image data of the opposite party.

[0126] In the following, the display method of the holographic image shown in FIG. 1 will be explained and described in detail in combination with FIGS. 2-7. Specifically:

[0127] In step S110, a video call request is sent to a second user end where a remote user corresponding to the local user is located, a call answering message corresponding to the video call request is fed back in response to the second user end, and a target remote model parameter corresponding to the remote user is determined.

[0128] Among them, the specific determination process of the target remote model parameter recorded here can be realized in the following way: the first remote model parameter corresponding to the remote user is queried in the first user end according to the remote user identifier of the remote user, and the second remote model parameter corresponding to the remote user is queried in the second user end; the first parameter version number of the first remote model parameter and the second parameter version number of the second remote model parameter are compared, and the target remote model parameter is selected from the first remote model parameter and the second remote model parameter based on the version number comparison result.

[0129] In the following, the specific establishment of the call request and the specific determination process of the target remote model parameter will be further explained and described. Specifically, in the actual application process, first, if a remote video call is to be implemented, the communication between the first user terminal and the second user terminal needs to be established first; that is, the first user terminal where the local user is located and the second user terminal where the remote user is located need to be in the state of holographic remote video communication; wherein the holographic remote video communication needs to be realized by the following way: first, the first user terminal and the second user terminal need to log in the system; for example, the local user logs in the system on the first user terminal system by using the local user identification number and the local user password, and the system obtains the latest version parameter of the local user stored in the local system; the remote user logs in the system on the second user terminal system by using the remote user identification and the remote user password, and the system obtains the latest version parameter of the remote user stored in the local system of the second user terminal; secondly, the first user terminal sends a video call request to the second user terminal where the remote user corresponding to the local user is located, and when the second user terminal receives the video call request, the remote user on the second user terminal answers to connect; further, the two systems will respectively query the algorithm model parameter version and the performance label corresponding to the user identification of the user on the opposite end stored in the local, that is, the first user terminal queries whether the remote user stores the algorithm model parameter in the local of the first user terminal, if yes, what is the parameter version and the parameter performance label (the parameter performance label can be generated according to PSNR and SSIM, the smaller the PSNR and SSIM are, the higher the level of the performance label is); the second user terminal queries whether the algorithm model is stored in the local of the second user terminal, if yes, what is the version and the performance label; further, the two systems respectively query whether the algorithm model parameter of the user on the opposite end exists in the opposite end system, that is, the first user terminal queries whether the algorithm model parameter with the updated version label and the stronger performance label exists in the local of the second user terminal, if yes, the first user terminal will obtain the algorithm model parameter, and obtain the target remote model parameter based on the algorithm model parameter; otherwise, no processing is performed; the B terminal also performs the same operation; when the two systems are queried and updated, the communication connection is completed.

[0130] In step S120, the original image processing model in the first user terminal is updated based on the target remote model parameter to obtain a target image processing model.

[0131] Specifically, the specific implementation process of the parameter update can be achieved by the following manner: calling the original image processing model in the first user terminal, and extracting the local model parameters in the original image processing model; replacing the local model parameters based on the target remote model parameters to obtain a target image processing model associated with the remote user. That is, in the actual application process, the local model parameters in the original image processing model can be directly replaced based on the target remote model parameters, so as to obtain the target image processing model corresponding to the remote user and having individualization. It should be noted that the image processing model arranged in each different user terminal has the same specific model structure and model depth, and the specific model parameter setting is also the same; the only different part is that the specific parameter value is set by each different user; therefore, the individualization of the model can be realized by exchanging the model parameters, so as to achieve the purpose of improving the accuracy of the obtained holographic image.

[0132] In step S130, a view point map of the remote user is determined according to the target image processing model and the remote image data of the remote user.

[0133] Specifically, in the holographic remote video communication system, the core algorithm (i.e., the image processing model) generates the images of the positions of the left and right eyes of the user, and displays them on the 3D screen through the layout algorithm, so that the user can watch the holographic (i.e., 3D) display effect. Therefore, the core algorithm, i.e., the generation algorithm of the new view point map, can refer to the specific algorithm scene implementation example diagram shown in FIG. 8. The input data of the algorithm is the color image and the depth image (the depth image is an optional input) collected by the camera of the remote terminal and the current and new view point positions, and the output is the new view point color image. The view point map recorded herein can include a two-view point map, or a nine-view point map, a 16-view point map, a 49-view point map, a 64-view point map, etc., and the present example does not specially limit this.

[0134] Further, referring to FIG. 9, the specific determination process of the view point map can include the following steps:

[0135] In step S910, the depth image data of the remote image data is determined according to the depth map calculation model, and the depth image data is fused according to the depth fusion module to obtain a fused depth image.

[0136] In the example embodiment, first, the depth image data of the remote image data is determined according to the depth map calculation model; specifically, it can be realized in the following way: the left eye feature map of the remote left eye image in the remote image data and the right eye feature map of the remote right eye image in the remote image data are determined based on the feature encoder, and the context feature of the remote left eye image is determined based on the context encoder; the correlation volume between the left eye feature map and the right eye feature map is calculated, and a multi-layer correlation pyramid is constructed based on the left eye feature map, the right eye feature map and the correlation volume; the correlation volume associated with the preset disparity parameter is indexed from the correlation pyramid according to the preset disparity parameter, and the correlation feature is constructed according to the correlation volume associated with the preset disparity parameter and the left eye feature map and the right eye feature map corresponding to the correlation volume; the correlation feature, the context feature and the preset disparity parameter are input into the disparity update module to obtain the disparity field of the remote left eye image, and the depth image data is determined according to the disparity field. Specifically, the depth map calculation described herein, that is, the depth map (invalid points in the depth map are marked as +∞) at one or more viewpoints is calculated from multiple color images, and a typical algorithm can be RaftStereo; in the depth map calculation model described in the example embodiment, in order to further reduce the system burden of the user end, the RaftStereo-tiny algorithm based on RaftStereo compression can be used to realize high-precision and fast depth map estimation; wherein the RaftStereo described herein is a classical depth learning-based algorithm model for calculating depth map from binocular color images, but the model is complex and has high parameter quantity, resulting in poor real-time performance, and based on the self-learning module described in the example embodiment of the present disclosure, the network model can be compressed under the premise of ensuring the quality of remote communication. The specific calculation process of the depth image data is shown in FIG. 10.

[0137] In the scene example diagram shown in FIG. 10, the feature encoder Feature Encoder can be applied to the left and right pictures, and each picture is mapped into a dense feature map, which is used to establish a correlation cost volume; that is, the left eye feature map of the remote left eye image in the remote image data and the right eye feature map of the remote right eye image in the remote image data can be determined based on the feature encoder; wherein the network used by the feature encoder can include a series of residual blocks and down-sampling layers; at the same time, the generated feature map is 1 / 4 or 1 / 8 of the input 256 channel image resolution, which depends on the number of down-sampling layers used in the experiment; and in the feature encoder Feature Encoder, instance normalization (Instance Normalization) is used; further, the context encoder Context Encoder described above can have the same model structure as the feature encoder, only the instance normalization (Instance Normalization) needs to be converted into batch normalization (Batch Normalization); at the same time, the context encoder only acts on the remote left eye image; its specific role is to initialize the hidden state of the update operator and inject into the disparity update module in each iteration of the update operator; further, in the correlation pyramid, first, a correlation volume Correlation Volume needs to be constructed, which can be determined based on the dot product between the feature vectors of the pixel points; then, the correlation pyramid is constructed based on the correlation volume; wherein in the construction process of the correlation pyramid, the k-th layer of the pyramid is obtained from the k-th layer of the correlation volume using 1-dimensional average pooling; in the process of average pooling, the size of the convolution kernel used can be 2, and the step can also be 2, through the average pooling processing, a new correlation volume can be obtained; and in the correlation pyramid, each layer of the pyramid includes a growing receptive field; at the same time, only the last dimension is pooled, which maintains the high-resolution representation of the information on the original picture, which makes it possible to recover very fine structures; finally, in the disparity update module, first, an estimate of the disparity d can be given, and a 1-dimensional grid of integer offsets is constructed around the current disparity estimate value, which is used to index from each layer in the correlation pyramid; then, when each correlation volume is indexed, linear interpolation is used to connect the recovered values into a feature map; and in the process of disparity update, a series of disparity fields can be predicted from a starting point (d=0), so that each pixel in the left eye view has a horizontal displacement; finally, based on the disparity field, the depth image data can be obtained.

[0138] In an example embodiment, the depth map determination model described above is obtained by sharing the first weight values of the feature encoder and the weight values of the context encoder in the binocular depth estimation model to obtain the depth map determination model, or compressing the first number of convolution kernels of the feature encoder and / or the second number of convolution kernels of the context encoder to obtain the depth map determination model, or compressing the first number of convolution layers of the feature encoder and / or the second number of convolution layers of the context encoder to obtain the depth map determination model. That is, the depth map calculation model described in the example embodiment of the present disclosure can compress the model parameter amount and operation amount in two ways of sharing weights and compressing parameters respectively: in the first compression method, the feature encoder and the context encoder share weights; specifically, in the actual application process, the network structure and parameters of the context encoder can be reused as part of the structure of the feature encoder, and a convolution layer is added based on the output of the context encoder to form the feature encoder; at the same time, the feature encoder and the context encoder share weights, which reduces the parameter amount and operation amount of the context encoder, thereby realizing the compression of the network model; in the second compression method, the parameter compression mainly embodies the compression of the number of convolution kernels and the number of convolution layers in the example embodiment of the present disclosure; the implementation principle of the number of convolution kernels is that the feature encoder and the context encoder are both composed of convolution layers, and the operation amount of each convolution layer is related to the size and number of the convolution kernel of the current layer, and the number of convolution kernels is reduced by half in the patent; the implementation principle of the number of convolution layers is that the feature encoder, the context encoder and the update module are composed of ordinary convolutional neural networks and recurrent neural networks, and the number of network layers is gradually reduced in the experiment to obtain a critical value that keeps the algorithm performance basically unchanged.

[0139] It should be noted that in the example embodiment of the present disclosure, each user terminal can correspond to one or more camera groups when setting the camera groups; the specific distribution of the camera groups can refer to FIG. 11; for example, the camera group shown in FIG. 11 can include four cameras, which can be divided into three groups of double camera systems, i.e., camera 1 and camera 2, camera 2 and camera 3, and camera 3 and camera 4; the three groups of double cameras can be connected in parallel through the RaftStereo-tiny structure to generate depth maps under the viewpoints of camera 2, camera 3 and camera 4; it should be noted that the specific number of cameras in the camera group can be determined according to actual needs, and this is only an example.

[0140] Further, after obtaining the depth image data, depth image data fusion is also needed; at the same time, the depth image data fusion is needed because the system can include multiple camera groups, so there will be multiple depth maps, and therefore depth map fusion is needed; in the process of depth map fusion, the depth map can be fused by projection to obtain the depth map under one or more specified viewpoints.

[0141] In step S920, the fusion depth image is projected to a new viewpoint pixel plane corresponding to the local user based on the projection transformation module to obtain a depth map projection result.

[0142] Specifically, the projection transformation and depth fusion module calculation process described herein can project the viewpoint with depth map and color to the new viewpoint pixel plane (resolution HxW) expected by the algorithm output, and the pixel value after projection is the pixel value (i.e. RGB value) of the color map; and one or more viewpoints can also be projected, and the image formed after projection is spliced in the channel dimension to form data with a dimension of HxWx3N (i.e. depth map projection result).

[0143] In step S930, the depth map projection result is image corrected based on the image correction model to obtain the viewpoint map of the remote user.

[0144] Specifically, the specific implementation process of image correction can be as follows: the depth map projection result is convoluted based on the multi-layer first convolution layer to obtain local features of different scales; the local features of different scales are down-sampled based on the multi-layer first down-sampling layer to obtain regional features of multiple local regions; the depth map projection result is image corrected based on the multi-layer first up-sampling layer and the regional features of the multiple local regions to obtain the viewpoint map of the remote user; specifically, in actual application, the depth map projection result can be viewpoint corrected based on the image correction model to obtain the final new viewpoint image (i.e. viewpoint map); further, in the process of image correction, the minimum network model that maintains the performance basically unchanged on single-person data can be obtained by reducing the number of network model layers, the number of convolution kernels, etc. The scene example diagram of the specific image correction process can be referred to in FIG. 12.

[0145] It should be noted here that based on the above-described image processing process, in the new viewpoint generation algorithm, the self-learning module, i.e. the deep learning-based module, is the depth calculation model and the image correction model; at the same time, due to the existence of the self-learning module, the minimum network model obtained under the condition of unchanged performance on single-person data can be published in the remote holographic communication system that can be widely used by users.

[0146] In step S140, the view point graph is subjected to image interleaving processing to obtain a target holographic image of the remote user, and the target holographic image is displayed.

[0147] In the example embodiment, first, the view point graph is subjected to image interleaving processing to obtain a target holographic image of the remote user; specifically, this can be achieved by calculating the interpolation range of the view point graph based on a vertex shader and calculating the interpolation content of the view point graph based on a fragment shader; and based on the interpolation range and interpolation content, the view point graph is subjected to image interleaving processing to obtain a target holographic image of the remote user. Specifically, in actual application, since the view point graph itself needs to be differentiated, the target left and right eye views need to be rendered before being processed using a shader. The shader described herein can include a vertex shader and a fragment shader. Specifically:

[0148] The vertex shader, also known as Vertshader, can be used to determine the display range of the view point graph, such as interleaving into 4K; in actual application, the vertex shader can act on each pixel vertex in the naked eye 3D resource and generate the final position of each pixel vertex; at the same time, the vertex shader is executed once for each pixel vertex to determine the final position of the pixel vertex; further, once the final position of each pixel vertex is determined, the GPU can assemble the set of visible vertices into points, straight lines and triangles to improve the speed of rendering scenes and models; and since the size of the view point graph is equivalent to half of the target overall view, the position of the pixel vertex in the view point graph needs to be determined before the multi-view point graph is interleaved.

[0149] The fragment shader, also known as PixShader, can be used to determine the interpolation content in the interleaving process. Specifically, since interleaving is equivalent to inserting RGB / RGBA values into a region according to a certain rule, the role of this shader is to insert the pixel values corresponding to the interleaving into the corresponding positions of the main display to cooperate with the 3D film for naked eye display, which depends on two factors: on the one hand, the RGB arrangement rule of the main display corresponding to the first user end; the RGB arrangement rule of the main display can include but is not limited to pixel spacing, screen line number and maximum offset, etc., and the RGB arrangement rule of the main display can be determined based on the extended display identification data of the main display; on the other hand, the view point interleaving rule; specifically, in actual application, the number of shader pixel matrices that need to be processed is different for 2-view points, the higher the number of view points, the higher the number of matrices, but the matrix size will decrease, because the length and width of a single view are getting smaller, so the shader pixel matrix needs to be calculated based on the fragment shader, and then the view point interleaving rule is determined based on the shader pixel matrix and the original pixel matrix.

[0150] Secondly, the target holographic image is displayed. Specifically, the following method can be used: determining a first user eye height of a local user and a second user eye height of a remote user; adjusting a display image height of the target holographic image at the first user end according to the first user eye height and the second user eye height, and displaying the target holographic image after the height adjustment.

[0151] In an example embodiment, the first user eye height of the local user and the second user eye height of the remote user can be determined by the following method: acquiring a first user image of the local user based on a first camera group associated with the first user end, and determining first left and right eye key point coordinates of the local user based on the first user image; determining the first user eye height of the left and right eyes of the local user on the first user end according to the first left and right eye key point coordinates; determining second pixel coordinates of the left and right eyes of the remote user on the target holographic image according to the target holographic image, and determining the second user eye height of the remote user on the first user end according to the second pixel coordinates.

[0152] In an example embodiment, the display image height of the target holographic image at the first user end can be adjusted according to the first user eye height and the second user eye height by the following method: determining an image height adjustment value of the target holographic image according to the first user eye height and the second user eye height, and adjusting the display image height of the target holographic image at the first user end according to the image height adjustment value.

[0153] In the following, the specific display process will be further explained and described. Specifically, in the actual application process, after obtaining the target holographic image, the target holographic image needs to be post-processed; the specific reason is that the screen installation height and the user seat height are different, so the height of the human body in the image collected by the camera is also different, which will lead to the difference between the eye height of the local user watching the screen and the eye height of the remote user in the screen, and cannot give the user the experience of visual interaction; therefore, the example embodiment of the disclosure dynamically adjusts the height of the human body (i.e. the height of the target holographic image) in the display picture through the detection of the eye key points to make the eye height of the local user consistent with the eye height of the remote user in the display. The specific implementation process is as follows:

[0154] First, the height estimation of the local user's eyes (the estimation of the first user's eye height of the local user); specifically, the local user images captured by the cameras 2 and 3 in the first camera group can be used to detect the pixel coordinates of the left and right eye key points, and then the three-dimensional coordinates of the left and right eyes in the camera 2 coordinate system are obtained through binocular distance measurement (taking the average of the two), and the three-dimensional coordinates are transformed into the three-dimensional coordinates in the screen center point coordinate system through the relative relationship between the camera and the screen determined during installation, and then the first user's eye height h1 of the local user on the screen of the first user terminal is obtained.

[0155] Second, the height estimation of the remote user's eyes (the estimation of the second user's eye height of the remote user); specifically, the resolution of the two generated new view images (the resolution in the example embodiment of the present disclosure can be set to 1024*1024) is known, and the pixel coordinates (u, v) of the left and right eyes of the remote user on the image are obtained using the eye detection algorithm (the average is also taken), and according to the physical size (h*w) of the screen and the resolution (H*W) of the screen, the second user's eye height of the remote user on the screen when displayed on the screen is v / 1024*h.

[0156] Finally, adjust the height of the display image (target holographic image); specifically, since the user's eyes and the remote user's eye height on the screen can be achieved by translating the image up and down, the height adjustment value h1-v / 1024*h can be calculated (usually moving upward), and the blank area generated by the translation can be covered by virtual conference tables, tea tables and other patterns, which enriches the scene picture and increases the contrast between the table and the human body, making the picture more three-dimensional and high.

[0157] At this point, the specific video call process has been fully implemented. In the actual application process, after the video call is over, further parameter adjustment needs to be made to the local original image processing model to ensure that the model parameters corresponding to each user are accurate. Specifically, the model parameter adjustment process can be realized in the following two ways:

[0158] The first implementation manner is: in response to the feedback result of the picture display quality corresponding to the target holographic image fed back by the second user terminal in response to the video call end instruction, it is determined whether the original image processing model needs to be optimized; when it is determined that the original image processing model needs to be optimized, the local image data of the local user is collected based on the first camera group, and the current local model parameters of the original image model are optimized based on the local image data, to obtain the original image model after optimization.

[0159] The second implementation manner is to determine a self-learning update frequency of the original image model, and perform optimization processing on current local model parameters of the original image model based on the self-learning update frequency and local image data of a local user collected by the first camera group, to obtain an original image model after optimization processing.

[0160] In the following, the specific adjustment process of the model parameters will be further explained and described. Specifically, when the user end detects a communication end signal, the video call is disconnected, and the local end system will remind the user whether to retain the model data of the remote end user; if yes, the user can directly click to save; if not, the user can ignore; further, after the communication ends, the system asks the two end users whether they are satisfied with the picture quality of the holographic remote video communication and scores it; at the same time, when the score of one end is lower than a threshold value, the system will prompt the user of the other end that the effect of the current version of the algorithm model parameters is not good, and suggest collecting a continuous video data for optimization. In the optimization process, a self-learning manner can be used to achieve it; wherein the self-learning function recorded in the example embodiments of the present disclosure can be divided into active self-learning and passive self-learning for the user. Among them:

[0161] (1) Active learning: that is, a learning process initiated by the user, which can be divided into two cases: 1) When the user logs in the system for the first time, the local stored image data is imported according to the system prompt or the continuous video of the specified action is recorded as the data source for training and optimization of the model parameters corresponding to the personal ID according to the system prompt; 2) After the user ends a call, the system prompts that the opposite user thinks that the quality of the call is not good, that is, the effect of the model parameters of the user is not good, the user agrees to record the video data as the training data for optimization and improvement of the self-learning module, the system will randomly or at a certain frequency extract image data from the images collected in the call process and generate new view point images to display to the user, and the user can reproduce the action according to the prompt for the system camera to collect; after the user uploads or collects the user data, the active learning starts the self-learning function of the core algorithm, and after the training and optimization are completed, the system will update the model parameters (version number + 1) stored in the local end of the user.

[0162] (2) Passive learning: that is, the system collects the image data of the local end user according to the user-selected self-learning module update frequency and other settings at a certain frame rate during the communication between the local end user and the remote end user, which is used as the training and optimization data source of the self-learning module; the passive learning will determine the start and stop and running time of the self-learning model according to the user-selected model parameter update frequency and other settings.

[0163] Thus far, the holographic image display method disclosed in the example embodiments of the present disclosure has been fully implemented. Based on the foregoing disclosure, it can be known that the holographic image display method disclosed in the example embodiments of the present disclosure can, on the one hand, solve the problem that the core algorithm of the holographic remote video communication system is limited by local hardware configuration (GPU power) and other factors, and cannot use a large model (model complexity, parameter order of magnitude, and data volume used for training) that takes into account generality and algorithm accuracy, so that the 3D display effect is poor (low algorithm model complexity, high real-time performance, but low algorithm accuracy, poor display effect) or the latency is large (high algorithm model complexity, high algorithm accuracy, good display effect, but high algorithm time consumption, low real-time performance), and the user experience is poor; on the other hand, it can also solve the problem that the core algorithm model cannot be customized for the user, does not have self-learning ability, and cannot improve the user experience; on the other hand, the holographic image display method provided by the example embodiments of the present disclosure can not only reduce the complexity and parameter amount of the core algorithm model deployed locally under the premise of ensuring the system display effect, improve the display frame rate, and reduce the stereoscopic display delay; but also optimize the performance of the algorithm model for each user to improve the user experience; on the other hand, the holographic image display method disclosed in the example embodiments of the present disclosure can also make full use of the idle time period of the system hardware to improve the hardware utilization rate of the entire system.

[0164] Finally, it also needs to be supplemented that, since the holographic remote video communication system will face tens of thousands of user groups in the future, the appearance, expression, and even action of each user are extremely personalized. Generally, a large model needs to consume a large amount of hardware resources (graphics cards, servers, etc.) and a time of several days or even months for a model parameter update based on user data; while the method provided by the example embodiments of the present disclosure can train and update a light network model based on small order data of a single user locally, and the required time is greatly reduced, which is calculated in hours. Therefore, the efficient update frequency can meet the timeliness requirement of the user for communication.

[0165] The following is a device embodiment of the present disclosure, which can be used to execute the method embodiments of the present disclosure. For details not disclosed in the device embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.

[0166] The example embodiments of the present disclosure also provide a holographic image display device configured in a first user terminal where a local user is located. Specifically, referring to FIG. 13, the holographic image display device can include a remote model parameter determination module 1310, a model parameter update module 1320, a view point graph determination module 1330, and a holographic image display module 1340. Among them:

[0167] The remote model parameter determination module 1310 can be configured to send a video call request to a second user terminal where a remote user corresponding to the local user is located, determine a target remote model parameter corresponding to the remote user in response to a call answering message corresponding to the video call request fed back by the second user terminal.

[0168] The model parameter updating module 1320 can be configured to perform parameter updating on an original image processing model in the first user terminal based on the target remote model parameter, to obtain a target image processing model.

[0169] The viewpoint map determination module 1330 can be configured to determine a viewpoint map of the remote user according to the target image processing model and remote image data of the remote user.

[0170] The holographic image display module 1340 can be configured to perform image interleaving processing on the viewpoint map to obtain a target holographic image of the remote user, and display the target holographic image.

[0171] In an example embodiment of the present disclosure, determining the target remote model parameter corresponding to the remote user comprises: querying a first remote model parameter corresponding to the remote user in the first user terminal according to a remote user identifier of the remote user, and querying a second remote model parameter corresponding to the remote user in the second user terminal; comparing a first parameter version number of the first remote model parameter and a second parameter version number of the second remote model parameter, and selecting the target remote model parameter from the first remote model parameter and the second remote model parameter based on a version number comparison result.

[0172] In an example embodiment of the present disclosure, performing parameter updating on the original image processing model in the first user terminal based on the target remote model parameter to obtain the target image processing model comprises: calling the original image processing model in the first user terminal, and extracting a local model parameter in the original image processing model; replacing the local model parameter based on the target remote model parameter to obtain a target image processing model associated with the remote user.

[0173] In an example embodiment of the present disclosure, the target image processing model comprises a depth map calculation model, a depth fusion module, a projection transformation module, and an image rectification model; wherein, according to the target image processing model and the remote image data of the remote user, the viewpoint map of the remote user is determined, comprising: determining the depth image data of the remote image data according to the depth map calculation model, and fusing the depth image data according to the depth fusion module to obtain a fused depth image; projecting the fused depth image to a new viewpoint pixel plane corresponding to the local user based on the projection transformation module to obtain a depth map projection result; and rectifying the depth map projection result based on the image rectification model to obtain the viewpoint map of the remote user.

[0174] In an example embodiment of the present disclosure, the depth map determination model comprises a feature encoder, a context encoder, a correlation pyramid, and a disparity update module; wherein, according to the depth map calculation model, the depth image data of the remote image data is determined, comprising: determining the left eye feature map of the remote left eye image in the remote image data and the right eye feature map of the remote right eye image in the remote image data based on the feature encoder, and determining the context feature of the remote left eye image based on the context encoder; calculating the correlation volume between the left eye feature map and the right eye feature map, and constructing a multi-layer correlation pyramid based on the left eye feature map, the right eye feature map, and the correlation volume; indexing the correlation volume associated with the preset disparity parameter from the correlation pyramid according to the preset disparity parameter, and constructing the correlation feature according to the correlation volume associated with the preset disparity parameter and the left eye feature map and the right eye feature map corresponding to the correlation volume; inputting the correlation feature, the context feature, and the preset disparity parameter into the disparity update module to obtain the disparity field of the remote left eye image, and determining the depth image data according to the disparity field.

[0175] In an example embodiment of the present disclosure, the depth map determination model is obtained by: sharing the first weight value of the feature encoder and the weight value of the context encoder in the binocular depth estimation model to obtain the depth map determination model; or compressing the first number of convolution kernels of the feature encoder and / or the second number of convolution kernels of the context encoder to obtain the depth map determination model; or compressing the first number of convolution layers of the feature encoder and / or the second number of convolution layers of the context encoder to obtain the depth map determination model.

[0176] In an example embodiment of the present disclosure, the image correction model comprises a plurality of first convolution layers, a plurality of first down-sampling layers, and a plurality of first up-sampling layers, the number of layers of the plurality of first down-sampling layers and the plurality of first up-sampling layers being the same; and the image correction of the depth map projection result based on the image correction model to obtain the viewpoint image of the remote user comprises: performing convolution processing on the depth map projection result based on the plurality of first convolution layers to obtain a plurality of local features of different scales; performing down-sampling on the plurality of local features of different scales based on the plurality of first down-sampling layers to obtain regional features of a plurality of local regions; and performing image correction on the depth map projection result based on the plurality of first up-sampling layers and the regional features of the plurality of local regions to obtain the viewpoint image of the remote user.

[0177] In an example embodiment of the present disclosure, the image interleaving processing of the viewpoint image to obtain the target holographic image of the remote user comprises: calculating an interpolation range of the viewpoint image based on a vertex shader, and calculating interpolation content of the viewpoint image based on a fragment shader; and performing image interleaving processing on the viewpoint image based on the interpolation range and the interpolation content to obtain the target holographic image of the remote user.

[0178] In an example embodiment of the present disclosure, the display of the target holographic image comprises: determining a first user eye height of a local user and a second user eye height of a remote user; adjusting a display image height of the target holographic image at a first user end according to the first user eye height and the second user eye height, and displaying the target holographic image after the height adjustment.

[0179] In an example embodiment of the present disclosure, the determination of the first user eye height of the local user and the second user eye height of the remote user comprises: capturing a first user image of the local user based on a first camera group associated with the first user end, and determining first left and right eye key point coordinates of the local user based on the first user image; determining a first user eye height of the local user on the first user end according to the first left and right eye key point coordinates; determining second pixel coordinates of the second left and right eyes of the remote user on the target holographic image according to the target holographic image, and determining a second user eye height of the remote user on the first user end according to the second pixel coordinates.

[0180] In an example embodiment of the present disclosure, the display image height of the target holographic image at the first user end is adjusted according to the first user eye height and the second user eye height, including: determining an image height adjustment value of the target holographic image according to the first user eye height and the second user eye height, and adjusting the display image height of the target holographic image at the first user end according to the image height adjustment value.

[0181] In an example embodiment of the present disclosure, the original image processing model is obtained by: obtaining a first historical user image of a local user, and pre-processing the first historical user image to obtain a first target user image; inputting the first target user image into a network model to be trained to obtain a first prediction result, and constructing a loss function according to a viewpoint label image corresponding to the first target user image and the first prediction result; adjusting parameters in the network model to be trained based on the loss function to obtain the original image processing model.

[0182] In an example embodiment of the present disclosure, the first historical user image is pre-processed to obtain a first target user image, including: performing face recognition on the first historical user image based on a preset face recognition algorithm to obtain a first historical face image corresponding to the local user; pre-processing the first historical face image to obtain the first target user image; wherein the pre-processing manner includes at least one of adding noise, color conversion, cropping and lifting.

[0183] In an example embodiment of the present disclosure, the original image model is updated in parameters by: in response to a feedback result of a picture display quality corresponding to the target holographic image fed back by the second user end in response to a video call end instruction, determining whether the original image processing model needs to be optimized; when it is determined that the original image processing model needs to be optimized, collecting local image data of the local user based on the first camera group, and optimizing the current local model parameters of the original image model based on the local image data to obtain an optimized original image model.

[0184] In an example embodiment of the present disclosure, the original image model is also updated in parameters by: determining a self-learning update frequency of the original image model, and optimizing the current local model parameters of the original image model based on the self-learning update frequency and the local image data of the local user collected by the first camera group to obtain an optimized original image model.

[0185] The specific details of each module of the holographic image display device have been described in detail in the corresponding holographic image display method, and thus will not be described here.

[0186] It should be noted that although several modules or units of the device for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. Indeed, according to embodiments of the disclosure, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into several modules or units embodied.

[0187] Moreover, although the various steps of the methods of the disclosure are described in a particular order in the figures, this is not required or implied as to the order of the steps, nor is it required that all of the steps be performed to achieve the desired result. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, one step can be broken into multiple steps, etc.

[0188] In an exemplary embodiment of the disclosure, an electronic device capable of implementing the above method is also provided.

[0189] Those skilled in the art can understand that various aspects of the disclosure can be implemented as a system, a method or a program product. Therefore, various aspects of the disclosure can be embodied as a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" here.

[0190] The electronic device 1400 according to such an embodiment of the disclosure will be described below with reference to FIG. 14. FIG. 14 shows the electronic device 1400 only as an example, and should not bring any limitation to the functions and use range of the embodiments of the disclosure.

[0191] As shown in FIG. 14, the electronic device 1400 is in the form of a general computing device. The components of the electronic device 1400 can include, but are not limited to, the at least one processing unit 1410 described above, the at least one storage unit 1420 described above, a bus 1430 connecting different system components (including the storage unit 1420 and the processing unit 1410), and a display unit 1440.

[0192] The storage unit stores program codes which can be executed by the processing unit 1410, so that the processing unit 1410 performs the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of the present specification. For example, the processing unit 1410 can perform the steps as shown in FIG. 1, such as step S110 of sending a video call request to a second user terminal where a remote user corresponding to the local user is located, responding to a call answering message corresponding to the video call request fed back by the second user terminal, and determining a target remote model parameter corresponding to the remote user; step S120 of performing parameter updating on an original image processing model in the first user terminal based on the target remote model parameter to obtain a target image processing model; step S130 of determining a view point map of the remote user according to the target image processing model and remote image data of the remote user; and step S140 of performing image interleaving processing on the view point map to obtain a target holographic image of the remote user, and displaying the target holographic image.

[0193] The storage unit 1420 can include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 14201 and / or a cache memory 14202, and can further include a read-only memory (ROM) 14203.

[0194] The storage unit 1420 can further include program / utilities 14204 having a set of (at least one) program modules 14205, which include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, and each of these examples or some combination thereof can include implementation of a network environment.

[0195] The bus 1430 can represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.

[0196] The electronic device 1400 can also communicate with one or more external devices 1500 such as a keyboard or pointing device, a Bluetooth device, or a device for reading media. Communication with one or more devices can enable a user to interact with the electronic device 1400 in order to use it or perform methods described herein. In some embodiments, the communication can be facilitated via an I / O interface 1450. Additionally, the electronic device 1400 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), or the Internet, through a network adapter 1460. As depicted, the network adapter 1460 can communicate with the other components of the electronic device 1400 through the bus 1430. It should be understood that, although not shown explicitly, other hardware and / or software components can be used in conjunction with the electronic device 1400. These components include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0197] From the above description of the embodiments, those skilled in the art will readily appreciate that the example embodiments described herein can be implemented by software and / or by hardware coupled with software. Accordingly, the technical solutions of the embodiments of the present disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash disk, or a mobile hard disk) or a network, and includes a number of instructions for causing a computing device (such as a personal computer, a server, a terminal device, or a network device) to perform the methods according to the embodiments of the present disclosure.

[0198] In the example embodiments of the present disclosure, a computer-readable storage medium is also provided, which stores a program product capable of implementing the above-described methods of the present disclosure. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program codes for causing a terminal device to perform the steps described in the above “Example Methods” section according to various example embodiments of the present disclosure when the program product is run on the terminal device.

[0199] The program product for implementing the above-described methods according to the embodiments of the present disclosure can take the form of a portable compact disc read-only memory (CD-ROM) and include program codes, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited to this, and in this document, a readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or apparatus.

[0200] The program product can employ any combination of one or more computer-readable media. The computer-readable media can be a computer-readable storage medium or a computer-readable signal medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include the following: an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0201] The computer-readable signal medium can include a computer-readable storage medium that is propagated as a carrier wave. The computer-readable signal medium can further be any computer-readable medium that is not a storage medium. The computer-readable signal medium can be a computer-readable storage medium that is a propagated signal on a computer-readable storage medium.

[0202] The program code embodied on the computer-readable media can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0203] The program code can be executed by one or more programmable processors, which can be implemented in one or more computer systems. In this context, a computer system generally includes a plurality of these programmable processors, which work in concert to perform a task. Additionally, the program code can be downloaded from an external source, including the internet, to the computer system.

[0204] In addition, the aforementioned diagrams are merely schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the aforementioned diagrams do not indicate or limit the time sequence of the processes. In addition, it is readily understood that the processes can be executed synchronously or asynchronously, for example, in a plurality of modules.

[0205] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features of the disclosure as set forth above. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the disclosure being indicated by the following claims.

Claims

1. A method for displaying a holographic image, characterized in that, The method for displaying the holographic image, configured on the first user terminal where the local user resides, includes: Send a video call request to the second user terminal where the remote user corresponding to the local user is located, respond to the call answer message corresponding to the video call request fed back by the second user terminal, and determine the target remote model parameters corresponding to the remote user; Based on the target remote model parameters, the original image processing model in the first user terminal is updated to obtain the target image processing model. Based on the target image processing model and the remote image data of the remote user, the viewpoint map of the remote user is determined; The viewpoint image is subjected to image interleaving processing to obtain the target holographic image of the remote user, and the target holographic image is displayed.

2. The method for displaying a holographic image according to claim 1, characterized in that, Determining the target remote model parameters corresponding to the remote user includes: Based on the remote user's remote user identifier, query the first remote model parameter corresponding to the remote user in the first user terminal, and query the second remote model parameter corresponding to the remote user in the second user terminal; The version number of the first parameter of the first remote model parameter and the version number of the second parameter of the second remote model parameter are compared, and the target remote model parameter is selected from the first remote model parameter and the second remote model parameter based on the version number comparison result.

3. The method for displaying a holographic image according to claim 1, characterized in that, Based on the target remote model parameters, the original image processing model in the first user terminal is updated to obtain the target image processing model, including: The original image processing model in the first user terminal is invoked, and the local model parameters in the original image processing model are extracted. The local model parameters are replaced based on the target remote model parameters to obtain the target image processing model associated with the remote user.

4. The method for displaying a holographic image according to claim 1, characterized in that, The target image processing model includes a depth map calculation model, a depth fusion module, a projection transformation module, and an image correction model; Specifically, the viewpoint map of the remote user is determined based on the target image processing model and the remote user's remote image data, including... The depth image data of the remote image data is determined according to the depth map calculation model, and the depth image data is fused according to the depth fusion module to obtain a fused depth image; Based on the projection transformation module, the fused depth image is projected onto a new viewpoint pixel plane corresponding to the local user to obtain the depth map projection result; Based on the image correction model, the depth map projection result is corrected to obtain the viewpoint map of the remote user.

5. The method for displaying a holographic image according to claim 4, characterized in that, The depth map determination model includes a feature encoder, a context encoder, an association pyramid, and a disparity update module; The depth image data determined according to the depth map calculation model includes: The feature encoder determines the left-eye feature map of the remote left-eye image in the remote image data and the right-eye feature map of the remote right-eye image in the remote image data, and the context encoder determines the context features of the remote left-eye image. Calculate the correlation volume between the left-eye feature map and the right-eye feature map, and construct a multi-layer correlation pyramid based on the left-eye feature map, the right-eye feature map, and the correlation volume; Based on the preset disparity parameters, the relevant volumes associated with the preset disparity parameters are indexed from the association pyramid, and the relevant features are constructed based on the relevant volumes associated with the preset disparity parameters and the left and right eye feature maps corresponding to the relevant volumes. The correlation features, context features, and preset disparity parameters are input into the disparity update module to obtain the disparity field of the remote left-eye image, and the depth image data is determined based on the disparity field.

6. The method for displaying a holographic image according to claim 4 or 5, characterized in that, The depth map determination model is obtained in the following way: The first weight value of the feature encoder and the weight value of the context encoder in the binocular depth estimation model are shared to obtain the depth map determination model; or The depth map determination model is obtained by compressing the number of first convolutional kernels of the feature encoder and / or the number of second convolutional kernels of the context encoder; or The depth map determination model is obtained by compressing the first convolutional layer number of the feature encoder and / or the second convolutional layer number of the context encoder.

7. The method for displaying a holographic image according to claim 4, characterized in that, The image correction model includes multiple first convolutional layers, multiple first downsampling layers, and multiple first upsampling layers, wherein the number of layers in the multiple first downsampling layers and the multiple first upsampling layers is the same. Specifically, the image correction process involves performing image correction on the depth map projection result based on the image correction model to obtain the viewpoint map of the remote user, including: The depth map projection result is convolved based on the multi-layer first convolutional layer to obtain multiple local features at different scales; Based on the multi-layer first downsampling layer, the local features at multiple different scales are downsampled to obtain the regional features of multiple local regions; Based on the regional features of the multi-layer first upsampling layer and the multiple local regions, the depth map projection result is image corrected to obtain the viewpoint map of the remote user.

8. The method for displaying a holographic image according to claim 1, characterized in that, The viewpoint map is subjected to image interleaving processing to obtain the target holographic image of the remote user, including: The interpolation range of the view graph is calculated based on the vertex shader, and the interpolation content of the view graph is calculated based on the fragment shader. Based on the interpolation range and interpolation content, the viewpoint map is subjected to image interleaving processing to obtain the target holographic image of the remote user.

9. The method for displaying a holographic image according to claim 1, characterized in that, Displaying the target holographic image includes: Determine the first user eye height for local users and the second user eye height for remote users; The height of the target holographic image displayed on the first user's terminal is adjusted according to the eye height of the first user and the eye height of the second user, and the target holographic image after height adjustment is displayed.

10. The method for displaying a holographic image according to claim 9, characterized in that, Determine the first user eye height for local users and the second user eye height for remote users, including: The first user image of the local user is acquired based on the first camera group associated with the first user terminal, and the coordinates of the first left and right eye key points of the local user are determined based on the first user image. Determine the height of the first user's first eye on the first user terminal based on the coordinates of the first left and right eye key points; The second pixel coordinates of the remote user's second left and right eyes on the target holographic image are determined based on the target holographic image, and the height of the remote user's second eye on the first user terminal is determined based on the second pixel coordinates.

11. The method for displaying a holographic image according to claim 9, characterized in that, Adjusting the display height of the target holographic image on the first user's end based on the eye height of the first user and the eye height of the second user includes: Based on the eye height of the first user and the eye height of the second user, the image height adjustment value of the target holographic image is determined, and the display image height of the target holographic image on the first user terminal is adjusted according to the image height adjustment.

12. The method for displaying a holographic image according to claim 1, characterized in that, The original image processing model was obtained in the following way: The first historical user image of the local user is obtained, and the first historical user image is preprocessed to obtain the first target user image; The first target user image is input into the network model to be trained to obtain the first prediction result, and a loss function is constructed based on the viewpoint label image corresponding to the first target user image and the first prediction result. The parameters in the network model to be trained are adjusted based on the loss function to obtain the original image processing model.

13. The method for displaying a holographic image according to claim 12, characterized in that, The first historical user image is preprocessed to obtain the first target user image, including: Based on a preset face recognition algorithm, face recognition is performed on the first historical user image to obtain the first historical face image corresponding to the local user; The first historical face image is preprocessed to obtain the first target user image; wherein the preprocessing method includes at least one of adding noise, color conversion, cropping, and stretching.

14. The method for displaying a holographic image according to claim 12, characterized in that, The original image model is updated with parameters in the following manner: In response to the feedback result of the second user terminal on the display quality of the screen corresponding to the target holographic image in response to the video call end command, it is determined whether the original image processing model needs to be optimized. When it is determined that the original image processing model needs to be optimized, local image data of the local user is collected based on the first camera group, and the current local model parameters of the original image model are optimized based on the local image data to obtain the optimized original image model.

15. The method for displaying a holographic image according to claim 12, characterized in that, The original image model is also updated in the following way: The self-learning update frequency of the original image model is determined, and based on the self-learning update frequency and the local image data of the local user collected by the first camera group, the current local model parameters of the original image model are optimized to obtain the optimized original image model.

16. A holographic image display device, characterized in that, The holographic image display device, configured on the first user terminal where the local user is located, includes: The remote model parameter determination module is used to send a video call request to the second user terminal where the remote user corresponding to the local user is located, respond to the call answer message corresponding to the video call request fed back by the second user terminal, and determine the target remote model parameters corresponding to the remote user. The model parameter update module is used to update the parameters of the original image processing model in the first user terminal based on the target remote model parameters to obtain the target image processing model. The viewpoint map determination module is used to determine the viewpoint map of the remote user based on the target image processing model and the remote image data of the remote user. The holographic image display module is used to perform image interlacing processing on the viewpoint map to obtain the target holographic image of the remote user, and to display the target holographic image.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for displaying holographic images according to any one of claims 1-15.

18. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method for displaying a holographic image according to any one of claims 1-15 by executing the executable instructions.

Citation Information

Patent Citations

  • Video call method and device, storage medium and program product

    CN113709401A

  • Image display method, image display device, equipment and storage medium

    CN114840165A

  • Video call method and device, electronic equipment and storage medium

    CN117917889A

  • Holographic image display method and device, computer readable storage medium and electronic equipment

    CN118869963A

  • Hybrid visual communication

    US20150042743A1

Cited By

  • Data processing method and device for naked eye 3D display

    CN122027779A