Video call method and apparatus, electronic device, and storage medium

By acquiring and adjusting the user model image of the second terminal and transmitting user feature information to simulate real-time status, the problems of video call stability and high cost are solved, and low-data-volume, high-stability video interaction is achieved.

CN117917889BActive Publication Date: 2025-11-18MOORE THREADS TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211295064.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2025-11-18
Estimated Expiration
2042-10-21

AI Technical Summary

Technical Problem

Existing video call technologies suffer from poor stability, large data volumes, high network speed requirements, and high encoding and decoding costs, making it difficult to achieve real-time video interaction.

Method used

By acquiring the user model of the second terminal and receiving its feature information, adjusting and displaying the model image, transmitting user feature information to simulate the user's real-time state, reducing data volume and network speed requirements, and using 3D rendering software for image processing.

Benefits of technology

It reduces data transmission volume and traffic costs, improves the stability and real-time performance of video calls, reduces encoding and decoding costs, and enhances video clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117917889B_ABST
    Figure CN117917889B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video call method, device, electronic equipment and storage medium, the video call method comprising: obtaining a user model corresponding to a second terminal; receiving user feature information sent by the second terminal during a video call process; and adjusting and displaying a model image corresponding to the user model according to the user feature information. In the present disclosure, the video call mode transmits user feature information, so compared with the video call mode of the related art which transmits image pixel information, the amount of data transmitted is smaller, the requirement for network speed is also smaller, which is conducive to reducing the cost of traffic and also conducive to realizing the video call function with better real-time performance. In addition, in the present disclosure, the clarity of the video is less related to the code rate, which is also conducive to reducing the coding and decoding cost of the developer of the video call function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of information processing, and particularly relates to a video call method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the development of the telecommunication field, more and more users begin to use video call. The video call can collect video in real time through the camera of an encoding terminal, and then send the video to a decoding terminal for video decoding, and finally display on the display screen of the decoding terminal to realize real-time video interaction. However, the video call function usually has poor stability, and therefore how to better realize the video call function is a technical problem to be solved by developers. SUMMARY

[0003] The present disclosure provides a video call technical solution.

[0004] According to an aspect of the present disclosure, a video call method is provided, applied to a first terminal, the first terminal communicates with a second terminal, and the video call method comprises: acquiring a user model corresponding to the second terminal; wherein the user model is used to represent a three-dimensional model corresponding to a target user in the second terminal for video call; in a video call process, receiving user feature information sent by the second terminal; and according to the user feature information, adjusting and displaying a model image corresponding to the user model.

[0005] In a possible implementation, the user feature information comprises sound information corresponding to the target user, and the adjusting and displaying of the model image corresponding to the user model according to the user feature information comprises: determining visual information corresponding to the target user according to the sound information corresponding to the target user; and adjusting and displaying the model image corresponding to the user model according to the visual information.

[0006] In a possible implementation, the user feature information comprises visual information corresponding to the target user, and the adjusting and displaying of the model image corresponding to the user model according to the user feature information comprises: adjusting and displaying the model image corresponding to the user model according to the visual information.

[0007] In a possible implementation, the adjusting and displaying the model image corresponding to the user model according to the visual information comprises: adjusting the user model according to the visual information; generating a model image corresponding to the adjusted user model, and determining first image features corresponding to the model image; obtaining a preset face image, and determining second image features corresponding to the preset face image; adjusting the second image features according to the visual information; performing feature fusion on the first image features and the adjusted second image features to obtain third image features; and adjusting and displaying the model image corresponding to the user model according to the third image features.

[0008] In a possible implementation, the preset face image is a front face image of the target user, and the second image features are feature maps corresponding to the front face image; and the adjusting the second image features according to the visual information comprises: determining pixel motion vectors according to the visual information; and adjusting positions of pixels in the feature maps according to the pixel motion vectors, and taking the adjusted positions of the pixels as the adjusted second image features.

[0009] In a possible implementation, the receiving the user feature information sent by the second terminal during the video call comprises: receiving, during the video call, user feature information compressed by the second terminal and sent by the second terminal; and the adjusting and displaying the model image corresponding to the user model according to the user feature information comprises: decompressing the compressed user feature information, and adjusting and displaying the model image corresponding to the user model according to the decompressed user feature information.

[0010] In a possible implementation, the video call method further comprises: obtaining a special effect parameter input or preset during the video call, and adjusting a user model corresponding to the first terminal or a user model corresponding to the second terminal; and the special effect parameter is used to adjust a display effect of the user model.

[0011] In a possible implementation, the visual information comprises at least one of head posture information and expression information of the target user.

[0012] According to an aspect of the present disclosure, a video call method is provided, which is applied to a second terminal, the second terminal communicates with a first terminal, and the video call method comprises: in a case of video call with the first terminal, collecting a to-be-processed image or a to-be-processed sound during the video call; determining user feature information corresponding to a target user according to the to-be-processed image or the to-be-processed sound, and sending the user feature information to the first terminal.

[0013] In a possible implementation, the sending of the user feature information to the first terminal comprises compressing the user feature information and sending the compressed user feature information to the first terminal.

[0014] According to an aspect of the present disclosure, a video call device is provided, which communicates with a second terminal, and comprises: a user model obtaining module, configured to obtain a user model corresponding to the second terminal; wherein the user model is used to represent a three-dimensional model corresponding to a target user who performs a video call through the second terminal; a feature information obtaining module, configured to receive user feature information or related information of the user feature information sent by the second terminal during a video call process; and a model adjusting module, configured to adjust and display a model image corresponding to the user model according to the user feature information or the related information.

[0015] According to an aspect of the present disclosure, a video call device is provided, which communicates with a first terminal, and comprises: a data collecting module, configured to collect a to-be-processed image or a to-be-processed sound during a video call process when performing a video call with the first terminal; and an information sending module, configured to determine user feature information corresponding to a target user according to the to-be-processed image, and send the user feature information to the first terminal, or send the to-be-processed sound as related information of the user feature information to the first terminal.

[0016] According to an aspect of the present disclosure, an electronic device is provided, which comprises: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored in the memory to perform the above-mentioned video call method.

[0017] According to an aspect of the present disclosure, a computer readable storage medium is provided, which stores computer program instructions, and the computer program instructions are executed by a processor to implement the above-mentioned video call method.

[0018] In the embodiments of the present disclosure, the user model corresponding to the second terminal can be acquired, and then in the process of the video call, the user feature information sent by the second terminal is received, and finally the model image corresponding to the user model is adjusted and displayed according to the user feature information. In the embodiments of the present disclosure, the user model corresponding to the second terminal displayed in the first terminal can be adjusted by collecting the user feature information, so as to simulate the real-time state of the user corresponding to the second terminal. Since the embodiments of the present disclosure transmit the user feature information, compared with the way of transmitting image pixel information in the related art, the transmission data amount of the embodiments of the present disclosure is smaller, and the requirement for network speed is also smaller, which is beneficial to reduce the traffic cost and also beneficial to realize the video call function with better real-time performance. In addition, in the embodiments of the present disclosure, the clarity of the video is less related to the code rate, which is also beneficial to the developers of the video call function to reduce the coding and decoding cost.

[0019] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, rather than limiting the present disclosure. Other features and aspects of the present disclosure will become apparent based on the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the technical solutions of the present disclosure.

[0021] Figure 1 A flowchart of a video call method according to an embodiment of the present disclosure is shown.

[0022] Figure 2 A flowchart of a video call method according to an embodiment of the present disclosure is shown.

[0023] Figure 3 A reference schematic diagram of a model structure according to an embodiment of the present disclosure is shown.

[0024] Figure 4 A flowchart of another video call method according to an embodiment of the present disclosure is shown.

[0025] Figure 5 A reference schematic diagram of a video call method according to an embodiment of the present disclosure is shown.

[0026] Figure 6 A block diagram of a video call device according to an embodiment of the present disclosure is shown.

[0027] Figure 7 A block diagram of another video call device according to an embodiment of the present disclosure is shown.

[0028] Figure 8A block diagram of an electronic device is shown according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0029] Various exemplary embodiments, features and aspects of the present disclosure will be explained in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar elements. Although various aspects of the embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically noted.

[0030] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0031] The term "and / or" used herein only means an association relationship of the associated objects, and means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality, for example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0032] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, elements and circuits that are well known to those skilled in the art are not described in detail in order to highlight the main idea of the present disclosure.

[0033] In the related art, the video call technology is based on the video coding technology, which needs to encode the image by the encoding terminal, and then send the encoded data to the decoding terminal. The decoding terminal decodes it through a preset algorithm, and then displays the image collected by the encoding terminal on the display screen of the decoding terminal. The process of encoding the image by the encoding terminal, since it is aimed at the pixel information in the whole image, the amount of data after encoding is large. In addition, the size of the picture quality depends on the size of the code rate, and the requirement for network speed is higher.

[0034] Therefore, the embodiment of the present disclosure provides a video call method, applied to a first terminal, the first terminal communicates with a second terminal, the video call method comprises: obtaining a user model corresponding to the second terminal, then receiving user feature information sent by the second terminal in a video call process, and finally adjusting and displaying a model image corresponding to the user model according to the user feature information. The embodiment of the present disclosure can adjust the user model corresponding to the second terminal displayed in the first terminal by collecting user feature information, so as to simulate the real-time state of the user corresponding to the second terminal. Since the embodiment of the present disclosure transmits user feature information, compared with the way of transmitting image pixel information in the related art, the transmission data amount of the embodiment of the present disclosure is smaller, and the requirement for network speed is also smaller, which is beneficial to reduce the traffic cost and realize the video call function with better real-time performance. In addition, the clarity of the video in the embodiment of the present disclosure is less related to the code rate, which is also beneficial to the developers of the video call function to reduce the coding and decoding cost.

[0035] Exemplarily, the first terminal in the embodiment of the present disclosure can be used as a decoding terminal, and the second terminal can be used as an encoding terminal. It should be understood that the above-mentioned encoding terminal and decoding terminal are only the functional naming of the encoding and decoding process. In actual scenarios, the decoding terminal can also be used as an encoding terminal, and the encoding terminal can also be used as a decoding terminal. The functions of the two can also be executed simultaneously. For example, in a video call, if the network speed is ideal, both users communicating can see the video taken by the other terminal in real time in their own terminals, that is, the terminal device itself can enable decoding function and also can enable encoding function at the same time. In this case, the first terminal or the second terminal can be called a decoding terminal and an encoding terminal. In the following, the embodiment of the present disclosure will take the first terminal as a decoding terminal and the second terminal as an encoding terminal to describe the interaction process, but it should not be limited to that the first terminal or the second terminal only has one of the decoding or encoding functions, or cannot execute the above two functions at the same time.

[0036] In a possible implementation, the first terminal or the second terminal can be a terminal device or other electronic device. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be realized by a processor invoking computer-readable instructions stored in a memory.

[0037] The embodiment of the present disclosure provides a video call method, applied to a first terminal, the first terminal communicates with a second terminal, referring to Figure 1 the drawing.Figure 1 A flow chart of a video call method provided by an embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the video call method comprises the following steps. Figure 1 As shown in FIG. 1, the video call method comprises the following steps. Step S100: obtaining a user model corresponding to the second terminal. The user model is used to represent a three-dimensional model corresponding to a target user who performs a video call through the second terminal. The three-dimensional model can be obtained by pre-modeling. For example, the user can collect at least one image with user facial information through a camera or an RGBD camera (a camera capable of shooting depth) of the second terminal, and then the second terminal or a server connected to the second terminal can convert the two-dimensional image of the user into a three-dimensional model (e.g., represented as a mesh, a skeleton, a texture, etc.) of the user through a modeling algorithm in the related art, and then the second terminal or the server can send related information (e.g., face mesh information, face skeleton information, etc.) of the three-dimensional model of the target user to the first terminal as the user model when the first terminal is about to perform a video call with the second terminal. The user model can have the features of the three-dimensional model. The establishment process of the three-dimensional model can be completed at any time and then saved by the target user, which can be called when the video call function is implemented. For example, each user can upload the corresponding user model to the server, and then directly call the user model when the user uses the video call function, i.e., the user does not need to re-establish the user model for each video call, which is beneficial to improve the convenience of the video call.

[0038] Step S200: receiving user feature information sent by the second terminal in the video call process. The user feature information is used to represent various features of the target user, which can be selected by a developer and is related to the loading effect of the video in the video call, which is not limited in the present disclosure.

[0039] Step S300, according to the user feature information, adjusting and displaying the model image corresponding to the user model. For example, the above-mentioned user feature information and the user model can be linked, for example: the user feature information can include expression information, which can be expressed as a blendshape parameter in the related art (a parameter representing an expression), which can be divided into multiple sub-parameters, each of which can be used to control a part of the facial expression, such as a sub-parameter responsible for controlling the opening degree of the left eye, a sub-parameter responsible for controlling the curvature of the lips, etc. The numerical range of the above-mentioned sub-parameters can be 0 to 1, which can represent the opening degree, curvature, etc. The blendshape parameter itself can be directly used in three-dimensional rendering software (for example: Unreal Engine, Uniy3D engine, etc.), which can directly adjust the expression of the user model through the expression information, so that the user model simulates the real-time expression of the target user. For another example: the user feature information can include head posture information, which can be expressed as a head rotation angle in the related art to represent the head angle, a head translation amount to represent the head position, etc. to represent the head posture of the user. In one example, the above-mentioned expression information and head posture information can also be expressed as other forms of data in the related art, which can be converted into parameters compatible with the user model through a data conversion algorithm in the related art. The specific data form of the expression information is not limited in the embodiment of the present disclosure. For example, the above-mentioned user model can be input into three-dimensional rendering software to obtain a model image corresponding to the user model (for example, by a camera in the three-dimensional rendering software to realize the collection of model images at different angles).

[0040] In a possible implementation, step S200 can include: receiving the user feature information compressed by the second terminal and sent by the second terminal during the video call. Step S300 can include: decompressing the compressed user feature information, and adjusting and displaying the model image corresponding to the user model according to the decompressed user feature information. For example, the above-mentioned user feature information can be compressed to further reduce the amount of data required during the video call, which is beneficial to reduce the cost of video call and improve the stability of the video call process. The specific way of compression is not limited in the embodiment of the present disclosure, and the developer can flexibly set according to the actual needs.

[0041] In a possible implementation, the user feature information includes sound information corresponding to the target user, and step S300 can include: determining visual information corresponding to the target user according to the sound information corresponding to the target user, and then adjusting and displaying a model image corresponding to the user model according to the visual information. For example, the sound information can be input into a trained machine learning model to output predicted visual information. The input of the machine learning model in the training stage can include training sound information and visual information corresponding to the training sound information, to iteratively guide the machine learning model, and after the final model converges, the training is considered to be completed. The model structure of the machine learning model and specific training details can be determined by the developer according to the actual situation, and the embodiments of the present disclosure do not limit this. In an example, the visual information can include at least one of head pose information and expression information corresponding to the target user. The embodiments of the present disclosure allow using sound as the feature information of the target user, and then the image can be adjusted according to the transmission of the feature information, which can further reduce the bandwidth requirement and the amount of data required for transmission.

[0042] In a possible implementation, the user feature information includes visual information corresponding to the target user. For example, Figure 2 As shown in the figure, Figure 2 A flowchart of a video call method is shown, which is provided by the embodiments of the present disclosure. Figure 2 As shown in the figure, step S300 can include: step S310, adjusting and displaying a model image corresponding to the user model according to the visual information. For example, the visual information can include any information related to the visual effect of the target user at the first terminal, such as head pose, expression, ambient light, etc. The visual information can be directly used as an adjustment parameter of the user model, so that the user model can simulate the current facial pose and expression of the user. Then, the user model can be collected by a three-dimensional rendering software to obtain a two-dimensional image, which is used as the model image.

[0043] In a possible implementation, step S310 can include: step S311, adjusting the user model according to the visual information. For example, the head posture of the user model can be adjusted according to the head posture information in the visual information, so that the head posture of the user model in the image captured by the camera (a tool in the three-dimensional rendering software for capturing and displaying a two-dimensional image) is similar or identical to the head posture of the user, so as to restore the shooting scene. For another example, the facial expression of the user model can be adjusted according to the expression information in the visual information, so that the facial expression of the user model in the image captured by the camera is similar or identical to the facial expression of the user. For example, in an example, the user of the first terminal can also adjust the shooting angle of the camera in the three-dimensional rendering software, that is, the user can arbitrarily change the display effect of the user of the second terminal in the first terminal to see the multi-angle image of the user of the second terminal, and the camera in the three-dimensional rendering software can also always maintain a preset angle, which is not limited in the embodiments of the present disclosure.

[0044] Step S312, generating a model image corresponding to the adjusted user model according to the adjusted user model, and determining a first image feature corresponding to the model image. For example, the two-dimensional image of the user model can be captured by the camera tool in the three-dimensional rendering software. In an example, the first image feature can be extracted by a feature extraction model in the related art, for example, a CNN (Convolutional Neural Network, convolutional neural network), a VGG (Visual Geometry Group, a deep convolutional neural network), a ResNet (Residual Network), a MobileNet (a lightweight network), and the like, which is not limited in the embodiments of the present disclosure.

[0045] Step S313, obtaining a preset face image, and determining a second image feature corresponding to the preset face image. For example, the second image feature can be extracted by a feature extraction model in the related art, which can be the same as or different from the feature extraction model used for extracting the first image feature, which is not limited in the embodiments of the present disclosure. The preset face image can include a front face image and a side face image of the user captured when the user model is established.

[0046] In a possible implementation, the preset face image is a front face image of the target user, and the second image feature is a feature map corresponding to the front face image. The front face image has a smaller adjustment range when adjusted to different side face postures, which can reduce the deformation degree in the adjustment process, and thus improve the authenticity of the adjusted second image feature, and is beneficial to improving the adjustment quality of the final model image.

[0047] At step S314, the second image feature is adjusted according to the visual information. For example, the second image feature can be adjusted according to the visual information, so that the adjusted second image feature can be similar to the image feature of the real image (or the actually photographed image) of the user under the visual information. For example, the mapping relationship between the multiple sets of head posture information and the pixel offset can be obtained, and then the position of the pixel in the second image feature is adjusted according to the mapping relationship. The present disclosure does not limit the specific adjustment manner between the visual information and the second image feature. The second image feature can represent that the real face image is offset according to the visual information.

[0048] In a possible implementation, step S314 can include: determining a pixel motion vector according to the visual information. For example, the visual information can include head posture information, expression information, and the like. Here, the head posture information is taken as an example. The head posture information can be represented as a deflection angle of the head. The deflection angle can correspond to a pixel motion vector. The corresponding relationship can be established by using multiple sets of deflection angles and pixel motion vectors, or by using other manners. The present disclosure does not limit the corresponding relationship. Then, the pixel motion vector can be directly obtained according to the corresponding relationship and the head posture information. Finally, the position of the pixel in the feature map is adjusted according to the pixel motion vector, and the adjusted second image feature is obtained. For example, the pixel motion vector can be mapped to a two-dimensional plane as a motion field by using related technologies, so as to adapt to the two-dimensional feature map (the number of channels of the feature map is not considered here), and the position of the pixel in the feature map is adjusted according to the motion field, so that the adjusted second image feature is obtained. The adjusted second image feature can represent the real image feature of the user under the specific visual information as a prediction result.

[0049] At step S315, the first image feature and the adjusted second image feature are fused to obtain a third image feature. For example, the image features can be fused by using add, concat, and the like in related technologies. The fused third image feature includes the image feature of the virtual model of the user adapted to the visual information, and the image feature of the actually photographed image of the user adapted to the visual information, so that the quality of the generated model image can be improved.

[0050] At step S316, the model image corresponding to the user model is adjusted and displayed according to the third image feature. For example, the step can be processed by using various deep learning decoders in related technologies, such as FCN (full convolutional neural network), Unet (a type of FCN), and the like. The present disclosure does not limit the deep learning decoders.

[0051] Exemplarily, the steps S312 to S316 can be processed by a GAN (Generative Adversarial Networks, generative model) model, for example, the GAN model can include a plurality of modules, such as a feature extraction module for extracting image features, a feature fusion module for performing feature fusion, a rendering module for adjusting model images, and the like. Developers can also add new modules according to actual needs, and the embodiments of the present disclosure do not limit this. Referring to Figure 3 , Figure 3 A reference schematic diagram of a model structure provided by an embodiment of the present disclosure is shown, as shown in Figure 3 The digital human picture (i.e., the model image described above), the vertex motion vector (i.e., the pixel point motion vector described above), and the pre-stored real human face picture (e.g., a front image of a real human face) (i.e., the preset face image described above) can be input into the model, and then the digital human image features (i.e., the first image features) and the real human face picture feature map (e.g., a front feature map of a real human face) (i.e., the feature map described above) are extracted. The real human face feature map with adjusted head angle and expression is obtained by adjusting the real human face picture feature map through the motion field corresponding to the fixed point motion vector. Finally, the digital human image features and the real human face feature map are fused through the add, concat, and the like, and a current face with high reality is rendered through the decoder. Each stage can be generated by the idea of GAN for adversarial training, that is, the generator is responsible for generating image features, feature fusion, face rendering, and the like and is input into the discriminator, and the discriminator is used to determine whether the obtained result is generated by the generator or is real existing data. Through multiple iterations of training, a generator with high reality can be obtained, and the generator is used as part of the GAN model used in the business process. The data used in the training process of the generator and the discriminator can be generated in batches by slicing the collected real speaker video, and the embodiments of the present disclosure do not limit this. The generator and the discriminator can be guided and trained by the loss function in the related art, for example, the pixel MSE loss (pixel mean square loss function), the pixel L1 loss (pixel mean absolute error loss), the real human face adversarial loss, and the like, and the embodiments of the present disclosure do not limit this.

[0052] In one possible implementation, the video call method further includes: acquiring input or preset special effects parameters during the video call, and adjusting the user model corresponding to the first terminal or the user model corresponding to the second terminal. The special effects parameters are used to adjust the display effect of the user model. For example, a user can adjust the special effects of their own user model or the user model corresponding to the user they are talking to, such as adding animation effects to achieve interactive effects, beautifying their user model, etc. This embodiment of the present disclosure does not impose any limitations. The aforementioned special effects parameters can be input by the user in real time, or saved by the user to the terminal's local storage or a server, so that the user model corresponding to the user can be automatically adjusted for each video call, reducing the number of manual adjustments required by the user. This embodiment of the present disclosure does not impose any limitations on the aforementioned special effects parameters, which may include skin smoothing, whitening, face slimming, background images, virtual clothing, etc., and developers can determine them according to actual needs. Furthermore, since the user model is pre-built in this embodiment of the present disclosure, it carries three-dimensional information and can be directly processed by the 3D rendering software mentioned above to achieve highly integrated program design, which also helps to reduce the time cost and data interaction cost of special effects adjustment.

[0053] This disclosure embodiment can adjust the user model on the decoding terminal by transmitting user feature information so that the head posture and expression of the user model are the same as or similar to the user's real posture, thereby achieving video calls with lower data interaction. In addition, the above video call method is less affected by bandwidth limitations, which is beneficial to improving the stability of the video call process.

[0054] See Figure 4 As shown, Figure 4 A flowchart of another video call method provided according to an embodiment of this disclosure is shown, such as... Figure 4 As shown in the embodiments of this disclosure, a video call method is also provided, applied to a second terminal that communicates with a first terminal. The video call method includes: step S600, in the case of a video call with the first terminal, acquiring an image or sound to be processed during the video call. Exemplarily, the image to be processed can be acquired using an image acquisition device of the second terminal, and the sound to be processed can be acquired using a sound acquisition device of the second terminal.

[0055] Step S700, determining the user feature information corresponding to the target user according to the to-be-processed image or the to-be-processed sound, and sending the user feature information to the first terminal. For example, the second terminal can perform image processing on the to-be-processed image by using a feature extraction model stored locally or in a server to obtain the visual feature described above and use the visual feature as the user feature information. In an example, the to-be-processed sound can be directly sent to the first terminal as the user feature information, and the first terminal can predict the visual feature. Alternatively, the second terminal can use the method of the first terminal described above to predict the visual feature, and then use the visual feature as the user feature information. The present embodiment does not limit the foregoing. For example, the user feature information corresponding to the to-be-processed image can be captured by using ARKit (an augmented reality tool) in a related technology to capture head posture and expression. The manner of obtaining the user feature information can refer to related technologies, and details are not described herein. In order to meet the real-time requirement of the video call, a faster feature extraction manner can be used to generate a result, and real-time interaction can be met.

[0056] In a possible implementation, the sending of the user feature information to the first terminal includes compressing the user feature information and sending the compressed user feature information to the first terminal. The present embodiment provides a compression manner for reference. The user feature information is mapped to an int8 space (only as an example, int32 and int64 can also be used), and then compressed by using an entropy coding algorithm (for example, Huffman coding or other coding manners) in a related technology, so as to obtain the compressed user feature information.

[0057] In combination with an actual application scenario, refer to Figure 5 , Figure 5 FIG. 1 shows a reference schematic diagram of a video call method according to an embodiment of the present disclosure. Figure 5As shown, the face 3D modeling process can be performed at any time before the video call, and a face picture of the speaker can be captured by the camera, and then a face 3D model is modeled. During the voice call, the encoding end (i.e., the encoding terminal described above) can capture the face posture and expression to obtain the posture (i.e., the head posture information described above) and the expression (i.e., the expression information described above), and then encode and compress the posture and expression, that is, the low code rate data (expression, head posture) can be transmitted to the decoding end (i.e., the decoding terminal described above). The decoding end can input the face mesh, the 3D format data (i.e., the user model described above) such as the skeleton, and the low code rate data into the 3D digital human rendering engine (i.e., the 3D rendering software described above) configured by the decoding end. The rendering engine inputs the digital human picture, the vertex motion vector into the model described above, and inputs the pre-stored real face picture used in the face 3D modeling stage into the model described above to obtain the rendered current face (for details, refer to the related description in the foregoing method embodiment). Figure 3

[0058] The quality of the video call method provided by the embodiments of the present disclosure is related to the high computing power of the decoding terminal and the low code rate, which can reduce the traffic cost. In addition, the embodiments of the present disclosure are beneficial to improving the realism of the finally rendered image by fusing the image features of the virtual model with the predicted features of the actual captured image of the user.

[0059] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Limited by the length, the present disclosure will not be repeated. Those skilled in the art can understand that in the above-mentioned method of the specific embodiment, the specific execution order of each step should be determined according to its function and possible internal logic.

[0060] In addition, the present disclosure also provides a video call device, an electronic device, a computer readable storage medium, and a program, all of which can be used to implement any one of the video call methods provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding description in the method part, and will not be repeated.

[0061] Referring to Figure 6 As shown, Figure 6 a block diagram of a video call device according to an embodiment of the present disclosure is shown, like Figure 6 ​As shown, the video call device 100 comprises a user model obtaining module 110, configured to obtain a user model corresponding to the second terminal; wherein the user model is configured to represent a three-dimensional model corresponding to a target user performing a video call through the second terminal; a feature information obtaining module 120, configured to receive user feature information or related information of the user feature information sent by the second terminal during a video call; and a model adjusting module 130, configured to adjust and display a model image corresponding to the user model according to the user feature information or the related information.

[0062] In a possible implementation, the user feature information comprises sound information corresponding to the target user, and the adjusting and displaying of the model image corresponding to the user model according to the user feature information comprises: determining visual information corresponding to the target user according to the sound information corresponding to the target user; and adjusting and displaying the model image corresponding to the user model according to the visual information.

[0063] In a possible implementation, the user feature information comprises visual information corresponding to the target user, and the adjusting and displaying of the model image corresponding to the user model according to the user feature information comprises: adjusting and displaying the model image corresponding to the user model according to the visual information.

[0064] In a possible implementation, the adjusting and displaying of the model image corresponding to the user model according to the visual information comprises: adjusting the user model according to the visual information; generating a model image corresponding to the adjusted user model according to the adjusted user model, and determining a first image feature corresponding to the model image; obtaining a preset face image, and determining a second image feature corresponding to the preset face image; adjusting the second image feature according to the visual information; performing feature fusion on the first image feature and the adjusted second image feature to obtain a third image feature; and adjusting and displaying the model image corresponding to the user model according to the third image feature.

[0065] In a possible implementation, the preset face image is a front face image of the target user, and the second image feature is a feature map corresponding to the front face image; the adjusting of the second image feature according to the visual information comprises: determining a pixel point motion vector according to the visual information; and adjusting the positions of pixel points in the feature map according to the pixel point motion vector, and taking the adjusted pixel points as the adjusted second image feature.

[0066] In a possible implementation, the receiving the user feature information sent by the second terminal during the video call process comprises: receiving, during the video call process, user feature information compressed by the second terminal and sent by the second terminal; and the adjusting and displaying the model image corresponding to the user model according to the user feature information comprises: decompressing the compressed user feature information, and adjusting and displaying the model image corresponding to the user model according to the decompressed user feature information.

[0067] In a possible implementation, the video call apparatus further comprises a display effect adjustment module configured to perform at least one of the following steps: obtaining a special effect parameter input or preset during the video call process, and adjusting the user model corresponding to the first terminal or the user model corresponding to the second terminal; wherein the special effect parameter is used to adjust the display effect of the user model.

[0068] In a possible implementation, the visual information comprises at least one of head posture information or expression information of the target user.

[0069] Referring to Figure 7 as shown, Figure 7 a block diagram of another video call apparatus provided by an embodiment of the present disclosure is shown, as Figure 7 As shown in FIG. 2, the video call apparatus 200 comprises: a data acquisition module 210 configured to acquire a to-be-processed image or a to-be-processed sound during a video call process when the video call is performed with the first terminal; and an information sending module 220 configured to determine user feature information corresponding to a target user according to the to-be-processed image, and send the user feature information to the first terminal, or send the to-be-processed sound as related information of the user feature information to the first terminal.

[0070] In a possible implementation, the sending the user feature information to the first terminal comprises: compressing the user feature information, and sending the compressed user feature information to the first terminal.

[0071] The method has specific technical correlation with the internal structure of the computer system, and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing data storage, reducing data transmission, improving hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system in accordance with the natural law.

[0072] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or contains modules which can be used to execute the methods described in the above method embodiments, and the specific implementation can be referred to the description of the above method embodiments. For brevity, it will not be described here again.

[0073] The embodiments of the present disclosure further provide a computer readable storage medium having stored thereon computer program instructions, which when executed by a processor, implement the method described above. The computer readable storage medium can be a volatile or non-volatile computer readable storage medium.

[0074] The embodiments of the present disclosure further provide an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to invoke the instructions stored by the memory to execute the method described above.

[0075] The embodiments of the present disclosure further provide a computer program product, comprising computer readable code, or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in a processor of an electronic device, the processor in the electronic device executes the method described above.

[0076] The electronic device can be provided as a terminal device or other forms of devices.

[0077] Figure 8 A block diagram of an electronic device 800 is shown according to an embodiment of the present disclosure. For example, the electronic device 800 can be a terminal device, such as a User Equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a Personal Digital Assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc.

[0078] Referring to Figure 8 The electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0079] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with displaying, making phone calls, data communications, camera operations and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the methods described above. In addition, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0080] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or nonvolatile memory, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disc, or optical disc.

[0081] The power supply component 806 supplies power for various components of the electronic device 800. The power supply component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0082] The multimedia component 808 includes a screen providing an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a back camera. The front camera and / or the back camera can receive external multimedia data when the electronic device 800 is in an operation mode, such as a photographing mode or a video mode. Each of the front camera and the back camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0083] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.

[0084] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0085] The sensor component 814 includes one or more sensors for providing status assessments for various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration / g-force and temperature of the electronic device 800. The sensor component 814 can include an optical sensor for detecting ambient light, a proximity sensor for detecting nearby objects without any physical touch, a light sensor such as a complementary metal-oxide semiconductor (CMOS) or charge coupled device (CCD) image sensor for use in imaging applications, or a combination thereof. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0086] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a corresponding communication standard, such as wireless fidelity (Wi-Fi), second generation (2G) cellular technology, third generation (3G) cellular technology, fourth generation (4G) cellular technology, long term evolution (LTE) of universal mobile telecommunications technology, fifth generation (5G) cellular technology, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 can further include a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technology.

[0087] In an example embodiment, the electronic device 800 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements for performing the above-described methods.

[0088] In an example embodiment, a non-transitory computer-readable storage medium, such as the memory 804 including computer program instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to implement the above-described methods.

[0089] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0090] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0091] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0092] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0093] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0094] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0095] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0096] The flow diagrams and the block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0097] The computer program product can be embodied in a tangible medium of

[0098] The above description of the various embodiments is intended to be illustrative in all aspects, rather than being restrictive. Those skilled in the art can refer to the description of the various embodiments to make modifications and / or improvements.

[0099] Those skilled in the art can understand that, in the above-described method of the specific embodiments, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible inherent logic.

[0100] If the technical solutions of the present application involve personal information, the product applying the technical solutions of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solutions of the present application involve sensitive personal information, the product applying the technical solutions of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his personal information, the individual's authorization is obtained under the condition that the device uses obvious mark / information to inform the individual of the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.

[0101] The above has described various embodiments of the present disclosure, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical application or improvement of technology in the market of the embodiments, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.

Claims

1. A video call method, characterized in that, The video call method is applied to a first terminal, which communicates with a second terminal, and includes: Obtain the user model corresponding to the second terminal; wherein, the user model is used to represent the three-dimensional model of the target user who is making a video call through the second terminal; During a video call, user feature information sent by the second terminal is received. The user feature information includes voice information or visual information corresponding to the target user. The voice information is used to determine the visual information. Adjusting and displaying the model image corresponding to the user model based on the user feature information includes: A first image feature corresponding to the model image is determined, wherein the model image is the adjusted model image corresponding to the user model, and the user model is adjusted according to the visual information. A preset facial image is acquired, and a second image feature corresponding to the preset facial image is determined. The preset facial image is a frontal image of the user's real face collected during the establishment of the user model. Based on the visual information, the second image feature is adjusted, wherein the position of the pixel in the second image feature is adjusted by the motion field corresponding to the pixel motion vector determined by the visual information, so that the adjusted second image feature represents that the frontal face image has been offset according to the visual information, and the adjusted second image feature represents the user's real image features under the visual information. The first image features and the adjusted second image features are fused to obtain the third image features; Based on the third image features, the model image corresponding to the user model is adjusted and displayed to achieve real-time video calls.

2. The video call method as described in claim 1, characterized in that, The user feature information includes the voice information corresponding to the target user. Adjusting and displaying the model image corresponding to the user model based on the user feature information includes: Based on the voice information corresponding to the target user, determine the visual information corresponding to the target user; Based on the visual information, adjust and display the model image corresponding to the user model.

3. The video call method as described in claim 1, characterized in that, The user feature information includes visual information corresponding to the target user, and adjusting and displaying the model image corresponding to the user model based on the user feature information includes: Based on the visual information, adjust and display the model image corresponding to the user model.

4. The video call method as described in claim 2 or 3, characterized in that, Based on the visual information, adjusting and displaying the model image corresponding to the user model includes: Adjust the user model based on the visual information; Based on the adjusted user model, generate a model image corresponding to the adjusted user model, and determine the first image feature corresponding to the model image; Obtain a preset facial image and determine the second image features corresponding to the preset facial image; The second image features are adjusted based on the visual information; The first image features and the adjusted second image features are fused to obtain the third image features; Based on the third image features, adjust and display the model image corresponding to the user model.

5. The video call method as described in claim 4, characterized in that, The preset facial image is a frontal image of the target user, and the second image feature is a feature map corresponding to the frontal image; the adjustment of the second image feature based on the visual information includes: Based on the visual information, determine the pixel motion vector; Based on the pixel motion vector, the position of the pixel in the feature map is adjusted and used as the adjusted second image feature.

6. The video call method as described in claim 1, characterized in that, The step of receiving user feature information sent by the second terminal during a video call includes: receiving user feature information compressed by the second terminal during a video call. The step of adjusting and displaying the model image corresponding to the user model based on the user feature information includes: The compressed user feature information is decompressed, and the model image corresponding to the user model is adjusted and displayed based on the decompressed user feature information.

7. The video call method as described in claim 1, characterized in that, The video call method also includes: During a video call, input or preset special effects parameters are obtained, and the user model corresponding to the first terminal or the second terminal is adjusted accordingly; wherein, the special effects parameters are used to adjust the display effect of the user model.

8. The video call method as described in claim 2, characterized in that, The visual information includes at least one of the following: head posture information and facial expression information corresponding to the target user.

9. A video call method, characterized in that, The video call method is applied to a second terminal, which communicates with the first terminal, and includes: In the case of a video call with the first terminal, during the video call, the image or sound to be processed is captured; Based on the image or sound to be processed, user feature information corresponding to the target user is determined, and the user feature information is sent to the first terminal. The user feature information includes sound information or visual information corresponding to the target user, wherein the sound information is used to determine the visual information; the user feature information is used to adjust and display the model image corresponding to the user model; adjusting and displaying the model image corresponding to the user model includes: A first image feature corresponding to the model image is determined, wherein the model image is the adjusted model image corresponding to the user model, and the user model is adjusted according to the visual information. A preset facial image is acquired, and a second image feature corresponding to the preset facial image is determined. The preset facial image is a frontal image of the user's real face collected during the establishment of the user model. Based on the visual information, the second image feature is adjusted, wherein the position of the pixel in the second image feature is adjusted by the motion field corresponding to the pixel motion vector determined by the visual information, so that the adjusted second image feature represents that the frontal face image has been offset according to the visual information, and the adjusted second image feature represents the user's real image features under the visual information. The first image features and the adjusted second image features are fused to obtain the third image features; Based on the third image features, the model image corresponding to the user model is adjusted and displayed to achieve real-time video calls.

10. The video call method as described in claim 9, characterized in that, Sending the user feature information to the first terminal includes: compressing the user feature information and sending it to the first terminal.

11. A video call device, characterized in that, The video calling device communicates with the second terminal, and the video calling device includes: The user model acquisition module is used to acquire the user model corresponding to the second terminal; wherein, the user model is used to represent the three-dimensional model of the target user who is making a video call through the second terminal; The feature information acquisition module is used to receive user feature information or related information of user feature information sent by the second terminal during a video call. The user feature information includes voice information or visual information corresponding to the target user. The voice information is used to determine the visual information. The model adjustment module is used to adjust and display the model image corresponding to the user model based on the user feature information or the relevant information, including: A first image feature corresponding to the model image is determined, wherein the model image is the adjusted model image corresponding to the user model, and the user model is adjusted according to the visual information. A preset facial image is acquired, and a second image feature corresponding to the preset facial image is determined. The preset facial image is a frontal image of the user's real face collected during the establishment of the user model. Based on the visual information, the second image feature is adjusted, wherein the position of the pixel in the second image feature is adjusted by the motion field corresponding to the pixel motion vector determined by the visual information, so that the adjusted second image feature represents that the frontal face image has been offset according to the visual information, and the adjusted second image feature represents the user's real image features under the visual information. The first image features and the adjusted second image features are fused to obtain the third image features; Based on the third image features, the model image corresponding to the user model is adjusted and displayed to achieve real-time video calls.

12. A video call device, characterized in that, The video calling device communicates with the first terminal, and the video calling device includes: The data acquisition module is used to acquire images or sounds to be processed during a video call with the first terminal. An information sending module is used to determine user feature information corresponding to a target user based on the image to be processed, and send the user feature information to the first terminal, or to use the sound to be processed as relevant information of user feature information and send the relevant information to the first terminal. The user feature information includes sound information or visual information corresponding to the target user, and the sound information is used to determine the visual information. The user feature information is used to adjust and display the model image corresponding to the user model. The adjustment and display of the model image corresponding to the user model includes: A first image feature corresponding to the model image is determined, wherein the model image is the adjusted model image corresponding to the user model, and the user model is adjusted according to the visual information. A preset facial image is acquired, and a second image feature corresponding to the preset facial image is determined. The preset facial image is a frontal image of the user's real face collected during the establishment of the user model. Based on the visual information, the second image feature is adjusted, wherein the position of the pixel in the second image feature is adjusted by the motion field corresponding to the pixel motion vector determined by the visual information, so that the adjusted second image feature represents that the frontal face image has been offset according to the visual information, and the adjusted second image feature represents the user's real image features under the visual information. The first image features and the adjusted second image features are fused to obtain the third image features; Based on the third image features, the model image corresponding to the user model is adjusted and displayed to achieve real-time video calls.

13. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the video call method according to any one of claims 1 to 10.

14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the video call method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Video call method and terminal device

    CN108881782A

  • Call method and device, terminal and storage medium

    CN110536095A

  • Social method, device and system, terminal equipment and storage medium

    CN110599359A