Virtual image processing method and device, equipment and medium

By obtaining media information of the target object to generate personalized virtual images, the problem of homogeneity of virtual images is solved, the user experience is improved and its application in media processing tasks is promoted.

CN120374802APending Publication Date: 2025-07-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410103100.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The virtual image in the existing market is seriously homogenized, with poor user experience and difficult to widely use.

Method used

By obtaining the media information of the target object, a personalized target virtual image is generated and saved. The characteristics of the target virtual image correspond to the target object and serve as material information and/or association information for the target media processing task.

Benefits of technology

It realizes personalized customization of virtual images, avoids similarities, improves user experience, and promotes the widespread application of virtual images in media processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374802A_ABST
    Figure CN120374802A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a virtual image processing method and device, equipment and a medium. The method comprises the steps that media information of a target object is acquired; obtaining a target virtual image corresponding to the target object based on the media information of the target object; wherein the target feature of the target virtual image corresponds to the target feature of the target object; storing the target virtual image of the target object; wherein the target virtual image is used as material information and / or associated information of the target media processing task. According to the embodiment of the invention, the personalized customization effect of the virtual image of the target object is realized, the situation that the virtual image is similar is effectively avoided, the target virtual image can be used as a material or associated information to be flexibly applied to a media processing task, and the virtual image can be promoted to be widely applied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method, apparatus, device, and medium for processing virtual avatars. Background Art

[0002] With the development of computer technology, virtual avatars such as digital humans have gradually become popular and are somewhat attractive to users. However, the inventors have found through research that virtual avatars on the existing market are highly homogenized and are basically platform assets, that is, users can only use a very limited number of fixed virtual avatars provided by the platform, resulting in a poor experience and making it difficult for virtual avatars to be widely applied. Summary of the Invention

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and medium for processing virtual avatars.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for processing a virtual avatar, the method including: obtaining media information of a target object; obtaining a target virtual avatar corresponding to the target object based on the media information of the target object, where target features of the target virtual avatar correspond to target features of the target object; and saving the target virtual avatar of the target object, where the target virtual avatar is used as material information and / or associated information for a target media processing task.

[0005] Optionally, the obtaining media information of the target object includes: collecting an image and / or video including the target object, and collecting audio of the target object when a virtual avatar customization request for the target object is received and the media information collection permission of the target object is obtained.

[0006] Optionally, the obtaining the target virtual avatar corresponding to the target object based on the media information of the target object includes: generating and displaying at least one candidate virtual avatar corresponding to the target object based on the media information of the target object; and obtaining the target virtual avatar corresponding to the target object based on the candidate virtual avatar selected by the first selection operation in response to the first selection operation for the at least one candidate virtual avatar.

[0007] Optionally, obtaining the target virtual image corresponding to the target object based on the candidate virtual image selected by the first selection operation includes: in response to an editing request for the candidate virtual image selected by the first selection operation, presenting a plurality of editing items; where different editing items are used to edit different virtual image features; in response to triggering of a target editing item among the plurality of editing items, performing feature editing on the candidate virtual image selected by the first selection operation based on the target editing item to obtain the target virtual image corresponding to the target object.

[0008] Optionally, the target media processing task includes a template synthesis task; the method further includes: when starting the template synthesis task, presenting a plurality of virtual image templates; where the virtual image templates are used to present media information of sample virtual images; in response to triggering of a target template among the plurality of virtual image templates, performing a generation process based on the target virtual image and the target template to obtain media information corresponding to the target virtual image.

[0009] Optionally, the target media processing task includes a media editing task; the method further includes: when starting the media editing task, obtaining media information to be edited, and obtaining an associated virtual image corresponding to the media information to be edited; identifying an object to be edited corresponding to the associated virtual image from the media information to be edited; performing an editing process on the object to be edited in the media information to be edited to obtain edited media information.

[0010] Optionally, obtaining the associated virtual image corresponding to the media information to be edited includes: presenting a first virtual image list; the first virtual image list presents at least one saved target virtual image; where different target virtual images correspond to different objects; in response to a second selection operation for a target virtual image in the first virtual image list, using the target virtual image selected by the second selection operation as the associated virtual image corresponding to the media information to be edited.

[0011] Optionally, the target virtual images in the virtual image list include a first virtual image and a second virtual image; where the first virtual image is the target virtual image of the target object, and the second virtual image is the target virtual image corresponding to an associated object of the target object.

[0012] Optionally, the method further includes: when it is monitored that there is a second virtual image in the virtual image list that meets a preset condition, setting the second virtual image that meets the preset condition to a non-selectable state, or deleting the second virtual image that meets the preset condition from the virtual image list.

[0013] Optionally, the preset condition includes one or more of the following: the duration of adding to the virtual image list is higher than a preset duration threshold, the image features are changed, and a cancellation association instruction is received.

[0014] Optionally, the target media processing task includes a media synthesis task; the method further includes: when starting the media synthesis task, collecting a media information stream and displaying the media information stream on an information preview interface; displaying a second virtual image list; the second virtual image list presents at least one saved target virtual image; wherein, different target virtual images correspond to different objects; in response to a third selection operation on a target virtual image in the second virtual image list, merging the target virtual image selected by the third selection operation with the media information stream and displaying them on the information preview interface.

[0015] Optionally, merging the target virtual image selected by the third selection operation with the media information stream and displaying them on the information preview interface includes: obtaining a target synthesis position of the target virtual image selected by the third selection operation in the media information stream; based on the target synthesis position, merging the target virtual image selected by the third selection operation with the media information stream and displaying them on the information preview interface.

[0016] Optionally, obtaining the target synthesis position of the target virtual image selected by the third selection operation in the media information stream includes: determining the target synthesis position of the target virtual image selected by the third selection operation based on the positions of the objects included in the media information stream; or, in response to detecting a position specifying operation, using the position corresponding to the position specifying operation as the target synthesis position of the target virtual image selected by the third selection operation.

[0017] Optionally, based on the target synthesis position, merging the target virtual image selected by the third selection operation with the media information stream and displaying them on the information preview interface includes: when the object in the media information stream is a full-body image and the target virtual image selected by the third selection operation is a half-body image, performing a completion process on the target virtual image selected by the third selection operation to obtain a completed target virtual image; based on the target synthesis position, merging the completed target virtual image with the media information stream and displaying them on the information preview interface.

[0018] Optionally, the target media processing task includes a conversation task; the method further includes: when starting the conversation task, obtaining a conversation form and presenting a third virtual image list; wherein, the conversation form includes a text form and / or a voice form; the third virtual image list presents at least one saved target virtual image; wherein, different target virtual images correspond to different objects; in response to a fourth selection operation on a target virtual image in the third virtual image list, using the target virtual image selected by the fourth selection operation as the current conversation image; obtaining target reference information to generate output information of the current conversation image based on the target reference information; wherein, the target reference information includes identity description information of the current conversation image and / or received user input information; presenting the output information of the current conversation image based on the conversation form.

[0019] Optionally, presenting the output information of the current conversation image based on the conversation form includes: when the conversation form is a voice form, using the current conversation image to express the output information of the current conversation image with a target voice feature and a target expression feature; wherein, the target voice feature is determined based on the voice feature of the object corresponding to the current conversation image, and the target expression feature is determined based on the expression feature of the object corresponding to the current conversation image.

[0020] In a second aspect, an embodiment of the present disclosure provides a virtual image processing device, including: an information acquisition module, configured to acquire media information of a target object; an image acquisition module, configured to acquire a target virtual image corresponding to the target object based on the media information of the target object; wherein, the target feature of the target virtual image corresponds to the target feature of the target object; an image storage module, configured to store the target virtual image of the target object; wherein, the target virtual image is used as material information and / or associated information for a target media processing task.

[0021] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device includes: a processor; a memory for storing executable instructions executable by the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the virtual image processing method provided by the embodiment of the present disclosure.

[0022] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to execute the virtual image processing method provided by the embodiment of the present disclosure.

[0023] The above technical solution provided by the embodiment of the present disclosure can obtain the target virtual image corresponding to the target object based on the media information of the target object, and then save the target virtual image to be used as the material information and / or associated information of the target media processing task. The target virtual image obtained by the above method has corresponding target features with the target object, thereby achieving the personalized customization effect of the virtual image of the target object. The user can customize the personalized virtual image according to the needs, which can effectively avoid the situation where the virtual image is the same, and the target virtual image can also be flexibly applied to the media processing task as the material or associated information, which helps to promote the widespread application of virtual images.

[0024] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0027] Figure 1 A flowchart of a method for processing a virtual image provided by an embodiment of the present disclosure;

[0028] Figure 2 A schematic diagram of a processing flow of a virtual image provided by an embodiment of the present disclosure;

[0029] Figure 3 A schematic diagram of a processing flow of a virtual image provided by an embodiment of the present disclosure;

[0030] Figure 4 A schematic diagram of the structure of a virtual image processing device provided by an embodiment of the present disclosure;

[0031] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0033] Numerous specific details are set forth in the following description to facilitate a full understanding of the present disclosure. However, the present disclosure may also be implemented in other ways different from those described herein. Obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0034] Figure 1 As shown in the flowchart of a method for processing a virtual image provided by an embodiment of the present disclosure, this method can be executed by a virtual image processing device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, this method mainly includes the following steps S102 to S106:

[0035] Step S102, obtaining media information of a target object.

[0036] The embodiments of the present disclosure do not limit the target object. For example, the target object can be a person or an animal, etc. In some implementation examples, the media information includes images and / or videos. Further, the media information can also include audio. That is, the obtained media information of the target object can be represented in a single media form such as an image or a video, or can be represented in a multimedia form such as an image, a video, and audio at the same time.

[0037] In some implementation manners, when a virtual image customization request for a target object is received and the permission to collect the media information of the target object is obtained, an image and / or video including the target object is collected, and the audio of the target object is collected. In some implementation examples, the target object can be a user who hopes to customize his own virtual image. The user can initiate a virtual image customization request through a user terminal such as a mobile phone. In practical applications, the user terminal can provide a trigger button for initiating a virtual image customization request on the user interaction interface. When it is detected that the trigger button is triggered, it is confirmed that the virtual image customization request of the target object is received. In addition, a prompt content for collecting the media information of the target object can be displayed on the user interaction interface, and an authorization control or a cancellation control is provided. When it is detected that the authorization control is triggered, it is confirmed that the permission to collect the media information of the target object is obtained. When a virtual image customization request for a target object is received and the permission to collect the media information of the target object is obtained, media collection devices such as the front camera of the user terminal can be turned on to collect an image and / or video of the user. For example, a user face image and / or a video including the user's face or the user's whole body can be collected. Further, the audio of the user can also be recorded. For example, the user can be allowed to speak randomly or read a specified text, so as to obtain the audio information of the user. In addition, in practical applications, if the video of the collected target object includes the voice of the target object, there is no need to collect the audio of the target object additionally.

[0038] Step S104: Obtain the target virtual image corresponding to the target object based on the media information of the target object; wherein, the target features of the target virtual image correspond to the target features of the target object. Specifically, the target features of the target virtual image are the same as or similar to the target features of the target object (such as the similarity is higher than a preset threshold).

[0039] The above-mentioned target features include but are not limited to facial features, body shape features, voice features, motion features, etc. These features can be further divided. For example, facial features can be specifically divided into facial feature features, and voice features can be specifically divided into timbre features, speech rate features, etc. There is no limitation here. In practical applications, any feature that hopes the target virtual image to be similar to the target object can be set as the target feature, and there is no limitation here.

[0040] In the embodiments of the present disclosure, the target features of the target object can be extracted based on the media information of the target object. The specific feature extraction algorithm used is not limited here; then, based on the target features of the target object, the target virtual image corresponding to the target object is generated; in other words, the target features of the target virtual image are determined based on the target features of the target object, and there is a strong correlation, which can effectively strengthen the connection between the target object and the target virtual image. For example, the target virtual image can be a 3D virtual image presenting the facial features and body shape features of the target object, and the voice emitted by the target virtual image can present the timbre features, speech rate features, etc. of the target object, so as to obtain a personalized customized image of the target object.

[0041] Step S106: Save the target virtual image of the target object; wherein, the target virtual image is used as the material information and / or associated information for the target media processing task. The embodiments of the present disclosure do not limit the target media processing task. Any processing event containing media information such as images and videos and requiring the use of the target virtual image can be used as the target media processing task. The target virtual image can be used as the material in the target media processing task. For example, it can be added to an image or video as the material; the target virtual image can also be used as the associated information (also called reference information) of the target media processing task, that is, the target media processing task needs to perform specific operations based on the target virtual image. For example, if the target media processing task is a photo retouching task, only the object corresponding to the target virtual image included in the image to be edited is beautified, and other objects remain unchanged. The embodiments of the present disclosure do not limit the specific application method of the target virtual image in the target media processing task.

[0042] The target virtual image obtained by the above method has the same or similar target features as the target object, thereby achieving a personalized customization effect of the virtual image of the target object. Users can customize personalized virtual images according to their needs, which can effectively avoid the situation where virtual images are similar. The target virtual image can also be flexibly applied to media processing tasks as material or related information, which is helpful to promote the widespread application of virtual images.

[0043] In some implementations, the above step S104, i.e., obtaining a target virtual image corresponding to the target object based on the media information of the target object, can be performed with reference to the following steps 1 and 2:

[0044] Step 1: Generate and display at least one candidate virtual image corresponding to the target object based on the media information of the target object. In practical applications, target features can be extracted based on the media information of the target object, and a neural network model such as a diffusion model can be used to generate at least one candidate virtual image with the target features. Different candidate virtual images have certain style differences, such as generating a specified number of real images, 3D images or cartoon images, etc., for users to choose from.

[0045] Step 2, in response to a first selection operation for at least one candidate virtual image, obtain a target virtual image corresponding to the target object based on the candidate virtual image selected by the first selection operation. In actual applications, the user can select the desired candidate virtual image through the first selection operation. The first selection operation can be, for example, a click, a long press, triggering a corresponding selection control, etc., which are not limited here. In some embodiments, the selected candidate virtual image can be directly used as the target virtual image corresponding to the target object. In other embodiments, the selected candidate virtual image can also be edited to obtain the target virtual image corresponding to the target object. Exemplarily, step 2 can be performed with reference to the following steps 2.1 and 2.2:

[0046] Step 2.1, in response to an edit request for the candidate avatar selected by the first selection operation, display multiple edit items; wherein different edit items are used to edit different avatar features. For example, the multiple edit items include hair edit items, eye edit items, nose edit items and other facial features edit items, body shape edit items, skin color edit items, etc. Each edit item can be used to edit one or more avatar features, such as a hair edit item can be used to edit hairstyle features and hair color features. Specifically, multiple hairstyle items or hair color items can be provided for user selection.

[0047] Step 2.2, in response to the triggering of a target editing item among a plurality of editing items, perform feature editing on the candidate virtual image selected by the first selection operation based on the target editing item to obtain a target virtual image corresponding to the target object. The user can select the target editing item for which feature editing is required according to the needs, so as to perform feature editing on the selected candidate virtual image by using the target editing item, so as to flexibly change the virtual image automatically generated by the neural network model according to the needs, and further improve the personalization effect of the target virtual image.

[0048] See Figure 2 The schematic diagram of a processing flow of a virtual image shown, which schematically shows the acquisition method and application method of the target virtual image. The main processes include: receiving a customization request for the virtual image, sampling facial information (for extracting facial features), sampling voice information (for extracting features such as timbre and speech rate), generating a plurality of candidate virtual images, and editing the selected candidate virtual image to generate a target virtual image. The target virtual image can be used as material information and / or associated information in a target media processing task. Among them, the above-mentioned sampling of facial information can be realized by collecting images or videos through a camera, and the above-mentioned sampling of voice information can be realized by recording audio through a recorder. The method of generating a virtual image can be realized by a neural network such as a pre-trained generation model, and the target media processing task can be flexibly set according to the needs.

[0049] See Figure 3 The schematic diagram of a processing flow of a virtual image shown further schematically shows that the target media processing task can be a template synthesis task, a media editing task, a media synthesis task, or a dialogue task. Among them, the target virtual image can be used as material information in the template synthesis task, the media synthesis task, and the dialogue task, and can also be used as associated information in the media editing task, that is, it is necessary to refer to the target virtual image to edit the media information to be edited. It should be noted that Figure 3 These are only several exemplary descriptions of the target media processing task and should not be regarded as limitations.

[0050] For the convenience of understanding, the following will be elaborated in detail for different target media processing tasks respectively.

[0051] In some embodiments, the target media processing task includes a template synthesis task; the method for processing a virtual image provided by the embodiments of the present disclosure further includes steps (1) and (2):

[0052] Step (1), when starting a template synthesis task, display multiple virtual image templates; wherein, the virtual image templates are used to present the media information of sample virtual images, and the media information of the sample virtual images can be an image containing the sample virtual image, or a dynamic video clip or animated GIF of the sample virtual image. In practical applications, multiple virtual image templates can be displayed in a preset manner (such as side-by-side display in the form of cards, dynamic switching display, etc.), and the template display form is not limited herein. By displaying multiple virtual image templates, users can clearly and intuitively know the existing templates currently.

[0053] Step (2), in response to a target template among the multiple virtual image templates being triggered, perform a generation process based on the target virtual image and the target template to obtain the media information corresponding to the target virtual image. Users can trigger the required target template according to their needs. On this basis, the target virtual image and the target template can be directly subjected to a generation process, and this generation process can be, for example, synthesizing the target virtual image and the target template, or generating media information with a target style based on the target virtual image and the target template. Thus, the media information corresponding to the target virtual image is generated, and the form of the media information corresponding to the target virtual image is the same as that of the media information of the sample virtual image. For example, if the media information of the sample virtual image is an image, the generated media information corresponding to the target virtual image is also an image; and if the media information of the sample virtual image is a short video, the generated media information corresponding to the target virtual image is also a short video.

[0054] The embodiments of the present disclosure do not limit the above generation processing method. In some specific implementation examples, performing a generation process based on the target virtual image and the target template includes: replacing the specified features of the sample virtual image in the target template with the specified features of the target virtual image. For example, the specified features can be one or more of facial features, body type features, and clothing features. Through the above method, the target virtual image can efficiently and conveniently achieve the effect presented by the target template. If the target template is an animated GIF of a girl (sample virtual image) turning her head, the facial features of the target virtual image can be used to replace the facial features of the girl in the target template, and finally the media information of the generated target virtual image is the animated GIF of the target virtual image turning its head.

[0055] In some implementation manners, the target media processing task includes a media editing task, which can specifically be an image editing task (such as a photo retouching task) or a video editing task, etc. On this basis, the method for processing virtual images provided by the embodiments of the present disclosure further includes the following steps a to c:

[0056] Step a, when starting a media editing task, obtain the media information to be edited and the associated virtual avatar corresponding to the media information to be edited. The media information to be edited can be an image or a video, which is not restricted here. In practical applications, users can download the media information to be edited through the network or select the media information to be edited from the local media library. The associated virtual avatar corresponding to the media information to be edited is the virtual avatar required for reference when performing the media editing task. For ease of understanding, the following is an exemplary description of the process of obtaining the associated virtual avatar:

[0057] In some specific implementation examples, obtaining the associated virtual avatar corresponding to the media information to be edited in step a includes: presenting a first virtual avatar list; the first virtual avatar list presents at least one saved target virtual avatar; wherein, different target virtual avatars correspond to different objects; in response to a second selection operation on the target virtual avatar in the first virtual avatar list, the target virtual avatar selected by the second selection operation is used as the associated virtual avatar corresponding to the media information to be edited.

[0058] In practical applications, the first virtual avatar list can be presented to the user on the interface, displaying one or more target virtual avatars. In some specific implementation examples, the target virtual avatars in the virtual avatar list include a first virtual avatar and a second virtual avatar; wherein, the first virtual avatar is the target virtual avatar of the target object, and the second virtual avatar is the target virtual avatar corresponding to the associated object of the target object. For example, the first virtual avatar can be a virtual avatar customized by the user himself / herself, and the second virtual avatar can be a virtual avatar shared by the user's associated object (such as a friend). The user can select the target virtual avatar as the associated virtual avatar according to the needs.

[0059] Further, to effectively guarantee the application experience of virtual avatars, the above method further includes: when it is detected that there is a second virtual avatar in the virtual avatar list that meets the preset conditions, setting the second virtual avatar that meets the preset conditions to a non-selectable state, or deleting the second virtual avatar that meets the preset conditions from the virtual avatar list. When the second virtual avatar is set to the non-selectable state, the virtual avatar cannot be triggered as an associated virtual object. Exemplarily, the preset conditions include one or more of the following: the duration of being added to the virtual avatar list is higher than the preset duration threshold, the image characteristics change, and a cancellation association instruction is received. That is, the second virtual avatar whose duration of being added to the virtual avatar list exceeds the preset duration threshold (such as six months) can be set to the non-selectable state or deleted, the second virtual avatar with changed image characteristics can be set to the non-selectable state or deleted, and the second virtual avatar that receives the cancellation association instruction can be set to the non-selectable state or deleted. The cancellation association instruction may come from the user himself or from the object corresponding to the second virtual avatar (such as the user's friend). By the above method, the effectiveness of the currently selectable virtual avatars can be effectively guaranteed.

[0060] Step b: Identify the object to be edited corresponding to the associated virtual avatar from the media information to be edited.

[0061] Taking the media information to be edited as an image as an example, assume that the image includes person A, person B, and person C, and the associated virtual avatars are the target virtual avatars corresponding to person A and person C respectively. Then, based on the obtained associated virtual avatars, person A and person C can be identified from the image as the objects to be edited.

[0062] Step c: Perform editing processing on the object to be edited in the media information to be edited to obtain the edited media information.

[0063] The embodiments of the present disclosure do not limit the specific manner of the editing process. For example, it can be beauty processing, body posture adjustment processing, special effect adding processing, etc. By the above method, the object to be edited can be conveniently and quickly identified based on the selected virtual avatar, which not only improves the media editing efficiency but also helps to better promote interaction between users. For example, users can quickly identify their friends from the image and help their friends retouch the pictures by receiving the virtual avatars of their friends. Users can also share their own virtual avatars with their friends so that their friends can help them retouch the pictures. The above method greatly improves the convenience and interest of media editing.

[0064] In some specific implementation examples, the target media processing task includes a media synthesis task; the method for processing virtual avatars provided by the embodiments of the present disclosure further includes steps A to C:

[0065] Step A, when the media synthesis task is started, the media information stream is collected and displayed on the information preview interface. The media information stream may be a video stream, for example, when the media synthesis task is detected to be triggered, the information preview interface is displayed, and the camera is called to collect the video stream in real time, so that the captured video stream is displayed on the information preview interface, and the user can adjust the shooting according to the content on the information preview interface.

[0066] Step B, displaying a second virtual image list; the second virtual image list presents at least one saved target virtual image; wherein different target virtual images correspond to different objects. It should be noted that the second virtual image list may be the same as or different from the first virtual image list, and relational terms such as "first" and "second" in the disclosed embodiment are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Similarly, the second virtual image list may also include the user's own virtual image and the virtual image shared by the user's friends. For details, please refer to the relevant description of the aforementioned first virtual image list, which will not be repeated here.

[0067] Step C, in response to the third selection operation for the target virtual image in the second virtual image list, the target virtual image selected by the third selection operation is combined with the media information stream and displayed on the information preview interface. In actual applications, users can flexibly select the desired target virtual image according to needs and merge it with the media information stream. For example, user A takes a group photo for user B and user C, but also hopes to be in the group photo, then can select his own target virtual image from the second virtual image list, at this time, the target virtual image of user A can be displayed on the preview interface together with user B and user C, so as to achieve the effect of user A, user B and user C taking a group photo. For another example, user D cannot be present due to something, but hopes to appear in the group photo, user D can send his corresponding virtual image to user A, after user A selects the virtual image corresponding to user D, the virtual image corresponding to user D can also be displayed on the preview interface together with user B and user C, so as to achieve the effect of user D, user B and user C taking a group photo. Through the above method, the limitations of time and space can be broken, and it is better applied to scenes where it is impossible to take a group photo due to various reasons.

[0068] In some specific implementation examples, step C may be performed with reference to the following steps C1 and C2:

[0069] Step C1, obtaining the target synthesis position of the target virtual image selected by the third selection operation in the media information stream. In one embodiment, the target synthesis position of the target virtual image selected by the third selection operation can be determined based on the position of the object contained in the media information stream; in other embodiments, in response to monitoring the position designation operation, the position corresponding to the position designation operation can be used as the target synthesis position of the target virtual image selected by the third selection operation. That is, the optimal synthesis position of the target virtual image in the media information stream can be automatically determined by the algorithm. For example, if user C is located on the right side of user B in the media information stream, the algorithm can automatically set the target virtual image to the left side of user B. In addition, the user can also flexibly specify the synthesis position of the target virtual image according to needs. In actual applications, the target synthesis position of the target virtual image can be displayed in real time on the information preview interface so that the user can clearly know or adjust the target synthesis position.

[0070] Step C2, based on the target synthesis position, the target virtual image selected by the third selection operation is combined with the media information stream and displayed on the information preview interface. In order to ensure the synthesis effect, in some specific implementations, step C2 can be performed with reference to the following steps C2.1 and C2.2:

[0071] Step C2.1, when the object in the media information stream is a full-body image, and the target virtual image selected by the third selection operation is a half-body image, the target virtual image selected by the third selection operation is completed to obtain a completed target virtual image; the completion processing is to complete the whole body of the target virtual image, specifically, the current virtual image library can be used to find the body part that matches the target virtual image, and it can be spliced with the existing half-body image of the target virtual image to obtain a full-body image of the target virtual image.

[0072] Step C2.2, based on the target synthesis position, the completed target virtual image is combined with the media information stream and displayed on the information preview interface. By completing the target virtual image, the merging effect of the target virtual image and the media information stream can be effectively guaranteed, so that the consistency between the target virtual image and the object in the media information stream is stronger, and the merging effect is more natural and realistic.

[0073] In some implementations, the method provided by the embodiment of the present disclosure further includes: in response to receiving a synthesis confirmation instruction, synthesizing the selected target virtual image and the media information stream based on the screen content currently displayed on the information preview interface to obtain the target media information. In this way, a synchronized image or synchronized video in which the target virtual image is integrated into the shooting screen can be directly obtained, providing the user with a convenient synchronized experience.

[0074] In some embodiments, the target media processing task includes a dialogue task; the method for processing a virtual avatar provided by the embodiments of the present disclosure further includes the following steps 1 to 4:

[0075] Step 1: When starting a dialogue task, obtain the dialogue form and display a list of third virtual avatars; wherein, the dialogue form includes a text form and / or a voice form; the list of third virtual avatars presents at least one saved target virtual avatar; wherein, different target virtual avatars correspond to different objects. The list of third virtual avatars may be the same as or different from the aforementioned list of first virtual avatars or second virtual avatars. For specific details, reference may be made to the relevant descriptions of the aforementioned virtual avatar list, which will not be elaborated herein.

[0076] Step 2: In response to a fourth selection operation on a target virtual avatar in the list of third virtual avatars, use the target virtual avatar selected by the fourth selection operation as the current dialogue avatar.

[0077] Step 3: Obtain target reference information to generate output information of the current dialogue avatar based on the target reference information; wherein, the target reference information includes identity description information of the current dialogue avatar and / or received user input information. The identity description information of the current object avatar can be pre-added and set by the user. For example, the identity description information of the current object avatar indicates that the current object avatar is the user's grandmother or the user's teacher. The user can input information through voice and / or text to facilitate dialogue interaction with the target virtual avatar.

[0078] Step 4: Present the output information of the current dialogue avatar based on the dialogue form.

[0079] When the dialogue form is a text form, the current dialogue avatar can be used as the background interface. When the dialogue form is a voice form, the output information of the current dialogue avatar is expressed through the current dialogue avatar using a target voice feature and a target expression feature; wherein, the target voice feature is determined based on the voice feature of the object corresponding to the current dialogue avatar, and the target expression feature is determined based on the expression feature of the object corresponding to the current dialogue avatar. For example, the target voice feature is the voice feature of the object corresponding to the current dialogue avatar, and the target expression feature is the expression feature of the object corresponding to the current dialogue avatar. Through the above method, the effect of having a real conversation with the object corresponding to the current dialogue avatar can be presented to the user, greatly improving the user experience. In practical applications, the above dialogue process can also be exported and saved as a video, which is not limited herein.

[0080] In summary, the virtual image processing method provided by the embodiments of the present disclosure can enable users to personalize the virtual image they need, and the virtual image has corresponding characteristics with the target object such as the user, and can also enhance the emotional connection between the user and the virtual image. Furthermore, the above-mentioned target virtual image can be edited, shared and interacted with, and can be flexibly used as material information or related information in media editing scenarios, which will help promote the widespread application of virtual images.

[0081] Corresponding to the aforementioned virtual image processing method, the present disclosure also provides a virtual image processing device. Figure 4 The structure diagram of a virtual image processing device provided by an embodiment of the present disclosure is shown in FIG. 1 . The device can be implemented by software and / or hardware and can generally be integrated in an electronic device, such as Figure 4 As shown, including:

[0082] Information acquisition module 402, used to acquire media information of a target object;

[0083] An image acquisition module 404 is used to acquire a target virtual image corresponding to the target object based on the media information of the target object; the target features of the target virtual image correspond to the target features of the target object;

[0084] The image saving module 406 is used to save the target virtual image of the target object; wherein the target virtual image is used as the material information and / or associated information of the target media processing task.

[0085] The target virtual image obtained by the above method has the same or similar target features as the target object, thereby achieving a personalized customization effect of the virtual image of the target object. Users can customize personalized virtual images according to their needs, which can effectively avoid the situation where virtual images are similar. The target virtual image can also be flexibly applied to media processing tasks as material or related information, which is helpful to promote the widespread application of virtual images.

[0086] In some embodiments, the information acquisition module 402 is specifically used to: upon receiving a virtual image customization request for a target object and obtaining media information collection authority for the target object, collect images and / or videos containing the target object, and collect audio of the target object.

[0087] In some embodiments, the image acquisition module 404 is specifically used to: generate and display at least one candidate virtual image corresponding to the target object based on the media information of the target object; and in response to a first selection operation for the at least one candidate virtual image, obtain a target virtual image corresponding to the target object based on the candidate virtual image selected by the first selection operation.

[0088] In some embodiments, the image acquisition module 404 is specifically configured to: in response to an editing request for a candidate virtual image selected by the first selection operation, display a plurality of editing items; wherein different editing items are used to edit different virtual image features; in response to a target editing item among the plurality of editing items being triggered, perform feature editing on the candidate virtual image selected by the first selection operation based on the target editing item to obtain a target virtual image corresponding to the target object.

[0089] In some embodiments, the target media processing task includes a template synthesis task; the apparatus further includes a first task execution module, configured to: when starting the template synthesis task, display a plurality of virtual image templates; wherein the virtual image templates are used to present media information of sample virtual images; in response to a target template among the plurality of virtual image templates being triggered, perform a generation process based on the target virtual image and the target template to obtain media information corresponding to the target virtual image.

[0090] In some embodiments, the first task execution module is specifically configured to: replace a specified feature of the sample virtual image in the target template with a specified feature of the target virtual image.

[0091] In some embodiments, the target media processing task includes a media editing task; the apparatus further includes a second task execution module, configured to: when starting the media editing task, acquire media information to be edited, and acquire an associated virtual image corresponding to the media information to be edited; identify an object to be edited corresponding to the associated virtual image from the media information to be edited; perform an editing process on the object to be edited in the media information to be edited to obtain edited media information.

[0092] In some embodiments, the second task execution module is specifically configured to: display a first virtual image list; the first virtual image list presents at least one saved target virtual image; wherein different target virtual images correspond to different objects; in response to a second selection operation on a target virtual image in the first virtual image list, use the target virtual image selected by the second selection operation as the associated virtual image corresponding to the media information to be edited.

[0093] In some embodiments, the target virtual images in the virtual image list include a first virtual image and a second virtual image; wherein the first virtual image is the target virtual image of the target object, and the second virtual image is the target virtual image corresponding to an associated object of the target object.

[0094] In some embodiments, the second task execution module is further configured to: when it is detected that there is a second virtual image in the virtual image list that meets the preset conditions, set the second virtual image that meets the preset conditions to a non-selectable state, or delete the second virtual image that meets the preset conditions from the virtual image list.

[0095] In some embodiments, the preset conditions include one or more of the following: the duration of being added to the virtual image list is higher than a preset duration threshold, the image features are changed, and a disconnection instruction is received.

[0096] In some embodiments, the target media processing task includes a media synthesis task; the apparatus further includes a third task execution module, configured to: when the media synthesis task is started, collect a media information stream and display the media information stream on an information preview interface; display a second virtual image list; the second virtual image list presents at least one saved target virtual image; wherein, different target virtual images correspond to different objects; in response to a third selection operation on a target virtual image in the second virtual image list, merge the target virtual image selected by the third selection operation with the media information stream and display it on the information preview interface.

[0097] In some embodiments, the third task execution module is specifically configured to: obtain a target synthesis position of the target virtual image selected by the third selection operation in the media information stream; based on the target synthesis position, merge the target virtual image selected by the third selection operation with the media information stream and display it on the information preview interface.

[0098] In some embodiments, the third task execution module is specifically configured to: determine a target synthesis position of the target virtual image selected by the third selection operation based on the position of the object included in the media information stream; or, in response to detecting a position specifying operation, use the position corresponding to the position specifying operation as the target synthesis position of the target virtual image selected by the third selection operation.

[0099] In some embodiments, the third task execution module is specifically configured to: when the object in the media information stream is a full-body image and the target virtual image selected by the third selection operation is a half-body image, perform a complement processing on the target virtual image selected by the third selection operation to obtain a complemented target virtual image; based on the target synthesis position, merge the complemented target virtual image with the media information stream and display it on the information preview interface.

[0100] In some embodiments, the target media processing task includes a conversation task; the apparatus further includes a fourth task execution module, configured to: when starting the conversation task, obtain a conversation form and display a list of third virtual avatars; wherein, the conversation form includes a text form and / or a voice form; the list of third virtual avatars presents at least one saved target virtual avatar; wherein different target virtual avatars correspond to different objects; in response to a fourth selection operation on a target virtual avatar in the list of third virtual avatars, use the target virtual avatar selected by the fourth selection operation as the current conversation avatar; obtain target reference information to generate output information of the current conversation avatar based on the target reference information; wherein, the target reference information includes identity description information of the current conversation avatar and / or received user input information; present the output information of the current conversation avatar based on the conversation form.

[0101] In some embodiments, the fourth task execution module is specifically configured to: when the conversation form is a voice form, use the current conversation avatar to express the output information of the current conversation avatar with a target voice feature and a target expression feature; wherein, the target voice feature is determined based on the voice feature of the object corresponding to the current conversation avatar, and the target expression feature is determined based on the expression feature of the object corresponding to the current conversation avatar.

[0102] The virtual avatar processing apparatus provided by the embodiments of the present disclosure can execute the virtual avatar processing method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0103] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the apparatus embodiments described above can refer to the corresponding process in the method embodiments, and will not be described in detail here.

[0104] The embodiments of the present disclosure provide an electronic device, which includes: a storage device on which a computer program is stored; a processing device configured to execute the computer program in the storage device to implement the steps of any one of the methods in the present disclosure.

[0105] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0106] As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0107] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be implemented or included alternatively.

[0108] Specifically, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.

[0109] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to execute the image processing method provided by the embodiments of the present disclosure. The computer program products may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0110] In addition, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the processor is caused to execute the method for processing virtual avatars provided by the embodiments of the present disclosure.

[0111] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0112] Embodiments of the present disclosure also provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the method for processing virtual avatars in the embodiments of the present disclosure.

[0113] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0114] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosed technical solution based on the prompt message.

[0115] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may, for example, be in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0116] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0117] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0118] The above are only specific implementation manners of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing a virtual image, characterized in that, Including: Obtain the media information of the target object; Based on the media information of the target object, obtain the target virtual image corresponding to the target object; wherein, the target features of the target virtual image correspond to the target features of the target object; Save the target virtual image of the target object; wherein, the target virtual image is used as the material information and / or associated information for the target media processing task.

2. The method according to claim 1, wherein The obtaining the media information of the target object includes: When receiving a virtual image customization request for the target object and obtaining the media information collection permission for the target object, collect the image and / or video containing the target object, and collect the audio of the target object.

3. The method according to claim 1, wherein The obtaining the target virtual image corresponding to the target object based on the media information of the target object includes: Based on the media information of the target object, generate and display at least one candidate virtual image corresponding to the target object; In response to a first selection operation for the at least one candidate virtual image, obtain the target virtual image corresponding to the target object based on the candidate virtual image selected by the first selection operation.

4. The method according to claim 3, characterized in that, The obtaining the target virtual image corresponding to the target object based on the candidate virtual image selected by the first selection operation includes: In response to an editing request for the candidate virtual image selected by the first selection operation, display a plurality of editing items; wherein, different editing items are used to edit different virtual image features; In response to the triggering of a target editing item among the plurality of editing items, perform feature editing on the candidate virtual image selected by the first selection operation based on the target editing item to obtain the target virtual image corresponding to the target object.

5. The method according to claim 1, characterized in that The target media processing task includes a template synthesis task; the method further includes: When starting the template synthesis task, display a plurality of virtual image templates; wherein, the virtual image templates are used to present the media information of the sample virtual image; In response to the triggering of a target template among the plurality of virtual image templates, perform a generation process based on the target virtual image and the target template to obtain the media information corresponding to the target virtual image.

6. The method according to claim 1, wherein The target media processing task includes a media editing task; the method further includes: When starting the media editing task, obtain the media information to be edited, and obtain the associated virtual image corresponding to the media information to be edited; Identify the object to be edited corresponding to the associated virtual image from the media information to be edited; Perform editing processing on the object to be edited in the media information to be edited to obtain the edited media information.

7. The method according to claim 6, characterized in that, The obtaining the associated virtual image corresponding to the media information to be edited includes: Display a first virtual image list; the first virtual image list presents at least one saved target virtual image; wherein, different target virtual images correspond to different objects; In response to a second selection operation for the target virtual image in the first virtual image list, use the target virtual image selected by the second selection operation as the associated virtual image corresponding to the media information to be edited.

8. The method according to claim 7, wherein The target virtual avatars in the virtual avatar list include a first virtual avatar and a second virtual avatar; wherein, the first virtual avatar is the target virtual avatar of the target object, and the second virtual avatar is the target virtual avatar corresponding to the associated object of the target object.

9. The method according to claim 8, wherein The method further includes: When it is detected that there is a second virtual avatar in the virtual avatar list that meets the preset conditions, setting the second virtual avatar that meets the preset conditions to a non-selectable state, or deleting the second virtual avatar that meets the preset conditions from the virtual avatar list.

10. The method according to claim 9, wherein The preset conditions include one or more of the following: the duration of being added to the virtual avatar list is higher than a preset duration threshold, the image features are changed, and a cancellation association instruction is received.

11. The method according to claim 1, wherein The target media processing task includes a media synthesis task; the method further includes: When starting the media synthesis task, collecting a media information stream and displaying the media information stream on an information preview interface. Displaying a second virtual avatar list; the second virtual avatar list presents at least one saved target virtual avatar; wherein, different target virtual avatars correspond to different objects. In response to a third selection operation on the target virtual avatar in the second virtual avatar list, merging the target virtual avatar selected by the third selection operation with the media information stream and displaying it on the information preview interface.

12. The method according to claim 11, wherein The merging the target virtual avatar selected by the third selection operation with the media information stream and displaying it on the information preview interface includes: Obtaining the target synthesis position of the target virtual avatar selected by the third selection operation in the media information stream. Based on the target synthesis position, merging the target virtual avatar selected by the third selection operation with the media information stream and displaying it on the information preview interface.

13. The method according to claim 12, wherein The obtaining the target synthesis position of the target virtual avatar selected by the third selection operation in the media information stream includes: Determining the target synthesis position of the target virtual avatar selected by the third selection operation based on the position of the object included in the media information stream. Or, in response to detecting a position specifying operation, using the position corresponding to the position specifying operation as the target synthesis position of the target virtual avatar selected by the third selection operation.

14. The method according to claim 12, characterized in that, The based on the target synthesis position, merging the target virtual avatar selected by the third selection operation with the media information stream and displaying it on the information preview interface includes: When the object in the media information stream is a full-body image and the target virtual avatar selected by the third selection operation is a half-body image, performing a complement processing on the target virtual avatar selected by the third selection operation to obtain a complemented target virtual avatar. Based on the target synthesis position, merging the complemented target virtual avatar with the media information stream and displaying it on the information preview interface.

15. The method according to claim 1, wherein The target media processing task includes a dialogue task; the method further includes: When starting the dialogue task, obtain the dialogue form and display a list of third virtual avatars; wherein, the dialogue form includes text form and / or voice form; the list of third virtual avatars presents at least one saved target virtual avatar; wherein, different target virtual avatars correspond to different objects. In response to a fourth selection operation on a target virtual avatar in the list of third virtual avatars, use the target virtual avatar selected by the fourth selection operation as the current dialogue avatar. Obtain target reference information to generate output information of the current dialogue avatar based on the target reference information; wherein, the target reference information includes identity description information of the current dialogue avatar and / or received user input information. Present the output information of the current dialogue avatar based on the dialogue form.

16. The method according to claim 15, wherein The presenting the output information of the current dialogue avatar based on the dialogue form includes: When the dialogue form is the voice form, use the current dialogue avatar to express the output information of the current dialogue avatar with target voice characteristics and target expression characteristics; wherein, the target voice characteristics are determined based on the voice characteristics of the object corresponding to the current dialogue avatar, and the target expression characteristics are determined based on the expression characteristics of the object corresponding to the current dialogue avatar.

17. An apparatus for processing an avatar, characterized in that Includes: An information acquisition module for acquiring media information of a target object. An avatar acquisition module for acquiring a target virtual avatar corresponding to the target object based on the media information of the target object; wherein, the target characteristics of the target virtual avatar correspond to the target characteristics of the target object. An avatar storage module for storing the target virtual avatar of the target object; wherein, the target virtual avatar is used as material information and / or associated information for a target media processing task.

18. An electronic device, characterized in that, The electronic device includes: A storage device on which a computer program is stored. A processing device for executing the computer program in the storage device to implement the steps of the virtual avatar processing method according to any one of claims 1-16.

19. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the virtual avatar processing method according to any one of the above claims 1-16.