Animation generation method, storage medium, electronic device, and computer program product
By selecting target objects from multimedia files and generating target migration data to adapt to virtual character animation, the problem of poor migration flexibility in AI animation generation is solved, and the diversity and flexibility of customized animation are achieved.
Patent Information
- Application Number
- PCT/CN2025/077193
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2025-02-13
- Publication Date
- 2025-09-25
AI Technical Summary
In the existing technology, AI has poor flexibility in migrating the overall data of character movements and expressions in video materials, resulting in a single animation generation effect and unable to meet players' customized animation generation and display needs.
By selecting target objects from multiple object dimensions, determining the content to be extracted from multimedia files, generating target migration data, and controlling the target character in the virtual scene to generate target animation that matches the target migration data, user-defined animation generation is supported.
It enables flexible generation of migration data in multiple object dimensions, meets user-defined virtual character animation requirements, and improves the diversity and flexibility of animation rendering methods.
Smart Images

Figure CN2025077193_25092025_PF_FP_ABST
Abstract
Description
Animation generation method, storage medium, electronic device and computer program product
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application number 202410338066.5, filed on March 22, 2024, entitled “Animation Generation Method, Storage Medium, Electronic Device and Computer Program Product”, the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0003] The present disclosure relates to the field of computer technology, and in particular to an animation generation method, a storage medium, an electronic device, and a computer program product. Background Art
[0004] With the rapid development of the internet and gaming industries, avatars have become an indispensable element in games. Game teams create numerous expressions and movements for avatars based on plot or gameplay requirements, making the dynamic effects of avatars more realistic, enhancing game performance and increasing player immersion. However, traditional techniques for creating avatar animations, either through handcrafted artistry or complex expression and motion capture equipment, incur significant costs and have high barriers to entry, preventing players from independently generating the desired animation content. With the rise of artificial intelligence (AI) technology, related technologies have proposed using AI to capture character movements and expressions from video footage, then mapping the captured results into the game, thereby improving the efficiency of avatar animation production. However, this capture method suffers from a relatively fixed dimension. Specifically, the captured movements and expressions from the video footage are directly transferred to the avatar during each animation generation. This not only consumes significant computing resources but also limits the flexibility of the transferred data, resulting in a monotonous animation generation effect. Furthermore, after players upload their video footage, they are unable to participate in the animation generation and display process for the avatar, making it difficult to meet players' customized animation generation and display needs.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0006] At least some embodiments of the present disclosure provide an animation generation method, storage medium, electronic device and computer program product to at least solve the technical problems of poor flexibility and single animation generation effect in the animation generation method provided in the related art, which uses AI to perform overall data migration of character movements and expressions in video materials.
[0007] According to one embodiment of the present disclosure, an animation generation method is provided, including: obtaining a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of a source object from multiple object dimensions; selecting a target focus object from the multiple object dimensions; determining content to be extracted from the multimedia file according to the target focus object; generating target migration data based on the content to be extracted; and controlling a target character in a virtual scene to generate a target animation adapted to the target migration data.
[0008] According to one embodiment of the present disclosure, an animation display method is provided, which provides a graphical user interface through a terminal device, and the content displayed by the graphical user interface includes object identifiers of a target virtual character and multiple objects of interest in a virtual scene. The method includes: receiving a multimedia file uploaded by a user; responding to an object selection operation for multiple object identifiers, and determining a target object of interest from multiple objects of interest; responding to an animation generation operation for the target virtual character, extracting target migration data from the multimedia file based on the target object of interest, and rendering the target virtual character based on the target migration data to generate a character animation that matches the target migration data; and playing the character animation.
[0009] According to one embodiment of the present disclosure, an animation generation device is also provided, including: an acquisition module, configured to execute acquisition of a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of a source object from multiple object dimensions; a selection module, configured to execute selection of a target focus object from multiple object dimensions; a determination module, configured to execute determination of content to be extracted from the multimedia file according to the target focus object; a generation module, configured to execute generation of target migration data based on the content to be extracted; and a control module, configured to execute control of a target character in a virtual scene to generate a target animation adapted to the target migration data.
[0010] According to one embodiment of the present disclosure, an animation display device is provided, which provides a graphical user interface through a terminal device, and the content displayed by the graphical user interface includes the object identification of a target virtual character and multiple objects of interest in a virtual scene. The device includes: a receiving module, configured to execute and receive multimedia files uploaded by a user; a determination module, configured to execute an object selection operation in response to multiple object identifications, and determine a target object of interest from multiple objects of interest; a processing module, configured to execute an animation generation operation in response to the target virtual character, extract target migration data from the multimedia file based on the target object of interest, and render the target virtual character based on the target migration data to generate a character animation matching the target migration data; and a playback module, configured to execute and play the character animation.
[0011] According to one embodiment of the present disclosure, a computer-readable storage medium is further provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned animation generation method in each embodiment of the present disclosure when running.
[0012] According to one embodiment of the present disclosure, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the above-mentioned animation generation method in each embodiment of the present disclosure.
[0013] According to another aspect of the embodiments of the present disclosure, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, the computer program implements the above-mentioned animation generation method in each embodiment of the present disclosure.
[0014] According to another aspect of an embodiment of the present disclosure, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned animation generation method in each embodiment of the present disclosure is implemented.
[0015] According to another aspect of the embodiments of the present disclosure, a computer program is further provided. When the computer program is executed by a processor, the above-mentioned animation generation method in each embodiment of the present disclosure is implemented.
[0016] In at least some embodiments of the present disclosure, by selecting target objects of interest from multiple object dimensions, determining the content to be extracted from the multimedia file to be identified according to the target objects of interest, and then generating target migration data based on the content to be extracted, so as to control the target character in the virtual scene to generate a target animation that is adapted to the target migration data, a variety of migration data can be obtained by selecting different target objects of interest, thereby achieving the purpose of flexibly generating migration data corresponding to the content to be extracted in multiple object dimensions to render virtual character animations that meet custom needs, thereby achieving the technical effect of improving the flexibility of the migration data generation method, the diversity of the animation rendering method, and meeting the user's custom animation generation needs, and thus solving the technical problems of poor flexibility and single animation generation effect of the animation generation method provided in the related technology using AI to perform overall data migration of character actions and expressions in video materials. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0018] FIG1 is a block diagram showing a hardware structure of a mobile terminal according to an animation generation method of the present exemplary embodiment;
[0019] FIG2 shows a flow chart of an animation generation method according to the present exemplary embodiment;
[0020] FIG3 is a schematic diagram showing an animation generation method according to the present exemplary embodiment;
[0021] FIG4 is a schematic diagram showing an animation generation method according to the present exemplary embodiment;
[0022] FIG5 is a flowchart showing an animation display method according to the present exemplary embodiment;
[0023] FIG6 is a schematic diagram showing an interface of an animation display method according to this exemplary embodiment;
[0024] FIG7 is a schematic diagram showing an interface of an animation display method according to this exemplary embodiment;
[0025] FIG8 is a schematic diagram showing an interface of an animation display method according to this exemplary embodiment;
[0026] FIG9 is a schematic diagram showing an interface of an animation display method according to this exemplary embodiment;
[0027] FIG10 is a schematic diagram showing an interface of an animation display method according to this exemplary embodiment;
[0028] FIG11 is a schematic diagram showing an expression animation according to the exemplary embodiment;
[0029] FIG12 is a schematic diagram showing an action animation according to the present exemplary embodiment;
[0030] FIG13 shows a structural block diagram of an animation generating device according to the present exemplary embodiment;
[0031] FIG14 shows a structural block diagram of an animation display device according to the present exemplary embodiment;
[0032] FIG. 15 is a schematic diagram showing an electronic device according to one exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] The animation generation method in one embodiment of the present disclosure can be run on a local terminal device or a server. When the animation generation method is run on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device.
[0036] In an optional embodiment, various cloud applications can be run under the cloud interaction system, such as cloud games. Taking cloud games as an example, cloud games refer to a gaming method based on cloud computing. In the cloud game operation mode, the operating body of the game program and the main body of the game screen presentation are separated. The storage and operation of the animation generation method are completed on the cloud game server. The role of the client device is to receive and send data and present the game screen. For example, the client device can be a display device with data transmission function close to the user side, such as a mobile terminal, TV, computer, PDA, etc.; but the cloud game server in the cloud is responsible for information processing. When playing the game, the player operates the client device to send operation instructions to the cloud game server. The cloud game server runs the game according to the operation instructions, encodes and compresses the game screen and other data, and returns it to the client device through the network. Finally, the client device decodes and outputs the game screen.
[0037] In an alternative embodiment, taking a game as an example, a local terminal device stores the game program and is used to present the game screen. The local terminal device is used to interact with the player via a graphical user interface (GUI), i.e., the game program is downloaded, installed, and run via a conventional electronic device. The local terminal device can present the GUI to the player in a variety of ways, such as rendering it on the terminal's display screen or presenting it to the player via holographic projection.
[0038] For example, the local terminal device can be a mobile terminal, a computer terminal or a similar computing device, which may include a display screen and a processor, wherein the display screen is used to present a graphical user interface, the graphical user interface includes a game screen, and the processor is used to run the game, generate the graphical user interface and control the display of the graphical user interface on the display screen.
[0039] Taking running on a mobile terminal as an example, the mobile terminal can be a smartphone, tablet computer, PDA, mobile Internet device, PAD, game console, or other terminal device. Figure 1 is a hardware structure block diagram of a mobile terminal for an animation generation method according to an embodiment of the present disclosure. As shown in Figure 1, the mobile terminal may include one or more (only one is shown in Figure 1) processors 102 and a memory 104 for storing data. Optionally, the mobile terminal may also include a transmission device 106, an input / output device 108, and a display device 110.
[0040] Those skilled in the art will appreciate that the structure shown in FIG1 is merely illustrative and does not limit the structure of the mobile terminal. For example, the mobile terminal may include more or fewer components than shown in FIG1 , or have a configuration different from that shown in FIG1 .
[0041] According to one embodiment of the present disclosure, an embodiment of an animation generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0042] FIG2 is a flow chart of an animation generation method according to one embodiment of the present disclosure. As shown in FIG2 , the method includes the following steps:
[0043] Step S21: obtaining a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions;
[0044] Step S22, selecting a target object of interest from multiple object dimensions;
[0045] Step S23, determining the content to be extracted from the multimedia file according to the target object of interest;
[0046] Step S24, generating target migration data based on the content to be extracted;
[0047] Step S25 , controlling the target character in the virtual scene to generate a target animation adapted to the target migration data.
[0048] The multimedia files to be identified may be files in the form of pictures, audio, video, etc., which are used to provide original materials. The original materials can describe the source object's emotional expression from multiple object dimensions, so as to understand its emotional state by analyzing the source object's expression, language, body movements, etc. in different situations. For example, a person's emotional state can be inferred by analyzing his facial expressions, tone, body posture, etc. in a photo, or the speaker's emotional expression can be understood by analyzing the tone, speaking speed, language selection, etc. in an audio file. Multimedia files can be obtained through search engines, professional multimedia platforms or social media platforms, or they can be shot or recorded by photography, video or audio equipment.
[0049] After acquiring the original footage, target subjects are flexibly selected from multiple dimensions. These include at least one of the following: facial expressions, body movements, voice, and dialogue. These multiple dimensions can help analyze a character's emotions, intentions, attitudes, and behaviors, leading to a more comprehensive understanding of the character. Facial expressions can reveal a character's emotional state, body movements can reveal their behavioral intentions, and voice and dialogue can convey their thoughts and feelings. Comprehensively considering these factors can lead to a deeper understanding of the character, leading to a more refined character portrayal in the animation.
[0050] For example, based on the user's different customized animation requirements, target objects of interest can be selected from multiple object dimensions to generate target animations, thereby enabling the generation of multiple animation effects for the same multimedia file. For example, if a user wants to generate an animation where a character needs to display various facial expressions, such as laughing, crying, or surprise, facial expressions can be selected as the target object of interest. This allows the character's facial expressions to be nuanced and varied in the target animation, thereby presenting a rich variety of emotions and feelings. If a user wants to generate an animation where a character needs to display various body movements, such as dancing, running, or punching, body movements can be selected as the target object of interest. This allows the character's body movements to be smoothly varied in the target animation, thereby presenting a lively and vivid image. If a user wants to generate an animation where a character needs to be dubbed and perform dialogue, voice can be selected as the target object of interest. This allows the character's lip movements and voice to be synchronized in the target animation, thereby presenting a realistic dialogue effect. If a user wants to generate an animation where a character needs to express specific lines or plot points, lines can be selected as the target object of interest. This allows the character's plot to be accurately interpreted in the target animation, thereby presenting a vivid storyline. It should be noted that, in actual applications, the target focus object can also be a combination of multiple object dimensions, which is flexibly determined according to the user's animation generation requirements. The embodiment of this disclosure only gives an example, but does not constitute a specific limitation.
[0051] Furthermore, the content to be extracted is determined from the multimedia file according to the target object of attention. In the multimedia file, the content to be extracted can be determined by facial expressions, body movements, voice, and lines. If the target object of attention is a facial expression, the relevant content can be extracted through the expression of the character in the video or photo; if it is a body movement, the relevant content can be extracted through the movement of the character in the video or photo; if it is a voice, the relevant content can be extracted through the sound in the audio file; if it is a line, the relevant content can be extracted through the dialogue content in the video or audio file. Different content to be extracted can be selected from the multimedia file according to different target objects of attention. In this way, target migration data is generated based on the content to be extracted, and the target character in the virtual scene is controlled to generate a target animation that is adapted to the target migration data.
[0052] Based on the above steps S21 to S25, by selecting the target focus object from multiple object dimensions, determining the content to be extracted from the multimedia file to be identified according to the target focus object, and then generating target migration data based on the content to be extracted, so as to control the target character in the virtual scene to generate a target animation that is adapted to the target migration data. Therefore, by selecting different target focus objects, various types of migration data can be obtained, thereby achieving the purpose of flexibly generating migration data corresponding to the content to be extracted in multiple object dimensions to render virtual character animations that meet custom needs, thereby achieving the technical effect of improving the flexibility of the migration data generation method, the diversity of the animation rendering method, and meeting the user's custom animation generation needs, and thus solving the technical problems of poor flexibility and single animation generation effect of the animation generation method provided in the related technology using AI to perform overall data migration of character actions and expressions in video materials.
[0053] The following further introduces the animation generation method in the embodiment of the present disclosure.
[0054] Optionally, in step S23, determining the content to be extracted from the multimedia file according to the target object of interest includes:
[0055] Step S231, in response to the target focused object including facial expressions and body movements, performing a confidence check on the display content of the source object in the multimedia file to obtain a check result, wherein the display content includes: the expression content and the movement content of the source object;
[0056] Step S232: In response to the detection result satisfying the first preset condition, obtaining the content to be extracted from the displayed content.
[0057] Specifically, when the target focus object includes facial expressions and body movements, the expression content and movement content of the source object displayed in the multimedia file are tested for confidence, thereby obtaining a detection result. For the expression content of the source object, it is necessary to detect whether there are situations such as facial occlusion, large-angle side face, blurred original material, etc. that make the expression content unreliable; for the movement content of the source object, it is necessary to detect whether the character is out of the frame, occluded, or wearing loose clothing, etc. The confidence of each key point of the human body is different in different situations, and the confidence of the final calculation will also be different. For example, when the confidence in the detection result is greater than a preset threshold, such as 80%, it is determined that the detection result meets the first preset condition, and then the content to be extracted is obtained from the displayed content.
[0058] Based on the above optional embodiment, by performing a confidence check on the displayed content of the source object in the multimedia file in response to the target focus object including facial expressions and body movements, obtaining a detection result, and then obtaining the content to be extracted from the displayed content in response to the detection result meeting the first preset condition, this can help identify and understand the emotional state of the source object, thereby more accurately extracting the content to be extracted, further improving the ability to understand and analyze the content of multimedia files.
[0059] Optionally, in step S231, a confidence test is performed on the display content of the source object in the multimedia file, and the test results include:
[0060] Step S2311, obtaining multiple limb skeleton key points corresponding to the source object;
[0061] Step S2312, based on the confidence values corresponding to the multiple limb skeleton key points and the confidence weights corresponding to the multiple limb skeleton key points, the action content is confidence tested to obtain the test results, wherein the confidence values are used to distinguish the degree of visualization and / or clarity of different limb parts corresponding to the multiple limb skeleton key points, and the confidence weights are used to distinguish the importance of different limb parts corresponding to the multiple limb skeleton key points in the limb structure.
[0062] Specifically, multiple limb skeleton key points corresponding to the source object can be obtained through image recognition technology. There are 26 limb skeleton key points. Each limb skeleton key point has a corresponding confidence value based on the degree of visualization and / or clarity of the limb part. The confidence value range is 0-100%. The confidence value can be obtained by labeling and training a large amount of human limb key data. For example, if the foot of the source object is completely off the screen, the confidence value of the limb skeleton key point related to the foot is basically 0. The confidence weights corresponding to multiple limb skeleton key points are used to distinguish the importance of different limb parts corresponding to the multiple limb skeleton key points in the limb structure. For example, among the 26 limb skeleton key points, limb skeleton key point 11, limb skeleton key point 12, and limb skeleton key point 19 are key points for controlling the center of mass of the human body. If the confidence values of these three points are low, it will seriously affect the stability of the entire character movement. Therefore, the confidence weights of these three points are relatively high.
[0063] For example, when the multimedia file to be identified is a video file, the comprehensive action confidence is obtained by weighting multiple video frames. For example, if the multimedia file to be identified is a 10s video with a total of 300 frames (30 frames per second), the comprehensive action confidence of this video can be obtained by weighting the confidence values of 26 limb skeleton key points in the 300 video frames.
[0064] Based on the above optional embodiment, a more accurate confidence detection of the action content can be performed by taking the confidence values and weights corresponding to multiple limb skeleton key points. By obtaining multiple limb skeleton key points corresponding to the source object, the details and characteristics of the action can be captured more comprehensively, thereby improving the accuracy and reliability of the detection. At the same time, by weighting the confidence values corresponding to multiple limb skeleton key points, the overall situation of the action can be better reflected, avoiding the impact of the error of a single key point on the overall detection result.
[0065] Optionally, in step S231, a confidence test is performed on the display content of the source object in the multimedia file, and the test results include:
[0066] Step S2313, obtaining multiple facial bone key points corresponding to the source object;
[0067] Step S2314, performing confidence detection on the expression content based on the confidence values corresponding to multiple facial bone key points to obtain a detection result, wherein the confidence values are used to distinguish the degree of visualization and / or clarity of different facial areas corresponding to multiple facial bone key points.
[0068] Specifically, multiple facial skeletal key points corresponding to the source object can be obtained through image recognition technology. There are 68 facial skeletal key points, each of which has a corresponding confidence value based on the degree of visualization and / or clarity of the limb part. The confidence value ranges from 0-100%. The confidence value can be obtained by annotating and training a large amount of human facial key data. For example, if the eyes of the source object are completely occluded, the confidence value of the facial skeletal key points related to the eyes is basically 0. The confidence weights corresponding to multiple facial skeletal key points are used to distinguish the importance of different facial parts corresponding to the multiple facial skeletal key points in the expression structure. For example, among the 68 facial skeletal key points, the movement and changes of facial skeletal key points such as the upper edge of the eyebrows, the center of the eyebrows, the outer corners of the eyes, the top of the nose bridge, the bottom of the nose, the outer corners of the mouth, and the tip of the jaw can significantly change facial expressions, thereby affecting people's emotions and communication. Therefore, the confidence weights of these facial skeletal key points are relatively high.
[0069] Exemplarily, when the multimedia file to be identified is a video file, the comprehensive expression confidence is obtained by weighting multiple video frames. For example, the multimedia file to be identified is a 10s video, totaling 300 frames (30 frames per second). The comprehensive expression confidence of this video can be obtained by weighting the confidence values of 26 facial skeleton key points in the 300 video frames.
[0070] Based on the above optional embodiment, by obtaining multiple facial bone key points of the source object and performing analysis based on the confidence values corresponding to these facial bone key points, the confidence of the expression content can be accurately judged, thereby realizing the detection and analysis of the expression content, which can help to more accurately understand the expression state of the source object and provide strong support for face recognition, emotion analysis and other fields.
[0071] Optionally, in step S24, generating target migration data based on the content to be extracted includes:
[0072] Step S241, determining limb skeleton key points corresponding to the limb action based on the content to be extracted;
[0073] Step S242, performing three-dimensional motion restoration on the limb skeleton key points corresponding to the limb motion to obtain three-dimensional motion data;
[0074] Step S243: Redirect the three-dimensional motion data using the body shape data of the target character to generate motion transition data.
[0075] Specifically, the limb skeleton key points corresponding to the limb movements are determined based on the content to be extracted. Three-dimensional motion restoration is then performed on these key points to generate 3D motion data. This is then restored to the character's 3D motion data through 2D key point prediction. Furthermore, the 3D motion data is reoriented based on the target character's body shape, such as their height, weight, and overall size, to generate motion transfer data.
[0076] The essence of redirection is to assign 3D motion data to the target character according to the pre-provided body data of the target character, so as to maintain the consistency of the animation. When using the body data of the target character to redirect the 3D motion data, AI can be used for automatic redirection to directly obtain the animation effects of different body shapes, rather than waiting for professional art animators to repair them. The characteristics of redirection are: (1) According to the fatness of the target character's body shape, the 3D motion data is adjusted so that the generated motion migration data will not have the motion of penetrating the model, and the unnatural feeling caused by the distance between the hands and feet and the torso is avoided. Specifically, AI can be used to detect whether the model has a collision; (2) According to the height of the target character's body shape, the 3D motion data is adjusted so that the generated motion migration data will not be attached to the air or sink into the ground; (3) Avoid sliding due to the body shape of the data migration, which can be calculated by the length of the legs and feet of the front and back body shapes.
[0077] Based on the above optional embodiments, the details and characteristics of limb movements can be accurately captured by extracting limb skeleton key points and restoring three-dimensional movements. By redirecting the three-dimensional movement data using the target character's body shape data, movement migration between characters of different body shapes can be achieved, making the movement performance between different characters more natural and realistic, thereby improving the efficiency and quality of motion capture and animation production.
[0078] Optionally, in step S24, generating target migration data based on the content to be extracted includes:
[0079] Step S244, determining initial expression data corresponding to the facial expression based on the content to be extracted;
[0080] Step S245 , generating expression migration data corresponding to the initial expression data using a preset corresponding relationship, wherein the preset corresponding relationship is obtained through a pre-trained expression coding, and the expression coding is used to compare the coding distance between expression data of different characters.
[0081] The initial expression data can be that of a real person, and the transferred expression data is used to determine the expression data to be rendered for the virtual character. The goal of expression encoding is to map the initial expression data into a low-dimensional vector space where similar expressions have a close encoding distance. Through expression encoding, the character's expression frame is detected and the real person's expression is transferred to the virtual character.
[0082] Specifically, a set of real-person expression images and a set of virtual-character expression images can be preselected. During training, one image is selected from each of the real-person expression images and one image is selected from each of the virtual-character expression images. The expression encoding distance between these two images is then compared. This comparison of the expression encoding distances between these different images ultimately allows for the identification of a correspondence between the real-person expression data and the virtual-character expression data. Through continuous iterative training, this predetermined correspondence is established between the real-person expression images and the virtual-character expression images.
[0083] Therefore, after the initial expression data corresponding to the facial expression is determined based on the content to be extracted, the above-mentioned preset corresponding relationship can be used to find the expression transition data with the closest coding distance to the initial expression data.
[0084] Based on the above optional embodiments, by determining the initial expression data corresponding to the facial expression based on the content to be extracted, and then using the preset correspondence relationship to generate expression migration data corresponding to the initial expression data, the migration of expressions between different characters can be achieved, thereby achieving richer and more diverse expression communication.
[0085] Optionally, the animation generation method in the embodiment of the present disclosure further includes:
[0086] Step S233: In response to the multimedia file being a video file, locking the outline information of at least one object displayed in the video file;
[0087] Step S234: Based on the locked contour information, count the occurrence frequency of the at least one object in the multiple frames of image material in the video file to obtain a statistical result;
[0088] Step S235 : In response to the statistical result satisfying the second preset condition, content to be extracted is determined from the multimedia file according to the target object of interest.
[0089] Specifically, after inputting a video file, the valid interval of the video file will be detected first, that is, the contour information of at least one object displayed in the video file will be locked. When multiple objects are displayed in the video file, one of the objects can be locked first. Usually, the object with the largest contour displayed in the video file will be locked. Based on the locked contour information, the frequency of occurrence of at least one object in the multi-frame image material in the video file is counted to obtain a statistical result. If it is determined based on the statistical result that the duration of the appearance of at least one object in the multi-frame image material in the video file accounts for more than 50% of the duration of the video file, then it is determined that the statistical result meets the second preset condition, and the content to be extracted is determined from the multimedia file according to the target object of attention, thereby optimizing problems such as inaccurate character recognition caused by multi-person scenes and recognition errors caused by characters leaving the screen area.
[0090] Based on the above optional embodiment, in response to the multimedia file being a video file, the contour information of at least one object displayed in the video file is locked, and then based on the locked contour information, the frequency of occurrence of at least one object in multiple frames of image material in the video file is counted to obtain statistical results. Finally, in response to the statistical results satisfying the second preset condition, the content to be extracted is determined from the multimedia file according to the target object of interest, thereby realizing automatic recognition and tracking of specific objects in the video file, thereby quickly and accurately obtaining relevant information of the specific objects in the video file, which provides convenience for subsequent animation generation.
[0091] Figure 3 is a schematic diagram of an animation generation method according to one embodiment of the present disclosure. As shown in Figure 3, when the multimedia file to be identified is a video file, a target object of interest is selected from multiple object dimensions. If the target object of interest includes the expression and body movements of the source object, the expression content and movement content are extracted from the video file, and then the expression migration data and movement migration data are generated; if the target object of interest includes the expression of the source object, the expression content and head movement data are extracted from the video file, wherein the head movement data is the effect guarantee data generated when there is no body content in order to make the expression animation effect more vivid, and then the expression migration data and head movement migration parameters are generated; if the target object of interest includes the body movements of the source object, the movement content is extracted from the video file, and then the movement migration data is generated.
[0092] For example, when a target object includes both the source object's facial expressions and body movements, after inputting a video file, the system first detects the valid intervals of the facial expressions and body movements. Specifically, the system locks the outline of at least one object displayed in the video file. Based on this locked outline, the system then calculates the frequency of occurrence of the at least one object within multiple frames of the video file to generate statistical results. If the statistical results determine that the at least one object appears within the multiple frames of the video file for at least 50% of the video file's duration, the system then determines the facial expressions and body movements to be extracted from the multimedia file based on the target object.
[0093] After confirming the feasibility of the valid range, the confidence level of the footage needs to be tested. The expression generation module detects facial occlusion, wide-angle side profiles, and blurred video to avoid unreliable expressions. The action generation module detects scenes such as people being out of frame, occluded, and wearing loose clothing. The confidence level of key points corresponding to the human limb skeleton varies in different situations, and the final calculated confidence level also varies. The final output is clips with a confidence level greater than 80%.
[0094] Next, when performing expression transfer, an expression encoder is pre-trained. The goal of expression encoding is to map the initial expression data to a low-dimensional vector space where similar expressions have a close expression encoding distance. Through expression encoding, the character's expression frame is detected, and the real person's expression is transferred to the virtual character, generating the corresponding expression transfer data. When performing motion transfer, the key points of the character's limb skeleton are first identified. Through 2D key point prediction, the 3D motion data of the character is restored and presented in the form of skeletal animation. Finally, the 3D motion data is redirected according to the different body shapes of the target character to generate motion transfer data to reduce the problems of clipping and sliding caused by different body shapes.
[0095] Finally, the target character in the virtual scene is controlled to generate a target animation that matches the expression transfer data and the action transfer data.
[0096] Figure 4 is a schematic diagram of another animation generation method according to one embodiment of the present disclosure. As shown in Figure 4, when the multimedia file to be identified is an image file, since the image is a single-frame input, there is no need to judge the valid interval. The other implementation processes are similar to those of video files and will not be repeated here.
[0097] FIG5 is a flow chart of an animation display method according to one embodiment of the present disclosure. A graphical user interface is provided by a terminal device, and the content displayed by the graphical user interface includes object identifiers of a target virtual character and multiple objects of interest in a virtual scene. As shown in FIG5 , the method includes the following steps:
[0098] Step S51, receiving multimedia files uploaded by users;
[0099] Step S52, in response to the object selection operation for the multiple object identifiers, determining a target object of interest from the multiple objects of interest;
[0100] Step S53, in response to the animation generation operation for the target virtual character, extracting target migration data from the multimedia file based on the target focus object, rendering the target virtual character based on the target migration data, and generating a character animation that matches the target migration data;
[0101] Step S54: play the character animation.
[0102] The multimedia files to be identified may be files in the form of pictures, audio, video, etc., which are used to provide original materials. The original materials can describe the source object's emotional expression from multiple object dimensions, so as to understand its emotional state by analyzing the source object's expression, language, body movements, etc. in different situations. For example, a person's emotional state can be inferred by analyzing his facial expressions, tone, body posture, etc. in a photo, or the speaker's emotional expression can be understood by analyzing the tone, speaking speed, language selection, etc. in an audio file. Multimedia files can be obtained through search engines, professional multimedia platforms or social media platforms, or they can be shot or recorded by photography, video or audio equipment.
[0103] In response to an object selection operation for multiple object identifiers, a target object of interest is determined from the multiple objects of interest. The target object of interest includes at least one of the following: facial expressions, body movements, voice, and dialogue. These multiple objects can help analyze a character's emotions, intentions, attitudes, and behaviors, leading to a more comprehensive understanding of the character. Facial expressions can reveal a character's emotional state, body movements can reveal their behavioral intentions, and voice and dialogue can convey their thoughts and feelings. Taking these factors into consideration can lead to a better understanding of the character and, consequently, a more refined character portrayal in the animation.
[0104] For example, based on the user's different customized animation requirements, a target focus object can be selected from multiple objects to generate a target animation, thereby generating multiple animation effects for the same multimedia file. For example, if a user wants to generate an animation in which a character needs to display various facial expressions, such as laughing, crying, or surprise, they can select facial expressions as the target focus object. This allows the character's facial expressions to be nuanced and varied in the target animation, thereby presenting a rich variety of emotions and feelings. If a user wants to generate an animation in which a character needs to display various body movements, such as dancing, running, or punching, they can select body movements as the target focus object. This allows the character's body movements to be smoothly varied in the target animation, thereby presenting a lively and vivid image. If a user wants to generate an animation in which a character needs to be dubbed and perform dialogue, they can select voice as the target focus object. This allows the character's lip movements and voice to be synchronized in the target animation, thereby presenting a realistic dialogue effect. If a user wants to generate an animation in which a character needs to express specific lines or plot points, they can select lines as the target focus object. This allows the character's plot to be accurately interpreted in the target animation, thereby presenting a vivid storyline. It should be noted that, in actual applications, the target focus object may also be a combination of multiple objects, which is flexibly determined according to the user's animation generation requirements. The embodiment of the present disclosure only gives an example, but does not constitute a specific limitation.
[0105] It should be noted that the object selection operation can be a touch operation in which a user touches the display screen of the terminal device with a finger and touches the terminal device. The touch operation can include single-point touch and multi-point touch, wherein the touch operation of each touch point can include clicking, long pressing, pressing hard, swiping, etc. The object selection operation can also be a touch operation achieved through input devices such as a mouse, keyboard, and game controller.
[0106] Based on the above steps S51 to S54, by responding to the object selection operation for the object identification of multiple attention objects, the target attention object is determined from the multiple attention objects, and the animation generation operation for the target virtual character is responded to. The target migration data is extracted from the multimedia file based on the target attention object, and the target virtual character is rendered based on the target migration data to generate and play the character animation matching the target migration data, thereby achieving the purpose of flexibly and personally editing the virtual character display animation on the mobile terminal, thereby achieving the technical effect of improving the animation display editing flexibility of the mobile terminal and meeting the user's customized animation display needs, and further solving the technical problem in the related technology that after the player uploads the video material, he cannot participate in the animation generation and display process of the virtual character, and thus it is difficult to meet the player's customized animation generation and display needs.
[0107] The following further introduces the animation display method in the embodiment of the present disclosure.
[0108] Optionally, in step S54, playing the character animation includes one of the following:
[0109] Step S541, in response to a single play operation for a facial expression, controlling the target character to generate an expression animation that matches the target migration data;
[0110] Step S542, in response to a single playback operation for a body movement, controlling the target character to generate a movement animation adapted to the target migration data;
[0111] Step S543 , in response to the linked play operation for facial expressions and body movements, controlling the target character to synchronously generate expression animation and movement animation that are adapted to the target migration data.
[0112] Based on the above optional embodiments, the facial expressions and body movements of the target character can be controlled and generated individually or in conjunction, so that the target character can generate corresponding expression animations and action animations according to the input target migration data, thereby achieving a more realistic and vivid character performance and improving the quality and realism of the character animation.
[0113] Optionally, the animation display method in the embodiment of the present disclosure further includes: responding to a single playback operation for a facial expression, while controlling the target character to separately generate an expression animation adapted to the target migration data, and synchronously playing a head animation associated with the expression animation.
[0114] Based on the above optional embodiment, a single playback operation for facial expressions can simultaneously play the head animation associated with the target character's facial expression animation while controlling the target character to generate an expression animation that matches the target migration data. This can make the character's expression more vivid and natural, enhancing the character's expressiveness and communication skills. It also improves the user experience, making the character more realistic and eye-catching in the virtual scene.
[0115] Figure 6 is a schematic diagram of an interface for an animation display method according to one embodiment of the present disclosure. As shown in Figure 6, the graphical user interface displays content including a target virtual character in a virtual scene. The user can click the smart generation control to enter the smart animation interface shown in Figure 7. Figure 7 is a schematic diagram of an interface for another animation display method according to one embodiment of the present disclosure. As shown in Figure 7, the user can choose to upload an image or video. Uploading an image generates a single-frame animation, while uploading a video generates an animation of the same length as the video. The frame rate can be flexibly adjusted as needed. Figure 7 also provides a video prompt control "?". Clicking the video prompt control enters a user prompt interface, which informs the user of more suitable content in the form of text and sample videos. Figure 7 also provides an upload control "+". Clicking the upload control allows the user to upload local images and videos, with pre-set size and format restrictions for videos and images. An object selection operation is then performed on multiple object identifiers to determine a target object of interest from multiple objects of interest. The multiple object identifiers are expressions, actions, voice, and dialogue. This allows the user to select the content to be extracted from the image / video, with the core being expression data, action data, or both. When selecting the video format input, additional voice and lines options are provided. Voice refers to obtaining the original sound of the video, and lines refers to translating the lines in the voice into text. In Figure 7, a generation control is also provided. Click the generation control to extract target migration data from the multimedia file based on the target focus object, and render the target virtual character based on the target migration data to generate a character animation that matches the target migration data. Figure 8 is an interface diagram of another animation display method according to one embodiment of the present disclosure. As shown in Figure 8, after clicking the generation control, the algorithm is called to generate the character animation. The user can click the cancel control at any time to interrupt the generation process.
[0116] Figure 9 is an interface diagram of another animation display method according to one embodiment of the present disclosure. As shown in Figure 9, after generating the character animation, clicking the replace control can reselect the original material and return to the interface shown in Figure 7. By clicking the expression generation control or the action generation control, the generated expression animation or action animation can be previewed and presented on the target character on the left. According to the preview effect, you can check the actual action to be adopted and save the adopted expression action to the user's custom action for long-term use. Clicking the apply control can adopt the current expression action and enter the interface shown in Figure 10.
[0117] Figure 10 is a schematic diagram of an interface for yet another animation display method according to one embodiment of the present disclosure. As shown in Figure 10 , clicking the smart generation control enters the smart animation interface shown in Figure 7 . Clicking the play expression control allows the user to play expression animations individually, clicking the play action control allows the user to play action animations individually, and clicking the one-click play control allows the user to play both expression and action animations simultaneously. Figure 10 also includes a loop playback control, which allows the user to flexibly control whether or not to loop character animations.
[0118] Figure 11 is a schematic diagram of an expression animation according to one embodiment of the present disclosure. As shown in Figure 11, an expression image uploaded by a user is received, an object selection operation for multiple object identifiers is responded to, a target focus object is determined as an expression from multiple focus objects, an animation generation operation for a target virtual character is responded to, expression migration data is extracted from a multimedia file based on the target focus object, and the target virtual character is rendered based on the expression migration data to generate an expression animation that matches the expression migration data, and then a single frame of expression animation can be played.
[0119] Figure 12 is a schematic diagram of an action animation according to one embodiment of the present disclosure. As shown in Figure 12, an action image uploaded by a user is received, an object selection operation for multiple object identifiers is responded to, a target focus object is determined as an action from multiple focus objects, an animation generation operation for a target virtual character is responded to, action migration data is extracted from a multimedia file based on the target focus object, and the target virtual character is rendered based on the action migration data to generate an action animation that matches the action migration data, and then a single-frame action animation can be played.
[0120] Compared to facial expression and motion capture devices, the animation generation method and animation display method in the disclosed embodiments do not require additional hardware equipment. The AI animation migration technology only uses images or videos as input, thereby reducing equipment cost and complexity. In addition, because the AI animation migration technology is real-time, it can be applied and presented instantly in games or applications to provide a more vivid user experience. Generally, the response time can be similar to the input video time.
[0121] In actual player experience, many materials that do not meet the usage scenarios are uploaded, such as missing characters, multi-person scenes, expressions / actions being blocked, out of frame, etc. The disclosed method embodiment can block materials that do not meet the scenario as much as possible and ensure that the materials that meet the requirements are stably output to the player, thus meeting the player's needs for custom animation generation.
[0122] In the disclosed embodiment, game players can generate actions that did not exist in the game before in real time. Through this method, players can find suitable film and television dramas, dance pictures, and video clips, and generate character animations (mainly expressions and actions) into virtual characters for use in making short video / picture materials, thereby improving the playability and dissemination of game content.
[0123] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the animation generation method described in each embodiment of the present disclosure.
[0124] In this embodiment, an animation generation device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. The details that have been described will not be repeated here. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0125] FIG13 is a structural block diagram of an animation generating device according to one embodiment of the present disclosure. As shown in FIG13 , the device includes:
[0126] The acquisition module 1301 is configured to acquire a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions;
[0127] The selection module 1302 is configured to select a target object of interest from multiple object dimensions;
[0128] The determination module 1303 is configured to determine the content to be extracted from the multimedia file according to the target object of interest;
[0129] A generating module 1304 is configured to generate target migration data based on the content to be extracted;
[0130] The control module 1305 is configured to control the target character in the virtual scene to generate a target animation adapted to the target migration data.
[0131] Optionally, the determination module 1303 is further configured to execute: in response to the target object of attention including facial expressions and body movements, perform a confidence check on the display content of the source object in the multimedia file to obtain a detection result, wherein the display content includes: the expression content and action content of the source object; in response to the detection result meeting the first preset condition, obtain the content to be extracted from the display content.
[0132] Optionally, the determination module 1303 is further configured to execute: obtaining multiple limb skeleton key points corresponding to the source object; performing confidence detection on the action content based on the confidence values corresponding to the multiple limb skeleton key points and the confidence weights corresponding to the multiple limb skeleton key points to obtain the detection results, wherein the confidence values are used to distinguish the degree of visualization and / or clarity of different limb parts corresponding to the multiple limb skeleton key points, and the confidence weights are used to distinguish the importance of different limb parts corresponding to the multiple limb skeleton key points in the limb structure.
[0133] Optionally, the determination module 1303 is also configured to execute: obtaining multiple facial bone key points corresponding to the source object; performing confidence detection on the expression content based on the confidence values corresponding to the multiple facial bone key points to obtain a detection result, wherein the confidence value is used to distinguish the degree of visualization and / or clarity of different facial areas corresponding to the multiple facial bone key points.
[0134] Optionally, the generation module 1304 is further configured to execute: determining the limb skeleton key points corresponding to the limb movement based on the content to be extracted; performing three-dimensional motion restoration on the limb skeleton key points corresponding to the limb movement to obtain three-dimensional motion data; and redirecting the three-dimensional motion data using the body shape data of the target character to generate motion migration data.
[0135] Optionally, the generation module 1304 is also configured to perform: determining initial expression data corresponding to the facial expression based on the content to be extracted; and generating expression migration data corresponding to the initial expression data using a preset correspondence relationship, wherein the preset correspondence relationship is obtained through a pre-trained expression coding, and the expression coding is used to compare the coding distance between expression data of different characters.
[0136] Optionally, the animation generating device further includes: a locking module 1306, configured to execute, in response to the multimedia file being a video file, locking the contour information of at least one object displayed in the video file; a statistical module 1307, configured to execute, based on the locked contour information, statistics on the frequency of occurrence of at least one object in multiple frames of image material in the video file to obtain statistical results; the determination module 1303 is further configured to execute, in response to the statistical result satisfying a second preset condition, determining the content to be extracted from the multimedia file according to the target object of interest.
[0137] In the embodiment of the present disclosure, by selecting target attention objects from multiple object dimensions, determining the content to be extracted from the multimedia file to be identified according to the target attention objects, and then generating target migration data based on the content to be extracted, so as to control the target character in the virtual scene to generate a target animation that is adapted to the target migration data, a variety of migration data can be obtained by selecting different target attention objects, thereby achieving the purpose of flexibly generating migration data corresponding to the content to be extracted in multiple object dimensions to render virtual character animations that meet custom needs, thereby achieving the technical effect of improving the flexibility of migration data generation methods, the diversity of animation rendering methods, and meeting the user's custom animation generation needs, and thus solving the technical problems of poor flexibility and single animation generation effect of the animation generation method provided in the related technology using AI to perform overall data migration of character actions and expressions in video materials.
[0138] FIG14 is a block diagram of an animation display device according to one embodiment of the present disclosure. A graphical user interface is provided through a terminal device. The content displayed on the graphical user interface includes object identifiers of a target virtual character and multiple objects of interest in a virtual scene. As shown in FIG14 , the device includes:
[0139] The receiving module 1401 is configured to receive multimedia files uploaded by users;
[0140] The determination module 1402 is configured to perform an object selection operation in response to the multiple object identifiers, and determine a target object of interest from the multiple objects of interest;
[0141] The processing module 1403 is configured to execute an animation generation operation in response to the target virtual character, extract target migration data from the multimedia file based on the target focus object, render the target virtual character based on the target migration data, and generate a character animation that matches the target migration data;
[0142] The playing module 1404 is configured to play the character animation.
[0143] Optionally, the playback module 1404 is also configured to perform: responding to a single playback operation for facial expressions, controlling the target character to separately generate an expression animation that is compatible with the target migration data; responding to a single playback operation for body movements, controlling the target character to separately generate an action animation that is compatible with the target migration data; responding to a linked playback operation for facial expressions and body movements, controlling the target character to synchronously generate an expression animation and an action animation that are compatible with the target migration data.
[0144] Optionally, the playback module 1404 is further configured to execute: in response to a single playback operation for a facial expression, while controlling the target character to separately generate an expression animation adapted to the target migration data, synchronously play the head animation associated with the expression animation.
[0145] In the embodiment of the present disclosure, by responding to the object selection operation for the object identification of multiple focus objects, determining the target focus object from the multiple focus objects, responding to the animation generation operation for the target virtual character, extracting the target migration data from the multimedia file based on the target focus object, and rendering the target virtual character based on the target migration data to generate and play the character animation matching the target migration data, thereby achieving the purpose of flexibly and personally editing the virtual character display animation on the mobile terminal, thereby achieving the technical effect of improving the animation display editing flexibility of the mobile terminal and meeting the user's customized animation display needs, and further solving the technical problem in the related art that after the player uploads the video material, he cannot participate in the animation generation and display process of the virtual character, and thus it is difficult to meet the player's customized animation generation and display needs.
[0146] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0147] An embodiment of the present disclosure further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when run.
[0148] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0149] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0150] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0151] S1, obtaining a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions;
[0152] S2, select the target object of attention from multiple object dimensions;
[0153] S3, determining the content to be extracted from the multimedia file according to the target object of interest;
[0154] S4, generating target migration data based on the content to be extracted;
[0155] S5, controlling the target character in the virtual scene to generate a target animation adapted to the target migration data.
[0156] Optionally, the above-mentioned computer-readable storage medium is also configured to store program code for executing the following steps: in response to the target object of attention including facial expressions and body movements, performing a confidence check on the display content of the source object in the multimedia file to obtain a detection result, wherein the display content includes: the expression content and action content of the source object; in response to the detection result meeting the first preset condition, obtaining the content to be extracted from the display content.
[0157] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: obtaining multiple limb skeletal key points corresponding to a source object; performing confidence detection on the action content based on the confidence values corresponding to the multiple limb skeletal key points and the confidence weights corresponding to the multiple limb skeletal key points to obtain a detection result, wherein the confidence values are used to distinguish the degree of visualization and / or clarity of different limb parts corresponding to the multiple limb skeletal key points, and the confidence weights are used to distinguish the importance of different limb parts corresponding to the multiple limb skeletal key points in the limb structure.
[0158] Optionally, the above-mentioned computer-readable storage medium is also configured to store program code for executing the following steps: obtaining multiple facial bone key points corresponding to the source object; performing confidence detection on the expression content based on the confidence values corresponding to the multiple facial bone key points to obtain a detection result, wherein the confidence value is used to distinguish the degree of visualization and / or clarity of different facial areas corresponding to the multiple facial bone key points.
[0159] Optionally, the computer-readable storage medium is further configured to store program codes for executing the following steps: determining the limb skeleton key points corresponding to the limb movement based on the content to be extracted; performing three-dimensional movement restoration on the limb skeleton key points corresponding to the limb movement to obtain three-dimensional movement data; and redirecting the three-dimensional movement data using the body shape data of the target character to generate movement migration data.
[0160] Optionally, the computer-readable storage medium is further configured to store program codes for executing the following steps: determining initial expression data corresponding to facial expressions based on the content to be extracted; and generating expression migration data corresponding to the initial expression data using a preset correspondence relationship, wherein the preset correspondence relationship is obtained through a pre-trained expression encoding, and the expression encoding is used to compare the encoding distance between expression data of different characters.
[0161] Optionally, the computer-readable storage medium is further configured to store program codes for executing the following steps: in response to the multimedia file being a video file, locking the contour information of at least one object displayed in the video file; based on the locked contour information, counting the frequency of occurrence of at least one object in multiple frames of image material in the video file to obtain statistical results; in response to the statistical results satisfying a second preset condition, determining the content to be extracted from the multimedia file according to the target object of interest.
[0162] In the computer-readable storage medium of this embodiment, by selecting target objects of interest from multiple object dimensions, determining the content to be extracted from the multimedia file to be identified according to the target objects of interest, and then generating target migration data based on the content to be extracted, so as to control the target character in the virtual scene to generate a target animation that is adapted to the target migration data, a variety of migration data can be obtained by selecting different target objects of interest, thereby achieving the purpose of flexibly generating migration data corresponding to the content to be extracted in multiple object dimensions to render virtual character animations that meet custom needs, thereby achieving the technical effect of improving the flexibility of migration data generation methods, the diversity of animation rendering methods, and meeting the user's custom animation generation needs, and thus solving the technical problems of poor flexibility and single animation generation effect of the animation generation method provided in the related technology using AI to perform overall data migration of character actions and expressions in video materials.
[0163] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0164] S1, receiving multimedia files uploaded by users;
[0165] S2, in response to an object selection operation for multiple object identifiers, determining a target object of interest from multiple objects of interest;
[0166] S3, in response to the animation generation operation for the target virtual character, extracting target migration data from the multimedia file based on the target focus object, rendering the target virtual character based on the target migration data, and generating a character animation that matches the target migration data;
[0167] S4, play character animation.
[0168] Optionally, the computer-readable storage medium is further configured to store program codes for executing the following steps: in response to a single playback operation for facial expressions, controlling the target character to separately generate an expression animation that matches the target migration data; in response to a single playback operation for body movements, controlling the target character to separately generate an action animation that matches the target migration data; in response to a linked playback operation for facial expressions and body movements, controlling the target character to synchronously generate an expression animation and an action animation that match the target migration data.
[0169] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: in response to a single playback operation for a facial expression, while controlling the target character to separately generate an expression animation that matches the target migration data, synchronously playing the head animation associated with the expression animation.
[0170] In the computer-readable storage medium of this embodiment, by responding to the object selection operation of the object identification of multiple focus objects, the target focus object is determined from the multiple focus objects, and the animation generation operation for the target virtual character is responded to, the target migration data is extracted from the multimedia file based on the target focus object, and the target virtual character is rendered based on the target migration data to generate and play the character animation matching the target migration data, thereby achieving the purpose of flexibly and personally editing the virtual character display animation on the mobile terminal, thereby achieving the technical effect of improving the animation display editing flexibility of the mobile terminal and meeting the user's customized animation display needs, and further solving the technical problem in the related technology that after the player uploads the video material, he cannot participate in the animation generation and display process of the virtual character, and thus it is difficult to meet the player's customized animation generation and display needs.
[0171] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0172] In exemplary embodiments of the present disclosure, a computer-readable storage medium stores a program product capable of implementing the aforementioned method of the present embodiment. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product comprising program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps described in the "Exemplary Method" section of the present embodiment according to various exemplary embodiments of the present disclosure.
[0173] According to an embodiment of the present disclosure, a program product for implementing the above-mentioned method can be a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the embodiment of the present disclosure is not limited thereto. In the embodiment of the present disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0174] The program product may be implemented in any combination of one or more computer-readable media. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (non-exhaustive) of computer-readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0175] It should be noted that the program code contained in the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any appropriate combination of the above.
[0176] An embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0177] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0178] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0179] S1, obtaining a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions;
[0180] S2, select the target object of attention from multiple object dimensions;
[0181] S3, determining the content to be extracted from the multimedia file according to the target object of interest;
[0182] S4, generating target migration data based on the content to be extracted;
[0183] S5, controlling the target character in the virtual scene to generate a target animation adapted to the target migration data.
[0184] Optionally, the above-mentioned processor can also be configured to perform the following steps through a computer program: in response to the target object of attention including facial expressions and body movements, performing a confidence check on the display content of the source object in the multimedia file to obtain a detection result, wherein the display content includes: the expression content and action content of the source object; in response to the detection result meeting the first preset condition, obtaining the content to be extracted from the display content.
[0185] Optionally, the above-mentioned processor can also be configured to perform the following steps through a computer program: obtaining multiple limb skeleton key points corresponding to the source object; performing confidence detection on the action content based on the confidence values corresponding to the multiple limb skeleton key points and the confidence weights corresponding to the multiple limb skeleton key points to obtain the detection results, wherein the confidence values are used to distinguish the degree of visualization and / or clarity of different limb parts corresponding to the multiple limb skeleton key points, and the confidence weights are used to distinguish the importance of different limb parts corresponding to the multiple limb skeleton key points in the limb structure.
[0186] Optionally, the above-mentioned processor can also be configured to perform the following steps through a computer program: obtaining multiple facial bone key points corresponding to the source object; performing confidence detection on the expression content based on the confidence values corresponding to the multiple facial bone key points to obtain a detection result, wherein the confidence value is used to distinguish the degree of visualization and / or clarity of different facial areas corresponding to the multiple facial bone key points.
[0187] Optionally, the above-mentioned processor can also be configured to perform the following steps through a computer program: determine the limb skeleton key points corresponding to the limb movement based on the content to be extracted; perform three-dimensional motion restoration on the limb skeleton key points corresponding to the limb movement to obtain three-dimensional motion data; and use the body shape data of the target character to redirect the three-dimensional motion data to generate motion migration data.
[0188] Optionally, the above-mentioned processor can also be configured to perform the following steps through a computer program: determining initial expression data corresponding to the facial expression based on the content to be extracted; generating expression migration data corresponding to the initial expression data using a preset correspondence relationship, wherein the preset correspondence relationship is obtained through a pre-trained expression coding, and the expression coding is used to compare the coding distance between expression data of different characters.
[0189] Optionally, the above-mentioned processor can also be configured to perform the following steps through a computer program: in response to the multimedia file being a video file, locking the contour information of at least one object displayed in the video file; based on the locked contour information, counting the frequency of occurrence of at least one object in multiple frames of image material in the video file to obtain statistical results; in response to the statistical results satisfying a second preset condition, determining the content to be extracted from the multimedia file according to the target object of interest.
[0190] In the electronic device of the embodiment of the present disclosure, by selecting target attention objects from multiple object dimensions, determining the content to be extracted from the multimedia file to be identified according to the target attention objects, and then generating target migration data based on the content to be extracted, so as to control the target character in the virtual scene to generate a target animation that is adapted to the target migration data. Therefore, by selecting different target attention objects, various types of migration data can be obtained, thereby achieving the purpose of flexibly generating migration data corresponding to the content to be extracted in multiple object dimensions to render virtual character animations that meet custom needs, thereby achieving the technical effect of improving the flexibility of migration data generation methods, the diversity of animation rendering methods, and meeting the user's custom animation generation needs, and thus solving the technical problems of poor flexibility and single animation generation effect of the animation generation method provided in the related technology using AI to perform overall data migration of character actions and expressions in video materials.
[0191] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0192] S1, receiving multimedia files uploaded by users;
[0193] S2, in response to an object selection operation for multiple object identifiers, determining a target object of interest from multiple objects of interest;
[0194] S3, in response to the animation generation operation for the target virtual character, extracting target migration data from the multimedia file based on the target focus object, rendering the target virtual character based on the target migration data, and generating a character animation that matches the target migration data;
[0195] S4, play character animation.
[0196] Optionally, the above-mentioned processor can also be configured to perform the following steps through a computer program: in response to a single playback operation for facial expressions, control the target character to separately generate an expression animation that is compatible with the target migration data; in response to a single playback operation for body movements, control the target character to separately generate an action animation that is compatible with the target migration data; in response to a linked playback operation for facial expressions and body movements, control the target character to synchronously generate an expression animation and an action animation that are compatible with the target migration data.
[0197] Optionally, the processor may also be configured to execute the following steps through a computer program: in response to a single playback operation for a facial expression, while controlling the target character to separately generate an expression animation that matches the target migration data, synchronously playing the head animation associated with the expression animation.
[0198] In the electronic device of this embodiment, by responding to the object selection operation of the object identification of multiple focus objects, the target focus object is determined from the multiple focus objects, and the animation generation operation for the target virtual character is responded to, the target migration data is extracted from the multimedia file based on the target focus object, and the target virtual character is rendered based on the target migration data to generate and play the character animation matching the target migration data, thereby achieving the purpose of flexibly and personally editing the virtual character display animation on the mobile terminal, thereby achieving the technical effect of improving the animation display editing flexibility of the mobile terminal and meeting the user's customized animation display needs, and further solving the technical problem in the related technology that after the player uploads the video material, he cannot participate in the animation generation and display process of the virtual character, and thus it is difficult to meet the player's customized animation generation and display needs.
[0199] FIG15 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG15 , the electronic device 1500 is merely an example and should not limit the functions and scope of use of the embodiment of the present disclosure.
[0200] As shown in FIG15 , electronic device 1500 is implemented as a general-purpose computing device. Components of electronic device 1500 may include, but are not limited to, at least one processor 1510 , at least one memory 1520 , a bus 1530 connecting various system components (including memory 1520 and processor 1510 ), and a display 1540 .
[0201] The memory 1520 stores program codes, which can be executed by the processor 1510 , so that the processor 1510 executes the steps described in the method section of the embodiment of the present disclosure according to various exemplary embodiments of the present disclosure.
[0202] The memory 1520 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 15201 and / or a cache memory unit 15202, and may further include a read-only memory unit (ROM) 15203, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
[0203] In some examples, memory 1520 may also include a program / utility 15204 having a set (at least one) of program modules 15205. Such program modules 15205 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Memory 1520 may further include memory remotely located relative to processor 1510. These remote memories may be connected to electronic device 1500 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0204] The bus 1530 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a local bus to the processor 1510, or any of a variety of bus architectures.
[0205] The display 1540 may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 1500 .
[0206] Optionally, the electronic device 1500 can also communicate with one or more external devices 1600 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 1500, and / or any device that enables the electronic device 1500 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 1550. Furthermore, the electronic device 1500 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 1560. As shown in FIG. 15 , the network adapter 1560 communicates with other modules of the electronic device 1500 via a bus 1530. It should be understood that, although not shown in FIG. 15 , other hardware and / or software modules can be used in conjunction with the electronic device 1500, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0207] The electronic device 1500 may further include: a keyboard, a cursor control device (such as a mouse), an input / output interface (I / O interface), a network interface, a power supply and / or a camera.
[0208] Those skilled in the art will understand that the structure shown in FIG15 is for illustration only and does not limit the structure of the electronic device described above. For example, the electronic device 1500 may also include more or fewer components than those shown in FIG15 , or have a configuration different from that shown in FIG1 . The memory 1520 may be used to store computer programs and corresponding data, such as the computer program and corresponding data corresponding to the animation generation method in the embodiment of the present disclosure. The processor 1510 executes various functional applications and data processing by running the computer program stored in the memory 1520, thereby implementing the above-mentioned animation generation method.
[0209] The embodiments of the present disclosure further provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, which implements the method provided in the above embodiment when executed by a processor.
[0210] Optionally, the computer program included in the computer program product is used by a processor to execute the following steps:
[0211] S1, obtaining a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions;
[0212] S2, select the target object of attention from multiple object dimensions;
[0213] S3, determining the content to be extracted from the multimedia file according to the target object of interest;
[0214] S4, generating target migration data based on the content to be extracted;
[0215] S5, controlling the target character in the virtual scene to generate a target animation adapted to the target migration data.
[0216] Optionally, the computer program included in the computer program product is used by a processor to execute the following steps:
[0217] S1, receiving multimedia files uploaded by users;
[0218] S2, in response to an object selection operation for multiple object identifiers, determining a target object of interest from multiple objects of interest;
[0219] S3, in response to the animation generation operation for the target virtual character, extracting target migration data from the multimedia file based on the target focus object, rendering the target virtual character based on the target migration data, and generating a character animation that matches the target migration data;
[0220] S4, play character animation.
[0221] Embodiments of the present disclosure also provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium that can be used to store a computer program that, when executed by a processor, implements the method provided in the above embodiments.
[0222] Optionally, the computer program stored in the non-volatile computer-readable storage medium is used by the processor to execute the following steps:
[0223] S1, obtaining a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions;
[0224] S2, select the target object of attention from multiple object dimensions;
[0225] S3, determining the content to be extracted from the multimedia file according to the target object of interest;
[0226] S4, generating target migration data based on the content to be extracted;
[0227] S5, controlling the target character in the virtual scene to generate a target animation adapted to the target migration data.
[0228] Optionally, the computer program stored in the non-volatile computer-readable storage medium is used by the processor to execute the following steps:
[0229] S1, receiving multimedia files uploaded by users;
[0230] S2, in response to an object selection operation for multiple object identifiers, determining a target object of interest from multiple objects of interest;
[0231] S3, in response to the animation generation operation for the target virtual character, extracting target migration data from the multimedia file based on the target focus object, rendering the target virtual character based on the target migration data, and generating a character animation that matches the target migration data;
[0232] S4, play character animation.
[0233] The embodiments of the present disclosure further provide a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the method provided in the above embodiment is implemented.
[0234] Optionally, the computer program is executed by a processor to:
[0235] S1, obtaining a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions;
[0236] S2, select the target object of attention from multiple object dimensions;
[0237] S3, determining the content to be extracted from the multimedia file according to the target object of interest;
[0238] S4, generating target migration data based on the content to be extracted;
[0239] S5, controlling the target character in the virtual scene to generate a target animation adapted to the target migration data.
[0240] Optionally, the computer program is executed by a processor to:
[0241] S1, receiving multimedia files uploaded by users;
[0242] S2, in response to an object selection operation for multiple object identifiers, determining a target object of interest from multiple objects of interest;
[0243] S3, in response to the animation generation operation for the target virtual character, extracting target migration data from the multimedia file based on the target focus object, rendering the target virtual character based on the target migration data, and generating a character animation that matches the target migration data;
[0244] S4, play character animation.
[0245] The serial numbers of the above-mentioned embodiments of the present disclosure are for description only and do not represent the advantages or disadvantages of the embodiments.
[0246] In the above embodiments of the present disclosure, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0247] In the several embodiments provided in the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0248] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0249] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0250] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0251] The above is only a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as within the scope of protection of the present disclosure.
Claims
1. A method for generating an animation, comprising: Acquire a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of the source object from multiple object dimensions; Selecting a target object of interest from the multiple object dimensions; Determining content to be extracted from the multimedia file according to the target object of interest; generating target migration data based on the content to be extracted; The target character in the virtual scene is controlled to generate a target animation adapted to the target migration data.
2. The animation generation method according to claim 1, wherein: Determining the content to be extracted from the multimedia file according to the target object of interest includes: In response to the target object of attention including facial expressions and body movements, performing a confidence test on the display content of the source object in the multimedia file to obtain a test result, wherein the display content includes: expression content and movement content of the source object; In response to the detection result satisfying a first preset condition, the content to be extracted is obtained from the displayed content.
3. The animation generation method according to claim 2, wherein: Performing a confidence test on the displayed content of the source object in the multimedia file to obtain the test result includes: Obtaining multiple limb skeleton key points corresponding to the source object; The action content is subjected to confidence detection based on the confidence values corresponding to the multiple limb skeleton key points and the confidence weights corresponding to the multiple limb skeleton key points to obtain the detection result, wherein the confidence values are used to distinguish the degree of visualization and / or clarity of different limb parts corresponding to the multiple limb skeleton key points, and the confidence weights are used to distinguish the importance of different limb parts corresponding to the multiple limb skeleton key points in the limb structure.
4. The animation generation method according to claim 2, wherein: Performing a confidence test on the displayed content of the source object in the multimedia file to obtain the test result includes: Obtaining multiple facial bone key points corresponding to the source object; The expression content is subjected to confidence detection based on the confidence values corresponding to the multiple facial bone key points to obtain the detection result, wherein the confidence values are used to distinguish the degree of visualization and / or clarity of different facial areas corresponding to the multiple facial bone key points.
5. The animation generation method according to claim 2, wherein: Generating the target migration data based on the content to be extracted includes: Determine the limb skeleton key points corresponding to the limb action based on the content to be extracted; Performing three-dimensional motion restoration on limb skeleton key points corresponding to the limb motion to obtain three-dimensional motion data; The three-dimensional motion data is redirected using the body shape data of the target character to generate motion migration data.
6. The animation generation method according to claim 2, wherein: Generating the target migration data based on the content to be extracted includes: Determining initial expression data corresponding to the facial expression based on the content to be extracted; The expression migration data corresponding to the initial expression data is generated by using a preset corresponding relationship, wherein the preset corresponding relationship is obtained by a pre-trained expression coding, and the expression coding is used to compare the coding distance between expression data of different characters.
7. The animation generation method according to claim 1, wherein: The animation generation method further includes: In response to the multimedia file being a video file, locking contour information of at least one object displayed in the video file; Based on the locked contour information, counting the occurrence frequency of the at least one object in the multiple frames of image material in the video file to obtain a statistical result; In response to the statistical result satisfying a second preset condition, the content to be extracted is determined from the multimedia file according to the target object of interest.
8. An animation display method, comprising providing a graphical user interface (GUI) via a terminal device, wherein the GUI displays content including a target virtual character and object identifiers of multiple objects of interest in a virtual scene, the animation display method comprising: Receive multimedia files uploaded by users; In response to an object selection operation for a plurality of object identifiers, determining a target object of interest from the plurality of objects of interest; In response to an animation generation operation for the target virtual character, target migration data is extracted from the multimedia file based on the target object of interest, and the target virtual character is rendered based on the target migration data to generate a character animation matching the target migration data; Play the character animation.
9. The animation display method according to claim 8, wherein: Playing the character animation includes one of the following: In response to a single play operation for a facial expression, controlling the target character to individually generate an expression animation adapted to the target migration data; In response to a single playback operation for a body movement, controlling the target character to separately generate a movement animation adapted to the target migration data; In response to the linked playback operation for the facial expression and the body movement, the target character is controlled to synchronously generate expression animation and movement animation adapted to the target migration data.
10. The animation display method according to claim 9, wherein: The animation display method further includes: In response to a single play operation for the facial expression, while controlling the target character to independently generate the expression animation adapted to the target migration data, the head animation associated with the expression animation is played synchronously.
11. An animation generating device, comprising: An acquisition module, configured to acquire a multimedia file to be identified, wherein the multimedia file is used to provide original material, and the original material is used to describe the emotional expression of a source object from multiple object dimensions; A selection module, configured to select a target object of interest from the multiple object dimensions; A determination module, configured to determine content to be extracted from the multimedia file according to the target object of interest; A generating module, configured to generate target migration data based on the content to be extracted; The control module is used to control the target character in the virtual scene to generate a target animation adapted to the target migration data.
12. A computer-readable storage medium having a computer program stored therein, wherein: The computer program is configured to execute the animation generating method described in any one of claims 1 to 7 or the animation display method described in any one of claims 8 to 10 when executed by a processor.
13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the animation generation method described in any one of claims 1 to 7 or the animation display method described in any one of claims 8 to 10.
14. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the animation generation method according to any one of claims 1 to 7 or the animation display method according to any one of claims 8 to 10.
Citation Information
Patent Citations
Three-dimensional attitude determination method and device, electronic equipment and storage medium
CN113887319A
Virtual character control method, storage medium and electronic device
CN115033106A
Virtual image animation generation method and device, computer equipment and storage medium
CN116958344A
Expression generation method and device of virtual image, computer equipment and storage medium
CN117475042A
Animation generation method, storage medium, electronic device and computer program product
CN118537458A