Virtual anchor synthesis method and apparatus, computer device, and readable storage medium
By pre-establishing a library of puppet movements and using AI voice synthesis technology, virtual anchors are automatically generated, solving the problems of complex generation, high cost, and insufficient physical movements in existing technologies, and achieving efficient and low-cost virtual anchor broadcasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI MEDIA TECH
- Filing Date
- 2022-12-26
- Publication Date
- 2026-06-26
AI Technical Summary
Existing virtual anchor technology has a complex generation process, cannot achieve automated production, has a long generation cycle, high cost, insufficient coordination of body movements, and imperfect lip-sync, which affects the viewing experience.
A pre-built puppet motion library stores multiple sets of 3D motion data packages of puppet motion forms. Combined with AI voice synthesis technology, it automatically generates voice packages for news texts. The motion script is then inserted through the editing interface, and animation is synthesized based on timestamps to generate a virtual anchor.
It has enabled the automated production of virtual anchors, simplified the operation process, shortened the generation cycle, reduced costs, and can meet the needs of daily, large-scale, and long-duration news broadcasting. It also enhances the emotional color and rhythm through body movements, thereby improving the audience experience.
Smart Images

Figure CN115984466B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and more particularly to a method, apparatus, computer device, and readable storage medium for synthesizing virtual anchors. Background Technology
[0002] Virtual anchors are virtual characters who appear in video programs using virtual avatars. Used in video programs, virtual anchors can replace human hosts in delivering news, weather forecasts, providing narration, or hosting. Virtual anchors can also be applied to live video streaming or customer service, replacing human interaction with viewers or customers. Replacing human anchors with virtual anchors can solve problems such as high human costs and inconsistent work quality.
[0003] In current technology, the appearance of virtual characters is mostly achieved through human-driven animation. For example, some well-known virtual singing and dancing characters launched in China include "Xiao Yang," the virtual host of Hunan TV, as well as "Luo Tianyi," "Liu Yexi," and "AYAYI." These types of virtual characters mainly showcase dance moves to the audience, and their facial expressions and lip-syncing during rapping are not yet realistic enough.
[0004] In existing technologies, virtual anchors (or virtual hosts) typically do not have physical movements; instead, they only have head movements and lip movements to accompany voice broadcasts, such as CCTV's "Kang Xiaohui" and People's Daily's "Guoguo."
[0005] Typically, generating a virtual anchor for each dynamic broadcast relies on a complex acquisition device and specific equipment. Common motion capture technologies include optical motion capture, inertial motion capture, and AR face capture. The animated image of the virtual anchor is synthesized through lip-syncing and animated facial expression generation techniques. For example, the Chinese invention patent "A Voice-Driven Method and System for Generating Facial Animation" discloses: "When a facial image is given in video form, since the facial image in the video already has natural head movements, facial animation can be directly generated based on a trained A2FGAN model." This technique, commonly known as "face swapping," replaces facial expressions in a pre-recorded head movement video to achieve lip-syncing. However, because the pre-recorded head movement video may not perfectly match the broadcast content—for example, the head may still be swaying left and right when the broadcast pauses, or the head may not move when a nod is expected—each generation of a virtual anchor requires manual post-editing and synthesis, making automatic generation impossible.
[0006] Current virtual news anchor technology suffers from the following technical problems: The generation of virtual anchors has not achieved automated production; the process is complex, the generation cycle is long, post-production requires significant labor costs, and production efficiency is low. Due to the lack of coordinated body movements, the broadcasting process appears somewhat stiff to viewers. Lip-sync facial animation technology is not yet perfect; during long broadcasts, inconsistencies between lip movements and speech inevitably occur, which, if noticed by viewers, negatively impact the viewing experience. Furthermore, virtual anchors struggle to match their facial expressions and movements to the emotional tone of the content, resulting in a perceived stiffness in viewers over extended periods. Therefore, current virtual news anchor technology is insufficient to meet the requirements of daily, high-volume, and long-duration news broadcasting. Summary of the Invention
[0007] This application provides a method, apparatus, computer device, and readable storage medium for synthesizing virtual anchors, which overcomes the problems existing in the prior art or at least partially solves the above-mentioned problems.
[0008] The first aspect of this application provides a method for synthesizing a virtual anchor, comprising: pre-establishing a puppet action library, wherein the puppet action library stores multiple three-dimensional action data packages corresponding one-to-one with multiple sets of puppet action forms and multiple names corresponding one-to-one with multiple sets of puppet action forms, each three-dimensional action data package being used to synthesize a corresponding set of puppet action forms; automatically generating a voice package from the entire news article, the voice package being used to synthesize the broadcast voice, the voice package carrying multiple first timestamps, the multiple first timestamps corresponding one-to-one with multiple second timestamps of each sentence in the entire news article; constructing an editing interface for inserting action scripts into the entire news article, and the window of the editing interface having a selector for inserting action scripts; when the editing interface receives a first instruction that the selector is selected, calling and displaying the virtual anchor based on the puppet action... The selection interface for multiple sets of puppet action forms generated from all 3D motion data packages in the library, upon receiving a second instruction that any set of puppet action forms from the multiple sets of puppet action forms be selected, inserts the selected set of puppet action forms into a specified position in the entire news article according to the instructions of the first instruction. A script is generated at the specified position, which instructs the naming of the selected set of puppet action forms and a third timestamp matching the corresponding sentence in the news article. This process continues until a set of puppet action forms from the multiple sets of puppet action forms is inserted into each sentence in the entire news article. Based on the voice package and the 3D motion data packages corresponding to all the puppet action forms inserted into the entire news article, the animation is synthesized according to the correspondence between the first and third timestamps to generate a virtual anchor.
[0009] In one optional approach, the step of pre-establishing a puppet action library specifically includes: establishing a virtual puppet model with a basic skeleton; setting multiple sets of different puppet action forms based on the basic skeleton of the virtual puppet model; establishing a one-to-one correspondence between the multiple sets of puppet action forms and multiple preset time periods, storing the one-to-one correspondence to form a puppet action library; collecting three-dimensional action data packets that correspond one-to-one with the multiple sets of puppet action forms according to the one-to-one correspondence, and each set of puppet action forms has a corresponding name; and storing all puppet action forms, three-dimensional action data packets, and names in a one-to-one correspondence manner into the puppet action library.
[0010] In one alternative approach, the step of collecting three-dimensional motion data packets corresponding one-to-one with multiple sets of puppet movement forms according to a one-to-one correspondence relationship specifically includes: collecting multiple sets of three-dimensional motion data packets within multiple preset time periods according to the one-to-one correspondence relationship and the motion capture method of the puppet's body, with each set of three-dimensional motion data packets including data on the movement changes of the main joints of the head, hands, feet, and torso.
[0011] In one alternative approach, the step of automatically generating a voice package for broadcasting from the entire news text specifically includes: establishing a virtual anchor AI voice synthesis timbre, and automatically generating a voice package for broadcasting from the entire news text using AI voice synthesis capabilities based on the established AI voice synthesis timbre.
[0012] In one alternative approach, if there is one or more sentences in the entire news article that do not include one set of puppet action forms from among multiple sets of puppet action forms, then during animation synthesis, the animation is also synthesized according to the correspondence between the first and third timestamps based on the default 3D action data package to generate a virtual anchor. The default 3D action data package is pre-stored in the puppet action library and is used to generate the default puppet action form.
[0013] In one alternative approach, when a third instruction to click a text box is received in the editing interface, the speech playback progress bar of the synthesized speech is jumped to the corresponding position based on the cursor's position when the text box is clicked.
[0014] In one alternative approach, when the editing interface receives a fourth instruction to click on the speech playback progress bar of the synthesized speech, the cursor used to click the text box is moved to the corresponding position in the entire news article based on the position of the clicked playback progress bar.
[0015] A second aspect of this application provides a virtual anchor synthesis device, comprising: a pre-establishment module for pre-establishing a puppet action library; a storage module for storing multiple three-dimensional action data packets corresponding one-to-one with multiple sets of puppet action forms and multiple names corresponding one-to-one with multiple sets of puppet action forms in the puppet action library, each three-dimensional action data packet being used to synthesize a corresponding set of puppet action forms; a generation module for automatically generating a voice package from the entire news article, the voice package being used to synthesize broadcast voice, the voice package carrying multiple first timestamps, the multiple first timestamps corresponding one-to-one with multiple second timestamps of each sentence in the entire news article; a construction module for constructing an editing interface for inserting action scripts into the entire news article, and the window of the editing interface having a selector for inserting action scripts; a receiving module for receiving a first instruction when the selector is selected in the editing interface; and a call and display module for calling and displaying multiple sets of puppet action forms generated based on all three-dimensional action data packets in the puppet action library when the first instruction when the selector is selected is received in the editing interface. The system includes: a selection interface for puppet action forms; a receiving module, which receives a second instruction from the selection interface to select any one of multiple sets of puppet action forms; an insertion module, which, upon receiving the second instruction from the selection interface to select any one of multiple sets of puppet action forms, inserts the selected set of puppet action forms into a specified position in the entire news article according to the instructions of the first instruction; a generation module, which, upon receiving the second instruction from the selection interface to select any one of multiple sets of puppet action forms, generates a script at a specified position, the script indicating the naming of the selected set of puppet action forms and a third timestamp matching a sentence of news article corresponding to the specified position, until a set of puppet action forms is inserted into each sentence of news article in the entire news article; and a synthesis module, which synthesizes animations according to the correspondence between the first timestamp and the third timestamp based on the voice packet and the 3D motion data packet corresponding to all puppet action forms inserted into the entire news article, to generate a virtual anchor.
[0016] A third aspect of this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the virtual anchor synthesis method of any of the above embodiments.
[0017] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the virtual anchor synthesis method of any of the above embodiments.
[0018] In summary, when synthesizing a virtual anchor from a complete news script, after inserting a set of puppet action forms into each sentence of the news script that requires puppet action forms, animation can be synthesized based on the voice package synthesized from the entire news script and the 3D motion data package corresponding to all the puppet action forms inserted into the entire news script. This generates a virtual anchor, achieving automated virtual anchor production. The implementation steps are relatively simple, the operation is convenient, and the generation cycle is short, meeting the requirements of daily, large-scale, and long-duration news broadcasting. Furthermore, for different news scripts, only one set of puppet action forms needs to be inserted into each sentence of the entire news script to automatically synthesize a virtual anchor and provide different virtual anchor broadcasting video services, thereby significantly reducing virtual anchor production time and costs. Moreover, the virtual anchor's action forms can be automatically associated with keywords in the news script to achieve the technical effect of using body movements to enhance the emotional tone of the broadcast content; body movements can also be used to distract the audience, preventing viewers from noticing inaccurate lip-syncing. In addition, the insertion of puppet movements can be used to control the rhythm and duration of voice broadcasts, thus better meeting the actual needs of news text video broadcasting.
[0019] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, the specific implementation methods of this application are described below. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram illustrating the process of pre-establishing a puppet action library in some embodiments of this application.
[0022] Figure 2 This is a flowchart of a virtual anchor synthesis method after a puppet action library has been pre-established in some embodiments of this application;
[0023] Figure 3 This is a schematic diagram of the structure of the virtual anchor synthesis device in some embodiments of this application.
[0024] Figure 4This is a schematic diagram of the structure of a computer device in some embodiments of this application. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and drawings of this application are intended to cover non-exclusive inclusion.
[0027] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0028] Furthermore, the terms "first," "second," etc., in the specification and claims of this application or in the aforementioned drawings are used to distinguish different objects rather than to describe a specific order, and may explicitly or implicitly include one or more of the features.
[0029] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "connected," "linked," or "communication connection" should be interpreted broadly. For example, "connected," "linked," or "communication connection" can refer not only to a physical connection, but also to an electrical or signal connection. For instance, it can be a direct connection (physical connection), or an indirect connection through at least one intermediate component, as long as the circuit is connected. It can also refer to the internal connection between two components. A signal connection can refer not only to a circuit connection but also to a signal connection through a media, such as radio waves. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0030] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0031] In the description of this application, unless otherwise stated, "multiple" and "at least two" mean two or more (including two), and similarly, "multiple groups" and "at least two groups" mean two or more (including two groups).
[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0033] Figure 1 This is a flowchart illustrating a virtual anchor synthesis method in an embodiment of this application. In the virtual anchor synthesis method, a puppet action library is pre-established. The puppet action library stores multiple three-dimensional action data packets that correspond one-to-one with multiple sets of puppet action forms, as well as multiple names that correspond one-to-one with multiple sets of puppet action forms. Each three-dimensional action data packet is used to synthesize a corresponding set of puppet action forms.
[0034] Specifically, the steps for pre-establishing a puppet action library include:
[0035] Step S101: A virtual doll model with a basic skeleton is established using computer 3D modeling. The basic skeleton includes a skeletal structure and movable joints, which are used to match the packaging material of the person. In one embodiment of this application, the packaging material includes facial features, face shape, hairstyle, clothing, decorations, etc.
[0036] Step S102: Based on the basic skeleton of the virtual puppet model, set multiple different puppet action forms. Among them, different puppet action forms correspond to different technical parameters, thus forming a 3D (Three Dimensions) mathematical model, i.e., a virtual puppet model, that can express different puppet action forms according to the input technical parameters.
[0037] Step S103: Establish one-to-one correspondences between multiple sets of puppet action forms and multiple preset time periods, and store the one-to-one correspondences to form a puppet action library.
[0038] Step S104: Collect three-dimensional motion data packages corresponding one-to-one with multiple sets of puppet action forms according to the one-to-one correspondence. Each set of puppet action forms has a corresponding name. In one embodiment of this application, the action form of the virtual puppet model with a serious expression and hands clenched into fists in front of its chest is named "clenched fist".
[0039] Based on a one-to-one correspondence, multiple sets of 3D motion data packets are collected within several preset time periods using motion capture techniques for the person controlling the virtual anchor (the person conducting the live broadcast). Each set of 3D motion data packets includes data on the movement changes of the head, hands, feet, and major joints of the torso. For example, based on the one-to-one correspondence, a set of 3D motion data packets for a puppet's movement form is collected in different preset time periods. These 3D motion data packets are used to generate the corresponding puppet's movement form. Taking one of the multiple preset time periods as an example, say the first preset time period, a set of puppet movement forms suitable for news broadcasting is collected in the first preset time period using motion capture techniques for the person controlling the virtual anchor. During motion capture, the person controlling the virtual anchor, wearing motion-sensing clothing, performs movements within a pre-set collection environment. The collected data for each set of puppet movement forms includes data on the movement changes of the head, hands, feet, and major joints of the torso. This data is recorded as a 3D motion data packet. Similarly, multiple 3D motion data packets corresponding to other sets of puppet movement forms within different preset time periods are recorded.
[0040] Step S105: Store all the puppet's motion shapes, 3D motion data packages, and names in the puppet motion library in a one-to-one correspondence manner.
[0041] The puppet's movement patterns corresponding to all the recorded 3D motion data packets can be saved as videos for preview demonstration. The 3D motion data packets can be used as basic data for seamless motion connection during the final video synthesis.
[0042] It should be noted that the pre-establishment of the puppet action library is a one-time creation action. After the puppet action library is pre-established, for different news articles, you only need to insert a set of puppet action forms for each sentence of the news article that requires the insertion of puppet action forms. This will automatically synthesize virtual anchors to provide different virtual anchor broadcasting video services, thereby significantly reducing the production time of virtual anchors and reducing production costs.
[0043] In one embodiment of this application, when the keyword "bring them to justice" appears in a sentence of a news article, a set of puppet actions named "clenched fist" is inserted.
[0044] After establishing a library of puppet actions in advance, the method for synthesizing virtual anchors can be as follows.
[0045] Step S201: Automatically generate a speech package from the entire news article. The speech package is used to synthesize the broadcast speech. The speech package has multiple first timestamps, which correspond one-to-one with multiple second timestamps of each sentence in the entire news article. For example, the second timestamp of each sentence in the entire news article refers to the start time and end time of that sentence, so that the generated speech speed is not too fast or too slow according to the one-to-one correspondence between the first and second timestamps.
[0046] Specifically, an AI-generated voice timbre for virtual anchors can be established. Based on this timbre, the entire news text can be automatically generated into a voice package for broadcasting using AI voice synthesis capabilities. The process of establishing the AI-generated voice timbre for virtual anchors supports adjustments to speech rate, tone, pauses, and polyphonic character replacement. The timbre is established through voice fitting training on specific human speech materials for a preset duration, such as ten or even dozens of hours, to obtain the aforementioned voice package. Based on this voice package, during the interactive synthesis of the AI-generated voice timbre for virtual anchors, the speech rate, tone, pauses, and polyphonic character replacement can be adjusted. In one embodiment of this application, to extend the broadcast time by 0.5 seconds, after the virtual anchor pronounces "brought to justice," a set of puppet actions named "clenched fist" is inserted and maintained for 0.5 seconds. After the "clenched fist" action is completed, the previous action is restored, and the subsequent broadcast continues.
[0047] Step S202: Build an editing interface for inserting action scripts into the entire news article, and the window containing the editing interface has a selector for inserting action scripts.
[0048] Step S203: When the first instruction of the selector being selected is received in the editing interface, the selection interface of multiple sets of puppet action forms generated according to all three-dimensional action data packages in the puppet action library is invoked and displayed.
[0049] In step S204, when the selection interface receives a second instruction that any one of the multiple sets of puppet action forms is selected, the selected set of puppet action forms is inserted into a specified position in the entire news article according to the instructions of the first instruction, and a script is generated at the specified position. The generated script is used to indicate the naming of the selected set of puppet action forms and a third timestamp that matches a sentence of news article corresponding to the specified position. The third timestamp corresponds one-to-one with the second timestamp of the sentence of news article corresponding to the specified position, and also includes the start and end times of the sentence of news article corresponding to the specified position.
[0050] This process continues until each sentence of the news article requires the insertion of one set of puppet action forms from multiple sets of puppet action forms.
[0051] Step S205: Based on the voice packet and the 3D motion data packets corresponding to all the puppet action forms inserted into the entire news article, animation is synthesized according to the correspondence between the first and third timestamps to generate a virtual anchor. It is worth mentioning that the virtual anchor's lip-syncing can be optimized and fitted using the ripple parameters generated by the AI-synthesized voice, thereby achieving automated lip-syncing generation for the virtual anchor.
[0052] Specifically, a matching relationship is established between the puppet's movement patterns and the speech generated from the voice pack based on the correspondence between the first and third timestamps. To reflect the timing of the speech in the news script, the generated voice pack includes the first timestamp of each sentence from the news script. For example, the start time is marked at the beginning of each sentence, and the end time is marked at the end. The duration of each sentence in the news script is calculated based on its start and end times. This allows for the calculation of the position in the news script where the speech will be broadcast when a motion script is inserted, based on the duration of the sentence where the cursor is located. This makes the synthesized virtual anchor more realistic.
[0053] During animation compositing, the server executing this virtual anchor compositing method uses animation state machine technology to sequentially connect the 3D motion data packets of each referenced puppet action form. This allows the adjacent puppet action forms of the generated virtual anchor to transition smoothly, greatly simplifying the animation design and enabling the virtual anchor to complete the interconnection of specified actions according to task requirements.
[0054] In another optional approach, if a sentence or more in the news article does not contain one of the multiple sets of puppet action forms, then during animation compositing, the animation is also synthesized according to the correspondence between the first and third timestamps based on the default 3D motion data package to generate a virtual anchor. The default 3D motion data package is pre-stored in the puppet action library and used to generate the default puppet action form. In practical applications, the default puppet action form of the virtual anchor can default to a sitting or standing posture, but it is not a static image. For example, it is usually accompanied by some periodic and subtle changes, such as the virtual anchor blinking intermittently, shaking its head, twisting its body, or swaying its hair, thus showing the realism and naturalness of the virtual anchor. The 3D motion data package of this default puppet action form can also be collected and stored in the aforementioned puppet action library using the method of human motion capture, which will not be elaborated here.
[0055] In another alternative approach, when the editing interface receives a third instruction to click the text box, the playback progress bar of the synthesized speech is moved to the corresponding position based on the cursor's location when the text box is clicked. When the editing interface receives a fourth instruction to click the playback progress bar of the synthesized speech, the cursor used to click the text box is moved to the corresponding position within the entire news article based on the position of the playback progress bar.
[0056] The technical solution of this application embodiment, when synthesizing a virtual anchor based on an entire news script, inserts a set of puppet action forms into each sentence of the news script that requires puppet action forms. Then, it can synthesize animation based on the voice package synthesized from the entire news script and the three-dimensional motion data package corresponding to all puppet action forms inserted into the entire news script, thereby generating a virtual anchor. This achieves the effect of automated virtual anchor production, with relatively simple implementation steps, convenient operation, and short generation cycle. It can meet the requirements of daily, large-scale, and long-term news broadcasting. Furthermore, for different news scripts, it is only necessary to insert a set of puppet action forms into each sentence of the news script that requires puppet action forms to automatically synthesize a virtual anchor to provide different virtual anchor broadcasting video services, thereby significantly reducing the production time and production costs of virtual anchors.
[0057] Please see Figure 3 , Figure 3This is a schematic diagram of a virtual anchor synthesis device according to an embodiment of this application. The virtual anchor synthesis device includes: a pre-establishment module 301, a storage module 302, a generation module 303, a construction module 304, a receiving module 305, a call and display module 306, an insertion module 307, and a synthesis module 308. Furthermore, the pre-establishment module 301, storage module 302, generation module 303, construction module 304, receiving module 305, call and display module 306, insertion module 307, and synthesis module 308 are communicatively connected, for example, through a bus.
[0058] Among them, the pre-establishment module 301 is used to pre-establish the puppet action library.
[0059] The storage module 302 is used to store multiple three-dimensional motion data packets corresponding to multiple sets of puppet motion forms and multiple names corresponding to multiple sets of puppet motion forms in the puppet motion library. Each three-dimensional motion data packet is used to synthesize a corresponding set of puppet motion forms.
[0060] The generation module 303 is used to automatically generate a voice package from the entire news article. The voice package is used to synthesize the broadcast voice. The voice package has multiple first timestamps, and the multiple first timestamps correspond one-to-one with the multiple second timestamps of each sentence in the entire news article.
[0061] Module 304 is used to build an editing interface for inserting action scripts throughout a news article, and the window containing the editing interface has a selector for inserting action scripts.
[0062] The receiving module 305 is used to receive the first instruction that the selector is selected in the editing interface.
[0063] The display module 306 is invoked to call and display the selection interface for multiple sets of puppet action forms generated based on all three-dimensional action data packages in the puppet action library when the first instruction of the selector is received in the editing interface.
[0064] The receiving module 305 is also used to receive a second instruction that any one of the multiple sets of puppet action forms is selected in the selection interface.
[0065] The insertion module 307 is used to insert the selected set of puppet action forms into a specified position in the entire news article according to the instructions of the first instruction when a second instruction is received from the selection interface that any set of puppet action forms among multiple sets of puppet action forms is selected.
[0066] The generation module 303 is also used to generate a script at a specified position when the selection interface receives a second instruction that any one of the multiple sets of puppet action forms is selected. The script is used to indicate the name of the selected set of puppet action forms and a third timestamp that matches a sentence of news text corresponding to the specified position, until a set of puppet action forms is inserted into each sentence of the news text that requires the insertion of a set of puppet action forms from the multiple sets of puppet action forms.
[0067] The synthesis module 308 is used to synthesize animations based on the voice packets and the three-dimensional motion data packets corresponding to all the puppet action forms inserted into the entire news article, according to the correspondence between the first and third timestamps, to generate a virtual anchor.
[0068] The functions or effects of each module in the virtual anchor synthesis device in this embodiment correspond to the virtual anchor synthesis method in any of the above embodiments. The relevant details mentioned in the above embodiments of the virtual anchor synthesis method can be used in the virtual anchor synthesis device of this embodiment, and will not be repeated here.
[0069] In summary, the technical solution of this application embodiment, when synthesizing a virtual anchor based on an entire news script, inserts a set of puppet action forms into each sentence of the news script that requires puppet action forms. Then, it can synthesize animation based on the voice package synthesized from the entire news script and the three-dimensional motion data package corresponding to all puppet action forms inserted into the entire news script, thereby generating a virtual anchor. This achieves the effect of automated virtual anchor production, with relatively simple implementation steps, convenient operation, and a short generation cycle. It can meet the requirements of daily, large-scale, and long-term news broadcasting. Furthermore, for different news scripts, it is only necessary to insert a set of puppet action forms into each sentence of the news script that requires puppet action forms to automatically synthesize a virtual anchor to provide different virtual anchor broadcasting video services, thereby significantly reducing the production time and production costs of virtual anchors.
[0070] Another embodiment of this application also provides a computer device, such as... Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of a computer device. The computer device includes a memory 41 and a processor 42. The memory 41 stores a computer program, and the processor 42 executes the computer program to implement the steps of the virtual anchor synthesis method as described in any of the above embodiments.
[0071] In the technical solution of this application embodiment, since the processor 42 can execute computer programs to implement the steps of the virtual anchor synthesis method as described in any of the above embodiments, the technical solution of this application embodiment, when synthesizing a virtual anchor based on a whole news article, inserts a set of puppet action forms into each sentence of the news article that requires the insertion of puppet action forms. Then, it can perform animation synthesis based on the voice package synthesized from the whole news article and the three-dimensional action data package corresponding to all puppet action forms inserted into the whole news article to generate a virtual anchor. This can achieve the effect of automated virtual anchor production. The implementation steps are relatively simple, the operation is relatively convenient, and the generation cycle is short. It can meet the requirements of daily, large-scale, and long-term news broadcasting. Furthermore, for different news articles, it is only necessary to insert a set of puppet action forms into each sentence of the news article that requires the insertion of puppet action forms to automatically synthesize a virtual anchor to provide different virtual anchor broadcasting video services, thereby significantly reducing the production time of virtual anchors and reducing production costs.
[0072] The memory 41 can be used to store software programs and modules, such as the program instructions / modules corresponding to the virtual anchor synthesis method and apparatus in the embodiments of this application. The processor 42 executes various functional applications and data processing by running the software programs and modules stored in the memory 41, thereby realizing the virtual anchor synthesis method of any of the above embodiments.
[0073] The memory 41 may include high-speed random access memory, and may also include non-volatile memory or volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories, such as flash memory, hard disks, multimedia cards, card-type memories (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disks, optical disks, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 41 may be an internal storage unit of a virtual anchor synthesis device, such as the hard disk or RAM of the virtual anchor synthesis device.
[0074] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the virtual anchor synthesis method as described in any of the above embodiments.
[0075] In the technical solution of this application embodiment, since the computer program can implement the steps of the virtual anchor synthesis method as described in any of the above embodiments when executed by the processor, the technical solution of this application embodiment, when synthesizing a virtual anchor based on the entire news script, inserts a set of puppet action forms into each sentence of the news script that requires the insertion of puppet action forms. Then, it can perform animation synthesis based on the voice package synthesized from the entire news script and the three-dimensional action data package corresponding to all puppet action forms inserted into the entire news script to generate a virtual anchor. This can achieve the effect of automated virtual anchor production, with relatively simple implementation steps, convenient operation, and short generation cycle. It can meet the requirements of daily, large-scale, and long-term news broadcasting. Furthermore, for different news scripts, it is only necessary to insert a set of puppet action forms into each sentence of the news script that requires the insertion of puppet action forms to automatically synthesize a virtual anchor to provide different virtual anchor broadcasting video services, thereby significantly reducing the production time and production costs of virtual anchors.
[0076] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, whereby the memory is used to store program code or instructions, including computer operation instructions, and the processor is used to execute the program code or instructions of the emergency rescue communication method stored in the memory.
[0077] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0078] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0079] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for synthesizing virtual anchors, characterized in that, The virtual anchor synthesis method includes: A puppet action library is pre-established. The puppet action library stores multiple three-dimensional action data packages that correspond one-to-one with multiple sets of puppet action forms, as well as multiple names that correspond one-to-one with the multiple sets of puppet action forms. Each three-dimensional action data package is used to synthesize a corresponding set of puppet action forms. The entire news article is automatically generated into a voice package, which is used to synthesize and broadcast the audio. The voice package has multiple first timestamps, which correspond one-to-one with multiple second timestamps of each sentence in the entire news article. An editing interface is set up for inserting action scripts into the entire news article, and the window containing the editing interface has a selector for inserting the action scripts; In response to receiving a first instruction that the selector is selected in the editing interface, the selection interface for multiple sets of puppet action forms generated based on all three-dimensional action data packages in the puppet action library is invoked and displayed. In response to receiving a second instruction from the selection interface that any one of the multiple sets of puppet action forms is selected, the selected set of puppet action forms is inserted into a specified position in the entire news article according to the instruction of the first instruction, and a script is generated at the specified position. The script is used to indicate the name of the selected set of puppet action forms and a third timestamp that matches a sentence of news article corresponding to the specified position. Repeat the above insertion steps until each sentence of the news article that requires the insertion of one of the multiple sets of puppet action forms has the corresponding puppet action form and its script inserted. Based on the voice package and the three-dimensional motion data package corresponding to all the puppet action forms inserted into the entire news article, the animation is synthesized according to the correspondence between the first timestamp and the third timestamp to generate a virtual anchor.
2. The virtual anchor synthesis method according to claim 1, characterized in that, The step of pre-establishing the puppet action library specifically includes: Establish a virtual puppet model with a basic skeleton; Based on the basic skeleton of the virtual puppet model, multiple sets of different puppet action forms are set; Establish one-to-one correspondences between multiple sets of puppet action forms and multiple preset time periods, and store the one-to-one correspondences to form a puppet action library; According to the one-to-one correspondence, three-dimensional motion data packets corresponding to multiple sets of puppet motion forms are collected, and each set of puppet motion forms corresponds to a name. All the puppet's movement forms, the three-dimensional movement data packets, and the names are stored in the puppet movement library in a one-to-one correspondence manner.
3. The virtual anchor synthesis method according to claim 2, characterized in that, The steps of collecting three-dimensional motion data packets corresponding one-to-one with multiple sets of puppet motion forms according to the one-to-one correspondence relationship specifically include: Based on the one-to-one correspondence, multiple sets of three-dimensional motion data packets are collected within the multiple preset time periods using the motion capture method. Each set of three-dimensional motion data packets includes data on the movement changes of the main joints of the head, hands, feet, and torso.
4. The virtual anchor synthesis method according to claim 1, characterized in that, The step of automatically generating a voice package for broadcasting from the entire news transcript specifically includes: A virtual anchor AI voice synthesis timbre is established, and based on the established AI voice synthesis timbre, the entire news text is automatically generated into a voice package for broadcasting using AI voice synthesis capabilities.
5. The virtual anchor synthesis method according to claim 1, characterized in that, If there is one or more sentences in the entire news article that do not include one of the multiple sets of puppet action forms, then during the animation synthesis, the animation is also synthesized according to the default three-dimensional action data package and the correspondence between the first timestamp and the third timestamp to generate a virtual anchor. The default three-dimensional action data package is pre-stored in the puppet action library and is used to generate the default puppet action form.
6. The virtual anchor synthesis method according to claim 1, characterized in that, When the editing interface receives a third instruction to click the text box, the audio playback progress bar of the synthesized speech is moved to the corresponding position according to the cursor's position when the text box is clicked.
7. The virtual anchor synthesis method according to claim 1, characterized in that, When the editing interface receives the fourth instruction to click the audio playback progress bar of the synthesized audio, the cursor used to click the text box will jump to the corresponding position in the entire news article according to the position of the clicked progress bar.
8. A virtual anchor synthesis device, characterized in that, The virtual anchor synthesis device includes: A pre-built module is used to pre-build a library of puppet movements; The storage module is used to store multiple three-dimensional motion data packets corresponding to multiple sets of puppet motion forms and multiple names corresponding to the multiple sets of puppet motion forms in the puppet motion library. Each three-dimensional motion data packet is used to synthesize a corresponding set of puppet motion forms. The generation module is used to automatically generate a voice package from the entire news article. The voice package is used to synthesize and broadcast the voice. The voice package has multiple first timestamps, and the multiple first timestamps correspond one-to-one with the multiple second timestamps of each sentence in the entire news article. A module is provided for building an editing interface for inserting action scripts into the entire news article, and the window containing the editing interface has a selector for inserting the action scripts. A receiving module is configured to receive a first instruction that the selector is selected in the editing interface; The display module is invoked to call and display a selection interface for multiple sets of puppet action forms generated based on all three-dimensional action data packages in the puppet action library when the editing interface receives the first instruction that the selector is selected. The receiving module is also used to receive a second instruction that any one of the multiple sets of puppet action forms is selected in the selection interface. The insertion module is used to insert the selected group of puppet action forms into a specified position in the entire news article according to the instructions of the first instruction when the selection interface receives a second instruction that any one of the multiple groups of puppet action forms is selected. The generation module is further configured to generate a script at a specified position when the selection interface receives a second instruction that any one of the multiple sets of puppet action forms is selected. The script is used to indicate the name of the selected set of puppet action forms and a third timestamp that matches a news article sentence corresponding to the specified position, until a set of puppet action forms is inserted into each sentence of the news article in which a set of puppet action forms from the multiple sets of puppet action forms needs to be inserted. The synthesis module is used to synthesize animations based on the voice package and the three-dimensional motion data packages corresponding to all the puppet action forms inserted into the entire news article, according to the correspondence between the first timestamp and the third timestamp, to generate a virtual anchor.
9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the virtual anchor synthesis method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the virtual anchor synthesis method as described in any one of claims 1 to 7.