Semantic collaboration virtual anchor video generation method and device, equipment and medium
By analyzing the action and text correlation characteristics in the virtual anchor template video, combining the emotional expansion and audio generation of user text, the multimodal fusion of virtual anchor videos is achieved, and the problem of insufficient consistency in the existing technology is solved, and the performance of virtual anchors is improved.
Patent Information
- Application Number
- CN202510448207.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-24
AI Technical Summary
In the existing virtual anchor technology, the text, audio and action consistency in the video is poor, resulting in the performance of the virtual anchors being not natural and vivid enough.
By obtaining the correlation characteristics between the actions and text in each frame of the virtual anchor template video, the content is expanded in combination with the multi-dimensional text emotion of the initial user text, the user text audio is generated, and the image correlation characteristics, the update user text and the user text audio are weighted and fused to obtain the virtual anchor features, and finally the virtual anchor template video is updated.
It significantly improves the consistency of text, audio and actions in the virtual anchor’s video, makes the performance of the virtual anchor more natural and vivid, and enhances the user’s interactive experience.
Smart Images

Figure CN120201261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech semantics, and particularly to a method, device, equipment and medium for generating a semantic collaborative virtual anchor video. Background Art
[0002] With the continuous progress of technology, virtual anchor technology has gradually become an important application tool, especially in customer service, marketing, education and training, etc., showing great potential. A virtual anchor is a customer service system generated based on advanced technologies such as video generation, natural language processing, speech synthesis and real-time rendering, which uses a virtual image to interact with customers in real time through video.
[0003] In the field of medical and health, a virtual anchor can simulate the interaction between a doctor and a patient in an online consultation platform. Through speech synthesis and natural language processing technologies, the virtual anchor can provide preliminary suggestions, answer questions according to the patient's symptom description, and guide the patient's next medical treatment process.
[0004] In the field of fintech business, in scenarios such as live sales, customer service, and interactive training, a virtual anchor can be integrated into an intelligent investment advisor system. Through deep learning and data mining technologies, it analyzes data such as users' investment preferences and risk tolerance, and provides customized investment advice.
[0005] In the prior art, although a virtual anchor can express some basic emotions and moods through motion capture and speech synthesis, the current emotion model is still difficult to simulate complex human emotion fluctuations and delicate expression changes. A virtual anchor usually interacts with the audience through speech synthesis technology. However, there are still some deficiencies in speech synthesis technology, especially in terms of fluency and naturalness. Even when using the most advanced speech synthesis technology, the speech of the virtual anchor may still seem too "machinelike", lacking the subtle intonations and tone changes in natural language.
[0006] Therefore, in the current technology, there is also a problem of poor consistency among the text, audio and virtual anchor actions in the virtual anchor video. Summary of the Invention
[0007] The present invention provides a method, device, equipment and medium for generating a semantic collaborative virtual anchor video, and its main purpose is to solve the problem of poor consistency among the text, audio and virtual anchor actions in the virtual anchor video.
[0008] In a first aspect, to achieve the above object, a method for generating a semantic collaborative virtual anchor video provided by the present invention includes:
[0009] Obtain a virtual anchor template video, analyze the correlation features between actions and texts in each frame image of the virtual anchor template video to obtain image correlation features;
[0010] Obtain an initial user text, identify the multi-dimensional text sentiment of the initial user text, and use the multi-dimensional text sentiment to expand the content of the initial user text to obtain an updated user text;
[0011] Generate a user text audio using the updated user text;
[0012] Perform weighted fusion on the image correlation features, the updated user text, and the user text audio to obtain virtual anchor features;
[0013] Use the virtual anchor features to update the virtual anchor template video to obtain a complete virtual anchor video.
[0014] In a second aspect, the present invention also provides a semantic collaborative virtual anchor video generation device, including:
[0015] A feature analysis module, configured to obtain a virtual anchor template video, analyze the correlation features between actions and texts in each frame image of the virtual anchor template video to obtain image correlation features;
[0016] A text expansion module, configured to obtain an initial user text, identify the multi-dimensional text sentiment of the initial user text, and use the multi-dimensional text sentiment to expand the content of the initial user text to obtain an updated user text;
[0017] An audio generation module, configured to generate a user text audio using the updated user text;
[0018] A feature fusion module, configured to perform weighted fusion on the image correlation features, the updated user text, and the user text audio to obtain virtual anchor features;
[0019] A video generation module, configured to use the virtual anchor features to update the virtual anchor template video to obtain a complete virtual anchor video.
[0020] In a third aspect, the present invention also provides an electronic device, the electronic device includes:
[0021] At least one processor; and,
[0022] A memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores a computer program that can be executed by the at least one processor. When the computer program is executed by the at least one processor, the at least one processor is enabled to execute the semantic collaborative virtual anchor video generation method described above.
[0024] In a fourth aspect, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores at least one computer program. The at least one computer program is executed by a processor in an electronic device to implement the semantic collaborative virtual anchor video generation method described above.
[0025] The present invention obtains a virtual anchor template video, analyzes the correlation features between actions and texts in each frame image of the virtual anchor template video to obtain image correlation features. By matching the facial change features of the image with the corresponding text descriptions, a more flexible and vivid virtual anchor performance can be achieved. The initial user text is obtained, the multi-dimensional text sentiment of the initial user text is identified, and the multi-dimensional text sentiment is used to expand the content of the initial user text to obtain an updated user text. By analyzing the multi-dimensional text sentiment of the initial user text, the user's emotional state and needs can be accurately grasped, ensuring that the generated video content matches the user's expectations in terms of tone, emotion, and information transmission. The updated user text is used to generate user text audio. By generating customized audio content according to the user's audio needs and adjusting and noise-canceling the audio, it helps to improve the audio quality and naturalness of the virtual anchor video. The image correlation features, the updated user text, and the user text audio are weighted and fused to obtain virtual anchor features. Fusing information from different modalities (image, text, audio) can ensure that the virtual anchor presents a more natural and consistent performance in terms of vision, language, and sound. Using the virtual anchor features to update the virtual anchor template video to obtain a complete virtual anchor video can effectively improve the consistency of text, audio, and virtual anchor actions in the virtual anchor video. Description of the Drawings
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0027] Figure 1 It is a schematic diagram of an application environment of a semantic collaborative virtual anchor video generation method in an embodiment of the present invention;
[0028] Figure 2Flow diagram of a semantic collaborative virtual anchor video generation method provided by an embodiment of the present invention;
[0029] Figure 3 Flow diagram of analyzing image correlation features in a semantic collaborative virtual anchor video generation method provided by an embodiment of the present invention;
[0030] Figure 4 Flow diagram of generating user text audio in a semantic collaborative virtual anchor video generation method provided by an embodiment of the present invention;
[0031] Figure 5 Flow diagram of generating a complete virtual anchor video in a semantic collaborative virtual anchor video generation method provided by an embodiment of the present invention;
[0032] Figure 6 Module diagram of a semantic collaborative virtual anchor video generation device provided by an embodiment of the present invention;
[0033] Figure 7 Structural diagram of an electronic device for implementing a semantic collaborative virtual anchor video generation method provided by an embodiment of the present invention;
[0034] Figure 8 Another structural diagram of an electronic device for implementing a semantic collaborative virtual anchor video generation method provided by an embodiment of the present invention.
[0035] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0036] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand how the present disclosure uses technical means to solve technical problems and achieve the corresponding technical effects, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0037] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0038] An embodiment of the present application provides a semantic collaborative virtual anchor video generation method. The execution subject of the semantic collaborative virtual anchor video generation method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the apparatus provided in the embodiments of the present application. In other words, the semantic collaborative virtual anchor video generation method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0039] An embodiment of the present invention provides a semantic collaborative virtual anchor video generation method, which can be applied to, for example Figure 1In the application environment. Among them, the client communicates with the server through the network. The server can obtain the virtual anchor template video through the client, analyze the correlation features between actions and texts in each frame of the virtual anchor template video to obtain image correlation features. By matching the facial change features of the image with the corresponding text descriptions, a more flexible and vivid virtual anchor performance can be achieved. Obtain the initial user text, identify the multi-dimensional text sentiment of the initial user text, and use the multi-dimensional text sentiment to expand the content of the initial user text to obtain the updated user text. By analyzing the multi-dimensional text sentiment of the initial user text, the user's emotional state and needs can be accurately grasped, ensuring that the generated video content matches the user's expectations in terms of tone, emotion, and information transmission. Generate the user text audio using the updated user text. By generating customized audio content according to the user's audio needs and adjusting and noise-canceling the audio, it helps to improve the audio quality and naturalness of the virtual anchor video. Weightedly fuse the image correlation features, the updated user text, and the user text audio to obtain the virtual anchor features. Fusing information from different modalities (image, text, audio) can ensure that the virtual anchor presents a more natural and consistent performance in terms of vision, language, and sound. Update the virtual anchor template video using the virtual anchor features to obtain the complete virtual anchor video, which can effectively improve the consistency of text, audio, and virtual anchor actions in the virtual anchor video. Finally, output and feedback the complete virtual anchor video back to the client. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0040] The following makes an explanatory note for the specification of the present invention. The present invention adopts a semantic collaborative virtual anchor video generation technology based on multi-modalities. By analyzing the correlation features between image actions and texts in the virtual anchor template video, combined with the expansion of the initial user text and the generated audio, weighted fusion of the image correlation features, the updated user text, and the user text audio is achieved, thereby updating the virtual anchor template video and generating a personalized complete virtual anchor video.
[0041] Refer to Figure 2 As shown in the figure, it is a schematic flowchart of a semantic collaborative virtual anchor video generation method provided by an embodiment of the present invention. In this embodiment, the semantic collaborative virtual anchor video generation method includes:
[0042] S1. Obtain the virtual anchor template video, analyze the correlation features between actions and texts in each frame of the virtual anchor template video, and obtain the image correlation features.
[0043] In an embodiment of the present invention, a virtual anchor template video is obtained. The virtual anchor template video includes: the actions of the virtual anchor and the text information associated with the actions. Usually, the actions of the virtual anchor include dynamic information such as facial expressions and lip-sync. Among them, the facial expressions include key points of eyes, eyebrows, nose, mouth, etc. The virtual anchor template video is extracted frame by frame, and each frame is used as an image unit. The CLIP model is used to analyze the correlation features between the actions and the text in each frame of the image, and image correlation features are obtained.
[0044] In a specific scenario of medical health, the virtual anchor can act as a health assistant and interact with patients. Analyzing the correlation between the actions and the text of the virtual anchor can improve the communication ability of the virtual assistant, making it more capable of "emotional" responses, and helping patients better understand their health conditions and treatment plans. For example, when the virtual assistant performs certain combinations of actions and text, patients can perceive emotional feedback, thereby enhancing the patients' trust and acceptance of the treatment process.
[0045] In a specific scenario of fintech, by analyzing the correlation between the actions and the text of the virtual anchor, more natural and user-emotion-compliant conversations can be generated. For example, when the virtual assistant handles customer inquiries, it not only responds to customers verbally but also enhances the communication effect through dynamic feedback such as facial expressions and actions. Combining the correlation features between the actions and the text can provide more accurate financial services for users, such as recommending personalized financial products or solutions for users. Based on the dynamic correlation of the facial expressions, actions, and text content of the virtual anchor, the financial platform can analyze the emotional state and psychological needs of customers, thereby providing more personalized recommendations. For example, when a financial advisor interacts with a customer, the virtual anchor can adjust the consultation content or recommend investment products according to the customer's emotional fluctuations, verbal feedback, and action performance, thereby enhancing the customer experience and satisfaction.
[0046] As Figure 3 shown, it is a schematic flowchart of analyzing image correlation features in a method for generating a semantic collaborative virtual anchor video provided by an embodiment of the present invention.
[0047] Specifically, the analysis of the correlation features between the actions and the text in each frame of the virtual anchor template video to obtain image correlation features includes:
[0048] Extracting the facial feature points of each frame of the virtual anchor template video;
[0049] Randomly selecting one frame of the virtual anchor template video as the target frame image, and forming a target group image with the target frame image and its next frame image;
[0050] Obtaining the key point coordinates of the facial feature points of each frame of the image;
[0051] Construct a facial change feature based on the key point coordinates corresponding to the facial feature points in the target group of images;
[0052] Obtain the text description of the target group of images, map the text description and the target group of images into a preset embedding space to obtain a text vector and a target group of image vectors;
[0053] Calculate the similarity between the text vector and the target group of image vectors;
[0054] Match the facial change feature and the text description according to the similarity to obtain an image association feature.
[0055] Specifically, input each frame image of the virtual anchor template video into the Dlib model. Dlib provides a face detector based on HOG features and a linear classifier, as well as a face key point detector based on regression trees. For each frame image, Dlib will perform corresponding calculations, such as face recognition, gesture detection, etc., and return the results. These results usually include the positions of facial feature points and key point coordinates detected in the virtual anchor template video.
[0056] Specifically, the text descriptions of the target frame image and the next frame image of the target frame image refer to the text content used to generate the target frame image and the next frame image. The text description and the facial dynamic changes are mapped into a preset embedding space through the CLIP model. Specifically, the text description is input into the text encoder of CLIP. The CLIP text encoder will map the text into an embedding space to obtain a vector representation of the text, that is, a vector text. CLIP converts the input text into a point in the embedding space through a two-tower architecture; the target group of images is input into the image encoder of CLIP, and the image encoder will map the target group of images into the same embedding space to obtain the corresponding image vector representation, that is, the target group of image vectors.
[0057] Specifically, calculate the similarity between the text vector and the target group of image vectors. The calculation formula is as follows:
[0058]
[0059] Among them, A represents the text vector, and B represents the target group of image vectors.
[0060] Analyze the matching degree of the text description and the target group of images in the embedding space by calculating the similarity between the text vector and the target group of image vectors. If the similarity is greater than the preset similarity threshold, it means that the matching degree of the text description and the target group of images is relatively high. If the similarity is less than or equal to the preset similarity threshold, it means that the matching degree of the text description and the target group of images is relatively low.
[0061] Specifically, the facial change features and the text description are matched according to the similarity to obtain image association features. If the similarity is greater than a preset similarity threshold, it indicates a high matching degree between the text description and the target group of images. The facial change features corresponding to the target group of images and the text description are then correlated to obtain the final image association features. For example: The text description is a happy expression, and the facial change features of the target group of images are that the opening amplitude of the upper and lower lips is small and the upward amplitude of the corners of the mouth is large. If the similarity is greater than the preset similarity threshold, it indicates that the happy expression corresponds to the small opening amplitude of the upper and lower lips and the large upward amplitude of the corners of the mouth. The final obtained image association features are that the text description is a happy expression, and the facial change features are that the opening amplitude of the upper and lower lips is small and the upward amplitude of the corners of the mouth is large.
[0062] In an embodiment of the present invention, the constructing the facial change features according to the key point coordinates corresponding to the facial feature points in the target group of images includes:
[0063] Randomly select a facial feature point of a target frame image in the target group of images as the target feature point;
[0064] Take the key point coordinates of the target feature point as the first coordinates;
[0065] Take the key point coordinates of the target feature point of the next frame image of the target frame image in the target group of images as the second coordinates;
[0066] Calculate the Euclidean distance between the first coordinates and the second coordinates;
[0067] Generate a change amplitude according to the Euclidean distance;
[0068] Generate facial change features according to the target feature point and the change amplitude.
[0069] Specifically, the Euclidean distance between the first coordinates and the second coordinates is calculated, and the calculation formula is as follows:
[0070]
[0071] where (a1, b1) represents the first coordinates and (a2, b2) represents the second coordinates.
[0072] Specifically, change features are generated according to the Euclidean distance. The Euclidean distance represents the amplitude of facial changes. The larger the Euclidean distance, the more obvious the movement of the facial feature points and the greater the facial change amplitude; the smaller the Euclidean distance, the less obvious the movement of the facial feature points and the smaller the facial change amplitude.
[0073] Specifically, facial dynamic changes are generated based on the target feature points and the change amplitude. Facial dynamic changes are generated through comprehensive analysis of the target feature points and the change amplitude. For example, the target feature points are the upper and lower eyelids of the eyes, and the change amplitude is the up and down movement amplitude of the upper and lower eyelids. The facial change feature is opening or closing the eyes; the target feature points are the upper and lower lips of the mouth, and the change amplitude is the opening amplitude of the upper and lower lips and the up and down amplitude of the corners of the mouth. The facial change feature is speaking or smiling, etc.
[0074] By analyzing the association between the changes in facial feature points in each frame of the image and the text description, the interaction between facial expressions and language can be accurately captured, which can improve the naturalness and authenticity of the virtual anchor when expressing emotions and conveying information. By matching the facial change features of the image with the corresponding text description, a more flexible and vivid performance of the virtual anchor can be achieved, making it more interactive and immersive. The performance effect of the virtual anchor can be optimized, making the virtual anchor more in line with the emotional needs and situational context of the audience, and improving the audience's sense of participation and interactive experience.
[0075] S2. Obtain the initial user text, identify the multi-dimensional text emotion of the initial user text, and use the multi-dimensional text emotion to expand the content of the initial user text to obtain the updated user text.
[0076] In the embodiment of the present invention, the initial user text includes: emotional tendency, important key information, etc. Identify the multi-dimensional text emotion of the initial user text, and use the multi-dimensional text emotion to expand the short text into a more detailed description through a large language model (such as the GPT series) to provide more information for subsequent speech generation.
[0077] In the specific scenario of medical and health, by expanding and analyzing the emotions of the preliminary text provided by the patient (such as medical records, consultation information, or health logs), the medical platform can track the emotional changes of the patient, timely discover potential psychological problems (such as depression, anxiety, etc.), and provide corresponding mental health support.
[0078] In the specific scenario of fintech, financial institutions can more accurately grasp the needs and emotions of customers through the emotional analysis of customer interaction texts, and then adjust service strategies. For example, if a customer expresses confusion or dissatisfaction, the intelligent customer service system can automatically push more help information or transfer to a human customer service to improve customer satisfaction.
[0079] In the embodiment of the present invention, the identifying the multi-dimensional text emotion of the initial user text includes:
[0080] Converting the initial user text into a text encoding vector through a preset encoder;
[0081] Obtain multi-dimensional sentiment labels, and perform multi-layer fully connected operations on the text encoding vector using the multi-dimensional sentiment labels to obtain sentiment polarity, sentiment intensity, and sentiment type;
[0082] Generate multi-dimensional text sentiment according to the sentiment polarity, the sentiment intensity, and the sentiment type.
[0083] Specifically, the sentiment polarity includes: positive, negative, neutral; the sentiment intensity includes: weak, medium, strong; and the sentiment type includes: joy, anger, sadness, etc.
[0084] Convert the initial user text into a text encoding vector through a preset encoder. The specific operation steps are as follows: Input the initial user text into the preset encoder. The preset encoder calculates the weight of each word in the context of the initial user text using internal parameters, converts each word into a vector representation according to the weight, and finally obtains the text encoding vector. Among them, the preset encoder includes: transformer encoder, etc.
[0085] Obtain multi-dimensional sentiment labels, and perform multi-layer fully connected operations on the text encoding vector using the multi-dimensional sentiment labels to obtain sentiment polarity, sentiment intensity, and sentiment type. The specific operation steps are as follows: Design three fully connected layers as output layers. Divide the training task into three categories according to the multi-dimensional sentiment labels. Among them, for the sentiment polarity task: the first output layer is a softmax classification layer, which is used to predict the sentiment polarity (positive, negative, neutral) of the text, and the task output is a 3-class classification problem. For the sentiment intensity task: the second output layer is a softmax classification layer, which is used to predict the sentiment intensity (weak, medium, strong), and the task output is a 3-class classification problem. For the sentiment type task: the third output layer is a softmax classification layer, which is used to predict the sentiment type (such as joy, anger, sadness, etc.). Assuming there are multiple sentiment types, the output category is a classification problem with the total number of sentiment types. Input the text encoding vector into the classification layer of each task to obtain the sentiment polarity, sentiment intensity, and sentiment type.
[0086] Specifically, the content expansion of the initial user text using the multi-dimensional text sentiment to obtain an updated user text includes:
[0087] Obtain the text context of the initial user text and the expansion requirements of the preset user;
[0088] Generate an expansion style according to the multi-dimensional text sentiment and the text context;
[0089] Perform content expansion on the initial user text according to the expansion requirements and the expansion style to obtain an updated user text.
[0090] Specifically, the text context represents the usage environment, language environment, etc. of the initial user text. The expansion requirements indicate the direction of expansion determined based on the multi-dimensional text sentiment of the initial user text and the user's intention. For example, positive sentiment may require enhancing positive information, while negative sentiment may require providing comfort or suggesting solutions. An expanded style is generated based on the multi-dimensional text sentiment and the text context, such as: conventional and conservative text, creative or emotionally intense text, etc.
[0091] Specifically, the initial user text is content-expanded according to the expansion requirements and the expanded style to obtain an updated user text. In a specific scenario, the initial user text: "Hello, let's talk about the latest technology trends today." The expanded style: creative text. The expansion requirements: introduce specific technology trends, provide some expert opinions, and make the tone more attractive and interactive. The updated user text after expanding the initial user text through a large language model is: "Hello everyone, welcome to today's technology exploration! What we are going to talk about today is the most exciting technology trends in 2025! From artificial intelligence to quantum computing, there are so many exciting innovations this year! Do you know? According to expert predictions, AI will revolutionize all aspects of our daily lives in the coming months, not only in mobile assistants, but also in applications in industries such as healthcare and finance reaching a new peak! Now let's take a look together at which technology trends will lead the future. Which one are you most looking forward to? Leave me a message and tell me! Let's discuss it together!"
[0092] By analyzing the multi-dimensional text sentiment of the initial user text, the user's emotional state and needs can be accurately grasped, ensuring that the generated video content matches the user's expectations in terms of tone, emotion, and information transmission. The user's expansion requirements can guide the expansion of the virtual anchor's content, avoiding one-sided or overly simplistic expressions and enhancing the diversity and richness of the content. The expanded user text enables the virtual anchor to generate more personalized, engaging, and interactive videos, improving the viewing experience and participation of the audience. Text expansion can also enable the virtual anchor to handle complex and changing user needs, thereby improving the accuracy of the generated videos and the satisfaction of the audience.
[0093] S3. Generate user text audio using the updated user text.
[0094] In the embodiment of the present invention, the updated user text obtained by text expansion is preprocessed, and the processed text is used to generate audio through a text-to-speech (TTS) tool. The speed, tone, volume, etc. of the generated audio are adjusted according to requirements to obtain user text audio.
[0095] In specific medical and health scenarios, patients are provided with daily health management services, such as medication reminders, diet recommendations, exercise plans, etc. For example, the user's input is "What medications do I need to take today?" The virtual health assistant generates answers through text analysis and feeds the answers back to the patient in the form of voice through text-to-speech technology. The virtual health assistant is presented in the image of a virtual anchor, has the ability to interact in real time and communicate emotionally, can provide personalized health services, and help patients manage daily health problems.
[0096] In specific FinTech scenarios, financial institutions (such as banks and securities companies) use virtual anchors to provide customers with 24 / 7 automated services. For example, virtual anchors can help customers check account balances, transaction records, exchange rate information, or recommend financial products to customers. Users enter query text: for example, "I want to know what my balance is?" The system generates a suitable answer through text analysis, and uses text-to-speech (TTS) technology to convert the answer into natural language audio. The virtual anchor presents the answer in the video through facial expressions and body movements. Such interactions can make customers feel more intimate and enhance their sense of trust through the image of the virtual anchor.
[0097] like Figure 4 The figure is a flow chart of generating user text audio in a semantic collaborative virtual anchor video generation method provided by one embodiment of the present invention.
[0098] In detail, the generating user text audio by using the updated user text includes:
[0099] Acquire the audio language of a preset user, and generate an initial text audio from the updated user text according to the audio language;
[0100] Acquire the emotional features of the updated user text, and generate audio speech speed, audio intonation and audio volume according to the emotional features;
[0101] Adjusting the initial text audio according to the audio speech speed, the audio intonation and the audio volume to obtain an updated text audio;
[0102] A noise sample of the update text audio is obtained, and noise is searched and eliminated on the update text audio according to the noise sample to obtain the user text audio.
[0103] Specifically, the audio requirements include: Audio language: the language the user wants to use (e.g., Chinese, English, French, etc.). Speech speed: the speed at which the audio is played (e.g., fast, normal, slow, etc.). Intonation: the voice adjustment of the audio (e.g., smooth, rising, falling, passionate, etc.). Volume: the volume at which the audio is played (e.g., standard volume, increased volume, reduced volume, etc.).
[0104] Specifically, input the updated user text into the TTS tool to obtain the initial text audio. Customize and adjust the initial text audio according to the audio speech rate, audio intonation, and audio volume. Use audio editing software to denoise the adjusted audio. Find a part that only contains noise (the blank part without speech) in the updated text audio as the "noise sample". Use the "noise cancellation" function on the updated text audio to remove the sound characteristics of the selected noise sample from the entire audio, obtaining the user text audio.
[0105] By generating customized audio content according to the user's audio requirements, and adjusting and canceling noise of the audio, it helps to improve the audio quality and naturalness of the virtual anchor video. By adjusting the speech rate, intonation, and volume, the voice of the virtual anchor can be made more suitable for different scenarios and audience needs, enhancing the expressiveness and immersion. At the same time, the noise cancellation step ensures that the audio is clear and interference-free, improving the audience's auditory experience, thereby enhancing the professionalism and attractiveness of the virtual anchor video.
[0106] S4. Perform weighted fusion on the image association feature, the updated user text, and the user text audio to obtain the virtual anchor feature.
[0107] In the embodiment of the present invention, obtain the emotional feature of the updated user text, and perform modality fusion through multiple attention modules (such as attention module 1, attention module 2, and attention module 3) to process the relationship between each modality. The role of each attention module is to weight and mutually fuse the features of different modalities, calculate the relationship between the features of different modalities, and determine the weight of each modality feature through the attention mechanism. Perform weighted fusion on the image association feature, the updated user text, the user text audio, and the emotional feature according to the weight to obtain the virtual anchor feature.
[0108] In the specific scenario of fintech, the virtual anchor can provide personalized investment advice according to the user's emotional features and behavioral data. For example, during periods of stock market volatility, the virtual anchor can provide targeted investment strategies by analyzing the user's emotional state (such as anxiety or optimism) to help customers make calm decisions.
[0109] In the specific scenario of medical and health, the virtual anchor can act as a health popularization role in hospitals, health management platforms or applications, and provide customized health education for different patient groups in combination with emotional analysis. For example, for elderly patients, the virtual anchor can communicate in a warm and gentle tone to convey health care knowledge.
[0110] Specifically, the performing weighted fusion on the image association feature, the updated user text, and the user text audio to obtain the virtual anchor feature includes:
[0111] Map the image-associated features, the updated user text, and the user text audio into a preset shared space;
[0112] Calculate the minimization of a preset objective function based on the image-associated features, the updated user text, and the preset shared space to obtain a first fusion coefficient;
[0113] Weightedly combine the image-associated features and the updated user text according to the first fusion coefficient to obtain a first fusion result;
[0114] Obtain a first weight matrix and generate a first attention coefficient based on the first fusion result and the first weight matrix;
[0115] Generate a first associated feature of the image-associated features and the updated user text according to the first attention coefficient;
[0116] Calculate the minimization of the preset objective function based on the image-associated features, the user text audio, and the preset shared space to obtain a second fusion coefficient;
[0117] Weightedly combine the image-associated features and the user text audio according to the second fusion coefficient to obtain a second fusion result;
[0118] Obtain a second weight matrix and generate a second attention coefficient based on the second fusion result and the second weight matrix;
[0119] Generate a second associated feature of the image-associated features and the user text audio according to the second attention coefficient;
[0120] Calculate the minimization of the preset objective function based on the updated user text, the user text audio, and the preset shared space to obtain a third fusion coefficient;
[0121] Weightedly combine the updated user text and the user text audio according to the third fusion coefficient to obtain a third fusion result;
[0122] Obtain a third weight matrix and generate a third attention coefficient based on the third fusion result and the third weight matrix;
[0123] Generate a third associated feature of the updated user text and the user text audio according to the third attention coefficient;
[0124] Generate a virtual anchor feature of the preset shared space based on the first associated feature, the second associated feature, and the third associated feature.
[0125] Specifically, map the image-associated features, the updated user text, and the user text audio to the same hidden space, i.e., a preset shared space, through a hidden layer. In the shared space, the image-associated features, the updated user text, and the user text audio can be aligned to improve the relevance of the virtual anchor features.
[0126] By mapping the image-associated features, the updated user text, and the user text audio to a shared space, it is ensured that the information of these three modalities can be integrated and compared in the same representation space. The generated virtual anchor can better understand and integrate information from different modalities (such as vision, text, and audio), improve the multi-modal perception ability, make the generated virtual anchor perform more richly and naturally, and be able to respond more accurately to different inputs.
[0127] Specifically, calculate the minimization of a preset objective function based on the image-associated features, the updated user text, and the preset shared space to obtain a first fusion coefficient. The calculation formula is as follows:
[0128] L(w1)=||F1 - S|| 2 -||F2 - S|| 2
[0129] Where, w1 represents the first fusion coefficient, F1 represents the image-associated features, F2 represents the updated user text, and S represents the preset shared space.
[0130] By minimizing the objective function, optimize the relationship between the image-associated features and the updated user text to obtain a suitable fusion coefficient, which helps the virtual anchor better integrate information of different modalities, ensure that the importance of each input modality is correctly measured, and make the output generated by the virtual anchor more coordinated and natural, and can reflect the balance between multi-modalities.
[0131] Further, weight and combine the image-associated features and the updated user text according to the first fusion coefficient to obtain a first fusion result. The calculation formula is as follows:
[0132] Combined Feature1=w1*(F1 + F2)
[0133] Where, w1 represents the first fusion coefficient, F1 represents the image-associated features, and F2 represents the updated user text.
[0134] Generate a first attention coefficient based on the first fusion result and the first weight matrix learned through the Softmax function. The calculation formula is as follows:
[0135] Attention1 = softmax(W1 * Combined Feature1)
[0136] Among them, W1 represents the first weight matrix, and Combined Feature1 represents the first fusion result.
[0137] The image features and text features are weighted and combined, enabling the virtual anchor to correctly fuse visual information and text information according to the first fusion coefficient when generating output, helping the virtual anchor generate content that better meets user needs based on visual and language information, such as understanding and expressing user intentions or the environment, and generating more realistic interactions. The attention coefficient generated by the first weight matrix helps the virtual anchor pay more attention to certain key features during the generation process. This attention mechanism enhances the decision-making ability of the virtual anchor, enabling it to make more accurate and reasonable responses when processing complex inputs, and improving the user interaction experience.
[0138] Generate the image correlation feature and the first correlation feature of the updated user text according to the first attention coefficient. The calculation formula is as follows:
[0139]
[0140] Among them, Attention1 represents the first attention coefficient, and F i represents the image correlation feature and the updated user text.
[0141] The generated first correlation feature effectively fuses the information of the image and text, enabling the virtual anchor to more accurately perform emotional expression, action generation, or speech synthesis through the interaction between visual and text features.
[0142] Specifically, calculate the minimization of the preset objective function according to the image correlation feature, the user text audio, and the preset shared space to obtain the second fusion coefficient. The calculation formula is as follows:
[0143] L(w2) = ‖F1 - S‖ 2 -‖F3 - S‖ 2
[0144] Among them, w2 represents the second fusion coefficient, F1 represents the image correlation feature, F3 represents the user text audio, and S represents the preset shared space.
[0145] By minimizing the preset objective function, an optimal fusion coefficient can be found according to the characteristics of the image correlation feature and the user text audio, optimizing the fusion of features and making the fused representation of the image and text audio as close as possible to the preset goal.
[0146] Further, the image-related features and the user text audio are weighted and combined according to the third fusion coefficient to obtain a second fusion result. The calculation formula is as follows:
[0147] Combined Feature2 = w2(F1 + F3)
[0148] Among them, w2 represents the second fusion coefficient, F1 represents the image-related features, and F3 represents the user text audio. Weighted combination of these two features can adjust the importance of different features according to the second fusion coefficient, so as to ensure that important information is given priority in the fused features. Through weighted combination, the complementarity of images and text audio can be effectively utilized to improve the accuracy and performance of the representation.
[0149] The second weight matrix is learned through the Softmax function. According to the second fusion result and the second weight matrix, a second attention coefficient is generated. The calculation formula is as follows:
[0150] Attention2 = softmax(W2 * Combined Feature2)
[0151] Among them, W2 represents the second weight matrix, and Combined Feature2 represents the second fusion result. By optimizing the weight matrix, the contribution of different features in the final task can be better adjusted, enhancing feature selectivity and task relevance. The attention mechanism dynamically focuses on different parts according to the relevance of the input features. Through the second attention coefficient, the model's focus can be dynamically adjusted for different aspects of the image and text audio. This not only improves the model's ability to extract key information but also enables it to perform more precisely when dealing with complex tasks.
[0152] A second associated feature of the user text audio and the image-related features is generated according to the second attention coefficient. The calculation formula is as follows:
[0153]
[0154] Among them, Attention2 represents the second attention coefficient, and F i represents the user text audio and the image-related features.
[0155] The second associated feature generated through the second attention coefficient can effectively capture the fine-grained association between the image and the text audio, be able to more fully understand the interaction between the image and the text audio in the task, optimize the final output of the task, and improve the overall effect.
[0156] Specifically, by minimizing the preset objective function based on the updated user text, the user text audio, and the preset shared space, a third fusion coefficient is obtained, and the calculation formula is as follows:
[0157] L(w3) = ||F2 - S|| 2 -||F3 - S|| 2
[0158] where w3 represents the third fusion coefficient, F2 represents the updated user text, F3 represents the user text audio, and S represents the preset shared space.
[0159] By minimizing the objective function, the system can find the most suitable fusion coefficient among the updated user text, the user text audio, and the shared space, which helps to improve the collaborative effect of the model among multiple modalities (text and audio).
[0160] According to the third fusion coefficient, the updated user text and the user text audio are weighted and combined to obtain a third fusion result, and the calculation formula is as follows:
[0161] Combined Feature3 = w3(F2 + F3)
[0162] where w3 represents the third fusion coefficient, F2 represents the updated user text, and F3 represents the user text audio. Through weighted combination, the relative importance of the user text and the audio can be adjusted according to the third fusion coefficient, enabling better complementarity of the two modal information.
[0163] Based on the third weight matrix learned through the Softmax function, a third attention coefficient is generated according to the third fusion result and the third weight matrix, and the calculation formula is as follows:
[0164] Attention3 = softmax(W3 * Combined Feature3)
[0165] where W3 represents the third weight matrix, and Combined Feature3 represents the third fusion result. The third weight matrix can further impose importance adjustment on the fused features after weighted combination. By calculating the third attention coefficient, the key parts in the combined text and audio features can be more flexibly focused on.
[0166] The third correlation feature of the updated user text and the user text audio is generated according to the third attention coefficient, and the calculation formula is as follows:
[0167]
[0168] Among them, Attention3 represents the third attention coefficient, and F i represents the user text audio and updates the user text. Through the third attention coefficient, the deep connection between the updated user text and audio can be captured, which helps to understand the mutual relationship between the text and audio in semantics and context, so as to generate virtual anchor features more accurately and enhance the naturalness and accuracy of the output.
[0169] Specifically, the virtual anchor features are generated according to the first association feature, the second association feature and the third association feature, and the calculation formula is as follows:
[0170] Virtual Feature=f(Feature1,Feature2,Feature3)
[0171] Among them, Feature1 represents the first association feature, Feature2 represents the second association feature, Feature3 represents the third association feature, and f represents an addition function. By combining the first, second, and third association features, the features of images, texts, and audios are comprehensively integrated, and finally a highly integrated, rich, and dynamic virtual anchor feature is generated.
[0172] By weighted-fusing the image association feature, the updated user text, and the user text audio, and generating virtual anchor features, the expressiveness and interactivity of the virtual anchor video can be significantly improved. Specifically, fusing information of different modalities (image, text, audio) can ensure that the virtual anchor presents a more natural and consistent performance in terms of vision, language, and sound. Combining the attention mechanism and feature weighted fusion, the virtual anchor can dynamically adjust its performance mode according to the context, so as to achieve more accurate emotional expression, context understanding, and user interaction. This multi-modal fusion method makes the virtual anchor video more vivid and personalized, enhancing the user experience and sense of participation.
[0173] S5. Update the virtual anchor template video by using the virtual anchor features to obtain a complete virtual anchor video.
[0174] In the embodiment of the present invention, the actions of the virtual anchor are generated by a decoder according to the features of the virtual anchor. The features of the virtual anchor include facial expressions, eye movements, lip movements, etc. Apply the generated actions of the virtual anchor to the corresponding positions in the virtual anchor template video to obtain a complete virtual anchor video.
[0175] In specific fintech scenarios, virtual hosts regularly release information such as stock market analysis, investment advice, and market dynamics. Virtual hosts can quickly update market data to help users keep abreast of changes in the financial market at any time. Based on users' investment needs and risk preferences, virtual hosts can provide customized financial management solutions. This service can be presented in multimedia forms such as voice and video, enhancing customers' sense of participation and interactivity.
[0176] In specific healthcare scenarios, virtual hosts can act as expert narrators to spread content such as popularizing disease knowledge, explaining treatment plans, and instructions for drug use through videos. This form can improve the efficiency of information dissemination. Especially in telemedicine and medical services in remote areas, virtual hosts can provide convenient health education.
[0177] As Figure 5 shown, it is a schematic flowchart of generating a complete virtual host video in a semantic collaborative virtual host video generation method provided by an embodiment of the present invention.
[0178] Specifically, the updating of the virtual host template video using the virtual host features to obtain a complete virtual host video includes:
[0179] Using a preset decoder to generate a virtual host action sequence from the virtual host features;
[0180] Obtaining video frames corresponding to the virtual host action sequence and adding the virtual host action sequence to the virtual host template video according to the video frames to obtain an initial virtual host video;
[0181] Obtaining the video text, video audio, and virtual host video actions of each frame of the initial virtual host video;
[0182] Analyzing the synchronization of the video text, the video audio, and the virtual host video actions to obtain a synchronization score;
[0183] Judging whether the synchronization score is greater than a preset score threshold;
[0184] If the synchronization score is less than or equal to the score threshold, return to the step of weighted fusion of the image association features, the updated user text, and the user text audio to obtain virtual host features;
[0185] If the synchronization score is greater than the score threshold, use the initial virtual host video as the complete virtual host video.
[0186] Specifically, a decoder, such as an LSTM / Transformer decoder, takes an input feature vector and outputs time-series action data. After inputting the virtual anchor features into the decoder, a virtual anchor action sequence is output.
[0187] By calculating the cosine similarity between the video text, video audio, and virtual anchor video actions pairwise, the synchronization of the video text, video audio, and virtual anchor video actions is analyzed to obtain a synchronization score. The calculation formula is as follows:
[0188]
[0189] Among them, V1 represents the video text, V2 represents the video audio, and V3 represents the virtual anchor video actions.
[0190]
[0191] Among them, Sim1 represents the similarity between the video text and the video audio, Sim2 represents the similarity between the video audio and the virtual anchor video actions, and Sim3 represents the similarity between the video text and the virtual anchor video actions.
[0192] By continuously optimizing the video content of the virtual anchor, the naturalness and fluency of its performance can be significantly improved. By synchronously analyzing the correlation between the video text, audio, and actions, the coordination between the actions of the virtual anchor and the speech and language content is ensured, and unnatural visual or auditory deviations are avoided. If the synchronization of the initial video is insufficient, the system will automatically perform feature weighted fusion to optimize the matching of text, audio, and image features, thereby improving the performance quality of the virtual anchor. When the synchronization reaches the preset standard, a complete video will be generated to ensure that the video of the virtual anchor maintains high-quality synchronization and consistency both visually and auditorily, enhancing the viewing experience of the audience.
[0193] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0194] As Figure 6 shown, it is a functional module diagram of a semantic collaborative virtual anchor video generation device provided by an embodiment of the present invention.
[0195] In the embodiments of the present disclosure, a semantic collaborative virtual anchor video generation device is provided. The semantic collaborative virtual anchor video generation device corresponds one-to-one with the semantic collaborative virtual anchor video generation method in the above embodiments. As Figure 6As shown in the figure, the semantic collaborative virtual anchor video generation device 100 can be installed in an electronic device. According to the functions implemented, the semantic collaborative virtual anchor video generation device 100 includes a feature analysis module 101, a text expansion module 102, an audio generation module 103, a feature fusion module 104, and a video generation module 105. The detailed description of each functional module is as follows:
[0196] The feature analysis module 101 is used to obtain a virtual anchor template video, analyze the correlation features between actions and texts in each frame image of the virtual anchor template video, and obtain image correlation features;
[0197] The text expansion module 102 is used to obtain an initial user text, identify the multi-dimensional text emotion of the initial user text, and use the multi-dimensional text emotion to expand the content of the initial user text to obtain an updated user text;
[0198] The audio generation module 103 is used to generate user text audio using the updated user text;
[0199] The feature fusion module 104 is used to perform weighted fusion on the image correlation features, the updated user text, and the user text audio to obtain virtual anchor features;
[0200] The video generation module 105 is used to update the virtual anchor template video using the virtual anchor features to obtain a complete virtual anchor video.
[0201] In one embodiment, when the feature analysis module 101 executes to obtain a virtual anchor template video, analyze the correlation features between actions and texts in each frame image of the virtual anchor template video, and obtain image correlation features, it is used for:
[0202] Extract the facial feature points of each frame image of the virtual anchor template video;
[0203] Randomly select a frame image in the virtual anchor template video as the target frame image, and form a target group image with the target frame image and its next frame image;
[0204] Obtain the key point coordinates of the facial feature points of each frame image;
[0205] Construct facial change features according to the key point coordinates corresponding to the facial feature points in the target group image;
[0206] Obtain the text description of the target group image, map the text description and the target group image to a preset embedding space, and obtain a text vector and a target group image vector;
[0207] Calculate the similarity between the text vector and the target group image vector;
[0208] Match the facial change features and the text description according to the similarity to obtain image association features.
[0209] In one embodiment, when the feature analysis module 101 executes to obtain the virtual anchor template video and analyze the association features between actions and texts in each frame image of the virtual anchor template video to obtain image association features, it is used for:
[0210] Randomly select a facial feature point of the target frame image in the target group of images as the target feature point;
[0211] Take the key point coordinates of the target feature point as the first coordinate;
[0212] Take the key point coordinates of the target feature point of the next frame image of the target frame image in the target group of images as the second coordinate;
[0213] Calculate the Euclidean distance between the first coordinate and the second coordinate;
[0214] Generate a change amplitude according to the Euclidean distance;
[0215] Generate facial change features according to the target feature point and the change amplitude.
[0216] In one embodiment, when the text extension module 102 executes to recognize the multi-dimensional text emotion of the initial user text, it is used for:
[0217] Convert the initial user text into a text encoding vector through a preset encoder;
[0218] Obtain multi-dimensional emotion labels, and perform multi-layer fully connected on the text encoding vector by using the multi-dimensional emotion labels to obtain emotion polarity, emotion intensity, and emotion type;
[0219] Generate multi-dimensional text emotion according to the emotion polarity, the emotion intensity, and the emotion type.
[0220] In one embodiment, when the audio generation module 103 executes to generate user text audio by using the updated user text, it is used for:
[0221] Obtain the audio language of a preset user, and generate initial text audio according to the audio language for the updated user text;
[0222] Obtain the emotion feature of the updated user text, and generate audio speech rate, audio tone, and audio volume according to the emotion feature;
[0223] Adjust the initial text audio according to the audio speech rate, the audio intonation, and the audio volume to obtain an updated text audio;
[0224] Obtain a noise sample of the updated text audio, search for and eliminate noise from the updated text audio according to the noise sample to obtain a user text audio.
[0225] In one embodiment, when the feature fusion module 104 performs weighted fusion of the image-associated feature, the updated user text, and the user text audio to obtain a virtual anchor feature, it is used for:
[0226] Map the image-associated feature, the updated user text, and the user text audio into a preset shared space;
[0227] Calculate the minimization of a preset objective function according to the image-associated feature, the updated user text, and the preset shared space to obtain a first fusion coefficient;
[0228] Perform weighted combination of the image-associated feature and the updated user text according to the first fusion coefficient to obtain a first fusion result;
[0229] Obtain a first weight matrix, and generate a first attention coefficient according to the first fusion result and the first weight matrix;
[0230] Generate a first associated feature of the image-associated feature and the updated user text according to the first attention coefficient;
[0231] Calculate the minimization of the preset objective function according to the image-associated feature, the user text audio, and the preset shared space to obtain a second fusion coefficient;
[0232] Perform weighted combination of the image-associated feature and the user text audio according to the second fusion coefficient to obtain a second fusion result;
[0233] Obtain a second weight matrix, and generate a second attention coefficient according to the second fusion result and the second weight matrix;
[0234] Generate a second associated feature of the image-associated feature and the user text audio according to the second attention coefficient;
[0235] Calculate the minimization of the preset objective function according to the updated user text, the user text audio, and the preset shared space to obtain a third fusion coefficient;
[0236] Perform weighted combination of the updated user text and the user text audio according to the third fusion coefficient to obtain a third fusion result;
[0237] Obtain a third weight matrix, and generate a third attention coefficient according to the third fusion result and the third weight matrix;
[0238] Generate a third correlation feature of the updated user text and the user text audio according to the third attention coefficient;
[0239] Generate a virtual anchor feature of the preset shared space according to the first correlation feature, the second correlation feature, and the third correlation feature.
[0240] In one embodiment, when the video generation module 105 executes to update the virtual anchor template video by using the virtual anchor feature to obtain a complete virtual anchor video, it is used for:
[0241] Generate a virtual anchor action sequence from the virtual anchor feature by using a preset decoder;
[0242] Obtain video frames corresponding to the virtual anchor action sequence, and add the virtual anchor action sequence to the virtual anchor template video according to the video frames to obtain an initial virtual anchor video;
[0243] Obtain the video text, video audio, and virtual anchor video actions of each frame of the initial virtual anchor video;
[0244] Analyze the synchronization of the video text, the video audio, and the virtual anchor video actions to obtain a synchronization score;
[0245] Determine whether the synchronization score is greater than a preset score threshold;
[0246] If the synchronization score is less than or equal to the score threshold, return to the step of weighted fusion of the image correlation feature, the updated user text, and the user text audio to obtain a virtual anchor feature;
[0247] If the synchronization score is greater than the score threshold, use the initial virtual anchor video as the complete virtual anchor video.
[0248] In the present invention, for a semantic collaborative virtual anchor video generation device, first, the present invention obtains a virtual anchor template video, analyzes the correlation features between actions and texts in each frame image of the virtual anchor template video to obtain image correlation features. By matching the facial change features of the image with the corresponding text descriptions, a more flexible and vivid virtual anchor performance can be achieved. Then, an initial user text is obtained, the multi-dimensional text sentiment of the initial user text is identified, and the initial user text is content-expanded using the multi-dimensional text sentiment to obtain an updated user text. By analyzing the multi-dimensional text sentiment of the initial user text, the emotional state and needs of the user can be accurately grasped, ensuring that the generated video content matches the user's expectations in terms of tone, emotion, and information transmission. A user text audio is generated using the updated user text. By generating customized audio content according to the user's audio requirements and adjusting and removing noise from the audio, it helps to improve the audio quality and naturalness of the virtual anchor video. Further, the image correlation features, the updated user text, and the user text audio are weighted and fused to obtain virtual anchor features. Fusing information from different modalities (image, text, audio) can ensure that the virtual anchor presents a more natural and consistent performance in terms of vision, language, and sound. Finally, the virtual anchor template video is updated using the virtual anchor features to obtain a complete virtual anchor video, which can effectively improve the consistency of the text, audio, and virtual anchor actions in the virtual anchor video. The specific limitations regarding a semantic collaborative virtual anchor video generation device can refer to the limitations regarding a semantic collaborative virtual anchor video generation method in the above text and will not be elaborated here. Each module in the above semantic collaborative virtual anchor video generation device can be implemented in whole or in part through software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form to facilitate the processor to call and execute the operations corresponding to the above modules.
[0249] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media, and internal memory. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with external clients through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a semantic collaborative virtual anchor video generation method.
[0250] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a semantic collaborative virtual anchor video generation method.
[0251] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0252] Obtain a virtual anchor template video, analyze the association features between actions and texts in each frame image of the virtual anchor template video to obtain image association features;
[0253] Obtain an initial user text, identify the multi-dimensional text sentiment of the initial user text, and use the multi-dimensional text sentiment to expand the content of the initial user text to obtain an updated user text;
[0254] Generate a user text audio using the updated user text;
[0255] Perform weighted fusion on the image association features, the updated user text, and the user text audio to obtain virtual anchor features;
[0256] Update the virtual anchor template video using the virtual anchor features to obtain a complete virtual anchor video.
[0257] In several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0258] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0259] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claim concerned.
[0260] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0261] In some embodiments of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.
[0262] The computer-readable storage medium described in the present invention stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:
[0263] Obtain a virtual anchor template video, analyze the association features between actions and texts in each frame image of the virtual anchor template video to obtain image association features;
[0264] Obtain an initial user text, identify the multi-dimensional text sentiment of the initial user text, and use the multi-dimensional text sentiment to expand the content of the initial user text to obtain an updated user text;
[0265] Generate a user text audio using the updated user text;
[0266] Perform weighted fusion on the image association features, the updated user text, and the user text audio to obtain virtual anchor features;
[0267] Update the virtual anchor template video using the virtual anchor features to obtain a complete virtual anchor video.
[0268] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0269] The computer-readable storage medium may also store at least one computer-executable program / instructions, such as computer-readable instructions. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, the non-transitory computer-readable storage medium may be connected to a computing device such as a computer. Then, when the computing device runs the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.
[0270] In addition, the computer device may also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (such as a keyboard, a mouse, a speaker, etc.).
[0271] The processor may communicate with external devices via the I / O bus through a wired or wireless network.
[0272] In one embodiment, the at least one computer-executable instruction may also be compiled into or form a software product / computer program product, and when one or more computer-executable instructions are run by the processor, the various functions and / or method steps in the embodiments described in the present technology are performed.
[0273] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0274] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0275] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0276] It should be noted that in the present disclosure, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, the element limited by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0277] The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included within the protection scope of the present invention.
[0278] It should be noted that if non-company software tools or components appear in the embodiments of this application, they are only used for illustrative introduction and do not represent actual use.
Claims
1. A semantic collaborative virtual anchor video generation method, characterized in that: The method comprises: Acquire a virtual anchor template video, analyze correlation features between actions and texts in each frame of the virtual anchor template video, and obtain image correlation features; Acquire an initial user text, identify a multidimensional text sentiment of the initial user text, and use the multidimensional text sentiment to expand the content of the initial user text to obtain an updated user text; generating user text audio using the updated user text; Performing weighted fusion on the image-related features, the updated user text, and the user text audio to obtain virtual anchor features; The virtual anchor template video is updated using the virtual anchor feature to obtain a complete virtual anchor video.
2. The semantic collaborative virtual anchor video generation method according to claim 1, characterized in that: The analyzing the correlation features between the action and the text in each frame of the virtual anchor template video to obtain the image correlation features includes: Extracting facial feature points of each frame of the virtual anchor template video; Randomly select a frame image in the virtual anchor template video as a target frame image, and form a target group image with the target frame image and its next frame image; Obtaining the key point coordinates of the facial feature points of each frame of image; Constructing facial change features according to the key point coordinates corresponding to the facial feature points in the target group image; Acquire a text description of the target group image, map the text description and the target group image into a preset embedding space, and obtain a text vector and a target group image vector; Calculating the similarity between the text vector and the target group image vector; The facial change feature and the text description are matched according to the similarity to obtain image-related features.
3. The semantic collaborative virtual anchor video generation method according to claim 2, characterized in that: The step of constructing facial change features according to the key point coordinates corresponding to the facial feature points in the target group image includes: Randomly selecting a facial feature point of a target frame image in the target group of images as a target feature point; Taking the key point coordinates of the target feature point as the first coordinates; Taking the key point coordinates of the target feature point of the next frame image of the target frame image in the target group image as the second coordinates; Calculating the Euclidean distance between the first coordinate and the second coordinate; generating a magnitude of change according to the Euclidean distance; Generate facial change features according to the target feature points and the change amplitude.
4. The semantic collaborative virtual anchor video generation method according to claim 1, characterized in that: The identifying the multi-dimensional text emotion of the initial user text comprises: Converting the initial user text into a text encoding vector by a preset encoder; Acquire a multidimensional emotion label, and use the multidimensional emotion label to perform multi-layer full connection on the text encoding vector to obtain emotion polarity, emotion intensity and emotion type; A multi-dimensional text sentiment is generated according to the sentiment polarity, the sentiment intensity and the sentiment type.
5. The semantic collaborative virtual anchor video generation method according to claim 1, characterized in that: The step of generating user text audio using the updated user text comprises: Acquire the audio language of a preset user, and generate an initial text audio from the updated user text according to the audio language; Acquire the emotional features of the updated user text, and generate audio speech speed, audio intonation and audio volume according to the emotional features; Adjusting the initial text audio according to the audio speech speed, the audio intonation and the audio volume to obtain an updated text audio; A noise sample of the update text audio is obtained, and noise is searched and eliminated on the update text audio according to the noise sample to obtain the user text audio.
6. The semantic collaborative virtual anchor video generation method according to claim 1, characterized in that: The step of weightedly fusing the image-related features, the updated user text, and the user text audio to obtain virtual anchor features includes: Mapping the image-associated features, the updated user text, and the user text audio into a preset shared space; Calculate the minimization of a preset objective function according to the image association feature, the updated user text and the preset shared space to obtain a first fusion coefficient; weightedly merging the image-related features and the updated user text according to the first fusion coefficient to obtain a first fusion result; Obtain a first weight matrix, and generate a first attention coefficient according to the first fusion result and the first weight matrix; Generate the image association feature and the first association feature of the updated user text according to the first attention coefficient; Calculate the minimization of the preset objective function according to the image association feature, the user text audio and the preset shared space to obtain a second fusion coefficient; The image-related features and the user text audio are weightedly combined according to the second fusion coefficient to obtain a second fusion result; Obtain a second weight matrix, and generate a second attention coefficient according to the second fusion result and the second weight matrix; Generate the second associated feature of the image associated feature and the user text audio according to the second attention coefficient; Calculate the minimization of the preset objective function according to the updated user text, the user text audio and the preset shared space to obtain a third fusion coefficient; weightedly merging the updated user text and the user text audio according to the third fusion coefficient to obtain a third fusion result; Obtaining a third weight matrix, and generating a third attention coefficient according to the third fusion result and the third weight matrix; Generating a third correlation feature of the updated user text and the user text audio according to the third attention coefficient; A virtual anchor feature of the preset shared space is generated according to the first association feature, the second association feature and the third association feature.
7. The semantic collaborative virtual anchor video generation method according to claim 1, characterized in that: The updating of the virtual anchor template video by using the virtual anchor feature to obtain the complete virtual anchor video includes: Using a preset decoder to generate a virtual anchor action sequence from the virtual anchor features; Acquire a video frame corresponding to the virtual anchor action sequence, and add the virtual anchor action sequence to the virtual anchor template video according to the video frame to obtain a virtual anchor initial video; Obtaining the video text, video audio and video action of the virtual anchor in each frame of the initial video of the virtual anchor; Analyzing the synchronization of the video text, the video audio, and the video action of the virtual anchor to obtain a synchronization score; Determining whether the synchronization score is greater than a preset score threshold; If the synchronization score is less than or equal to the score threshold, returning to the step of weighted fusion of the image-related features, the updated user text, and the user text audio to obtain virtual anchor features; If the synchronization score is greater than the score threshold, the virtual anchor's initial video is used as the virtual anchor's complete video.
8. A semantic collaborative virtual anchor video generation device, characterized in that: The device comprises: A feature analysis module is used to obtain a virtual anchor template video, analyze the correlation features between actions and texts in each frame of the virtual anchor template video, and obtain image correlation features; A text expansion module, used to obtain an initial user text, identify the multi-dimensional text sentiment of the initial user text, and use the multi-dimensional text sentiment to expand the content of the initial user text to obtain an updated user text; An audio generation module, used to generate user text audio using the updated user text; A feature fusion module, used for weighted fusion of the image-related features, the updated user text and the user text audio to obtain virtual anchor features; The video generation module is used to update the virtual anchor template video by using the virtual anchor characteristics to obtain a complete virtual anchor video.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute a semantic collaborative virtual anchor video generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a semantic collaborative virtual anchor video generation method as described in any one of claims 1 to 7 is implemented.