Multi-agent-based news generation and digital people broadcasting method and system
By using multi-agent interaction generation and multimodal data processing, the problems of poor content quality and fragmented workflow in news broadcasting are solved, achieving efficient and professional news broadcasting effects, and is applicable to news generation and digital human broadcasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies suffer from poor content quality, fragmented workflows, and poor digital human broadcasting effects in news broadcasting, especially in multilingual TTS systems which lack professional voice features and digital human video generation systems which struggle to support the multimodal fusion needs of text and image news.
A multi-agent-based news generation method is adopted, which generates news text sequences through multiple rounds of iteration between generating and detecting agents. It then combines ChatTTS for speech synthesis to achieve prosodic control and noise reduction, performs multimodal data alignment, generates lip movement sequences and fuses facial features, and finally generates news broadcast videos that meet broadcast standards through a robust video keying model and subtitle processing.
It improves the factual accuracy, logic, and professionalism of news generation, ensures natural and clear audio, and enhances the stability and standardization of video, thereby increasing the system's generation efficiency and adaptability to meet the needs of news scenarios.
Smart Images

Figure CN121815041A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence, and particularly relates to a news generation and digital person broadcasting method and system based on multiple agents. BACKGROUND
[0002] At present, the comprehensive penetration of artificial intelligence technology in the field of news dissemination is a phenomenon-level development. With the change of social and cultural environment, the traditional news production mode presents the disadvantages of long production cycle, high input cost and insufficient digitalization, and the demand for media intelligent transformation development is urgent. However, the existing general generative AI still has obvious shortcomings in the production of news works: it only focuses on a single link in the news editing and broadcasting process, lacks overall coordination ability throughout the news production process, and the digital person system lacks vertical professional news dissemination scene driving, and cannot meet the requirements of accuracy and audio-visual effect in news production. In addition, there are still problems in the multi-modal fusion technology such as text, voice and video in digital person broadcasting.
[0003] Current domestic news broadcasting technology research focuses on algorithm development and system construction. In the field of text-to-speech (TTS), the Fish Speech system proposed by scholars innovatively uses large language models (LLM) and dual aggregation (Dual-AR) architecture to replace traditional grapheme-to-phoneme (G2P) methods, significantly improving the naturalness and scalability of Chinese, English and Japanese speech synthesis. The ChatTTS system (2024) developed by domestic teams breaks through the difficulty of processing mixed Chinese and English contexts, realizing smooth synthesis of cross-language speech. In the aspect of digital person generation theory, the ParaLip non-autoregressive lip shape synchronization model proposed by domestic researchers reduces the inference delay by 67% by improving the decoding mechanism, while maintaining the accuracy of lip shape generation. In addition, a deep learning-based full-automatic portrait matting algorithm has been proposed, which combines semantic segmentation branch, detail branch and hybrid branch learning methods to realize full-automatic and high-precision matting of characters in images. At the commercial application level, the "Yanyan" 3D digital person system has realized the function of video text-driven broadcasting, but it still has the defect of insufficient scene customization ability.
[0004] Foreign scholars systematically discussed the transformative impact of deep learning on the audio generation paradigm, providing theoretical support for speech synthesis technology. In the practice of TTS, the Orpheus-TTS system developed by foreign scholar teams implements emotional speech generation based on the Llama-3b architecture, but its training data is mainly based on English daily conversations, and has not yet covered the news broadcast scenario. Some scholars have developed a Wav2Lip lip synchronization model that can produce more accurate lip synchronization effects in dynamic and unconstrained speaker face videos. In addition, the NewsRobot system designed by foreign researchers achieves automatic broadcast of news articles through API integration, but its video synthesis module relies on web crawling materials and does not combine digital human technology to generate anchor news video. In addition, the librosa algorithm developed by some researchers can provide audio signal processing, feature extraction, rhythm analysis and other functions, and is widely used in audio processing, music analysis, speech recognition and other fields.
[0005] In summary, the existing technology has the following limitations: multi-language TTS systems mainly serve interactive entertainment scenarios and lack the professional speech features required for news broadcasting; digital human video generation systems are difficult to support the multi-modal fusion needs of graphic news; there is currently a lack of complete workflows that combine the fields of research at home and abroad, and the workflows are fragmented.
[0006] Therefore, there is an urgent need for a method to address the issues of poor content quality of news works, fragmented workflows, and poor digital human news broadcasting effects. SUMMARY
[0007] To solve the above technical problems, the present application provides a multi-agent based news generation and digital human broadcasting method, comprising the following steps:
[0008] Step S1: According to the user input news abstract, after multiple iterations of the generation agent and the detection agent, the news text sequence is generated in combination with the preset rules ;
[0009] Step S2: Construct a speech synthesis network based on ChatTTS to convert into synthesized speech , and realize rhythm control, pause optimization and noise reduction processing during synthesis;
[0010] Step S3: Based on the user input news broadcast video and , perform multi-modal data alignment to generate a lip action sequence, and fuse and enhance the facial features and background images of the news broadcast video to obtain the final video ;
[0011] Step S4: Convert After dynamic segmentation by a robust video matting model, synthesis is performed with the target background input by the user, and post-processing such as subtitle creation alignment, bottom strip and logo addition is performed to generate a news broadcast piece that meets the broadcast standards.
[0012] The present application provides a news generation and digital person broadcasting method based on multiple agents, which has the following advantages:
[0013] (1) Iterative agent interaction generation: multiple rounds of interaction between generation and detection agents optimize news text, improve factualness, logicality and coherence, and reduce semantic jumps and expression confusion.
[0014] (2) Multi-dimensional automatic evaluation: through multiple indicators such as BLEU and METEOR, comprehensive measurement from morphology, semantics to subjective expression is realized to accurately identify defects and ensure news professionalism and readability.
[0015] (3) In the speech synthesis link, a joint technology route of text semantic coding, prosody prediction, acoustic modeling and speech enhancement is introduced to restore the rhythm and stress of news broadcasting, and through spectral reconstruction and noise reduction, the synthesized speech is natural, clear and professional.
[0016] (4) Digital person video optimization: combining phoneme alignment, lip shape prediction, face enhancement and robust matting technology, audio-visual synchronization is realized, edge processing and subtitle packaging are optimized, and picture stability and standardization are improved.
[0017] (5) Multi-agent collaboration: content generation, quality evaluation and multi-modal processing are divided and cooperated to enhance system generation efficiency, quality stability and self-adaptive adjustment ability, and adapt to the needs of news scenes. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 A news generation and digital person broadcasting method based on multiple agents according to the present application is shown in the flowchart;
[0019] Figure 2 A method flowchart of the present application is shown;
[0020] Figure 3 A structural block diagram of a news generation and digital person broadcasting system based on multiple agents according to the present application is shown. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0022] Embodiment one:
[0023] As Figure 1 shown, the embodiment of the application provides a news generation and digital person broadcasting method based on multiple agents, which comprises the following steps:
[0024] Step S1: generating a news text sequence according to a user input news abstract, through multiple iterations of a generation agent and a detection agent, and combining a preset rule ;
[0025] Step S2: constructing a speech synthesis network based on ChatTTS to convert the news text sequence into synthesized speech , and realizing prosody control, pause optimization and noise reduction processing in the synthesis process;
[0026] Step S3: based on a user input news broadcasting video and , performing multi-modal data alignment to generate a lip action sequence, and fusing and enhancing the face features and background images of the news broadcasting video to obtain a final video ;
[0027] Step S4: after dynamic segmentation by a robust video cutout model, synthesizing with a user input target background, and then through post-processing such as subtitle creation alignment, bottom strip and logo addition, a news broadcasting piece meeting the broadcasting standards is generated.
[0028] In one embodiment, the above step S1: generating a news text sequence according to a user input news abstract, through multiple iterations of a generation agent and a detection agent, and combining a preset rule , specifically comprises:
[0029] Step S11: constructing a news data set :
[0030] ;
[0031] Among them, is the unique identifier of the news; is the news abstract; is the news written by a human; is the news generated by the agent based on ; is used to save the news generated by the agent in the last round, which is initially empty; is the news generated by the agent in the current round; is used to save the evaluation of the news generated by the agent in the last round; saves the of the 5 news with the largest cosine similarity with the current news.
[0032] Step S12: detecting the intelligent agent represented as:
[0033] ;
[0034] ;
[0035] ;
[0036] wherein, the news of the first round of evaluation results, and stored to field of the news, is a function of detecting the intelligent agent to evaluate the quality of the news, index concatenation of the news of the first round of evaluation results and the five most similar news to it
[0037] is the evaluation standard of the evaluation task, including title, lead, body, and conclusion; news target, including , , ; for concatenating , , , the comparison result of the current round and the last round of evaluation matrix, the single sample information block based on the optimization direction prompt of the comparison result;
[0038] and other cosine similarity function of set of functions of the five most similar news semantic vector of the news, semantic vector inner product, is the L2 norm, is the function of taking out the maximum similarity 5 news .
[0039] The invention calculates the semantic vector of the news summary by the SentenceTransformer model, and screens out the 5 most similar samples to the target news by combining the cosine similarity , and integrates these similar samples as examples into the prompt words. This mechanism enables the model to refer to more relevant cases when evaluating and generating news, improving the relevance of the prompt words and the generation / evaluation quality. To improve the running rate, fill in the similar sample number in the JSON data file for subsequent iteration process.
[0040] Step S13: generating an agent representation as:
[0041] ;
[0042] ;
[0043] wherein, is is the first round of generation results of the news , and stored to is the field of the news , is a function of generating news generated by the agent, is is the index of the news splicing agent prompt word;
[0044] is the generation task instruction head, including the learning goal and example explanation; is used to splice is the , (if empty, use ), field, comparison result of the current round and the last round evaluation matrix, single sample information block based on the optimization direction prompt of the comparison result; is the target generation instruction, including is the of the news and format requirements.
[0045] Step S14: After each iteration, the news generated by the agent and the human news are compared and evaluated using the evaluation index matrix, which can be represented as:
[0046] ;
[0047] ;
[0048] ;
[0049] ;
[0050] ;
[0051] ;
[0052] wherein, is used to measure the n-gram overlap between the candidate text and the reference text; is the n-gram precision, is the n-gram weight, is the length penalty factor; is the maximum order of n-gram;
[0053] is used to fuse stem matching, synonym matching, and chunk matching; is the harmonic mean of precision and recall, is the stem matching rate, is the coverage rate of chunk, and are adjustment parameters;
[0054] is used to focus on the longest common subsequence LCS between the candidate text and the reference text; is the recall, i.e., the length of the longest common subsequence LCS between the candidate text and the reference text as a proportion of the length of the reference text; is the precision, i.e., the length of the LCS as a proportion of the length of the candidate text; is the harmonic mean of recall and precision;
[0055] is the word embedding similarity based on a pre-trained language model; is the token sequence of the candidate text, is the token sequence of the reference text, is used to calculate the cosine similarity, is the token embedding output by the model; is the first token in the token sequence of the candidate text a token; is the reference text token sequence; is the reference text token sequence;
[0056] The score of the LLM output is normalized to the interval [0, 1]; The score of the LLM output is normalized to the interval [0, 1]; is the average of multiple scores, is the maximum value of the score scale.
[0057] After the end of a round of iteration process, the evaluation index matrix result is automatically stored in JSON format, and then the data set is updated for the next round of iteration.
[0058] The evaluation index matrix designed by the application adopts indexes including traditional NLP indexes (BLEU, METEOR, ROUGE-L), semantic similarity indexes (BERTScore) and semantic subjective scoring indexes (G-Eval), compares and evaluates the generated agent news and human news, comprehensively measures the news generation quality from multiple dimensions, and realizes the continuous iteration and improvement of the generation effect.
[0059] The application processes the news generation and evaluation tasks of the generated agent and the detection agent in parallel through a thread pool, adopts atomic operations and thread locks, ensures the consistency of JSON file reading and writing in a multi-threaded environment through file locks and thread locks, and avoids data competition.
[0060] In terms of robustness, the application combines a token truncation mechanism to control the input length:
[0061] ;
[0062] So as to ensure that the input does not exceed the maximum token limit of the model. For API rate limiting, an exponential backoff retry strategy is adopted, and the waiting time formula is:
[0063] ;
[0064] Wherein, is the number of retries, and the maximum number of retries is set to 5, is the initial waiting time.
[0065] The core of step S1 is to build a multi-round closed-loop iteration framework of "evaluation, generation and update", continuously optimize the news generation quality through the dynamic interaction of the generated agent and the detection agent, and simultaneously consider the running speed and robustness design, and the generated news after multiple rounds of iteration optimization meets the requirements of factual accuracy, logical coherence and compliance with news genre specifications.
[0066] In one embodiment, the step S2 of constructing a speech synthesis network based on ChatTTS is converted to into synthesized speech , and realizes prosody control, pause optimization and noise reduction processing during synthesis, specifically including:
[0067] Step S21: Let the news text sequence be:
[0068] ;
[0069] wherein, represents the th word in the text, is the length of the text.
[0070] The news text is subjected to character normalization (such as unified punctuation format), numerical time regularization (such as standardization processing of "2025 April 10"), punctuation correction and language recognition to determine the language category , representing Chinese and English respectively.
[0071] Step S22: Generate phoneme sequence and prosody boundary sequence through word segmentation, part-of-speech tagging and syntax analysis:
[0072] ;
[0073] ;
[0074] wherein, represents the th phoneme, represents the prosody boundary corresponding to the phoneme, is the length of the phoneme sequence.
[0075] Step S23: Use the text encoder to map , to the hidden table:
[0076] ;
[0077] wherein, is the semantic representation of the text in high-dimensional space, thereby providing context information for subsequent prosody modeling.
[0078] Step S24: According to , predict phoneme duration, fundamental frequency, energy and pause information, and establish a prosody prediction model :
[0079] ;
[0080] wherein, predicted duration of each phoneme; predicted fundamental frequency curve; predicted energy value; predicted pause information.
[0081] To ensure the prediction accuracy, the prosody loss is constructed as:
[0082] ;
[0083] wherein, represents the loss weight; represents the mean square error; represents the cross-entropy loss; is the true label.
[0084] The prosody loss is used to ensure that the generated speech has natural rhythm, reasonable pauses, and is closer to the real human speech pattern.
[0085] Step S25: introducing language embedding and style embedding to form the input vector of the acoustic model :
[0086] ;
[0087] wherein, represents the vector concatenation operation; sufficiently fuses the language features and style control information.
[0088] Using the acoustic model , the input vector is mapped to the target mel spectrum :
[0089] ;
[0090] Then, using the neural vocoder , the preliminary speech waveform is generated:
[0091] ;
[0092] The acoustic loss is defined as:
[0093] ;
[0094] wherein, is the true mel spectrum, is the true speech waveform, represents the short-time Fourier transform, is the spectrum constraint weight, used to control the contribution proportion of each loss term.
[0095] The acoustic loss is optimized by a joint optimization model of the spectral reconstruction error and the time-frequency domain consistency, so as to ensure that the generated speech is highly consistent with the real speech in timbre, intelligibility and details.
[0096] Step S26: using the speech enhancement model to perform noise reduction on the synthesized speech ;
[0097] ;
[0098] wherein, is the speech enhancement model.
[0099] The noise reduction loss is defined as:
[0100] ;
[0101] wherein, is the clean reference speech, is a scale-invariant signal-to-noise ratio evaluation function, is a weight coefficient.
[0102] Step S27: constructing a total loss function :
[0103] ;
[0104] wherein, the weight coefficient is used to ensure that the output synthesized speech is balanced between content accuracy, prosody naturalness, pause rationality and voice clarity.
[0105] The present application trains the speech synthesis network by using the total loss function, so as to ensure that the generated synthesized speech is consistent and stable in voice quality, prosody and style.
[0106] In one embodiment, the above step S3: based on the user input news broadcast video and , multi-modal data alignment is performed, the lip action sequence is generated, and the face features, background images of the news broadcast video are fused and enhanced to obtain the final video , specifically comprising:
[0107] Step S31: extracting the face features , background images , audio features in the user input news broadcast video; and aligning the multi-modal information according to the time stamp:
[0108] ;
[0109] wherein, the speech recognition function will The time sequence converted into phonemes; For converting news text sequence Into speech; For phoneme alignment function, for aligning synthesized speech And ; For the aligned synthesized speech information, containing phoneme level timestamp information, so as to eliminate the error of neural network due to pause, adjusting phoneme duration;
[0110] The speech synchronization prediction lip parameter sequence function; Then, according to the audio features The predicted and Speech synchronized lip parameter sequence.
[0111] Step S32: fuse With facial features , and then repair and enhance through Model:
[0112] ;
[0113] Wherein, The fusion function is The model is used for repairing, deblurring and improving the resolution of the fused face.
[0114] The present application intelligently fuses the facial features Of the user input news broadcast video (including the expression and posture style of the model) with the predicted lip movement sequence To generate facial animation data that not only retains the model style but also synchronizes the speech. Then use the GFPGAN image quality enhancement function to repair, deblur and improve the resolution of the fused face , realize "picture enhancement".
[0115] Step S33: synthesize And , and render into the final video frame:
[0116] ;
[0117] Wherein, The final video Frame , the synthesis and rendering function is .
[0118] Step S3 designs a voice synthesis network based on ChatTTS. The network can receive the audio stream in real time, automatically extract the core information for driving the digital human generation, including three elements of facial features, audio features, and background content. Among them, in the visual aspect, the key facial feature points of the digital human are extracted from the user-uploaded recording video (or the built-in default video), such as lip deformation, eye movement, and facial contour points, to provide structural basis for subsequent lip shape and expression generation; in the audio aspect, phoneme features are extracted from the audio stream, including the pronunciation patterns of vowels and consonants, and rhythm information such as speech rate and pause are analyzed to support phoneme-level synchronization; in the scene aspect, the system reads the background image or video stream selected by the user to construct the environment for the final presentation of the digital human, thereby realizing the driving control of multi-modal fusion.
[0119] In one embodiment, the above step S4: replacing the background of the digital human in the video stream with the target background input by the user includes: After dynamic segmentation by the robust video matting RVM model, the target background input by the user is synthesized, and post-processing such as subtitle creation alignment, bottom strip and logo addition is performed to generate a news broadcast piece that meets the broadcast standards, specifically including:
[0120] Step S41: inputting the video stream of the digital human and the target background input by the user into the RVM model to obtain a preliminary green screen video. Input the robust video matting RVM pre-training model, and replace the target background input by the user Adjust to the resolution of and complete format adaptation.
[0121] Step S42: the RVM pre-training model aggregates the inter-frame temporal information through the ConvGRU structure in the recurrent decoder, and outputs a temporary green screen video including the digital human foreground and green background. The state update formula is:
[0122]
[0123] Among them, is the multi-scale feature of the current frame extracted by the encoder, is the temporal information of the previous frame, and the update gate is the gate variable calculated by the sigmoid function, and the reset gate controls information retention and forgetting through the sigmoid function, is the candidate state, which is generated by the hyperbolic tangent function, represents convolution operation, represents element-wise multiplication; , , respectively represent the pre-training bias terms of the update gate, the reset gate, and the candidate hidden state calculation link, which are used to adjust the output , , The baseline value, together with the other two, improves the gating mechanism's accuracy in matting fine structures such as the edges and hair strands of digital humans; , , These represent the current frame input features of the update gate, reset gate, and candidate hidden state calculation stages, respectively. Pre-trained convolutional kernel parameters, along with the three components, capture the spatial local features of video frames through convolution operations. This is crucial for ConvGRU to achieve inter-frame temporal information aggregation and ensure cross-frame consistency of RVM keying. , , These represent the hidden states of the previous frame in the update gate, reset gate, and candidate hidden state calculation stages, respectively. Pre-trained convolutional kernel parameters incorporate inter-frame temporal correlations into gating decisions through convolutional transformations of historical hidden states, ensuring cross-frame consistency of video keying and accurate segmentation of fine structures.
[0124] Step S42 can resolve the ambiguity issues of hair strands and semi-transparent edges in single-frame keying.
[0125] Step S43: The temporary green screen video is optimized for edge details using the DGF module; then, the green screen is detected in the HSV color space, green samples are extracted, a dynamic threshold is calculated to generate a mask, and the mask is optimized using morphological closing operations and Gaussian blur to obtain a foreground containing only the digital human. Then, blend according to the following pixel-level formula. With target background :
[0126] ;
[0127] in, This is the result of normalizing the mask after green screen inspection and morphological optimization, showing the value of each pixel. Strictly between 0 and 1, in the green screen background area, , making the participation background Participating in integration in the digital human prospect area, Enhancing the prospects of digital humans Almost no blending background .
[0128] Finally extract The audio, and with The video is then re-composited to obtain the base video with the background replaced.
[0129] Step S44: Read the news text sequence , the first sentence is automatically split into title and the remaining content into subtitle materials; the background replaced base video is loaded, the news icon and specified font are synchronously loaded; then the bottom is made, the icon is scaled and superimposed on the white block to the lower left corner, and a gradient transparent blue title bar is made on the right side for displaying the title; then the display time of each sentence is allocated according to the total time of the video and the number of characters of the subtitle, the subtitle is generated and aligned according to the time axis; finally, the base video, the icon white block, the title bar and the subtitle element are superimposed, and the news report is output as a film by using H.264 and AAC encoding.
[0130] Figure 2 The schematic diagram of the whole process of the method is shown.
[0131] The multi-agent based news generation and digital person reporting method of the application ensures the quality of news text through agent interaction, realizes natural reporting effect with the help of advanced speech synthesis and lip shape generation technology, and finally outputs standardized films through video synthesis packaging, with high automation degree and high quality generation effect, and provides an efficient technical solution for the field of news reporting.
[0132] Embodiment two:
[0133] As shown in Figure 3 , the embodiment of the application provides a multi-agent based news generation and digital person reporting system, which comprises the following modules:
[0134] The agent interaction news generation module 51 is used for generating a news text sequence according to a user input news abstract through multiple iterations of a generated agent and a detection agent, in combination with a preset rule ;
[0135] The text to speech module 52 is used for constructing a speech synthesis network based on ChatTTS, converting into synthesized speech , and realizing rhythm control, pause optimization and noise reduction processing in the synthesis process;
[0136] The digital person lip shape generation module 53 is used for performing multi-modal data alignment based on a user input news reporting video and , generating a lip action sequence, and fusing and enhancing the face features and background images of the news reporting video to obtain a final video ;
[0137] The video synthesis and packaging module 54 is used for dynamically segmenting after robust video matting model, synthesizing with a target background input by the user, and then performing post-processing such as subtitle creation alignment, bottom strip and icon addition to generate a news reporting film that meets the broadcast standard.
[0138] A multi-agent based news generation and digital human broadcasting device, comprising one or more electronic devices, wherein the one or more electronic devices are configured to implement a multi-agent based news generation and digital human broadcasting method.
[0139] An electronic device, comprising: one or more processors; memory storing one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement a multi-agent based news generation and digital human broadcasting method.
[0140] A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to implement a multi-agent based news generation and digital human broadcasting method.
[0141] A non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements a multi-agent based news generation and digital human broadcasting method.
[0142] The above description is merely that of the specific embodiments of the application and as such is not to be taken in a limiting sense but is made merely for the purpose of providing some exemplary embodiments of the application. Various modifications to these embodiments can and will be apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without the use of the inventive faculty. Thus, all such modifications are intended to be included within the scope of the application. The application is not to be restricted, except within the scope of the claims and their equivalents.
Claims
1. A method for news generation and digital human broadcasting based on multi-agent intelligence, characterized in that, include: Step S1: Based on the news summary input by the user, and through multiple rounds of iteration between the generating agent and the detecting agent, a news text sequence is generated in accordance with preset rules. ; Step S2: Construct a speech synthesis network based on ChatTTS, and... Convert to synthesized speech Furthermore, prosody control, pause optimization, and noise reduction are implemented during the synthesis process. Step S3: Based on user input, news broadcast videos and Multimodal data alignment is performed to generate a lip movement sequence, which is then fused and enhanced with the facial features and background image of the news broadcast video to obtain the final video. ; Step S4: After dynamic segmentation using a robust video keying model, the image is composited with the target background input by the user. Then, through post-processing such as subtitle creation and alignment, and the addition of bottom stripes and markers, a news broadcast that meets broadcast standards is generated.
2. The news generation and digital human broadcasting method based on multi-agent technology according to claim 1, characterized in that, Step S1: Based on the news summary input by the user, and through multiple iterations of the generating agent and the detecting agent, a news text sequence is generated in combination with preset rules. Specifically, it includes: Step S11: Construct a news dataset : ; in, It is the sole identifier of news; It is a news summary; It's news written by humans; Based on News generated by intelligent agents; This item is used to store the news generated by the agent in the previous round; it is initially empty. This is the news generated by the current round of intelligent agents; Used to store evaluations of the news generated by the agent in the previous round; Save the 5 news items with the highest cosine similarity to the current news. ; Step S12: The detection agent is represented as follows: ; ; ; in, for for News No. Round evaluation results, and store them to for News Fields, It is a function for detecting how intelligent agents evaluate the quality of news. Based on for News The index is concatenated with the 5 most similar elements. for A function for detecting intelligent prompts in news articles, where + indicates a concatenation operation; It is the evaluation criteria for evaluating the task, including the title, introduction, body, and conclusion; To be evaluated for News objectives, including , , ; For splicing for Similar news , , The evaluation matrix comparison results between the current round and the previous round, and a single-sample information block providing optimization direction suggestions based on the comparison results; For calculation for News Other for News Use the cosine similarity function and construct the top 5 news items with the highest similarity. Functions of sets; for for The semantic vector of news Represents semantic vectors and semantic vectors Inner product, It is the L2 norm. It extracts the 5 news items with the highest similarity. The function; Step S13: Generate the agent representation as follows: ; ; in, for for News No. The round generates the results and stores them. for News Fields, It is the function that generates news by the intelligent agent. Based on for News A function for generating agent prompts by concatenating indexes; It generates a task instruction header, which includes learning objectives and example descriptions; For splicing for Similar news , , Fields, comparison results of the evaluation matrix between the current round and the previous round, and a single-sample information block with optimization direction suggestions based on the comparison results; It is a target generation instruction, containing for News and format requirements; Step S14: After each iteration, the news generated by the agent and human news are compared and evaluated using an evaluation index matrix, and this matrix is applied to the operation of the agent. This matrix can be represented as: ; ; ; ; ; ; in, Used to measure the n-gram overlap between candidate text and reference text; It is the accuracy of n-grams. It is an n-gram weight. It is a length penalty factor; It is the largest order of an n-gram; Used to combine stemming matching, synonym matching, and chunk matching; It is the harmonic mean of precision and recall. It is the stem matching rate. It's the coverage of the chunk. and To adjust the parameters; The longest common subsequence (LCS) used to focus on candidate and reference texts; Recall is the ratio of the length of the longest common subsequence (LCS) between the candidate text and the reference text to the length of the reference text. Precision is the ratio of LCS length to candidate text length. It is the harmonic mean of recall and precision. It is based on word embedding similarity from a pre-trained language model; For candidate text token sequences, For reference text token sequence, Used to calculate cosine similarity Embed the token output by the model; The first in the candidate text token sequence One token; For the reference text token sequence, the first One token; Normalize the LLM output score to interval; It is the average of multiple ratings. This represents the maximum value of the rating scale.
3. The news generation and digital human broadcasting method based on multi-agent technology according to claim 2, characterized in that, Step S2: Construct a speech synthesis network based on ChatTTS, and Convert to synthesized speech Furthermore, prosody control, pause optimization, and noise reduction are implemented during the synthesis process, specifically including: Step S21: Let the news text sequence be: ; in, Indicates the first in the text One word, The length of the text; The news text undergoes character normalization, digital time regularization, punctuation correction, and language recognition to determine the language category. , representing Chinese and English respectively; Step S22: Generate phoneme sequences and prosodic boundary sequences through word segmentation, part-of-speech tagging, and syntactic analysis. ; ; in, Indicates the first One phoneme, This indicates the prosodic boundary corresponding to the phoneme. The length of the phoneme sequence; Step S23: Use a text encoder Will , Mapped to a hidden table: ; in, The semantic representation of text in high-dimensional space; Step S24: According to Predict phoneme duration, fundamental frequency, energy, and pause information to establish a prosodic prediction model. : ; in, This indicates the predicted duration of each phoneme; This represents the predicted fundamental frequency curve; Indicates the predicted energy value; Indicates the predicted pause information; To ensure accurate predictions, the prosodic loss is constructed as follows: ; in, Indicates the loss weight; Indicates mean square error; Represents cross-entropy loss; This is a true label; Step S25: Introduce language embeddings With style embedding Input vectors that form the acoustic model : ; in, This represents a vector concatenation operation; Using acoustic models input vector Mapped to target Mel spectrum : ; Then, using a neural vocoder Generate preliminary speech waveform : ; The acoustic loss is defined as: ; in, It is the true Mel spectrum. It is a real speech waveform. Represents the short-time Fourier transform. For spectral constraint weights; Step S26: Use a speech enhancement model to... Noise reduction: ; in, For speech enhancement models; The noise reduction loss is defined as: ; in, For clean reference speech, This is the scale-invariant signal-to-noise ratio evaluation function. These are the weighting coefficients; Step S27: Construct the total loss function : ; Among them, the weighting coefficient .
4. The news generation and digital human broadcasting method based on multi-agent technology according to claim 3, characterized in that, Step S3: Based on the news broadcast video input by the user and Multimodal data alignment is performed to generate a lip movement sequence, which is then fused and enhanced with the facial features and background image of the news broadcast video to obtain the final video. Specifically, it includes: Step S31: Extract facial features from the news broadcast video input by the user. Background image Audio features Align multimodal information according to timestamps: ; Among them, speech recognition function Will Time series converted into phonemes; Used to sequence news text Perform text-to-speech conversion; This is a phoneme alignment function used to align synthesized speech. and ; The aligned synthesized speech information includes phoneme-level timestamp information; Speech-synchronous prediction of lip shape parameter sequence function; This is based on audio features Predicted sum A sequence of lip-shape parameters for speech synchronization; Step S32: With facial features To merge, and then through The model is repaired and enhanced: ; in, For fusion function, The model is used to repair, deblur, and improve the resolution of the fused face. Step S33: ... and The components are then composited and rendered into the final video frames. ; in, This is the final video. The frame, These are the functions for compositing and rendering.
5. The news generation and digital human broadcasting method based on multi-agent technology according to claim 4, characterized in that, Step S4: ... After dynamic segmentation using a robust video keying model, the image is composited with the target background input by the user. Post-processing, including subtitle creation and alignment, and the addition of bottom stripes and markers, generates a news broadcast that meets broadcast standards. Specifically, this includes: Step S41: ... Input the robust video matting RVM pre-trained model and input the target background to be replaced from the user input. Adjust to The resolution was adjusted and format adaptation was completed. Step S42: The RVM pre-trained model aggregates inter-frame temporal information through the ConvGRU structure in the recurrent decoder, outputting a temporary green screen video including a digital human foreground and a green background. Its state update formula is: ; in, The multi-scale features extracted by the encoder for the current frame. Update the gate based on the temporary information from the previous frame. The gate variable is calculated using the sigmoid function, and the gate is reset. The sigmoid function controls information retention and forgetting. Candidate states are generated using the hyperbolic tangent function. This represents the convolution operation. Represents element-wise product; , , These represent the pre-training bias terms for the update gate, reset gate, and candidate hidden state calculation stages, respectively. , , These represent the current frame input features of the update gate, reset gate, and candidate hidden state calculation stages, respectively. Pre-trained convolutional kernel parameters; , , These represent the hidden states of the previous frame in the update gate, reset gate, and candidate hidden state calculation stages, respectively. Pre-trained convolutional kernel parameters; Step S43: The temporary green screen video is optimized for edge details using the DGF module; then, the green screen is detected in the HSV color space, green samples are extracted, a dynamic threshold is calculated to generate a mask, and then the mask is optimized by morphological closing operation and Gaussian blur to obtain a foreground containing only the digital human. Then, blend according to the following pixel-level formula. With the target background : ; in, The result is the normalized version of the mask after green screen inspection and morphological optimization. Finally extract The audio, and with The composite video is then re-generated to obtain the base video with the background replaced. Step S44: Read the news text sequence The process automatically splits the first line into a title and the remaining content into sentences for subtitles; it loads the base video with a replaced background, and simultaneously loads the news icon and specified font; then it creates the bottom, scaling the icon and overlaying it with a white block in the lower left corner, and adds a gradient transparent blue title bar on the right to display the title; it then allocates the display duration of each sentence according to the total video length and the number of subtitle characters, generates subtitles and aligns them according to the timeline; finally, it overlays the base video, icon white block, title bar and subtitle elements, and outputs the news broadcast using H.264 and AAC encoding.
6. A news generation and digital human broadcasting system based on multi-agent intelligence, characterized in that, Includes the following modules: The intelligent agent interactive news generation module is used to generate a news text sequence based on the news summary input by the user, through multiple rounds of iteration between the generating and detecting intelligent agents, and in accordance with preset rules. ; The text-to-speech module is used to build a speech synthesis network based on ChatTTS, and to... Convert to synthesized speech Furthermore, prosody control, pause optimization, and noise reduction are implemented during the synthesis process. A digital human lip-shape generation module for news broadcast videos based on user input and... Multimodal data alignment is performed to generate a lip movement sequence, which is then fused and enhanced with the facial features and background image of the news broadcast video to obtain the final video. ; The video compositing and packaging module is used to... After dynamic segmentation using a robust video keying model, the image is composited with the target background input by the user. Then, through post-processing such as subtitle creation and alignment, and the addition of bottom stripes and markers, a news broadcast that meets broadcast standards is generated.
7. A news generation and digital human broadcasting device based on multi-agent intelligence, characterized in that, It includes one or more electronic devices, wherein the one or more electronic devices are used to implement the method of any one of claims 1 to 5.
8. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method of any one of claims 1 to 5.
9. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, cause the processor to perform the method described in any one of claims 1 to 5.
10. A non-transitory computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Video synthesis method and device, equipment and storage medium
CN112866586A
Rhythm migration speech synthesis method and system
CN115910026A
Internet news content automatic generation method based on big data
CN116483990A
Digital human mouth broadcast video generation method, system and device and medium
CN120475233A