Method and system for automatically generating news video based on guided reconstruction to maintain original style

By using a multidimensional emotion vector and valence constraint-based guidance reconstruction method, the problems of coarse public opinion guidance control and inconsistent style in news video generation are solved. This method achieves fine and controllable public opinion guidance and factual authenticity, and the generated video is consistent with the original style, making it suitable for automatic news video generation.

CN121151653BActive Publication Date: 2026-04-07CHENGDOU HUAQIYUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing news video generation technology is crude in controlling public opinion, making it difficult to achieve continuous and proportional fine-tuning. Furthermore, the generated videos have inconsistent styles and lack end-to-end optimization, resulting in difficulty in ensuring the authenticity of news facts and low credibility of dissemination.

Method used

By employing a multidimensional emotion vector and valence constraint-based guidance reconstruction method, and through multimodal element analysis, dynamic emotion scoring, rhetorical substitution, and visual emotion mapping, a new video with a style consistent with the original news video is generated, ensuring the accurate control of factual authenticity and public opinion guidance.

Benefits of technology

It achieves precise and controllable guidance of public opinion, and the generated videos are visually and aurally indistinguishable from their source, ensuring the authenticity of news facts. Furthermore, the parameters of each module are self-consistent and unified, achieving end-to-end optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121151653B_ABST
    Figure CN121151653B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of video generation, in particular to a news video automatic generation method and system based on guidance reconstruction to maintain original style. The present application realizes automatic, fine and concealed controllable reconstruction of public opinion guidance. First, the original news video is inputted and the continuously adjustable target guidance parameter is set. Then, the video is analyzed in multiple modes to separate the unchangeable fact triplets and the adjustable rhetoric word set. The large language model is used to dynamically calculate the segment-level emotion vector, and the target guidance parameter is distributed to each paragraph through the constraint optimization algorithm to generate new scripts conforming to the target emotion. At the same time, the prosody parameters of the original audio are copied to drive the text-to-speech, and the visual parameters such as the subtitle backplane are adjusted based on the emotion mapping function. Finally, by replacing the anchor, synthesizing new voice and subtitles, a news video product is outputted, which is highly consistent with the original video in fact, picture style and reporting rhythm, but the public opinion guidance has been accurately reconstructed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video generation, in particular to a news video automatic generation method and system based on guided reconstruction to maintain original style. BACKGROUND

[0002] As the core carrier of information dissemination, news videos play a crucial role in public opinion guidance and public cognition shaping. With the development of artificial intelligence technology, the automatic production and editing technology of news content has made significant progress.

[0003] Currently, the existing technologies related to the automatic production and editing technology of news content mainly cover the following two directions:

[0004] 1. Automatic generation and sentiment rewriting of news text: This type of technology usually uses large language models to automatically rewrite or generate news articles. Existing methods can control the sentiment tendency of the text to some extent, for example, adjusting the overall emotional tone of the text to one of the three categories: "positive", "negative", or "neutral".

[0005] However, this sentiment control is discrete and rough, and cannot achieve continuous, proportional fine-tuning (for example, generating a text with "negative 70%" or "positive 40% with fear color"). More importantly, existing text rewriting methods often struggle to strictly guarantee that the original news facts are not changed when adjusting the sentiment tendency, and are prone to introduce false or distorted information in the rewriting process, posing a risk of misleading the audience.

[0006] 2. Digital anchor and text-to-speech technology: This type of technology is based on deep learning and computer vision, and can drive virtual human images and combine text-to-speech systems to generate news broadcast videos based on input text.

[0007] However, existing virtual anchor systems mainly focus on the realism of the image and the synchronization of the lip movements, and the generated broadcast videos have a general style that is difficult to match the unique style of a specific original news video.

[0008] Specifically, the speech rhythm (such as speech rate, pauses, and pitch fluctuations) cannot accurately reproduce the original anchor's broadcast rhythm, and the visual style of the subtitles and backdrops (such as background color, font, size, position, transparency, etc.) cannot be extracted from the original video and restored equivalently. This leads to the generated video being easily recognized by the audience as a "copy" or "synthetic product", which weakens its credibility and concealment in dissemination.

[0009] Furthermore, existing technical solutions typically target only a single stage in the news production process (such as text, audio, or images), lacking an end-to-end, end-to-end optimization system. When multiple versions with different public opinion orientations need to be generated quickly from the same news material, the discrete modules cannot share unified parameter constraints, resulting in a lack of consistency and coordination between the text, audio, visuals, and subtitles in the final video, leading to poor overall quality.

[0010] In summary, existing technologies suffer from shortcomings such as crude control of public opinion guidance, difficulty in guaranteeing factual authenticity, inconsistent video styles, and lack of end-to-end optimization. These shortcomings make it difficult to meet the urgent need for rapid, accurate, covert, and factually accurate news videos in the field of public opinion dissemination. Summary of the Invention

[0011] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for automatically generating news videos based on guidance reconstruction while maintaining the original style. It aims to provide a one-stop generation solution that ensures precise control of public opinion guidance, guarantees factual authenticity, maintains consistent video style, and has end-to-end optimization, thereby meeting the urgent need for fast, accurate, covert, and factually accurate news videos in the fields of public opinion dissemination.

[0012] To achieve the above objectives, this application proposes a method for automatically generating news videos that maintains the original style based on guided reconstruction, comprising the following steps:

[0013] Step S1: Input the original news video and set the guidance intensity parameter; wherein, the original news video includes the original anchor screen, the original audio, and the original subtitle background, and the guidance intensity parameter includes a multi-dimensional emotion vector and the corresponding valence constraint;

[0014] Step S2: Perform multimodal element analysis on the original audio of the original news video to extract speech text, sentence structure, fact triples and rhetorical word set;

[0015] Step S3: Dynamically calculate the sentiment vector of each rhetorical word in the rhetorical word set based on the large language model, and aggregate them into a segment-level sentiment vector;

[0016] Step S4: Based on the guidance intensity parameter, assign a paragraph-level target sentiment vector to each paragraph using a constraint optimization algorithm, and perform rhetorical replacement based on the sentiment difference between the paragraph-level sentiment vector and the paragraph-level target sentiment vector to generate new copy.

[0017] Step S5: Extract the prosodic parameters of the original audio, and generate new prosodic parameters based on the prosodic parameters adjusted according to the emotional difference; use text-to-speech technology to generate new speech based on the new text and the new prosodic parameters; generate a new broadcaster's image based on the broadcaster's lip movements in the speech-driven library and the new speech.

[0018] Step S6: Parse and extract the subtitle style parameters of the original subtitle background, and adjust the subtitle style parameters according to the emotion difference through the visual-emotion mapping mechanism to generate a new subtitle background;

[0019] Step S7: Generate a news video with a redesigned direction and consistent style by replacing the new anchor screen, overlaying the new voice and new subtitle background.

[0020] As a further solution, in step S1, the guidance intensity parameter is defined by a multi-dimensional emotion vector, and valence constraints are established by combining a valence mapping vector and a weight matrix; wherein,

[0021] The valence constraint is used for a continuous mapping from unidimensional guidance intensity to a multidimensional emotion distribution, expressed as:

[0022]

[0023] in, Represents a valence mapping vector. Represents the weight matrix. Represents a multidimensional emotion vector. This represents the guiding strength parameter, and the one-dimensional guiding polarity is set to... -1 = negative 100%, 0 = neutral, +1 = positive 100%.

[0024] As a further solution, in step S2, multimodal element analysis is performed through the following steps:

[0025] The original news video's audio was processed using a speech recognition model to convert the audio into text, resulting in audio-text.

[0026] The paragraph set is divided according to semantic boundaries using natural language processing algorithms to obtain the sentence structure; wherein, the sentence structure includes sentence division and time boundary;

[0027] F fact triples F are extracted through dependency parsing and used as an immutable fact semantic base;

[0028] The set of rhetorical words for each paragraph is identified using attention mechanisms and an emotion dictionary.

[0029] As a further solution, in step S3, a large language model is used to dynamically calculate the multidimensional sentiment intensity of rhetorical words in context, generating word-level sentiment vectors, and then the word-level sentiment vectors are aggregated into segment-level sentiment vectors; wherein, the segment-level sentiment vector is represented as:

[0030]

[0031] in, Represents segment-level sentiment vectors. Represents word-level sentiment vectors. Paragraph A collection of rhetorical devices.

[0032] As a further solution, in step S4, the segment-level target sentiment vector is generated using the following formula:

[0033] ,

[0034] in, Indicates the extreme emotion of the target segment. Indicates segment weight, Indicates paragraph number, This represents a multidimensional emotion vector.

[0035] As a further solution, in step S4, new copy is generated under constraints by using fact-based locking and rhetorical rewriting, combined with LLM; wherein, sentiment distribution control is embedded in the prompt of the LLM.

[0036] As a further solution, in step S5, the prosodic parameters include speech rate, average pitch, pauses, and fluctuations.

[0037] As a further solution, in step S6, a new subtitle background is generated through the following steps:

[0038] The system detects and identifies subtitle regions in each frame, and merges the same subtitle regions in frames that are time-continuous to form a set of subtitle segments.

[0039] Calculate and extract subtitle style parameters for each subtitle region;

[0040] Find the template with the most similarity to the subtitle area from the template library as the new subtitle background: if the minimum distance between the template and the subtitle area in the template library is greater than a threshold, then call the AI ​​generation module to generate a new template based on the subtitle area;

[0041] By associating segment-level emotion vectors with subtitle style parameters and achieving cross-modal emotion expression consistency through a visual-emotion mapping mechanism, new subtitle style parameters are obtained.

[0042] As a further solution, in step S7, based on the original video scene structure, the system intelligently distinguishes between the broadcast segments and non-broadcast segments, and performs video replacement and subtitle overlay respectively; among which,

[0043] If a frame of the anchor is detected, the most similar background template is matched from the background template library based on the studio characteristics of the original news video as the new background, and then merged with the anchor frame to replace the entire image.

[0044] If no anchor's image frame is detected, the original main content of the video remains unchanged, and only the generated audio track and the corresponding new subtitle background and new text are superimposed.

[0045] Among them, the background style and text format of the new subtitle backdrop are automatically adjusted according to the new subtitle style parameters to ensure that the visuals and text are consistent in terms of emotional guidance;

[0046] The video footage is spliced ​​together in the order of the audio, and transition effects are added at the splicing points to make the transitions more natural.

[0047] On the other hand, the present invention also provides an automatic news video generation system based on guided reconstruction to maintain the original style, employing a method for automatic news video generation based on guided reconstruction to maintain the original style as described in any of the preceding claims, including:

[0048] Input and guidance setting module: Input the original news video and set the guidance intensity parameter; wherein, the original news video includes the original anchor screen, original audio, and original subtitle background, and the guidance intensity parameter includes a multi-dimensional emotion vector and the corresponding valence constraint;

[0049] News Multimodal Element Analysis Module: Performs multimodal element analysis on the original audio of the original news video to extract speech text, sentence structure, fact triples, and rhetorical word set;

[0050] Dynamic emotion scoring module: Based on a large language model, dynamically calculate the emotion vector of each rhetorical word in the rhetorical word set and aggregate them into segment-level emotion vectors;

[0051] Guidance and copywriting reconstruction module: Based on the guidance strength parameter, a paragraph-level target sentiment vector is assigned to each paragraph through a constraint optimization algorithm, and rhetorical replacement is performed based on the sentiment difference between the paragraph-level sentiment vector and the paragraph-level target sentiment vector to generate new copywriting;

[0052] Voice and Anchor Generation Module: Extracts prosodic parameters from the original audio and generates new prosodic parameters based on the prosodic parameters adjusted according to the emotional difference; generates new voice based on the new text and new prosodic parameters using text-to-speech technology; and generates new anchor visuals based on the anchor lip movements in the voice-driven library and the new voice.

[0053] Visual layer parsing and subtitle generation module: Parses and extracts the subtitle style parameters of the original subtitle background, and adjusts the subtitle style parameters according to the emotion difference through the visual-emotion mapping mechanism to generate a new subtitle background;

[0054] Video reconstruction and compositing module: By replacing the anchor's image with a new one, overlaying the new voice and new subtitle background, a news video with a consistent style and a reconstructed direction is generated.

[0055] Compared with related technologies, the automatic generation method and system for news videos based on guided reconstruction while maintaining the original style provided by this invention has the following advantages:

[0056] 1. By introducing continuous polarity parameters and multidimensional emotion vectors, this invention can achieve complex guidance control such as "70% negative" and "40% positive with fear", which far exceeds the existing conventional three-classification method and achieves precise and controllable guidance of public opinion.

[0057] 2. This invention adopts a method of decoupling facts and rhetoric, only adjusting the rhetoric and narrative order to ensure that the reported facts remain unchanged. As a result, the generated video not only conforms to the goal of public opinion guidance, but also does not constitute false reporting, thus ensuring the authenticity of the news facts.

[0058] 3. This invention replicates prosody parameters and restores subtitle background parameters, making the generated video consistent with the original video in terms of the anchor's speaking rhythm, timbre rhythm, and subtitle style, making it difficult to distinguish the source from the visual and auditory perspective;

[0059] 4. The parameters of each module in this invention are mapped and constrained through derivation formulas to form a complete algorithm chain, ensuring that the text, voice, video and subtitles remain consistent and unified after the guidance is adjusted, thus achieving end-to-end link optimization. Attached Figure Description

[0060] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0061] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0062] Figure 1 A schematic diagram illustrating the steps of an automatic news video generation method based on guided reconstruction to maintain the original style, provided by this invention;

[0063] Figure 2 This invention provides a schematic diagram of the structure of an automatic news video generation system based on guided reconstruction to maintain the original style.

[0064] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0066] Example 1

[0067] Please see Figure 1 This embodiment provides a method for automatically generating news videos that maintain the original style based on guided reconstruction, including the following steps:

[0068] Step S1: Input the original news video and set the guidance intensity parameter; wherein, the original news video includes the original anchor screen, the original audio, and the original subtitle background, and the guidance intensity parameter includes a multi-dimensional emotion vector and the corresponding valence constraint;

[0069] Step S2: Perform multimodal element analysis on the original audio of the original news video to extract speech text, sentence structure, fact triples and rhetorical word set;

[0070] Step S3: Dynamically calculate the sentiment vector of each rhetorical word in the rhetorical word set based on the large language model, and aggregate them into a segment-level sentiment vector;

[0071] Step S4: Based on the guidance intensity parameter, assign a paragraph-level target sentiment vector to each paragraph using a constraint optimization algorithm, and perform rhetorical replacement based on the sentiment difference between the paragraph-level sentiment vector and the paragraph-level target sentiment vector to generate new copy.

[0072] Step S5: Extract the prosodic parameters of the original audio, and generate new prosodic parameters based on the prosodic parameters adjusted according to the emotional difference; use text-to-speech technology to generate new speech based on the new text and the new prosodic parameters; generate a new broadcaster's image based on the broadcaster's lip movements in the speech-driven library and the new speech.

[0073] Step S6: Parse and extract the subtitle style parameters of the original subtitle background, and adjust the subtitle style parameters according to the emotion difference through the visual-emotion mapping mechanism to generate a new subtitle background;

[0074] Step S7: Generate a news video with a redesigned direction and consistent style by replacing the new anchor screen, overlaying the new voice and new subtitle background.

[0075] It should be noted that with the rapid development of digital information technology, news is an important medium for the public to obtain information. The same facts require us to select and set the news orientation according to actual needs in order to achieve a good public opinion effect. The core technical challenge is: how to accurately and controllably change the perspective of news and public opinion (for example, automatically converting 70% positive news to 70% negative or neutral news) without changing the facts or the video presentation, and quickly generate new video clips that are difficult for users to distinguish but have a significant guiding effect.

[0076] Therefore, an automated, parameterized, and imperceptible method for reconstructing the narrative direction is needed. However, news videos contain both factual and rhetorical layers, and redirection is not simply a matter of changing a few words. It requires ensuring that the overall guiding strength and paragraph polarity distribution conform to the set parameters, while simultaneously restoring the video's subtitles, backdrop, and narration rhythm as equivalently as possible. Existing methods cannot precisely adjust the rhetoric and narrative logic while keeping the facts unchanged, specifically manifested in the following ways:

[0077] 1. Existing manual writing or simple emotional substitution may alter the facts; in addition, news requires timeliness, and the method of rewriting and recording cannot meet the needs of pushing out news within a few hours.

[0078] 2. Existing video reproduction methods: re-recording the anchor or directly using AI digital humans, but the styles are obviously different and easily recognized; in addition, the three-category orientation (positive / negative / neutral) cannot be scaled up.

[0079] Therefore, the news video generation method based on public opinion guidance reconstruction proposed in this embodiment constructs an end-to-end guidance reconstruction algorithm chain, which includes: input → news element analysis → guidance and text reconstruction → voice and anchor generation → subtitle and background restoration → video synthesis output.

[0080] Segment-level polarity modeling: Divide the original news video into segments and calculate the polarity score for each segment;

[0081] Goal-oriented parameterization: Input goal-oriented parameters are allocated to each segment through constraint optimization to obtain the goal polarity;

[0082] Rhetoric and sequence restructuring: Maintain the news facts, for example, replace "positive cooperation" with "over-reliance", and in terms of sequence, mention the risks first and then the benefits;

[0083] Copy generation and duration alignment: Through multi-objective optimization, ensure that the copy's direction changes but the sentence length and rhythm remain consistent with the original video;

[0084] Prosody transfer TTS: Generates new voice with a rhythm highly consistent with the original broadcaster, but the content is an adjusted script;

[0085] Anchor replacement and lip-syncing: Replace segments in the original video that contain the anchor's image, but maintain the original composition and camera language;

[0086] Subtitle background restoration: Invert the original subtitle background parameters to render new subtitle text, maintaining a consistent style while ensuring the mood aligns with the new direction.

[0087] Furthermore, in step S1, the input is the original news video. Analyzing it yields: the original streamer's footage Original audio Original subtitle backdrop ;

[0088] The upper-level strategy engine sets the public opinion guidance intensity parameter G to indicate the positive or negative degree of the overall stance, and also defines a multi-dimensional sentiment vector. Through the valence mapping vector (symmetric vector s) and the weight matrix A valence constraint is established, which realizes a continuous mapping from unidimensional guidance intensity to multidimensional emotion distribution;

[0089] Specifically, the guidance strength parameter is defined by a multi-dimensional emotion vector, and valence constraints are established by combining a valence mapping vector and a weight matrix; wherein,

[0090] The valence constraint is used for a continuous mapping from unidimensional guidance intensity to a multidimensional emotion distribution, expressed as:

[0091]

[0092] in, Represents a valence mapping vector. Represents the weight matrix. This is used to control the contribution of each dimension to polarity. Represents a multidimensional emotion vector. Indicates the guiding strength parameter;

[0093] The guidance setting parameters include:

[0094] Single-dimensional directional polarity:

[0095] ;

[0096] Multidimensional sentiment vector (target public opinion distribution):

[0097] ;

[0098] in, These represent emotional scores across eight dimensions: anger, fear, joy, trust, sadness, disgust, anticipation, and surprise.

[0099] Valence mapping vector: .

[0100] Furthermore, in step S2, multimodal element analysis is performed through the following steps:

[0101] The original news video's audio was processed using a speech recognition model to convert the audio into text, resulting in audio-text. ;Right now: ;

[0102] Natural language processing algorithms are used to segment paragraphs based on semantic boundaries, resulting in sentence structures. The clause structure includes clauses and time boundaries; that is: ; This indicates the start and end positions of the current paragraph k, where K is the total number of paragraphs.

[0103] F fact triples F are extracted through dependency parsing and used as an immutable fact semantic base; that is: In this context, sub represents the subject, pred represents the predicate, and obj represents the object. State the facts;

[0104] The set of rhetorical words R for each paragraph is identified using attention mechanisms and an emotion dictionary; that is: ;in, This refers to the rhetorical devices in the current paragraph k.

[0105] Furthermore, in step S3, the multidimensional emotion intensity of rhetorical words is dynamically calculated in context using a large language model to generate word-level emotion vectors, and then the word-level emotion vectors are aggregated into a segment-level emotion vector; wherein,

[0106] Word-level sentiment vector calculation (generated by LLM)

[0107]

[0108] in: Rhetorical devices, Indicates the context of a paragraph. The output is an 8-dimensional emotion vector: ;

[0109] LLM generates a score of [0,1] for each dimension using a natural language prompt. For example: "Please rate the phrase 'significant opportunity' on each of the eight dimensions: anger, fear, joy, trust, sadness, disgust, anticipation, and surprise (0-1)."

[0110] The segment-level sentiment vector is represented as follows:

[0111]

[0112] The global sentiment distribution of the original video is represented as follows:

[0113]

[0114] in, Represents segment-level sentiment vectors. Represents word-level sentiment vectors. Paragraph A collection of rhetorical devices.

[0115] Furthermore, in step S4, a target sentiment is generated for each segment:

[0116] The segment-level target sentiment vector is generated using the following formula:

[0117] ,

[0118] in, Indicates the extreme emotion of the target segment. Indicates the segment weight (which can be taken as sentence length or word count). Indicates paragraph number, This represents a multidimensional emotion vector.

[0119] Rhetorical substitution rules: based on the difference Emotional intensification or weakening corresponds to the following replacement: Public expression:

[0120]

[0121] For example: if Replace positive words with "warning," "concern," etc., if Replace words of joy with "brief" or "superficial".

[0122] Copywriting generation optimization: Based on fact locking and rhetorical rewriting, constrained generation is performed using LLM (embedded sentiment distribution control in the prompt) and a new copy is solved. :

[0123]

[0124] in: This indicates that the facts remain unchanged. To maintain a journalistic tone, Expressing emotions and goals match.

[0125] Furthermore, in step S5, prosodic parameters are extracted from the original audio. ;in, These represent speech rate, average pitch, pauses, and fluctuations, respectively.

[0126] When performing text-to-speech conversion, the following expression is used:

[0127] =TTS(

[0128] in, , For the new copywriting.

[0129] Then, based on the lip movements of the broadcaster in the generated voice-driven library, a broadcast video of the broadcaster is generated; Example: ("positive" → "negative 70%"); Original sentence:

[0130] "Economic cooperation between country A and country B continues to heat up, which experts believe is a major opportunity to promote the development of enterprises in country B."

[0131] Original level emotion (dynamically calculated by LLM):

[0132] (0.10, 0.10, 0.70, 0.60, 0.10, 0.00, 0.50, 0.10)

[0133] Target emotion (70% negative):

[0134] (0.28, 0.52, 0.28, 0.24, 0.34, 0.18, 0.30, 0.09)

[0135] Difference:

[0136] (+0.18,+0.42,-0.42,-0.36,+0.24,+0.18,-0.20,-0.01)

[0137] →Increase fear / anger / sadness / disgust, decrease joy / trust.

[0138] Rhetorical substitution results:

[0139] "Continued escalation" → "Over-reliance"

[0140] "Significant Opportunities" → "Potential Risks"

[0141] Generated text: The cooperation between country A and country B seems to be going well, but experts warn that over-reliance by companies in country B may bring long-term risks.

[0142] The new copy was validated by LLM sentiment analysis: Valence = -0.7, which is completely in line with the goal orientation.

[0143] Furthermore, in step S6, a new subtitle background is generated through the following steps:

[0144] The system detects and identifies subtitle regions in each frame (e.g., using Segment-Anything detection), and merges temporally consecutive subtitle regions from different frames to form a set of subtitle segments; where a subtitle region is represented as... , These represent the horizontal and vertical coordinates, as well as the width and height parameters, respectively.

[0145] Calculate and extract subtitle style parameters for each subtitle region. The expression is:

[0146] (

[0147] in, Indicates the color of the back panel. Indicates font, Indicates font color; its font contrast: ; This indicates whether the font size is large or small, estimated from the height of the text border. The location is calculated from the center of the region; Transparency is indicated by the difference in brightness. This indicates a shadow / outline, determined by the edge gradient intensity.

[0148] Furthermore, the template most similar to the subtitle area is found in the template library as the new subtitle background: if the minimum distance between the template and the subtitle area in the template library is greater than a threshold, the AI ​​generation module is called to generate a new template based on the subtitle area.

[0149] By associating segment-level emotion vectors with subtitle style parameters and achieving cross-modal emotion expression consistency through a visual-emotion mapping mechanism, new subtitle style parameters are obtained.

[0150] Specifically, segment-level emotions With subtitle style parameters This association enables consistency in cross-modal emotion expression.

[0151]

[0152] in, A paragraph-level sentiment vector representing a segment; This represents the target sentiment vector after guidance optimization; Indicates the direction and intensity of emotional changes; This represents the visual modulation weight matrix that maps this "emotional change" to "visual parameter change"; , These represent the original and new subtitle style parameters (background color, font color, transparency, shadow, etc.).

[0153] The source and construction method are based on empirical modeling, as detailed in Table 1;

[0154] Table 1. Empirical Modeling Table

[0155]

[0156] For example:

[0157] If fear or sadness increases: the back panel color will change. Adjust to dark blue / gray tone; font color Adjust to cool colors (blue, gray, white);

[0158] If joy is added: the background color will be adjusted to a warm color (orange, yellow); the font color will be adjusted to high saturation and brightness.

[0159] If the angle is increased: the background will be adjusted to a red tone; the shadow intensity and font weight will be increased.

[0160] Furthermore, after generating the new text, new voice, new digital anchor visuals, and new subtitle style parameter Lnew, the system enters the video reconstruction and final synthesis stage.

[0161] In step S7, based on the original video scene structure, the system intelligently distinguishes between the broadcast segments and non-broadcast segments, and performs video replacement and subtitle overlay respectively; among which,

[0162] If a frame of the anchor is detected, the most similar background template is matched from the background template library based on the studio characteristics of the original news video as the new background, and then merged with the anchor frame to replace the entire image.

[0163] If no anchor's image frame is detected, the original main content of the video remains unchanged, and only the generated audio track and the corresponding new subtitle background and new text are superimposed.

[0164] Among them, the background style and text format of the new subtitle background are automatically adjusted according to the new subtitle style parameters to ensure that the visuals and text are consistent in terms of emotional guidance;

[0165] The video footage is spliced ​​together in the order of the audio, and transition effects are added at the splicing points to make the transition more natural.

[0166] The final video maintains consistency with the original video in terms of overall picture structure, studio style and rhythm, but fully matches the new public opinion guidance parameters in terms of spoken text, tone of voice and visual subtitle style, achieving the effect of automatic news re-synthesis with "unchanged content structure and controllable presentation guidance".

[0167] Example 2

[0168] Please see Figure 2 Based on Embodiment 1, this embodiment also provides an automatic news video generation system that maintains the original style through guided reconstruction, including:

[0169] Input and guidance setting module: Input the original news video and set the guidance intensity parameter; wherein, the original news video includes the original anchor screen, original audio, and original subtitle background, and the guidance intensity parameter includes a multi-dimensional emotion vector and the corresponding valence constraint;

[0170] News Multimodal Element Analysis Module: Performs multimodal element analysis on the original audio of the original news video to extract speech text, sentence structure, fact triples, and rhetorical word set;

[0171] Dynamic emotion scoring module: Based on a large language model, dynamically calculate the emotion vector of each rhetorical word in the rhetorical word set and aggregate them into segment-level emotion vectors;

[0172] Guidance and copywriting reconstruction module: Based on the guidance strength parameter, a paragraph-level target sentiment vector is assigned to each paragraph through a constraint optimization algorithm, and rhetorical replacement is performed based on the sentiment difference between the paragraph-level sentiment vector and the paragraph-level target sentiment vector to generate new copywriting;

[0173] Voice and Anchor Generation Module: Extracts prosodic parameters from the original audio and generates new prosodic parameters based on the prosodic parameters adjusted according to the emotional difference; generates new voice based on the new text and new prosodic parameters using text-to-speech technology; and generates new anchor visuals based on the anchor lip movements in the voice-driven library and the new voice.

[0174] Visual layer parsing and subtitle generation module: Parses and extracts the subtitle style parameters of the original subtitle background, and adjusts the subtitle style parameters according to the emotion difference through the visual-emotion mapping mechanism to generate a new subtitle background;

[0175] Video reconstruction and compositing module: By replacing the anchor's image with a new one, overlaying the new voice and new subtitle background, a news video with a consistent style and a reconstructed direction is generated.

[0176] It should be noted that this embodiment achieves controllable regeneration of the expression orientation of news videos under the premise that the facts remain unchanged by constructing a multimodal algorithm chain of "fact-rhetoric separation + dynamic emotion scoring + guidance vector optimization + emotion visual mapping + segmented video reconstruction". Its innovation lies in the collaborative control mechanism of multimodal features and the video synthesis method driven by guidance vector, which is an original technical solution in the field of intelligent generation of public opinion expression.

[0177] This embodiment uses a multimodal element parsing and hierarchical extraction mechanism to simultaneously extract text content and visual subtitle style parameters (color, font, position, etc.) from the original video, and structures the two into a unified parameter space to establish a multimodal correspondence between the semantic layer and the visual layer.

[0178] This embodiment constructs a dynamic sentiment scoring and guidance optimization algorithm driven by a large language model. The large language model is used as a dynamic sentiment scorer to calculate the multidimensional sentiment vector of each rhetorical word. The continuous and controllable adjustment of public opinion guidance is achieved through a constrained optimization formula.

[0179] This embodiment also constructs a fact-rhetoric separation and controlled reconstruction mechanism, using dependency syntax and sentiment word localization algorithms to separate news fact triples. With rhetorical expression Separate; in the reconstruction phase, only the rhetorical part is modified, while the factual part remains completely unchanged, thus changing the direction while preserving the content.

[0180] This embodiment establishes a mapping function between emotion vectors and visual caption style parameters. Adjusting the weight matrix visually The system maps directional changes (such as a negative 70%) to changes in parameters such as subtitle background color, brightness, and font, achieving visually consistent emotional expression. When the original video subtitle template cannot match a template in the library, the system calls a generative model to automatically generate a new subtitle background, ensuring that different video sources can be re-composited with a consistent style.

[0181] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for automatically generating news videos while preserving the original style based on guided reconstruction, characterized in that, Includes the following steps: Step S1: Input the original news video and set the guidance intensity parameter; wherein, the original news video includes the original anchor screen, the original audio, and the original subtitle background, and the guidance intensity parameter includes a multi-dimensional emotion vector and the corresponding valence constraint; Step S2: Perform multimodal element analysis on the original audio of the original news video to extract speech text, sentence structure, fact triples and rhetorical word set; Step S3: Dynamically calculate the sentiment vector of each rhetorical word in the rhetorical word set based on the large language model, and aggregate them into a segment-level sentiment vector; Step S4: Based on the guidance intensity parameter, assign a paragraph-level target sentiment vector to each paragraph using a constraint optimization algorithm, and perform rhetorical replacement based on the sentiment difference between the paragraph-level sentiment vector and the paragraph-level target sentiment vector to generate new copy. Step S5: Extract the prosodic parameters of the original audio, and generate new prosodic parameters based on the prosodic parameters adjusted according to the emotional difference; use text-to-speech technology to generate new speech based on the new text and the new prosodic parameters; generate a new broadcaster's image based on the broadcaster's lip movements in the speech-driven library and the new speech. Step S6: Parse and extract the subtitle style parameters of the original subtitle background, and adjust the subtitle style parameters according to the emotion difference through the visual-emotion mapping mechanism to generate a new subtitle background; Step S7: By replacing the new anchor's screen with the new voice and the new subtitle background, a news video with a redesigned direction and consistent style is generated; In step S1, the guidance strength parameter is defined by a multi-dimensional emotion vector, and valence constraints are established by combining the valence mapping vector and the weight matrix; wherein, The valence constraint is used for a continuous mapping from unidimensional guidance intensity to a multidimensional emotion distribution, expressed as: in, Represents a valence mapping vector. Represents the weight matrix. Represents a multidimensional emotion vector. This represents the guiding strength parameter, and the one-dimensional guiding polarity is set to... -1 = negative 100%, 0 = neutral, +1 = positive 100%.

2. The method for automatically generating news videos based on guided reconstruction while maintaining the original style, as described in claim 1, is characterized in that... In step S2, multimodal element analysis is performed through the following steps: The original news video's audio was processed using a speech recognition model to convert the audio into text, resulting in audio-text. The paragraph set is divided according to semantic boundaries using natural language processing algorithms to obtain the sentence structure; wherein, the sentence structure includes sentence division and time boundary; F fact triples F are extracted through dependency parsing and used as an immutable fact semantic base; The set of rhetorical words for each paragraph is identified using attention mechanisms and an emotion dictionary.

3. The method for automatically generating news videos based on guided reconstruction while maintaining the original style, as described in claim 1, is characterized in that... In step S3, the multidimensional sentiment intensity of rhetorical words is dynamically calculated in context using a large language model to generate word-level sentiment vectors. These word-level sentiment vectors are then aggregated into segment-level sentiment vectors. The segment-level sentiment vector is represented as follows: in, Represents segment-level sentiment vectors. Represents word-level sentiment vectors. Paragraph A collection of rhetorical devices.

4. The method for automatically generating news videos based on guided reconstruction while maintaining the original style, as described in claim 1, is characterized in that... In step S4, the segment-level target sentiment vector is generated using the following formula: , in, Indicates the extreme emotion of the target segment. Indicates segment weight, Indicates paragraph number, This represents a multidimensional emotion vector.

5. The method for automatically generating news videos based on guided reconstruction while maintaining the original style, as described in claim 1, is characterized in that... In step S4, new copy is generated under constraints by using fact-based locking and rhetorical rewriting, combined with LLM; wherein, sentiment distribution control is embedded in the prompt of the LLM.

6. The method for automatically generating news videos based on guided reconstruction while maintaining the original style, as described in claim 1, is characterized in that... In step S5, the prosodic parameters include speech rate, average pitch, pauses, and fluctuations.

7. The method for automatically generating news videos based on guided reconstruction while maintaining the original style, as described in claim 1, is characterized in that... In step S6, a new subtitle background is generated through the following steps: The system detects and identifies subtitle regions in each frame, and merges the same subtitle regions in frames that are time-continuous to form a set of subtitle segments. Calculate and extract subtitle style parameters for each subtitle region; Find the template with the most similarity to the subtitle area from the template library as the new subtitle background: if the minimum distance between the template and the subtitle area in the template library is greater than a threshold, then call the AI ​​generation module to generate a new template based on the subtitle area; By associating segment-level emotion vectors with subtitle style parameters and achieving cross-modal emotion expression consistency through a visual-emotion mapping mechanism, new subtitle style parameters are obtained.

8. The method for automatically generating news videos based on guided reconstruction while maintaining the original style, as described in claim 7, is characterized in that... In step S7, based on the original video scene structure, the system intelligently distinguishes between the broadcast segments and non-broadcast segments, and performs video replacement and subtitle overlay respectively; among which, If a frame of the anchor is detected, the most similar background template is matched from the background template library based on the studio characteristics of the original news video as the new background, and then merged with the anchor frame to replace the entire image. If no anchor's image frame is detected, the original main content of the video remains unchanged, and only the generated audio track and the corresponding new subtitle background and new text are superimposed. Among them, the background style and text format of the new subtitle backdrop are automatically adjusted according to the new subtitle style parameters to ensure that the visuals and text are consistent in terms of emotional guidance; The video footage is spliced ​​together in the order of the audio, and transition effects are added at the splicing points to make the transitions more natural.

9. A news video automatic generation system based on guided reconstruction while preserving the original style, employing the news video automatic generation method based on guided reconstruction while preserving the original style as described in any one of claims 1 to 8, characterized in that, include: Input and guidance setting module: Input the original news video and set the guidance intensity parameter; wherein, the original news video includes the original anchor screen, original audio, and original subtitle background, and the guidance intensity parameter includes a multi-dimensional emotion vector and the corresponding valence constraint; News Multimodal Element Analysis Module: Performs multimodal element analysis on the original audio of the original news video to extract speech text, sentence structure, fact triples, and rhetorical word set; Dynamic emotion scoring module: Based on a large language model, dynamically calculate the emotion vector of each rhetorical word in the rhetorical word set and aggregate them into segment-level emotion vectors; Guidance and copywriting reconstruction module: Based on the guidance strength parameter, a paragraph-level target sentiment vector is assigned to each paragraph through a constraint optimization algorithm, and rhetorical replacement is performed based on the sentiment difference between the paragraph-level sentiment vector and the paragraph-level target sentiment vector to generate new copywriting; Voice and Anchor Generation Module: Extracts prosodic parameters from the original audio and generates new prosodic parameters based on the prosodic parameters adjusted according to the emotional difference; generates new voice based on the new text and new prosodic parameters using text-to-speech technology; and generates new anchor visuals based on the anchor lip movements in the voice-driven library and the new voice. Visual layer parsing and subtitle generation module: Parses and extracts the subtitle style parameters of the original subtitle background, and adjusts the subtitle style parameters according to the emotion difference through the visual-emotion mapping mechanism to generate a new subtitle background; Video reconstruction and compositing module: By replacing the new anchor's image with the new voice and the new subtitle background, a news video with a consistent style and direction is generated; In the input and guidance setting module, the guidance intensity parameter is defined by a multi-dimensional emotion vector, and valence constraints are established by combining a valence mapping vector and a weight matrix; wherein, The valence constraint is used for a continuous mapping from unidimensional guidance intensity to a multidimensional emotion distribution, expressed as: in, Represents a valence mapping vector. Represents the weight matrix. Represents a multidimensional emotion vector. This represents the guiding strength parameter, and the one-dimensional guiding polarity is set to... -1 = negative 100%, 0 = neutral, +1 = positive 100%.

Citation Information

Patent Citations

  • Method, electronic device, and computer program product for video processing

    US20240070956A1

  • KR20250053282A