Content generation method and apparatus, electronic device, and storage medium

By acquiring the associated data of the target entity, using a large model to generate content scripts that conform to the preset narrative structure, and combining preset tools to acquire media materials, the problems of content generation logic and multi-platform adaptation in existing technologies are solved, and efficient and professional multimodal content generation is achieved.

CN122133620APending Publication Date: 2026-06-02BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-02-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing content generation technologies lack logic and authenticity, cannot meet the requirements of multiple platform formats, involve a high degree of human intervention, and are difficult to automate with high reliability.

Method used

By acquiring the associated data of the target entity, a large model is used to generate a content script that conforms to the preset narrative structure. Media materials are acquired by combining preset tools, and multimodal content in various output formats is generated according to the preset narrative structure.

Benefits of technology

It enhances the logical rigor, professionalism, and factual credibility of the content, achieves automated adaptation of content formats across multiple platforms, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133620A_ABST
    Figure CN122133620A_ABST
Patent Text Reader

Abstract

This disclosure provides a content generation method, apparatus, electronic device, and storage medium, relating to the field of artificial intelligence technology, particularly to the fields of large-scale models, deep learning, natural language processing, and computer vision. The specific implementation scheme includes: acquiring associated data of a target entity, wherein the associated data includes news data related to the target entity and entity attribute data of the target entity obtained from a database; generating a content script conforming to a preset narrative structure using a large-scale model based on the news data and entity attribute data, the content script including content fragments and rendering instructions corresponding to the content fragments; calling preset tools to obtain corresponding media materials according to the rendering instructions; and generating multimodal content in multiple output formats according to the preset narrative structure based on the content fragments and media materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of large models, deep learning, natural language processing, and computer vision. More specifically, this disclosure provides a content generation method, apparatus, electronic device, storage medium, and computer program product. Background Technology

[0002] Currently, content creation mainly relies on keyword matching of common images and other materials, resulting in content with weak narrative logic and a lack of authenticity. Furthermore, the current production model typically employs a single approach, failing to accommodate the format requirements of multiple platforms, and involves high levels of human intervention, making it difficult to achieve highly reliable automated production. Summary of the Invention

[0003] This disclosure provides a content generation method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to the first aspect, a content generation method is provided, the method comprising: acquiring associated data of a target entity, wherein the associated data includes news data related to the target entity and entity attribute data of the target entity obtained from a database; generating a content script conforming to a preset narrative structure using a large model based on the news data and entity attribute data, the content script including content fragments and rendering instructions corresponding to the content fragments; invoking preset tools to obtain corresponding media materials according to the rendering instructions; and generating multimodal content in multiple output formats according to the preset narrative structure based on the content fragments and media materials.

[0005] According to the second aspect, a content generation apparatus is provided, comprising: an acquisition module for acquiring associated data of a target entity, wherein the associated data includes news data related to the target entity and entity attribute data of the target entity obtained from a database; a first generation module for generating a content script conforming to a preset narrative structure using a large model based on the news data and entity attribute data, the content script including content fragments and rendering instructions corresponding to the content fragments; a processing module for acquiring corresponding media materials by calling preset tools according to the rendering instructions; and a second generation module for generating multimodal content in multiple output formats according to the content fragments and media materials and a preset narrative structure.

[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to the present disclosure.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the methods provided in this disclosure.

[0008] According to a fifth aspect, a computer program product is provided, comprising a computer program stored on at least one of a readable storage medium and an electronic device, wherein the computer program, when executed by a processor, implements the method provided in this disclosure.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 This is an exemplary system architecture diagram of a content generation method and apparatus applicable according to an embodiment of the present disclosure;

[0012] Figure 2 This is a flowchart of a content generation method according to an embodiment of the present disclosure;

[0013] Figure 3 This is a flowchart of a content generation method according to another embodiment of the present disclosure;

[0014] Figure 4 This is a flowchart of a content generation method according to another embodiment of the present disclosure;

[0015] Figure 5 This is a flowchart of a content generation method according to another embodiment of the present disclosure;

[0016] Figure 6 This is a flowchart of a content generation method according to another embodiment of the present disclosure;

[0017] Figure 7 This is a flowchart of a content generation method according to another embodiment of the present disclosure;

[0018] Figure 8 This is a block diagram of a content generation apparatus according to an embodiment of the present disclosure;

[0019] Figure 9 This is a block diagram of a content generation apparatus according to another embodiment of the present disclosure;

[0020] Figure 10 This is a block diagram of a content generation apparatus according to another embodiment of the present disclosure; and

[0021] Figure 11 This is a block diagram of an electronic device using a content generation method according to an embodiment of the present disclosure. Detailed Implementation

[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0023] The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.

[0024] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0025] Currently, in content generation scenarios such as generating text and images and short videos, the common solutions are "template filling" or "simple text-to-video conversion". Text-to-video conversion breaks down the article into sentences, searches for a generic image for each sentence based on keywords, and then creates a simple slideshow with accompanying audio. Template filling uses a preset template, replacing only the text fields and background image.

[0026] The aforementioned solutions suffer from several problems: weak logic, monotonous and inaccurate visual formats, lack of refined arrangement, and inadequate risk control. Weak logic means an inability to understand complex narrative structures (such as a "four-part" financial analysis), resulting in videos that are merely chronological accounts lacking coherence. Monotonous and inaccurate visual formats mean an inability to handle multi-source, heterogeneous data. For example, when the text mentions "net profit growth of 20%", it might typically match a generic image of a gold coin instead of automatically generating a realistic financial statement screenshot or dynamic chart. Lack of refined arrangement means a lack of control over the video's pacing, failing to implement complex rendering logic such as "typewriter effects," "picture-in-picture," and "split-screen comparison" that are automatically triggered based on content type. Inadequate risk control means a lack of automated verification processes for sensitive text and key values ​​(years, financial data) in the generated content.

[0027] Figure 1 This is a schematic diagram of an exemplary system architecture for a content generation method and apparatus applicable according to an embodiment of this disclosure. It should be noted that... Figure 1The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0028] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices, including but not limited to smartphones, tablets, laptops, etc.

[0030] Server 105 can be a server providing various services, such as a backend management server supporting websites browsed by users using terminal devices 101, 102, and 103 (this is just an example). The backend management server can analyze and process received user requests and other data, and feed back the processing results (e.g., generating multimodal content in various output formats according to a preset narrative structure based on content fragments and media materials) to the terminal devices. Server 105 can be deployed with a trained pre-defined model, which can be a Large Language Model (LLM). This model can generate content scripts that conform to a preset narrative structure. The content scripts can include content fragments and rendering instructions corresponding to the content fragments.

[0031] The content generation method provided in this disclosure can generally be executed by server 105. Correspondingly, the content generation apparatus provided in this disclosure can generally be located in server 105. The content generation method provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the content generation apparatus provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0032] Figure 2 This is a flowchart of a content generation method according to an embodiment of the present disclosure.

[0033] like Figure 2As shown, the content generation method 200 includes operations S210 to S240.

[0034] In operation S210, the associated data of the target entity is obtained, including news data related to the target entity and entity attribute data of the target entity obtained from the database.

[0035] In this embodiment, the target entity can specifically refer to an enterprise, institution, or specific brand. News data can refer to timely external public information, such as social media trending topics, news reports, and industry dynamics, used to provide the narrative background of the content. The database can refer to an internal enterprise database, specifically a database obtained by collecting and organizing publicly available enterprise data, such as annual reports and business registration information. Entity attribute data can be business registration data or financial statement data obtained from this internal enterprise database. Specifically, the associated data of the target entity can be obtained through application programming interface (API) calls, database queries, or internal system pushes.

[0036] In operation S220, based on news data and entity attribute data, a large model is used to generate a content script that conforms to a preset narrative structure. The content script includes content fragments and rendering instructions corresponding to the content fragments.

[0037] In this embodiment, the preset narrative structure can refer to the logical framework of the generated content, and the large model refers to a deep learning model with natural language processing capabilities, responsible for understanding, summarizing, reorganizing, and logically reasoning about multi-source data. The content script is, for example, a four-part script, where content segments refer to content units broken down from the script. For example, the first segment might be the opening title, the second segment a presentation with effects such as a typewriter, the third segment screenshots of internal company information, and the fourth segment a concluding summary.

[0038] The script also includes rendering instructions for each segment. For example, the rendering instruction for the second segment could be to generate a presentation listing viewpoints, and the rendering engine can then generate the corresponding presentation based on the rendering instruction. Another example is that rendering instructions can be JSON-formatted instructions that define the visual presentation method, such as "display a financial statement bar chart in the current segment," "trigger a typewriter effect," or "replace the background with a specific business snapshot."

[0039] When operating the S230, the preset tools are invoked to obtain the corresponding media materials according to the rendering instructions.

[0040] In this embodiment of the disclosure, the preset tool may refer to an automated background plugin or service, such as a webpage long screenshot tool, a dynamic chart generation engine, a video material library indexing program, etc., and the corresponding media material may refer to information that is highly matched with the rendering instruction. For example, if the rendering instruction is "display a financial statement bar chart in the current segment", then a dynamic financial statement bar chart can be generated using a dynamic chart generation engine.

[0041] When operating the S240, multimodal content in various output formats is generated based on content fragments and media materials according to a preset narrative structure.

[0042] In this embodiment of the disclosure, multiple output formats can refer to differentiated media forms such as videos or text and image notes. After generating the content script, the media materials can be processed according to the content fragments and media materials in accordance with the preset narrative structure, thereby generating multimodal content such as videos or text and image notes.

[0043] This disclosure breaks down data silos by deeply integrating news data related to the target entity and entity attribute data of the target entity obtained from the database, and generates content based on data evidence, thereby enhancing the logical rigor, professionalism and factual credibility of the content.

[0044] In addition, generating multimodal content in multiple output formats can solve the problem of high labor costs caused by the separation of text and video production processes in current content generation scenarios. It enables the generation of multiple content formats that are compatible with different platform specifications for the same data source.

[0045] According to embodiments of this disclosure, obtaining associated data of a target entity includes: obtaining a target topic and obtaining news data from multiple information sources based on the target topic; determining the target entity and news data related to the target entity based on entity words extracted from the news data; and obtaining entity attribute data of the target entity from a database based on the entity identifier corresponding to the target entity.

[0046] In this embodiment of the disclosure, the Uniform Resource Locator (URL), text, theme image, and video material of external news can be obtained by calling the API, resulting in massive news data. Then, the Named Entity Recognition (NER) model can be used to extract entity words from the news data, such as company abbreviations or brand names. The target entity can be determined based on the entity words. For example, the company corresponding to the abbreviation can be determined based on the company abbreviation. The company is taken as the target entity, and news data related to the target entity can be further determined.

[0047] The target entity is bound to the unique identifier of the internal enterprise database. That is, the target entity has a unique entity identifier in the database, so the entity attribute data of the target entity can be obtained from the database using the entity identifier.

[0048] By introducing an entity identifier binding mechanism, precise correlation was achieved between unstructured news data and structured enterprise databases. This broke down the silos between external trending topics and internal professional data, improving the automation level of content generation and data accuracy.

[0049] According to embodiments of this disclosure, obtaining a target topic includes: filtering target statements of a preset category from search statements that meet popularity criteria; and semantically expanding the target statements using a large model to obtain the target topic.

[0050] In this embodiment, the popularity criterion can be a quantitative indicator measuring whether a search term or topic has dissemination value. For example, it can be achieved by monitoring news hot lists or search hotspots to obtain real-time search volume, click-through rate, or topic growth rate data. When the growth rate of a search term exceeds a preset threshold (e.g., a month-on-month increase of 50%) or enters the top N of the hot list, it is determined to meet the popularity criterion. Search terms that meet the popularity criterion are then filtered to obtain target terms of a preset category, where the preset category can refer to corporate news. The target terms of the preset category can be obtained through a trained semantic classification model; that is, the semantic classification model can be used to filter out target terms of the preset category from search terms that meet the popularity criterion.

[0051] After filtering out the target statements in the preset categories, the target statements can be expanded using a large model. For example, if the target statement is a company's financial report, the expanded target topic could be an in-depth analysis of the company's 2025 financial report.

[0052] By filtering target statements based on popularity criteria and preset categories, high-value topics can be automatically located from massive amounts of social media and search data.

[0053] According to embodiments of this disclosure, obtaining associated data of a target entity includes: for the target entity, obtaining entity attribute data of the target entity from a database, and obtaining news data of a preset time range associated with the target entity from multiple data sources.

[0054] In this embodiment of the disclosure, the entity attribute data may be data such as business registration and shareholder composition obtained from the database. The data source may be multiple external Internet information websites or self-media platforms that provide real-time information, social trends or hot topics. For the target entity, news data within a preset time range can be obtained through multiple data sources, such as the most recent month.

[0055] By acquiring news data within a preset time range associated with the target entity, it can be ensured that the generated content accurately utilizes the latest changes of the target entity, avoiding data lag.

[0056] According to embodiments of this disclosure, generating a content script conforming to a preset narrative structure using a large model based on news data and entity attribute data includes: generating a title and body text based on news data and entity attribute data; rewriting the title and body text into layout data conforming to the preset narrative structure, the layout data including multiple content fragments; and determining rendering instructions corresponding to the content fragments based on their types.

[0057] In this embodiment, after acquiring news data and entity attribute data, a large model can be used to understand and refine the news data and entity attribute data, generating attractive titles and logically structured body content. The titles and body content are then rewritten according to a narrative structure to obtain layout data that conforms to front-end specifications. The layout data clearly defines the content's arrangement in visual space and along the timeline; for example, it specifies where the title is, where the illustrations are, and other layout information. After obtaining layout data including multiple content fragments, rendering instructions corresponding to the content fragments can be determined based on their types.

[0058] According to embodiments of this disclosure, determining the rendering instruction corresponding to the content segment based on the type of the content segment includes: in response to the content segment being a start segment or an end segment, determining the rendering instruction as an image acquisition instruction; in response to the content segment being a middle segment, determining the rendering instruction as at least one of an image acquisition instruction, a chart generation instruction, a screenshot instruction, an animation effect instruction, and a speech synthesis instruction.

[0059] In this embodiment of the disclosure, the type of content segment may include a start segment, a middle segment, and an end segment. In the start segment or the end segment, the rendering instruction may be an image acquisition instruction. Based on the image acquisition instruction, a corresponding preset tool can be triggered to search for and retrieve a high-quality background image that matches the theme.

[0060] In the middle section, which typically needs to convey core facts, in-depth analysis, or key data, rendering instructions can be at least one of the following: image acquisition instructions, chart generation instructions, screenshot instructions, animation effects instructions, and speech synthesis instructions. For example, for segments involving financial growth, the rendering instruction could be a chart generation instruction, which can instantly render dynamic bar charts; for segments involving business evidence, the rendering instruction could be a screenshot instruction, which can retrieve real snapshots; at the same time, rendering instructions can also be combined with animation effects instructions (such as text jumping and transition effects) and speech synthesis instructions to achieve rich audiovisual expression.

[0061] This segment-type-based differentiated instruction strategy ensures that visual materials closely follow the narrative rhythm, avoiding the problem of materials being disconnected from the text. By introducing empirical materials such as screenshots and charts in the core middle paragraphs, the factual persuasiveness and professional depth of the generated content can be enhanced, making the final video or text notes have a stronger sense of logical hierarchy.

[0062] According to embodiments of this disclosure, invoking a preset tool to obtain corresponding media materials according to a rendering instruction includes: in response to an image acquisition instruction, invoking a cropping tool to obtain a main image from news data; in response to a screenshot instruction, invoking a screenshot tool to extract attribute information images from entity attribute data; and in response to an animation effect instruction, invoking an effects tool to convert the fragment content into a presentation containing animation effects.

[0063] In this embodiment, if the rendering instruction is an image acquisition instruction, a cropping tool integrating an intelligent cropping algorithm can be invoked. This tool can automatically perform semantic analysis on the original images attached to the news data, accurately identify and locate the core subjects in the image, such as news figures, released products, or specific logos. By automatically removing background noise and optimizing the composition, a "first-screen main image" adapted to different terminals (such as mobile vertical screens) is generated, ensuring the visual effect of the visual material on the first screen.

[0064] If the rendering command is a screenshot command, a backend snapshot screenshot service can be invoked. Specifically, based on the unique identifier of the target entity, the corresponding entity details page or data interface can be automatically accessed, and attribute information images can be selectively captured, such as automatically capturing the latest equity structure diagram, business registration snapshot, or key financial statements. Using the real-time rendered snapshot as video material can solve the problem of the authenticity of the generated content, directly providing empirical support for the generated videos or text / images.

[0065] If the rendering command is an animation effects command, the effects generation engine can be invoked to transform plain text fragments into slideshow components with dynamic visual effects. For example, based on the logical hierarchy of the fragment content, a list of viewpoints with typewriter effects or visual cards with panning and zooming effects can be automatically generated. This method transforms static text narration into dynamic audiovisual language, resulting in content that maintains professional depth while possessing greater interactivity and visual appeal.

[0066] According to embodiments of this disclosure, the method further includes: obtaining historical entity attribute data from a database; generating trend data and explanatory text based on the historical entity attribute data using a large model; and generating a trend display video based on the trend data and explanatory text.

[0067] In this embodiment of the disclosure, the backend engine can retrieve historical entity attribute data of the target entity from the database, such as the company's annual revenue over the past five years, stock price fluctuation records, or brand search popularity values. Then, a large model is used to perform in-depth analysis and aggregation of this historical data, and based on this, accurate explanatory text and structured data for describing the changing trends are automatically generated.

[0068] After generating trend data, the front-end rendering engine is driven to dynamically draw visual charts, such as smooth stock price trend curves or 3D financial growth bar charts. These dynamic charts can be recorded as high-quality video streams and aligned with the audio narration generated by the large model. Then, the dynamic video stream, automatically generated narration, and preset dynamic cover are automatically packaged to obtain a trend display video that intuitively showcases the target entity.

[0069] Figure 3 This is a flowchart of a content generation method according to another embodiment of the present disclosure.

[0070] like Figure 3 As shown, the content generation method 300 includes operations S310 to S340.

[0071] In operation S310, historical entity attribute data of the target entity is obtained.

[0072] For example, retrieve long-term attribute data of the target entity from the backend database, such as the company's stock price fluctuations, revenue changes, or market popularity indicators over the past five years.

[0073] In operation S320, historical entity attribute data is aggregated.

[0074] For example, historical data can be cleaned and aggregated to extract core trend features that reflect growth logic or turning points.

[0075] When operating the S330, the front end dynamically renders and generates trend charts and records them as video streams, while the large model generates narration and titles and generates a cover.

[0076] Dynamic "trend charts" (such as smooth stock price trend curves) are generated based on aggregated data and recorded as video streams in real time. Large models are used to semantically interpret the trends, automatically write narration and titles that match the rhythm of the video, and generate visual covers.

[0077] When operating the S340, the backend assembles the cover and video stream to generate a trend chart video.

[0078] The cover image, the recorded trend chart video stream, and the synthesized audio narration are packaged together to produce a complete trend display video.

[0079] According to embodiments of this disclosure, generating multimodal content in multiple output formats based on content fragments and media materials according to a preset narrative structure includes: combining content fragments and media materials according to a preset narrative structure to generate at least one of video data and graphic data.

[0080] For example, in the scenario of generating video, the first screen can be a composite of the main news image and a sensational / suspenseful headline, while the second screen can display a snapshot of the company's details page automatically captured using the backend snapshot function, serving as video material for in-depth analysis. In the scenario of generating text and image data, based on the aforementioned generated title and body text, the rendering engine can dynamically adjust the font size, margins, and theme color according to the JSON layout data, compositing the cover image and body text into a long image.

[0081] Figure 4 This is a flowchart of a content generation method according to another embodiment of the present disclosure.

[0082] like Figure 4 As shown, the content generation method 400 includes operations S401 to S414.

[0083] Using S401, retrieve the main text of news materials.

[0084] Unstructured raw news text can be collected from external data sources via an interface as the initial information input for the entire content generation process.

[0085] When operating S402, input the main text of the news material into the large model.

[0086] The collected original news material is transmitted to a large model, where natural language understanding capabilities are used to perform semantic analysis and extract key information.

[0087] In operation S403, the target entity is identified.

[0088] Large models can be used to extract core entities (such as specific company names) from news articles using entity recognition algorithms, which can then be used as target entities.

[0089] When operating S404, determine the title and body of the video.

[0090] Large models can be used to write sensational or suspenseful headlines that conform to the dissemination patterns of short videos, and generate accompanying explanatory text.

[0091] When operating S405, publish the title and description.

[0092] In operation S406, the identifier of the target entity is determined.

[0093] Entity links map the identified company name to a unique entity identifier in the internal database. The entity identifier is an index that retrieves attribute data from the database.

[0094] When operating S407, call the backend internal data interface to obtain the basic information of the target entity.

[0095] By using entity identifiers to access the internal structured database, in-depth attribute data such as the company's establishment date, registered capital, and business status can be obtained.

[0096] When operating S408, the backend snapshot screenshot function is invoked to obtain a snapshot screenshot related to the target entity.

[0097] Based on the entity identifier, the backend tool is automatically driven to access the enterprise details page and instantly capture real snapshot images with empirical value, such as equity structure diagrams and financial statement data tables.

[0098] When operating S409, the large model summary yields the second screen of description information.

[0099] For example, the data obtained from operating S407 can be further refined to generate logically rigorous in-depth interpretation text, which serves as the core information flow for the second screen of the video.

[0100] When operating the S410, the screenshot and description information are combined to obtain the second screen video.

[0101] For example, the real snapshot generated by operation S408 and the interpretive text generated by operation S409 are synthesized in visual space to obtain a second screen video clip.

[0102] Using S411, obtain news source images.

[0103] For example, relevant image materials can be extracted from the original news source as the initial image input for the video visual stream.

[0104] When operating S412, the news material image is input into the image understanding model for analysis, and the intelligent cropping algorithm is called to obtain the main image of the first screen.

[0105] When operating S413, the video title and main image are aggregated to obtain the first screen of video.

[0106] For example, an attractive title can be overlaid on a cropped main image to create a high-click-through-rate video "first screen" (i.e., cover screen).

[0107] When operating S414, the first screen video and the second screen video are combined to obtain the large-character poster video.

[0108] For example, following a pre-set narrative logic, the first screen video, which serves as a cover, and the second screen video, which serves as in-depth empirical evidence, are sequentially linked on the timeline and packaged into a multimodal poster video.

[0109] It is understandable that the above operations S401 to S414 are the pipeline process for generating the poster video, which is a dual-stream parallel architecture. Operations S411 to S412 are information streams used to obtain the first screen main image, and operations S401 to S409 are visual streams used to obtain the second screen description information. After obtaining the first screen main image and the second screen description information, the poster video is obtained through a synthesis stream.

[0110] For example, in a scenario where a news URL for "XX Technology completes Series B financing" is detected, the news image and text can be retrieved. An image model can be used to intelligently crop the news image, preserving the main subject. Then, an LLM (Local Management Model) can be used to rewrite and generate a short, suspenseful headline, such as "Raising hundreds of millions?". The entity link is used to find the company's entity identifier, and the backend automatically calls a screenshot service to capture a real-time snapshot of its "shareholder information" page. The rendering engine then sequentially synthesizes the cropped news image, the headline, and the shareholder screenshot into a video, quickly generating a visually impactful 15-second video suitable for short video platforms.

[0111] For example, in a scenario where the input is "XX financial institution", its number of branches and financial report data can be retrieved. Then, an LLM script can be used to generate a four-segment script, and the instruction corresponding to the "many branches" segment is marked as needing to call the map screenshot component, while the instruction corresponding to the "inclusive finance" segment is needing to generate a presentation of viewpoints. The rendering engine, based on the instructions, mixes voice, background music, dynamic map screenshots, and presentation animations to generate a 1-minute professional financial interpretation video, which includes real data screenshots and a logically clear list of viewpoints.

[0112] Figure 5 This is a flowchart of a content generation method according to another embodiment of the present disclosure.

[0113] like Figure 5 As shown, the content generation method 500 includes operations S510 to S540.

[0114] Using S510 to obtain enterprise information data.

[0115] The core attribute information of the target entity can be retrieved from the internal structured database. This data may include the company's basic business registration information, core business qualifications, or key financial statement values.

[0116] When operating S520, input enterprise information data into the large model, generate titles and body text, and then rewrite them in a stylized manner.

[0117] By using a large model to process the input enterprise data and extract the core selling points, the professional data is transformed into attractive original copy through semantic understanding. Then, it can be stylized according to the preset audience needs, such as rewriting it into "product recommendation copy" that conforms to the characteristics of social media, thus producing the final title and body content.

[0118] When operating the S530, JSON layout data conforming to front-end specifications is generated and rendered by the front-end rendering engine.

[0119] The generated text is further converted into JSON format layout data that conforms to front-end typesetting specifications. The layout data can include specific typesetting logic and typesetting information. After receiving the JSON data, the front-end rendering engine will automatically and dynamically adjust visual parameters such as font size, margins, and theme color.

[0120] When operating the S540, the cover image and the body text are combined to create a long image.

[0121] For example, a preset or generated cover image can be combined with the formatted text content to output a complete, high-quality multimodal long image note that can be distributed on image-text social platforms.

[0122] According to embodiments of this disclosure, the method further includes: performing factual consistency verification on time information and numerical information in the content segment; and performing quality inspection and content compliance verification on the media material.

[0123] In this embodiment of the disclosure, after generating the content fragment, the time information and numerical information in the content fragment can also be verified for factual consistency. The time information can refer to the year, date or a specific time. Specifically, the time information appearing in the content fragment can be compared with the underlying knowledge base to prevent illusion problems such as fabricated years. The numerical information can include financial report data and percentages to ensure that the numerical information is accurate. At the same time, content relevance verification can also be performed to confirm that the generated content fragment has not deviated from the topic, thereby ensuring the rigor and objectivity of the content fragment.

[0124] Furthermore, after acquiring media materials, quality checks and content compliance verification can be performed. For example, the materials can be checked for low-quality images such as blurry or stretched images, and the content can be checked for non-compliant elements. Materials that fail the verification can then be deleted. Additionally, Contrastive Language-Image Pre-training (CLIP) models can be used to ensure semantic consistency between the media materials and content segments. These verification methods can help avoid logical errors or compliance risks in the generated content.

[0125] Figure 6 This is a flowchart of a content generation method according to another embodiment of the present disclosure.

[0126] like Figure 6 As shown, after obtaining the content to be published 610, if the content is an article or text, the article / text risk control pipeline needs to be executed. The content to be published is input into the first verification module 620 for time verification, numerical accuracy verification, and content relevance verification. If the content to be published is an image, the image risk control pipeline needs to be executed. The content to be published is input into the second verification module 630 for image sensitivity verification, image quality verification, and image-text consistency verification. After obtaining the verification results, they are input into the comprehensive judgment module 640 for comprehensive judgment. If the comprehensive judgment result is passed, the passed content to be published can be put into the publishing queue to wait for publication. If the comprehensive judgment result is failed, it means that it needs to be blocked. The specific handling method can be that the content to be published is manually reviewed, or the content to be published is regenerated.

[0127] According to embodiments of this disclosure, entity attribute data includes at least one of enterprise basic information and financial statement data.

[0128] In this disclosed embodiment, basic enterprise information may refer to information reflecting the identity characteristics and qualifications of an entity, specifically including business registration information, equity structure, qualifications, establishment time, registered capital, etc. Financial data may refer to structured values ​​reflecting the entity's operating status and financial results, specifically including core indicators such as operating revenue, net profit, asset-liability ratio, and cash flow.

[0129] Figure 7 This is a flowchart of a content generation method according to another embodiment of the present disclosure.

[0130] like Figure 7 As shown, after obtaining the verified content to be published, the content is published through the operation and publishing module 710. Specifically, the verified images or videos 712 are combined with the account matrix management module 711 for unified management. The account matrix management module 711 is responsible for maintaining the account resource library of multiple platforms, recording data such as the vertical field and activity level of each account, and then sending it to the publishing module 713. The publishing module 713 is responsible for receiving the generated finished content (images or videos), realizing one-click access to multiple platforms through API interface, and the intelligent scheduling module 714 is responsible for intelligently selecting accounts to publish the corresponding content based on account quality and content quality.

[0131] After content is published, a data feedback process is conducted through the data closed-loop module 720. Specifically, the data collection module 721 is responsible for capturing raw performance data from various platforms, while the core metric monitoring module 722 is responsible for monitoring key metrics reflecting content quality, such as click-through rate, completion rate, and fan interaction (comments / likes). The performance analysis module 723 is responsible for conducting in-depth analysis from both account and content perspectives, identifying viral content based on the monitored key metrics, determining the corresponding theme for the viral content, generating another form of content using that theme, and publishing it to another account.

[0132] Figure 8 This is a block diagram of a content generation apparatus according to an embodiment of the present disclosure, such as Figure 8 As shown, the content generation device includes a data acquisition layer 810, an intelligent arrangement layer 820, a material processing and rendering layer 830, and a risk control output layer 840.

[0133] The data acquisition layer 810 is responsible for acquiring external unstructured data, acquiring internal structured data from the database, cleaning and aggregating the data, and binding entity words in unstructured text with unique identifiers in the database.

[0134] The Intelligent Orchestration Layer 820 is responsible for using large models to perform logical reasoning and construct narrative structures, determining the focus and distribution direction of generated content based on business needs and content popularity, transforming generated content into layout data adapted to social media, generating content scripts that include subtitles, visual instructions and rendering parameters, and defining the change logic and narration of dynamic charts based on historical aggregated data.

[0135] The material processing and rendering layer 830 can include multiple components such as backend snapshot screenshot, presentation animation generation, intelligent cropping, and speech synthesis and recognition. It is responsible for processing the materials and synthesizing them into content to be published.

[0136] The risk control output layer 840 is responsible for performing risk control verification on the content to be published, deciding whether to block it or distribute it to different platforms based on the verification results, and collecting feedback on the results.

[0137] Figure 9 This is a block diagram of a content generation apparatus according to another embodiment of the present disclosure.

[0138] like Figure 9As shown, the data acquisition layer 910 includes an on-site enterprise data acquisition module 911 and an online trending news acquisition module 912. The on-site enterprise data acquisition module 911 obtains on-site enterprise data from the internal database, while the online trending news acquisition module 912 obtains online trending news from the external internet. The intelligent arrangement layer 920 includes a large model 921 and preset tools 922. The material processing and rendering layer 930 includes a text and image note generation module 931, a trend video generation module 932, and a large-character poster video generation module 933, used to generate text and image notes, trend videos, and large-character poster videos, respectively. The risk control output layer 940 includes a risk verification module 941, a platform distribution module 942, and an effect feedback module 943. The risk control verification module 941 performs consistency and compliance verification on the content to be published, eliminating AI illusions and non-compliant content. The platform distribution module 942 is responsible for delivering the content to the terminal platform. The effect feedback module 943 determines the viral content and the corresponding themes by monitoring core quality indicators, using them as target themes. The effect feedback module 943 can feed back the target theme to the data acquisition layer 910. After processing by the target acquisition layer 910, intelligent arrangement layer 920, material processing and rendering layer 930, and risk control output layer 940, new content can be generated. For example, similar content in a different form than the viral content can be generated based on the target theme and published to another account, thus replicating the viral content.

[0139] Figure 10 This is a block diagram of a content generation apparatus according to another embodiment of the present disclosure.

[0140] like Figure 10 As shown, the content generation device 1000 includes an acquisition module 1010, a first generation module 1020, a processing module 1030, and a second generation module 1040.

[0141] The acquisition module 1010 is used to acquire the associated data of the target entity, wherein the associated data includes news data related to the target entity and entity attribute data of the target entity obtained from the database.

[0142] The first generation module 1020 is used to generate a content script that conforms to a preset narrative structure based on news data and entity attribute data using a large model. The content script includes content fragments and rendering instructions corresponding to the content fragments.

[0143] The processing module 1030 is used to call preset tools to obtain the corresponding media materials according to the rendering instructions.

[0144] The second generation module 1040 is used to generate multimodal content in multiple output formats according to content fragments and media materials and a preset narrative structure.

[0145] According to an embodiment of this disclosure, the acquisition module 1010 includes a first acquisition submodule, a first determination module, and a second acquisition submodule. The first acquisition submodule is used to acquire a target topic and acquire news data from multiple information sources based on the target topic. The first determination module is used to determine a target entity and news data related to the target entity based on entity words extracted from the news data. The second acquisition submodule is used to acquire entity attribute data of the target entity from the database based on the entity identifier corresponding to the target entity.

[0146] According to an embodiment of this disclosure, the first acquisition submodule includes a filtering module and an expansion module. The filtering module is used to filter out target statements of a preset category from search statements that meet the popularity criteria. The expansion module is used to semantically expand the target statements using a large model to obtain the target topic.

[0147] According to an embodiment of this disclosure, the acquisition module 1010 further includes a third acquisition submodule, which is used to acquire entity attribute data of the target entity from the database and acquire news data of a preset time range associated with the target entity from multiple data sources.

[0148] According to an embodiment of this disclosure, the first generation module 1020 includes a first generation submodule, a rewriting module, and a second determination module. The first generation submodule is used to generate a title and body text based on news data and entity attribute data. The rewriting module is used to rewrite the title and body text into layout data that conforms to a preset narrative structure. The layout data includes multiple content fragments. The second determination module is used to determine the rendering instructions corresponding to the content fragments based on the type of the content fragments.

[0149] According to embodiments of this disclosure, the second determining module includes a first determining submodule and a second determining submodule. The first determining submodule is used to determine the rendering instruction as an image acquisition instruction in response to the fragment content being a start segment or an end segment. The second determining submodule is used to determine the rendering instruction as at least one of an image acquisition instruction, a chart generation instruction, a screenshot instruction, an animation effect instruction, and a speech synthesis instruction in response to the fragment content being a middle segment.

[0150] According to embodiments of this disclosure, the processing module 1030 includes a first processing submodule, a second processing submodule, and a third processing submodule. The first processing submodule is used to respond to the rendering instruction as an image acquisition instruction and call a cropping tool to acquire the main image from the news data. The second processing submodule is used to respond to the rendering instruction as a screenshot instruction and call a screenshot tool to extract attribute information images from the entity attribute data. The third processing submodule responds to the rendering instruction as an animation effect instruction and calls an effects tool to convert the fragment content into a presentation containing animation effects.

[0151] According to embodiments of this disclosure, the content generation apparatus 1000 further includes a historical entity attribute data acquisition module, a third generation module, and a fourth generation module. The historical entity attribute data acquisition module is used to acquire historical entity attribute data from a database; the third generation module is used to generate trend data and explanatory text based on the historical entity attribute data using a large model; and the fourth generation module is used to generate a trend display video based on the trend data and explanatory text.

[0152] According to embodiments of this disclosure, the second generation module 1040 is further configured to combine content fragments and media materials according to a preset narrative structure to generate at least one of video data and graphic data.

[0153] According to embodiments of this disclosure, the content generation device 1000 further includes a first verification module and a second verification module. The first verification module is used to verify the factual consistency of time information and numerical information in the content segment; the second verification module is used to perform quality detection and content compliance verification on the media material.

[0154] According to embodiments of this disclosure, entity attribute data includes at least one of enterprise basic information and financial statement data.

[0155] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0156] Figure 11 A schematic block diagram of an example electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0157] like Figure 11 As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded into random access memory (RAM) 1103 from storage unit 1108. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0158] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0159] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as content generation methods. For example, in some embodiments, the content generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the content generation method described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the content generation method by any other suitable means (e.g., by means of firmware).

[0160] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0164] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0165] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0166] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0167] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A content generation method, comprising: Obtain the associated data of the target entity, wherein the associated data includes news data related to the target entity and entity attribute data of the target entity obtained from the database; Based on the news data and the entity attribute data, a content script conforming to a preset narrative structure is generated using a large model. The content script includes content fragments and rendering instructions corresponding to the content fragments. According to the rendering instructions, the preset tools are invoked to obtain the corresponding media materials; Based on the content fragments and media materials, multimodal content in various output formats is generated according to the preset narrative structure.

2. The method according to claim 1, wherein, The acquisition of the associated data of the target entity includes: Obtain the target topic, and retrieve news data from multiple information sources based on the target topic; Based on the entity words extracted from the news data, the target entity and the news data related to the target entity are determined; Based on the entity identifier corresponding to the target entity, the entity attribute data of the target entity is obtained from the database.

3. The method according to claim 2, wherein, The target topic to be obtained includes: Filter target statements of preset categories from search statements that meet the popularity criteria; The target statement is semantically expanded using a large model to obtain the target topic.

4. The method according to claim 1, wherein, The acquisition of the associated data of the target entity includes: For a target entity, entity attribute data of the target entity is obtained from the database, and news data associated with the target entity within a preset time range is obtained from multiple data sources.

5. The method according to claim 1, wherein, The step of generating a content script conforming to a preset narrative structure using a large model based on the news data and the entity attribute data includes: Based on the news data and the entity attribute data, generate the title and body text; The title and body text are rewritten into layout data that conforms to the preset narrative structure, and the layout data includes multiple content fragments; Based on the type of the content fragment, determine the rendering instructions corresponding to the content fragment.

6. The method according to claim 5, wherein, The step of determining the rendering instruction corresponding to the content fragment based on the type of the content fragment includes: In response to the content segment being either a start segment or an end segment, the rendering instruction is determined to be an image acquisition instruction; In response to the content segment being of type intermediate segment, the rendering instruction is determined to be at least one of image acquisition instruction, chart generation instruction, screenshot instruction, animation effect instruction, and speech synthesis instruction.

7. The method according to claim 6, wherein, The step of calling a preset tool to obtain the corresponding media material according to the rendering instruction includes: In response to the rendering instruction being an image acquisition instruction, a cropping tool is invoked to acquire the main image from the news data; In response to the rendering command being a screenshot command, a screenshot tool is invoked to extract an image of the attribute information from the entity attribute data; In response to the rendering instruction being an animation effect instruction, the effects tool is invoked to convert the fragment content into a presentation containing animation effects.

8. The method according to claim 1, further comprising: Retrieve historical entity attribute data from the database; The large model is used to generate trend data and explanatory text based on the historical entity attribute data; A trend display video is generated based on the trend data and the narration.

9. The method according to claim 1, wherein, The step of generating multimodal content in multiple output formats according to the content fragments and media materials and the preset narrative structure includes: The content fragments and media materials are combined according to the preset narrative structure to generate at least one of video data and text / image data.

10. The method according to claim 1, further comprising: Perform factual consistency verification on the time and numerical information in the content segment; The media materials were subjected to quality testing and content compliance verification.

11. The method according to claim 1, wherein, The entity attribute data includes at least one of the enterprise's basic information and financial statement data.

12. A content generation apparatus, comprising: The acquisition module is used to acquire the associated data of the target entity, wherein the associated data includes news data related to the target entity and entity attribute data of the target entity acquired from the database; The first generation module is used to generate a content script that conforms to a preset narrative structure based on the news data and the entity attribute data using a large model. The content script includes content fragments and rendering instructions corresponding to the content fragments. The processing module is used to call a preset tool to obtain the corresponding media material according to the rendering instructions; The second generation module is used to generate multimodal content in multiple output formats according to the content fragments and media materials and the preset narrative structure.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method according to any one of claims 1 to 11 when executed by a processor.