An AI intelligent manuscript creation method and system based on historical manuscripts

By adopting a genre-based feature collection strategy and building a feature library of historical news articles, the problem of low historical data correlation in the AI ​​writing system was solved, and the depth and timeliness of news content were improved.

CN120561355BActive Publication Date: 2025-09-19苏州日报社
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511044649.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-09-19
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

The existing AI writing system cannot effectively link data from multiple time nodes and lacks historical data mining, resulting in serious homogeneity of news content and a lack of in-depth and timely insights.

Method used

By obtaining the genre attributes of the topic, matching the corresponding element collection strategy, dynamically generating creation instructions, controlling the language model to generate a pre-arranged framework, and using the BERT-News, GPT-StyleCLIP and ResNet-152 models to extract the semantic, style and visual features of historical news articles, a historical news article feature library is constructed to achieve cross-modal retrieval and multi-level association mining.

Benefits of technology

It significantly improves the accuracy and efficiency of factor collection, realizes multi-level correlation mining of historical news data, and enhances the depth and timeliness of news content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561355B_ABST
    Figure CN120561355B_ABST
Patent Text Reader

Abstract

This application discloses an AI intelligent manuscript creation method and system based on historical manuscripts to solve the technical problem of low correlation of historical data. Among them, an AI intelligent manuscript creation solution based on historical manuscripts can significantly improve the accuracy and efficiency of factor collection by matching different factor collection strategies according to different genre attributes. And according to the factor data and genre attributes, a number of historical news articles are matched as reference prompt instructions, realizing multi-level correlation mining of historical news data and greatly improving the correlation depth of historical data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text creation technology, and in particular to an AI intelligent manuscript creation method and system based on historical manuscripts. Background Art

[0002] In the Internet age, the acquisition and transmission of information are becoming increasingly convenient, and people's demand for the timeliness, richness of content and accuracy of news information is growing. The traditional news production line that relies on editors to write full texts can no longer adapt to the current pace of life.

[0003] Currently, language models are being used to assist news article creation. Specifically, by using predefined rules and fill-in templates, they can generate short, concise, and efficient articles. However, the articles currently generated with the assistance of language models are highly homogeneous in content and lack in-depth analysis and flexible expression.

[0004] In the process of implementing the prior art, the inventors found that:

[0005] Existing technologies already have AI writing systems and generative AI tools based on template matching, but these systems or tools generally suffer from the defect of insufficient historical data mining, making it impossible to effectively associate data from multiple time nodes and capture the patterns of cross-cycle events. For example, it is difficult to compare three-year trends in sports event reports, and financial news lacks the ability to dynamically reason about historical economic indicators. In addition, the generated content is limited to shallow semantic retrieval and static knowledge base calls, resulting in a single analysis dimension and logical gaps, which weakens the depth and timeliness of news content.

[0006] Therefore, it is necessary to provide an AI intelligent manuscript creation solution based on historical manuscripts to solve the technical problem of low correlation of historical data. Summary of the Invention

[0007] The embodiments of the present application provide an AI intelligent manuscript creation method and system based on historical manuscripts, which are used to solve the technical problem of low correlation of historical data.

[0008] Specifically, an AI intelligent manuscript creation method based on historical manuscripts includes the following steps:

[0009] Get the genre attributes of the topic;

[0010] According to the genre attributes of the topic, match the element collection strategy corresponding to the genre attributes;

[0011] Obtain updated feature data based on feature collection strategies;

[0012] Dynamically generate creative instructions based on genre attributes and element data;

[0013] According to the creative instructions, the language model is controlled to generate a pre-arranged framework consisting of several outlines;

[0014] Obtain a first draft framework based on pre-edited framework updates;

[0015] According to the draft framework, the language model is controlled to generate a draft text consisting of several paragraphs;

[0016] The genre attribute of the topic includes at least one of news genre, correspondence genre, in-depth report genre or commentary genre;

[0017] Based on the genre attributes of the topic, the element collection strategies that match the corresponding genre attributes include:

[0018] The genre attribute of the selected topic is selected as news genre or communication genre, matching the real-time factor data collection strategy;

[0019] The genre attribute of the selected topic is in-depth report genre or commentary genre, matching the historical element data retrieval strategy;

[0020] Based on genre attributes and feature data, creative instructions are dynamically generated, including:

[0021] According to the element data and genre attributes, several historical news articles are matched as reference prompt instructions;

[0022] Generate creation instructions including element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and reference prompt instructions;

[0023] When no historical news item is matched, a creation instruction is generated including creation instructions of element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, and workflow instructions.

[0024] Furthermore, the method further comprises:

[0025] Based on the first draft text, perform semantic vector conversion to determine the core events, emotional tendencies, and scene keywords as text features;

[0026] According to the draft text or text features, the image matching model is controlled to match images corresponding to the draft text or paragraph text as illustrations.

[0027] Furthermore, the method further comprises:

[0028] Obtain historical news release data;

[0029] Use the BERT-News model to extract semantic vectors from historical news article data;

[0030] Use the GPT-StyleCLIP extraction model to extract style vectors from historical news release data;

[0031] Analyze the hierarchical relationship between the title, introduction, and body of historical news articles data as the structural vector of the historical news articles data;

[0032] Use the ResNet-152 model to extract visual features of historical news release data;

[0033] The semantic vectors of the historical news article data, the style vectors of the historical news article data, the structural vectors of the historical news article data, and the visual features of the historical news article data are stored as the historical news article mixed coding objects;

[0034] Several mixed coding objects of historical news articles constitute the historical news article feature library.

[0035] Furthermore, the matching of several historical news articles based on element data and genre attributes specifically includes:

[0036] Filter the first similar historical news article collection based on genre attributes:

[0037] Update cross-modal search formulas based on feature data;

[0038] Using the updated cross-modal search formula, the second most similar historical news article set is filtered from the historical news article feature library:

[0039] Based on the manuscript quality evaluation coefficient, a preset number of historical news articles are selected from the first similar historical news article set or the second similar historical news article set as a reference manuscript set;

[0040] Among them, the cross-modal search expression is as follows:

[0041] Similarity = α·BERT(text_embed) + β·CLIP(img_embed) + γ·TimeDecay(t)

[0042] Where BERT(text_embed) represents the text semantic features extracted by the BERT model, α represents the semantic feature similarity requirement; CLIP(img_embed) represents the cross-modal feature alignment of image and text by the CLIP model, β represents the correlation requirement between image and text, and γ·TimeDecay(t) represents the time decay factor.

[0043] Furthermore, according to the creative instructions, the language model is controlled to generate a pre-arranged framework consisting of several outlines, specifically including:

[0044] Match factual content in several historical news articles based on factor data;

[0045] Based on the sentence features, paragraph structure, and term usage frequency of several historical news articles, a pre-arranged framework of similar style is generated.

[0046] Furthermore, the method further comprises:

[0047] Obtain selected text based on the draft text, as well as update instructions or update materials corresponding to the selected text;

[0048] According to the update instructions or update materials corresponding to the selected text, text is generated through the language model, and the generated text is used as the replacement text;

[0049] Wherein, the update instruction includes at least one of a polish instruction, a modification instruction, an expansion instruction, and a rewrite instruction;

[0050] The update material includes at least one of field, style, emotion, timeliness, and word count.

[0051] The embodiment of the present application also provides an AI intelligent manuscript creation system based on historical manuscripts.

[0052] Specifically, an AI intelligent manuscript creation system based on historical manuscripts includes:

[0053] Topic planning module, used to obtain the genre attributes of the topic;

[0054] The material retrieval module is used to match the feature collection strategy corresponding to the genre attribute of the selected topic; it is also used to obtain the feature data updated based on the feature collection strategy;

[0055] A framework generation module is used to dynamically generate creative instructions based on genre attributes and element data; it is also used to control the language model to generate a pre-arranged framework consisting of several outlines based on the creative instructions;

[0056] A content filling module is used to obtain a draft framework updated based on the pre-arranged framework; and is also used to control the language model to generate a draft text consisting of several paragraphs based on the draft framework;

[0057] The genre attribute of the topic includes at least one of news genre, correspondence genre, in-depth report genre or commentary genre;

[0058] The material retrieval module matches the element collection strategy corresponding to the genre attribute according to the genre attribute of the selected topic, specifically including:

[0059] The genre attribute of the selected topic is selected as news genre or communication genre, matching the real-time factor data collection strategy;

[0060] The genre attribute of the selected topic is in-depth report genre or commentary genre, matching the historical element data retrieval strategy;

[0061] The framework generation module dynamically generates creation instructions based on genre attributes and element data, specifically including:

[0062] According to the element data and genre attributes, several historical news articles are matched as reference prompt instructions;

[0063] Generate creation instructions including element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and reference prompt instructions;

[0064] When no historical news item is matched, a creation instruction is generated including creation instructions of element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, and workflow instructions.

[0065] Furthermore, the system also includes an image matching module for performing semantic vector conversion based on the draft text to determine core events, emotional tendencies, and scene keywords as text features;

[0066] It is also used to control the image matching model to match images corresponding to the draft text or paragraph text as illustrations based on the draft text or text features.

[0067] Furthermore, the system also includes a historical news article feature library;

[0068] The steps of constructing the historical news article feature database include:

[0069] Obtain historical news release data;

[0070] Use the BERT-News model to extract semantic vectors from historical news article data;

[0071] Use the GPT-StyleCLIP extraction model to extract style vectors from historical news release data;

[0072] Analyze the hierarchical relationship between the title, introduction, and body of historical news articles data as the structural vector of the historical news articles data;

[0073] Use the ResNet-152 model to extract visual features of historical news release data;

[0074] The semantic vectors of the historical news article data, the style vectors of the historical news article data, the structural vectors of the historical news article data, and the visual features of the historical news article data are stored as the historical news article mixed coding objects;

[0075] Several mixed coding objects of historical news articles constitute the historical news article feature library.

[0076] Furthermore, the framework generation module matches several historical news articles based on the element data and genre attributes, including:

[0077] Filter the first similar historical news article collection based on genre attributes:

[0078] Update cross-modal search formulas based on feature data;

[0079] Using the updated cross-modal search formula, the second most similar historical news article set is filtered from the historical news article feature library:

[0080] Based on the manuscript quality evaluation coefficient, a preset number of historical news articles are selected from the first similar historical news article set or the second similar historical news article set as a reference manuscript set;

[0081] Among them, the cross-modal search expression is as follows:

[0082] Similarity = α·BERT(text_embed) + β·CLIP(img_embed) + γ·TimeDecay(t)

[0083] Where BERT(text_embed) represents the text semantic features extracted by the BERT model, α represents the semantic feature similarity requirement; CLIP(img_embed) represents the cross-modal feature alignment of image and text by the CLIP model, β represents the correlation requirement between image and text, and γ·TimeDecay(t) represents the time decay factor.

[0084] The technical solutions provided in the embodiments of the present application have at least the following beneficial effects:

[0085] By matching different feature collection strategies based on different genre attributes, the accuracy and efficiency of feature collection can be significantly improved. Furthermore, by matching several historical news articles based on feature data and genre attributes as reference prompts, multi-level association mining of historical news data is achieved, significantly increasing the depth of historical data associations. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0087] Figure 1A flowchart of an AI intelligent manuscript creation method based on historical manuscripts provided in an embodiment of the present application;

[0088] Figure 2 A flowchart of an AI intelligent manuscript creation method based on historical manuscripts provided in an embodiment of the present application;

[0089] Figure 3 A structural diagram of an AI intelligent manuscript creation system based on historical manuscripts provided in an embodiment of the present application.

[0090] Description of reference numerals:

[0091] 100-AI intelligent manuscript creation system based on historical manuscripts; 11-Topic planning module; 12-Material retrieval module; 13-Framework generation module; 14-Content filling module; 15-Style calibration module; 16-Illustration module. DETAILED DESCRIPTION

[0092] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0093] Please refer to Figure 1 and Figure 2 In order to solve the technical problem of low correlation of historical data, this application provides an AI intelligent manuscript creation method based on historical manuscripts, including the following steps:

[0094] S110: Obtain the genre attribute of the selected topic.

[0095] S120: According to the genre attributes of the selected topic, an element collection strategy corresponding to the genre attributes is matched.

[0096] S130: Acquire the updated feature data based on the feature collection strategy.

[0097] It is understandable that this application is based on the actual news production workflow and adopts node-based editing technology to break down news creation into five core nodes: topic planning, material retrieval, framework generation, content filling, and style calibration. Step S110 corresponds to the topic planning node, and step S120 corresponds to the material retrieval node.

[0098] In specific application scenarios, the genre attributes of selected topics include at least one of news, correspondence, feature, in-depth report, or commentary. Different genre attributes also have different requirements for element collection. Specifically, news or correspondence (breaking news) genres require higher real-time performance, while in-depth report or commentary genres require in-depth content mining and cross-period relevance.

[0099] Therefore, in a specific embodiment provided by the present application, step S120 includes, based on the genre attributes of the selected topic, matching the element collection strategy corresponding to the genre attributes:

[0100] The genre attribute of the selected topic is selected as news genre or communication genre, matching the real-time factor data collection strategy;

[0101] The genre attribute of the selected topic is in-depth reporting genre or commentary genre, matching the historical element data retrieval strategy.

[0102] The feature collection strategy here can be expressed in specific application scenarios as a real-time feature data collection strategy (such as web crawler information capture tools such as spider Tool), a historical feature data retrieval strategy (such as historical data retrieval tools such as searcher tool), or a collection of such material collection tools.

[0103] Specifically, for news or correspondence genres with high real-time requirements, real-time data collection tools (spider tools) are preferred for collecting real-time news materials. For in-depth reporting or commentary genres that require deep exploration and cross-period correlations, historical data search tools (searcher tools) are preferred, with historical news release data as the primary source of material.

[0104] Of course, the element collection strategy here also includes preset element guidance templates, which, in the form of templates or questionnaires, guide users to enter specific element data. Specifically, this application utilizes dynamic form technology to automatically load the corresponding element guidance template based on the selected genre. For example, the news element guidance template will guide users to verify the completeness of the 5W1H (Who, What, When, Where, Why, How) matrix, while the commentary element guidance template will guide users to construct a point of view and position tree. Based on the genre attributes of the topic selected by the user in step S110, the user is guided to fill in the corresponding element guidance template. Different news genres have corresponding specific element templates developed by journalism experts (for example: the news element guidance template will guide users to complete the entry of element data such as the 5W1H (Who, What, When, Where, Why, How); the news commentary element guidance template will guide users to complete the entry of element data such as the theme news event or social issue, the author's point of view and position; the investigative report element guidance template will guide users to complete the entry of different element data such as the theme, relevant evidence and facts).

[0105] In step S120, the genre attributes of the previously selected topic, the subject of the manuscript, and the genre element information selected by the user and submitted by the user will be combined with a pre-established standard prompt word template to form a large model data collection tool, and the scenario prompt words of the calling task will be submitted to the general large model. The model will automatically select the specific tool to execute the material collection link based on the provided materials, perform tool calling (function calling), execute material collection, and return the results.

[0106] Here, the element data input by the user according to the element guidance template and the element data collected by the real-time element data collection strategy or the historical element data retrieval strategy are used as the element data updated based on the element collection strategy in step S130.

[0107] S140: Dynamically generate creation instructions based on genre attributes and element data.

[0108] S150: According to the creation instruction, the language model is controlled to generate a pre-arranged framework consisting of a plurality of outlines.

[0109] After acquiring updated feature data based on the feature collection strategy, creative instructions are dynamically generated based on genre attributes and feature data. These creative instructions can be understood as sets of prompt word instructions that serve as input to the language model to improve the quality of the control language model output.

[0110] This application utilizes an expert panel comprised of senior editors and journalists from the newspaper to develop pre-defined creative instructions for different news genres, focusing on the characteristics and styles of different news genres, such as news, correspondence, interviews, and commentary. The panel will then identify common writing priorities and styles, along with key writing points and outline templates for each genre. For example, news stories focus on clearly and concisely conveying news facts; correspondence stories can delve deeper into the details of an event; features focus on vividly depicting a scene or character; and commentary stories require analysis and evaluation of the news event.

[0111] The pre-set creative instructions include at least three main outlines: the introduction, the main body, and the conclusion. Different genres may require different writing elements for the main body, and these will be discussed by the expert group to determine a universal and feasible plan.

[0112] for example:

[0113] The genre is news, and the outline elements include title, introduction, body, background, and conclusion:

[0114] 1. Title: concise and clear, summarizing the core content of the news;

[0115] 2. Introduction: briefly introduce the news event to attract readers' attention;

[0116] 3. Body: Detail the events and provide key details;

[0117] 4. Background: Supplementary background information related to the incident;

[0118] 5. Conclusion: Summarize the events and propose follow-up prospects.

[0119] The genre is review. The outline elements include title, introduction, body, and conclusion:

[0120] 1. Title: Clearly express your views and position.

[0121] 2. Introduction: Introduce the subject of review and the viewpoint.

[0122] 3. Main body of the paper: Support your viewpoint through arguments, citing examples and data.

[0123] 4. Conclusion: Summarize views and make suggestions or prospects.

[0124] In addition to the genre elements and outline summaries developed by experts in different genres, the expert group also uniformly developed a large-scale model outline and summary writing prompt template as a creative instruction for the writing of outlines and summaries of different genres.

[0125] After the creation instructions are input into the language model, the language model will output a pre-arranged framework consisting of several outlines based on the creation instructions.

[0126] It's understandable that the language model here is a generative pre-trained model. This type of model learns linguistic patterns from massive amounts of text and can automatically complete or generate coherent, diverse text based on input prompts. With creative commands, users can fine-tune the output style and content by adjusting template commands.

[0127] In a specific embodiment provided in this application, based on genre attributes and element data, dynamically generating creation instructions includes:

[0128] According to the element data and genre attributes, several historical news articles are matched as reference prompt instructions;

[0129] Generate creation instructions including element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and reference prompt instructions;

[0130] When no historical news item is matched, a creation instruction is generated including creation instructions of element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, and workflow instructions.

[0131] Furthermore, the matching of several historical news articles based on element data and genre attributes specifically includes:

[0132] Filter the first similar historical news article collection based on genre attributes:

[0133] Update cross-modal search formulas based on feature data;

[0134] Using the updated cross-modal search formula, the second most similar historical news article set is filtered from the historical news article feature library:

[0135] Based on the manuscript quality evaluation coefficient, a preset number of historical news articles are selected from the first similar historical news article set or the second similar historical news article set as a reference manuscript set;

[0136] Among them, the cross-modal search expression is as follows:

[0137] Similarity = α·BERT(text_embed) + β·CLIP(img_embed) + γ·TimeDecay(t)

[0138] Where BERT(text_embed) represents the text semantic features extracted by the BERT model, α represents the semantic feature similarity requirement; CLIP(img_embed) represents the cross-modal feature alignment of image and text by the CLIP model, β represents the correlation requirement between image and text, and γ·TimeDecay(t) represents the time decay factor.

[0139] For example, this application can retrieve the top 10-20 excellent historical articles related to the topic semantics from the historical news article library based on the different news genres selected by the user and the specific elements for different genres filled in by the user through genre, elements and main points.

[0140] Of course, in addition to extracting the top 10 to 20 historical news articles by relevance as small sample drafts and submitting them to the language model, users can also select an article from the retrieved historical excellent manuscripts that best meets the current writing scenario or style requirements as a special reference sample and provide it to the language model.

[0141] Here, according to the preset creative instructions corresponding to different genres, the retrieved excellent historical manuscripts are input into a specific language model as small sample instances to carry out AIGC (Artificial Intelligence Generated Content) manuscript writing.

[0142] The following takes the creation instructions for generating the corresponding message genre as an example.

[0143] The specific content is news, and the genre elements are as follows:

[0144] The subject of the message is "Environmental Conference is held", and the specific materials include: time (when): XX / XX / 2024, location (where): conference center, people (who): environmental heads of various countries, what (what): signing an environmental agreement, why (why): responding to climate change.

[0145] The specific message creation instructions are as follows:

[0146] - Role: News writing expert and experienced journalist

[0147] - Background (Scene prompt): Users are required to write a news outline for the topic "{{Environmental Conference Held}}." They provide specific material, including the time, location, people involved, key events, and background information. {{Time (when): X month XX day, 2024; Location (where): Conference center; Who (who): Environmental officials from various countries; What (what): Signing of an environmental agreement; Why (why): Addressing climate change.}}

[0148] These materials are the basis for writing news-type news. They need to be integrated into the title, lead, body, background, conclusion and other parts in accordance with the standards of news writing.

[0149] - Profile (Role prompts): You are an experienced news writing expert who has been engaged in news writing for a long time. You have a deep understanding of the structure and elements of news and rich practical experience. You can quickly transform materials into outlines that meet news writing standards.

[0150] - Skills: You are proficient in the basic principles of news writing, including how to extract core facts, how to construct news structures, and how to express complex information in concise and clear language. You also have excellent logical thinking and information integration skills.

[0151] - Goals (goal prompt instructions): Based on the message theme "Environmental Conference Held" and specific materials provided by the user, write a clear and complete news writing outline. Ensure that the outline covers elements such as the title, introduction, main body, background, and conclusion, and reasonably integrate specific materials such as when, where, who, what, and why.

[0152] - Constraints: The outline should follow the writing standards of news reports, use concise and clear language, avoid lengthy and complex sentence structures, and ensure the accuracy and objectivity of the information.

[0153] - OutputFormat (output rule instructions): The outline should include a title, introduction, body, background, conclusion and other parts. Each part should have a concise description of the key points or a brief content of each part should be written based on the provided elements.

[0154] - Workflow (workflow instructions):

[0155] 1. Based on the message subject "{{Environmental Conference Held}}" provided by the user, determine the core facts and focus of the news.

[0156] 2. Integrate specific materials such as when, where, who, what, and why to build the basic framework of the news.

[0157] 3. Follow the structural requirements of news reports and write the title, introduction, main body, background, conclusion and other parts, ensuring that each part is logically coherent and the information is complete.

[0158] - Examples (reference prompt instructions):

[0159] - Example 1: {{The content of the relevant news article 1 in the historical article library through semantic retrieval}}

[0160] - Example 2: {{The content of Good Draft 2 of the relevant news in the historical manuscript database through semantic retrieval}}

[0161] - Example n: ...

[0162] It's important to note that when dynamically generating creative instructions, the template will search for semantically related news articles of the same genre in the historical library based on other parameters such as genre type, and place them in the "Examples" section. If no suitable reference articles are found, the "Examples" section will not appear in the template.

[0163] When using a specific system prompt word template developed by an expert group for the same genre, you only need to replace the content about the theme and genre elements in the template. In the Examples section, use the semantics of the theme and element content to retrieve excellent A drafts with similar or related message semantics and scenarios from the historical database as small sample reference examples. As creative instructions, submit them to the designated language model to write the outline of the message manuscript and the summary content of each part.

[0164] Among them, the news genre writing standards and elements formulated by the expert group must be scientific and standardized, and will be continuously optimized and iterated in actual applications, and adjusted and modified after a lot of practical applications. In addition, this application will also retrieve relevant reference examples from the historical data manuscript library based on semantics. After data governance, this application will classify and store the historical manuscript library according to news genres, and perform semantic vectorization and embedding storage on the text to improve the accuracy of semantic retrieval; it will also classify and label historical news articles of different genres, such as labeling multiple weights of excellent, representative and paradigm reference value with different coefficients. In actual retrieval and recall, in addition to semantic sorting, other parameters are dynamically adjusted for different application scenarios, so that the historical manuscripts that are most suitable as reference samples can be recalled as important reference samples for large-scale model-assisted writing.

[0165] The following are the steps for building a historical news release feature database:

[0166] Obtain historical news release data;

[0167] Use the BERT-News model to extract semantic vectors from historical news article data;

[0168] Use the GPT-StyleCLIP extraction model to extract style vectors from historical news release data;

[0169] Analyze the hierarchical relationship between the title, introduction, and body of historical news articles data as the structural vector of the historical news articles data;

[0170] Use the ResNet-152 model to extract visual features of historical news release data;

[0171] The semantic vectors of the historical news article data, the style vectors of the historical news article data, the structural vectors of the historical news article data, and the visual features of the historical news article data are stored as the historical news article mixed coding objects;

[0172] Several mixed coding objects of historical news articles constitute the historical news article feature library.

[0173] Specifically, this application constructs a five-level classification tag library based on the news genre classification system for historical news article data:

[0174] Basic genres: News / Correspondence / Feature / In-depth Report / Commentary

[0175] Field tags: Current Affairs / Finance / Sports / Technology

[0176] Style tags: Data-intensive / storytelling / investigative

[0177] Emotional tags: neutral reporting / critical / advocacy

[0178] Timeliness tags: News / Follow-up reports / Yearbook summary

[0179] In addition, hybrid coding technology is used to decompose historical manuscripts into semantic vectors (using the BERT-News model), style vectors (extracted based on GPT-StyleCLIP), and structural vectors (parsing the hierarchical relationship between title-lead-text).

[0180] During the draft generation stage, the top 50 high-click-rate articles of the same genre can be retrieved from the knowledge base based on the user-specified genre tag (such as "financial in-depth report") as a style reference.

[0181] In a specific embodiment provided in this application, according to the creation instruction, the language model is controlled to generate a pre-arranged framework consisting of several outlines, specifically including:

[0182] Match factual content in several historical news articles based on factor data;

[0183] Based on the sentence features, paragraph structure, and term usage frequency of several historical news articles, a pre-arranged framework of similar style is generated.

[0184] Specifically, the language model extracts the writing elements of a specific genre input by the user and similar genre samples retrieved from historical news materials according to the pre-arranged workflow, and performs pre-arranged workflow writing creation.

[0185] In this step, the present application develops a dual-channel generator based on the retrieval-enhanced generation RAG architecture, including a content generation channel and a style simulation channel.

[0186] The content generation channel leverages retrieved historical manuscripts to provide factual support (e.g., comparisons of data from previous events). The style simulation channel extracts sentence features (average sentence length, distribution of rhetorical devices), paragraph structure (inverted pyramid / hourglass), and terminology frequency from reference manuscripts, and fine-tunes the generation model through comparative learning.

[0187] S160: Obtaining a draft framework updated based on the pre-arranged framework.

[0188] S170: Based on the draft framework, control the language model to generate a draft text consisting of several paragraphs.

[0189] It should be noted that the "preliminary draft framework" updated based on the pre-arranged framework here refers to the "preliminary draft framework" article outline generated by the user after adding, modifying, deleting, or adjusting the order of the pre-arranged framework (i.e., the article outline) generated in step S160. Subsequently, based on the element data in the preliminary draft framework, the language model can expand each article outline to generate a preliminary draft text consisting of several paragraphs.

[0190] To improve user experience, in a specific implementation provided in this application, the method further includes:

[0191] Obtain selected text based on the draft text, as well as update instructions or update materials corresponding to the selected text;

[0192] According to the update instructions or update materials corresponding to the selected text, text is generated through the language model, and the generated text is used as the replacement text;

[0193] Wherein, the update instruction includes at least one of a polish instruction, a modification instruction, an expansion instruction, and a rewrite instruction;

[0194] The update material includes at least one of field, style, emotion, timeliness, and word count.

[0195] It is understandable that the present application may expand each paragraph according to the writing outline. When the expansion is completed, the present application may also polish and revise the entire text paragraph by paragraph.

[0196] Specifically, the expansion or polishing here refers to paragraph-level expansion based on the news article writing outline and the brief content of each part in the first draft text, after the user adds, deletes, or adjusts the outline and related parts. After the user selects the paragraph to be expanded with the cursor (it can be the entire text or a partial paragraph), a menu will pop up. After the user selects the "Expand Instructions" function, a dialog box will pop up. The user can fill in specific "expand instructions" or prompt words in this dialog box, or add supplementary materials to the prompt words to supplement the content. The language model will then combine the expansion instructions and supplementary materials filled in by the user in the dialog box based on the previous news genre, writing theme and genre elements, as well as the writing outline and summary content generated in the previous step, to synthesize the expansion prompt word template content and submit it to the general writing language model for content expansion of the specific paragraph.

[0197] The specific instruction requirements of the expansion instruction, such as the number of words, style, emotion, etc., are filled in by the user in this pop-up dialog box. In addition, this application will also provide some reference samples of expansion prompt words for different genres and paragraphs for users to refer to and adopt.

[0198] The operating logic of the polish instruction, modify instruction, and rewrite instruction is the same as that of the expansion instruction, and will not be repeated here.

[0199] In addition, this application also includes a review and proofreading of the first draft text. Specifically, the language model will check and filter the generated text for sensitive words based on a self-built sensitive word library and provide revision suggestions. The language model will then further check sentences and paragraphs for typos and commonly used standard terms, and provide revision suggestions.

[0200] The language model can also check the overall text coherence and text semantic logic of the draft text, propose modification suggestions, and the user will check and confirm the modifications and output the updated draft text.

[0201] Based on the above-mentioned updated draft text, this application allows the user to further perform segmented adjustments, polishing, modification, expansion, rewriting, and other operations on the updated draft text, and then finalize it.

[0202] Furthermore, in a preferred embodiment provided in the present application, the method further comprises:

[0203] Based on the first draft text, perform semantic vector conversion to determine the core events, emotional tendencies, and scene keywords as text features;

[0204] According to the draft text or text features, the image matching model is controlled to match images corresponding to the draft text or paragraph text as illustrations.

[0205] Specifically, if the user does not need to add pictures to the draft text, the draft text can be directly output as the final draft. If the user needs to add pictures to the draft text, the first draft text intelligent picture matching stage will be entered.

[0206] If you need to intelligently match pictures for the generated article, the user can select the picture matching mode. You can use global matching, that is, matching pictures for the entire article, or you can choose to match local semantic intelligent pictures for specific paragraphs in the article.

[0207] In the full-text semantic image matching process, this application will intelligently retrieve matching images from private image libraries based on the full-text genre, style, or narrative events.

[0208] In the local semantic image matching process, this application will only target the semantics and scene factors of a specific paragraph and intelligently retrieve matching images from a private image library.

[0209] Specifically, when a specific entity (such as "garden") is detected, garden-related images are directly retrieved from the local image library.

[0210] If no suitable picture is found, or the style, emotion, or composition does not meet the requirements, users can also use Wenshengtu to generate pictures intelligently by AI for illustration.

[0211] This application uses a multimodal vector model to vectorize and embed the text and images of historical news articles, encode the content of historical news articles into 512-dimensional semantic vectors, and focus on capturing core events (60% weight), emotional tendencies (20% weight), and scene keywords (20% weight).

[0212] This application also constructs an image feature library, and the construction process is as follows:

[0213] Perform deep feature extraction on private image libraries, using ResNet-152 to extract visual features and the CLIP model to extract image-text matching features;

[0214] A three-level index is established: subject object (face / building / natural scenery) - scene category (conference site / disaster site) - emotional tone (serious / cheerful / sad).

[0215] When a specific entity (such as "airport") is detected, the entity association library is directly called.

[0216] Abstract concepts can also be mapped into visual elements through cross-modal alignment technology (such as "garden visitor volume" matching the upward arrow chart + the crowd ticket purchasing scene diagram).

[0217] In addition, the illustration style can be automatically selected based on the article genre (data charts are preferred for news articles, and high-resolution scene images are used for feature articles).

[0218] It is also possible to integrate review nodes to perform copyright verification on candidate images (matching known copyright libraries), filter sensitive content (using self-built sensitive words to filter sensitive content), and adapt resolution (automatically crop according to the publishing platform).

[0219] Finally, in the image-text matching test phase, the model sorts the images by relevance based on the full text, or the semantic retrieval of local paragraphs / the AI-generated image list. Ultimately, the user determines and selects the most suitable image to insert into the corresponding position of the article, completing the intelligent image matching and outputting the finalized article with both text and images.

[0220] In summary, this application significantly improves the accuracy and efficiency of feature collection by matching different feature collection strategies based on different genre attributes. Furthermore, by matching several historical news articles based on feature data and genre attributes as reference prompts, it achieves multi-level association mining of historical news data, significantly increasing the depth of historical data associations.

[0221] Please refer to Figure 3 In order to support the AI ​​intelligent manuscript creation method based on historical manuscripts, this application also provides an AI intelligent manuscript creation system 100 based on historical manuscripts.

[0222] The article creation system 100 has a three-tiered process architecture, comprising a task decomposition layer, a model scheduling layer, and a dynamic orchestration layer. The task decomposition layer utilizes node-based orchestration technology to break down news creation into five core nodes: topic planning, source material retrieval, framework generation, content filling, and style calibration. Each node deploys a dedicated fine-tuning model (e.g., the topic planning node uses DeepSeek to analyze hot trends, while the style calibration node deploys the BERT text style transfer model). A dynamic path planning mechanism is introduced to automatically adjust the node order based on event type (e.g., breaking news prioritizes real-time data collection nodes, while in-depth reports activate historical data analysis nodes).

[0223] The model scheduling layer adopts a hybrid model invocation strategy, with the main process driven by a general-purpose large model and key links embedded in domain-specific models. After selecting a specific news genre, the route is routed to a genre-specific manuscript assistant, which completes manuscript generation based on the characteristics and style of each genre.

[0224] The dynamic orchestration layer, based on batch node technology, supports the parallel generation of multiple genres (e.g., simultaneous generation of news flashes and in-depth features). It also employs a reinforcement learning mechanism to optimize node parameter configurations through feedback from historical article generation results (e.g., adjusting the weight of descriptive nodes in feature articles).

[0225] Specifically, an AI intelligent manuscript creation system 100 based on historical manuscripts includes:

[0226] The topic planning module 11 is used to obtain the genre attributes of the topic;

[0227] The material retrieval module 12 is used to match the element collection strategy corresponding to the genre attribute of the selected topic; and is also used to obtain the element data updated based on the element collection strategy;

[0228] The framework generation module 13 is used to dynamically generate creative instructions based on genre attributes and element data; and is also used to control the language model to generate a pre-arranged framework consisting of a plurality of outlines according to the creative instructions;

[0229] The content filling module 14 is used to obtain a draft framework updated based on the pre-arranged framework; and is also used to control the language model to generate a draft text consisting of several paragraphs based on the draft framework;

[0230] The genre attribute of the topic includes at least one of news genre, correspondence genre, in-depth report genre or commentary genre;

[0231] The material retrieval module 12 matches the element collection strategy corresponding to the genre attribute according to the genre attribute of the selected topic, specifically including:

[0232] The genre attribute of the selected topic is selected as news genre or communication genre, matching the real-time factor data collection strategy;

[0233] The genre attribute of the selected topic is in-depth report genre or commentary genre, matching the historical element data retrieval strategy;

[0234] The framework generation module 13 dynamically generates creation instructions based on genre attributes and element data, specifically including:

[0235] According to the element data and genre attributes, several historical news articles are matched as reference prompt instructions;

[0236] Generate creation instructions including element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and reference prompt instructions;

[0237] When no historical news item is matched, a creation instruction is generated including creation instructions of element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, and workflow instructions.

[0238] It is understandable that this application is based on the actual news production workflow and adopts node-based editing technology to break down news creation into five core nodes: topic planning, material retrieval, framework generation, content filling, and style calibration. The topic planning module 11 corresponds to the topic planning node, the material retrieval module 12 corresponds to the material retrieval node, the framework generation module 13 corresponds to the framework generation node, and the content filling module 14 corresponds to the content filling node.

[0239] In specific application scenarios, the genre attributes of the topics acquired by the topic planning module 11 include at least one of news, correspondence, feature, in-depth report, or commentary. Different genre attributes also have different requirements for element collection. Specifically, news or correspondence (breaking news) genres have higher requirements for real-time performance, while in-depth report or commentary genres require in-depth content mining and cross-period relevance.

[0240] Therefore, in a specific embodiment provided by the present application, the material retrieval module 12 matches the element collection strategy corresponding to the genre attribute according to the genre attribute of the selected topic, including:

[0241] The genre attribute of the selected topic is selected as news genre or communication genre, matching the real-time factor data collection strategy;

[0242] The genre attribute of the selected topic is in-depth reporting genre or commentary genre, matching the historical element data retrieval strategy.

[0243] The feature collection strategy here can be expressed in specific application scenarios as a real-time feature data collection strategy (such as web crawler information capture tools such as spider Tool), a historical feature data retrieval strategy (such as historical data retrieval tools such as searcher tool), or a collection of such material collection tools.

[0244] Specifically, for news or correspondence genres with high real-time requirements, real-time data collection tools (spider tools) are preferred for collecting real-time news materials. For in-depth reporting or commentary genres that require deep exploration and cross-period correlations, historical data search tools (searcher tools) are preferred, with historical news release data as the primary source of material.

[0245] Of course, the element collection strategy here also includes preset element guidance templates, which, in the form of templates or questionnaires, guide users to enter specific element data. Specifically, the material retrieval module 12 utilizes dynamic form technology to automatically load the corresponding element guidance template based on the selected genre. For example, the news element guidance template guides users to verify the completeness of the 5W1H (Who, What, When, Where, Why, How) matrix, while the commentary element guidance template guides users to construct a point of view and position tree. The material retrieval module 12 guides users to fill out the corresponding element guidance template based on the genre attributes of the user-selected topic. Different news genres have corresponding specific element templates developed by journalism experts (for example, the news element guidance template guides users to complete the entry of 5W1H (Who, What, When, Where, Why, How) element data; the news commentary element guidance template guides users to complete the entry of element data such as the topic news event or social issue, the author's point of view and position; and the investigative report element guidance template guides users to complete the entry of different element data such as the theme and the presentation of relevant evidence and facts).

[0246] The material retrieval module 12 will use the genre attributes of the previous topic, the theme of the manuscript, and the user's selected news genre and submitted genre element information to form a large model data collection tool using a pre-established standardized prompt word template, and submit the scenario prompt words of the calling task to the general large model. The model will automatically select the specific tool to execute the material collection link based on the provided materials, perform tool calling (function calling), execute material collection, and return the results.

[0247] Here, the material retrieval module 12 uses the element data input by the user according to the element guidance template and the element data collected by the real-time element data collection strategy or the historical element data retrieval strategy as the element data updated based on the element collection strategy.

[0248] The framework generation module 13 then dynamically generates creative instructions based on genre attributes and element data, and controls the language model to generate a pre-arranged framework consisting of several outlines according to the creative instructions.

[0249] After acquiring the updated element data based on the element collection strategy, the framework generation module 13 will dynamically generate creation instructions based on the genre attributes and element data. The creation instructions here can be understood as a set of multiple prompt word instructions used as input to the language model to improve the output quality of the control language model.

[0250] This application utilizes an expert panel comprised of senior editors and journalists from the newspaper to develop pre-defined creative instructions for different news genres, focusing on the characteristics and styles of different news genres, such as news, correspondence, interviews, and commentary. The panel will then identify common writing priorities and styles, along with key writing points and outline templates for each genre. For example, news stories focus on clearly and concisely conveying news facts; correspondence stories can delve deeper into the details of an event; features focus on vividly depicting a scene or character; and commentary stories require analysis and evaluation of the news event.

[0251] The pre-set creative instructions include at least three main outlines: the introduction, the main body, and the conclusion. Different genres may require different writing elements for the main body, and these will be discussed by the expert group to determine a universal and feasible plan.

[0252] In addition to the genre elements and outline summaries developed by experts in different genres, the expert group also uniformly developed a large-scale model outline and summary writing prompt template as a creative instruction for the writing of outlines and summaries of different genres.

[0253] After the creation instruction is input into the language model, the framework generation module 13 will output a pre-arranged framework consisting of several outlines according to the creation instruction.

[0254] It is understood that the framework generation module 13 here uses a language model to generate the pre-arranged framework. A language model is a generative pre-trained model. This type of model learns linguistic patterns from massive amounts of text and can automatically complete or generate coherent and diverse text based on input prompts. With creative instructions, users can fine-tune the output style and content by adjusting the template instructions.

[0255] In a specific embodiment provided in this application, based on genre attributes and element data, dynamically generating creation instructions includes:

[0256] According to the element data and genre attributes, several historical news articles are matched as reference prompt instructions;

[0257] Generate creation instructions including element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and reference prompt instructions;

[0258] When no historical news item is matched, a creation instruction is generated including creation instructions of element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, and workflow instructions.

[0259] Furthermore, the matching of several historical news articles based on element data and genre attributes specifically includes:

[0260] Filter the first similar historical news article collection based on genre attributes:

[0261] Update cross-modal search formulas based on feature data;

[0262] Using the updated cross-modal search formula, the second most similar historical news article set is filtered from the historical news article feature library:

[0263] Based on the manuscript quality evaluation coefficient, a preset number of historical news articles are selected from the first similar historical news article set or the second similar historical news article set as a reference manuscript set;

[0264] Among them, the cross-modal search expression is as follows:

[0265] Similarity = α·BERT(text_embed) + β·CLIP(img_embed) + γ·TimeDecay(t)

[0266] Where BERT(text_embed) represents the text semantic features extracted by the BERT model, α represents the semantic feature similarity requirement; CLIP(img_embed) represents the cross-modal feature alignment of image and text by the CLIP model, β represents the correlation requirement between image and text, and γ·TimeDecay(t) represents the time decay factor.

[0267] For example, this application can retrieve the top 10-20 excellent historical articles related to the topic semantics from the historical news article library based on the different news genres selected by the user and the specific elements for different genres filled in by the user through genre, elements and main points.

[0268] Of course, in addition to extracting the top 10 to 20 historical news articles by relevance as small sample drafts and submitting them to the language model, users can also select an article from the retrieved historical excellent manuscripts that best meets the current writing scenario or style requirements as a special reference sample and provide it to the language model.

[0269] Here, according to the preset creative instructions corresponding to different genres, the retrieved excellent historical manuscripts are input into a specific language model as small sample instances to carry out AIGC (Artificial Intelligence Generated Content) manuscript writing.

[0270] It's important to note that when dynamically generating creative instructions, the template will search for semantically related news articles of the same genre in the historical library based on other parameters such as genre type, and include them as style references in the reference prompt instructions. If no suitable reference articles are retrieved, the template will not include the reference prompt instructions.

[0271] When using a specific system prompt word template developed by an expert group for the same genre, you only need to replace the content about the theme and genre elements in the template, and refer to the prompt instructions. Through the semantics of the theme and element content, retrieve excellent A drafts with similar or related message semantics and scenarios from the historical database as small sample reference examples, and submit them to the designated language model as creation instructions to write the outline of the message manuscript and the summary content of each part.

[0272] Among them, the news genre writing standards and elements formulated by the expert group must be scientific and standardized, and will be continuously optimized and iterated in actual applications, and adjusted and modified after a lot of practical applications. In addition, this application will also retrieve relevant reference examples from the historical data manuscript library based on semantics. After data governance, this application will classify and store the historical manuscript library according to news genres, and perform semantic vectorization and embedding storage on the text to improve the accuracy of semantic retrieval; it will also classify and label historical news articles of different genres, such as labeling multiple weights of excellent, representative and paradigm reference value with different coefficients. In actual retrieval and recall, in addition to semantic sorting, other parameters are dynamically adjusted for different application scenarios, so that the historical manuscripts that are most suitable as reference samples can be recalled as important reference samples for large-scale model-assisted writing.

[0273] Furthermore, in a specific embodiment provided by the present application, the system further includes a historical news article feature library;

[0274] The steps of constructing the historical news article feature database include:

[0275] Obtain historical news release data;

[0276] Use the BERT-News model to extract semantic vectors from historical news article data;

[0277] Use the GPT-StyleCLIP extraction model to extract style vectors from historical news release data;

[0278] Analyze the hierarchical relationship between the title, introduction, and body of historical news articles data as the structural vector of the historical news articles data;

[0279] Use the ResNet-152 model to extract visual features of historical news release data;

[0280] The semantic vectors of the historical news article data, the style vectors of the historical news article data, the structural vectors of the historical news article data, and the visual features of the historical news article data are stored as the historical news article mixed coding objects;

[0281] Several mixed coding objects of historical news articles constitute the historical news article feature library.

[0282] Specifically, this application constructs a five-level classification tag library based on the news genre classification system for historical news release data.

[0283] In addition, hybrid coding technology is used to decompose historical manuscripts into semantic vectors (using the BERT-News model), style vectors (extracted based on GPT-StyleCLIP), and structural vectors (parsing the hierarchical relationship between title-lead-text).

[0284] During the draft generation stage, the top 50 high-click-rate articles of the same genre can be retrieved from the knowledge base based on the user-specified genre tag (such as "financial in-depth report") as a style reference.

[0285] In a specific embodiment provided in this application, according to the creation instruction, the language model is controlled to generate a pre-arranged framework consisting of several outlines, specifically including:

[0286] Match factual content in several historical news articles based on factor data;

[0287] Based on the sentence features, paragraph structure, and term usage frequency of several historical news articles, a pre-arranged framework of similar style is generated.

[0288] Specifically, the language model extracts the writing elements of a specific genre input by the user and similar genre samples retrieved from historical news materials according to the pre-arranged workflow, and performs pre-arranged workflow writing creation.

[0289] The framework generation module 13 generates the RAG architecture based on retrieval enhancement and develops a dual-channel generator, including a content generation channel and a style simulation channel.

[0290] The content generation channel leverages retrieved historical manuscripts to provide factual support (e.g., comparisons of data from previous events). The style simulation channel extracts sentence features (average sentence length, distribution of rhetorical devices), paragraph structure (inverted pyramid / hourglass), and terminology frequency from reference manuscripts, and fine-tunes the generation model through comparative learning.

[0291] Afterwards, the content filling module 14 obtains the draft framework updated based on the pre-arranged framework, and controls the language model to generate a draft text consisting of several paragraphs based on the draft framework.

[0292] It should be noted that the updated draft framework obtained by content filling module 14 based on the pre-arranged framework refers to the "draft framework" article outline generated by the user after adding, modifying, deleting, or adjusting the order of the pre-arranged framework (i.e., the article outline) generated by framework generation module 13. Subsequently, content filling module 14 expands each article outline based on the element data in the draft framework using a language model, thereby generating a draft text consisting of several paragraphs.

[0293] To improve user experience, in a specific embodiment provided in this application, the system further includes a style calibration module 15 corresponding to the style calibration node. The style calibration module 15 is used to:

[0294] Obtain selected text based on the draft text, as well as update instructions or update materials corresponding to the selected text;

[0295] According to the update instructions or update materials corresponding to the selected text, text is generated through the language model, and the generated text is used as the replacement text;

[0296] Wherein, the update instruction includes at least one of a polish instruction, a modification instruction, an expansion instruction, and a rewrite instruction;

[0297] The update material includes at least one of field, style, emotion, timeliness, and word count.

[0298] It is understandable that the style calibration module 15 can expand each paragraph according to the writing outline. After the expansion is completed, the style calibration module 15 can also polish and modify the entire text paragraph by paragraph.

[0299] Specifically, the expansion or polishing here refers to paragraph-level expansion based on the news article writing outline and the brief content of each part in the first draft text, after the user adds, deletes, or adjusts the outline and related parts. After the user selects the paragraph to be expanded with the cursor (it can be the entire text or a partial paragraph), a menu will pop up. After the user selects the "Expand Instructions" function, a dialog box will pop up. The user can fill in specific "expand instructions" or prompt words in this dialog box, or add supplementary materials to the prompt words to supplement the content. The language model will then combine the expansion instructions and supplementary materials filled in by the user in the dialog box based on the previous news genre, writing theme and genre elements, as well as the writing outline and summary content generated in the previous step, to synthesize the expansion prompt word template content and submit it to the general writing language model for content expansion of the specific paragraph.

[0300] The specific instruction requirements of the expansion instruction, such as the number of words, style, emotion, etc., are filled in by the user in this pop-up dialog box. In addition, this application will also provide some reference samples of expansion prompt words for different genres and paragraphs for users to refer to and adopt.

[0301] The operating logic of the polish instruction, modify instruction, and rewrite instruction is the same as that of the expansion instruction, and will not be repeated here.

[0302] Furthermore, the style calibration module 15 is used to review and proofread the first draft text. Specifically, the style calibration module 15 will filter the generated text for sensitive words based on its own sensitive word library and provide revision suggestions. The style calibration module 15 will then use the language model to further check for typos and commonly used standard terms in sentences and paragraphs, and provide revision suggestions.

[0303] The style calibration module 15 can also check the overall text coherence and text semantic logic of the draft text, propose modification suggestions, and the user can check and confirm the modification, and output the updated draft text.

[0304] Based on the updated draft text, the style calibration module 15 allows the user to perform segmented adjustments, polishing, modification, expansion, rewriting, and other operations on the updated draft text again, and then finalize it.

[0305] Furthermore, in a preferred embodiment provided by the present application, the system further includes a diagram matching module 16 for performing semantic vector conversion based on the draft text to determine core events, emotional tendencies, and scene keywords as text features;

[0306] It is also used to control the image matching model to match images corresponding to the draft text or paragraph text as illustrations based on the draft text or text features.

[0307] Specifically, if the user does not need to add pictures to the draft text, the draft text can be directly output as the final draft. If the user needs to add pictures to the draft text, the first draft text intelligent picture matching stage will be entered.

[0308] If it is necessary to intelligently match pictures for the generated article, the user can select a picture matching mode. The picture matching module 16 uses global matching, that is, matching pictures for the entire article, or can choose to perform local semantic intelligent matching for specific paragraphs of the article.

[0309] In the full-text semantic image matching process, the image matching module 16 will intelligently retrieve matching images from the private image library based on the full-text genre, style or narrative events.

[0310] In the local semantic image matching process, the image matching module 16 will intelligently retrieve matching images from the private image library based on the semantics and scene factors of a specific paragraph.

[0311] Specifically, when the image matching module 16 detects a specific entity (such as "garden"), it directly retrieves garden-related images from the local image library.

[0312] If no suitable picture is found, or the style, emotion, or composition does not meet the requirements, users can also use Wenshengtu to generate pictures intelligently by AI for illustration.

[0313] The illustration module 16 uses a multimodal vector model to vectorize and embed the text and pictures of historical news articles, encoding the content of historical news articles into 512-dimensional semantic vectors, focusing on capturing core events (60% weight), emotional tendencies (20% weight), and scene keywords (20% weight).

[0314] This application also constructs an image feature library, and the construction process is as follows:

[0315] Perform deep feature extraction on private image libraries, using ResNet-152 to extract visual features and the CLIP model to extract image-text matching features;

[0316] A three-level index is established: subject object (face / building / natural scenery) - scene category (conference site / disaster site) - emotional tone (serious / cheerful / sad).

[0317] When a specific entity (such as "airport") is detected, the entity association library is directly called.

[0318] The image matching module 16 can also map abstract concepts into visual elements through cross-modal alignment technology (such as "garden visitor volume" matching the upward arrow chart + the crowd ticket purchasing scene diagram).

[0319] In addition, the illustration module 16 can automatically select the illustration style based on the genre of the manuscript (data charts are preferred for news articles, and high-resolution scene pictures are used for feature articles).

[0320] The picture matching module 16 can also integrate review nodes to perform copyright verification on candidate pictures (matching known copyright libraries), sensitive content filtering (using self-built sensitive words to filter sensitive content), and resolution adaptation (automatically cropping according to the publishing platform).

[0321] Finally, in the image-text matching test phase, the image matching module 16 sorts the images by relevance based on the full text, or the semantic retrieval of local paragraphs / the AI-generated image matching list. The user ultimately determines and selects the most suitable image to insert into the corresponding position of the article, completing the intelligent image matching and outputting the finalized article with both text and images.

[0322] The AI-powered intelligent manuscript creation system 100, based on historical manuscripts, provided in this application, features a spatiotemporal data fusion engine. Leveraging news event knowledge graph technology, it constructs a spatiotemporal correlation model covering 50 years of historical news data. By dynamically annotating event causal relationships (e.g., "increased infrastructure investment → raw material price fluctuations → linkage to manufacturing costs"), it automatically links data across multiple cycles. Testing has shown that this engine can increase the depth of data correlation in financial reports by 40%.

[0323] Furthermore, the AI-powered intelligent article creation system 100, based on historical articles, features a three-step intelligent creation process, including a historical feature extraction layer, a dynamic event modeling layer, and a multimodal verification mechanism. Specifically, the historical feature extraction layer uses NLP technology to identify the intent of historical articles and extract high-frequency semantic frameworks (e.g., the "match results comparison - player development history - impact of technological innovation" structure in sports event reports). The dynamic event modeling layer incorporates machine learning technology to analyze the deviation of emergencies from historical patterns in real time and automatically generates a trend prediction module. The multimodal verification mechanism fact-checks and reinforces the opinions of AI-generated content, ensuring that the output adheres to journalistic ethics and introducing human review for final approval.

[0324] Unlike traditional static templates, the AI-powered manuscript creation system 100, based on historical manuscripts, is equipped with a self-evolutionary algorithm module. Each generated manuscript is fed back into the historical database, and the creation path is optimized through reinforcement learning.

[0325] The AI ​​intelligent manuscript creation system 100 based on historical manuscripts has the following advantages:

[0326] 1. Dynamic knowledge evolution is achieved throughout the entire news creation process through the spatiotemporal correlation of historical data, collaborative optimization of generative models, and multimodal semantic matching. By establishing a deep human-machine collaborative mechanism, human intervention channels are retained at key nodes (outline confirmation, sensitive word filtering, and image selection), forming a virtuous cycle of "AI generation - manual verification - system evolution." This improves the utilization of historical data, transforming old data into new assets, improving the efficiency of article image matching, and significantly reducing intellectual property risks associated with images.

[0327] 2. Build a knowledge graph of news events covering 75 years of historical data, using dynamic annotation technology to establish causal chains (e.g., "increased infrastructure investment → raw material price fluctuations → manufacturing cost linkage"). A spatiotemporal decay factor algorithm is used to link data across time periods. This overcomes the fragmented use of historical data in existing systems and significantly increases the depth of historical data association. This enables automatic cross-period data linkage for financial reports, and automatically links historical performance of past athletes for sports reports, creating in-depth content with valuable longitudinal comparisons.

[0328] 3. Utilizing node-based editing technology, news creation is broken down into core nodes, including topic planning, source material retrieval, framework generation, content filling, content security verification, and style calibration. Human interaction is required at each core node. Compared to existing writing tools, this invention implements multi-level correlation mining of historical news data. The suggested data used during the collaborative editing phase is derived from the correlation information generated by graph-based historical data processing. A human-computer collaborative interaction mechanism allows non-technical editors to interact naturally with the AI ​​system. This deep human-computer collaboration mechanism retains human intervention at key nodes (e.g., outline confirmation, sensitive word filtering, and image selection), creating a virtuous cycle of "AI generation-human verification-system evolution." Intelligent image matching automatically matches text semantics, style, genre, setting, and sentiment, enabling coordinated human-computer interaction. A real-time suggestion engine generates intelligent optimization solutions simultaneously with manual modifications (e.g., automatically presenting a library of classic sentence patterns from previous reports on similar events when a user adjusts the "event introduction").

[0329] 4. Develop a multimodal matching algorithm based on CLIP-ViT (MatchScore = Σcos(Ti,Ij)·e^(-λ|tT-tI|)), constructing a three-level image indexing system encompassing subject, scene category, and emotional tone, and incorporating a reinforcement learning feedback loop. This algorithm addresses the semantic gap between images and text in traditional image matching systems, significantly improving the accuracy of image matching. When reporting on "Garden Visits," it automatically matches an upward trend chart with an image of a ticket purchase scene. After manual correction of the generated content, subsequent image matching strategies can be automatically optimized, forming a continuously evolving creative system.

[0330] It should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, product, or apparatus. In the absence of further limitations, the phrase "comprises a..." to define an element does not preclude the presence of other identical elements in the process, method, product, or apparatus that includes the element.

[0331] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0332] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. An AI intelligent manuscript creation method based on historical manuscripts, characterized in that: The following steps are involved: Get the genre attributes of the topic; According to the genre attributes of the topic, match the element collection strategy corresponding to the genre attributes; Obtain updated feature data based on feature collection strategies; Dynamically generate creative instructions based on genre attributes and element data; According to the creative instructions, the language model is controlled to generate a pre-arranged framework consisting of several outlines; Obtain a first draft framework based on pre-edited framework updates; According to the draft framework, the language model is controlled to generate a draft text consisting of several paragraphs; The genre attribute of the topic includes at least one of news genre, correspondence genre, in-depth report genre or commentary genre; Based on the genre attributes of the topic, the element collection strategies that match the corresponding genre attributes include: The genre attribute of the selected topic is selected as news genre or communication genre, matching the real-time factor data collection strategy; The genre attribute of the selected topic is in-depth report genre or commentary genre, matching the historical element data retrieval strategy; Based on genre attributes and feature data, creative instructions are dynamically generated, including: According to the element data and genre attributes, several historical news articles are matched as reference prompt instructions; Generate creation instructions including element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and reference prompt instructions; When no historical news item is matched, a creation instruction is generated including creation instructions of element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, and workflow instructions.

2. The AI ​​intelligent manuscript creation method based on historical manuscripts according to claim 1 is characterized in that: The method further comprises: Based on the first draft text, perform semantic vector conversion to determine the core events, emotional tendencies, and scene keywords as text features; According to the draft text or text features, the image matching model is controlled to match images corresponding to the draft text or paragraph text as illustrations.

3. The AI ​​intelligent manuscript creation method based on historical manuscripts according to claim 1 is characterized in that: The method further comprises: Obtain historical news release data; Use the BERT-News model to extract semantic vectors from historical news article data; Use the GPT-StyleCLIP extraction model to extract style vectors from historical news release data; Analyze the hierarchical relationship between the title, introduction, and body of historical news articles data as the structural vector of the historical news articles data; Use the ResNet-152 model to extract visual features of historical news release data; The semantic vectors of the historical news article data, the style vectors of the historical news article data, the structural vectors of the historical news article data, and the visual features of the historical news article data are stored as the historical news article mixed coding objects; Several mixed coding objects of historical news articles constitute the historical news article feature library.

4. The AI ​​intelligent manuscript creation method based on historical manuscripts according to claim 3 is characterized in that: The matching of several historical news articles based on element data and genre attributes specifically includes: Filter the first similar historical news article collection based on genre attributes: Update cross-modal search formulas based on feature data; Using the updated cross-modal search formula, the second most similar historical news article set is filtered from the historical news article feature library: Based on the manuscript quality evaluation coefficient, a preset number of historical news articles are selected from the first similar historical news article set or the second similar historical news article set as a reference manuscript set; Among them, the cross-modal retrieval expression is Similarity = α·BERT(text_embed) + β·CLIP(img_embed) + γ·TimeDecay(t) Where BERT(text_embed) represents the text semantic features extracted by the BERT model, α represents the semantic feature similarity requirement; CLIP(img_embed) represents the cross-modal feature alignment of image and text by the CLIP model, β represents the correlation requirement between image and text, and γ·TimeDecay(t) represents the time decay factor.

5. The AI ​​intelligent manuscript creation method based on historical manuscripts according to claim 1 is characterized in that: According to the creative instructions, the language model is controlled to generate a pre-arranged framework consisting of several outlines, including: Match factual content in several historical news articles based on factor data; Based on the sentence features, paragraph structure, and term usage frequency of several historical news articles, a pre-arranged framework of similar style is generated.

6. The AI ​​intelligent manuscript creation method based on historical manuscripts according to claim 1, characterized in that: The method further comprises: Obtain selected text based on the draft text, as well as update instructions or update materials corresponding to the selected text; According to the update instructions or update materials corresponding to the selected text, text is generated through the language model, and the generated text is used as the replacement text; Wherein, the update instruction includes at least one of a polish instruction, a modification instruction, an expansion instruction, and a rewrite instruction; The update material includes at least one of field, style, emotion, timeliness, and word count.

7. An AI intelligent manuscript creation system based on historical manuscripts, characterized by: include: Topic planning module, used to obtain the genre attributes of the topic; The material retrieval module is used to match the feature collection strategy corresponding to the genre attribute of the selected topic; it is also used to obtain the feature data updated based on the feature collection strategy; A framework generation module is used to dynamically generate creative instructions based on genre attributes and element data; it is also used to control the language model to generate a pre-arranged framework consisting of several outlines based on the creative instructions; A content filling module is used to obtain a draft framework updated based on the pre-arranged framework; and is also used to control the language model to generate a draft text consisting of several paragraphs based on the draft framework; The genre attribute of the topic includes at least one of news genre, correspondence genre, in-depth report genre or commentary genre; The material retrieval module matches the element collection strategy corresponding to the genre attribute according to the genre attribute of the selected topic, specifically including: The genre attribute of the selected topic is selected as news genre or communication genre, matching the real-time factor data collection strategy; The genre attribute of the selected topic is in-depth report genre or commentary genre, matching the historical element data retrieval strategy; The framework generation module dynamically generates creation instructions based on genre attributes and element data, specifically including: According to the element data and genre attributes, several historical news articles are matched as reference prompt instructions; Generate creation instructions including element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, workflow instructions, and reference prompt instructions; When no historical news item is matched, a creation instruction is generated including creation instructions of element data, background prompt instructions, role prompt instructions, ability prompt instructions, target prompt instructions, restriction rule instructions, output rule instructions, and workflow instructions.

8. The AI ​​intelligent manuscript creation system based on historical manuscripts according to claim 7, characterized in that: The system also includes a diagram module for converting the text into semantic vectors based on the first draft text, and determining core events, emotional tendencies, and scene keywords as text features; It is also used to control the image matching model to match images corresponding to the draft text or paragraph text as illustrations based on the draft text or text features.

9. The AI ​​intelligent manuscript creation system based on historical manuscripts according to claim 7, characterized in that: The system also includes a historical news release feature library; The steps of constructing the historical news article feature database include: Obtain historical news release data; Use the BERT-News model to extract semantic vectors from historical news article data; Use the GPT-StyleCLIP extraction model to extract style vectors from historical news release data; Analyze the hierarchical relationship between the title, introduction, and body of historical news articles data as the structural vector of the historical news articles data; Use the ResNet-152 model to extract visual features of historical news release data; The semantic vectors of the historical news article data, the style vectors of the historical news article data, the structural vectors of the historical news article data, and the visual features of the historical news article data are stored as the historical news article mixed coding objects; Several mixed coding objects of historical news articles constitute the historical news article feature library.

10. The AI ​​intelligent manuscript creation system based on historical manuscripts according to claim 9, characterized in that: The framework generation module matches several historical news articles based on element data and genre attributes, including: Filter the first similar historical news article collection based on genre attributes: Update cross-modal search formulas based on feature data; Using the updated cross-modal search formula, the second most similar historical news article set is filtered from the historical news article feature library: Based on the manuscript quality evaluation coefficient, a preset number of historical news articles are selected from the first similar historical news article set or the second similar historical news article set as a reference manuscript set; Among them, the cross-modal retrieval expression is Similarity = α·BERT(text_embed) + β·CLIP(img_embed) + γ·TimeDecay(t) Where BERT(text_embed) represents the text semantic features extracted by the BERT model, α represents the semantic feature similarity requirement; CLIP(img_embed) represents the cross-modal feature alignment of image and text by the CLIP model, β represents the correlation requirement between image and text, and γ·TimeDecay(t) represents the time decay factor.

Citation Information

Patent Citations

  • News influence calculation method and device, computer equipment and storage medium

    CN111552882A

  • Hotspot mining system and method based on Internet financial information, and medium

    CN119537665A