A multi-platform content generation method and system based on brand corpus atlas and human setup proportion constraint
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明旨在提供一种基于品牌语料图谱与人设占比约束的多平台内容生成方法及系统,以解决现有人工智能生成内容中品牌表达不稳定、人设比例不可控、多平台适配效率低及生成结果难以闭环修正的问题
[0065] This invention constructs a brand corpus graph consisting of a semantic knowledge base, a relationship graph library, and a strategy action library. This transforms scattered brand corpus, product corpus, user profile corpus, activity corpus, and platform rule corpus into recallable, inferable, and combinable content evidence units, ensuring that generated content has clear brand basis and semantic origin. Simultaneously, by calculating the target persona proportion vector T and the actual persona proportion vector A, the invention quantifies the expression ratio of user profiles, brand profiles, product profiles, activity scenarios, and platform styles in candidate content, and identifies dimensions of excessive or insufficient expression based on the deviation vector D. Furthermore, by combining compliance review, platform format review, and quality scoring results, the invention automatically routes to the corresponding rewrite node in the state graph workflow, performing targeted corrections on generated prompts or content evidence units. This improves the compliance pass rate, platform adaptation efficiency, and content output stability of multi-platform content generation while ensuring brand consistency and persona controllability.
Smart Images

Figure CN122311149B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence content generation and digital marketing technology, specifically involving a multi-platform content generation method and system based on brand corpus graph and persona proportion constraints. Background Technology
[0002] With the development of generative artificial intelligence, large language models, and knowledge graph technologies, brands are increasingly adopting automated content generation systems to produce content across multiple platforms, including product seeding, event promotion, user Q&A, and brand education. Existing technologies typically collect brand information, product selling points, historical copy, user profiles, or platform rules, combining them with prompt templates or large language models to generate marketing copy, thereby improving content production efficiency. For example, prior art document CN119515558A discloses a method for generating personalized copy. It segments customer groups based on historical customer characteristic data, obtains keyword combinations for copy based on historical copy for each customer group, constructs customer group profiles, and then combines target customer characteristics, customer group profiles, and product recommendation schemes to generate personalized copy through a large copy generation model, thus improving the targeting of copywriting.
[0003] However, existing content generation solutions still have the following shortcomings: First, brand corpora, product corpora, user profile corpora, and platform rule corpora are mostly stored in scattered forms such as text libraries, tag libraries, or template libraries. They lack a unified corpus graph that can simultaneously express semantic concepts, entity relationships, and content action strategies, leading to unclear evidence sources and unstable brand expression in the generated content. Second, existing user profiles or customer group profiles are usually used to determine target audiences or generate recommendation directions, rarely converting the proportions of "user profile, brand profile, product profile, activity scenario, platform style," etc., into calculable target proportion constraints. Therefore, the generated results are prone to problems such as product selling point stacking, insufficient brand tone, weak scenario sense, or unbalanced platform style. Third, multi-platform content generation often relies on rewriting templates from different platforms separately, lacking a mechanism for unified adaptation, review, and traceability based on the same content evidence unit under different platform rules. Fourth, existing systems mostly perform simple sensitive word filtering or manual modification after generation, failing to automatically route to the corresponding rewriting node for targeted correction based on the deviation between the actual content and the target persona proportion, compliance review results, platform format review results, and quality score results.
[0004] Therefore, there is an urgent need for a content generation method and system that can integrate brand corpus graphs, persona proportion constraints, multi-platform adaptation, and closed-loop rewriting mechanisms to improve the consistency, controllability, compliance, and platform adaptation efficiency of brand content generation. Summary of the Invention
[0005] This invention aims to provide a multi-platform content generation method and system based on brand corpus graph and persona proportion constraints, in order to solve the problems of unstable brand expression, uncontrollable persona proportion, low multi-platform adaptation efficiency, and difficulty in closed-loop correction of generation results in existing AI-generated content.
[0006] To achieve the above-mentioned technical objectives, the present invention provides the following technical solutions.
[0007] In a first aspect, the present invention provides a multi-platform content generation method based on brand corpus graphs and persona proportion constraints, comprising the following steps:
[0008] S1. Obtain basic brand corpus, product corpus, user profile corpus, activity corpus, and platform rule corpus, and construct a brand corpus graph consisting of a semantic knowledge base, a relationship graph library, and a strategy action library;
[0009] S2. Receive content generation task, generate target persona proportion vector T and allowable deviation ε according to target platform, user tags, brand identity, product identity and activity scenario. The target persona proportion vector T includes at least user profile proportion, brand profile proportion, product profile proportion, activity scenario proportion and platform style proportion.
[0010] S3. Based on the content generation task, recall semantic elements, reasoning paths and strategy actions from the brand corpus graph, merge reasoning paths with confidence levels that meet the threshold into content evidence units, and compile the content evidence units, the target persona proportion vector T and the target platform configuration into generation prompts.
[0011] S4. Call the generation model to generate candidate content, extract the semantic features of the candidate content and map them to each persona dimension, normalize to obtain the actual persona proportion vector A, calculate the deviation vector D=AT, and perform compliance review, platform format review and quality scoring.
[0012] S5. When the absolute value of any dimension in the deviation vector D is greater than the allowable deviation ε, or when any of the compliance review, platform format review, or quality score fails to meet the standard, the corresponding rewrite node is routed in the state diagram workflow according to the failure type. The generated prompt or content evidence unit is corrected and regenerated. When the requirements are met, the content adapted to the target platform and the generated traceability record are output.
[0013] Specifically, in step S1, the semantic knowledge base is used to store brand semantic concept records. Each brand semantic concept record includes an industry field, a primary category field, a secondary category field, a keyword field, a synonym field, a standard definition field, an AI reasoning instruction field, an entity type field, an associated keyword field, and a risk level field.
[0014] The relation graph library is used to store relation triples. Each relation triple includes a subject term, a subject type, a relation type, an object term, an object type, a relation category, a confidence level, and an explanatory text. The relation type includes one or more of the following: support relation, causal relation, improvement relation, inclusion relation, applicable object relation, and prohibition relation.
[0015] The strategy action library is used to store strategy action records. Each strategy action record includes trigger intent, intent keywords, trigger conditions, execution action, action type, response template, association graph reasoning path, and compliance checkpoint.
[0016] Specifically, in step S2, the target persona proportion vector T is represented as:
[0017] T = (tu, tb, tp, ta, ts);
[0018] Where tu represents the target percentage of user profile, tb represents the target percentage of brand profile, tp represents the target percentage of product profile, ta represents the target percentage of activity scenario, and ts represents the target percentage of platform style, and tu+tb+tp+ta+ts=1;
[0019] The allowable deviation ε is the maximum allowable deviation of the actual persona proportion vector A from the target persona proportion vector T in a single dimension.
[0020] When the target platform is a content community platform, increase the values of tu and ts; when the target platform is a WeChat official account platform, increase the values of tb and tp; when the target platform is a short video platform, increase the values of ta and ts.
[0021] When content generation tasks contain strong marketing objectives, limit TP (Title, Message, Page) to no higher than the preset product mention limit to avoid over-marketing of candidate content.
[0022] Specifically, in step S3, based on the content generation task, semantic elements, reasoning paths, and strategy actions are retrieved from the brand corpus graph, including:
[0023] The content generation task is analyzed to obtain the target industry, target product, target user tags, target scenario, and target platform;
[0024] Retrieve keywords, synonyms, standard definitions, and AI reasoning instructions that match the target product, target user tags, and target scenario from the semantic knowledge base to form a set of candidate semantic elements;
[0025] In the relation graph database, path retrieval is performed starting from candidate semantic elements to obtain at least one candidate reasoning path consisting of subject words, relation types, and object words;
[0026] In the strategy action library, corresponding strategy actions are matched according to the target scenario and user tags. The strategy actions include at least one of the following: educational explanation, selling point introduction, scenario resonance, comparative explanation, and interactive guidance.
[0027] The candidate semantic element set, candidate reasoning path, and policy action are associated to form an intermediate candidate evidence set for generating prompts and compiling.
[0028] Specifically, in step S3, reasoning paths that meet the confidence threshold are merged into content evidence units, including:
[0029] Candidate inference paths are divided into core paths, supplementary paths, and reference paths according to their confidence levels. The confidence level of the core path is not lower than the first confidence threshold, the confidence level of the supplementary path is lower than the first confidence threshold but not lower than the second confidence threshold, and the confidence level of the reference path is lower than the second confidence threshold.
[0030] The core path is used as the main source of arguments for the candidate content, the supplementary path is used as an auxiliary source of expression for the candidate content, and the reference path is used only to prompt the model to avoid semantic deviation and is not directly written into the main text.
[0031] When multiple candidate reasoning paths have conflicting conclusions, the candidate reasoning path with higher confidence is retained first; when the confidence levels are the same, they are selected according to the priority of the relation category; when the conflict still cannot be resolved, the corresponding semantic elements are written into the manual verification mark or the direct referencing prohibition mark.
[0032] The selected core path, supplementary path, corresponding semantic elements, strategic actions, and compliance checkpoints are combined into content evidence units.
[0033] Wherein, the first confidence threshold is the lower confidence limit for determining the core path, and the second confidence threshold is the lower confidence limit for determining the supplementary path.
[0034] Specifically, in step S4, semantic features of the candidate content are extracted and mapped to each persona dimension, and normalized to obtain the actual persona proportion vector A, including:
[0035] The candidate content is segmented into sentences to obtain multiple content fragments;
[0036] For each content fragment, entity recognition, sentiment recognition, scene word recognition, brand word recognition, product selling point recognition, platform style word recognition, and risk word recognition are performed to obtain a set of semantic features;
[0037] Based on the preset feature-dimension mapping table, the semantic features in the semantic feature set are mapped to the user profile dimension, brand profile dimension, product profile dimension, activity scenario dimension, and platform style dimension, respectively.
[0038] Calculate the matching score for each persona dimension, and normalize each matching score to obtain the actual persona proportion vector A;
[0039] The actual persona proportion vector A is represented as A=(au, ab, ap, aa, as), where au is the actual proportion of user profile, ab is the actual proportion of brand profile, ap is the actual proportion of product profile, aa is the actual proportion of activity scenario, and as is the actual proportion of platform style, and au+ab+ap+aa+as=1.
[0040] Specifically, in step S4, the deviation vector D=AT is calculated, including:
[0041] Calculate the user profile deviation du=au-tu, brand profile deviation db=ab-tb, product profile deviation dp=ap-tp, activity scenario deviation da=aa-ta, and platform style deviation ds=as-ts respectively;
[0042] Combine du, db, dp, da, and ds to form the deviation vector D = (du, db, dp, da, ds);
[0043] When any deviation value is greater than the allowable deviation ε, the persona dimension corresponding to that deviation value is marked as an excessive dimension.
[0044] When any deviation value is less than -ε, the persona dimension corresponding to that deviation value is marked as an insufficient dimension.
[0045] Rewrite control instructions are generated based on excessive and insufficient dimensions. The rewrite control instructions include one or more of the following: enhancing the expression of insufficient dimensions, weakening the expression of excessive dimensions, replacing the corresponding semantic elements, and adjusting the corresponding strategy actions.
[0046] Specifically, in step S5, routing to the corresponding rewrite node in the state graph workflow based on the failure type includes:
[0047] When the failure type is a deviation in the persona proportion, the route is routed to the persona rewriting node. The persona rewriting node determines the excessive and insufficient dimensions based on the deviation vector D, and adjusts the number of semantic elements, expression intensity, or fragment position in the content evidence unit corresponding to each persona dimension.
[0048] When the failure type is "compliance review not met", the route is routed to the compliance rewrite node. The compliance rewrite node writes the prohibited words, prohibited expressions, prohibited relationships or high-risk semantic elements into the negative constraint segment of the generated prompt, and deletes or replaces the corresponding semantic elements from the content evidence unit.
[0049] When the failure type is "platform format review not met", the route is routed to the platform format rewriting node. The platform format rewriting node corrects the title length, body text length, paragraph structure, interactive guidance, hot word injection method or disables elements according to the target platform configuration.
[0050] When the failure type is that the quality score does not meet the standard, the route is routed to the quality rewrite node, which adjusts the generation parameters, content structure requirements or strategy action sequence.
[0051] If the number of rewrites reaches the preset limit and still fails to meet the target, the route is sent to a manual review node.
[0052] More specifically, the target platform configuration includes at least the platform identifier, maximum title length, body text length range, content style tags, platform hot word sources, prohibited elements, interactive guidance rules, content structure template, and review rules;
[0053] When the same content generation task corresponds to multiple target platforms, different target platform configurations are applied based on the same content evidence unit and target persona proportion vector T to generate candidate content for multiple platforms.
[0054] For each candidate content on the platform, calculate the actual persona proportion vector A and the deviation vector D, and perform compliance review, platform format review and quality scoring respectively;
[0055] The generated traceability record includes at least the content generation task identifier, target platform, target persona percentage vector T, actual persona percentage vector A, deviation vector D, adopted reasoning path, strategy action, compliance review result, platform format review result, quality score result, number of rewrites, manual review status, and final output content.
[0056] Secondly, the present invention also provides a multi-platform content generation system based on brand corpus graphs and persona proportion constraints, for implementing the method described in the first aspect, the system comprising:
[0057] The corpus graph module is used to build and manage a brand corpus graph consisting of a semantic knowledge base, a relation graph library, and a strategy action library, and to generate semantic elements, inference paths, and strategy actions based on the content.
[0058] The character constraint module is used to generate the target character proportion vector T and allowable deviation ε, and calculate the actual character proportion vector A and deviation vector D based on the candidate content;
[0059] The prompt compilation module is used to compile the content evidence unit, the target persona proportion vector T, and the target platform configuration into a generated prompt.
[0060] The content generation module is used to call the generation model to generate candidate content;
[0061] The review and scoring module is used to perform compliance review, platform format review, and quality scoring on candidate content;
[0062] The state diagram workflow module is used to route candidate content to the corresponding rewrite node, manual review node, or result saving node based on the deviation vector D, compliance review results, platform format review results, and quality score results.
[0063] The multi-platform adaptation module is used to output the final content for the corresponding platform based on the configuration of different target platforms;
[0064] The traceability and recording module is used to record the content generation task, reasoning path, target persona proportion vector T, actual persona proportion vector A, deviation vector D, review and scoring results, rewriting process, and final output content.
[0065] This invention constructs a brand corpus graph consisting of a semantic knowledge base, a relationship graph library, and a strategy action library. This transforms scattered brand corpus, product corpus, user profile corpus, activity corpus, and platform rule corpus into recallable, inferable, and combinable content evidence units, ensuring that generated content has clear brand basis and semantic origin. Simultaneously, by calculating the target persona proportion vector T and the actual persona proportion vector A, the invention quantifies the expression ratio of user profiles, brand profiles, product profiles, activity scenarios, and platform styles in candidate content, and identifies dimensions of excessive or insufficient expression based on the deviation vector D. Furthermore, by combining compliance review, platform format review, and quality scoring results, the invention automatically routes to the corresponding rewrite node in the state graph workflow, performing targeted corrections on generated prompts or content evidence units. This improves the compliance pass rate, platform adaptation efficiency, and content output stability of multi-platform content generation while ensuring brand consistency and persona controllability. Attached Figure Description
[0066] Figure 1 This is a schematic diagram of the overall architecture of the multi-platform content generation system of the present invention.
[0067] Figure 2 This is a flowchart illustrating the multi-platform content generation method of the present invention.
[0068] Figure 3 This is a schematic diagram illustrating the relationship between the semantic knowledge base, the relational graph library, and the strategy action library in this invention, which collaboratively generate content evidence units.
[0069] Figure 4 This is a schematic diagram illustrating the calculation and correction relationship between the target persona proportion vector, the actual persona proportion vector, and the deviation vector in this invention.
[0070] Figure 5 This is a schematic diagram illustrating the process of routing and rewriting according to failure type in the state diagram workflow of this invention. Detailed Implementation
[0071] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0072] I. Overview of Implementation Methods
[0073] This embodiment provides a multi-platform content generation method and system based on brand corpus graph and persona proportion constraints. It is applicable to scenarios such as brand owners, marketing service platforms, content operation systems, e-commerce content generation systems, and private domain operation systems. It is used to automatically generate marketing content, product seeding content, Q&A content, event promotion content, or brand popularization content that is adapted to multiple target platforms based on brand, product, user profile, event scenario, and platform rules.
[0074] This invention does not simply generate text directly using a general generation model, nor does it merely concatenate product selling points, platform rules, and user tags into a prompt. Instead, it first constructs a brand corpus graph consisting of a semantic knowledge base, a relationship graph library, and a strategy action library. Then, it generates a target persona proportion vector based on the target platform and content task. Subsequently, it extracts semantic elements, reasoning paths, and strategy actions from the brand corpus graph, integrates them to form content evidence units, and compiles these content evidence units, the target persona proportion vector, and the target platform configuration into a generated prompt. After candidate content is generated, it calculates the actual persona proportion vector through semantic feature extraction, compares it with the target persona proportion vector to obtain a deviation vector, and combines it with compliance review, platform format review, and quality scoring results. Through a state graph workflow, it routes the deviation vector to the corresponding rewrite node to achieve targeted correction of the generated content.
[0075] Therefore, this invention can avoid the problems common in existing content generation solutions, such as unclear corpus sources, uncontrollable character expression, unstable brand tone, stacking of product selling points, low efficiency of multi-platform adaptation, and large amount of manual modification.
[0076] In this embodiment, the generation model can be a pre-trained general natural language generation model or a Chinese content generation model deployed privately by an enterprise. Semantic feature extraction can employ one or more combinations of rule dictionaries, named entity recognition models, text classification models, sentiment recognition models, and semantic similarity models. The technical solution of this invention does not involve retraining the basic generation model, but rather controls the input basis, generation boundary, platform format, and output correction process of the generation model through a brand corpus graph, target persona proportion vector, actual persona proportion vector, and a state graph closed-loop rewriting mechanism, enabling it to stably output content that conforms to brand tone, persona proportion, platform rules, and compliance requirements.
[0077] II. Terminology Explanation
[0078] To facilitate understanding of this implementation method, the relevant terms are explained below.
[0079] Brand corpus graph refers to a data system composed of a semantic knowledge base, a relationship graph library, and a strategy action library. It is used to express the semantic, inference, and content action relationships between brands, products, user profiles, activity scenarios, platform rules, and compliance rules.
[0080] A semantic knowledge base is a data set used to store brand-related semantic concepts. Its recorded content includes industry, classification, keywords, synonyms, standard definitions, model reasoning instructions, entity types, related keywords, and risk levels.
[0081] A relation graph library refers to a data collection that stores relationships between entities in the form of relation triples. Each relation triple includes at least a subject term, a relation type, and an object term, and further includes relation category, confidence level, and explanatory text. Examples include "core component—support—user experience improvement" and "platform rules—prohibit—absolute expression."
[0082] The strategy action library refers to a data collection used to store content generation strategies. Each strategy action record includes the trigger intent, intent keywords, trigger conditions, execution action, action type, response template, association graph reasoning path, and compliance checkpoints.
[0083] Generate prompts are structured text instructions input into the generative model, including brand evidence, content objectives, persona proportion requirements, platform format requirements, compliance restrictions, and output format requirements. Generate prompts guide the generative model to output candidate content that meets the task requirements.
[0084] A sequence generation model based on a self-attention mechanism refers to a model structure that can automatically model contextual relationships and generate natural language text based on an input text sequence. In this embodiment, this model serves as a candidate content generation tool, and its specific model structure does not constitute the sole limitation of this invention.
[0085] The target persona proportion vector T refers to a multi-dimensional expression proportion vector set before content generation based on the target platform, user tags, brand identity, product identity, and event scenario. It is used to constrain the expression proportion of user profile, brand profile, product profile, event scenario, and platform style in candidate content.
[0086] The actual persona proportion vector A refers to the actual expression proportion vector obtained after semantic feature extraction, dimension mapping, and normalization of candidate content.
[0087] The deviation vector D is the difference between the actual persona proportion vector A and the target persona proportion vector T, i.e., D=AT. It is used to identify whether there is over-expression or under-expression in each persona dimension.
[0088] Content evidence units refer to the structured evidence generated by the fusion of semantic elements, reasoning paths, strategic actions, and compliance checkpoints. They serve as an important component of the generation prompts and are input into the generation model.
[0089] A state diagram workflow is a process control structure consisting of data source acquisition nodes, prompt compilation nodes, content generation nodes, character proportion detection nodes, compliance review nodes, platform format review nodes, quality scoring nodes, rewriting nodes, manual review nodes, and result saving nodes. It executes conditional routing based on different review statuses and failure types.
[0090] III. System Structure and Implementation
[0091] See Figure 1 The system in this embodiment includes a corpus graph module, a character constraint module, a prompt compilation module, a content generation module, a review and scoring module, a state graph workflow module, a multi-platform adaptation module, and a traceability and recording module.
[0092] The corpus graph module is used to build and manage the semantic knowledge base, relation graph library, and strategy action library. This module can be implemented using a combination of relational databases, graph databases, and vector retrieval databases. The semantic knowledge base stores standardized brand concepts, product selling points, user profile terms, activity scenario terms, and platform rule terms; the relation graph library stores support relationships, causal relationships, applicable object relationships, and prohibited expression relationships between various entities; the strategy action library stores content organization actions under different platforms, different user tags, and different content objectives.
[0093] The persona constraint module generates a target persona proportion vector T based on the target platform, user tags, brand identifier, product identifier, and activity scenario. After candidate content generation, it calculates the actual persona proportion vector A and the deviation vector D. This module includes a dimension configuration unit, a target vector generation unit, a semantic feature extraction unit, an actual proportion calculation unit, and a deviation determination unit.
[0094] The prompt compilation module is used to compile content evidence units, target persona proportion vector T, target platform configuration, and generation parameters into generation prompts. These prompts can include system role segments, brand evidence segments, persona proportion constraint segments, platform format constraint segments, compliance negative constraint segments, and output format segments. Through these segmented prompts, the generation model can generate candidate content with clearly defined evidence sources, expression proportions, and output boundaries.
[0095] The content generation module is used to invoke a generative model to generate candidate content. The generative model can be a locally deployed model or a remote model invoked through a model service interface. The content generation module can receive parameters such as sampling temperature, maximum output length, repetition penalty factor, target language, and output format.
[0096] The review and scoring module is used to perform compliance review, platform format review, and quality scoring on candidate content. Compliance review may include detection of prohibited words, sensitive expressions, expressions that pose risks under advertising laws, expressions making medical or efficacy promises, absolute terms, and prohibited words on the platform. Platform format review may include heading length, body text length, number of paragraphs, use of emojis, hashtags, interactive guidance, and rules regarding prohibited brand exposure. Quality scoring may include assessments of personality appeal, brand consistency, logical completeness, content readability, marketing value, and innovation.
[0097] The state diagram workflow module is used to route candidate content to the corresponding rewrite node, manual review node, or result saving node based on the deviation vector D, compliance review results, platform format review results, and quality score results. This module is the key module for achieving closed-loop correction in this invention.
[0098] The multi-platform adaptation module is used to adapt title length, body text length, expression style, platform hot keywords, prohibited elements, and interactive guidance methods according to different platform configurations. For example, content community platforms emphasize authentic sharing and conversational expression; WeChat official account platforms emphasize complete structure, clear logic, and professional credibility; short video platforms emphasize attractive openings, rhythm, and interactive guidance.
[0099] The traceability and recording module is used to record content generation tasks, target platforms, target persona proportion vector T, actual persona proportion vector A, deviation vector D, adopted reasoning paths, strategy actions, compliance review results, platform format review results, quality score results, number of rewrites, manual review status, and final output content, thereby supporting quality backtracking, problem localization, version review, and subsequent corpus optimization.
[0100] IV. Detailed Implementation of the Method and Flow
[0101] See Figure 2 The method of this invention includes: first, acquiring corpora such as brand, product, user profile, activities, and platform rules to construct a brand corpus graph; then, generating a target persona proportion vector T and an allowable deviation ε based on the content generation task; next, recalling semantic elements, reasoning paths, and strategic actions from the brand corpus graph, fusing them to form content evidence units, and compiling them to generate prompts; subsequently, calling the generation model to generate candidate content, calculating the actual persona proportion vector A and deviation vector D, and performing compliance review, platform format review, and quality scoring; finally, determining whether the method meets the standards based on the detection and review results. If it does not meet the standards, it enters the state graph workflow for targeted rewriting; if it meets the standards, it outputs multi-platform adapted content and generates traceability records. The following details each step of the method of this invention.
[0102] S1. Constructing a brand corpus graph
[0103] This step provides a structured, reusable, and reasonable foundation of brand content evidence for subsequent content generation. Compared to directly inputting text data into the generation model, this step avoids problems such as scattered corpora, lack of evidence sources, inconsistent expression, and difficulties in platform migration.
[0104] In one embodiment, the following data is first acquired: brand basic data, product data, user profile data, activity data, and platform rule data. Brand basic data includes brand name, brand positioning, brand philosophy, prohibited expressions, standard brand statements, brand visual style, and historical content. Product data includes product name, product specifications, product ingredients, product selling points, target audience, usage scenarios, prohibited expressions, and risk warnings. User profile data includes user identity, age group, consumption stage, lifestyle, language style, pain points, purchase obstacles, and topics of interest. Activity data includes activity theme, activity time, discount rules, featured products, communication goals, and activity scenarios. Platform rule data includes the target platform's title length, body text length, hot word rules, prohibited words, content format, interactive guidance methods, and platform review rules.
[0105] Then, the above corpus undergoes cleaning and structuring processing. Cleaning includes removing duplicate text, standardizing brand name spelling, standardizing product aliases, eliminating expired promotional information, marking risky terms, and correcting obvious errors. Structuring processing includes entity extraction, keyword classification, synonym merging, standard definition generation, risk level labeling, and applicable scenario labeling.
[0106] Table 1 illustrates the basic field composition of each brand semantic concept record in the semantic knowledge base. Through industry fields, classification fields, keyword fields, standard definition fields, model inference instruction fields, and risk level fields, the scattered brand, product, user, and platform rule corpora are transformed into searchable, callable, and auditable structured data.
[0107] Each brand semantic concept record in the semantic knowledge base is transformed into structured data that can be searched, accessed, and audited. A typical example is shown in Table 1.
[0108] Table 1 Brand Semantic Concept Record Fields Table
[0109]
[0110] In the relationship graph library, each relationship record can adopt a triple structure: subject term—relationship type—object term. For example, "core component—support—user experience improvement," "user pain point—applicability—scene resonance expression," and "platform rules—prohibition—absolute terms." Relationship types can include support relationships, causal relationships, improvement relationships, inclusion relationships, applicable object relationships, and prohibited expression relationships. Each relationship further stores fields such as relationship category, confidence level, explanatory text, source identifier, and update time.
[0111] In the strategy action library, each strategy action record is used to express "how to organize content under what content intent and scenario conditions". For example, when the user tag is "working mom", the activity scenario is "commuting with child", and the content intent is "product recommendation", the strategy action can include "first introduce the pain point with the commuting scenario, then introduce the product selling points with gentle expression, and finally conclude with non-coercive interaction guidance".
[0112] S2. Generate the target persona proportion vector T and allowable deviation ε
[0113] This step involves transforming the abstract concept of "persona consistency" into a calculable, comparable, and correctable multidimensional proportion constraint.
[0114] In one embodiment, the system receives a content generation task. The task fields include task identifier, brand identifier, product identifier, platform identifier, user tags, activity identifier, content objective, tone intensity, marketing intensity, and output language. The platform identifier indicates the target platform, such as a content community platform, a public account platform, or a short video platform; the user tags indicate the target user group, such as working mothers, new mothers, users with sensitive skin, or fitness enthusiasts; and the content objective indicates the purpose of content generation, such as new product recommendation, brand education, activity conversion, or user Q&A.
[0115] The target persona proportion vector T is represented as:
[0116] T = (tu, tb, tp, ta, ts)
[0117] Where tu represents the target percentage for user profiles, tb represents the target percentage for brand profiles, tp represents the target percentage for product profiles, ta represents the target percentage for activity scenarios, and ts represents the target percentage for platform style, and satisfies the following:
[0118] tu+tb+tp+ta+ts=1.
[0119] On different platforms, the system can automatically adjust the target vector according to the platform configuration. For example, content community platforms emphasize authentic sharing and user resonance, and can be configured with T=(0.32, 0.16, 0.22, 0.18, 0.12); WeChat official account platforms emphasize brand credibility and complete product information, and can be configured with T=(0.18, 0.26, 0.28, 0.14, 0.14); short video platforms emphasize scene hooks and rhythm, and can be configured with T=(0.22, 0.14, 0.20, 0.24, 0.20).
[0120] The allowable deviation ε is used to determine whether the actual expression deviates from the target. For brand content with high consistency requirements, ε can be set to 0.05; for creative content, ε can be relaxed to 0.08; for promotional content, a separate deviation can be set for the product profile dimension to prevent excessive product expression from causing platform review risks or user aversion.
[0121] S3, recall semantic elements, reasoning paths, and strategic actions are integrated into a content evidence unit.
[0122] like Figure 3 As shown, this step does not directly splice the brand text into the generated prompt. Instead, it retrieves semantic elements, reasoning paths, and strategic actions from the brand corpus graph and merges the reasoning paths that meet the confidence threshold to form structured content evidence units.
[0123] In one embodiment, the system first parses the content generation task to obtain the target industry, target product, target user tags, target scenario, and target platform. For example, the task is to "generate content to promote Brand A infant formula on a content community platform, targeting working mothers who commute to work, with the activity scenario being commuting with children in the morning and evening, and promoting the main selling points as gentle absorption and daily comfort."
[0124] The system retrieves semantic elements such as gentle absorption, daily comfort experience, commuting mothers, genuine sharing feeling, and rushing in the morning and evening from the semantic knowledge base, and reads their standard definitions, synonyms, model inference instructions, and risk levels. If a keyword has a high risk level, such as involving expressions like "treatment" or "most effective," a compliance checkpoint is marked during the recall.
[0125] Then, the system performs path retrieval in the relational graph database, starting from the recalled semantic elements. For example, candidate reasoning paths are retrieved as shown in Table 2.
[0126] Table 2 Candidate Inference Paths and Confidence Table
[0127]
[0128] The system groups paths based on a first confidence threshold and a second confidence threshold; for example, the first confidence threshold is 0.90 and the second confidence threshold is 0.80. Paths with a confidence level not lower than the first confidence threshold are designated as core paths; paths with a confidence level greater than or equal to the second confidence threshold but less than the first confidence threshold are designated as supplementary paths; and paths with a confidence level lower than the second confidence threshold are designated as reference paths. Core paths can serve as the main basis for the main text; supplementary paths can serve as auxiliary explanations; and reference paths are not directly included in the main text but are only used to prompt the generation model to avoid semantic deviation.
[0129] When multiple candidate paths conflict, for example, one path points to "suitable for all users" while another path prohibits absolute expressions like "all" based on compliance rules, the compliance or risk control path should be retained first, and "all users" should be replaced with more prudent expressions such as "users who are concerned about the corresponding scenario" or "some users can choose according to the actual situation".
[0130] The strategy action library matches actions based on content tasks. For example, in the context of "content community platform seeding + working mother + commuting scenario," the strategy actions could be: the first part uses scenario resonance, the second part introduces user pain points, the third part uses the core path to explain the selling points, and the fourth part uses a light interaction to conclude. The system combines semantic elements, core paths, supplementary paths, strategy actions, and compliance checkpoints to form content evidence units.
[0131] Content evidence units can be represented using structured key-value data, including core paths, supplementary paths, user profile elements, brand profile elements, product profile elements, strategy actions, and compliance checkpoints. These content evidence units serve as crucial inputs for subsequent prompt generation, enabling the generation model to generate content based on explicit evidence and rules.
[0132] S4. Generate candidate content and calculate the actual persona proportion vector A and the deviation vector D.
[0133] like Figure 4 As shown, this step not only generates candidate content, but also performs quantifiable character proportion detection on the generated results, thereby providing a basis for targeted rewriting of S5.
[0134] In one embodiment, the prompt compilation module compiles the content evidence unit, the target persona proportion vector T, and the target platform configuration into a generated prompt. The generated prompt includes at least the following parts:
[0135] First, the system role segment: the generation model is limited to serving as a brand content generation assistant and is required to comply with brand corpus and platform rules.
[0136] Second, the content evidence section: includes the core path, supplementary paths, strategy actions, and usable expressions.
[0137] Third, the persona proportion constraint segment: write the target persona proportion vector T, for example, user profile about 30%, brand profile about 18%, product profile about 24%, activity scenario about 16%, and platform style about 12%.
[0138] Fourth, the platform format constraints section: This section includes the following: title length, body text length, style requirements, use of hot words, and prohibited elements.
[0139] Fifth, the negative compliance constraint section: This includes prohibiting promises of treatment, absolute terms, false claims of efficacy, and direct exaggeration.
[0140] Sixth, output format section: Requires output of title, body text, topic tags, interactive prompts, etc.
[0141] The content generation module calls the generation model to obtain candidate content. The generation model's parameters can include sampling temperature, candidate word sampling range, maximum output length, duplicate expression penalty factor, and existence penalty factor. If the subsequent quality score is low or compliance review fails, the sampling temperature can be lowered to improve output stability.
[0142] After candidate content is generated, the system segments it into sentences. For example, the main text is split into title segments, opening segments, scene segments, selling point segments, brand segments, and interactive segments. Then, entity recognition, keyword matching, sentiment recognition, scene word recognition, brand word recognition, product selling point recognition, platform style word recognition, and risk word recognition are performed on each segment.
[0143] The actual persona proportion vector A is represented as:
[0144] A = (au, ab, ap, aa, as)
[0145] Where au represents the actual percentage of user profiles, ab represents the actual percentage of brand profiles, ap represents the actual percentage of product profiles, aa represents the actual percentage of activity scenarios, and as represents the actual percentage of platform style, and satisfies au+ab+ap+aa+as=1.
[0146] In a specific calculation method, a matching score is calculated for each content fragment across various dimensions. If a content fragment contains user profile terms, first-person expressions, lifestyle identity terms, or user pain point terms, the user profile dimension score is increased; if it contains brand philosophy, brand standard expressions, or brand tone terms, the brand profile dimension score is increased; if it contains product names, ingredients, specifications, or selling points, the product profile dimension score is increased; if it contains time, location, activities, or scene actions, the activity scene dimension score is increased; and if it contains platform hot words, interactive expressions, emoticons, or expressions that conform to the platform style, the platform style dimension score is increased.
[0147] For example, if a candidate content is analyzed and the scores for each dimension are as follows: User profile dimension score 32, Brand profile dimension score 12, Product profile dimension score 30, Activity scenario dimension score 16, Platform style dimension score 10, and the total score is 100, then A = (0.32, 0.12, 0.30, 0.16, 0.10).
[0148] If the target persona proportion vector T = (0.30, 0.20, 0.25, 0.15, 0.10), then the deviation vector D = AT = (0.02, -0.08, 0.05, 0.01, 0.00). When the allowable deviation ε = 0.05, the brand profile dimension is lower than the target value and exceeds the allowable deviation, and is marked as an insufficient dimension; the product profile dimension is at the boundary, and can be slightly weakened according to platform rules. Based on this, the system generates rewrite control instructions: enhance brand tone expression, reduce the stacking of direct product selling points, and maintain user scenario expression.
[0149] Meanwhile, the review and scoring module performs compliance review, platform format review, and quality scoring. Compliance review results can be categorized as pass, fail, or require manual confirmation; platform format review checks for title length, body text length, prohibited symbols, number of topic tags, etc.; quality scoring uses a 100-point scale, and triggers quality rewriting when the score falls below a preset threshold.
[0150] S5, State Diagram Workflow Routing, Directed Rewriting, and Output
[0151] like Figure 5 As shown, this step identifies the failure type based on the deviation vector D, compliance review results, platform format review results, and quality score results, and routes it to the corresponding rewriting node in the state graph workflow, giving the rewriting a clear direction. The rewriting in this invention does not involve completely regenerating the candidate content; instead, it locates the excessive and insufficient dimensions of expression based on the deviation vector D, and only adjusts the semantic elements, strategy actions, or generated prompt fragments associated with the corresponding dimensions, ensuring a one-to-one correspondence between the rewriting direction and the cause of failure.
[0152] In one embodiment, the state diagram workflow includes a data source acquisition node, a prompt compilation node, a content generation node, a response parsing node, a persona proportion detection node, a compliance review node, a platform format review node, a quality scoring node, a routing judgment node, a persona rewriting node, a compliance rewriting node, a platform format rewriting node, a quality rewriting node, a manual review node, and a result saving node.
[0153] When the failure type is a deviation in the user persona proportion, the system enters the persona rewriting node. This node reads the deviation vector D and identifies excessive and insufficient dimensions. If the user persona is insufficient, it adds first-person experience, user pain points, and life scenarios; if the brand persona is insufficient, it adds brand tone, brand philosophy, and brand credibility expression; if the product persona is excessive, it reduces product name repetition, weakens parameter stacking, and changes to scenario-based expression; if the platform style is insufficient, it adds platform hot words or interactive expressions.
[0154] When the failure type is "compliance review not met," the system enters the compliance rewrite node. This node writes prohibited words, forbidden expressions, prohibited relationships, or high-risk semantic elements into the negative constraint segment of the generated prompt, and deletes or replaces the corresponding semantic elements from the content evidence unit. For example, it replaces "treat a problem" with "pay attention to changes in daily experience," and "most effective" with "more suitable for the direction that some users are concerned about."
[0155] When the failure type is "platform format not meeting review standards," the system enters the platform format rewriting node. For example, if the title of a content community platform exceeds the preset character limit, the title will be compressed; if the body text of a public account is less than the preset length, structured paragraphs will be expanded; if the short video copy lacks interactive guidance, an interactive sentence will be added at the end.
[0156] When the failure type is a quality score that does not meet the standard, the system enters the quality rewrite node. This node can reduce generation randomness, increase structure requirements, adjust the order of strategy actions, or add story details.
[0157] If the system fails to meet the target after reaching the preset limit for rewrite attempts, it will transfer the candidate content, failure reason, deviation vector D, and rewrite history to the manual review node. The preset limit can be 3 times. If the candidate content meets the deviation, compliance, platform format, and quality requirements, it will proceed to the result saving node, output the final content, and generate a traceability record.
[0158] The generated traceability record includes at least: task identifier, brand identifier, platform identifier, target persona percentage vector, actual persona percentage vector, deviation vector, graph inference path used, strategy actions, compliance review results, platform format review results, quality score results, number of rewrites, rewrite history, manual review status, final title, final text, and output time. This traceability record is used for subsequent quality analysis, debriefing, model parameter optimization, and brand corpus updates.
[0159] V. Model Structure
[0160] The technical solution of this invention involves generative models and semantic feature extraction models, but this invention does not require retraining of the basic generative model as a necessary condition. For feasibility, the following disclosure is provided.
[0161] 1. Generative Model
[0162] The generative model can employ a sequence generation model based on a self-attention mechanism. The input is the generation prompts generated by the prompt compilation module, and the output includes candidate titles, body text, topic tags, and interactive guidance. The generative model can be a locally deployed model or a remote model called through a model service interface. Table 3 shows a parameter setting table for a generative model.
[0163] Table 3. Parameter Settings for Generated Model
[0164]
[0165] The generative model does not require retraining for each brand, but brand knowledge can be injected through brand corpus graphs and cue compilation. If a company needs to further enhance a specific brand style, it can also use historical compliant content for supervised fine-tuning.
[0166] 2. Semantic Feature Extraction Model
[0167] Semantic feature extraction models are used to extract entities, sentiments, scenarios, brands, products, platform styles, and risk words from candidate content. A combined structure of "rule dictionary + text classification model + semantic similarity model" can be adopted.
[0168] The rule dictionary includes a brand dictionary, a product dictionary, a platform hot word dictionary, a risk word dictionary, and a scenario word dictionary. The text classification model can use a pre-trained Chinese language model to determine which category (user profile, brand profile, product profile, activity scenario, or platform style) a content fragment belongs to. The semantic similarity model is used to determine the semantic similarity between candidate content fragments and user profile descriptions, brand tone descriptions, and product selling point descriptions.
[0169] 3. Training dataset
[0170] The semantic feature extraction model can be trained using the datasets shown in Table 4.
[0171] Table 4 Training Data Table for Semantic Feature Extraction Model
[0172]
[0173] In one optional embodiment, the training sample size can be 20,000 content fragments, including 4,000 user profile fragments, 3,500 brand profile fragments, 4,500 product profile fragments, 3,500 activity scenario fragments, 3,500 platform style fragments, and 1,000 risk expression fragments. The training set, validation set, and test set are divided in an 8:1:1 ratio. The model output consists of five-dimensional category probabilities and risk category probabilities. The training objective can employ the cross-entropy loss function. The model is trained for 5 to 20 epochs, with a learning rate of 2×10⁻⁵ to 5×10⁻⁵, and the number of training samples per batch is 16 to 64. If a rule-based dictionary approach is used, the model can be skipped, and the actual proportions can be calculated directly through keyword matching, semantic similarity, and a manually configured weight table.
[0174] 4. Compliance Audit Model
[0175] Compliance audits can employ one or more of the following methods: sensitive word matching, regular expression rules, text classification models, and third-party content security services. Risk categories include absolute terms, exaggerated efficacy claims, medical promises, false comparisons, unauthorized brand exposure, platform-banned words, and sensitive expressions.
[0176] 5. Quality Scoring Model
[0177] Quality scoring can employ a combination of rule-based scoring and scoring models. Scoring dimensions include role immersion, natural expression, emotional expression, detailed description, language style, brand consistency, content completeness, marketing value, and innovation. Each item can be scored out of 100 or weighted separately, and the final scores are aggregated into a comprehensive quality score. Specific Implementation
[0178] Example 1: Content Generation for a Maternal and Infant Brand Content Community Platform
[0179] This embodiment takes a certain maternal and infant brand's milk powder product as the object, the target platform as a content community platform, the target user tag as commuting working mothers, the activity scenario as commuting with children in the morning and evening, the content goal as product recommendation, and the main semantic elements promoted include gentle absorption, daily comfortable experience, and ease of use.
[0180] The system constructs a target persona proportion vector T = (0.30, 0.18, 0.24, 0.16, 0.12), with an allowable deviation ε = 0.05. The system retrieves semantic elements such as "gentle absorption," "comfortable daily experience," "commuting mother," "authentic sharing," and "rushing to work in the morning and evening" from the semantic knowledge base; it retrieves paths such as "core components → support → improved user experience" and "product features → suitable for working mothers" from the relationship graph library; and it retrieves strategic actions such as "scene resonance—selling point introduction—light interactive conclusion" from the strategy action library. These are then integrated to form content evidence units.
[0181] After the initial generation, the system calculates the actual persona proportion vector A1=(0.28, 0.11, 0.34, 0.17, 0.10) and the deviation vector D1=(-0.02, -0.07, 0.10, 0.01, -0.02). Insufficient brand profile and excessive product profile trigger the persona rewriting node. The rewriting node generation instruction is: reduce the stacking of product selling points and increase brand credibility and a gentle, supportive expression. After the second generation, A2=(0.31, 0.17, 0.25, 0.16, 0.11), and all deviations are within the allowable range; compliance review is passed, platform format review is passed, the quality score is 91 points, and the system outputs the final content.
[0182] Example 2: Generation of Science Popularization Content for WeChat Official Accounts
[0183] This embodiment targets the WeChat Official Account platform, with the target content being brand education and product knowledge explanations. The system configuration is T=(0.18, 0.28, 0.30, 0.10, 0.14), ε=0.06. Compared to content community platforms, the WeChat Official Account platform increases the weight of brand and product profiles and requires complete text structure, clear logic, titles not exceeding 64 characters, and text length of 800-1500 words.
[0184] During the initial generation, the platform's format review found that the main text length was only 620 characters, which did not meet the length requirements for WeChat official account main text, triggering a platform format rewrite node. This node added the platform's configured requirements of "main text length range of 800-1500 characters, complete structure, and including subheadings" to the generation prompts, and required the addition of a four-paragraph structure: "Problem Background - Ingredient Explanation - Applicable Scenarios - Selection Suggestions". After rewriting, the main text length was 1020 characters, the platform's format review was passed, and the quality score improved from the initial 78 points to 89 points.
[0185] Example 3: Generation of in-person scripts for short video platforms
[0186] In this embodiment, the target platform is a short video platform, and the target content is the short video title and the accompanying text. The system configuration is T=(0.20, 0.12, 0.20, 0.28, 0.20), ε=0.06. The system increases the proportion of activity scenes and platform style, requiring titles to be no more than 55 characters and body text to be 100-300 characters, using short sentences, strong openings, and interactive guidance structures.
[0187] During the initial generation, the compliance review identified "immediate results" as a risky expression, triggering a compliance rewrite. The system added a negative constraint to "immediate results" and removed the corresponding risky expression from the content evidence unit, replacing it with "more focused on changes in daily experience." After the rewrite, the compliance review was passed, the platform's format review was passed, and the quality score was 87 points.
[0188] VII. Comparative Example
[0189] Comparative Example 1 uses a general generation model to directly generate content, only requiring the input of the prompt "Please generate copy for a certain brand's content platform". It does not use brand corpus graphs, persona proportion vectors, or state diagram closed-loop rewriting.
[0190] Comparative Example 2 uses a fixed template filling method, filling the brand name, product name, selling points and platform fields into a preset template, without performing graph reasoning path fusion and actual character proportion detection.
[0191] Comparative Example 3 uses a knowledge base retrieval-enhanced generation method, which retrieves text from the brand knowledge base and directly inputs it into the generation model, but does not perform target / actual persona ratio vector calculation or bias-driven rewriting.
[0192] The embodiment adopts the solution of this invention. The test task involves generating a total of 300 pieces of content for the same brand on content community platforms, WeChat official account platforms, and short video platforms, with 100 pieces on each platform. Each piece of content is automatically generated by the system and manually reviewed. The content production efficiency, persona consistency, compliance pass rate, platform format first-time pass rate, and average manual modification time are statistically analyzed. The results are shown in Table 5. Valid content refers to content that simultaneously meets the requirements of persona consistency, compliance review, platform format review, and quality score, and can be published without major modifications after manual review.
[0193] Table 5 Comparison of Content Generation Effects of Different Schemes
[0194]
[0195] As shown in Table 5, while Comparative Example 1 offers flexible generation, it lacks structured brand evidence and character proportion correction, resulting in low character consistency and compliance pass rates. Comparative Example 2, although having a more stable platform format, suffers from obvious template-based approach, lacking naturalness and compelling character appeal. Comparative Example 3 improves factual consistency through knowledge base enhancement, but still cannot accurately control the expression proportions of various character dimensions in the content, nor can it perform targeted rewriting based on failure types. In contrast, this invention forms content evidence units through a brand corpus graph, calculates the deviation vector D by comparing the target character proportion vector T with the actual character proportion vector A, and performs targeted rewriting using a state graph workflow, significantly improving content consistency, compliance pass rates, and platform adaptation efficiency.
[0196] The above embodiments are merely illustrative of preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Without departing from the core concept of the present invention, those skilled in the art can deploy the semantic knowledge base, relation graph library, and strategy action library in different types of databases, and can also use different generation models, text classification models, content security models, or platform configuration methods to achieve the same functionality. As long as the technical solution still generates content evidence units through brand corpus graphs, calculates the deviation vector D through the target persona proportion vector T and the actual persona proportion vector A, and performs targeted rewriting and closed-loop output in the state graph workflow based on deviation, compliance, platform format, and quality results, it should fall within the scope of protection of the present invention.
Claims
1. A multi-platform content generation method based on brand corpus graph and persona proportion constraints, characterized in that, Includes the following steps: S1. Obtain basic brand corpus, product corpus, user profile corpus, activity corpus, and platform rule corpus, and construct a brand corpus graph consisting of a semantic knowledge base, a relationship graph library, and a strategy action library; S2. Receive content generation task, generate target persona proportion vector T and allowable deviation ε according to target platform, user tags, brand identity, product identity and activity scenario. The target persona proportion vector T includes at least user profile proportion, brand profile proportion, product profile proportion, activity scenario proportion and platform style proportion. S3. Based on the content generation task, retrieve semantic elements, reasoning paths and strategy actions from the brand corpus graph, integrate the reasoning paths with confidence levels that meet the threshold, the corresponding semantic elements, strategy actions and compliance checkpoints to form a content evidence unit, and compile the content evidence unit, the target persona proportion vector T and the target platform configuration into a generation prompt. S4. Call the generation model to generate candidate content, extract the semantic features of the candidate content and map them to each persona dimension, normalize to obtain the actual persona proportion vector A, calculate the deviation vector D=AT, and perform compliance review, platform format review and quality scoring. S5. When the absolute value of any dimension in the deviation vector D is greater than the allowable deviation ε, or when any of the compliance review, platform format review, or quality score fails to meet the standard, the corresponding rewrite node is routed in the state diagram workflow according to the failure type. The generated prompt or content evidence unit is corrected and regenerated. When the requirements are met, the content adapted to the target platform and the generated traceability record are output.
2. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 1, characterized in that, In step S1, the semantic knowledge base is used to store brand semantic concept records. Each brand semantic concept record includes an industry field, a primary category field, a secondary category field, a keyword field, a synonym field, a standard definition field, an AI reasoning instruction field, an entity type field, a related keyword field, and a risk level field. The relation graph library is used to store relation triples. Each relation triple includes a subject term, a subject type, a relation type, an object term, an object type, a relation category, a confidence level, and an explanatory text. The relation type includes one or more of the following: support relation, causal relation, improvement relation, inclusion relation, applicable object relation, and prohibition relation. The strategy action library is used to store strategy action records. Each strategy action record includes trigger intent, intent keywords, trigger conditions, execution action, action type, response template, association graph reasoning path, and compliance checkpoint.
3. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 1, characterized in that, In step S2, the target persona proportion vector T is represented as: T = (tu, tb, tp, ta, ts); Where tu represents the target percentage of user profile, tb represents the target percentage of brand profile, tp represents the target percentage of product profile, ta represents the target percentage of activity scenario, and ts represents the target percentage of platform style, and tu+tb+tp+ta+ts=1; The allowable deviation ε is the maximum allowable deviation of the actual persona proportion vector A from the target persona proportion vector T in a single dimension. When the target platform is a content community platform, increase the values of tu and ts; when the target platform is a WeChat official account platform, increase the values of tb and tp; when the target platform is a short video platform, increase the values of ta and ts. When content generation tasks contain strong marketing objectives, limit TP (Title, Message, Page) to no higher than the preset product mention limit to avoid over-marketing of candidate content.
4. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 2, characterized in that, In step S3, based on the content generation task, semantic elements, reasoning paths, and strategy actions are retrieved from the brand corpus graph, including: The content generation task is analyzed to obtain the target industry, target product, target user tags, target scenario, and target platform; Retrieve keywords, synonyms, standard definitions, and AI reasoning instructions that match the target product, target user tags, and target scenario from the semantic knowledge base to form a set of candidate semantic elements; In the relation graph database, path retrieval is performed starting from candidate semantic elements to obtain at least one candidate reasoning path consisting of subject words, relation types, and object words; In the strategy action library, corresponding strategy actions are matched according to the target scenario and user tags. The strategy actions include at least one of the following: educational explanation, selling point introduction, scenario resonance, comparative explanation, and interactive guidance. The candidate semantic element set, candidate reasoning path, and policy action are associated to form an intermediate candidate evidence set for generating prompts and compiling.
5. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 4, characterized in that, In step S3, reasoning paths that meet the confidence threshold are merged into content evidence units, including: Candidate inference paths are divided into core paths, supplementary paths, and reference paths according to their confidence levels. The confidence level of the core path is not lower than the first confidence threshold, the confidence level of the supplementary path is lower than the first confidence threshold but not lower than the second confidence threshold, and the confidence level of the reference path is lower than the second confidence threshold. The core path is used as the main source of arguments for the candidate content, the supplementary path is used as an auxiliary source of expression for the candidate content, and the reference path is used only to prompt the model to avoid semantic deviation and is not directly written into the main text. When multiple candidate reasoning paths have conflicting conclusions, the candidate reasoning path with higher confidence is retained first; when the confidence levels are the same, they are selected according to the priority of the relation category; when the conflict still cannot be resolved, the corresponding semantic elements are written into the manual verification mark or the direct referencing prohibition mark. The selected core path, supplementary path, corresponding semantic elements, strategic actions, and compliance checkpoints are combined into content evidence units. Wherein, the first confidence threshold is the lower confidence limit for determining the core path, and the second confidence threshold is the lower confidence limit for determining the supplementary path.
6. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 3, characterized in that, In step S4, semantic features of the candidate content are extracted and mapped to each persona dimension, and normalized to obtain the actual persona proportion vector A, including: The candidate content is segmented into sentences to obtain multiple content fragments; For each content fragment, entity recognition, sentiment recognition, scene word recognition, brand word recognition, product selling point recognition, platform style word recognition, and risk word recognition are performed to obtain a set of semantic features; Based on the preset feature-dimension mapping table, the semantic features in the semantic feature set are mapped to the user profile dimension, brand profile dimension, product profile dimension, activity scenario dimension, and platform style dimension, respectively. Calculate the matching score for each persona dimension, and normalize each matching score to obtain the actual persona proportion vector A; The actual persona proportion vector A is represented as A=(au, ab, ap, aa, as), where au is the actual proportion of user profile, ab is the actual proportion of brand profile, ap is the actual proportion of product profile, aa is the actual proportion of activity scenario, and as is the actual proportion of platform style, and au+ab+ap+aa+as=1.
7. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 6, characterized in that, In step S4, the deviation vector D=AT is calculated, including: Calculate the user profile deviation du=au-tu, brand profile deviation db=ab-tb, product profile deviation dp=ap-tp, activity scenario deviation da=aa-ta, and platform style deviation ds=as-ts respectively; Combine du, db, dp, da, and ds to form the deviation vector D = (du, db, dp, da, ds); When any deviation value is greater than the allowable deviation ε, the persona dimension corresponding to that deviation value is marked as an excessive dimension. When any deviation value is less than -ε, the persona dimension corresponding to that deviation value is marked as an insufficient dimension. Rewrite control instructions are generated based on excessive and insufficient dimensions. The rewrite control instructions include one or more of the following: enhancing the expression of insufficient dimensions, weakening the expression of excessive dimensions, replacing the corresponding semantic elements, and adjusting the corresponding strategy actions.
8. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 7, characterized in that, In step S5, routing to the corresponding rewrite node in the state graph workflow based on the failure type includes: When the failure type is a deviation in the persona proportion, the route is routed to the persona rewriting node. The persona rewriting node determines the excessive and insufficient dimensions based on the deviation vector D, and adjusts the number of semantic elements, expression intensity, or fragment position in the content evidence unit corresponding to each persona dimension. When the failure type is "compliance review not met", the route is routed to the compliance rewrite node. The compliance rewrite node writes the prohibited words, prohibited expressions, prohibited relationships or high-risk semantic elements into the negative constraint segment of the generated prompt, and deletes or replaces the corresponding semantic elements from the content evidence unit. When the failure type is "platform format review not met", the route is routed to the platform format rewriting node. The platform format rewriting node corrects the title length, body text length, paragraph structure, interactive guidance, hot word injection method or disables elements according to the target platform configuration. When the failure type is that the quality score does not meet the standard, the route is routed to the quality rewrite node, which adjusts the generation parameters, content structure requirements or strategy action sequence. If the number of rewrites reaches the preset limit and still fails to meet the target, the request is routed to a manual review node.
9. The multi-platform content generation method based on brand corpus graph and persona proportion constraints according to claim 8, characterized in that, The target platform configuration includes at least the platform identifier, maximum title length, body text length range, content style tags, platform hot word sources, prohibited elements, interactive guidance rules, content structure template, and review rules. When the same content generation task corresponds to multiple target platforms, different target platform configurations are applied based on the same content evidence unit and target persona proportion vector T to generate candidate content for multiple platforms. For each candidate content on the platform, calculate the actual persona proportion vector A and the deviation vector D, and perform compliance review, platform format review and quality scoring respectively; The generated traceability record includes at least the content generation task identifier, target platform, target persona percentage vector T, actual persona percentage vector A, deviation vector D, adopted reasoning path, strategy action, compliance review result, platform format review result, quality score result, number of rewrites, manual review status, and final output content.
10. A multi-platform content generation system based on brand corpus graph and persona proportion constraints, characterized in that, The system for implementing the method according to any one of claims 1 to 9, the system comprising: The corpus graph module is used to build and manage a brand corpus graph consisting of a semantic knowledge base, a relation graph library, and a strategy action library, and to generate semantic elements, inference paths, and strategy actions based on the content. The character constraint module is used to generate the target character proportion vector T and allowable deviation ε, and calculate the actual character proportion vector A and deviation vector D based on the candidate content; The prompt compilation module is used to compile the content evidence unit, the target persona proportion vector T, and the target platform configuration into a generated prompt. The content generation module is used to call the generation model to generate candidate content; The review and scoring module is used to perform compliance review, platform format review, and quality scoring on candidate content; The state diagram workflow module is used to route candidate content to the corresponding rewrite node, manual review node, or result saving node based on the deviation vector D, compliance review results, platform format review results, and quality score results. The multi-platform adaptation module is used to output the final content for the corresponding platform based on the configuration of different target platforms; The traceability and recording module is used to record the content generation task, reasoning path, target persona proportion vector T, actual persona proportion vector A, deviation vector D, review and scoring results, rewriting process, and final output content.
Citation Information
Patent Citations
Personalized copywriting generation method and device, computer equipment and storage medium
CN119515558A
Marketing content generation method and system based on big data
CN121834691A
Digital culture creative content generation system based on artificial intelligence
CN121936570A