Cross-modal alignment type intelligent marketing strategy generation method and system based on strategy intention guidance

By aligning the generated strategy intent vector and multimodal content vector, and combining them with user behavior data, a systematic and precise marketing strategy generation was achieved. This solved the problems of unquantifiable strategies, unalignable content, and uncontrollable pace in existing systems, and generated an executable marketing strategy plan.

CN121981759BActive Publication Date: 2026-07-31GUANGZHOU YUNZHIDACHUANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU YUNZHIDACHUANG TECH CO LTD
Filing Date
2026-04-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing marketing strategy generation systems struggle to integrate brand marketing intent, influencer content characteristics, and user behavior patterns, resulting in strategies that are not quantifiable, content that is not aligned, and pace that is not controllable, making it impossible to generate high-quality and precise marketing strategies.

Method used

By obtaining user marketing task briefings to generate strategy intent vectors, using visual encoding networks and language modeling networks to extract multimodal content vectors, and aligning them with the strategy space through cross-modal dynamic fusion models, a structured deployment plan is generated by combining user behavior data.

Benefits of technology

It enables system-level generation of marketing strategies from text input to executable deployment, solving problems such as rough strategy parsing, fragmented content matching, and lack of unified standards for execution rhythm, and generating high-quality and precise marketing strategy plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981759B_ABST
    Figure CN121981759B_ABST
Patent Text Reader

Abstract

This invention relates to the field of marketing strategy technology, and more particularly to a method and system for generating cross-modal aligned intelligent marketing strategies based on strategy intent guidance. The method involves: acquiring task briefing fields submitted by users when creating marketing tasks on a platform; preprocessing the task briefing fields to generate sentence vectors; inputting the sentence vectors into a strategy encoder to generate strategy intent vectors; constructing image representation vectors, video representation vectors, and text representation vectors, and fusing them into a content representation vector; constructing a fusion matching function; outputting a structure alignment score through the vector matching function; inputting the structure alignment score into a scoring function; outputting a final combination score through the scoring function; sorting and filtering based on the final combination score; and outputting a candidate combination set. The method also includes configuring the platform for each pair of combinations in the candidate combination set, calculating the preferred time period and delivery frequency coefficient after determining the platform configuration, and generating a structured deployment plan based on the combination, platform configuration, preferred time period, and delivery frequency coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marketing strategy technology, and more particularly to a method and system for generating cross-modal aligned intelligent marketing strategies based on strategic intent guidance. Background Technology

[0002] With new media content marketing becoming mainstream, brand promotion on platforms like Douyin and Xiaohongshu has become a crucial growth strategy. However, as the content ecosystem rapidly expands, marketing strategy formulation is evolving from relatively simple influencer selection and material selection into a complex decision-making task relying on multi-source information, cross-modal content, and real-time user behavior. Brand marketing briefs often include multiple elements such as product selling points, expression style, and platform preferences, but their textual expression lacks standardization, making it difficult to systematically analyze and quantify marketing intent. On the content side, influencer creation formats are diverse, encompassing text and images, short videos, and verbal announcements, and the same influencer may present completely different tones in different materials. This makes traditional selection methods relying on tags and historical metrics inaccurate in understanding the fit between content style and marketing objectives.

[0003] Meanwhile, platform users' active behavior exhibits clear temporal and phased characteristics. Users show significant differences in their engagement with the same type of content at different conversion stages. However, existing systems struggle to link these user phase changes with content strategy matching, leading to difficulties in uniformly managing ad placement timing and content presentation. Most existing SaaS marketing tools remain at the level of static data presentation and rule-based filtering, lacking the ability to comprehensively integrate brief intent, influencer content characteristics, user behavior patterns, and platform rhythm, thus failing to achieve an automated closed loop from strategy generation to executable deployment. With the continuous evolution of content formats, increasingly rich platform mechanisms, and rapidly changing user behavior patterns, traditional manual operation models are insufficient to support the generation of high-quality, systematic, and precise marketing strategies. Therefore, there is an urgent need for an intelligent strategy generation method that can uniformly analyze brief intent, construct multimodal content representations, complete task-oriented content matching, and generate executable campaign plans. This would address the problems of unquantifiable strategies, misaligned content, and uncontrollable rhythm in existing technologies. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides a method and system for generating cross-modal aligned intelligent marketing strategies based on strategic intent guidance.

[0005] To achieve the above objectives, this invention proposes a cross-modal aligned intelligent marketing strategy generation method based on strategy intent guidance, comprising: The task briefing field submitted by the user when creating a marketing task on the platform is obtained, and the task briefing field is preprocessed to generate sentence vectors. The sentence vectors are then input into the strategy encoder to generate strategy intent vectors. Sub-vectors related to keywords are extracted from word vectors and projected into the tag space to generate tone supplement vectors. Image representation vectors, video representation vectors, and text representation vectors are extracted from image, video, and text information using a visual coding network, an image frame averaging network, and a language modeling network. These vectors are then fused into a content representation vector using a cross-modal dynamic fusion model. This model dynamically assigns weights to single-modal features using a weighting function, introduces a vector constructor with policy offset compensation to align modal features with the policy space, and incorporates a tone-preserving regularization term. This term ensures that the image style does not deviate from the tone complement vector in the policy objective. The tone-preserving regularization term includes features such as the image tone extraction projection matrix, image representation vector, tone complement vector, tone dimension, L2 norm, and regularization coefficients. A vector matching function is constructed based on the strategy intent vector and content representation vector. The vector matching function outputs a structure alignment score. The structure alignment score is input into a scoring function with a key dimension bias penalty. The scoring function outputs the final combined score. At the same time, a content diversity balance term based on the maximum mean difference is introduced. The final combined score and the content diversity balance term are used for sorting and filtering, and a candidate combination set is output. For each pair of candidate combinations, a platform is configured for calculation. After the platform configuration is determined, based on the user's full-cycle time-series active behavior data, the maximum active period of the target audience is located by the sliding time window clustering algorithm as the preferred time period for combination delivery. Then, the personalized delivery frequency coefficient is calculated by the frequency control function that integrates combination scores and cross-vector deviation distance. A structured deployment plan is generated based on the combination, platform configuration, preferred time period, and delivery frequency coefficient.

[0006] In some embodiments, the preprocessing of the task briefing fields to generate sentence vectors specifically includes: The task summary fields are standardized in terms of characters and format, and common punctuation, abnormal and redundant tags and platform template sentences are removed. Dictionary-based error correction and unified replacement are completed to generate standard field text. Standard field text is input into a language coding model based on a transformer structure. The context vector representation is extracted through the language coding model to generate sentence vectors.

[0007] In some embodiments, the visual coding network includes a convolutional layer, which includes a convolutional kernel, a normalization module, and a max pooling operation.

[0008] In some embodiments, the expression for the vector constructor is: ; in Content representation vector, For content mapping matrix, For modal fusion vectors, For the strategy intent vector, The output dimension of the fused modality; This is the mapping matrix from mode to policy space. This is a constant term.

[0009] In some embodiments, the vector matching function includes the dimensional components of the policy intent vector, the dimensional components corresponding to the content representation vector, and the weight ratios of the task policy factors, wherein the expression of the vector matching function is: ; in, For structural alignment scores, The components of the corresponding dimension of the strategy intent vector. The dimension components corresponding to the content representation vector. This represents the weighting of the task strategy factor.

[0010] In some embodiments, the scoring function includes a penalty coefficient, a dimensional gating flag, a policy intent vector, and a content representation vector, wherein the expression of the scoring function is: ; in, Score the final combination. For structural alignment scores, The penalty coefficient is... As a dimensional gating indicator, The components of the corresponding dimension of the strategy intent vector. The dimension component is the one that represents the content vector.

[0011] In some embodiments, the candidate combination set includes a final combination score, a content diversity balance item, a content ID, and an influencer ID.

[0012] In some embodiments, the content diversity balancing factor includes the number of candidate content items and diversity control parameters.

[0013] In some embodiments, the frequency control function includes the final combined score, the absolute deviation distance between the strategy intent vector and the content representation vector, the adjustment term coefficient, and the user's activity level in each time period, wherein the expression of the frequency control function is: ; in This is the final delivery frequency coefficient. For strategy matching score, The absolute deviation distance between the strategy intent vector and the content representation vector. For adjustment term coefficients, This represents the user's activity level at different times.

[0014] To achieve the above objectives, another aspect of the present invention proposes a cross-modal aligned intelligent marketing strategy generation system based on strategy intent guidance, comprising: The strategy intent vector construction module is used to obtain the task briefing field submitted by users when creating marketing tasks on the platform, preprocess the task briefing field to generate sentence vectors, input the sentence vectors into the strategy encoder to generate strategy intent vectors, extract sub-vectors related to keywords from word vectors, and perform tag space projection to generate tone supplement vectors. The content representation vector construction module is used to extract image representation vectors, video representation vectors, and text representation vectors from image, video, and text information through a visual encoding network, an image frame averaging network, and a language modeling network. It then fuses these vectors into a content representation vector using a cross-modal dynamic fusion model. This model dynamically assigns weights to single-modal features using a weighting function, introduces a vector constructor with policy offset compensation to align modal features with the policy space, and adds a tone-preserving regularization term. This term ensures that the image style does not deviate from the tone complement vector in the policy objective. The tone-preserving regularization term includes features such as the image tone extraction projection matrix, image representation vector, tone complement vector, tone dimension, L2 norm, and regularization coefficients. The vector fusion and combination module is used to construct a vector matching function based on the strategy intent vector and the content representation vector. The vector matching function outputs a structure alignment score, which is then input into a scoring function with a key dimension bias penalty. The scoring function outputs the final combination score. At the same time, a content diversity balance term based on the maximum mean difference is introduced. The final combination score and the content diversity balance term are used for sorting and filtering, and a candidate combination set is output. The deployment plan generation module is used to configure the computing platform for each pair of combinations in the candidate combination set. After the platform configuration is determined, based on the user's full-cycle time-series active behavior data, the maximum active segment of the target audience is located by the sliding time window clustering algorithm as the preferred time period for combination delivery. Then, the personalized delivery frequency coefficient is calculated by the frequency control function that integrates combination score and cross-vector deviation distance. A structured deployment plan is generated based on the combination, platform configuration, preferred time period and delivery frequency coefficient.

[0015] The beneficial effects of this invention are as follows: This invention first transforms unstructured brief text into a strategic intent vector containing tone, audience, expression, platform preference, and conversion goals to address the difficulty in standardizing marketing intent. Second, it utilizes visual encoding networks, video frame modeling networks, and text semantic extraction networks to construct multimodal content vectors of images, videos, and text. Through a strategy-guided structural offset mechanism and tone alignment mechanism, the content vector and strategy vector are completely unified in a six-dimensional structural space. Subsequently, a scoring structure combining weighted structural similarity, key dimension deviation gating, and content diversity control is employed to achieve task-oriented influencer content combination screening, transforming content selection from "similar content recommendation" to "strategy-driven matching." Finally, by combining content scoring, strategic intent vectors, and user activity behavior, a strategy execution structure including platform selection, time period configuration, and delivery frequency is generated, transforming the screening results into a practically executable deployment plan. This invention establishes a unified strategy representation space and multimodal content space, and designs a task-oriented cross-modal alignment mechanism, enabling strategy intent, content expression, and user behavior to work collaboratively in the same logical link. This overcomes the problems of coarse strategy parsing, fragmented content matching, and lack of unified standards for execution rhythm in existing technologies, and realizes the system-level generation capability of marketing strategies from text input to executable schedules. Attached Figure Description

[0016] Figure 1 This is a flowchart of a cross-modal aligned intelligent marketing strategy generation method based on strategy intent guidance in a specific embodiment of the present invention; Figure 2 This is a block diagram of a cross-modal aligned intelligent marketing strategy generation system based on strategy intent guidance, as described in a specific embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 As shown, the present invention relates to a method for generating cross-modal aligned intelligent marketing strategies based on strategy intent guidance, comprising: S1: Obtain the task summary field submitted by the user when creating a marketing task on the platform, preprocess the task summary field to generate sentence vectors, input the sentence vectors into the strategy encoder to generate strategy intent vectors, extract sub-vectors related to keywords from word vectors, and perform tag space projection to generate tone supplement vectors, specifically including: This step converts the text briefing information input in the marketing task into a structured strategy intent vector, serving as the foundation for subsequent multimodal content matching and strategy generation. The input data originates from the task briefing fields submitted by users when creating marketing tasks on the platform, passed through web forms or API interfaces, and stored uniformly in raw string format, denoted as [vector name missing]. Its content typically includes the name of the product being promoted, the target audience, the desired content style, reference content, and platform requirements, for example: "Promoting a brand of wedding candy, targeting unmarried women born after 1995, the content should emphasize a high-end and modern feel, preferably in a Vlog style, to create a seeding effect on a certain platform." After receiving the text, the system first performs character standardization and format cleaning, removing common punctuation errors, redundant tags, platform template sentences, etc., and performs dictionary-based error correction and unified replacement (such as merging "vlog" and "Vlog") to ensure semantic input consistency. The cleaned text is then fed into a language coding model based on a transformer structure to extract its context vector representation. This model adopts a bidirectional structure, consisting of 12 stacked layers, each containing a self-attention mechanism and a feedforward module, and outputs a fixed-dimensional sentence vector, denoted as . ,in .

[0019] To extract six core factors for strategy judgment: content tone, audience profile, expression form, style, platform preference, and conversion goal, sentence vectors are used. The input is a policy encoder consisting of two linear transformation layers. The first layer is used to compress dimensionality, and after activation, it enters the second layer structure to map to a six-dimensional output space. The process is expressed as follows: ; in , , This represents the activation function. and For the corresponding bias term, This is the intermediate dimension. Output the policy intent vector. The six components represent the six types of policy information mentioned above. Each component ranges from [0,1], indicating the importance or preference intensity of the policy item in the current task.

[0020] The accuracy of the tone dimension is particularly crucial. To enhance its compatibility with text, images, and video content, a tone subspace projection mechanism is further introduced. The system maintains a set of static word vector matrices containing ten tone tags. Each row corresponds to a word vector with a tone tag (such as "high-end", "fresh", "authentic", "everyday life", etc.), and the dimension is... .from Extract sub-vectors that are context-dependent with keywords such as "style", "aesthetics", and "tone". Perform label space projection to generate tone complement vectors: ; In this formula, This indicates the similarity calculation between the tone dictionary and the text tone vector. Normalization is performed to generate the label weight distribution, and then a new tone vector is obtained through weighted summation. This vector is used for replacement. The components at mid-tone positions are used as structural references for alignment with image style vectors.

[0021] For example, when mentioning "natural," "sophisticated," and "avoiding overly commercial content" in a user task briefing, The focus will be on activating vectors tagged with "advanced" and "lifestyle". This will ultimately construct a strategic intent vector. It not only possesses a semantic distribution structure across six policy dimensions, but its tonal component also exhibits cross-modal comparability. This enables the system to achieve semantic mapping with policy intent when subsequently processing image composition, filter colors, video editing rhythm, and other content. The final output... The vector will be used as the representation of multimodal content in the next step. The matching criteria directly influence the process of content selection, influencer pairing, and platform distribution strategy generation.

[0022] S2: Image representation vectors, video representation vectors, and text representation vectors are extracted from image, video, and text information through a visual encoding network, an image frame averaging network, and a language modeling network. These vectors are then fused into a content representation vector using a cross-modal dynamic fusion model. This model dynamically assigns weights to single-modal features using a weighting function, introduces a vector constructor with policy offset compensation to align modal features with the policy space, and incorporates a tone-preserving regularization term. This term ensures that the image style does not deviate from the tone complement vector in the policy objective. The tone-preserving regularization term includes features such as the image tone extraction projection matrix, image representation vector, tone complement vector, tone dimension, L2 norm, and regularization coefficients. Specifically, it includes: This step aims to integrate the images, videos, and text information in influencer content into a unified content representation vector. Used to compare with the strategy intent vector generated in the previous step Structured semantic alignment is performed. This vector construction not only needs to consider the general representation of multimodal content, but also must actively adapt to the tone, form, and target matching of the marketing brief given in the invention scenario. This requirement determines that this step is not only a multimodal representation extraction process, but also a content vector construction mechanism constrained by task objectives and tone; its core design is "target-driven multimodal semantic projection modeling." The input content is historical content data published by platform influencers, which has been accessed and stored in a structured database through an interface in the system. Each piece of content has a unique identifier, represented by an image. Video frame sequence With text description The components correspond to the cover image / text, video clip sequence, note title, and descriptive text, respectively. As a requirement for connecting with the previous step, this step introduces strategy vectors. As a priori control, the content representation vector is required. Ultimately, in terms of tone, form, platform preference, and other dimensions, it is consistent with The structure remains consistent. In particular, the tone subvectors... After a specialized projection is constructed, the tonal modeling of the image modality must be directly applied in this step.

[0023] First, image modal features are extracted using a visual coding network based on a convolutional structure. This network consists of five convolutional layers, each containing... Convolution kernels, normalization modules, and max pooling operations ultimately output a set of dimensions. Image representation vector Video modality processing employs an image frame averaging strategy network, which extracts one frame every two seconds, feeds it into the averaging strategy network for processing, and then performs temporal mean pooling to obtain the video representation. The text modality is encoded using the same language modeling network as in the previous step, resulting in a dimension of [dimensional value missing]. Text representation vector These three types of vectors will form the modal basic feature set. Next, a cross-modal dynamic fusion model will be used to unify the three types of vectors into a content representation vector. .

[0024] The key innovation of modality fusion lies in the introduction of a task-aware term, ensuring that the fused content vector not only retains multimodal features but also aligns structurally with the policy objective. First, a set of static weights is used to weight the modalities; the specific weighting function is as follows: ; in This is the weighted modal fusion vector. These are modality fusion coefficients, pre-defined based on content type or obtained through data-driven learning. Based on this, enhancements and strategy structures are implemented. To achieve matching, a vector constructor is introduced, which in turn introduces a structure offset term. ,in Let be the mapping matrix from mode to policy space, where the expression for the vector constructor is: ; in The final content representation vector, For content mapping matrix, The output dimension of the fused modality; and Same structure, specifically designed for constructing "strategy-guided content representation"; The constant term used to adjust the guiding strength of this structure is typically set to a value of [value to be filled in]. The introduction of this structure breaks through the passive characteristic of "content self-representation" in traditional modal fusion methods, transforming it into "task-guided content expression," reflecting the essential characteristics of the task-driven generation mechanism in the scenario of this invention.

[0025] Considering the crucial role of tone matching in campaign conversion, a tone preservation regularization term is further introduced to ensure that the image style does not deviate from the tone preservation regularization term in the strategy objectives. Specifically: ; in Extract the projection matrix for image tonality, and convert the image vector... Mapped to the tonal space; The tone vector generated in the previous step, As for the tonality dimension, Represents the L2 norm; This is the regularization coefficient, usually set to 0.1. This item is used to constrain the image's tonal expression to closely match the policy objective, preventing high-quality but stylistically incompatible content from being selected.

[0026] S3: Construct a vector matching function based on the strategy intent vector and content representation vector. The vector matching function outputs a structure alignment score. This score is then input into a scoring function with a penalty for key dimension bias. The scoring function outputs the final combined score. Simultaneously, a content diversity balancing term based on the maximum mean difference is introduced. The final combined score and the content diversity balancing term are used for sorting and filtering, and a candidate combination set is output, specifically including: This step, based on the output of the previous two steps, completes the policy intent vector. With content representation vector The semantic matching and combination generation aims to select the most relevant influencer content pairs within a task-driven semantic space, ultimately forming a structurally consistent and goal-oriented combination of deployable materials. In the entire invention process, this step plays a crucial bridging role in "converting intent into actual candidate content." It is not a general recommendation module, but rather a targeted filter strictly serving the task strategy, particularly demonstrating the invention's unique strategy consistency control requirements in its tone consistency and structural deviation penalty mechanisms.

[0027] The input for this step includes the content representation vector generated in the previous step. and the strategy intent vector generated in the first step Both vectors reside in a six-dimensional space of structural alignment, where each dimension corresponds to six factors: content tone, target audience, content format, expression style, platform preference, and conversion goal. During the matching process, two dimensions must be considered simultaneously: firstly, the overall similarity of content and strategy in numerical structure; and secondly, the degree of difference in key strategic points (such as tone and platform) within specific dimensions. Therefore, a comprehensive scoring structure integrating structural alignment and dimensional difference control is constructed.

[0028] Consider the strategy intent vector With content representation vector In the matching degree of the six-dimensional semantic space, the policy weight vector is first defined. This vector reflects the priority of different strategy factors during task execution and can be automatically adjusted based on whether a strategy objective is explicitly mentioned in the brief. For example, if the brief explicitly states "it needs to be launched on a certain platform," the platform dimension weight will be significantly increased. Based on this, a vector matching function is defined to output a structure alignment score, where the structure alignment score... For weighted cosine similarity, the specific vector matching function is expressed as follows: ; in, and For the components of the corresponding dimension, This represents the weighting of the task strategy factors. This similarity term measures the directional consistency of the two strategies in their overall structure and is the basic matching score. To avoid ignoring deviations in key strategy points based solely on the overall structure, a scoring function is designed, incorporating a gating mechanism for dimensional penalty terms. This is used to adjust the overall score when there are significant deviations in certain strategy dimensions. The specific expression for the scoring function is as follows: ; in, Score the final combination. The penalty coefficient controls the overall tolerance for deviation and is typically set at... between; For dimensional gating flags, when the first When the dimension is a policy item explicitly mentioned in the Brief, take Otherwise take This innovative structure enables "explicit task goal priority control": if the brief only states "high-level tone", then even if the content deviates slightly in other dimensions, it will not be penalized; however, if the tone deviates, it will be penalized, thus lowering the overall score.

[0029] Furthermore, to further improve the stability of the system's strategy expression and prevent score fluctuations due to overfitting of content samples, this step also introduces a content diversity balancing term during the calculation process. Considering that some influencers may have multiple content samples matching the strategy goals simultaneously, which could easily lead to content concentration and insufficient coverage, the diversity regularization term is defined as follows: ; in, This represents the number of candidate content pieces from the same influencer selected under the current strategy. This is a parameter for diversity control. When multiple pieces of content come from the same influencer, increasing the regularization value will cause the system to automatically lower the score of additional content from that influencer in the ranking, thereby increasing the exposure opportunities for other influencers. This achieves broad coverage of material supply and helps avoid the risk of "single hit content".

[0030] Based on the above scoring structure and regularization control mechanism, the system scores each candidate content-expert pair in the matching phase. and combined with regular terms Sort and filter, and finally output. The highest-rated content in the group - the expert combination, is recorded as follows: Each set of results includes information such as structural score, key dimension deviation value, content ID and influencer ID.

[0031] S4: Configure the platform for each pair of combinations in the candidate combination set. After determining the platform configuration, based on the user's full-cycle time-series active behavior data, use a sliding time window clustering algorithm to locate the maximum active segment of the target audience as the preferred time period for combination delivery. Then, calculate the personalized delivery frequency coefficient by integrating combination scores and cross-vector deviation distance frequency control functions. Generate a structured deployment plan based on the combination, platform configuration, preferred time period, and delivery frequency coefficient, specifically including: This step is based on the candidate combination set output from the previous stage. content With experts Matching pairs and their strategy matching scores , combined with the strategic intention vector extracted from the marketing briefing and the platform audience behavior data, construct a structured strategic plan that can be directly deployed and executed, including the configuration results of the content delivery platform, time period, and delivery frequency. The core task of this step is to convert the "content-influencer combination with the highest matching score" into scheduling information on "when, where, and with what intensity" to spread, ensuring that the system can not only understand and match tasks, but also ultimately implement them into automated delivery actions.

[0032] The input data mainly includes two parts: one is the combination list output from the previous step , where is the structure matching score, is the specific content vector, is the corresponding influencer identifier, and the content comes from the multi-modal fusion vector generated in S2 ; the other is the user behavior statistical data of the platform, which is used to construct the activity distribution of the target population at different time periods, denoted as . The three dimensions respectively represent the average active proportion of the target population in the morning (8:00–11:00), afternoon (12:00–17:00), and evening (18:00–22:00), which is derived from the windowed statistics of the user browsing, liking, forwarding and other behavior logs of the system in the past 90 days.

[0033] In addition, the platform preference dimension in the strategic intention vector is also required. This value comes from the structured briefing information extracted in the first step and is usually a weight distribution with a length of 2, corresponding to the priorities of "Platform 1" and "Platform 2" for delivery respectively. For example, when it is clearly stated in the briefing that "mainly invest in Platform 2", the system will adjust the weight value of "Platform 2" in this dimension to be higher than that of "Platform 1". These information form the basis for scheduling the platform and time arrangements in this step.

[0034] The system first calculates the target platform configuration for each pair . The platform preference vector is denoted as , where represents the weight of "Platform 1", represents the weight of "Platform 2". At the same time, the content source platform is identified by the field , where indicates that the content comes from "Platform 1", indicates that it comes from "Platform 2". The final platform configuration is determined by the following formula: ; where For indicator functions, when The time value is Otherwise The platform's selection mechanism is based on practical engineering experience, avoiding the risk of content style mismatch caused by inconsistent content delivery across platforms. For example, a product recommendation post with a primarily text-and-image style published on Platform 2 will be prioritized for execution on Platform 2.

[0035] After the platform is selected, the system needs to determine the time period for each piece of content to be displayed. With delivery frequency coefficient Time period Choose from the three types of windows: Indicates morning. Indicates noon. This indicates evening. It is based on user activity levels at different times. (in (As indicated by the aforementioned platform number), the system selects the most active period as the preferred time slot for ad delivery, i.e.: ; For example, when the target users are on the platform during the evening hours The system will automatically select the person with the highest activity level. corresponding time period Set to 3 to indicate evening delivery.

[0036] A more innovative design lies in the system's control over delivery frequency. Each piece of content should not be repeatedly pushed indefinitely simply because of a high matching score; instead, frequency compression should be implemented based on strategy deviations. Therefore, the system introduces the following frequency control function: ; in This is the final delivery frequency coefficient. For strategy matching score, This represents the absolute deviation between the strategic intent and the content structure. It is the adjustment term coefficient (generally taken as...). to This frequency function is used to encourage the selection of content with higher structural relevance, thereby increasing its exposure probability. It considers both the target user's behavioral window and policy consistency as an important moderating factor, thus achieving a balance between content personalization and policy control.

[0037] The final system generates a structured deployment plan. Each record corresponds to a complete deployment configuration for a single influencer's content combination, including platform, time period, and frequency. All configuration outputs are in JSON format, which can be directly parsed by the deployment system for content scheduling.

[0038] Further reference Figure 2 As shown, to achieve the above objectives, another aspect of this invention proposes a cross-modal aligned intelligent marketing strategy generation system based on strategy intent guidance, comprising: The strategy intent vector construction module is used to obtain the task briefing field submitted by users when creating marketing tasks on the platform, preprocess the task briefing field to generate sentence vectors, input the sentence vectors into the strategy encoder to generate strategy intent vectors, extract sub-vectors related to keywords from word vectors, and perform tag space projection to generate tone supplement vectors. The content representation vector construction module is used to extract image representation vectors, video representation vectors, and text representation vectors from image, video, and text information through a visual encoding network, an image frame averaging network, and a language modeling network. It then fuses these vectors into a content representation vector using a cross-modal dynamic fusion model. This model dynamically assigns weights to single-modal features using a weighting function, introduces a vector constructor with policy offset compensation to align modal features with the policy space, and adds a tone-preserving regularization term. This term ensures that the image style does not deviate from the tone complement vector in the policy objective. The tone-preserving regularization term includes features such as the image tone extraction projection matrix, image representation vector, tone complement vector, tone dimension, L2 norm, and regularization coefficients. The vector fusion and combination module is used to construct a vector matching function based on the strategy intent vector and the content representation vector. The vector matching function outputs a structure alignment score, which is then input into a scoring function with a key dimension bias penalty. The scoring function outputs the final combination score. At the same time, a content diversity balance term based on the maximum mean difference is introduced. The final combination score and the content diversity balance term are used for sorting and filtering, and a candidate combination set is output. The deployment plan generation module is used to configure the computing platform for each pair of combinations in the candidate combination set. After the platform configuration is determined, based on the user's full-cycle time-series active behavior data, the maximum active segment of the target audience is located by the sliding time window clustering algorithm as the preferred time period for combination delivery. Then, the personalized delivery frequency coefficient is calculated by the frequency control function that integrates combination score and cross-vector deviation distance. A structured deployment plan is generated based on the combination, platform configuration, preferred time period and delivery frequency coefficient.

[0039] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for generating cross-modal aligned intelligent marketing strategies based on strategy intent guidance, characterized in that: include: The task briefing field submitted by the user when creating a marketing task on the platform is obtained, and the task briefing field is preprocessed to generate sentence vectors. The sentence vectors are then input into the strategy encoder to generate strategy intent vectors. Sub-vectors related to keywords are extracted from word vectors and projected into the tag space to generate tone supplement vectors. Image representation vectors, video representation vectors, and text representation vectors are extracted from image, video, and text information using a visual coding network, an image frame averaging network, and a language modeling network. These vectors are then fused into a content representation vector using a cross-modal dynamic fusion model. This model dynamically assigns weights to single-modal features using a weighting function, introduces a vector constructor with policy offset compensation to align modal features with the policy space, and incorporates a tone-preserving regularization term. This term ensures that the image style does not deviate from the tone complement vector in the policy objective. The tone-preserving regularization term includes features such as the image tone extraction projection matrix, image representation vector, tone complement vector, tone dimension, L2 norm, and regularization coefficients. A vector matching function is constructed based on the strategy intent vector and content representation vector. The vector matching function outputs a structure alignment score. The structure alignment score is input into a scoring function with a key dimension bias penalty. The scoring function outputs the final combined score. At the same time, a content diversity balance term based on the maximum mean difference is introduced. The final combined score and the content diversity balance term are used for sorting and filtering, and a candidate combination set is output. For each pair of candidate combinations, a platform is configured for calculation. After the platform configuration is determined, based on the user's full-cycle time-series active behavior data, the maximum active period of the target audience is located by the sliding time window clustering algorithm as the preferred time period for combination delivery. Then, the personalized delivery frequency coefficient is calculated by the frequency control function that integrates combination scores and cross-vector deviation distance. A structured deployment plan is generated based on the combination, platform configuration, preferred time period, and delivery frequency coefficient.

2. The policy intent-based guided cross-modal alignment type intelligent marketing strategy generation method according to claim 1, characterized in that, The preprocessing of the task briefing fields to generate sentence vectors specifically includes: The task summary fields are standardized in terms of characters and format, and common punctuation, abnormal and redundant tags and platform template sentences are removed. Dictionary-based error correction and unified replacement are completed to generate standard field text. Standard field text is input into a language coding model based on a transformer structure. The context vector representation is extracted through the language coding model to generate sentence vectors. 3.The policy intent-based cross-modal alignment intelligent marketing strategy generation method according to claim 1, characterized in that, The visual coding network includes convolutional layers, which in turn include convolutional kernels, normalization modules, and max pooling operations. 4.The policy intent-based cross-modal alignment intelligent marketing strategy generation method according to claim 1, characterized in that, The expression for the vector constructor is: ; in Content representation vector, For content mapping matrix, For modal fusion vectors, For the strategy intent vector, The output dimension of the fused modality; This is the mapping matrix from mode to policy space. This is a constant term. 5.The policy intent-based cross-modal alignment intelligent marketing strategy generation method according to claim 1, characterized in that, The vector matching function includes the dimensional components of the policy intent vector, the dimensional components corresponding to the content representation vector, and the weight ratios of the task policy factors. The expression of the vector matching function is as follows: ; in, For structural alignment scores, The components of the corresponding dimension of the strategy intent vector. The dimension components corresponding to the content representation vector. This represents the weighting of the task strategy factor. 6.The policy intent-based cross-modal alignment intelligent marketing strategy generation method according to claim 1, wherein, The scoring function includes a penalty coefficient, a dimensional gating flag, a policy intent vector, and a content representation vector, wherein the expression of the scoring function is: ; wherein, is the final combined score, is the structure alignment score, is the penalty coefficient, is the dimension level gating flag, is the component of the policy intent vector corresponding to the dimension, is the component of the content representation vector corresponding to the dimension. 7.The policy intent-based cross-modal alignment intelligent marketing strategy generation method according to claim 1, characterized in that, The candidate combination set includes the final combination score, content diversity balance item, content ID, and influencer ID. 8.The policy intent-based cross-modal alignment intelligent marketing strategy generation method according to claim 1, wherein, The content diversity balancing factor includes the number of candidate content items and diversity control parameters. 9.The policy intent-based cross-modal alignment intelligent marketing strategy generation method according to claim 1, wherein, The frequency control function includes the final combined score, the absolute deviation distance between the strategy intent vector and the content representation vector, the adjustment coefficient, and the user's activity level in each time period. The expression for the frequency control function is as follows: ; wherein is a final drop frequency coefficient, is a strategy matching score, is an absolute deviation distance between a strategy intent vector and a content representation vector, is an adjustment term coefficient, is a user's activity level in each time period.

10. A cross-modal alignment-based intelligent marketing strategy generation system based on a policy intention guide, characterized in that, include: The strategy intent vector construction module is used to obtain the task briefing field submitted by users when creating marketing tasks on the platform, preprocess the task briefing field to generate sentence vectors, input the sentence vectors into the strategy encoder to generate strategy intent vectors, extract sub-vectors related to keywords from word vectors, and perform tag space projection to generate tone supplement vectors. The content representation vector construction module is used to extract image representation vectors, video representation vectors, and text representation vectors from image, video, and text information through a visual encoding network, an image frame averaging network, and a language modeling network. It then fuses these vectors into a content representation vector using a cross-modal dynamic fusion model. This model dynamically assigns weights to single-modal features using a weighting function, introduces a vector constructor with policy offset compensation to align modal features with the policy space, and adds a tone-preserving regularization term. This term ensures that the image style does not deviate from the tone complement vector in the policy objective. The tone-preserving regularization term includes features such as the image tone extraction projection matrix, image representation vector, tone complement vector, tone dimension, L2 norm, and regularization coefficients. The vector fusion and combination module is used to construct a vector matching function based on the strategy intent vector and the content representation vector. The vector matching function outputs a structure alignment score, which is then input into a scoring function with a key dimension bias penalty. The scoring function outputs the final combination score. At the same time, a content diversity balance term based on the maximum mean difference is introduced. The final combination score and the content diversity balance term are used for sorting and filtering, and a candidate combination set is output. The deployment plan generation module is used to configure the computing platform for each pair of combinations in the candidate combination set. After the platform configuration is determined, based on the user's full-cycle time-series active behavior data, the maximum active segment of the target audience is located by the sliding time window clustering algorithm as the preferred time period for combination delivery. Then, the personalized delivery frequency coefficient is calculated by the frequency control function that integrates combination score and cross-vector deviation distance. A structured deployment plan is generated based on the combination, platform configuration, preferred time period and delivery frequency coefficient.