Content marketing effect evaluation and optimization method and system based on deep learning

By building a deep learning dual-tower neural network model, the shortcomings of traditional content marketing evaluation and optimization methods are solved, accurate matching of content and scenarios and precise allocation of resources are achieved, and marketing effectiveness and return on investment are improved.

CN120450744BActive Publication Date: 2025-10-14BEIJING GALAXY GRAVITY TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510535940.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-10-14
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

Traditional content marketing effectiveness evaluation methods cannot fully and accurately measure the actual influence and value of content. Existing optimization technologies are difficult to adapt to the rapidly changing market environment and user needs. Resource allocation strategies are rigid and cannot achieve optimal resource utilization.

Method used

A dual-tower neural network model based on deep learning is constructed. The content tower and scenario tower are used to process data to generate a representation set. The matching score is calculated in combination with the interaction layer. A resource allocation algorithm with nonlinear mapping and progressive two-stage optimization is designed. Iterative optimization is performed to determine the optimal matching combination and resource input ratio.

Benefits of technology

It improves the accuracy and reliability of content marketing effectiveness evaluation, significantly enhances the generalization and robustness of the model, achieves precise allocation of marketing resources, and significantly improves the overall effect and return on investment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450744B_ABST
    Figure CN120450744B_ABST
Patent Text Reader

Abstract

The application provides a content marketing effect evaluation and optimization method and system based on deep learning, relates to the technical field of artificial intelligence, and comprises the following steps: obtaining historical data of content marketing; constructing a double-tower deep neural network model, processing the content data through a content tower to generate a content representation set, processing the scene data through a scene tower to generate a scene representation set, and calculating the matching score of the content representation set and the scene representation set through an interaction layer; the double-tower deep neural network model is trained through a contrast learning framework of hierarchical scene perception and difficult sample mining, and is iteratively optimized in combination with a content-scene cross-attention collaborative representation mechanism; based on the matching score, a resource allocation algorithm of non-linear mapping and gradual two-stage optimization is designed to generate a resource allocation optimization scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence technology, and in particular to a method and system for evaluating and optimizing content marketing effectiveness based on deep learning. Background Art

[0002] Content marketing, as a key digital marketing strategy, attracts and retains target audiences by creating and sharing valuable, relevant, and consistent content, ultimately driving customer behavior. Traditional content marketing effectiveness evaluation methods rely primarily on simple metrics such as click-through rate and pageviews, which fail to fully and accurately measure the actual impact and value of content. This approach overlooks the complex relationship between content and specific scenarios and fails to capture the differences in content performance across different contexts.

[0003] Existing content optimization technologies often rely on static rules or simple machine learning models, making them difficult to adapt to rapidly changing market environments and user needs. These methods lack a deep understanding of the connections between content and context, resulting in limited optimization effectiveness.

[0004] Resource allocation strategies play a key role in content marketing, but existing methods typically employ linear or fixed-ratio allocations, failing to adapt flexibly to the dynamic matching relationship between content and context. This rigid resource allocation approach makes it difficult to achieve optimal resource utilization, hindering overall marketing effectiveness. Summary of the Invention

[0005] The embodiments of the present invention provide a content marketing effectiveness evaluation and optimization method and system based on deep learning, which can solve the problems in the existing technology.

[0006] According to a first aspect of the embodiments of the present invention,

[0007] Provides content marketing effectiveness evaluation and optimization methods based on deep learning, including:

[0008] Obtain historical data on content marketing, including content data and scenario data;

[0009] Constructing a dual-tower deep neural network model, processing the content data through the content tower to generate a content representation set, processing the scene data through the scene tower to generate a scene representation set, and calculating the matching score between the content representation set and the scene representation set through the interaction layer; the dual-tower deep neural network model is trained using a contrastive learning framework of hierarchical scene perception and difficult sample mining, and is iteratively optimized in combination with a content-scene cross-attention collaborative representation mechanism;

[0010] Based on the matching score, a resource allocation algorithm of non-linear mapping and gradual two-stage optimization is designed, and the optimal matching combination of content and scene and the resource input proportion of each matching combination are determined through iterative optimization to generate a resource allocation optimization scheme;

[0011] The resource allocation optimization scheme is executed, and feedback data is collected to update the double-tower deep neural network model.

[0012] In an optional implementation,

[0013] A double-tower deep neural network model is constructed, a content representation set is generated by processing the content data through a content tower, a scene representation set is generated by processing the scene data through a scene tower, and a matching score of the content representation set and the scene representation set is calculated through an interaction layer, including:

[0014] The content tower adopts a multi-head self-attention mechanism to process the content data, and an attraction score of the content element is calculated through attention pooling, the content elements are weighted and summed based on the attraction score, and a content representation set is generated;

[0015] The scene data is processed through the scene tower, the scene tower encodes the launch time, channel characteristics and target audience characteristics respectively, and captures the mutual influence between scene factors through a self-attention mechanism to generate a scene representation set;

[0016] In the interaction layer, a multi-level interaction matrix of content and scene is calculated, and a hierarchical specialization enhancement operation is applied: the high layer uses an attention mechanism to extract semantic alignment, the middle layer uses a bilinear mapping to capture nonlinear relationship, and the bottom layer uses a local convolution to match the pattern; The hierarchical weight is calculated through a context-based evaluation function, and the matching score is generated after weighted fusion through an information bottleneck network and a dimension reduction network;

[0017] Based on the matching score, the content-scene combination is sorted; the matching score, the content global representation and the scene global representation of the double-tower deep neural network model are output as the input of the downstream resource allocation optimization task.

[0018] In an optional implementation,

[0019] In the interaction layer, the matching score is calculated, including:

[0020] Each layer of representation in the content representation set and each layer of representation in the scene representation set are interactively calculated to obtain a multi-level interaction matrix; a hierarchical specialized feature enhancement operation is applied to the multi-level interaction matrix, wherein the high layer semantic representation adopts an attention mechanism to calculate the semantic alignment degree, the middle layer relationship representation adopts a bilinear mapping to capture the nonlinear relationship, and the bottom layer feature representation adopts a local convolution operation to capture the local pattern matching, to generate an enhanced interaction matrix set;

[0021] construct a hierarchical importance evaluation function based on the content representation, the scene representation and the marketing context information, and calculate the importance weight of each enhanced interaction matrix in the enhanced interaction matrix set through an attention network;

[0022] perform dynamic weighted fusion on the enhanced interaction matrix set according to the importance weight to obtain a fused interaction matrix, input the fused interaction matrix into an attention-guided information bottleneck network to extract matching information, and generate a matching score after preserving the topological structure of the matching information through a progressive dimension reduction network.

[0023] In an optional implementation,

[0024] The dual-tower deep neural network model is trained through a contrastive learning framework of hierarchical scene perception and difficult sample mining, and iteratively optimized through a content-scene cross-attention collaborative representation mechanism, including:

[0025] construct a contrastive learning mechanism of hierarchical scene perception, the contrastive learning mechanism sets a contrastive learning weight for high-level semantic representation, middle-level relationship representation and bottom-level feature representation in the dual-tower deep neural network model, dynamically adjusts the contrastive learning of different hierarchical representations according to the hierarchical weight calculated by the interaction layer, and generates content representation and scene representation with hierarchical weight;

[0026] construct a sample difficulty measurement function based on the interaction matrix of the content representation and the scene representation, the sample difficulty measurement function calculates the matching pattern in the interaction matrix, selects a negative sample from the distance measurement result of the interaction matrix, and inputs the negative sample into a temperature parameter calculation module of a contrastive learning loss function to obtain an adjusted temperature parameter;

[0027] use the adjusted temperature parameter for content-scene collaborative representation learning, make the content tower obtain scene information to generate new content representation through a cross-attention mechanism, make the scene tower obtain content information to generate new scene representation, and input the new content representation and the new scene representation into the contrastive learning mechanism for iterative training.

[0028] In an optional implementation,

[0029] based on the matching score, design a resource allocation algorithm of non-linear mapping and progressive two-stage optimization, determine the optimal matching combination of content and scene and the resource input proportion of each matching combination through iterative optimization, and generate a resource allocation optimization scheme including:

[0030] construct a time sequence adjustment factor containing a periodic fluctuation term and a time decay term;

[0031] Constructing an S-shaped nonlinear mapping function based on the matching score, the S-shaped nonlinear mapping function including an elastic influence parameter and a marginal effect parameter, and generating a dynamic input-output efficiency of the content and scenario combination in combination with the timing adjustment factor;

[0032] Based on the dynamic input-output efficiency, a marginal benefit matrix is ​​constructed, and an initial resource allocation matrix is ​​generated through iterative calculation. In each round of iteration, the marginal benefits of all content and scenario combinations are calculated and the combination with the largest marginal benefit is selected for resource allocation;

[0033] Constructing a multi-objective constraint optimization function, wherein the multi-objective constraint optimization function includes an input-output efficiency term, an entropy regularization term, and a coverage constraint term, wherein the coverage constraint term sets a minimum coverage threshold based on the number of content and the number of scenes;

[0034] Based on the multi-objective constraint optimization function, the initial resource allocation matrix is ​​subjected to progressive two-stage iterative optimization, a random perturbation term is introduced in the exploration stage, and a final resource allocation matrix is ​​obtained based on gradient update in the utilization stage. A resource allocation optimization scheme including priority sorting and resource allocation ratio is generated according to the final resource allocation matrix.

[0035] In an optional embodiment,

[0036] Constructing a timing adjustment factor including periodic fluctuation term and time decay term includes:

[0037] Inputting historical marketing effect data into a wavelet transform network to obtain a multi-scale coefficient sequence, performing empirical mode decomposition on the multi-scale coefficient sequence to obtain an internal mode function sequence, calculating the information entropy value and the variance ratio of the internal mode function sequence, calculating a periodic characteristic value based on the information entropy value and the variance ratio, screening a periodic component based on the periodic characteristic value, calculating a time-varying amplitude coefficient, a time-varying period coefficient, and a time-varying phase coefficient based on the periodic component, and generating a periodic fluctuation term through a kernel weighted polynomial regression combination;

[0038] Calculate the change rate sequence of historical marketing effect data, use the Bayesian method to detect decay rate change points based on the change rate sequence, input the content feature vector and the scene feature vector into the neural network to obtain segmented decay rate parameters, and generate a time decay term based on the decay rate change points and the segmented decay rate parameters;

[0039] Multiplying the periodic fluctuation term by the time decay term to obtain a basic adjustment factor, calculating a residual sequence of historical marketing effect data, inputting the residual sequence into an attention network to generate a residual weight vector, and performing a weighted combination of the basic adjustment factor and the residual weight vector to generate a timing adjustment factor;

[0040] A prediction error of the timing adjustment factor is calculated, and the timing adjustment factor is corrected by a smooth transition function when the prediction error exceeds a preset error threshold.

[0041] In an alternative embodiment,

[0042] The progressive two-stage iterative optimization of the resource allocation initial matrix based on the multi-objective constraint optimization function comprises:

[0043] An initial target value of the multi-objective constraint optimization function is calculated.

[0044] A diversity index, an uncertainty index and a historical improvement rate are calculated according to the resource allocation initial matrix, and a disturbance coefficient is calculated, and a random disturbance term is generated according to the disturbance coefficient.

[0045] The random disturbance term is applied to the resource allocation initial matrix to generate a candidate resource allocation matrix group, and the candidate resource allocation matrix group is input into a multi-objective constraint optimization function to calculate a current target value, the current target value is evaluated based on the initial target value, and a candidate resource allocation matrix with a better evaluation result than the initial target value is selected.

[0046] The candidate resource allocation matrix is stored in an exploration memory bank, and an entropy change rate of the candidate resource allocation matrix is calculated, and when the entropy change rate is lower than a preset change rate threshold, the disturbance coefficient of the random disturbance term is reduced, and a gradient optimization stage is entered.

[0047] A gradient direction of the multi-objective constraint optimization function with respect to the candidate resource allocation matrix is calculated, a learning rate is combined with the dynamic input-output efficiency calculation, and the candidate resource allocation matrix is iteratively updated according to the gradient direction and the learning rate.

[0048] An improvement amplitude of the target value of the multi-objective constraint optimization function is calculated, and when the improvement amplitude of the target value is lower than a termination threshold for a plurality of times, a final resource allocation matrix is output.

[0049] In a second aspect of the embodiment of the application, a content marketing effect evaluation and optimization system based on deep learning is provided, comprising:

[0050] A first unit is configured to acquire historical data of content marketing, including content data and scene data.

[0051] The second unit is used to build a dual-tower deep neural network model, process the content data through the content tower to generate a content representation set, process the scene data through the scene tower to generate a scene representation set, and calculate the matching score of the content representation set and the scene representation set through the interaction layer; the dual-tower deep neural network model is trained through a contrastive learning framework of hierarchical scene perception and difficult sample mining, and is iteratively optimized in combination with a content-scene cross-attention collaborative representation mechanism;

[0052] The third unit is used to design a resource allocation algorithm based on matching scores, combining nonlinear mapping and progressive two-stage optimization. Through iterative optimization, the algorithm determines the optimal matching combination of content and scenarios and the resource investment ratio for each matching combination, generating a resource allocation optimization plan.

[0053] The fourth unit is used to execute the resource allocation optimization plan and collect feedback data to update the dual-tower deep neural network model.

[0054] According to a third aspect of the embodiments of the present invention,

[0055] An electronic device is provided, comprising:

[0056] processor;

[0057] a memory for storing processor-executable instructions;

[0058] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0059] According to a fourth aspect of the embodiments of the present invention,

[0060] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0061] This application achieves efficient processing and representation of content data and scene data by constructing a dual-tower deep neural network model, which can accurately capture the complex relationship between content and scenes, thereby improving the accuracy and reliability of content marketing effectiveness evaluation.

[0062] This application adopts a contrastive learning framework of hierarchical scene perception and difficult sample mining for model training, and combines iterative optimization with the content-scene cross-attention collaborative representation mechanism, which significantly improves the generalization ability and robustness of the model, enabling it to adapt to diverse marketing scenarios.

[0063] This application designs a resource allocation algorithm with nonlinear mapping and progressive two-stage optimization. Through iterative optimization, it determines the optimal matching combination and resource input ratio, realizes the precise allocation of marketing resources, and significantly improves the overall effect and return on investment of content marketing. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 A flowchart of the content marketing effect evaluation and optimization method based on deep learning of the embodiment of the present application is shown in

[0065] Figure 2 A hierarchical representation dynamic weight distribution diagram for different scenarios is shown in

[0066] Figure 3 A performance stability comparison diagram for different data scales is shown in

[0067] Figure 4 A comparison learning loss reduction comparison diagram for different methods is shown in DETAILED DESCRIPTION

[0068] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0069] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in some embodiments.

[0070] Figure 1 A flowchart of the content marketing effect evaluation and optimization method based on deep learning of the embodiment of the present application is shown in Figure 1 As shown in the figure, the method comprises:

[0071] Obtaining historical data of content marketing, including content data and scenario data;

[0072] Building a double-tower deep neural network model, generating a content representation set by processing the content data through a content tower, generating a scenario representation set by processing the scenario data through a scenario tower, and calculating a matching score of the content representation set and the scenario representation set through an interaction layer; the double-tower deep neural network model is trained through a contrastive learning framework of hierarchical scenario perception and difficult sample mining, and iteratively optimized through a content-scenario cross-attention collaborative representation mechanism;

[0073] Based on the matching score, a resource allocation algorithm of non-linear mapping and gradual two-stage optimization is designed, and the optimal matching combination of content and scenario and the resource input proportion of each matching combination are determined through iterative optimization to generate a resource allocation optimization scheme;

[0074] The resource allocation optimization scheme is executed, and feedback data is collected to update the double-tower deep neural network model.

[0075] In an optional implementation, a double-tower deep neural network model is constructed, a content representation set is generated by processing the content data through a content tower, a scene representation set is generated by processing the scene data through a scene tower, and a matching score of the content representation set and the scene representation set is calculated through an interaction layer, including:

[0076] The content tower processes the content data using a multi-head self-attention mechanism, calculates an attraction score of a content element through attention pooling, and performs weighted summation on the content elements based on the attraction score to generate the content representation set.

[0077] The scene data is processed through the scene tower, the scene tower encodes the delivery time, channel characteristics, and target audience characteristics respectively, and captures the mutual influence between scene factors through a self-attention mechanism to generate the scene representation set.

[0078] In the interaction layer, a multi-level interaction matrix of content and scene is calculated, and a hierarchical specialization enhancement operation is applied: the high layer uses an attention mechanism to extract semantic alignment, the middle layer uses a bilinear mapping to capture nonlinear relationships, and the bottom layer uses a local convolution to match patterns; the hierarchical weights are calculated through a context-based evaluation function, and the matching score is generated through an information bottleneck network and a dimension reduction network after weighted fusion.

[0079] The content-scene combinations are ranked based on the matching score; and the matching score, the content global representation, and the scene global representation of the double-tower deep neural network model are output as inputs for a downstream resource allocation optimization task.

[0080] For example, a double-tower deep neural network model is constructed, including a content tower, a scene tower, and an interaction layer. The content tower is used to process content data, the scene tower is used to process scene data, and the interaction layer is used to calculate a matching score of content representation and scene representation.

[0081] For the content tower, a multi-head self-attention mechanism is used to process the content data. Specifically, the content data is encoded into a vector sequence, and each vector represents a content element. Eight attention heads are used, and each head has a dimension of 64. Query, key, and value matrices are generated through linear transformation, attention weights are calculated and the value matrix is weighted and aggregated to obtain attention output. The outputs of the 8 heads are spliced and passed through a linear layer to obtain the output of the multi-head attention. Attention pooling is performed on the output of the multi-head attention to calculate the attraction score of the content element. A two-layer feedforward neural network is used, with an input dimension of 512, a hidden layer dimension of 256, and an output dimension of 1. A scalar score is calculated for each content element, and the attraction score is obtained through softmax normalization. Based on the attraction score, the content elements are weighted and summed to generate the content representation set.

[0082] For the scenario tower, the delivery timing, channel characteristics, and target audience characteristics are encoded separately. The delivery timing uses time embedding to map the timestamp to a 128-dimensional vector. Channel characteristics are encoded using one-hot encoding and mapped to a 64-dimensional vector through the embedding layer. Target audience characteristics are encoded using multi-hot encoding and mapped to a 256-dimensional vector through the embedding layer. These three feature vectors are concatenated to obtain a 448-dimensional scenario vector. The self-attention mechanism is used to capture the mutual influence between scenario factors. Single-head self-attention is used, with an attention dimension of 448. Query, key, and value matrices are generated through linear transformation, and the attention weights are calculated and weighted aggregated value matrices are used to obtain the attention output. Finally, after passing through a layer of feedforward neural network, the output dimension is 512, resulting in a set of scenario representations.

[0083] In the interaction layer, matching scores are calculated and content-scene combinations are ranked based on the matching scores. For example, for 10 content and 100 scene combinations, 1,000 matching scores are calculated, ranked from high to low, and the top N combinations are selected as the final result. The dual-tower deep neural network model outputs the matching scores, content global representation, and scene global representation. The content global representation is the output of the last layer of the content tower, with a dimension of 512. The scene global representation is the output of the last layer of the scene tower, also with a dimension of 512. These outputs can serve as input for downstream resource allocation optimization tasks.

[0084] This application uses a dual-tower deep neural network to achieve precise matching of content and scenarios. The content tower uses a multi-head attention mechanism to extract content attractiveness features, and the scenario tower integrates multi-dimensional information on delivery timing, channel characteristics, and audience characteristics. The interactive layer realizes multi-level feature enhancement and dynamic weight fusion. Combined with the timing adjustment factor and the progressive two-stage optimization strategy, it significantly improves the delivery accuracy and resource utilization efficiency of content marketing, and maximizes the marketing effect.

[0085] In an optional embodiment, calculating the matching score in the interaction layer includes:

[0086] Performing interactive calculations on each layer of representation in the content representation set and each layer of representation in the scene representation set to obtain a multi-level interaction matrix; applying a hierarchically specialized feature enhancement operation to the multi-level interaction matrix, wherein the high-level semantic representation uses an attention mechanism to calculate semantic alignment, the middle-level relationship representation uses a bilinear mapping to capture nonlinear relationships, and the low-level feature representation uses a local convolution operation to capture local pattern matching, thereby generating an enhanced interaction matrix set;

[0087] Constructing a hierarchical importance evaluation function based on content representation, scene representation, and marketing context information, and calculating the importance weight of each enhanced interaction matrix in the enhanced interaction matrix set through an attention network;

[0088] The enhanced interaction matrix set is dynamically weighted and fused according to the importance weights to obtain a fused interaction matrix, the fused interaction matrix is ​​input into an attention-guided information bottleneck network to extract matching information, and a matching score is generated after retaining the topological structure of the matching information through a progressive dimensionality reduction network.

[0089] For example, each layer of representation in the content representation set is interactively calculated with each layer of representation in the scene representation set to obtain a multi-level interaction matrix. Specifically, cosine similarity can be used to calculate the similarity between the two representation vectors to generate the interaction matrix. For example, for a 512-dimensional content representation vector and a 512-dimensional scene representation vector, a 512x512-dimensional interaction matrix is ​​calculated.

[0090] Hierarchically specialized feature enhancement operations are applied to the multi-level interaction matrix. For high-level semantic representations, an attention mechanism is used to calculate semantic alignment. A multi-head attention mechanism can be used, with eight attention heads, each with a dimension of 64. By performing linear transformations and softmax normalization on the query matrix, key matrix, and value matrix, attention weights are calculated and weighted to form an aggregate value matrix to obtain an enhanced high-level semantic interaction matrix. For mid-level relational representations, a bilinear mapping is used to capture nonlinear relationships. Specifically, two fully connected layers are used to map the content representation and scene representation to a 128-dimensional space, respectively. The outer product is then calculated to obtain a 128x128 bilinear interaction matrix. This approach effectively captures the nonlinear cross-representation between the two representations. For low-level feature representations, local convolution operations are used to capture local pattern matching. A 3x3 convolution kernel with a stride of 1, padding of 1, and 64 output channels can be used. This convolution operation extracts matching patterns in local regions and enhances the interaction information of low-level features.

[0091] A hierarchical importance evaluation function is constructed based on content representation, scenario representation, and marketing context information. Specifically, the three types of information are concatenated and passed through a two-layer fully connected network, with the output dimension matching the number of augmented interaction matrices. The softmax function is then used for normalization to obtain the importance weights for each augmented interaction matrix.

[0092] Based on the calculated importance weights, the set of augmented interaction matrices is dynamically weighted and fused. Assuming there are three augmented interaction matrices with corresponding weights of [0.4, 0.3, 0.3], the three matrices are multiplied by their corresponding weights and then summed to obtain the fused interaction matrix. The fused interaction matrix is ​​then fed into an attention-guided information bottleneck network to extract matching information. This network consists of multiple layers of self-attention modules and a feedforward network. It can be configured as four layers, each with a hidden dimension of 256. This multi-layer attention mechanism enables the network to adaptively focus on important matching information. A progressive dimensionality reduction network preserves the topological structure of the matching information and generates a matching score. This network consists of multiple fully connected layers, progressively reducing the feature dimension from 256 to 128, 64, 32, and 16, ultimately outputting a scalar matching score. During the dimensionality reduction process, residual connections can be added to preserve the original information.

[0093] Figure 2 This figure shows the dynamic weight distribution of hierarchical representations in different scenarios, demonstrating the importance weights automatically assigned to each layer of representations by the present invention based on different content marketing scenarios. In e-commerce scenarios, high-level semantic representations received the highest weight of 0.446, followed by mid-level relational representations at 0.352 and low-level feature representations at 0.202. In social media scenarios, the weight distribution was more balanced, at 0.382, 0.363, and 0.255, respectively. In cross-media marketing scenarios, the weight of high-level semantic representations was even more prominent, reaching 0.472, while mid-level relational representations received 0.321 and low-level feature representations received 0.207.

[0094] Figure 3This is a comparison chart of performance stability under different data scales. As shown in the figure, in a small sample scenario (5,000 samples), the F1 score of the method of the present application still reaches 0.872, which is significantly higher than the deep cross network (0.642) and cosine similarity (0.613); as the sample size increases to 5 million, the F1 score of the method of the present application stabilizes at above 0.935, and the performance improvement curve tends to be stable early, verifying that the hierarchical specialization processing and dynamic fusion strategy adopted by the present application can effectively alleviate the data sparsity problem and can also extract effective matching features in small sample scenarios. In particular, when the data scale increases from 250,000 to 5 million, the performance improvement of the method of the present application is only 1.7%, while the comparison method still has 4%-7% room for improvement, proving that the method of the present application is close to the optimal performance under smaller data scales. This advantage stems from the efficient extraction and retention capabilities of the multi-level interaction matrix and the attention-guided information bottleneck network. This application achieves accurate content-scene matching through multi-level interactive calculation and feature enhancement, combined with dynamic weighting and information bottleneck networks. Each step of the method is carefully designed to capture matching information at different levels and can be flexibly adjusted according to actual application scenarios. Through experimental verification and interpretable analysis on large-scale datasets, the advantages of this method in improving matching accuracy and understanding matching mechanisms are demonstrated.

[0095] In an optional embodiment, the dual-tower deep neural network model is trained through a contrastive learning framework of hierarchical scene perception and difficult sample mining, and iteratively optimized in combination with a content-scene cross-attention collaborative representation mechanism, including:

[0096] Constructing a contrastive learning mechanism for hierarchical scene perception. The contrastive learning mechanism sets contrastive learning weights for high-level semantic representations, mid-level relational representations, and low-level feature representations in the dual-tower deep neural network model. The contrastive learning of representations at different levels is dynamically adjusted based on the hierarchical weights calculated at the interaction layer to generate content representations and scene representations with hierarchical weights.

[0097] constructing a sample difficulty metric function based on an interaction matrix between the content representation and the scene representation, the sample difficulty metric function calculating matching patterns in the interaction matrix, selecting a negative sample from a distance measurement result based on the interaction matrix, and inputting the negative sample into a temperature parameter calculation module of a contrastive learning loss function to obtain an adjusted temperature parameter;

[0098] The adjusted temperature parameters are used to perform content-scene collaborative representation learning. The content tower obtains scene information through a cross-attention mechanism to generate a new content representation, and the scene tower obtains content information to generate a new scene representation. The new content representation and the new scene representation are input into the contrastive learning mechanism for iterative training.

[0099] For example, during the construction of a contrastive learning mechanism for hierarchical scene perception, contrastive learning weights are set for the three levels of representation in the dual-tower model. Specifically, the initial weight of the high-level semantic representation is set to 0.5, the initial weight of the mid-level relational representation is set to 0.3, and the initial weight of the low-level feature representation is set to 0.2. The input content information and scene information are processed through the content tower and the scene tower, respectively, to obtain representations at each level. For example, in an e-commerce recommendation scenario, when the product description "thin and portable 13-inch laptop" is input as content information and the user browsing history "office equipment, digital products" is input as scene information, the content tower extracts the low-level feature representation (such as keyword features: thin and light, portable, 13 inches, laptop), the mid-level relational representation (such as the relationship between product category and attribute: computer - thin and light), and the high-level semantic representation (such as the overall semantics of the product: mobile office equipment). Similarly, the scene tower extracts three-level representations of scene information.

[0100] The weights of representations at different levels are calculated through the interaction layer. The interaction layer receives the output of the content tower and the scene tower and calculates the correlation between the two. Specifically, the interaction layer calculates the similarity matrix of the representations of each layer of the content tower and the scene tower to perform importance scoring. If the similarity of the high-level semantic representation is high (for example, the similarity is 0.85), the weight of the high-level semantic representation is increased to 0.6; if the similarity of the middle-level relationship representation is moderate (for example, the similarity is 0.65), the weight is maintained at 0.3; if the similarity of the underlying feature representation is low (for example, the similarity is 0.4), the weight is reduced to 0.1. Through this dynamic adjustment, content representations and scene representations with hierarchical weights are generated.

[0101] A sample difficulty measurement function is constructed based on the interaction matrix of content representation and scene representation. The interaction matrix is ​​a two-dimensional matrix, in which each element represents the correlation between a dimension in the content representation and a dimension in the scene representation. The sample difficulty measurement function calculates the matching pattern in the interaction matrix. In specific implementation, the distribution of each element in the interaction matrix is ​​calculated, including the mean, variance, and kurtosis. If the elements of the interaction matrix are distributed relatively evenly (variance is less than 0.1), it means that the matching degree between the content and the scene is low, and the sample is judged as a difficult sample, and its difficulty value is set to 0.8; if the elements of the interaction matrix are concentrated and the peak is obvious (kurtosis is greater than 3), it means that the matching degree between the content and the scene is high, and the difficulty value of the sample is set to 0.2.

[0102] Negative samples are selected based on the distance measurement results of the interaction matrix. In the training batch, assume that there are 32 samples, each containing a pair of content-scene information. For the i-th sample, its content information and the scene information of the j-th sample (j≠i) form a negative sample pair. By calculating the content-scene representation distance of all possible negative sample pairs, the 5 with the smallest distance are selected as difficult negative samples. For example, if the cosine distance between the content representation of the 3rd sample and the scene representation of the 12th sample is 0.75, which is higher than the average distance of 0.35, then the negative sample pair is identified as a difficult negative sample. The selected negative samples are input into the temperature parameter calculation module of the contrastive learning loss function. The temperature parameter controls the smoothness of the similarity distribution between sample pairs in contrastive learning. In specific implementation, the base temperature parameter is set to 0.07. For samples with a difficulty value of 0.8, the temperature parameter is adjusted to 0.05 (base temperature multiplied by (1-0.8×0.5)) to obtain a higher gradient in the loss calculation; for samples with a difficulty value of 0.2, the temperature parameter is adjusted to 0.065 (base temperature multiplied by (1-0.2×0.5)).

[0103] The adjusted temperature parameters are used for content-scene collaborative representation learning. Through the cross-attention mechanism, the content tower obtains scene information and the scene tower obtains content information. In specific implementation, the attention weight matrix of the content representation and the scene representation is calculated, and the matrix size is the content representation dimension multiplied by the scene representation dimension. Through this matrix, the content representation is weighted according to the scene representation to generate a new content representation that takes scene information into account; similarly, the scene representation is weighted according to the content representation to generate a new scene representation that takes content information into account. The new content representation and the new scene representation are input into the contrastive learning mechanism for iterative training. In each iteration cycle, the similarity of the positive sample pair (matching content-scene pair) and the similarity of the negative sample pair (mismatching content-scene pair) are calculated, and the contrast loss is calculated using the adjusted temperature parameters. For example, in the first iteration, the similarity of the positive sample pairs is 0.75, the similarity of the difficult negative sample pairs is 0.65, and the contrast loss calculated using the temperature parameter 0.05 is 1.25; after the second iteration, the similarity of the positive sample pairs increases to 0.85, the similarity of the difficult negative sample pairs decreases to 0.55, and the contrast loss decreases to 0.85.

[0104] Figure 4The figure is a comparison of the contrastive learning loss decrease of different methods. The horizontal axis in the figure represents the number of training iterations, and the vertical axis represents the contrastive learning loss value. The loss value of the method of this application (blue line) is 0.724 at the beginning of training (10 iterations), drops to 0.409 in the middle (40 iterations), and converges to 0.191 in the end (100 iterations), showing the fastest decline rate and the lowest convergence value. In contrast, the standard contrastive learning method (orange line) converges slowly, with a final loss value of 0.650; the performance of the hard negative sample mining method (green line) is between the two, with a final loss value of 0.435. This shows that the contrastive learning framework of hierarchical scene perception and difficult sample mining of this application is superior to the contrast method in both optimization efficiency and final performance, confirming its effectiveness in accelerating model convergence and improving representation quality. Through the above iterative training process, the dual-tower deep neural network model gradually learns the complex relationship between content information and scene information, improving the model's representation ability and recommendation accuracy. The traditional content-scene matching model does not fully consider the difficulty of the samples during the training process, which makes it difficult for the model to capture complex matching patterns, especially when dealing with long-tail scenes and sparse features. This application proposes a contrastive learning framework for hierarchical scene perception and difficult sample mining. By designing a dynamic weight mechanism at three levels: high-level semantics, middle-level relationships, and low-level features, it achieves adaptive fusion of feature representation. At the same time, the sample difficulty metric and temperature parameter adjustment strategy based on the interaction matrix are introduced to enhance the model's learning ability for difficult samples. In addition, an innovative content-scene cross-attention collaborative representation mechanism is designed to enable the two feature towers to perceive each other and enhance each other's representation capabilities.

[0105] In an optional embodiment, based on the matching score, a resource allocation algorithm combining nonlinear mapping and progressive two-stage optimization is designed. The optimal matching combination of content and scenario and the resource investment ratio of each matching combination are determined through iterative optimization. The generated resource allocation optimization solution includes:

[0106] Construct a timing adjustment factor including periodic fluctuation term and time decay term;

[0107] Constructing an S-shaped nonlinear mapping function based on the matching score, the S-shaped nonlinear mapping function including an elastic influence parameter and a marginal effect parameter, and generating a dynamic input-output efficiency of the content and scenario combination in combination with the timing adjustment factor;

[0108] Based on the dynamic input-output efficiency, a marginal benefit matrix is ​​constructed, and an initial resource allocation matrix is ​​generated through iterative calculation. In each round of iteration, the marginal benefits of all content and scenario combinations are calculated and the combination with the largest marginal benefit is selected for resource allocation;

[0109] Constructing a multi-objective constraint optimization function, wherein the multi-objective constraint optimization function includes an input-output efficiency term, an entropy regularization term, and a coverage constraint term, wherein the coverage constraint term sets a minimum coverage threshold based on the number of content and the number of scenes;

[0110] Based on the multi-objective constraint optimization function, the initial resource allocation matrix is ​​subjected to progressive two-stage iterative optimization, a random perturbation term is introduced in the exploration stage, and a final resource allocation matrix is ​​obtained based on gradient update in the utilization stage. A resource allocation optimization scheme including priority sorting and resource allocation ratio is generated according to the final resource allocation matrix.

[0111] For example, a timing adjustment factor is constructed that includes a periodic fluctuation term and a time decay term. An S-shaped nonlinear mapping function is constructed, with the slope reaching its maximum within the matching score range of 0.3 to 0.7, reflecting the marginal effect of resource input. When the matching score is less than 0.3, the return from increased investment is minimal; when the matching score is greater than 0.7, the marginal effect of further investment decreases. Combined with the timing adjustment factor, dynamic input-output efficiency can be obtained. For example, the input-output efficiency of a certain content can be increased by 25% during the holiday season.

[0112] A marginal benefit matrix is ​​constructed based on dynamic input-output efficiency. For example, assuming there are 10 content items and 100 scenarios, a marginal benefit matrix with 1,000 combinations is formed. Resources are allocated iteratively, with the combination with the highest marginal benefit selected in each round. For example, in the first round, a combination with a marginal benefit of 0.85 is allocated 15% of resources. In the second round, a combination with a marginal benefit of 0.78 is allocated 12%, and so on.

[0113] When constructing a multi-objective constrained optimization function, the entropy regularization term is used to control the degree of dispersion of resource allocation and avoid excessive concentration of resources. The coverage constraint term ensures that a sufficient number of content and scenarios are covered to prevent blind spots in delivery. For example, at least 30% of scenarios and 60% of content are required to receive resource allocation. The input-output efficiency weight is set to 0.5, the entropy regularization weight is set to 0.3, and the coverage constraint weight is set to 0.2. For 100 scenarios, the minimum coverage threshold is set to 30%, which means that at least 30 scenarios need to be covered. For 10 content, the minimum coverage threshold is set to 60%, which means that at least 6 content needs to be delivered.

[0114] A progressive two-stage optimization is performed to obtain the final resource allocation matrix. A specific delivery plan is generated based on the resource allocation matrix. For example, for a combination of ten contents and one hundred scenarios, the matching scores are sorted from high to low: content A accounts for 18% of the delivery in scenario 1, content B accounts for 15% of the delivery in scenario 3, content C accounts for 12% of the delivery in scenario 5, and so on, until the resources are allocated.

[0115] The present invention constructs a resource allocation algorithm based on timing adjustment and S-shaped nonlinear mapping, combines a multi-objective optimization function of input-output efficiency, entropy regularization and coverage constraints, and adopts a progressive two-stage iterative strategy. It can effectively capture the periodic fluctuations and time decay characteristics of marketing effects, avoid excessive concentration of resources, and ensure reasonable coverage of content and scenarios. At the same time, through random perturbations in the exploration stage and gradient optimization in the utilization stage, the global optimization of the resource allocation plan is achieved, which significantly improves the accuracy of marketing delivery and resource utilization efficiency.

[0116] In an optional implementation, constructing a timing adjustment factor including a periodic fluctuation term and a time decay term includes:

[0117] Inputting historical marketing effect data into a wavelet transform network to obtain a multi-scale coefficient sequence, performing empirical mode decomposition on the multi-scale coefficient sequence to obtain an internal mode function sequence, calculating the information entropy value and the variance ratio of the internal mode function sequence, calculating a periodic characteristic value based on the information entropy value and the variance ratio, screening a periodic component based on the periodic characteristic value, calculating a time-varying amplitude coefficient, a time-varying period coefficient, and a time-varying phase coefficient based on the periodic component, and generating a periodic fluctuation term through a kernel weighted polynomial regression combination;

[0118] Calculate the change rate sequence of historical marketing effect data, use the Bayesian method to detect decay rate change points based on the change rate sequence, input the content feature vector and the scene feature vector into the neural network to obtain segmented decay rate parameters, and generate a time decay term based on the decay rate change points and the segmented decay rate parameters;

[0119] Multiplying the periodic fluctuation term by the time decay term to obtain a basic adjustment factor, calculating a residual sequence of historical marketing effect data, inputting the residual sequence into an attention network to generate a residual weight vector, and performing a weighted combination of the basic adjustment factor and the residual weight vector to generate a timing adjustment factor;

[0120] A prediction error of the timing adjustment factor is calculated, and when the prediction error exceeds a preset error threshold, the timing adjustment factor is corrected using a smooth transition function.

[0121] For example, constructing the cyclical fluctuation term requires inputting historical marketing effectiveness data. For example, for an e-commerce platform, daily sales data for the past 24 months was collected as historical marketing effectiveness data. This data was then fed into a wavelet transform network, where a four-layer decomposition using the Daubechies wavelet basis was performed. This yielded a multi-scale coefficient sequence, including both approximate and detailed coefficients. These coefficients reflect the characteristics of the sales data at different time scales.

[0122] An empirical mode decomposition (EMD) is performed on the multi-scale coefficient sequence. Specifically, through an iterative screening process, the sales data is decomposed into 10 IMF sequences and one residual term. Each IMF represents an oscillation mode within a specific frequency range in the data. For each IMF sequence, its information entropy and variance ratio are calculated. The information entropy is obtained by calculating the probability distribution of the sequence, and the variance is calculated by calculating the sum of the squared deviations of the sequence from its mean. For example, for the first IMF, its information entropy is 0.85, its variance is 2.3, and its ratio is 0.37.

[0123] The periodic eigenvalue is calculated based on the information entropy and variance ratio. The periodic eigenvalue is obtained by identifying the primary frequency components through Fourier transform and weighting them with the information entropy and variance ratio. For example, the periodic eigenvalue of the first internal mode function is 0.42, and the second is 0.65. A threshold of 0.5 is set based on the periodic eigenvalues ​​to filter the periodic components. Internal mode functions with eigenvalues ​​greater than the threshold are selected as periodic components. In this example, the second, third, and fifth internal mode functions are selected as periodic components.

[0124] Based on the selected periodic components, the time-varying amplitude coefficient, time-varying period coefficient, and time-varying phase coefficient are calculated. The time-varying amplitude coefficient is obtained by Hilbert transforming the instantaneous amplitude. For example, the amplitude coefficient of the second internal mode function is [1.2, 1.3, 1.1...]. The time-varying period coefficient is obtained by analyzing the instantaneous frequency variation. For example, the period coefficient is [7, 7.2, 6.9...] days. The time-varying phase coefficient is obtained by calculating the instantaneous phase. For example, the phase coefficient is [0.2π, 0.3π, 0.25π...]. These coefficients are input into the kernel-weighted polynomial regression model and weighted using a Gaussian kernel function to generate a periodic fluctuation term, a time-varying periodic function of the form "baseline value + amplitude × sin(phase + time / period)".

[0125] To construct the time decay term, we first calculate the rate of change sequence of historical marketing effectiveness data. For example, if sales volume changes from 1,000 to 900 units for two consecutive days, the rate of change is -0.1. Based on the rate of change sequence, a Bayesian approach is used to detect change points in the decay rate. Specifically, a Bayesian online change point detection algorithm is used to identify change points by calculating the posterior probability distribution. In this example, significant changes in the decay rate are detected on days 30, 90, and 180. Product content feature vectors (e.g., price 200 yuan, weight 0.5 kg, color red) and scenario feature vectors (e.g., holiday 1, promotion 2, season spring, etc.) are collected. These feature vectors are input into a deep neural network consisting of two hidden layers, with 64 and 32 neurons in each layer, respectively, using the ReLU activation function. The network outputs segmented decay rate parameters, for example, 0.05 / day for the first segment, 0.02 / day for the second segment, 0.01 / day for the third segment, and 0.03 / day for the fourth segment. Generate a time decay term based on the decay rate change point and the segmented decay rate parameters.

[0126] Multiply the cyclical fluctuation term by the time decay term to obtain the base adjustment factor. For example, if the cyclical fluctuation term is 1.2 and the time decay term is 0.8 on a particular day, the base adjustment factor is 0.96. Calculate the residual sequence between the historical marketing performance data and the predicted value using the base adjustment factor. This residual sequence is input into a self-attention network, which uses a multi-head attention mechanism with 8 heads and a hidden layer dimension of 256. The network outputs a residual weight vector, such as [0.05, 0.08, 0.03...]. The base adjustment factor and the residual weight vector are weighted together, i.e., "base adjustment factor × (1 + residual weight)", to generate the final time series adjustment factor.

[0127] The forecast error of the time series adjustment factor is calculated as the mean absolute percentage error between the sales volume predicted using the adjustment factor and the actual sales volume. If the forecast error exceeds a preset error threshold of 5%, the time series adjustment factor is adjusted using a smoothing function. Specifically, a sigmoid function is used to achieve smoothing and suppress abnormal fluctuations. For example, if the adjustment factor for a particular day is 1.5, it will be adjusted to 1.35, bringing the predicted value closer to the actual value.

[0128] This application proposes a time series adjustment method based on multi-scale decomposition and adaptive fusion. By combining wavelet transform with empirical mode decomposition, it extracts multi-level periodic features from marketing effectiveness data. It innovatively introduces Bayesian decay rate change point detection and neural network-driven piecewise decay modeling to accurately characterize temporal decay characteristics. Furthermore, through the design of an attention mechanism and a smooth transition function, the model's adaptability to abnormal fluctuations is improved.

[0129] The method proposed in this application accurately captures the cyclical fluctuations and time-attenuation characteristics of marketing effectiveness, significantly improving forecasting accuracy. It demonstrates enhanced modeling capabilities, particularly when dealing with complex scenarios such as multiple overlapping periods and dynamic attenuation. By introducing residual learning and a smooth transition mechanism, the model demonstrates enhanced robustness to unexpected events and abnormal fluctuations, providing reliable decision support for the precise allocation of marketing resources.

[0130] In an optional implementation, performing a progressive two-stage iterative optimization on the initial resource allocation matrix based on the multi-objective constraint optimization function includes:

[0131] Calculating an initial objective value of the multi-objective constrained optimization function;

[0132] Calculating a diversity index, an uncertainty index, and a historical improvement rate according to the initial resource allocation matrix and obtaining a disturbance coefficient, and generating a random disturbance term according to the disturbance coefficient;

[0133] Applying the random perturbation term to the initial resource allocation matrix to generate a candidate resource allocation matrix group, inputting the candidate resource allocation matrix group into a multi-objective constraint optimization function to calculate a current target value, evaluating the current target value based on the initial target value, and selecting a candidate resource allocation matrix having an evaluation result that is better than the initial target value;

[0134] Storing the candidate resource allocation matrix in an exploration memory, calculating the entropy change rate of the candidate resource allocation matrix, and when the entropy change rate is lower than a preset change rate threshold, reducing the perturbation coefficient of the random perturbation term and entering the gradient optimization stage;

[0135] Calculating the gradient direction of the multi-objective constraint optimization function with respect to the candidate resource allocation matrix, calculating a learning rate in combination with the dynamic input-output efficiency, and iteratively updating the candidate resource allocation matrix according to the gradient direction and the learning rate;

[0136] The target value improvement range of the multi-objective constraint optimization function is calculated, and a final resource allocation matrix is ​​output when the target value improvement range is lower than a termination threshold for multiple consecutive times.

[0137] Exemplarily, the initial target value of the multi-objective constraint optimization function is calculated. Specifically, the initial resource allocation matrix can be input into a predefined multi-objective constraint optimization function to obtain an initial target value. For example, assuming the initial target value is 100.

[0138] The diversity index, uncertainty index, and historical improvement rate are calculated according to the resource allocation initial matrix, and a perturbation coefficient is calculated. The diversity index can be measured by calculating the number of different elements in the matrix. The uncertainty index can be represented using the variance of the matrix elements. The historical improvement rate can be determined according to the target value changes of the previous iterations. By considering these indicators comprehensively, a perturbation coefficient such as 0.1 can be obtained. Then, a random perturbation term is generated according to the perturbation coefficient, which can be achieved by multiplying a normally distributed random number by the perturbation coefficient.

[0139] The random perturbation term is applied to the resource allocation initial matrix to generate a candidate resource allocation matrix group. Specifically, each element of the initial matrix can be added to the corresponding random perturbation value to obtain multiple candidate matrices. For example, 10 candidate matrices are generated. These candidate resource allocation matrix groups are input into the multi-objective constraint optimization function to calculate the current target value. Based on the initial target value, the current target value is evaluated, and the candidate resource allocation matrix with a better evaluation result than the initial target value is selected. For example, assuming that the target values of the three candidate matrices are 105, 102, and 98, respectively, the two matrices with target values of 105 and 102 are selected.

[0140] The selected candidate resource allocation matrix is stored in the exploration memory bank. The entropy change rate of these candidate resource allocation matrices is calculated, which can be obtained by calculating the information entropy of the matrix element distribution and comparing it with the last iteration result. When the entropy change rate is lower than the preset change rate threshold, for example, lower than 0.01, the perturbation coefficient of the random perturbation term is reduced, such as halved, and the gradient optimization phase is entered.

[0141] In the gradient optimization phase, the gradient direction of the multi-objective constraint optimization function with respect to the candidate resource allocation matrix is first calculated. Numerical differentiation method can be used to make a small perturbation to each element in the matrix, calculate the change of the target function, and thus obtain the gradient vector. The learning rate is calculated in combination with the dynamic input-output efficiency. The dynamic input-output efficiency can be measured by calculating the target value improvement brought by unit resource input. For example, if increasing 1% of resource input brings 0.5% of target value improvement, the learning rate can be set to 0.005. The candidate resource allocation matrix is iteratively updated according to the gradient direction and the learning rate. Specifically, each element in the matrix can be moved along the gradient direction by a distance of the learning rate size.

[0142] The target value improvement amplitude of the multi-objective constraint optimization function is calculated, which can be compared with the target value of the last iteration. When the target value improvement amplitude is continuously lower than the termination threshold for multiple times, the final resource allocation matrix is output. For example, if the target value improvement amplitude of 5 consecutive iterations is less than 0.1%, it can be considered that the optimization process has converged, and the current resource allocation matrix is output as the final result.

[0143] In practical applications, the above process can be adjusted appropriately based on the characteristics of the specific problem. For example, for a resource allocation problem with 10 resources and 20 tasks, the initial matrix can be a 10x20 random matrix. The multi-objective constrained optimization function can include multiple objectives, such as maximizing overall efficiency, minimizing resource waste, and balancing task load. During the optimization process, the maximum number of iterations can be set to 1000, generating 50 candidate matrices each time, and selecting the 10 with the best objective values ​​for the next optimization step. The perturbation coefficient can start at 0.1 and be reduced by 10% each time until the gradient optimization phase begins. The learning rate can start at 0.01 and be dynamically adjusted based on the improvement in the objective value. Ultimately, when the objective value improvement for 10 consecutive iterations is less than 0.01%, the final resource allocation solution is output.

[0144] In the field of resource allocation optimization, traditional methods mainly use linear programming or greedy algorithms for solution, such as phased iteration, simulated annealing and other methods. These methods are prone to fall into local optimality when facing large-scale, multi-objective, and dynamically changing resource allocation problems in content marketing scenarios, and cannot effectively balance the relationship between efficiency and coverage. This application proposes a resource allocation optimization method based on two-stage iteration, which innovatively introduces diversity indicators, uncertainty indicators and historical improvement rates to dynamically adjust the disturbance coefficient. By combining random disturbances in the exploration phase with gradient optimization in the utilization phase, it achieves a balance between global search and local optimization. At the same time, it improves optimization efficiency by adaptively controlling the algorithm phase transition through the entropy change rate and adjusting the learning rate in combination with dynamic input-output efficiency.

[0145] A second aspect of an embodiment of the present invention provides a content marketing effectiveness evaluation and optimization system based on deep learning, the system comprising:

[0146] The first unit is used to obtain historical data of content marketing, including content data and scenario data;

[0147] The second unit is used to build a dual-tower deep neural network model, process the content data through the content tower to generate a content representation set, process the scene data through the scene tower to generate a scene representation set, and calculate the matching score of the content representation set and the scene representation set through the interaction layer; the dual-tower deep neural network model is trained through a contrastive learning framework of hierarchical scene perception and difficult sample mining, and is iteratively optimized in combination with a content-scene cross-attention collaborative representation mechanism;

[0148] The third unit is used to design a resource allocation algorithm based on matching scores, combining nonlinear mapping and progressive two-stage optimization. Through iterative optimization, the algorithm determines the optimal matching combination of content and scenarios and the resource investment ratio for each matching combination, generating a resource allocation optimization plan.

[0149] The fourth unit is used to execute the resource allocation optimization plan and collect feedback data to update the dual-tower deep neural network model.

[0150] According to a third aspect of the embodiments of the present invention,

[0151] An electronic device is provided, comprising:

[0152] processor;

[0153] a memory for storing processor-executable instructions;

[0154] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0155] According to a fourth aspect of the embodiments of the present invention,

[0156] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0157] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A content marketing effectiveness evaluation and optimization method based on deep learning, characterized by: include: Obtain historical data on content marketing, including content data and scenario data; A dual-tower deep neural network model is constructed, wherein the content data is processed by the content tower to generate a content representation set, the scene data is processed by the scene tower to generate a scene representation set, and the matching score of the content representation set and the scene representation set is calculated by the interaction layer, including: the content tower uses a multi-head self-attention mechanism to process the content data, and calculates the attractiveness score of the content elements through attention pooling, and performs weighted summation on the content elements based on the attractiveness score to generate a content representation set; the scene data is processed by the scene tower, and the scene tower encodes the delivery timing, channel characteristics and target audience characteristics respectively, and captures the mutual influence between scene factors through the self-attention mechanism to generate a scene representation set; in the interaction layer, each layer of representation in the content representation set is interactively calculated with each layer of representation in the scene representation set to obtain a multi-level interaction matrix; and a hierarchical specialized feature enhancement operation is applied to the multi-level interaction matrix. The high-level semantic representation uses the attention mechanism to calculate the semantic alignment, the middle-level relationship representation uses bilinear mapping to capture nonlinear relationships, and the low-level feature representation uses local convolution operations to capture local pattern matching to generate a set of enhanced interaction matrices; a hierarchical importance evaluation function is constructed based on content representation, scene representation and marketing context information, and the importance weight of each enhanced interaction matrix in the set of enhanced interaction matrices is calculated through the attention network; the set of enhanced interaction matrices is dynamically weighted and fused according to the importance weight to obtain a fused interaction matrix, and the fused interaction matrix is ​​input into the attention-guided information bottleneck network to extract matching information, and a matching score is generated after retaining the topological structure of the matching information through the progressive dimensionality reduction network; content-scene combinations are ranked based on the matching score; the matching score, content global representation and scene global representation of the dual-tower deep neural network model are output as input to the downstream resource allocation optimization task; The dual-tower deep neural network model is trained through a contrastive learning framework of hierarchical scene perception and difficult sample mining, and is iteratively optimized in combination with a content-scene cross-attention collaborative representation mechanism; Based on the matching score, a resource allocation algorithm with nonlinear mapping and progressive two-stage optimization is designed. Through iterative optimization, the optimal matching combination of content and scenarios and the resource investment ratio of each matching combination are determined to generate a resource allocation optimization plan. Execute resource allocation optimization plan and collect feedback data to update the dual-tower deep neural network model.

2. The method according to claim 1, characterized in that The dual-tower deep neural network model is trained through a contrastive learning framework of hierarchical scene perception and difficult sample mining, and is iteratively optimized by combining a content-scene cross-attention collaborative representation mechanism. Constructing a contrastive learning mechanism for hierarchical scene perception. The contrastive learning mechanism sets contrastive learning weights for high-level semantic representations, mid-level relational representations, and low-level feature representations in the dual-tower deep neural network model. The contrastive learning of representations at different levels is dynamically adjusted based on the hierarchical weights calculated at the interaction layer to generate content representations and scene representations with hierarchical weights. constructing a sample difficulty metric function based on an interaction matrix between the content representation and the scene representation, the sample difficulty metric function calculating matching patterns in the interaction matrix, selecting a negative sample from a distance measurement result based on the interaction matrix, and inputting the negative sample into a temperature parameter calculation module of a contrastive learning loss function to obtain an adjusted temperature parameter; The adjusted temperature parameters are used to perform content-scene collaborative representation learning. The content tower obtains scene information through a cross-attention mechanism to generate a new content representation, and the scene tower obtains content information to generate a new scene representation. The new content representation and the new scene representation are input into the contrastive learning mechanism for iterative training.

3. The method according to claim 1, characterized in that Based on the matching score, a resource allocation algorithm with nonlinear mapping and progressive two-stage optimization is designed. Through iterative optimization, the optimal matching combination of content and scenarios and the resource investment ratio of each matching combination are determined. The generated resource allocation optimization solution includes: Construct a timing adjustment factor including periodic fluctuation term and time decay term; Constructing an S-shaped nonlinear mapping function based on the matching score, the S-shaped nonlinear mapping function including an elastic influence parameter and a marginal effect parameter, and generating a dynamic input-output efficiency of the content and scenario combination in combination with the timing adjustment factor; Based on the dynamic input-output efficiency, a marginal benefit matrix is ​​constructed, and an initial resource allocation matrix is ​​generated through iterative calculation. In each round of iteration, the marginal benefits of all content and scenario combinations are calculated and the combination with the largest marginal benefit is selected for resource allocation; Constructing a multi-objective constraint optimization function, wherein the multi-objective constraint optimization function includes an input-output efficiency term, an entropy regularization term, and a coverage constraint term, wherein the coverage constraint term sets a minimum coverage threshold based on the number of content and the number of scenes; Based on the multi-objective constraint optimization function, the initial resource allocation matrix is ​​subjected to progressive two-stage iterative optimization, a random perturbation term is introduced in the exploration stage, and a final resource allocation matrix is ​​obtained based on gradient update in the utilization stage. A resource allocation optimization scheme including priority sorting and resource allocation ratio is generated according to the final resource allocation matrix.

4. The method according to claim 3, characterized in that Constructing a timing adjustment factor including periodic fluctuation term and time decay term includes: Inputting historical marketing effect data into a wavelet transform network to obtain a multi-scale coefficient sequence, performing empirical mode decomposition on the multi-scale coefficient sequence to obtain an internal mode function sequence, calculating the information entropy value and the variance ratio of the internal mode function sequence, calculating a periodic characteristic value based on the information entropy value and the variance ratio, screening a periodic component based on the periodic characteristic value, calculating a time-varying amplitude coefficient, a time-varying period coefficient, and a time-varying phase coefficient based on the periodic component, and generating a periodic fluctuation term through a kernel weighted polynomial regression combination; Calculate the change rate sequence of historical marketing effect data, use the Bayesian method to detect decay rate change points based on the change rate sequence, input the content feature vector and the scene feature vector into the neural network to obtain segmented decay rate parameters, and generate a time decay term based on the decay rate change points and the segmented decay rate parameters; Multiplying the periodic fluctuation term by the time decay term to obtain a basic adjustment factor, calculating a residual sequence of historical marketing effect data, inputting the residual sequence into an attention network to generate a residual weight vector, and performing a weighted combination of the basic adjustment factor and the residual weight vector to generate a timing adjustment factor; A prediction error of the timing adjustment factor is calculated, and when the prediction error exceeds a preset error threshold, the timing adjustment factor is corrected using a smooth transition function.

5. The method according to claim 3, characterized in that Performing a progressive two-stage iterative optimization on the initial resource allocation matrix based on the multi-objective constraint optimization function includes: Calculating an initial objective value of the multi-objective constrained optimization function; Calculating a diversity index, an uncertainty index, and a historical improvement rate according to the initial resource allocation matrix and obtaining a disturbance coefficient, and generating a random disturbance term according to the disturbance coefficient; Applying the random perturbation term to the initial resource allocation matrix to generate a candidate resource allocation matrix group, inputting the candidate resource allocation matrix group into a multi-objective constraint optimization function to calculate a current target value, evaluating the current target value based on the initial target value, and selecting a candidate resource allocation matrix having an evaluation result that is better than the initial target value; Storing the candidate resource allocation matrix in an exploration memory, calculating the entropy change rate of the candidate resource allocation matrix, and when the entropy change rate is lower than a preset change rate threshold, reducing the perturbation coefficient of the random perturbation term and entering the gradient optimization stage; Calculating the gradient direction of the multi-objective constraint optimization function with respect to the candidate resource allocation matrix, calculating a learning rate in combination with the dynamic input-output efficiency, and iteratively updating the candidate resource allocation matrix according to the gradient direction and the learning rate; The target value improvement range of the multi-objective constraint optimization function is calculated, and a final resource allocation matrix is ​​output when the target value improvement range is lower than a termination threshold for multiple consecutive times.

6. A content marketing effectiveness evaluation and optimization system based on deep learning, used to implement the method of any one of claims 1 to 5, characterized in that: include: The first unit is used to obtain historical data of content marketing, including content data and scenario data; The second unit is used to build a dual-tower deep neural network model, process the content data through the content tower to generate a content representation set, process the scene data through the scene tower to generate a scene representation set, and calculate the matching score between the content representation set and the scene representation set through the interaction layer; The dual-tower deep neural network model is trained through a contrastive learning framework of hierarchical scene perception and difficult sample mining, and is iteratively optimized in combination with a content-scene cross-attention collaborative representation mechanism; The third unit is used to design a resource allocation algorithm based on matching scores, combining nonlinear mapping and progressive two-stage optimization. Through iterative optimization, the algorithm determines the optimal matching combination of content and scenarios and the resource investment ratio for each matching combination, generating a resource allocation optimization plan. The fourth unit is used to execute the resource allocation optimization plan and collect feedback data to update the dual-tower deep neural network model.

7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Online course recommendation method based on double-tower graph convolutional neural network

    CN116992151A

  • Intelligent retrieval method and system for unstructured asset content based on large model

    CN119646243A