Advertisement data classification method and system

By adopting asymmetric autoencoding neural networks and generative artificial intelligence big models in advertising data classification, the limitations of traditional methods when processing high-dimensional sparse data are solved, and more efficient advertising data feature extraction and classification effects are achieved.

CN120145159AActive Publication Date: 2025-06-13HANGZHOU GUIYI INTELLIGENT TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510623723.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing advertising data classification methods have limitations when processing high-dimensional sparse data, making it difficult to accurately capture the implicit information in advertising copy, and model training is prone to local optimality and convergence speed is slow.

Method used

Asymmetric autoencoding neural network is used, and the connection weight matrix between the encoder and the decoder is not shared, and sparse connections are used in the encoder. The advertising data is analyzed and feature extracted in combination with a generative artificial intelligence model. The training asymmetric autoencoding neural network converts high-dimensional vectors into low-dimensional feature representations, and then inputs to the classifier for classification.

Benefits of technology

It improves the adaptability to high-dimensional advertising data, enhances the ability to identify complex patterns of advertising data, improves feature retention rate, shortens the convergence time of model training, and improves classification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145159A_ABST
    Figure CN120145159A_ABST
Patent Text Reader

Abstract

The invention provides an advertisement data classification method and system, and the method comprises the following steps: collecting advertisement data from a plurality of sources, and carrying out the preprocessing of the collected advertisement data; analyzing the preprocessed advertisement data by using a generative artificial intelligence large model, and converting the analyzed text data into a high-dimensional vector; the method comprises the following steps: constructing an asymmetric self-encoding neural network in which connection weight matrixes of an encoder and a decoder are not shared, the encoder adopts sparse connection and the decoder is full connection, and constructing a loss function to train the asymmetric self-encoding neural network; and converting the high-dimensional vector into low-dimensional feature representation by using the trained asymmetric self-encoding neural network, and then inputting the low-dimensional feature representation into a classifier for classification. Compared with the prior art, the method provided by the invention can realize more effective advertisement classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an advertisement data classification method and system. Background Art

[0002] In the digital advertising industry, precise advertisement placement depends on the effective classification of advertisement data. However, there are still many challenges in existing advertisement classification methods. First of all, the sources of advertisement data are complex, involving multiple channels such as social media, search engines, website advertisement systems, etc. It is difficult to guarantee the authenticity of data, and it is easily affected by data tampering or false advertisements. Secondly, the advertisement content is diverse, including various formats such as text, images, and videos. Traditional machine learning methods have limitations in feature extraction and semantic understanding, and it is difficult to accurately capture the implicit information in advertisement copywriting. Especially in the processing of high-dimensional sparse data, traditional dimensionality reduction methods such as principal component analysis (PCA), symmetric autoencoder neural networks, etc. are difficult to achieve a balance between feature compression and information retention, resulting in limited classification effects. In addition, current advertisement classification models generally adopt fixed learning rates or traditional gradient descent optimization strategies. Facing the complex distribution of high-dimensional advertisement data, model training is prone to falling into local optima and has a slow convergence speed.

[0003] For example, a Chinese invention patent with the publication number CN111191445A proposes an advertisement text classification method, including: obtaining the text to be classified; calculating the word vector of the text to be classified by using a trained word vector model; calculating the similarity with similar words corresponding to a given category according to the word vector of the text to be classified to obtain the similarity score of the text to be classified with each given category; and configuring the given category with the highest similarity score with the text to be classified as the classification result of the text to be classified. This method can automatically classify the text to be classified into the corresponding subcategory with the highest similarity score, so as to achieve a high degree of matching between text classification and the subcategories given by the business, and thus achieve accurate classification of advertisement text.

[0004] For another example, a Chinese invention patent with the publication number CN113220966A proposes an advertising creative classification and display method, including the following steps: obtaining advertising creatives from preset media channels and regions through web crawler technology, and constructing an original database of channel creatives; standardizing each channel creative in the original database of channel creatives to construct a standard database of creatives; parsing the advertising creative data in the standard database of creatives, performing back-tagging processing on the advertising creatives, and obtaining the advertising creative data after back-tagging processing; sorting the advertising creative data after back-tagging processing based on a preset evaluation criterion to obtain the sorted advertising creative data; and indexing based on the sorted advertising creative data according to the received retrieval instruction and returning the retrieved advertising creative data for advertising creative classification and display. This method can improve the efficiency of obtaining advertising creatives and can achieve high-quality creative recommendations.

[0005] However, the existing technologies including the above method still have the following deficiencies:

[0006] 1) In the advertising data classification task, the traditional autoencoder neural network has limitations. The lack of a symmetric structure limits its adaptability to high-dimensional sparse data.

[0007] 2) In the advertising data classification task, the traditional model has insufficient sparse control ability. Traditional L1 / L2 regularization cannot effectively control the feature sparsity.

[0008] 3) In the advertising data classification task, the traditional fixed learning rate method is inefficient, cannot adapt to different data distributions, and is prone to falling into local optima or gradient explosion. Summary of the Invention

[0009] Based on the above background, the present invention proposes an advertising data classification method and system to overcome at least one of the deficiencies existing in the above-mentioned existing technologies. The following technical solutions are specifically adopted:

[0010] The first aspect of the present invention provides an advertising data classification method, including the following steps:

[0011] Collecting advertising data from multiple sources and preprocessing the collected advertising data;

[0012] Using a generative artificial intelligence large model to parse the preprocessed advertising data and converting the parsed text data into high-dimensional vectors;

[0013] Constructing an asymmetric autoencoder neural network with non-shared connection weight matrices between the encoder and the decoder, where the encoder uses sparse connections and the decoder is fully connected, and constructing a loss function to train the asymmetric autoencoder neural network;

[0014] The trained asymmetric auto - encoding neural network is used to convert the high - dimensional vector into a low - dimensional feature representation, and then it is input into a classifier for classification.

[0015] Furthermore, the pre - processing of the collected advertising data includes:

[0016] Data cleaning, including removing invalid data, removing duplicate data, and / or filling in missing values;

[0017] Data transformation, including normalizing and standardizing the data and encoding categorical variables.

[0018] Furthermore, using the generative artificial intelligence large - model to parse the pre - processed advertising data and convert the parsed text data into high - dimensional vectors includes:

[0019] Using the generative artificial intelligence large - model to parse the advertising text, extracting the theme, sentiment tendency, target audience, and / or market positioning information of the advertisement;

[0020] Based on natural language processing technology, performing semantic understanding on the advertising text to identify the core intention of the advertising content;

[0021] Converting the parsed text data into high - dimensional vectors.

[0022] Furthermore, the asymmetric auto - encoding neural network uses learnable latent variables to strengthen the expression of complex patterns, and the specific calculation method is expressed as:

[0023]

[0024] In the formula, is the feature representation after dimensionality reduction; is the weight matrix, representing the connection strength between neurons in the asymmetric auto - encoding neural network; is the input data, representing the vectorized advertising copy data; is the interaction - term weighting coefficient, representing the strength of the influence of latent variables; is the latent variable matrix, representing the automatic learning of complex patterns or associated features; is the bias term of the asymmetric auto - encoding neural network; is the non - linear activation function, representing the non - linear transformation of the mapping result, and its calculation method is expressed as:

[0025]

[0026] In the formula, is the exponential function, representing the transformation of exponentially amplifying the input value; is the input signal of the current neuron.

[0027] Furthermore, a loss function is constructed by combining the reconstruction error and the sparsity regularization term to train the asymmetric autoencoder neural network. The expression of the constructed loss function is as follows:

[0028]

[0029] In the formula, is the total loss function of the asymmetric autoencoder neural network; is the total number of samples, representing the scale of the copywriting samples available for training; is the th input data; represents the L2 distance between the th data sample in the input and output; is the regularization strength parameter; is the th power absolute value of the th dimensionality-reduced feature, is a positive integer; is the regularization coefficient, representing the influence of the additional entropy term in the sparsity constraint; represents the entropy-based sparsity measure; is the variance of the current batch of data, representing the degree of dispersion of the distribution of the batch of advertising copywriting vectors; is the skewness adjustment factor, representing the correction strength of the distribution skewness; is the sparsity regularization term, and its calculation method is expressed as:

[0030]

[0031] In the formula, is the number of features after dimensionality reduction, representing the dimension of the final low-dimensional space.

[0032] Furthermore, training the asymmetric autoencoder neural network includes:

[0033] Set the number of network layers and the number of neurons in each layer of the asymmetric autoencoder neural network, and perform parameter initialization operations;

[0034] During the training process, gradually calculate the loss function and update the asymmetric autoencoder neural network according to its gradients of the weight matrix, bias term, and latent variable matrix, and adaptively decay the learning rate with reference to the simulated annealing process;

[0035] After each training epoch ends, adjust the parameter update direction using a dynamic correction mechanism;

[0036] Repeat the above steps iteratively until the preset stop iteration condition is met, and complete the training.

[0037] Further, the parameter initialization operation is performed in the following manner:

[0038]

[0039] wherein, represents a symbol subject to a specific distribution; is a normal distribution with a mean of 0 and a variance of characterizing the random generation of the initial weights; is an identity matrix, characterizing the basic linear transformation that does not change the vector direction.

[0040] Further, during the training process, the weight matrix of the asymmetric autoencoder neural network is updated based on the momentum mechanism and combined with gradient accumulation, and the calculation method is expressed as:

[0041]

[0042] wherein, is the corrected weight update amount, characterizing the additional correction to the current weight; is the weight matrix of the th iteration; is the weight matrix of the th iteration; is the gradient of the loss function with respect to the weight matrix, characterizing the update direction of the advertisement copy reconstruction loss and the sparse regularization for the asymmetric autoencoder neural network; is the momentum term weighting coefficient, characterizing the influence proportion of the momentum on this update; is the momentum term of the th iteration, and its update calculation method is expressed as:

[0043]

[0044] wherein, is the momentum decay coefficient, characterizing the retention proportion of the historical gradient; is the momentum term of the th iteration.

[0045] Further, during the training process, the calculation method for iteratively updating the bias term of the asymmetric autoencoder neural network is:

[0046]

[0047] wherein, is the bias vector of the th iteration; is the bias vector of the th iteration; is the learning rate of the asymmetric autoencoder neural network; It is the gradient of the loss function of the asymmetric auto-encoding neural network with respect to the bias, representing the optimization direction of this bias term in the dimensionality reduction objective.

[0048] Furthermore, the calculation method for iteratively updating the latent variable matrix of the asymmetric auto-encoding neural network during the training process is expressed as:

[0049]

[0050] In the formula, is the latent variable matrix at the -th iteration; is the latent variable matrix at the -th iteration; is the learning rate of the asymmetric auto-encoding neural network, representing the update step size of the parameters of the asymmetric auto-encoding neural network for each iteration; is the gradient of the loss function of the asymmetric auto-encoding neural network with respect to the latent variable matrix, representing the optimization direction of the vectorized advertisement copy data in the latent space.

[0051] Furthermore, the calculation method for adaptively decaying the learning rate with reference to the simulated annealing process during the training process is:

[0052]

[0053] In the formula, is the learning rate at the -th iteration; is the diagonal approximation of the Hessian matrix; is the learning rate adjustment coefficient, representing the decay speed of the learning rate with the number of iterations.

[0054] Furthermore, the calculation method for adjusting the parameter update direction by using a dynamic correction mechanism is:

[0055]

[0056] In the formula, is the sign function, representing symbolizing the weight direction to enhance the non-linear adjustment; is the decay function based on the time step, representing gradually reducing the correction intensity during the iteration process; is the correction term adjustment coefficient.

[0057] Furthermore, the classifier is any one of logistic regression, support vector machine, random forest, or decision tree.

[0058] The second aspect of the present invention provides an advertisement data classification system for implementing the advertisement data classification method as described in the first aspect above, including:

[0059] A data acquisition module, which is used to collect advertisement data from multiple sources and preprocess the collected advertisement data;

[0060] A data processing module, which is used to parse the preprocessed advertisement data using a generative artificial intelligence large model, and convert the parsed text data into high-dimensional vectors; and use a trained asymmetric autoencoder neural network to convert the high-dimensional vectors into low-dimensional feature representations, and then input them into a classifier for classification;

[0061] A blockchain evidence storage and verification module, which is used to store the collected data on the blockchain to ensure the credibility of the model training data, and perform verifiable storage on the model inference results;

[0062] And a system interaction and visualization module, which is used to provide a user interface, display the classification results, and support users to query and adjust the model parameters.

[0063] The beneficial technical effects of the present invention are as follows:

[0064] 1) The adopted asymmetric autoencoder neural network adopts an asymmetric coding structure, enhancing the adaptability to high-dimensional advertisement data.

[0065] 2) The adopted asymmetric autoencoder neural network combines a latent variable matrix, improving the ability to recognize complex patterns in advertisement data.

[0066] 3) The adopted asymmetric autoencoder neural network combines sparsity-induced regularization and adopts a composite regularization strategy, improving the retention rate of key features in advertisement data.

[0067] 4) The adopted asymmetric autoencoder neural network adopts adaptive learning rate optimization, combines local curvature perception and dynamic adjustment strategies, improving the convergence speed and stability. Description of the Drawings

[0068] Figure 1 It is a schematic flowchart of an embodiment of the advertisement data classification method of the present invention.

[0069] Figure 2 It is a schematic diagram of the result of Verification Experiment 1 in the embodiment of the present invention.

[0070] Figure 3 It is a schematic diagram of the result of Verification Experiment 2 in the embodiment of the present invention.

[0071] Figure 4 It is a schematic diagram of the result of Verification Experiment 3 in the embodiment of the present invention.

[0072] Figure 5 It is a schematic diagram of the result of Verification Experiment 4 in the embodiment of the present invention.

[0073] Figure 6 This is a schematic diagram of the results of Verification Experiment 5 in the embodiments of the present invention. Detailed implementation manners

[0074] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0075] Refer to Figure 1 , embodiments of the present invention provide an advertising data classification method, including the following steps:

[0076] S1. Collect advertising data from multiple sources and preprocess the collected advertising data.

[0077] Specifically, advertising data can be collected in real time from channels such as social media platforms, search engine advertisements, and website advertising systems through web crawler technology and API interfaces, and at the same time, blockchain technology is used to store evidence of the data to ensure the traceability and immutability of the data source.

[0078] As a preferred implementation manner, in this embodiment, preprocessing the collected advertising data includes:

[0079] Data cleaning, including removing invalid data, removing duplicate data, and / or filling in missing values;

[0080] Data transformation, including normalizing and standardizing the data and encoding categorical variables.

[0081] Furthermore, keyword extraction, sentiment analysis, user behavior feature construction, etc. can also be performed through feature engineering to enhance the usability and effectiveness of the data.

[0082] S2. Use a generative artificial intelligence large model to parse the preprocessed advertising data and convert the parsed text data into high-dimensional vectors.

[0083] As a preferred implementation manner, in this embodiment, it specifically includes the following processes:

[0084] Use a generative artificial intelligence large model to parse the advertising text, extract the theme, sentiment tendency, target audience, and / or market positioning information of the advertisement;

[0085] Based on natural language processing technology, perform semantic understanding on the advertising text to identify the core intention of the advertising content;

[0086] Convert the parsed text data into high-dimensional vectors.

[0087] Since it is an existing technical means in the field to process and convert text data using generative artificial intelligence large models, no specific elaboration will be made here.

[0088] S3. Construct an asymmetric autoencoder neural network where the connection weight matrices of the encoder and decoder are not shared, the encoder uses sparse connections (e.g., no connections between some neurons), the decoder is a fully connected one, and construct a loss function to train the asymmetric autoencoder neural network.

[0089] In this embodiment, the asymmetric autoencoder neural network refers to, different from the traditional autoencoder neural network (composed of symmetric encoder and decoder), for the complex distribution modeling requirements of high-dimensional sparse data such as advertising copy vectors. For example, advertising copy data usually has high-dimensional sparse features (such as the bag-of-words model). The asymmetric structure can specifically design the encoder to capture potential complex patterns in the data (such as the implicit associations of keyword combinations) through a deeper network, and the decoder is simplified to avoid overfitting. The asymmetric autoencoder neural network breaks the symmetric structure, and its asymmetry is reflected in two aspects:

[0090] 1) Structural asymmetry: The number of layers, the number of neurons, or the connection methods of the encoder and decoder are different. In a specific example, the encoder may be deeper (e.g., 5 layers), while the decoder is shallower (e.g., 4 layers) to adapt to the compression characteristics from high dimension to low dimension.

[0091] 2) Parameter asymmetry: The weight matrices of the encoder and decoder are not shared, and the encoder uses additional latent variables (latent variable matrix ) to enhance the expression ability of non-linear relationships.

[0092] As a preferred implementation, in this embodiment, the sparse connection method is realized through a sparsity matrix. In the initialization stage, a matrix with the same dimension as the weight matrix of the asymmetric autoencoder neural network is generated, that is, the sparsity matrix , the elements of which follow a Bernoulli distribution. The sparsity matrix forces some weights to be 0 through element-wise multiplication , and at the same time, under the constraint of the sparsity regularization term, key features are screened out, and the model is enabled to adaptively adjust the sparse pattern during the training process.

[0093] As a preferred implementation, in this embodiment, the asymmetric autoencoder neural network uses learnable latent variables to strengthen the expression of complex patterns, and the specific calculation method is expressed as:

[0094]

[0095] In the formula, is the feature representation after dimensionality reduction; is the weight matrix, representing the connection strength between neurons in the asymmetric autoencoder neural network; is the input data, representing the vectorized advertisement copy data; is the interaction term weighting coefficient, representing the strength of the influence of latent variables. Preferably, is set to 0.1; is the latent variable matrix, representing the automatic learning of complex patterns or associated features; is the bias term of the asymmetric autoencoder neural network; is the non-linear activation function, representing the non-linear transformation of the mapping result, and its calculation method is expressed as:

[0096]

[0097] In the formula, is the exponential function, representing the transformation of exponentially amplifying the input value; is the input signal of the current neuron.

[0098] In this embodiment, the latent variable matrix is inspired by the latent variable idea in the probabilistic graphical model and is used to model the potential associations in the data. As a learnable low-rank matrix, it captures the implicit relationships between advertisement copies (such as the correlation between "price" and "promotion") through iterative updates. By directly incorporating the latent variable as an optimizable parameter (rather than a random variable) into the forward calculation of the autoencoder, the training process is simplified.

[0099] In this embodiment, in order to fully model the non-linear relationships in the advertisement copy vectors, a non-linear activation function based on the principle of matter expansion in physics is adopted, enabling the activation function to maintain the smoothness of the output while capturing the irregular feature distributions of the advertisement copy data, and enhancing the modeling ability for the long-tail distribution (dominated by a few keywords) in the advertisement copy.

[0100] Furthermore, the asymmetric autoencoder neural network is trained by combining the reconstruction error and the sparsity regularization term to construct a loss function. The expression of the constructed loss function is as follows:

[0101]

[0102] In the formula, is the total loss function of the asymmetric autoencoder neural network; is the total number of samples, representing the scale of the copy samples available for training; is the th input data; represents the th L2 distance between the input and output of the is the regularization strength parameter. Preferably, it is set to 0.3; is the power absolute value of the th dimensionality reduction feature. Preferably, it is set to 2, where is a positive integer; is the regularization coefficient, representing the influence of the additional entropy term in the sparsity constraint. Preferably, it is set to 0.5; represents the entropy-based sparsity measure; is the variance of the current batch of data, representing the degree of dispersion of the distribution of the advertisement copy vectors in this batch; is the skewness adjustment factor, representing the correction strength for the distribution skewness. Preferably, it is set to 2, and is used to dynamically adjust the loss weights under different data distributions; is the sparsity regularization term, which is used to select the most discriminative features for advertisement copy classification or key descriptions, thereby suppressing redundant and noise features and improving the generalization performance and interpretability of the dimensionality reduction model. The calculation method is expressed as:

[0103]

[0104] In the formula, is the number of features after dimensionality reduction, representing the dimension of the final low-dimensional space.

[0105] As a preferred implementation, in this embodiment, training the asymmetric autoencoder neural network includes:

[0106] Setting the number of network layers and the number of neurons in each layer of the asymmetric autoencoder neural network, and performing parameter initialization operations;

[0107] During the training process, gradually calculate the loss function and update the asymmetric autoencoder neural network according to its gradients of the weight matrix, bias term, and latent variable matrix, and adaptively decay the learning rate with reference to the simulated annealing process;

[0108] After each training epoch ends, adjust the parameter update direction using a dynamic correction mechanism;

[0109] Repeat the above steps iteratively until the preset stop iteration condition is met, and complete the training.

[0110] Furthermore, the parameter initialization operation is performed in the following manner:

[0111]

[0112] In the formula, A symbol subject to a specific distribution; is a normal distribution with a mean of 0 and a variance of , representing the random generation of the initial weights; is the identity matrix, representing the basic linear transformation that does not change the direction of the vector.

[0113] Furthermore, during the training process, the weight matrix of the asymmetric autoencoder neural network is updated according to the momentum mechanism and combined with gradient accumulation, and the calculation method is expressed as:

[0114]

[0115] In the formula, is the corrected weight update amount, representing the additional correction to the current weight; is the weight matrix of the th iteration; is the weight matrix of the th iteration; is the gradient of the loss function with respect to the weight matrix, representing the update direction of the advertisement copy reconstruction loss and the sparse regularization to the asymmetric autoencoder neural network; is the momentum term weighting coefficient, representing the influence ratio of the momentum on this update. Preferably, is set to 0.4; is the momentum term of the th iteration, and its update calculation method is expressed as:

[0116]

[0117] In the formula, is the momentum decay coefficient, representing the retention ratio of the historical gradient. Preferably, is set to 0.3; is the momentum term of the th iteration.

[0118] Furthermore, during the training process, the calculation method for iteratively updating the bias term of the asymmetric autoencoder neural network is:

[0119]

[0120] In the formula, is the bias vector of the th iteration; is the bias vector of the th iteration; is the learning rate of the asymmetric autoencoder neural network; is the gradient of the loss function of the asymmetric autoencoder neural network with respect to the bias, representing the optimization direction of this bias term in the dimensionality reduction target.

[0121] Furthermore, the calculation method for iteratively updating the latent variable matrix of the asymmetric auto-encoding neural network during the training process is expressed as:

[0122]

[0123] In the formula, is the latent variable matrix at the -th iteration; is the latent variable matrix at the -th iteration; is the learning rate of the asymmetric auto-encoding neural network, representing the update step size of the parameters of the asymmetric auto-encoding neural network in each iteration. Preferably, is set to 0.01; is the gradient of the loss function of the asymmetric auto-encoding neural network with respect to the latent variable matrix, representing the optimization direction of the vectorized advertising copy data in the latent space.

[0124] Furthermore, during the training process, the calculation method for adaptively decaying the learning rate with reference to the simulated annealing process is:

[0125]

[0126] In the formula, is the learning rate at the -th iteration; is the diagonal approximation of the Hessian matrix; is the learning rate adjustment coefficient, representing the decay speed of the learning rate with the number of iterations.

[0127] Furthermore, after each cycle of the asymmetric auto-encoding neural network training, a dynamic correction mechanism is adopted to adjust the parameter update direction. By combining the current gradient and historical error information, precise control is exerted over the update path in the high-dimensional copy feature space to prevent getting stuck in local optima and reduce oscillations. The calculation method is expressed as:

[0128]

[0129] In the formula, is the sign function, representing the symbolic processing of the weight direction to enhance the non-linear adjustment; is the decay function based on the time step, representing the gradual reduction of the correction intensity during the iteration process; is the correction term adjustment coefficient.

[0130] Repeat the above steps iteratively until the preset iteration stop condition is met, which indicates that the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0131] S4. Use the trained asymmetric autoencoder neural network to convert the high-dimensional vector into a low-dimensional feature representation, and then input it into a classifier for classification.

[0132] Optionally, in this embodiment, the classifier is any one of logistic regression, support vector machine, random forest, or decision tree. Since classifying low-dimensional features using a classifier is a conventional technical means in the art, it will not be specifically elaborated here.

[0133] The following further illustrates the technical advantages of the technical solution disclosed in the present invention (referred to as the present technology in the figure) through several experimental data.

[0134] Experiment 1

[0135] This experiment aims to verify the improvement effect of the asymmetric structure design on the feature selection ability of advertising copywriting, and compare the performance of conventional dimensionality reduction methods such as principal component analysis, traditional symmetric autoencoders, and manifold learning algorithms and the solution of the present invention in terms of the discrimination degree of feature importance. See Figure 2 , the experimental results show that the method of the present invention has significantly better discrimination in deep features such as semantic coherence than other methods, especially showing stronger pattern recognition ability when capturing the implicit associations of keyword combinations in advertising copywriting. The asymmetric structure in the present invention effectively strengthens the adaptability to high-dimensional sparse data distribution through the cooperation of deep modeling of the encoder and simplified design of the decoder, enabling the reduced-dimensional feature space to more accurately reflect the core elements of advertising semantics, while traditional methods are difficult to balance the relationship between feature compression and information retention due to the limitations of the symmetric structure.

[0136] Experiment 2

[0137] This experiment focuses on comparing the influence of sparsity-induced regularization and traditional sparse constraint methods on feature selection, and comparing the retention effects of key features under different regularization strengths. See Figure 3 , the experimental results show that the composite regularization strategy proposed by the method of the present invention significantly improves the retention rate of high-value information such as core keywords and sentiment intensity while suppressing noise features. The single sparse constraint adopted by traditional methods easily leads to excessive feature pruning, while the present invention realizes the coordinated control of feature sparsity and distribution irregularity through entropy constraint and dynamic weighting mechanism, enabling rare keywords with long-tailed distributions in advertising copywriting to be effectively expressed during the dimensionality reduction process, enhancing the practicality of the model in real business scenarios.

[0138] Experiment 3

[0139] This experiment compares the convergence process of the adaptive learning rate strategy of the present invention with that of the traditional fixed learning rate training to verify the role of local curvature perception and dynamic adjustment mechanism in improving training efficiency. Figure 4 , experimental data show that the method of the present invention is superior to the traditional method in terms of loss reduction speed and stability, especially in the early stage of training, it quickly avoids the high gradient oscillation area through curvature perception, and achieves smooth convergence in the middle and late stages by combining learning rate decay. Traditional methods are prone to fall into local optimality or gradient explosion due to their lack of adaptability to the distribution characteristics of high-dimensional data. The present invention uses a simulated annealing parameter update strategy to enable the model to autonomously adjust the optimization path according to the geometric characteristics of the feature space, significantly shortening the training cycle required to achieve stable convergence.

[0140] Experiment 4

[0141] This experiment uses heat map visualization to compare the distribution differences of reconstruction errors of different methods in high-dimensional space, and verifies the effect of dynamic correction mechanism on enhancing the generalization ability of the model. Figure 5 , Experimental data show that the reconstruction error of the method of the present invention is more uniform in spatial distribution and has a lower overall level, indicating that it can better maintain the topological structure of the original data. The traditional method lacks a correction mechanism in the parameter update direction, resulting in abnormal peaks in the reconstruction error in a specific dimension. The present invention effectively suppresses the direction deviation in the parameter update process through the dynamic fusion of gradient sign correction and historical error feedback, allowing the model to steadily optimize along the optimal path of the loss surface, thereby achieving more accurate low-dimensional mapping on complex high-dimensional advertising data.

[0142] Experiment 5

[0143] This experiment analyzes the role of the dynamic correction mechanism in improving the efficiency of high-dimensional parameter space exploration by visually comparing the spatial characteristics of parameter optimization paths. Figure 6 Compared with the disordered oscillations and path wandering exhibited by traditional gradient descent methods in parameter space, the optimization trajectory of the method of the present invention shows obvious direction correction and path focusing characteristics, especially in areas where the curvature changes dramatically, it can autonomously adjust the update step size and direction. The experimental results show that the parameter update trajectory of the present invention stably approaches the optimal area in the form of spiral convergence in three-dimensional space, while the trajectory of the traditional method has a large number of redundant exploration paths due to the lack of perception of the geometric characteristics of the parameter space. The experimental results show that the dynamic correction mechanism, by integrating historical gradient information and curvature perception, gives the model the ability to autonomously plan the optimal path in a complex parameter space, effectively avoiding the convergence delay problem caused by direction offset in traditional methods.

[0144] The second embodiment of the present invention also provides an advertising data classification system for implementing the advertising data classification method as described in the first embodiment above, including:

[0145] A data collection module for collecting advertising data from multiple sources and preprocessing the collected advertising data;

[0146] A data processing module for parsing the preprocessed advertising data using a generative artificial intelligence large model and converting the parsed text data into high-dimensional vectors; and converting the high-dimensional vectors into low-dimensional feature representations using a trained asymmetric autoencoder neural network, and then inputting them into a classifier for classification;

[0147] A blockchain evidence storage and verification module for storing the collected data on the blockchain to ensure the credibility of the model training data and storing the model inference results in a verifiable manner;

[0148] And a system interaction and visualization module for providing a user interface, displaying the classification results, and supporting users to query and adjust the model parameters.

[0149] Specifically, the functions of the blockchain evidence storage and verification module include data traceability, operation record storage, and result verification. Among them, data traceability enables data to be traceable by storing data fingerprints on the chain; operation record storage is used to record all key steps such as data collection, preprocessing, large model parsing, and machine learning modeling to ensure the reliability of system operations; result verification ensures the integrity and consistency of the inference results through the blockchain consensus mechanism.

[0150] The functions of the system interaction and visualization module include data query, result display, and model tuning interface. Among them, data query supports users to retrieve historical data based on the blockchain; result display presents the advertising classification results and credibility evaluation in a visual manner; the model tuning interface allows users to adjust hyperparameters, update training data, and observe changes in model performance.

[0151] It should be noted that the method of the embodiment of the present invention can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of these multiple devices can only execute one or more steps of the method of the embodiment of the present invention, and these multiple devices will interact with each other to complete the described method.

[0152] The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for classifying advertisement data, characterized in that: The steps include: Collect advertising data from various sources and pre-process the collected advertising data; Use a generative AI big model to parse the pre-processed advertising data and convert the parsed text data into a high-dimensional vector; Constructing a connection weight matrix between the encoder and the decoder that is not shared, and the encoder adopts sparse connection, the decoder is a fully connected asymmetric autoencoder neural network, and constructing a loss function to train the asymmetric autoencoder neural network; The high-dimensional vector is converted into a low-dimensional feature representation using a trained asymmetric autoencoder neural network, and then input into a classifier for classification.

2. The advertisement data classification method according to claim 1, characterized in that: The preprocessing of the collected advertising data includes: Data cleaning, including removing invalid data, removing duplicate data, and / or filling in missing values; Data transformation, including normalization and standardization of data and encoding of categorical variables.

3. The advertisement data classification method according to claim 1, characterized in that: The method of using a generative artificial intelligence model to parse the preprocessed advertisement data and converting the parsed text data into a high-dimensional vector includes: Use generative AI big models to parse ad texts and extract ad themes, sentiment, target audiences, and / or market positioning information; Based on natural language processing technology, the advertising text is semantically understood to identify the core intent of the advertising content; Convert the parsed text data into a high-dimensional vector.

4. The advertisement data classification method according to claim 1, characterized in that: The asymmetric autoencoder neural network uses learnable hidden variables to enhance the expression of complex patterns. The specific calculation method is expressed as: In the formula, is the feature representation after dimensionality reduction; is the weight matrix, which represents the connection strength between neurons in the asymmetric autoencoder neural network; The input data represents the vectorized advertising copy data; is the weighted coefficient of the interaction term, which represents the strength of the influence of the latent variable; is a latent variable matrix, representing the automatic learning of complex patterns or associated features; is the bias term of the asymmetric autoencoder neural network; is a nonlinear activation function, which represents the nonlinear transformation of the mapping result. Its calculation method is expressed as: In the formula, is an exponential function, representing a transformation that exponentially amplifies the input value; is the input signal of the current neuron.

5. The advertisement data classification method according to claim 4, characterized in that: The loss function is constructed by combining the reconstruction error and the sparsity regularization term to train the asymmetric autoencoder neural network. The constructed loss function expression is as follows: In the formula, is the total loss function of the asymmetric autoencoder neural network; is the total number of samples, representing the size of the copywriting samples that can be used for training; For the Input data; Characterization The L2 distance between the input and output of a data sample; is the regularization strength parameter; For the dimensionality reduction features The absolute value of the power, is a positive integer; is the regularization coefficient, which represents the influence of the additional entropy term in the sparsity constraint; Characterizing entropy-based sparsity measures; is the variance of the current batch of data, representing the degree of dispersion of the distribution of the batch of advertising copy vectors; is the skewness adjustment factor, which represents the correction strength of the distribution skewness; is the sparsity regularization term, and its calculation method is expressed as: In the formula, is the number of features after dimensionality reduction, representing the dimension of the final low-dimensional space.

6. The advertisement data classification method according to claim 5, characterized in that: Training the asymmetric autoencoder neural network includes: Setting the number of network layers and the number of neurons in each layer of the asymmetric autoencoder neural network, and performing parameter initialization operations; During the training process, the loss function is calculated step by step and the asymmetric autoencoder neural network is updated according to its gradient to the weight matrix, bias term and latent variable matrix, and the learning rate is adaptively decayed with reference to the simulated annealing process; After each training cycle, a dynamic correction mechanism is used to adjust the parameter update direction; Repeat the above steps until the preset stop iteration condition is met and the training is completed.

7. The advertisement data classification method according to claim 6, characterized in that: The parameter initialization operation is performed in the following manner: In the formula, Symbols that indicate compliance with a particular distribution; The mean is 0 and the variance is The normal distribution of represents the random generation of initial weights; is the identity matrix, representing basic linear transformations that do not change the direction of a vector.

8. The advertisement data classification method according to claim 6, characterized in that: During the training process, the weight matrix of the asymmetric autoencoder neural network is updated based on the momentum mechanism and combined with gradient accumulation. The calculation method is expressed as: In the formula, is the modified weight update amount, representing the additional correction to the current weight; For the The weight matrix of the iteration; For the The weight matrix of the iteration; is the gradient of the loss function with respect to the weight matrix, representing the update direction of the asymmetric autoencoder neural network by the advertisement reconstruction loss and sparse regularization; is the weighted coefficient of the momentum term, which represents the influence ratio of momentum on this update; For the The momentum term of the iteration is updated as follows: In the formula, is the momentum decay coefficient, which represents the retention ratio of the historical gradient; For the The momentum term for the iteration.

9. The advertisement data classification method according to claim 6, characterized in that: The calculation method for iteratively updating the bias term of the asymmetric autoencoder neural network during training is: In the formula, For the The bias vector for the iteration; For the The bias vector for the iteration; is the learning rate of the asymmetric autoencoder neural network; It is the gradient of the loss function of the asymmetric autoencoder neural network with respect to the bias, which represents the optimization direction of the bias term in the dimensionality reduction objective.

10. The advertisement data classification method according to claim 6, characterized in that: The calculation method for iteratively updating the latent variable matrix of the asymmetric autoencoder neural network during training is expressed as: In the formula, For the The latent variable matrix of the iteration; For the The latent variable matrix of the iteration; is the learning rate of the asymmetric autoencoder neural network, representing the update step size of the asymmetric autoencoder neural network parameters in each iteration; It is the gradient of the loss function of the asymmetric autoencoder neural network with respect to the latent variable matrix, representing the optimization direction of the vectorized advertising copy data in the latent space.

11. The advertisement data classification method according to claim 6, characterized in that: The calculation method for adaptively decaying the learning rate during training is as follows: In the formula, For the The learning rate at the iteration; is the diagonal approximation of the Hessian matrix; is the learning rate adjustment coefficient, which represents the decay rate of the learning rate with the number of iterations.

12. The advertisement data classification method according to claim 8, characterized in that: The calculation method of using the dynamic correction mechanism to adjust the parameter update direction is: In the formula, is the symbolic function, Representation symbolizes the weight direction to enhance nonlinear regulation; is a decay function based on the time step, representing the gradual reduction of the correction strength during the iteration process; is the correction term adjustment coefficient.

13. The advertisement data classification method according to claim 1, characterized in that: The classifier is any one of logistic regression, support vector machine, random forest or decision tree.

14. An advertisement data classification system, used to implement the advertisement data classification method according to any one of claims 1 to 13, characterized in that: include: A data collection module, used to collect advertising data from various sources and pre-process the collected advertising data; A data processing module is used to parse the pre-processed advertising data using a generative artificial intelligence model and convert the parsed text data into a high-dimensional vector; and to convert the high-dimensional vector into a low-dimensional feature representation using a trained asymmetric autoencoder neural network, and then input it into a classifier for classification; The blockchain evidence storage and verification module is used to store the collected data on the blockchain, ensure the credibility of the model training data, and store the model reasoning results in a verifiable manner; And the system interaction and visualization module is used to provide a user interface, display classification results, and support users to query and adjust model parameters.

Citation Information

Patent Citations

  • Advertisement text classification method and device

    CN111191445A

  • Advertisement creative classification display method and system, equipment and readable storage medium

    CN113220966A

  • Multi-omics and phenotype association mining method based on interpretable auto-encoder

    CN115691677A

  • Data compression method and related equipment

    CN116095183A

  • Real-time steam generator state monitoring method based on artificial intelligence

    CN119150079A