Method, device, computer storage medium and terminal for determining visitor search volume

By constructing generator and discriminator models and combining them with reverse verification, the problem of attribution analysis of independent visitor numbers for soft and hard advertising channels was solved, enabling accurate analysis of various information delivery channels and improving the coverage and accuracy of attribution analysis.

CN115936780BActive Publication Date: 2026-05-01BEIJING MININGLAMP SOFTWARE SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING MININGLAMP SOFTWARE SYST CO LTD
Filing Date
2022-10-09
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively attribute the number of unique visitors to soft and hard advertising channels. Market mix models are applicable to soft advertising channels but perform poorly in hard advertising channels.

Method used

By constructing generator and discriminator models and using reverse validation, combined with relevant variable data and interaction conversion data of information delivery, an attribution analysis model suitable for both soft and hard advertising is built to determine the number of unique visitors for each information delivery channel.

Benefits of technology

It enables accurate analysis of unique visitor counts for both soft and hard advertising channels, improving the coverage and accuracy of attribution analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115936780B_ABST
    Figure CN115936780B_ABST
Patent Text Reader

Abstract

The application discloses a method, device, computer storage medium and terminal for determining visitor search volume. Embodiments of the application construct a model suitable for attribution analysis of soft advertising and hard advertising by a reverse verification method based on first related variable data, first interactive conversion data and a first actual value of independent visitors, and realize analysis of the number of independent visitors of each information delivery channel based on the constructed model for attribution analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, devices, computer storage media and terminals for determining visitor search volume Technical Field

[0001] This article relates to, but is not limited to, data analysis techniques, particularly a method, apparatus, computer storage medium, and terminal for determining visitor search volume. Background Technology

[0002] In recent years, with the popularization of the internet, internet advertising, which closely connects consumers, brands, and products, has become increasingly popular among advertisers and has experienced rapid development. From 2010 to 2018, the overall size of the digital marketing market nearly increased sixfold. Advertisers place ads across multiple channels (TPs, also known as Touch Points), including hard advertising methods such as TV commercials, ad inserts, and video interstitials, as well as soft advertising methods such as social media platforms like Douyin and Xiaohongshu, and inviting key opinion leaders (KOLs) like live streamers to promote products. After seeing the ad, users may visit the corresponding flagship store on e-commerce platforms (such as JD.com, Tmall, Vipshop, etc.) to search for the advertised product. After obtaining the Search Unique Visitors (UV) for each e-commerce platform, advertisers want to evaluate the growth value of each channel's Search UV for the store, i.e., attribution analysis of omnichannel unique visitors.

[0003] Attribution analysis algorithms in related technologies use user browsing data (landscape) to build time series models, thereby estimating the contribution of each channel. In addition, the Marketing Mix Model is another method for estimating the contribution of each channel; this method is mostly used in attribution analysis of soft advertising channels and is not applicable to attribution analysis of hard advertising channels, which differ from soft advertising.

[0004] In summary, how to achieve attribution analysis of unique visitors that can cover both soft and hard advertising has become an unsolved problem. Summary of the Invention

[0005] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.

[0006] This invention provides a method, apparatus, and computer storage medium for determining visitor search volume, which can construct a model for attribution analysis applicable to both soft and hard advertising, and realize the determination of the number of unique visitors.

[0007] This invention provides a method for determining visitor search volume, comprising:

[0008] Input the pre-built training dataset containing the input data and the actual value of the first independent visitor count into the preset generator to obtain the reconstructed data of the input data and the predicted value of the first independent visitor count;

[0009] Each input data and each reconstructed data is treated as a sample to construct a sample dataset. The constructed data sample dataset is then input into a preset discriminator to obtain judgment information on whether the samples in the prediction data sample dataset are real samples.

[0010] Based on the training dataset, the constructed data sample set, the predicted number of first independent visitors, and the judgment information, the model for attribution analysis, consisting of a generator and a discriminator, is determined by reverse verification.

[0011] Based on the established model used for attribution analysis, determine the number of unique visitors for each information delivery channel;

[0012] The input data includes: first relevant variable data for information delivery and first interactive conversion data for information delivery.

[0013] On the other hand, embodiments of the present invention also provide a computer storage medium storing a computer program, which, when executed by a processor, implements the above-described method for determining visitor search volume.

[0014] Furthermore, embodiments of the present invention also provide a terminal, comprising: a memory and a processor, wherein the memory stores a computer program; wherein,

[0015] The processor is configured to execute computer programs in memory;

[0016] When the computer program is executed by the processor, it implements the method for determining visitor search volume as described above.

[0017] Furthermore, embodiments of the present invention also provide an apparatus for determining visitor search volume, comprising: a first unit, a second unit, a determining unit, and a processing unit; wherein,

[0018] The first unit is set to: input a pre-constructed training dataset containing input data and the actual value of the first independent visitor count into a preset generator to obtain the reconstructed data of the input data and the predicted value of the first independent visitor count;

[0019] The second unit is configured to: treat each input data and each reconstructed data as a sample, construct a sample dataset, and input the constructed data sample set into a preset discriminator to obtain judgment information on whether the samples in the predicted data sample set are real samples;

[0020] The determination unit is set as follows: based on the training dataset, the constructed data sample set, the first independent visitor number prediction value and judgment information, the model for attribution analysis, consisting of a generator and a discriminator, is determined by reverse verification method.

[0021] The processing unit is set to determine the number of unique visitors for each information delivery channel based on the established model used for attribution analysis.

[0022] The input data includes first relevant variable data for information delivery and first interactive conversion data for information delivery.

[0023] The technical solution of this application includes: inputting a pre-constructed training dataset containing input data and actual values ​​of the first independent visitor count into a preset generator to obtain reconstructed data of the input data and predicted values ​​of the first independent visitor count; constructing a sample dataset by treating each input data and each reconstructed data as a sample, and inputting the constructed data sample set into a preset discriminator to obtain judgment information on whether the samples in the predicted data sample set are real samples; determining a model for attribution analysis composed of the generator and the discriminator using a reverse verification method based on the training dataset, the constructed data sample set, the predicted values ​​of the first independent visitor count, and the judgment information; determining the independent visitor count information of each information delivery channel based on the determined model for attribution analysis; wherein, the input data includes: first relevant variable data of information delivery and first interaction conversion data of information delivery. This embodiment of the invention, based on the first relevant variable data, the first interaction conversion data, and the actual values ​​of the first independent visitor count, constructs an attribution analysis model suitable for both soft and hard advertising using a reverse verification method, and based on the constructed model for attribution analysis, realizes the analysis of the independent visitor count of each information delivery channel.

[0024] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description, claims, and drawings. Attached Figure Description

[0025] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of the present invention and do not constitute a limitation on the technical solutions of the present invention.

[0026] Figure 1 is a flowchart of a method for determining visitor search volume according to an embodiment of the present invention;

[0027] Figure 2 is a structural block diagram of the device for determining visitor search volume according to an embodiment of the present invention;

[0028] Figure 3 is a flowchart of the method for an application example of the present invention;

[0029] Figure 4 is a system schematic diagram of attribution analysis in this application example. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

[0031] The steps illustrated in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases the steps shown or described may be performed in a different order than that presented here.

[0032] Figure 1 is a flowchart of a method for determining visitor search volume according to an embodiment of the present invention. As shown in Figure 1, it includes:

[0033] Step 101: Input the pre-constructed training dataset containing the input data and the actual value of the first independent visitor count into the preset generator to obtain the reconstructed data of the input data and the predicted value of the first independent visitor count; wherein, the input data includes: the first relevant variable data of information delivery and the first interaction conversion data of information delivery.

[0034] Step 102: Treat each input data and each reconstructed data as a sample to construct a sample dataset, and input the constructed data sample set into a preset discriminator to obtain judgment information on whether the samples in the prediction data sample set are real samples;

[0035] Step 103: Based on the training dataset, the constructed data sample set, the predicted number of first independent visitors, and the judgment information, determine the model for attribution analysis consisting of a generator and a discriminator using the reverse verification method.

[0036] It should be noted that the first relevant variable data, the first interaction conversion data, and the actual value of the first number of unique visitors in the embodiments of the present invention refer to the data in the training dataset used for model training.

[0037] In one exemplary instance, information delivery in this embodiment of the invention includes advertising delivery.

[0038] In one exemplary instance, the information delivery channels of this invention include more than one, and the advertising delivery channels include, but are not limited to, soft advertising and hard advertising; in one exemplary instance, hard advertising in this invention includes TV commercials, pre-roll ads, video interstitials, etc.; in one exemplary instance, soft advertising in this invention includes social media such as Douyin and Xiaohongshu, where key opinion leaders (KOLs) such as anchors conduct live-streaming sales, etc.

[0039] Step 104: Based on the established model for attribution analysis, determine the number of unique visitors for each information delivery channel;

[0040] Based on the first relevant variable data, the first interaction conversion data, and the actual value of the first number of unique visitors, this embodiment of the invention constructs an attribution analysis model applicable to both soft and hard advertising through reverse verification. Based on the constructed attribution analysis model, the analysis of the number of unique visitors for each information delivery channel is realized.

[0041] In one exemplary instance, embodiments of the present invention determine the number of unique visitors for each information delivery channel, including:

[0042] The second relevant variable data and the second interactive conversion data collected according to the preset period are input into the determined model for attribution analysis. The model calculates the weight of the predicted value of the second independent visitor count of each information delivery channel to the sum of the predicted values ​​of the second independent visitor count of all information delivery channels.

[0043] Based on the calculated weights and the actual number of second independent visitors collected within the preset period, the actual number of third independent visitors for each information delivery channel is calculated.

[0044] In this embodiment of the invention, the second relevant variable data, the second interaction conversion data, and the actual value of the second number of independent visitors refer to the data collected according to a preset period for attribution analysis.

[0045] In one exemplary instance, the generator in this embodiment of the invention includes a first preset number of first contact distribution packaging layers, a linear connection layer, and a second preset number of second contact distribution packaging layers;

[0046] The final layer, the first touchpoint distribution packaging layer, outputs the first independent visitor count growth value, which is then input to the linear connection layer and the first layer, the second touchpoint distribution packaging layer. The linear connection layer obtains the first independent visitor count prediction value based on the first independent visitor count growth value. The first touchpoint distribution packaging layer contains a preset number of fully connected layers, which perform a fully connected operation on the input data. The second touchpoint distribution packaging layer contains a preset number of fully connected layers, which are used to concatenate the first independent visitor count growth value and the actual first independent visitor count value to obtain reconstructed data.

[0047] In one exemplary instance, the splicing process of this embodiment of the invention includes: serially concatenating the first independent visitor count growth value and the first independent visitor count actual value.

[0048] In one exemplary instance, the generator of this embodiment of the invention includes two parts: an encoder and a decoder; wherein, a first contact distribution packaging layer and a linear connection layer located after the first contact distribution packaging layer constitute the encoder; and a second preset number of second contact distribution packaging layers constitute the decoder.

[0049] In this embodiment of the invention, the first touchpoint distribution (TPDistributed) wrapping layer can apply the fully connected layer to the feature vectors of n advertising channels respectively. Each feature vector composed of m-dimensional features is fully connected to map the m-dimensional features to m′-dimensional features. That is, the neurons of the first TPDistributed wrapping layer of each layer are fully connected only on their respective channels, and the neurons of each channel are not fully connected to each other. Finally, the output of the first TPDistributed wrapping layer of each layer is n×m′.

[0050] In one exemplary instance, before inputting the pre-built training dataset into a preset generator, the method of this embodiment of the invention further includes:

[0051] A training dataset was constructed based on the collected advertising data;

[0052] The advertising data includes: the first relevant variable data for ad placement, the first interaction conversion data for ad placement, and the actual value of the first unique visitor count.

[0053] In one exemplary instance, step 103 of this embodiment of the invention determines a model for attribution analysis, consisting of a generator and a discriminator, using a reverse verification method, including:

[0054] Calculate the loss function based on the training dataset, data sample set, predicted number of first independent visitors, and judgment information;

[0055] When the calculated loss function satisfies the preset convergence condition, the generator and discriminator are combined into a model for attribution analysis.

[0056] If the calculated loss function does not meet the preset convergence condition, the parameters of the generator and discriminator are adjusted. Based on the adjusted generator and discriminator, the data sample set, the predicted value of the first independent visitor count, and the judgment information are updated. Based on the training dataset, the updated data sample set, the predicted value of the first independent visitor count, and the judgment information, the loss function is recalculated. If the recalculated loss function meets the preset convergence condition, the generator and discriminator with adjusted parameters are combined into a model for attribution analysis. If the recalculated loss function does not meet the preset convergence condition, the parameters of the generator and discriminator are adjusted again, and the loss function is recalculated until the recalculated loss function meets the convergence condition. Then, the generator and discriminator that meet the convergence condition are combined into a model for attribution analysis.

[0057] In one exemplary instance, the present invention performs the above-described iterative training on the generator and discriminator according to the loss function; in another exemplary instance, the present invention performs iterative training on the generator and discriminator, including: performing iterative training by minimizing the loss function through Adam optimization.

[0058] In this embodiment of the invention, the input data and reconstructed data are judged by a loss function, and the accuracy of the impact of the information mined by the generator on the number of unique visitors on different channels is verified by reverse verification. Based on the accuracy obtained by verification, the attribution analysis model is constructed.

[0059] In one exemplary instance, the loss function in this embodiment of the invention is L, which is calculated according to the following formula:

[0060] L = L SUV +L G +L D ;

[0061] In the formula, L SUV =MSE(y real ,y pred ), y pred y represents the predicted number of first unique visitors. real L represents the actual number of first independent visitors; MSE represents the mean squared error of the calculation. G =MSE(X true ,X recst )+KL(X true ,X recst ), X recst Indicates reconstructed data, X true Indicates the input data, KL(X) true ,X recst ) represents the KL divergence between the reconstructed data and the input data; X represents the binary cross-entropy between the judgment information of whether a sample in the predicted data sample set is a real sample and the real label of the input sample; i Let z represent the predicted label of the i-th sample. i Let L represent the true label of the i-th sample. D This is the binary cross-entropy between the label predicted by the discriminator as whether the input sample is a real sample and the actual label of the input sample. Real samples are the input data contained in the training dataset; when the input sample is not a real sample, it means that the samples in the constructed data sample set are reconstructed data.

[0062] In one exemplary instance, the discriminator of this invention may refer to related technologies and consist of a convolutional layer, a fully connected layer, and a logistic regression layer.

[0063] In one exemplary instance, the method of this embodiment of the invention further includes: determining the hyperparameters of the model using a grid search method.

[0064] It should be noted that the determination of model hyperparameters using the grid search method can be based on relevant principles and will not be elaborated here.

[0065] In one exemplary instance, before inputting a pre-constructed training dataset containing input data and actual values ​​of the first number of independent visitors into a preset generator, the method of this embodiment further includes obtaining the input data through the following processing:

[0066] Identify categorical and continuous data in the first relevant variable data and the first interactive conversion data;

[0067] Numerical encoding is performed on the identified categorical data;

[0068] Perform feature standardization on continuous data and numerically coded categorical data;

[0069] Input data is obtained by concatenating continuous data that has undergone feature standardization and categorical data that has been numerically encoded.

[0070] In one exemplary instance, the splicing method of this invention includes: serially connecting continuous data that has undergone feature standardization processing and categorical data that has been numerically encoded.

[0071] In one exemplary instance, the numerical encoding in this embodiment of the invention may include: sequence number encoding or one-hot encoding.

[0072] The embodiments of the present invention perform feature standardization on the data, providing data support for the accuracy of model construction; in an exemplary instance, the embodiments of the present invention can use the standard deviation standardization method to process the data, and the data after feature standardization follows a standard normal distribution with a mean of 0 and a variance of 1.

[0073] This invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the above-described method for determining visitor search volume.

[0074] This invention also provides a terminal, comprising: a memory and a processor, wherein the memory stores a computer program; wherein,

[0075] The processor is configured to execute computer programs in memory;

[0076] The computer program implements the above method for determining visitor search volume when executed by a processor.

[0077] Figure 3 is a structural block diagram of the device for determining visitor search volume according to an embodiment of the present invention. As shown in Figure 3, it includes: a first unit, a second unit, a determining unit, and a processing unit; wherein,

[0078] The first unit is set to: input a pre-constructed training dataset containing input data and the actual value of the first independent visitor count into a preset generator to obtain the reconstructed data of the input data and the predicted value of the first independent visitor count;

[0079] The second unit is set up as follows: each input data and each reconstructed data is treated as a sample to construct a sample dataset, and the constructed data sample set is input into a preset discriminator to obtain judgment information on whether the samples in the prediction data sample set are real samples.

[0080] The determination unit is set as follows: based on the training dataset, the constructed data sample set, the first independent visitor number prediction value and judgment information, the model for attribution analysis, consisting of a generator and a discriminator, is determined by reverse verification method.

[0081] The processing unit is set to determine the number of unique visitors for each information delivery channel based on the established model used for attribution analysis.

[0082] The input data includes the first relevant variable data for information delivery and the first interactive conversion data for information delivery.

[0083] Based on the first relevant variable data, the first interaction conversion data, and the actual value of the first number of unique visitors, this embodiment of the invention constructs an attribution analysis model applicable to both soft and hard advertising through reverse verification. Based on the constructed attribution analysis model, the analysis of the number of unique visitors for each information delivery channel is realized.

[0084] In one exemplary embodiment, the processing unit of this invention includes: an input module and a calculation module; wherein,

[0085] The input module is set to input the second relevant variable data and the second interactive conversion data collected according to a preset period into a determined model for attribution analysis, and calculate the weight of the predicted value of the second independent visitor count of each information delivery channel to the sum of the predicted values ​​of the second independent visitor count of all information delivery channels through the model.

[0086] The calculation module is set to calculate the actual number of third independent visitors for each information delivery channel based on the calculated weight and the actual number of second independent visitors collected within a preset period.

[0087] In one exemplary instance, the generator in this embodiment of the invention includes a first preset number of first contact distribution packaging layers, a linear connection layer, and a second preset number of second contact distribution packaging layers;

[0088] The final layer, the first touchpoint distribution packaging layer, outputs the first independent visitor count growth value, which is then input to the linear connection layer and the first layer, the second touchpoint distribution packaging layer. The linear connection layer obtains the first independent visitor count prediction value based on the first independent visitor count growth value. The first touchpoint distribution packaging layer contains a preset number of fully connected layers, which perform a fully connected operation on the input data. The second touchpoint distribution packaging layer contains a preset number of fully connected layers, which are used to concatenate the first independent visitor count growth value and the actual first independent visitor count value to obtain reconstructed data.

[0089] In one exemplary instance, the determining unit of this embodiment of the invention is configured as follows:

[0090] Calculate the loss function based on the training dataset, data sample set, predicted number of first independent visitors, and judgment information;

[0091] When the calculated loss function satisfies the preset convergence condition, the generator and discriminator are combined into a model for attribution analysis.

[0092] If the calculated loss function does not meet the preset convergence condition, the parameters of the generator and discriminator are adjusted. Based on the adjusted generator and discriminator, the data sample set, the predicted value of the first independent visitor count, and the judgment information are updated. Based on the training dataset, the updated data sample set, the predicted value of the first independent visitor count, and the judgment information, the loss function is recalculated. If the recalculated loss function meets the preset convergence condition, the generator and discriminator with adjusted parameters are combined into a model for attribution analysis. If the recalculated loss function does not meet the preset convergence condition, the parameters of the generator and discriminator are adjusted again, and the loss function is recalculated until the recalculated loss function meets the convergence condition. Then, the generator and discriminator that meet the convergence condition are combined into a model for attribution analysis.

[0093] In one exemplary instance, the loss function in this embodiment of the invention is L, which is calculated according to the following formula:

[0094] L = L SUV +L G +L D ;

[0095] In the formula, L SUV =MSE(y real ,y pred ), y pred y represents the predicted number of first unique visitors. real L represents the actual number of first independent visitors; MSE represents the mean squared error of the calculation. G =MSE(X true ,X recst )+KL(X true ,X recst ), X recst Indicates reconstructed data, X true Indicates the input data, KL(X) true ,X recst ) represents the KL divergence between the reconstructed data and the input data; X represents the binary cross-entropy between the judgment information of whether a sample in the predicted data sample set is a real sample and the real label of the input sample; i Let z represent the predicted label of the i-th sample. i This represents the true label of the i-th sample.

[0096] In one exemplary instance, the device determining unit of this embodiment is further configured to: determine the hyperparameters of the model by a grid search method.

[0097] In one exemplary embodiment, the apparatus of the present invention further includes a preprocessing unit, configured to:

[0098] Identify categorical and continuous data in the first relevant variable data and the first interactive conversion data;

[0099] Numerical encoding is performed on the identified categorical data;

[0100] Perform feature standardization on continuous data and numerically coded categorical data;

[0101] Input data is obtained by concatenating continuous data that has undergone feature standardization and categorical data that has been numerically encoded.

[0102] The following application examples briefly illustrate the embodiments of the present invention. These application examples are only used to describe the embodiments of the present invention and are not intended to limit the scope of protection of the present invention.

[0103] Application Examples

[0104] The following uses information delivery as an example of advertising delivery to briefly describe the embodiments of the present invention. Figure 3 is a flowchart of the method of the application example of the present invention, as shown in Figure 3, including:

[0105] Step 301: Collect advertising data for ads placed on various advertising channels; wherein, the advertising data includes: the first relevant variable data, the first interaction conversion data, and the actual value of the first unique visitor count for each advertising channel;

[0106] The first relevant variable data includes one or any combination of the following: advertising type (soft or hard advertising), advertising platform, advertising influence (number of followers, number of likes, etc.), advertising duration, and advertising format (interstitial or embedded ads).

[0107] The first interactive conversion data includes one or any combination of the following: number of likes, number of shares, number of comments, number of impressions, number of clicks, etc.

[0108] In this application example, the number of unique visitors refers to the actual number of visitors who entered the store within a given period.

[0109] In one exemplary instance, the advertising data in this application example can be collected from an advertising monitoring system in the related technology, such as the Admonitor lite system;

[0110] Step 302: Preprocess the collected advertising data.

[0111] The advertising data collected in this application example mainly includes categorical and continuous data. Categorical data includes advertising placement type and placement format. For categorical data, this application example uses relevant technologies to perform numerical encoding, converting categories into numerical quantities; the converted numerical quantities are then subjected to feature standardization processing. The numerical encoding in this application example includes: ordinal encoding or one-hot encoding. This application example ensures the accuracy of subsequent models through standardization processing. The standardization processing in this application example includes standard deviation standardization, which makes the obtained data follow a standard normal distribution with a mean of 0 and a variance of 1.

[0112] Step 303: Construct a training dataset based on the preprocessed advertising data;

[0113] In one exemplary instance, this application example constructs a training dataset by: concatenating first relevant variable data with first interaction conversion data to obtain input data, determining the actual values ​​of independent visitors as label data, and using the input data and label data as data in the training dataset;

[0114] In one exemplary instance, the input data in this application example can be a 2D matrix obtained by concatenating the first relevant variable data and the first interactive transformation data.

[0115] Step 304: Construct a model for attribution analysis based on reverse validation; the model used for attribution analysis in this application example refers to the attribution analysis model for the number of unique visitors across all channels.

[0116] Figure 4 is a system diagram illustrating the determination of visitor search volume in this application example. As shown in Figure 4, advertising data is processed through numerical quantification and feature standardization to obtain input data and tag data; the input data and tag data are then fed into the generator and discriminator of this application example; where...

[0117] The generator consists of an encoder and a decoder, used to generate reconstructed data. The encoder and decoder each contain a corresponding touch point distribution (TPDistributed, Touch Point Distributed) wrapper layer.

[0118] In one exemplary instance, the input to the encoder in this application example is the input data of the training dataset. The encoder's input size is n×m, where n is the number of channels and m is the number of features. The encoder contains two or more first TPDistributed wrapper layers. The output of the penultimate first TPDistributed wrapper layer is the first unique visitor growth value corresponding to each advertising channel, with a size of n×1. The last layer is a linear connection layer used to determine the predicted value of the first unique visitor count based on the first unique visitor growth value. The first unique visitor growth value output by the penultimate layer of the encoder is concatenated with the actual value of the first unique visitor count to obtain reconstructed data of size (n+1)×1.

[0119] The discriminator takes as input a set of data samples consisting of input data from the training dataset and reconstructed data generated by the decoder. The output is the label of whether a sample in the data sample set is part of the input data. If a sample in the data sample set is part of the input data, the label of the sample is 1 in the output. If a sample in the data sample set is part of the reconstructed data, the label of the sample is 0 in the output. The discriminator is constructed from multiple convolutional layers, fully connected layers, and logistic regression layers.

[0120] This application example trains an attribution analysis model using one or any combination of the following loss functions L: L = L SUV +L G +L D ;

[0121] In the formula, L SUV =MSE(y real ,y pred ), y pred y represents the predicted number of first unique visitors. real L represents the actual number of first independent visitors; MSE represents the mean squared error of the calculation. G =MSE(X true ,X recst )+KL(X true ,X recst ), X recst Indicates reconstructed data, X true Indicates the input data, KL(X) true ,X recst ) represents the KL divergence between the reconstructed data and the input data; X represents the binary cross-entropy between the judgment information of whether a sample in the predicted data sample set is a real sample and the real label of the input sample; i Let z represent the predicted label of the i-th sample. i This represents the true label of the i-th sample.

[0122] This application example uses a model for attribution analysis. Iterative training to minimize the loss function using Adam optimization is employed to determine the model parameters. Simultaneously, the model's hyperparameters, such as the number of layers and neurons per layer in the generator, and the number of layers and convolutional kernels in the discriminator, can be determined through grid search training. It's important to note that grid search is an exhaustive search method for specified parameter values. It optimizes the learning algorithm by cross-validating the parameters of the estimated function. Specifically, it involves permuting and combining all possible parameter values ​​to generate a "grid." These combinations are then used for SVM training, and cross-validation is used to evaluate performance. After trying all parameter combinations for the fitted function, a suitable classifier is returned, automatically adjusting to the optimal parameter combination. This application example, after determining the parameters and hyperparameters, yields a model for attribution analysis of independent visitor counts.

[0123] Step 305: Based on the obtained model for attribution analysis, perform attribution analysis on the number of unique visitors for each advertising channel;

[0124] After obtaining the model for attribution analysis, this application example performs attribution analysis on unique visitors across all channels through the following processes: calculating the weight of unique visitors for each advertising channel and recalculating the unique visitor count for each channel; wherein,

[0125] The weight of unique visitors for each advertising channel: The second relevant variable data and the second interaction conversion data collected according to the preset period are input into the pre-built model for attribution analysis. The model calculates the weight of the predicted value of the second unique visitors for each advertising channel relative to the sum of the predicted values ​​of the second unique visitors for all advertising channels. In other words, this application example calculates the weight of unique visitors for each advertising channel through the encoder of the attribution analysis model of unique visitors for all channels based on reverse validation. The second relevant variable data of advertising within the period and the corresponding second interaction conversion data are input into the encoder. The weight of unique visitors for each advertising channel is calculated based on the ratio of the output of the second-to-last layer of the encoder.

[0126] The calculation of the number of unique visitors for each advertising channel is as follows: Based on the determined weights and the actual value of the second number of unique visitors collected within the preset period, the actual value of the number of unique visitors for each advertising channel is calculated. In other words, this application example calculates the actual value of the number of unique visitors for each advertising channel by multiplying the number of unique visitors collected within the period by the weight of the number of unique visitors for each channel and the actual value of the second number of unique visitors collected within the period.

[0127] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

Claims

1. A method for determining visitor search volume, comprising: A pre-constructed training dataset containing input data and the actual value of the first independent visitor count is input into a preset generator to obtain reconstructed data of the input data and a predicted value of the first independent visitor count. The generator includes a first preset number of first touchpoint distribution wrapping layers, a linear connection layer, and a second preset number of second touchpoint distribution wrapping layers. The last first touchpoint distribution wrapping layer outputs the first independent visitor count growth value, which is input to the linear connection layer and the first second touchpoint distribution wrapping layer. The linear connection layer obtains the predicted value of the first independent visitor count based on the growth value. The first touchpoint distribution wrapping layer includes a preset number of fully connected layers, which perform a fully connected operation on the input data. The second touchpoint distribution wrapping layer includes... A predetermined number of fully connected layers are used to concatenate the growth value of the first independent visitor count and the actual value of the first independent visitor count to obtain the reconstructed data. Each input data and each reconstructed data are treated as a sample to construct a data sample set, and the constructed data sample set is input into a predetermined discriminator to obtain judgment information on whether the samples in the predicted data sample set are real samples. Based on the training dataset, the constructed data sample set, the predicted value of the first independent visitor count, and the judgment information, a model for attribution analysis consisting of a generator and a discriminator is determined by reverse verification. Based on the determined model for attribution analysis, the independent visitor count information of each information delivery channel is determined. The input data includes: the first relevant variable data of information delivery and the first interaction conversion data of information delivery.

2. The method according to claim 1, characterized in that, The determination of the number of unique visitors for each information delivery channel includes: inputting the second relevant variable data and the second interaction conversion data collected according to a preset period into a determined model for attribution analysis; calculating the weight of the predicted value of the number of unique visitors for each information delivery channel relative to the sum of the predicted values ​​of the number of unique visitors for all information delivery channels; and calculating the actual value of the number of unique visitors for each information delivery channel based on the calculated weight and the actual value of the number of unique visitors collected within the preset period.

3. The method according to claim 1, characterized in that, The method of determining the model for attribution analysis, consisting of a generator and a discriminator, using reverse validation includes: calculating a loss function based on the training dataset, the data sample set, the first predicted number of independent visitors, and the judgment information; when the calculated loss function satisfies a preset convergence condition, combining the generator and the discriminator into the model for attribution analysis; when the calculated loss function does not satisfy the preset convergence condition, adjusting the parameters of the generator and the discriminator; and updating the data sample set, the first predicted number of independent visitors, and the judgment information based on the parameter-adjusted generator and discriminator. Based on the training dataset, the updated data sample set, the first predicted number of independent visitors, and the judgment information, the loss function is recalculated. When the recalculated loss function satisfies a preset convergence condition, the generator and discriminator with adjusted parameters are combined to form the model for attribution analysis. When the recalculated loss function does not satisfy the preset convergence condition, the parameters of the generator and discriminator are adjusted again, and the loss function is recalculated until the recalculated loss function satisfies the convergence condition. Then, the generator and discriminator that satisfy the convergence condition are combined to form the model for attribution analysis.

4. The method according to claim 3, characterized in that, The loss function is L, which is calculated according to the following formula: L = L SUV +L G +L D In the formula, L SUV =MSE(y real ,y pred ), y pred y represents the predicted number of first unique visitors. real L represents the actual number of first independent visitors; MSE represents the mean squared error of the calculation. G =MSE(X true ,X recst )+KL(X true ,X recst ), X recst Indicates data reconstruction, X true Indicates the input data, KL(X) true ,X recst ) represents the KL divergence between the reconstructed data and the input data; X represents the binary cross-entropy between the judgment information of whether a sample in the predicted data sample set is a real sample and the real label of the input sample; i Let z represent the predicted label of the i-th sample. i Let represent the true label of the i-th sample, n represent the number of samples in the training dataset, and 2n represent the number of samples in the data sample set.

5. The method according to any one of claims 1-4, characterized in that, The method further includes determining the hyperparameters of the model using a grid search method.

6. The method according to any one of claims 1-4, characterized in that, Before inputting the pre-constructed training dataset containing input data and the actual value of the first number of independent visitors into the preset generator, the method further includes obtaining the input data through the following processes: determining categorical and continuous data in the first relevant variable data and the first interaction conversion data; numerically encoding the determined categorical data; performing feature standardization processing on the continuous data and the numerically encoded categorical data; and concatenating the continuous data that has undergone feature standardization processing and the numerically encoded categorical data to obtain the input data.

7. A computer storage medium storing a computer program that, when executed by a processor, implements the method for determining visitor search volume as described in any one of claims 1-6.

8. A terminal, comprising: A memory and a processor, wherein the memory stores a computer program; wherein the processor is configured to execute the computer program in the memory; and the computer program, when executed by the processor, implements the method for determining visitor search volume as described in any one of claims 1-6.

9. An apparatus for determining visitor search volume, comprising: The system comprises a first unit, a second unit, a determining unit, and a processing unit; wherein the first unit is configured to: input a pre-constructed training dataset containing input data and the actual value of the first independent visitor count into a preset generator to obtain reconstructed data of the input data and a predicted value of the first independent visitor count; the generator includes a first preset number of first touchpoint distribution wrapping layers, a linear connection layer, and a second preset number of second touchpoint distribution wrapping layers; wherein the last first touchpoint distribution wrapping layer outputs the first independent visitor count growth value, and the first independent visitor count growth value is input to the linear connection layer and the first second touchpoint distribution wrapping layer; the linear connection layer obtains the first independent visitor count prediction value based on the first independent visitor count growth value; the first touchpoint distribution wrapping layer includes a preset number of fully connected layers, and the fully connected layers perform a fully connected operation on the input data; the second touchpoint distribution wrapping layer... The first unit comprises a preset number of fully connected layers for concatenating the growth value of the first independent visitor count and the actual value of the first independent visitor count to obtain the reconstructed data. The second unit is configured to construct a data sample set by treating each input data and each reconstructed data as a sample, and input the constructed data sample set into a preset discriminator to obtain judgment information on whether the samples in the predicted data sample set are real samples. The determining unit is configured to determine the model for attribution analysis composed of the generator and the discriminator by using a reverse verification method based on the training dataset, the constructed data sample set, the predicted value of the first independent visitor count, and the judgment information. The processing unit is configured to determine the independent visitor count information of each information delivery channel based on the determined model for attribution analysis. The input data includes the first relevant variable data of information delivery and the first interaction conversion data of information delivery.

Citation Information

Patent Citations

  • Multi-dimensional advertisement effect evaluation method and system, electronic equipment and storage medium

    CN113947435A

  • Method and device for detecting advertisement traffic, electronic equipment and storage medium

    CN114581148A