An advertisement effect evaluation method and system based on reinforcement learning
By combining multi-source data evaluation methods such as millimeter-wave radar, WiFi probes, and anonymous QR codes, a state space is constructed and deduplication is performed. Reinforcement learning models are used to optimize advertising effectiveness evaluation, thus solving the risk of privacy data evaluation and achieving high-precision advertising effectiveness evaluation.
Patent Information
- Application Number
- CN202511483368.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing technologies that use private data to evaluate advertising effectiveness pose privacy and security risks, necessitating an evaluation method that does not require private data.
By combining millimeter-wave radar, WiFi probes, and anonymous QR codes, the system acquires multi-source data on users watching ads, constructs a state space, calculates state vectors, outputs the optimal action and comprehensive effect score, and uses a reinforcement learning model to optimize ad performance evaluation, while also performing deduplication and privacy protection.
It achieves high-precision advertising effectiveness evaluation, avoids privacy data leakage, and provides accurate advertising effectiveness evaluation results.
Smart Images

Figure CN120996875B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and particularly relates to a method for evaluating the effect of advertisement delivery by using a reinforcement learning method. BACKGROUND
[0002] Advertisement has penetrated into all aspects of people's life, from watches and mobile phones to cars and houses, and advertisement can increase exposure and thus greatly help the sales of products.
[0003] The delivery of advertisement needs to be balanced in many aspects, and the position, time and quantity of delivery need to be analyzed in advance so that each delivered advertisement can obtain a high conversion rate as much as possible. Therefore, after the delivery of advertisement, the effect of advertisement needs to be evaluated according to the feedback of the audience, and thus the delivery strategy in the later stage can be provided. For example, CN117408757A provides an intelligent evaluation system for monitoring the effect of advertisement delivery. In the face of the scene of outdoor advertisement which cannot obtain user feedback, the system collects the facial images of people when they watch the advertisement, and obtains the interest degree of people to the advertisement after analysis.
[0004] However, the facial image belongs to private data, and there is a risk of privacy security when using these data. Therefore, it is necessary to provide a method for determining the effect of advertisement without using private data. SUMMARY
[0005] The embodiments of the present application provide a method and system for evaluating the effect of advertisement based on reinforcement learning, to solve the risk problem caused by using private data of people to evaluate the effect of advertisement in the prior art.
[0006] In one aspect, the embodiments of the present application provide a method for evaluating the effect of advertisement based on reinforcement learning, comprising:
[0007] obtaining multi-source data generated by a user when watching an advertisement, the multi-source data comprising millimeter wave radar data, WiFi probe data and anonymous two-dimensional code data;
[0008] constructing a state space;
[0009] substituting the multi-source data into the state space to obtain a corresponding state vector;
[0010] outputting an optimal action by an intelligent agent under the state vector, the optimal action comprising the weight of each data in the multi-source data;
[0011] calculating a comprehensive effect score based on the weight and the single-dimensional effect score of each data in the multi-source data;
[0012] determining a corresponding immediate reward according to the comprehensive effect score;
[0013] The state vector of the current moment, the optimal action, the immediate reward and the state vector of the next moment are stored in the experience replay pool, the network parameters are updated by the gradient descent method, and the reinforcement learning model is optimized;
[0014] The multi-source data to be processed is input into the reinforcement learning model to obtain an advertising effect evaluation result;
[0015] After obtaining the multi-source data, the multi-source data is subjected to a deduplication process, and the deduplication method includes:
[0016] For each time window, a radar user set, a WiFi user set and a two-dimensional code user set are constructed;
[0017] The intersection of the radar user set, the WiFi user set and the two-dimensional code user set is calculated U 重叠1 The intersection of the radar user set and the WiFi user set U 重叠2 The intersection of the radar user set and the two-dimensional code user set U 重叠3 The intersection of the WiFi user set and the two-dimensional code user set U 重叠4 :
[0018]
[0019]
[0020]
[0021]
[0022] The final intersection after adjustment U 重叠 is expressed as:
[0023]
[0024] wherein, U R the radar user set is U W the WiFi user set is U Q the two-dimensional code user set is
[0025] The union of the radar user set, the WiFi user set and the two-dimensional code user set is subtracted from the intersection to obtain an effective user set U 有效 :
[0026] .
[0027] In another aspect, the embodiment of the present application also provides an advertisement effect evaluation system based on reinforcement learning, comprising:
[0028] a data acquisition module, configured to acquire multi-source data generated by a user when watching an advertisement, the multi-source data comprising millimeter wave radar data, WiFi probe data and anonymous two-dimensional code data;
[0029] a state space construction module, configured to construct a state space;
[0030] a state vector calculation module, configured to substitute the multi-source data into the state space to obtain a corresponding state vector;
[0031] an action output module, configured to output an optimal action by an agent under the state vector, the optimal action comprising a weight of each kind of data in the multi-source data;
[0032] a score calculation module, configured to calculate a comprehensive effect score based on the weight and a single-dimension effect score of each kind of data in the multi-source data;
[0033] a reward calculation module, configured to determine a corresponding immediate reward according to the comprehensive effect score;
[0034] a model optimization module, configured to store the state vector at a current moment, the optimal action, the immediate reward and the state vector at a next moment into an experience replay pool, update network parameters through a gradient descent method and optimize the reinforcement learning model;
[0035] an advertisement evaluation module, configured to input the multi-source data to be processed into the reinforcement learning model to obtain an advertisement effect evaluation result.
[0036] In another aspect, the embodiment of the present application also provides a computer storage medium, which stores a plurality of computer instructions, and the computer instructions are used to make a computer execute the method described above.
[0037] The advertisement effect evaluation method and system based on reinforcement learning in the present application have the following advantages:
[0038] The millimeter wave radar+WiFi probe+anonymous two-dimensional code combination method not only does not have privacy risks, but also has high detection accuracy, and thus more accurate advertisement effect evaluation results are obtained. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description only constitute some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0040] Figure 1 A flowchart of an advertisement effect evaluation method based on reinforcement learning provided by an embodiment of the present application. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0042] Figure 1 A flowchart of an advertisement effect evaluation method based on reinforcement learning provided by an embodiment of the present application. The present application provides an advertisement effect evaluation method based on reinforcement learning, which comprises:
[0043] S100, acquiring multi-source data generated by a user when watching an advertisement, wherein the multi-source data comprises millimeter wave radar data, WiFi probe data and anonymous two-dimensional code data.
[0044] Exemplarily, the collection time window of each kind of multi-source data is 1 minute / time, so in the whole time T, a plurality of time windows can be divided, denoted as t=1, 2, …, T.
[0045] Further, the millimeter wave radar data is acquired R t At this time, the distance between the i-th user and the advertisement screen in the t-th time window is collected by the millimeter wave radar t i d t,i i =1, 2, …, N R When the distance is less than a set distance threshold and the user stays for more than a time threshold, for example, 3 seconds, the user is recorded as an effective user, and the staying time of the effective user r t,i and the corresponding distance d t,i are recorded as the millimeter wave radar data, so the millimeter wave radar data R t can be represented as:
[0046]
[0047] in, N R The number of effective users detected by millimeter-wave radar.
[0048] Obtain WiFi probe data W t At that time, the first t Within minutes, the user carried the first j The MAC (Media Access Control) address of an electronic device is hashed to de-identify it. Specifically, salted hashing can be used, where the salt value is random and not stored, to obtain an anonymous device identifier. m t,j , j =1,2,…, N W It also obtains the time that the user's electronic device stays within 50 meters of the advertising screen, and identifies anonymous devices. m t,j and equipment dwell time s t,j As WiFi probe data W t Therefore, WiFi probe data W t It can be represented as:
[0049]
[0050] in, N W This represents the number of valid users collected via WiFi probes.
[0051] Obtain anonymous QR code data Q t At that time, obtain the feedback value of the k-th user within the t-th minute after scanning the anonymous QR code displayed on the advertising screen using an electronic device. f t,k It also obtains the user's scanning time. t t,k , k =1,2,…, N Q , feedback value f t,k and scanning time t t,k As anonymous QR code data Q t Therefore, anonymous QR code data Q t It can be represented as:
[0052]
[0053] wherein, N Q the number of effective users scanning the anonymous two-dimensional code.
[0054] Further, the anonymous two-dimensional code in the embodiments of the present application is updated every 10 minutes, and the user can input 1 and 0 to represent the feedback value after scanning the code, wherein 1 represents satisfaction and 0 represents dissatisfaction.
[0055] Further, since the same user can be detected by multiple devices at the same time, for example, captured by the millimeter wave radar and detected by the WiFi probe, and the code scanning feedback, it is necessary to identify the overlapping users through the space-time matching method to avoid repeated calculation of data. After obtaining the multi-source data, the multi-source data is de-duplicated, and if there is an intersection among the stay time period of the user in the millimeter wave radar data, the stay time period of the electronic device, and the code scanning time, i.e., overlap ≥ 2 seconds, it can be determined as a potential same user.
[0056] Specifically, the method of de-duplication processing includes:
[0057] 1. For each time window, a radar user set, a WiFi user set and a two-dimensional code user set are constructed.
[0058] Radar user set U R is represented as:
[0059]
[0060] wherein, is the number of the N R th user identified by the millimeter wave radar.
[0061] WiFi user set U W is represented as:
[0062]
[0063] wherein, is the number of the N W th device identified by the WiFi probe.
[0064] Two-dimensional code user set U Q is represented as:
[0065]
[0066] wherein, is the number of theN Q The ID of the user who scanned the anonymous QR code.
[0067] 2. Calculate the intersections of the radar user set, the WiFi user set, and the QR code user set, as well as the intersections between any two of them. U 重叠 .
[0068] The intersection of radar user set, WiFi user set, and QR code user set U 重叠1 Represented as:
[0069]
[0070] The intersection of the radar user set and the WiFi user set U 重叠2 for:
[0071]
[0072] Similarly, the intersection of the radar user set and the QR code user set can also be obtained. U 重叠3 :
[0073]
[0074] The intersection of the WiFi user set and the QR code user set U 重叠4 :
[0075]
[0076] Therefore, the final intersection after adjustment U 重叠 Represented as:
[0077] .
[0078] Intersection U 重叠 Subsequently, based on the dwell time period in the millimeter-wave radar data, the dwell time period of the electronic device, and the scanning time, the same user can be identified by merging their numbers together to mark them as the same user. However, the numbers in each user set are still retained after merging to facilitate subsequent processing.
[0079] 3. Subtract the intersection from the union of the radar user set, the WiFi user set, and the QR code user set to obtain the effective user set. U 有效 :
[0080] .
[0081] After obtaining the valid user set, combine it with the QR code user set. U Q The intersection of these values serves as the identifier of the valid user who scanned the anonymous QR code after deduplication. U Q有效 Combine the effective user set with the WiFi user set U W The intersection of these values serves as the identifier of the valid user detected by the WiFi probe after deduplication. U W有效 The effective user set and the radar user set U R The intersection of these values serves as the identifier of the valid user detected by the millimeter-wave radar after deduplication. U R有效 .
[0082] S110, construct the state space.
[0083] For example, the state space includes time period type, real-time pedestrian density, average dwell time, historical performance rating and weather type, and the average dwell time is the average of the dwell time of effective users and the dwell time of devices in the deduplicated multi-source data.
[0084] S120, substitute the multi-source data into the state space to obtain the corresponding state vector.
[0085] For example, the state vector is formed by substituting the data from multiple sources into the corresponding time period types. s 1. Real-time crowd density s 2. Average length of stay s 3. Historical performance rating s 4 and weather type s The vector obtained after step 5.
[0086] Specifically, time period type s The value 1 has three meanings: 0 indicates off-peak weekdays, 1 indicates peak weekdays, and 2 indicates weekends. Real-time pedestrian density. s 2 is 100m 2 The number of people inside can be statistically analyzed using millimeter-wave radar and WiFi probes. Historical performance rating. s 4. You can take the first 5 time windows. S 总 The mean. Weather type. s The value of 5 also has three possibilities: 0 represents sunny, 1 represents cloudy, and 2 represents rain.
[0087] Therefore, the state vector S It can be represented as:
[0088] .
[0089] S130 is the optimal action output by the agent in the reinforcement learning framework under the state vector. The optimal action contains the weight of each data in the multi-source data.
[0090] For example, the weight of each data point in the multi-source data is obtained based on the sample size, the number of valid users included in each data point, and the total number of users.
[0091] Specifically, the weights include radar weights. w R WiFi weight w W and QR code weight w Q The calculation formulas are expressed as follows:
[0092]
[0093]
[0094]
[0095] in, α This is the sample size weighting coefficient, with a value between 0.6 and 0.8. N 总有效 = N R有效 + N W有效 + N Q有效 This represents the total number of active users. N R有效 This represents the number of valid users detected by the millimeter-wave radar after deduplication. N W有效 This represents the number of valid users detected by the WiFi probe after deduplication. N Q有效 This represents the number of valid users who scanned the anonymous QR code after deduplication. d avg The average distance to the user detected by millimeter-wave radar. N 总人流 This represents the total foot traffic around the advertising screen during the current time period.
[0096] After obtaining the weights, the weights are adjusted based on real-time pedestrian density and the usage scenario of the advertising screen. For example, during peak hours, such as morning and evening rush hours, WiFi probes have wider coverage, so the WiFi weight is increased. During low-traffic periods, proactive QR code scanning feedback is more reliable, so the QR code weight is increased. In close-range scenarios, such as shopping mall screens, millimeter-wave radar data is more accurate, so the radar weight is increased.
[0097] Furthermore, the optimal action also includes selecting the best display version from multiple different ad versions based on real-time pedestrian density. a content .
[0098] Specifically, different ad versions include short, long, and interactive versions. You can choose the appropriate version based on the real-time crowd density in the current scene. For example, choose the 15-second short version when the crowd is dense, choose the 30-second long version when the crowd is sparse, and choose the interactive version when the crowd is in the middle.
[0099] Therefore, the optimal action A It can be represented as:
[0100] .
[0101] S140 calculates the overall effect score based on the weights and the single-dimensional effect score of each data point in the multi-source data.
[0102] For example, the single-dimensional performance evaluation includes radar performance evaluation, WiFi performance evaluation, and QR code performance evaluation. The radar performance evaluation is calculated using the number of effective users, the dwell time of effective users, and the distance contained in the millimeter-wave radar data. The WiFi performance evaluation is calculated using the number of effective users and the dwell time of devices contained in the WiFi probe data. The QR code performance evaluation is calculated using the number of effective users and the feedback value contained in the anonymous QR code data.
[0103] Specifically, radar performance rating S R Based on a weighted average of dwell time and distance, the following was obtained:
[0104]
[0105] in, N R有效 This represents the number of valid users detected by the millimeter-wave radar after deduplication. U R有效 This refers to the ID of a valid user detected by the millimeter-wave radar after deduplication.
[0106] WiFi performance rating S W Represented as:
[0107]
[0108] in, N W有效 This represents the number of valid users detected by the WiFi probe after deduplication. U W有效This refers to the ID of the valid user detected by the WiFi probe after deduplication.
[0109] QR code performance rating S Q Represented as:
[0110]
[0111] in, N Q有效 This represents the number of valid users who scanned the anonymous QR code after deduplication. U Q有效 This is the ID of a valid user who scanned the anonymous QR code after deduplication.
[0112] Overall performance score S 总 Represented as:
[0113] .
[0114] S150: The corresponding instant reward is determined based on the overall performance score.
[0115] Exemplary, in this application embodiment, instant rewards R It includes two parts: performance rating and privacy and security, as detailed below:
[0116]
[0117] in, Lambda This is the weighting coefficient, with a value between 0.7 and 0.8. C 隐私 This is the privacy compliance coefficient, which is fixed at 1. If excessive data collection is detected, such as excessively high WiFi probe scanning frequency, it will be reduced to 0.5 to enforce privacy protection.
[0118] S160, the current state vector S Optimal action A Instant rewards R and the state vector at the next time step S 'Store the experience replay pool in the reinforcement learning framework, update the network parameters through gradient descent, and optimize the reinforcement learning model.'
[0119] For example, the reinforcement learning model used in this application embodiment is DQN (Deep Q-Network), and the agent learns the optimal policy through DQN. The goal is to maximize cumulative rewards:
[0120]
[0121] in, Gt The cumulative reward starting from time t. Gamma This is a discount factor, with a value between 0.9 and 0.95. R t+k’+1 The instant reward starting from time t+1. In the state s Choose the action a The corresponding strategy.
[0122] S170: Input the multi-source data to be processed into the reinforcement learning model to obtain the advertising effectiveness evaluation results.
[0123] For example, steps S100-S160 above are all training processes, while step S170 is the actual usage process. In the actual usage process, after obtaining the multi-source data to be processed, it is necessary to first establish the corresponding state vector, then the agent outputs the optimal action, and finally obtains the comprehensive effect score. This process is similar to steps S120-S140 above, and will not be described again here.
[0124] This application also provides an advertising effectiveness evaluation system based on reinforcement learning, the system comprising:
[0125] The data acquisition module is used to acquire multi-source data generated by users when they watch advertisements. The multi-source data includes millimeter-wave radar data, WiFi probe data, and anonymous QR code data.
[0126] The state space construction module is used to construct the state space, which includes time period type, real-time pedestrian density, average stay duration, historical performance score, and weather type.
[0127] The state vector calculation module is used to substitute the multi-source data into the time period type in the state space. s 1. Real-time crowd density s 2. Average length of stay s 3. Historical performance rating s 4 and weather type s 5. Obtain the corresponding state vector;
[0128] The action output module is used by the agent to output the optimal action under the state vector. The optimal action includes the weight of each data in the multi-source data.
[0129] The scoring calculation module is used to calculate a comprehensive performance score based on the weights and single-dimensional performance scores of each data point in the multi-source data. The single-dimensional performance scores include radar performance scores, WiFi performance scores, and QR code performance scores. The comprehensive performance score... S 总 Represented as:
[0130]
[0131] in, w R For radar weights, w W For WiFi weight, w Q For QR code weight, S R , S W and S Q These are the radar performance score, the WiFi performance score, and the QR code performance score, respectively.
[0132] The reward calculation module is used to determine the corresponding immediate reward based on the overall performance score; the immediate reward R It includes two parts: performance rating and privacy and security, as shown below:
[0133]
[0134] in, Lambda These are the weighting coefficients. C 隐私 Privacy compliance coefficient;
[0135] The model optimization module stores the current state vector, optimal action, immediate reward, and the next state vector into the experience replay pool. It updates the network parameters using gradient descent to optimize the reinforcement learning model and learns the optimal policy using DQN. The goal is to maximize cumulative rewards:
[0136]
[0137] in, G t From t Accumulated rewards starting from a certain moment Gamma As a discount factor, R t+k’+1 From t Instant rewards starting at +1 moment. In the state s Choose the action a Corresponding strategies;
[0138] The advertising evaluation module is used to input multi-source data to be processed into the reinforcement learning model to obtain advertising effectiveness evaluation results.
[0139] This application also provides a computer storage medium storing a plurality of computer instructions for causing a computer to execute the above-described method.
[0140] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0141] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for evaluating advertising effectiveness based on reinforcement learning, characterized in that, include: Acquire multi-source data generated by users while they are watching advertisements, including millimeter-wave radar data, WiFi probe data, and anonymous QR code data; Construct the state space; Substituting the multi-source data into the state space yields the corresponding state vector; An agent in the reinforcement learning framework outputs an optimal action under the state vector, and the optimal action includes the weight of each data point in the multi-source data; the weight of each data point in the multi-source data is obtained based on the sample size, the number of valid users in each data point, and the total number of users; Based on the weights and the single-dimensional effect scores of each data point in the multi-source data, a comprehensive effect score is calculated. The corresponding immediate reward is determined based on the comprehensive performance score. The current state vector, the optimal action, the immediate reward, and the state vector at the next moment are stored in the experience replay pool of the reinforcement learning framework. The network parameters are updated by gradient descent to optimize the reinforcement learning model based on the deep Q-network. The multi-source data to be processed is input into the reinforcement learning model to obtain the advertising effect evaluation results; Specifically, when acquiring the millimeter-wave radar data, the distance between the user and the advertising screen is collected by the millimeter-wave radar. When the distance is less than a set distance threshold and the user's dwell time exceeds a time threshold, the user is recorded as a valid user, and the dwell time of the valid user and the corresponding distance are used as the millimeter-wave radar data. When acquiring the WiFi probe data, the MAC address of the user's electronic device is hashed and anonymized to obtain an anonymous device identifier. At the same time, the time the user's electronic device stays within 50 meters of the advertising screen is also acquired. The anonymous device identifier and the device stay time are used as the WiFi probe data. When obtaining the anonymous QR code data, the feedback value after the user scans the QR code displayed on the advertising screen through an electronic device is obtained, and the user's scanning time is also obtained. The feedback value and the scanning time are used as the anonymous QR code data. The QR code displayed on the advertising screen is updated every 10 minutes. After acquiring the multi-source data, the multi-source data is deduplicated. The deduplication method includes: For each time window, construct a radar user set, a WiFi user set, and a QR code user set; Calculate the intersection of the radar user set, the WiFi user set, and the QR code user set. U 重叠1 The intersection between the radar user set and the WiFi user set U 重叠2 The intersection of the radar user set and the QR code user set U 重叠3 and the intersection of the WiFi user set and the QR code user set. U 重叠4 : The final intersection after adjustments U 重叠 Represented as: in, U R For the radar user set, U W For the set of WiFi users, U Q The set of users associated with the QR code; Subtracting the intersection from the union of the radar user set, the WiFi user set, and the QR code user set yields the effective user set. U 有效 : 。 2. The advertising effectiveness evaluation method based on reinforcement learning according to claim 1, characterized in that, The state vector includes time period type, real-time pedestrian density, average dwell time, historical performance rating and weather type. The average dwell time is the average of the dwell time of effective users and the dwell time of devices in the multi-source data.
3. The advertising effectiveness evaluation method based on reinforcement learning according to claim 1, characterized in that, The single-dimensional performance evaluation includes radar performance evaluation, WiFi performance evaluation, and QR code performance evaluation. The radar performance evaluation is calculated based on the number of effective users, the dwell time of effective users, and the distance contained in the millimeter-wave radar data. in, S R Rate the radar performance. N R有效 This represents the number of valid users detected by the millimeter-wave radar after deduplication. U R有效 The ID of the valid user detected by the millimeter-wave radar after deduplication. r t,i For the first t Within minutes i The dwell time of each effective user d t,i For the first t Within minutes i The distance between each user and the advertising screen; The WiFi performance score is calculated based on the number of valid users contained in the WiFi probe data and the device dwell time. in, S W Rate the WiFi performance. N W有效 This represents the number of valid users detected by the WiFi probe after deduplication. U W有效 This refers to the ID of a valid user detected by the WiFi probe after deduplication. s t,j For the first t Within minutes, the user carried the first j The duration of time spent on each electronic device; The QR code effectiveness score is calculated using the number of valid users contained in the anonymous QR code data and the feedback value: in, S Q Rate the effectiveness of the QR code. N Q有效 This represents the number of valid users who scanned the anonymous QR code after deduplication. U Q有效 This refers to the ID of a valid user who scanned the anonymous QR code after deduplication. f t,k For the first t Within minutes k Feedback value from users scanning anonymous QR codes displayed on advertising screens using their electronic devices.
4. The advertising effectiveness evaluation method based on reinforcement learning according to claim 1, characterized in that, After obtaining the weights, the weights are adjusted according to the real-time pedestrian density and the usage scenario of the advertising screen.
5. The advertising effectiveness evaluation method based on reinforcement learning according to claim 1, characterized in that, The optimal action also includes selecting the best display version from multiple different ad versions based on real-time pedestrian density.
6. A system applying the reinforcement learning-based advertising effectiveness evaluation method according to any one of claims 1-5, characterized in that, include: The data acquisition module is used to acquire multi-source data generated by users when watching advertisements. The multi-source data includes millimeter-wave radar data, WiFi probe data, and anonymous QR code data. The state space construction module is used to construct the state space, which includes time period type, real-time pedestrian density, average stay duration, historical performance score, and weather type. The state vector calculation module is used to substitute the multi-source data into the time period type in the state space. s 1. Real-time crowd density s 2. Average length of stay s 3. Historical performance rating s 4 and weather type s 5. Obtain the corresponding state vector; An action output module is used for the agent to output the optimal action under the state vector, wherein the optimal action includes the weight of each data in the multi-source data; The scoring calculation module is used to calculate a comprehensive performance score based on the weights and the single-dimensional performance scores of each data point in the multi-source data; the single-dimensional performance scores include radar performance scores, WiFi performance scores, and QR code performance scores; the comprehensive performance score... S 总 Represented as: in, w R For radar weights, w W For WiFi weight, w Q For QR code weight, S R , S W and S Q These are the radar performance score, the WiFi performance score, and the QR code performance score, respectively. The reward calculation module is used to determine the corresponding instant reward based on the comprehensive performance score; the instant reward R It includes two parts: performance rating and privacy and security, as shown below: in, λ These are the weighting coefficients. C 隐私 Privacy compliance coefficient; The model optimization module stores the current state vector, the optimal action, the immediate reward, and the state vector at the next time step into the experience replay pool, updates the network parameters using gradient descent, optimizes the reinforcement learning model, and learns the optimal policy using DQN. The goal is to maximize cumulative rewards: in, G t From t Accumulated rewards starting from a certain moment γ As a discount factor, R t+k’+1 From t Instant rewards starting at +1 moment. In the state s Choose the action a Corresponding strategies; The advertising evaluation module is used to input the multi-source data to be processed into the reinforcement learning model to obtain the advertising effectiveness evaluation results.
7. A computer storage medium, characterized in that, The computer storage medium stores a plurality of computer instructions, which are used to cause the computer to perform the method described in any one of claims 1-5.
Citation Information
Patent Citations
Large-size screen targeted advertising system and method based on multi-source heterogeneous data analysis
CN108428158A
Advertisement putting method, device and system and storage medium
CN110706030A
Advertisement delivery tracking method and device and terminal equipment
CN110751502A
Method and device for pushing advertising value of advertisement screen, and advertisement screen
CN113674024A
Advertisement marketing system for data diversity identification
CN120198177A