Advertisement effect evaluation method and system based on reinforcement learning
By combining multi-source data processing from millimeter-wave radar, WiFi probes, and anonymous QR codes, a reinforcement learning model was constructed to solve the privacy and security issues in evaluating advertising effectiveness using private data, thus achieving high-precision advertising effectiveness evaluation.
Patent Information
- Application Number
- CN202511483368.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-17
AI Technical Summary
Existing technologies that use private data to evaluate advertising effectiveness pose privacy and security risks, necessitating an evaluation method that does not require private data.
By combining millimeter-wave radar, WiFi probes, and anonymous QR codes, the system acquires multi-source user data, constructs a state space, calculates state vectors, outputs the optimal action, calculates a comprehensive score based on weights and effect scores, and optimizes the results through a reinforcement learning model to obtain advertising effectiveness evaluation results.
It achieves high-precision advertising effectiveness evaluation, avoids privacy risks, and provides accurate evaluation results.
Smart Images

Figure CN120996875A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, in particular to a method for evaluating the effect of advertisement delivery by using reinforcement learning method. BACKGROUND
[0002] Advertisement has penetrated into every aspect of people's life, from watches and mobile phones to cars and houses, all of which can increase exposure through advertisement and thus provide great help to product sales.
[0003] The delivery of advertisement needs to be balanced in many aspects, and the position, time and quantity of delivery need to be analyzed in advance to make each delivered advertisement obtain as high conversion rate as possible. Therefore, after the delivery of advertisement, the effect of advertisement needs to be evaluated according to the feedback of the audience to provide guidance for the later delivery strategy. For example, CN117408757A provides an intelligent evaluation system for monitoring the effect of advertisement delivery. In the face of the scene of outdoor advertisement which cannot obtain user feedback, the system collects facial images of people when they watch the advertisement and obtains the degree of interest of people in the advertisement after analysis.
[0004] However, facial images belong to private data, and there is a risk of privacy security in using these data. Therefore, it is necessary to provide a method for determining the effect of advertisement without using private data. SUMMARY
[0005] The embodiments of the present application provide a method and system for evaluating the effect of advertisement based on reinforcement learning, to solve the problem of risk caused by using private data of people to evaluate the effect of advertisement in the prior art.
[0006] In one aspect, the embodiments of the present application provide a method for evaluating the effect of advertisement based on reinforcement learning, comprising: obtaining multi-source data generated by a user when watching an advertisement, the multi-source data comprising millimeter wave radar data, WiFi probe data and anonymous two-dimensional code data; constructing a state space; substituting the multi-source data into the state space to obtain a corresponding state vector; outputting an optimal action by an agent under the state vector, the optimal action comprising the weight of each data in the multi-source data; calculating a comprehensive effect score based on the weight and the single-dimensional effect score of each data in the multi-source data; determining a corresponding immediate reward according to the comprehensive effect score; storing the state vector at the current time, the optimal action, the immediate reward and the state vector at the next time into an experience replay pool, updating network parameters by gradient descent method and optimizing the reinforcement learning model; Input the multi-source data to be processed into the reinforcement learning model to obtain an advertisement effect evaluation result. Wherein, after obtaining the multi-source data, the multi-source data is subjected to deduplication processing, and the deduplication processing method comprises: For each time window, a radar user set, a WiFi user set and a two-dimensional code user set are constructed; The intersection between the radar user set, the WiFi user set and the two-dimensional code user set is calculated U 重叠1 The intersection between the radar user set and the WiFi user set U 重叠2 The intersection between the radar user set and the two-dimensional code user set U 重叠3 The intersection between the WiFi user set and the two-dimensional code user set U 重叠4 :
[0007]
[0008]
[0009]
[0010] The final intersection after adjustment U 重叠 is represented as:
[0011] Wherein, U R the radar user set is U W the WiFi user set is U Q the two-dimensional code user set is The union set of the radar user set, the WiFi user set and the two-dimensional code user set is subtracted by the intersection set to obtain an effective user set U 有效 : .
[0012] On the other hand, the embodiment of the application further provides an advertisement effect evaluation system based on reinforcement learning, comprising: A data acquisition module is configured to acquire multi-source data generated by a user when watching an advertisement, wherein the multi-source data comprises millimeter wave radar data, WiFi probe data and anonymous two-dimensional code data; A state space construction module is configured to construct a state space; The state vector calculation module is configured to substitute the multi-source data into a state space to obtain a corresponding state vector. The action output module is configured to output an optimal action by the agent under the state vector, and the optimal action contains a weight of each data in the multi-source data. The score calculation module is configured to calculate a comprehensive effect score based on the weight and a single-dimension effect score of each data in the multi-source data. The reward calculation module is configured to determine a corresponding immediate reward according to the comprehensive effect score. The model optimization module is configured to store the state vector at the current moment, the optimal action, the immediate reward, and the state vector at the next moment into an experience replay pool, update network parameters by a gradient descent method, and optimize the reinforcement learning model. The advertisement evaluation module is configured to input the multi-source data to be processed into the reinforcement learning model to obtain an advertisement effect evaluation result.
[0013] In another aspect, the application also provides a computer storage medium, which stores a plurality of computer instructions for causing a computer to execute the above method.
[0014] The advertisement effect evaluation method and system based on reinforcement learning in the application have the following advantages: The combination of millimeter wave radar, WiFi probe, and anonymous two-dimensional code not only does not have privacy risks, but also has high detection accuracy, and thus more accurate advertisement effect evaluation results are obtained. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute some embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings.
[0016] Figure 1 A flowchart of the advertisement effect evaluation method based on reinforcement learning provided by the embodiments of the application. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments only constitute some embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.
[0018] Figure 1 A flowchart illustrating an advertising effectiveness evaluation method based on reinforcement learning, provided as an embodiment of this application. This application provides an advertising effectiveness evaluation method based on reinforcement learning, comprising: S100 acquires multi-source data generated by users when they watch advertisements. The multi-source data includes millimeter-wave radar data, WiFi probe data, and anonymous QR code data.
[0019] For example, the acquisition time window for each type of multi-source data is 1 minute / time. Therefore, the entire duration T can be divided into multiple time windows, denoted as t=1,2,…,T.
[0020] Furthermore, acquire millimeter-wave radar data. R t At that time, the data was collected by millimeter-wave radar. t Within minutes i Distance between individual users and advertising screens d t,i , i =1,2,…, N R When the distance is less than the set distance threshold and the user's dwell time exceeds the time threshold, such as 3 seconds, the user is recorded as a valid user, and the dwell time of the valid user is recorded. r t,i and the corresponding distance d t,i As millimeter-wave radar data, therefore millimeter-wave radar data R t It can be represented as:
[0021] in, N R The number of effective users detected by millimeter-wave radar.
[0022] Obtain WiFi probe data W t At that time, the first t Within minutes, the user carried the first j The MAC (Media Access Control) address of an electronic device is hashed to de-identify it. Specifically, salted hashing can be used, where the salt value is random and not stored, to obtain an anonymous device identifier. m t,j , j =1,2,…, N W It also obtains the time that the user's electronic device stays within 50 meters of the advertising screen, and identifies anonymous devices. m t,j and equipment dwell time st,j WiFi probe data W t WiFi probe data W t WiFi probe data
[0023] WiFi probe data N W WiFi probe data
[0024] Anonymized QR code data Q t Anonymized QR code data f t,k Anonymized QR code data t t,k Anonymized QR code data k =1,2,…, N Q Anonymized QR code data f t,k Anonymized QR code data t t,k Anonymized QR code data Q t Anonymized QR code data Q t Anonymized QR code data
[0025] Anonymized QR code data N Q Anonymized QR code data
[0026] Further, the anonymized QR code in the embodiments of the present application is updated every 10 minutes, and the user can input 1 and 0 to represent the feedback value after scanning the code, wherein 1 represents satisfaction and 0 represents dissatisfaction.
[0027] Further, since the same user can be detected by multiple devices at the same time, for example, both captured by the millimeter wave radar and detected by the WiFi probe, and the scanning feedback, it is necessary to identify the overlapping users through the space-time matching method to avoid repeated calculation of data. After obtaining the multi-source data, the multi-source data is processed to remove the duplicate data. If there is an intersection among the stay time period of the user in the millimeter wave radar data, the stay time period of the electronic device, and the scanning time, i.e., the overlap is greater than or equal to 2 seconds, it can be determined that it is a potential same user.
[0028] Specifically, the method for removing the duplicate data includes: 1. For each time window, construct the radar user set, the WiFi user set and the QR code user set.
[0029] Radar user set U R is denoted as:
[0030] wherein, is the number of the i-th user identified by the millimeter wave radar. N R
[0031] WiFi user set U W is denoted as:
[0032] wherein, is the number of the i-th device identified by the WiFi probe. N W
[0033] QR code user set U Q is denoted as:
[0034] wherein, is the number of the i-th user of the scanned anonymous QR code. N Q
[0035] 2. Calculate the intersection of the radar user set, the WiFi user set and the QR code user set, and the intersection between any two of them U 重叠 .
[0036] Intersection of the radar user set, the WiFi user set and the QR code user set U 重叠1 is denoted as:
[0037] Intersection of the radar user set and the WiFi user set U 重叠2 is:
[0038] Similarly, the intersection of the radar user set and the QR code user set can also be obtained U 重叠3 :
[0039] Intersection of the WiFi user set and the two-dimensional code user set U 重叠4 :
[0040] Thus, the final intersection after adjustment U 重叠 is expressed as: .
[0041] The intersection is obtained U 重叠 After that, according to the same user determined by the stay period in the millimeter wave radar data, the stay period of the electronic device, and the code scanning time, the numbers can be merged together to mark the same user, but the numbers in each user set are still retained after merging to facilitate subsequent processing.
[0042] 3. Subtract the intersection from the union of the radar user set, the WiFi user set, and the two-dimensional code user set to obtain an effective user set U 有效 : .
[0043] After obtaining the effective user set, the intersection of the effective user set and the two-dimensional code user set U Q is used as the number of effective users who scan the anonymous two-dimensional code after deduplication processing U Q有效 The intersection of the effective user set and the WiFi user set U W is used as the number of effective users detected by the WiFi probe after deduplication processing U W有效 The intersection of the effective user set and the radar user set U R is used as the number of effective users detected by the millimeter wave radar after deduplication processing U R有效 .
[0044] S110, constructing a state space.
[0045] Exemplarily, the state space includes a period type, a real-time crowd density, an average stay duration, a historical effect score, and a weather type, and the average stay duration is an average value of the stay duration of the effective user in the deduplicated multi-source data and the device stay duration.
[0046] S120, substituting the multi-source data into the state space to obtain a corresponding state vector.
[0047] For example, the state vector is formed by substituting the data from multiple sources into the corresponding time period types. s 1. Real-time crowd density s 2. Average length of stay s 3. Historical performance rating s 4 and weather type s The vector obtained after step 5.
[0048] Specifically, time period type s The value 1 has three meanings: 0 indicates off-peak weekdays, 1 indicates peak weekdays, and 2 indicates weekends. Real-time pedestrian density. s 2 is 100m 2 The number of people inside can be statistically analyzed using millimeter-wave radar and WiFi probes. Historical performance rating. s 4. You can take the first 5 time windows. S 总 The mean. Weather type. s The value of 5 also has three possibilities: 0 represents sunny, 1 represents cloudy, and 2 represents rain.
[0049] Therefore, the state vector S It can be represented as: .
[0050] S130 is the optimal action output by the agent in the reinforcement learning framework under the state vector. The optimal action contains the weight of each data in the multi-source data.
[0051] For example, the weight of each data point in the multi-source data is obtained based on the sample size, the number of valid users included in each data point, and the total number of users.
[0052] Specifically, the weights include radar weights. w R WiFi weight w W and QR code weight w Q The calculation formulas are expressed as follows:
[0053]
[0054]
[0055] in, α This is the sample size weighting coefficient, with a value between 0.6 and 0.8. N 总有效 = N R有效 + N W有效 + NQ有效 , represents the total number of effective users, N R有效 is the number of effective users detected by the millimeter wave radar after deduplication processing, N W有效 is the number of effective users detected by the WiFi probe after deduplication processing, N Q有效 is the number of effective users scanned by the anonymous two-dimensional code after deduplication processing, d avg is the average distance of the users detected by the millimeter wave radar, N 总人流 is the total flow of people around the advertising screen in the current period.
[0056] After obtaining the weights, the weights are adjusted according to the real-time flow density and the use scenario of the advertising screen. For example, in a high flow period such as rush hour, the coverage of the WiFi probe is wider, so the WiFi weight is increased, in a low flow period, the active code scanning feedback is more reliable, so the two-dimensional code weight is increased, in a close-range scenario, such as a shopping mall screen, the millimeter wave radar data is more accurate, so the radar weight is increased.
[0057] Further, the optimal action further includes selecting an optimal display version from a plurality of different advertising versions according to the real-time flow density a content .
[0058] Specifically, the different advertising versions include a short version, a long version, and an interactive version, and the corresponding version can be selected according to the real-time flow density in the current scenario, for example, a 15-second short version is selected in a high flow, a 30-second long version is selected in a low flow, and an interactive version is selected in a medium flow.
[0059] Therefore, the optimal action A can be represented as: .
[0060] S140, based on the weights and the single-dimensional effect score of each data in the multi-source data, a comprehensive effect score is calculated.
[0061] Exemplarily, the single-dimensional effect score includes a radar effect score, a WiFi effect score, and a two-dimensional code effect score, the radar effect score is calculated by the number of effective users, the dwell time of effective users, and the distance contained in the millimeter wave radar data, the WiFi effect score is calculated by the number of effective users and the device dwell time contained in the WiFi probe data, and the two-dimensional code effect score is calculated by the number of effective users and the feedback value contained in the anonymous two-dimensional code data.
[0062] Specifically, the radar effect score S RBased on the length of stay and distance weighting:
[0063] Wherein, N R有效 is the number of effective users detected by the millimeter wave radar after deduplication processing, U R有效 is the number of effective users detected by the millimeter wave radar after deduplication processing.
[0064] WiFi effect score S W is expressed as:
[0065] Wherein, N W有效 is the number of effective users detected by the WiFi probe after deduplication processing, U W有效 is the number of effective users detected by the WiFi probe after deduplication processing.
[0066] QR code effect score S Q is expressed as:
[0067] Wherein, N Q有效 is the number of effective users scanned by the anonymous QR code after deduplication processing, U Q有效 is the number of effective users scanned by the anonymous QR code after deduplication processing.
[0068] Comprehensive effect score S 总 is expressed as: .
[0069] S150, according to the comprehensive effect score to determine the corresponding instant reward.
[0070] Exemplarily, the instant reward in the embodiment of the application R contains two parts of effect score and privacy security, which is specifically expressed as follows:
[0071] Wherein, Lambda is a weight coefficient, the value is 0.7-0.8, C 隐私 is a privacy compliance coefficient, its value is fixed at 1, if the data collection is detected to be out of limit, for example, the WiFi probe scanning frequency is too high, then it is reduced to 0.5, in order to force the constraint of privacy protection.
[0072] S160, store the state vector of the current moment S , the optimal action A , the immediate reward R and the state vector of the next moment in the experience replay pool in the reinforcement learning framework, update the network parameters by gradient descent method, and optimize the reinforcement learning model. S
[0073] Exemplarily, the reinforcement learning model adopted by the embodiment of the application is DQN (Deep Q Network), and the agent learns the optimal policy through the DQN, and the goal is to maximize the cumulative reward:
[0074] wherein, G t is the cumulative reward starting from the t moment, Gamma is a discount factor, and the value is 0.9-0.95, R t+k’+1 is the immediate reward starting from the t+1 moment, is the action s corresponding to the policy selected under the state a .
[0075] S170, input the multi-source data to be processed into the reinforcement learning model to obtain the advertising effect evaluation result.
[0076] Exemplarily, the steps S100-S160 are all training processes, and the step S170 is an actual use process. In the actual use process, after obtaining the multi-source data to be processed, the corresponding state vector needs to be established first, then the optimal action is output by the agent, and finally the comprehensive effect score is obtained. The process is similar to the steps S120-S140, and will not be repeated here.
[0077] The embodiment of the application also provides an advertising effect evaluation system based on reinforcement learning, which comprises: a data acquisition module configured to acquire multi-source data generated by a user when watching an advertisement, wherein the multi-source data comprises millimeter wave radar data, WiFi probe data and anonymous two-dimensional code data; a state space construction module configured to construct a state space, wherein the state space comprises a time period type, a real-time crowd density, an average stay duration, a historical effect score and a weather type; a state vector calculation module configured to correspondingly substitute the multi-source data into the time period type s 1. real-time crowd density s 2. average stay duration s 3. historical effect scores 4 and weather type s 5, to obtain a corresponding state vector; an action output module, configured to output, by the agent, an optimal action under the state vector, the optimal action containing a weight of each data in the multi-source data; a score calculation module, configured to calculate a comprehensive effect score based on the weight and a single-dimension effect score of each data in the multi-source data; the single-dimension effect score includes a radar effect score, a WiFi effect score and a two-dimensional code effect score, and the comprehensive effect score S 总 is represented as:
[0078] wherein, w R is a radar weight, w W is a WiFi weight, w Q is a two-dimensional code weight, S R , S W and S Q are the radar effect score, the WiFi effect score and the two-dimensional code effect score, respectively; a reward calculation module, configured to determine a corresponding immediate reward according to the comprehensive effect score; the immediate reward R contains two parts of an effect score and privacy security, and is represented as follows:
[0079] wherein, Lambda is a weight coefficient, C 隐私 is a privacy compliance coefficient; a model optimization module, configured to store the state vector at a current moment, the optimal action, the immediate reward and the state vector at a next moment into an experience replay pool, update network parameters through a gradient descent method, optimize the reinforcement learning model, and learn the optimal strategy through DQN , the goal of which is to maximize the cumulative reward:
[0080] wherein, G t is a cumulative reward starting from t moment, Gamma is a discount factor, R t+k’+1 is an immediate reward starting from t +1 moment, is a states The selected action a The corresponding strategy; The advertisement evaluation module is configured to input the multi-source data to be processed into the reinforcement learning model to obtain an advertisement effect evaluation result.
[0081] The embodiments of the present application further provide a computer storage medium, which stores a plurality of computer instructions for causing a computer to execute the method described above.
[0082] Although the preferred embodiments of the present application have been described, those skilled in the art who understand the basic inventive concept can make additional changes and modifications to the embodiments once they are aware of the basic inventive concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0083] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A method for evaluating advertisement effect based on reinforcement learning, characterized in that, The method comprises: acquiring multi-source data generated by a user when watching an advertisement, the multi-source data comprising millimeter wave radar data, WiFi probe data and anonymous two-dimensional code data; constructing a state space; substituting the multi-source data into the state space to obtain a corresponding state vector; outputting an optimal action by an agent in a reinforcement learning framework under the state vector, the optimal action comprising a weight of each kind of data in the multi-source data; calculating a comprehensive effect score based on the weight and a single-dimension effect score of each kind of data in the multi-source data; determining a corresponding instant reward according to the comprehensive effect score; storing the state vector at a current time, the optimal action, the instant reward and the state vector at a next time into an experience replay pool in the reinforcement learning framework, updating network parameters by a gradient descent method and optimizing a reinforcement learning model based on a deep Q network; inputting to-be-processed multi-source data into the reinforcement learning model to obtain an advertisement effect evaluation result; wherein, after acquiring the multi-source data, the multi-source data is subjected to deduplication processing, and the deduplication processing method comprises: for each time window, constructing a radar user set, a WiFi user set and a two-dimensional code user set; computing an intersection between the set of radar users, the set of WiFi users, and the set of QR code users U 重叠1 an intersection between the set of radar users and the set of WiFi users U 重叠2 an intersection between the set of radar users and the set of QR code users U 重叠3 and an intersection between the set of WiFi users and the set of QR code users U 重叠4 : adjusted final intersection U 重叠 is represented as: wherein, U R for the set of radar users, U W for the set of WiFi users, U Q for the set of QR code users; subtracting the intersection from the union of the set of radar users, the set of WiFi users, and the set of QR code users, to obtain a set of valid users U 有效 : 。 2.The method of claim 1, wherein, when acquiring the millimeter wave radar data, collecting the distance between a user and an advertisement screen by a millimeter wave radar, and when the distance is less than a set distance threshold and the user stays for more than a time threshold, recording the user as an effective user, and taking the staying time of the effective user and the corresponding distance as the millimeter wave radar data; when acquiring the WiFi probe data, performing hash desensitization processing on the MAC address of an electronic device carried by a user to obtain an anonymous device identifier, and also acquiring the time during which the electronic device carried by the user stays within 50 meters of an advertisement screen, taking the anonymous device identifier and the device staying time as the WiFi probe data; when acquiring the anonymous two-dimensional code data, acquiring the feedback value after a user scans a two-dimensional code displayed on an advertisement screen by an electronic device, and also acquiring the code scanning time of the user, taking the feedback value and the code scanning time as the anonymous two-dimensional code data. 3.The method of claim 2, wherein, The state vector comprises a time period type, a real-time crowd density, an average staying time, a historical effect score and a weather type, and the average staying time is an average value of the staying time of an effective user and the device staying time in the multi-source data. 4.The method of claim 2, wherein, The single-dimension effect score comprises a radar effect score, a WiFi effect score and a two-dimensional code effect score, the radar effect score being calculated by the number of effective users, the staying time of the effective users and the distance contained in the millimeter wave radar data: wherein, S R is the radar effect score, N R有效 is the number of valid users detected by the millimeter wave radar after deduplication processing, U R有效 is the number of valid users detected by the millimeter wave radar after deduplication processing, r t,i is the stay duration of the t th valid user in the i th minute, d t,i is the distance between the t th user and the advertising screen in the i th minute. the WiFi effect score being calculated by the number of effective users and the device staying time contained in the WiFi probe data: wherein, S W a WiFi effect score, N W有效 a number of valid users detected by the WiFi probe after deduplication, U W有效 a number of valid users detected by the WiFi probe after deduplication, s t,j a first t minute dwell time of a first j electronic device carried by the user. the two-dimensional code effect score being calculated by the number of effective users and the feedback value contained in the anonymous two-dimensional code data: wherein, S Q a score for the two-dimensional code effect, N Q有效 a valid user number of the de-duplicated scanned anonymous two-dimensional code, U Q有效 a number of the valid user of the de-duplicated scanned anonymous two-dimensional code, f t,k a first t minute feedback value of a first k code scanning user through an electronic device after scanning an anonymous two-dimensional code displayed on an advertising screen. 5.The method of claim 1, wherein, The weight of each kind of data in the multi-source data is obtained according to a sample size, the number of effective users contained in each kind of data and a total number of users. 6.The method of claim 1, wherein, After obtaining the weight, the weight is adjusted according to the real-time crowd density and the use scenario of the advertisement screen. 7.The method of claim 1, wherein, The optimal action further includes an optimal display version selected from a plurality of different advertisement versions according to a real-time crowd density.
8. A system for applying the method for evaluating the effect of advertising based on reinforcement learning according to any one of claims 1 to 7, characterized in that, Comprise: A data acquisition module for acquiring multi-source data generated by a user when watching an advertisement, the multi-source data including millimeter wave radar data, WiFi probe data, and anonymous two-dimensional code data; A state space construction module for constructing a state space; the state space includes time period type, real-time crowd density, average stay duration, historical effect score, and weather type; a state vector calculation module, configured to correspondingly substitute the multi-source data into a time period type in the state space s 1. Real-time crowd density s 2. Average stay duration s 3. Historical effect score s 4. Weather type s 5. Obtain the corresponding state vector; An action output module for outputting an optimal action by an intelligent agent under the state vector, the optimal action containing a weight of each type of data in the multi-source data; The score calculation module is configured to calculate a comprehensive effect score based on the weight and a single-dimension effect score of each data in the multi-source data; the single-dimension effect score includes a radar effect score, a WiFi effect score, and a two-dimensional code effect score, and the comprehensive effect score S 总 is represented as: wherein, w R is a radar weight, w W is a WiFi weight, w Q is a QR code weight, S R , S W and S Q are the radar effect score, the WiFi effect score, and the QR code effect score, respectively. The reward calculation module is configured to determine a corresponding instant reward according to the comprehensive effect score; and the instant reward R The effect score and the privacy security are combined, and are expressed as follows: wherein, Lambda is a weight coefficient, C 隐私 is a privacy compliance coefficient; A model optimization module is configured to store the state vector at the current time, the optimal action, the immediate reward, and the state vector at the next time into an experience replay pool, update network parameters by using a gradient descent method, optimize the reinforcement learning model, and learn an optimal policy by using DQN The objective is to maximize the cumulative reward: wherein, G t is the cumulative reward from t time step, Gamma is a discount factor, R t+k’+1 is the immediate reward from t time step, is the action s chosen under the policy a corresponding policy; An advertisement evaluation module for inputting the multi-source data to be processed into the reinforcement learning model to obtain an advertisement effect evaluation result.
9. A computer storage medium, characterized in that The computer storage medium stores a plurality of computer instructions, and the plurality of computer instructions are used to make a computer execute the method in any one of claims 1-7.
Citation Information
Patent Citations
Advertisement management method utilizing fresh or takeaway food delivery service, storage medium, server and system
CN107358470A
Large-size screen targeted advertising system and method based on multi-source heterogeneous data analysis
CN108428158A
Advertisement putting method, device and system and storage medium
CN110706030A
Advertisement delivery tracking method and device and terminal equipment
CN110751502A
Method and device for pushing advertising value of advertisement screen, and advertisement screen
CN113674024A