A method for constructing a macro-law event prediction model based on reinforcement learning
Through a macro-law event prediction model based on reinforcement learning, combined with the Naive Bayes algorithm and Monte Carlo tree search model, the problem of lack of effective event prediction methods in the existing technology is solved, and high accuracy prediction of future event behavior is achieved.
Patent Information
- Application Number
- CN202210184289.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-02-28
AI Technical Summary
At this stage, my country lacks effective social behavior event prediction methods, and it is difficult to predict possible future events through macro historical events.
The macro-law event prediction model construction method based on reinforcement learning is adopted, and event classification is performed by statistics and cleaning event data, and the Naive Bayes algorithm is used to perform event classification, and combined with reinforcement learning Monte Carlo tree search model to simulate individual activities and predict future event behavior.
A high-accuracy prediction of possible event behaviors in future scenarios is achieved, a solidified model is formed, and social efficiency is improved.
Smart Images

Figure CN114743258B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and particularly to a method for constructing a macro-law event prediction model based on reinforcement learning. Background Art
[0002] At the present stage in China, with the development of the times and the progress of science and technology, various emerging things have emerged, and the social behaviors of humans are also changing day by day. How to predict various behaviors in society through macro historical events based on real time, space, and social environment information, so as to predict possible events in advance, can greatly improve social efficiency.
[0003] Currently, there are few event prediction methods for social behaviors in society, and there is a gap in this regard. Therefore, we propose a method for constructing a macro-law event prediction model based on reinforcement learning. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0005] A method for constructing a macro-law event prediction model based on reinforcement learning of the present invention includes the following steps:
[0006] S1. Stat all event data, clean the data, and remove invalid values and missing values to ensure the accuracy of the data;
[0007] S2. Conduct data statistics based on the region where the event occurs, count the events, people, and vehicle flows belonging to the region division, classify different events, and count the individual characteristics of the personnel corresponding to the events;
[0008] S3. Based on the Naive Bayes algorithm, classify different types of events to obtain classifier A, and predict the tendencies of different individuals in the crowd towards different events and predict the event types;
[0009] S4. Conduct Naive Bayes algorithm classification based on the event spatio-temporal data and types to obtain classifier B, and calculate the event occurrence probability of personnel in a real specific time period;
[0010] S5. According to the personnel attribute distribution in the background environment, randomly generate a large number of personnel, classify the behavior tendencies of the crowd based on classifier A, and find out the crowds where different events occur;
[0011] S6. Use the Monte Carlo Tree Search model MCTS of reinforcement learning to simulate events for the crowd in S5, and generate event probability predictions based on time periods using the classifier B regarding the impact of time periods on behaviors in S4. Collaborate with the eigenvalue reward function of the category to which the personnel behavior tendency belongs, and calculate the search of this Monte Carlo Tree Search model MCTS of reinforcement learning to achieve the prediction of the behavior situation in this scenario.
[0012] As a preferred technical solution of the present invention, the calculation formula of the Naive Bayes algorithm in S3 and S4 is:
[0013] As a preferred technical solution of the present invention, the event spatio-temporal data in S4 includes the season where the event is located, whether it is a legal rest day, whether it is a legal holiday or before and after, the event time point, the regional pedestrian flow, and the number of monitors in the region.
[0014] As a preferred technical solution of the present invention, the operation process of the Monte Carlo Tree Search model MCTS of reinforcement learning in S6 includes:
[0015] a. Selection: First, select un-explored child nodes. If all have been searched, then select the child node with the largest UCB value;
[0016] b. Expansion: Take a step in the selected child node to create a new child node;
[0017] c. Simulation: Start simulating the created node until the entire search tree reaches the leaf node and then end, and then calculate the total score of this expanded node;
[0018] d. Backpropagation: Feed back the score of the expansion to all the previous parent nodes, and update the quality value Q(v′) and the access times N(v′) of these nodes to facilitate the subsequent calculation of the UCB value.
[0019] As a preferred technical solution of the present invention, the UCB calculation formula in the selection process is as follows:
[0020]
[0021] Where v′ represents the current tree node, v represents the parent node, Q represents the weighted value of the cumulative reward function of this tree node, N represents the number of times this tree node is selected, and c is a constant.
[0022] The beneficial effects of the present invention are:
[0023] The method for constructing a macro-law event prediction model based on reinforcement learning first conducts data analysis on historical events, classifies the characteristic attributes based on different populations using a classification algorithm, simulates individual activities using a reinforcement learning model in real time and real environment, combines the event characteristic data in a specific period and specific environment, aims at a high prediction accuracy rate, scores the personnel activities within a certain period to train their behavior choices, forms a solidified model, and then can predict the event behaviors that may occur in future scenarios based on event laws. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0025] Figure 1 is a flowchart of a method for constructing a macro-law event prediction model based on reinforcement learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0027] Embodiment 1
[0028] As Figure 1 shown, the method for constructing a macro-law event prediction model based on reinforcement learning according to the present invention first conducts data analysis on historical events, classifies the characteristic attributes based on different populations using a classification algorithm, simulates individual activities using a reinforcement learning model in real time and real environment, combines the event characteristic data in a specific period and specific environment, aims at a high prediction accuracy rate, scores the personnel activities within a certain period to train their behavior choices, forms a solidified model, and predicts the event behaviors that may occur in future scenarios. The specific steps are as follows:
[0029] S1. Statistically analyze all event data, clean the data, and remove invalid values and missing values to ensure the accuracy of the data;
[0030] S2. Conduct data statistics based on the regional division to which the event belongs, count the events, people, and vehicle flows belonging to the regional division, classify different events, and count the individual characteristics of the personnel corresponding to the events;
[0031] S3. Classify different types of events based on the Naive Bayes algorithm to obtain a classifier. Its main calculation formula is:
[0032]
[0033] By using this classifier, it is possible to predict the occurrence tendencies of different individuals in a population towards different events and simultaneously predict the types of those events.
[0034] S4. To simulate an objective and real environment and explore the influence of different times on the probability of event occurrence, based on event spatio-temporal data, including the season in which the event occurs, whether it is a legal rest day, whether it is a legal holiday or around it, the event time point (morning, noon, afternoon, evening), the regional pedestrian flow, the number of monitors in the region, etc., naive Bayes classification is performed based on the event type to obtain a classifier, and its calculation formula is the same as that in S3.
[0035] S5. According to the distribution of personnel attributes in the background environment, a large number of personnel are randomly generated, and the behavior tendencies of the population are classified based on the classifier in S3 to identify the populations in which different events occur.
[0036] S6. To simulate the event behaviors of different populations in a real environment in step 5, the reinforcement learning Monte Carlo tree search model (hereinafter referred to as MCTS) is used to simulate events for the population in S5 under the condition that events can occur for a long time. Based on the classifier in S4 regarding the influence of time periods on behaviors, a time period-based event probability prediction is generated, and the MCTS search for this time is calculated in collaboration with the eigenvalue reward function of the category to which the personnel behavior tendency belongs.
[0037] Among them, the main operation process of MCTS includes:
[0038] a. Selection: Find a best node worthy of search in the entire tree. Generally, the strategy is to first select un-explored child nodes. If all have been searched, then select the child node with the largest UCB value.
[0039] b. Expansion: Take a step in the previously selected child node to create a new child node. Generally, the strategy is to randomly perform an operation that cannot be repeated with the previous child nodes.
[0040] c. Simulation: Start simulation from the newly expanded node until the entire search tree reaches the leaf node and ends. In this way, the total score of this expanded node can be calculated.
[0041] d. Backpropagation: The scores obtained from the previous expansion are fed back to all the previous parent nodes to update the quality values Q(v′) and the number of visits N(v′) of these nodes to facilitate subsequent calculation of the UCB value.
[0042] Among them, the UCB calculation formula in the selection process is as follows:
[0043]
[0044] Among them, v' represents the current tree node, v represents the parent node, Q represents the weighted value of the cumulative reward function of this tree node, N represents the number of times this tree node is selected, and c is a constant.
[0045] By selecting the behavior with the highest probability of personnel in a specific period and specific environment through MCTS, the behavior situation in this scenario is predicted accordingly.
[0046] Embodiment 2
[0047] Taking the event to be analyzed "Event confrontation prediction in a certain responsibility area of a certain city" as an example for description:
[0048] 1. First, based on the existing historical event data in a certain responsibility area of a certain city, data cleaning is carried out, and the cleaned data is as follows:
[0049] Table 1 Historical event table of a certain responsibility area
[0050]
[0051] 2. Based on the vehicle flow and pedestrian flow in this responsibility area, statistics are carried out, and the data is as follows:
[0052] Table 2 Flow statistics table of a certain responsibility area by time period
[0053] Monitoring Number Total Flow Morning Noon Afternoon Evening Night ********* 1663 438 292 846 86 1 ********* 714 242 102 362 7 1 ********* 1577 382 262 753 178 2 ********* 1089 321 170 497 99 2
[0054] 3. Based on the Naive Bayes classifier, the following different personnel attribute data is classified to generate a classifier:
[0055] Table 3 Personnel record table
[0056]
[0057] 4. Based on the event information and regional data, generate classifiers for the event occurrence time and space based on different events;
[0058] 5. Based on the actual distribution of personnel attributes in a certain city according to the process, a total of 10,000 personnel are randomly generated;
[0059] Table 4 Random personnel generation simulation table
[0060]
[0061] 6. Classify the above personnel through the classifier generated in S3 to obtain a group of people with specific behavior tendencies, and perform Monte Carlo tree search on this group of people in January 2022. Based on the classifier generated in step 4, predict the behavior probability during this period. Use The UCB formula calculates each node 1,000 times to obtain the behaviors of the population with specific behavioral tendencies within that month. By determining the specific time periods of the population that has performed a certain behavior within that month, the prediction of this event in the responsible area in January 2022 is achieved. The simulation results are shown in Table 5:
[0062] Table 5 Simulation Results of a Certain Event Based on MCTS
[0063]
[0064] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a prediction model of macroscopic regular events based on reinforcement learning, characterized in that, it includes the following steps: S1. Statistically analyze all event data, clean the data, remove invalid values and missing values, and ensure the accuracy of the data; S2. Conduct data statistics based on the regional division to which the event belongs. Statistically analyze the events, people, and vehicle flows belonging to the regional division, classify different events, and statistically analyze the individual characteristics of the personnel corresponding to the events; S3. Based on the Naive Bayes algorithm, classify different types of events to obtain classifier A, and predict the tendencies of different individuals in the crowd towards different events and the predicted event types; S4. Conduct Naive Bayes algorithm classification based on event spatio-temporal data and types to obtain classifier B, and calculate the probability of events occurring for personnel during specific real-time periods; S5. According to the distribution of personnel attributes in the background environment, randomly generate a large number of personnel, classify the behavior tendencies of the crowd based on classifier A, and find out the crowds where different events occur; S6. Use the reinforcement learning Monte Carlo tree search model MCTS to simulate events for the crowd in S5, and generate event probability predictions based on time periods based on classifier B regarding the influence of time periods on behavior in S4. Collaborate with the eigenvalue reward function of the category to which the personnel behavior tendency belongs, and calculate the search of this reinforcement learning Monte Carlo tree search model MCTS to achieve the prediction of the behavior situation in this scenario.
2. The method for constructing a prediction model of macroscopic regular events based on reinforcement learning according to claim 1, characterized in that, The calculation formula of the Naive Bayes algorithm in S3 and S4 is as follows:
3. The method for constructing a prediction model of macroscopic regular events based on reinforcement learning according to claim 1, characterized in that, the event spatio-temporal data in S4 includes the season in which the event is located, whether it is a legal rest day, whether it is a legal holiday or before and after, the event time point, the regional pedestrian flow, and the number of monitors in the region.
4. The method for constructing a prediction model of macroscopic regular events based on reinforcement learning according to claim 1, characterized in that, the operation process of the reinforcement learning Monte Carlo tree search model MCTS in S6 includes: a. Selection: First, select un-explored child nodes. If all have been searched, then select the child node with the largest UCB value; b. Expansion: Take a step in the selected child node to create a new child node; c. Simulation: Start simulating the created node until the entire search tree reaches the leaf node and end, and then the total score of this expanded node can be calculated; d. Backpropagation: Feed back the score of the expansion to all the previous parent nodes, and update the quality value Q(v′) and the number of visits N(v′) of these nodes to facilitate subsequent calculation of the UCB value.
5. The method for constructing a prediction model of macroscopic regular events based on reinforcement learning according to claim 4, characterized in that, the UCB calculation formula during the selection process is as follows: where v′ represents the current tree node, v represents the parent node, Q represents the weighted value of the cumulative reward function of this tree node, N represents the number of times this tree node is selected, and c is a constant.
Citation Information
Patent Citations
Abnormal event processing method and device based on Monte Carlo tree search
CN112700005A
Public digital life scene rule model prediction and early warning method based on deep Bayesian network
CN113010572A