Emotional infection method based on parameterized reinforcement learning
By employing an emotion contagion method based on parametric reinforcement learning, this method identifies crowd groups and individual social identities, calculates the emotional intensity and behavioral weights of individuals, and addresses the lack of dynamic adaptability in existing crowd emotion contagion methods. It achieves the capture of complex behavioral decisions and environmental adaptation, thereby improving the accuracy of simulating crowd behavior.
Patent Information
- Application Number
- CN202511024699.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-07
AI Technical Summary
Existing methods for infecting crowds with emotions lack dynamic adaptability to the individual's environment, fail to fully capture the complex behavioral decisions driven by emotions, and lack learning ability, making it difficult to adapt to complex scene changes in different environmental layouts and high-density crowds.
The emotion contagion method based on parametric reinforcement learning acquires real-life videos of crowds, identifies crowd groups and individual social identities, calculates the emotional intensity and behavioral weights of individuals, constructs an individual emotion contagion model using Gaussian mixture models and OCEAN models, and combines parametric reinforcement learning methods to simulate the evacuation behavior of individuals under different emotional intensities and environments.
It enables individuals to dynamically adapt to the environment, captures complex behavioral decisions, improves learning ability, adapts to complex scene changes in different environmental layouts and high-density populations, and improves the realism and accuracy of simulated crowd behavior.
Smart Images

Figure CN120913148A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of crowd behavior research, and particularly relates to an emotion contagion method based on parameterized reinforcement learning. BACKGROUND
[0002] In emergency scenarios, the behavior of a crowd is highly influenced by individual emotions and environmental factors (the presence of an emergency event source and the surrounding neighbor individuals). With the increasing frequency of emergency events, how to simulate the behavior of a crowd in emergency situations has become an important issue for safety management, urban planning, and emergency evacuation design. The study of crowd simulation involves the intersection of multiple disciplines such as artificial intelligence, psychology, and mathematical modeling to understand how individuals make decisions in a crowded and chaotic environment in emergency situations.
[0003] Traditional crowd behavior simulation methods are mainly divided into crowd behavior simulation based on micro-level and crowd behavior simulation based on macro-level. Among them, crowd behavior simulation based on micro-level focuses on describing the interaction between individuals in a crowd, and is simulated based on individuals, but usually relies on pre-set behavior rules and lacks dynamic response to the emotional state of individuals; unlike this, crowd behavior simulation based on macro-level focuses on the movement pattern of the whole crowd, but cannot reflect the decision-making differences of individuals. Since in emergency scenarios, emotion contagion (or emotion transmission) in a crowd is an important factor leading to changes in group behavior, traditional methods of emotion contagion in a crowd mostly use rules or simple linear propagation mechanisms to simulate the propagation of emotions among individuals in a crowd. Such emotion contagion models usually assume that each neighbor individual (or neighbor individual) of an individual has the same influence on the propagation of the individual's emotions, but ignore the influence of complex social psychological factors such as individual social identity, empathy, and attention enhancement.
[0004] In order to consider the influence of social psychological factors on the propagation of emotions among individuals in a crowd, existing methods of crowd emotion contagion combine emotion contagion models with mutual speed obstacle models, so that the movement of an individual is affected by its own emotional state, in order to more realistically reproduce the behavior of a crowd in an emergency scenario.
[0005] However, existing methods of crowd emotion contagion have deficiencies: they only simulate the emotion contagion of individuals in a crowd based on rules or static parameters, but lack dynamic adaptability to the environment of the individual in the crowd, cannot fully capture the complex behavior decisions of individuals driven by emotions (such as aggressive behavior and group separation behavior under panic emotions), and also lack learning ability, making it difficult to adapt to changes in complex scenarios such as different environmental layouts and high-density crowds. SUMMARY
[0006] The technical problem to be solved by the present application is to provide an emotion infection method based on parameterized reinforcement learning with stronger dynamic adaptive capability and complex scene adaptability.
[0007] The technical solution adopted by the present application to solve the above technical problem is an emotion infection method based on parameterized reinforcement learning, characterized by comprising the following steps:
[0008] Step 1: obtaining a real crowd video with social identity attributes;
[0009] Step 2: performing recognition processing on the obtained real crowd video to obtain all groups constituting the real crowd and the social identity of each individual;
[0010] Step 3: calculating the emotional intensity of each individual based on the influence of the emergency event source in the surrounding environment of each individual and the emotional influence of the surrounding neighbor individuals;
[0011] Step 4: calculating the behavior weight of the behavior corresponding to each individual under different emotional intensities based on the obtained emotional intensity of each individual; wherein the emotional intensity of an individual corresponds to the behavior of the individual one-to-one;
[0012] Step 5: obtaining the evacuation behavior of each individual under different emotional intensities and environmental conditions based on the same reward function and the behavior weight of all individuals.
[0013] In the improved emotion infection method based on parameterized reinforcement learning, in step 2, a Gaussian mixture model is used to perform recognition processing on the real crowd video to obtain all groups of the real crowd and the social identity of each individual; wherein the recognition processing process comprises the following steps:
[0014] Step a1: collecting the historical trajectory information of each individual in the real crowd video; wherein the historical trajectory information of an individual is the movement information of the individual in the real crowd video within the time period corresponding to the real crowd video, and the movement information includes the motion path, position coordinates and movement direction;
[0015] Step a2: performing normalization processing on the collected historical trajectory information of each individual to obtain the normalized historical trajectory information of each individual;
[0016] Step a3: pre-calculating the weight of each group and evaluating the optimal group number of the real crowd in theory and the number of individuals in each group based on the contour coefficient; wherein the weight value of a group is the mixing coefficient of each group in the Gaussian mixture model;
[0017] Step a4, the crowd in the real crowd video is processed by spatiotemporal consistency clustering using a Gaussian mixture model and evaluating the optimal number of groups obtained, to obtain a plurality of groups composed of different individuals; wherein the group and the cluster center are in one-to-one correspondence;
[0018] Step a5, the attention weight of each individual to each group is calculated using the Mahalanobis distance of each individual to each cluster center and the direction similarity between the motion direction of each individual and the direction from the individual position to the cluster center; wherein:
[0019] χ ij = αexp(MDIS ij ) + (1-α)Sim ij ; α = 0.5;
[0020]
[0021] wherein χ ij is the attention weight of individual i to the jth group, α is a relative importance index for controlling the Mahalanobis distance and the direction similarity, MDIS ij is the Mahalanobis distance between individual i and the cluster center j of the jth group, Sim ij indicates the direction similarity between the motion direction of individual i and the direction from the position of individual i to the cluster center j; x i indicates the position of individual i, σ j indicates the position of the cluster center j, σ j -x i indicates the vector from the position of individual x i to the position of the cluster center σ j ; v i indicates the motion velocity vector of individual i; Sim ij = 1 indicates that the motion direction of individual i is completely directed to the cluster center j, Sim ij = 0 indicates that the motion direction of individual i is completely deviated from the cluster center j;
[0022] Step a6, the social identity strength of each individual to each group is calculated respectively; wherein:
[0023]
[0024] wherein, is the social identity strength of individual i to the jth group, κ ij is the degree of belonging of individual i to the jth group.
[0025] Further, in the emotion infection method based on parameterized reinforcement learning, the individual is affected by the emergency event source in the surrounding environment of the individual to cause the panic emotion change of the individual; wherein:
[0026]
[0027] wherein, represents the emergency event source H a the panic emotion change caused by the emotion of the individual i, P represents the emergency event source H a any position in the danger field formed, represents the emergency event source H a at the position of the danger field, dis sv represents the distance perceived by the individual i from the emergency event source H a , and U represents the duration of the emergency event source.
[0028] Further improvement, in the present application, the emotion infection method based on parameterized reinforcement learning further comprises: constructing an individual emotion infection model after being affected by socialization; wherein, the construction process of the individual emotion infection model comprises the following steps:
[0029] Step b1, taking the OCEAN model as a measuring tool for measuring individual personality; wherein, in the OCEAN model, Φ=(ψ 0 , ψ C , ψ E , ψ A , ψ N ); Φ is the individual personality, and the personality includes five mutually orthogonal factors of openness ψ 0 , conscientiousness ψ C , extraversion ψ E , agreeableness ψ A and neuroticism ψ N ;
[0030] Step b2, constructing an individual initial empathy model of the individual affected by the individual personality; wherein:
[0031] ε i =(0.35ψ 0 +0.177ψ C +0.135ψ E +0.312ψ A +0.021ψ N +1) / 2;
[0032] wherein, ε i is the individual initial empathy level of the individual i;
[0033] Step b3, calculating the social identity degree of each individual according to the obtained social identity strength of the individual to each group; wherein:
[0034]
[0035] wherein g i,social_identity represents the social identity of individual i, ε i,C represents the level of empathy of individual i to the members of extreme social identity; ε i,U represents the level of transcendental empathy of individual i to the members without obvious relationship or social identity relationship; represents the exponential adjustment value of the difference between individual identity and collective identity of individual i; ε i,C is greater than ε i,U , indicating that the influence of collective identity on the emotion of individual i increases; ε i,C is less than ε i,U , indicating that the influence of collective identity on the emotion of individual i decreases;
[0036] Step b4, constructing an empathy breakdown model of individual empathy affected by emotion in an emergency scene; wherein:
[0037]
[0038] wherein ε I represents the critical threshold of empathy of individual in the case of extreme irrationality, ε I is taken from a Gaussian distribution with a mean of 0.7-0.4ε and a standard deviation of (0.7-0.4ε) / 10; panT represents the critical threshold of empathy of individual affected by panic emotion, and the critical threshold panT is negatively correlated with the neuroticism factor of individual, and the critical threshold panT is taken from a Gaussian distribution with a mean of 0.5-0.5ψ N and a standard deviation of (0.5-0.5ψ N ) / 10; variable E represents the emotion intensity of individual, variable ε represents the empathy level of individual, and variable E and variable ε are positively correlated;
[0039] Step b5, obtaining the revised empathy level of individual according to the individuality, social identity and empathy breakdown model of individual; wherein:
[0040] ε i,modify =f(OCEAN,E)*g i,social_identity ;
[0041] wherein ε i,modify represents the revised empathy level of individual i.
[0042] In order to calculate the emotion intensity of individual affected by the emotions of all neighbors around it, the emotion infection method based on parameterized reinforcement learning of the application further comprises steps c1-c4:
[0043] Step c1, constructing an emotion infection intensity model of all neighbor individuals around individual to the individual to represent the emotion influence of surrounding neighbor individuals on individual; wherein:
[0044]
[0045]
[0046] wherein, is the emotional infection strength of all the neighboring individuals on individual i, is the connection strength of neighboring individual j, E j is the emotional strength value of neighboring individual j, γ ij is the connection strength between individual i and neighboring individual j, R i is the sum of connection strength between individual i and all the neighbors within the perception range of individual i, N is the total number of all the neighboring individuals within the perception range of individual i;
[0047] is the connection strength between individual i and neighboring individual j, γ ij is calculated as follows:
[0048]
[0049] wherein, θ j is the potential attention of individual i to neighboring individual j within the perception range of individual i, is the basic attention of individual i to neighboring individual j within the perception range of individual i (or the emotional infection weight), α' is the basic attention coefficient between individual i and neighboring individual j, is the social distance between individual i and neighboring individual j; N i is the total number of all the neighboring individuals within the perception range of individual i, m is the number of one of all the neighboring individuals of individual i;
[0050] E j is the emotional value of neighboring individual j, A j is taken from a Gaussian distribution with a mean of 0.6-0.6ψ N and a standard deviation value of (0.6-0.6ψ N ) / 10; μ i is the emotional preference of individual i;
[0051] Step c2, a mapping relationship between the OCEAN personality model and the emotional preference of individual i is established; wherein:
[0052]
[0053] ω NP + ω NP = 1; ω NP ∈(0, 1), ω CP ∈(0, 1);
[0054]
[0055] wherein, denotes the neurotic trait factor value of individual i, denotes the conscientiousness trait factor value of individual i, ω NP denotes the neurotic trait weight factor, ω CP denotes the conscientiousness trait weight factor;
[0056] Step c3, calculate the emotional intensity of each individual according to the emergency event source influence in the individual's surrounding environment and the emotional influence of the surrounding neighbors on the individual; wherein the emotional intensity of individual i is marked as E i :
[0057]
[0058] wherein, denotes the emergency event source H a the emotional infection intensity of individual i, denotes the emotional infection intensity of all neighboring individuals around individual i on individual i;
[0059] Step c4, calculate the real-time emotional intensity of the individual based on the obtained emotional intensity of the individual; wherein the real-time emotional intensity of individual i at time t is marked as
[0060]
[0061] wherein, is the real-time emotional intensity of individual i at time t-1, β is a decay proportion factor and is proportional to the neurotic trait.
[0062] Further improved, in the present application, the emotional infection method based on parameterized reinforcement learning further comprises: using a parameterized reinforcement learning method to simulate the behavior-driven simulation of the simulated heterogeneous population with diverse behaviors; wherein the parameterized reinforcement learning method includes a training phase and a runtime simulation phase; wherein:
[0063] The training phase includes:
[0064] Configure multiple agents in an emergency scene with the same emergency event source and different environmental conditions;
[0065] Let all agents share the same reward function configuration;
[0066] Based on the panic emotional value of each agent caused by the emergency event source, discretize to generate four types of emotional states of the agent;
[0067] Make dynamic adjustments to the relative importance of different behavior reward signals of the agent;
[0068] The runtime simulation stage includes:
[0069] According to the current emotion classification of the calculated agent, the evacuation agent is configured with the behavior corresponding to the emotion mapping in real time, so that the agent can dynamically respond to the changing emotional state during the simulation process.
[0070] Further, in the emotion infection method based on parameterized reinforcement learning, the training strategy adopted in the training stage is set as follows:
[0071] Obtain the observable state of each agent observing other agents around itself; wherein the observable state includes the position of the agent, the speed of the agent and the radius of the agent;
[0072] Let each agent evaluate the state gap between the current state and the desired state, and adjust the agent's own movement according to its real-time evaluation of the environment and the behavior of other agents; wherein the state of the agent at the current time t is denoted as The observed state of the neighbor agent j at the current time t is denoted as The combined input state of the agent is denoted as
[0073] Map the current state of the agent to its action behavior in a preset policy control manner to obtain the maximum expected return; wherein the preset policy control manner is as follows:
[0074]
[0075] Wherein, R(s t ,a t ) is the reward obtained by the agent at time t, a t is the action taken by the agent at time t, ρ is the discount factor, ρ △t is the discount factor in the time period △t, v pref is the normalization term in the discount factor ρ, P(s t ,a t ,s t +△t ) is the transition probability from time t to time t+△t, V * is the optimal value function.
[0076] Compared with the prior art, the emotion infection method based on parameterized reinforcement learning has the advantages that: the emotion infection method based on parameterized reinforcement learning of the application obtains all groups and the social identity of each individual constituting the real crowd by identifying the obtained real crowd video with the social identity attribute, calculates the emotion intensity of each individual based on the influence of the emergency event source in the surrounding environment of each individual and the emotion influence of the surrounding neighbor individuals, and calculates the behavior weight of the behavior corresponding to each individual under different emotion intensities based on the obtained emotion intensity of each individual, and then obtains the evacuation behavior of each individual under different emotion intensities and environmental conditions based on the same reward function and the behavior weight of all individuals. In this way, the real situation that an individual is influenced by the emergency event source in the environment and the emotion influence of the surrounding neighbor individuals is considered, the dynamic adaptability of the individual to the crowd environment is realized, each individual has a dynamic attribute, the complex behavior decision of the individual under the emotion driving can be fully captured, the learning ability is improved, and the complex scene changes of different environmental layouts and high-density crowds are well adapted. BRIEF DESCRIPTION OF DRAWINGS
[0077] Figure 1 FIG. 1 is a flowchart of the emotion infection method based on parameterized reinforcement learning in the embodiment of the application;
[0078] Figure 2 FIG. 3 is a flowchart of the identification process of the real crowd video by using the Gaussian mixture model in the embodiment of the application. DETAILED DESCRIPTION
[0079] The application will be further described in detail below with reference to the embodiments of the drawings.
[0080] The embodiment provides an emotion infection method based on parameterized reinforcement learning. Specifically, referring to FIG. 1, the emotion infection method based on parameterized reinforcement learning of the embodiment includes the following steps: Figure 1
[0081] Step 1, obtaining a real crowd video with a social identity attribute;
[0082] Step 2, identifying the obtained real crowd video to obtain all groups constituting the real crowd and the social identity of each individual;
[0083] Step 3, calculating the emotion intensity of each individual based on the influence of the emergency event source in the surrounding environment of each individual and the emotion influence of the surrounding neighbor individuals;
[0084] Step 4, calculating the behavior weight of the behavior corresponding to each individual under different emotion intensities based on the obtained emotion intensity of each individual; wherein the emotion intensity of the individual corresponds to the behavior of the individual one-to-one;
[0085] The behavior weight is specially configured to simulate the psychological state and behavior result of the individual in the evacuation process. For example, when the panic level is high, the weight of the action related to the rapid escape towards the exit direction can be increased; while when the panic is low, the weight can be more inclined to promote more calm action;
[0086] Step 5, based on the same reward function and the behavior weight of all individuals obtained, the evacuation behavior of each individual under different emotional intensity and environmental conditions is obtained.
[0087] Specifically, referring to Figure 2 As shown in step 2 of the embodiment, the real crowd video is identified by using a Gaussian mixture model to obtain all groups of the real crowd and the social identity of each individual; wherein the identification process includes the following steps a1-a6:
[0088] Step a1, collecting the historical trajectory information of each individual in the real crowd video; wherein the historical trajectory information of the individual is the movement information of the individual in the real crowd video within the time period corresponding to the real crowd video, and the movement information includes the motion path, the position coordinates and the moving direction;
[0089] Step a2, normalizing the historical trajectory information of each individual collected to obtain the normalized historical trajectory information of each individual; wherein normalizing the historical trajectory information of the individual is a conventional technical means in the art, which will not be described here;
[0090] Step a3, pre-calculating the weight of each group, and evaluating the optimal group number of the real crowd in theory and the number of individuals in each group based on the contour coefficient; wherein the weight value of the group is the mixing coefficient of each group in the Gaussian mixture model; the Gaussian mixture model here is a conventional technical means known to those skilled in the art, which will not be described here;
[0091] Step a4, using the Gaussian mixture model and the optimal group number evaluated to perform spatiotemporal consistency clustering on the crowd in the real crowd video to obtain a plurality of groups composed of different individuals; wherein the group and the cluster center are in one-to-one correspondence; the number of groups obtained is equal to the optimal group number;
[0092] Step a5, calculating the attention weight of each individual to each group by using the Mahalanobis distance between each individual and each cluster center and the direction similarity between the motion direction of each individual and the direction from the individual position to the cluster center;
[0093] χ ij = αexp(MDIS ij ) + (1-α)Sim ij ; α = 0.5;
[0094]
[0095] where χ ij is the attention weight of individual i to the jth group, a is the relative importance index for controlling the Mahalanobis distance and the direction similarity, MDIS ij is the Mahalanobis distance of individual i to the cluster center j of the jth group, Sim ij represents the direction similarity between the moving direction of individual i and the direction from the position of individual i to the cluster center j; x i represents the position of individual i, σ j represents the position of the cluster center j, σ j -x i represents the vector from the position of individual x i to the position of the cluster center σ j ; v i represents the moving velocity vector of individual i; Sim ij = 1 indicates that the moving direction of individual i is completely directed to the cluster center j, Sim ij = 0 indicates that the moving direction of individual i is completely deviated from the cluster center j.
[0096] The attention weight of individual to group reflects the importance of the group in the overall social identity of the individual; even if the individual has a certain degree of membership to multiple groups, not all groups have equal importance to the identity of the individual;
[0097] Step a6, respectively calculate the social identity strength of each individual to each group; wherein:
[0098]
[0099] wherein, is the social identity strength of individual i to the jth group, κ ij is the degree of belonging (or membership) that individual i feels to the jth group. It should be noted that the social identity strength of individual to group represents the degree of association between individual and the social group; the higher the membership of individual, the closer the connection between individual and social group, and the more likely to accept and internalize the values and behavior standards of the social group.
[0100] It should be noted that in this embodiment, in step 3, the individual is affected by the emergency event source in the individual's surrounding environment to cause the panic emotional change of the individual, so as to image the emotional strength of the individual; wherein:
[0101]
[0102] wherein, represents the emergency event source Ha P represents the panic emotion change caused by the emotional infection of the individual i, and H represents the emergency event source a Any position in the danger field, H represents the emergency event source a dis represents the position in the danger field, sv dis represents the distance perceived by the individual i between itself and the emergency event source H a U represents the duration of the emergency event source.
[0103] As an improvement measure, the emotional infection method based on parameterized reinforcement learning of the embodiment further includes: constructing an individual emotional infection model after being affected by socialization; wherein the construction process of the individual emotional infection model includes steps b1-b5 as follows:
[0104] Step b1, using the OCEAN model as a tool for measuring individual personality; wherein in the OCEAN model, Φ=(ψ 0 ,ψ C ,ψ E ,ψ A ,ψ N ); Φ is the individual personality, and the personality includes five mutually orthogonal factors: openness ψ 0 , conscientiousness ψ C , extraversion ψ E , agreeableness ψ A and neuroticism ψ N ; wherein the value of each factor is limited between [-1, 1], and the greater the negative value represents the more negative influence, and vice versa. For example, (0.236, -0.267, -0.401, 0.452, 0.98) represents that the neuroticism factor of the corresponding individual is relatively strong;
[0105] Step b2, constructing an individual initial empathy model of the individual affected by its own personality; wherein:
[0106] ε i =(0.35ψ 0 +0.177ψ C +0.135ψ E +0.312ψ A +0.021ψ N +1) / 2;
[0107] wherein ε i is the individual initial empathy level of the individual i;
[0108] Specifically, in this embodiment, the individual initial empathy level here is calculated by the OCEAN model (measuring five traits of individuals, openness, conscientiousness, extraversion, agreeableness and neuroticism). The individual initial empathy level as a basic parameter is used to calculate the social identity degree in the subsequent step b3 (which combines the individual's initial empathy level and the individual's sense of belonging to the group). In addition, the individual initial empathy level is also used to build the empathy breakdown model in the subsequent step b4, and to calculate the individual's empathy critical threshold in the emergency scene combined with the panic emotion and the neurotic factor, thereby affecting the individual's revised empathy level.
[0109] Finally, this revised empathy level will affect the calculation of the individual emotional contagion intensity and behavior weight in step 5, and determine the individual's dynamic response in the evacuation behavior. Therefore, the initial empathy level in this step b2 is the core starting point of the entire emotional contagion model, which affects the individual's perception of the environment and the neighbor's emotion and behavior decision.
[0110] Step b3, according to the social identity strength of the individual to each group, the social identity degree of each individual is calculated respectively; wherein:
[0111]
[0112] ε i,C = ε i · χ ij ;
[0113] Wherein, g i,social_identity represents the social identity degree of individual i, ε i,C represents the empathy level of individual i to the extreme social identity member; ε i,U represents the transcendental empathy level of individual i to the member without obvious relationship or social identity relationship, which refers to the empathy ability of individual to those members without obvious relationship or without obvious social identity relationship; represents the exponential adjustment value of the difference between individual identity and collective identity of individual i; ε i,C > ε i,U , indicating that the influence of collective identity on the emotion of individual i increases; ε i,C < ε i,U , indicating that the influence of collective identity on the emotion of individual i decreases;
[0114] In this embodiment, the empathy level of individual i to the extreme social identity member ε i,C and the transcendental empathy level of individual i to the social identity relationship member ε i,Uis calculated by adjusting the initial empathy level of the individual. Specifically, the initial empathy level is taken as the basis, and the influence of social identity strength and social identity relationship is combined to adjust:
[0115] The empathy level of extreme social identity members is calculated according to the initial empathy level of the individual, multiplied by an adjustment factor reflecting the social identity strength; the adjustment factor mentioned here is the attention weight of the individual to each group.
[0116] The transcendental empathy level is calculated by combining the initial empathy level with an adjustment factor based on the social identity relationship; the adjustment factor here is the social identity strength. This way ensures that the empathy level of extreme social identity members and the transcendental empathy level of social identity relationship members are both related to the initial empathy level of the individual, while reflecting the dynamic role of social identity.
[0117] Step b4, constructing an empathy collapse model of individual empathy in an emergency situation affected by emotions; wherein:
[0118]
[0119] wherein, ε I represents the critical threshold of the individual's empathy in extreme irrational situations, and ε I is taken from a Gaussian distribution with a mean of 0.7-0.4ε and a standard deviation of (0.7-0.4ε) / 10; panT represents the critical threshold (or performance threshold) of the individual's own influence on empathy due to panic emotions, and the critical threshold panT is negatively correlated with the individual's own neuroticism factor, and the critical threshold panT is taken from a Gaussian distribution with a mean of 0.5-0.5ψ N and a standard deviation of (0.5-0.5ψ N ) / 10; variable E represents the emotional intensity of the individual, variable ε represents the empathy level of the individual, and variable E and variable ε are positively correlated; in this embodiment, the greater the variable ε, the stronger the individual's empathy level, and emotions will also affect the individual's empathy level;
[0120] Empathy collapse refers to the phenomenon that the individual's empathy level significantly decreases in extreme or high-pressure situations. In emergency situations, disasters or high-pressure environments, the individual's own emotional response and survival instinct override the ability to pay attention to and respond to the emotions of others;
[0121] In this embodiment, the empathy critical threshold ε IAs the boundary between susceptible state and infected state, it is negatively correlated with the individual's initial empathy level. That is to say, the stronger the individual's initial empathy level, the more likely he / she is to be affected by the emotional state of the neighbor;
[0122] The performance threshold panT is the boundary of the individual's ability to spread emotions, and only the infected individual whose emotional state exceeds the performance threshold panT has the ability to spread emotions to the outside. The performance threshold panT is related to the individual's extroversion; the neighbor individual must be within the individual's perception range and be an infected individual who reaches the performance threshold;
[0123] Step b5, according to the individual's personality, social identity and empathy collapse model, the individual's modified empathy level is obtained; wherein:
[0124] ε i,modify = f (OCEAN, E) * g i,social_identity ;
[0125] Wherein, ε i,modify represents the modified empathy level of individual i.
[0126] In order to calculate the emotional intensity of the individual after being affected by the emotions of all neighbors, the emotion infection method based on parameterized reinforcement learning of the embodiment further comprises steps c1-c4:
[0127] Step c1, an emotional infection intensity model of all neighbor individuals around the individual to the individual is constructed to represent the emotional influence of the individual by the emotions of all neighbor individuals around the individual; wherein:
[0128]
[0129] Wherein, is the emotional infection intensity of all neighbor individuals around individual i to individual i, is the connection strength of neighbor individual j, E j is the emotional intensity value of neighbor individual j, γ ij is the connection strength between individual i and neighbor individual j, R i is the connection strength sum value between individual i and all neighbors within the individual's perception range, and N is the total number of all neighbor individuals within the individual's perception range.
[0130] Specifically, the connection strength γ ij between individual i and neighbor individual j is calculated as follows:
[0131]
[0132] Wherein, θ j is the potential attention of individual i to neighbor individual j within the individual's perception range, is the basic attention degree (or called emotional contagion weight) of individual i to its neighbor individual j within its own perception range, and a' is the basic attention coefficient between individual i and neighbor j, is the social distance between individual i and neighbor individual j; N i is the total number of neighbor individuals within the perception range of individual i, and m is the number of a neighbor individual among all neighbor individuals of individual i;
[0133] E j is the emotion value of neighbor individual j, A j is taken from a Gaussian distribution with a mean of 0.6-0.6ψ N and a standard deviation of (0.6-0.6ψ N ) / 10; μ i is the emotion preference of individual i;
[0134] If μ i <0.5, it indicates that individual i is more susceptible to negative emotions such as panic; on the contrary, if μ i >0.5, it indicates that individual i has more attention to positive emotions and is more susceptible to neighbors with strong positive emotions;
[0135] The empathy level measures the willingness and ability of an individual to receive emotions;
[0136] The emotional contagion weight represents the weighted attention share of an individual to the emotional expression of surrounding neighbors;
[0137] The social distance controls the spatial intensity of emotional transmission;
[0138] Step c2, establish the mapping relationship between the OCEAN personality model and the emotion preference of individual; wherein:
[0139] ω NP + ω NP = 1; ω NP ∈(0, 1), ω CP ∈(0, 1);
[0140]
[0141] wherein, represents the neuroticism trait factor value of individual i, represents the conscientiousness trait factor value of individual i, ω NP represents the neuroticism trait weight factor, ω CP represents the conscientiousness trait weight factor;
[0142] Step c3, calculate the emotional intensity of each individual according to the influence of emergency event source in the individual's surrounding environment and the emotional influence of surrounding neighbors on the individual; wherein the emotional intensity of individual i is marked as E i :
[0143] wherein, H a represents the emotional infection intensity of individual i, represents the emotional infection intensity of individual i on individual i by all neighboring individuals around individual i;
[0144] Step c4, calculate the real-time emotional intensity of the individual based on the obtained emotional intensity of the individual; wherein the real-time emotional intensity of individual i at time t is marked as
[0145] wherein, is the real-time emotional intensity of individual i at time t-1, and β is a decay proportion factor and is proportional to the neuroticism feature. Over time, the emotional intensity of the individual gradually decays and eventually recovers to the baseline state.
[0146] In order to evaluate the accuracy of the constructed emotional infection model of the individual affected by the influence of the emergency event source in the individual's surrounding environment and the emotional influence of the surrounding neighbor individuals, the emotional infection method based on parameterized reinforcement learning of this embodiment further comprises: using a parameterized reinforcement learning method to perform behavior-driven simulation on the simulated heterogeneous crowd with diverse behaviors; wherein the parameterized reinforcement learning method comprises a training phase and a runtime simulation phase; wherein:
[0147] The training phase comprises:
[0148] Step d1, a plurality of agents are configured in an emergency scene with the same emergency event source and different environmental conditions; wherein the agent here is used to simulate an individual in a real crowd video;
[0149] In the art, it is well known to those skilled in the art that an agent (Agent) refers to an agent that can perceive the environment and take actions to achieve a specific goal, and it has autonomy, adaptability and interaction ability; the agent perceives the changes in the environment (such as through sensors or data input), makes judgments and decisions according to the knowledge and algorithms learned by itself, and then performs actions to affect the environment or achieve the predetermined goal; the core of the agent is autonomous learning and continuous evolution to better complete the task and adapt to complex environment.
[0150] The agent mainly has autonomy, reactivity, initiative, sociality and evolution; wherein:
[0151] Autonomy is one of the most fundamental characteristics of an agent, referring to the ability of an agent to perceive the environment, make decisions, and perform actions independently without continuous human intervention or guidance. Autonomy enables agents to work independently in dynamic and unpredictable environments, adapting to changes and adjusting their behavior. For example, an autonomous vehicle is a highly autonomous agent that can perceive surrounding vehicles and pedestrians in a complex traffic environment, autonomously plan a path, control speed, and make obstacle avoidance decisions.
[0152] Reactivity refers to the ability of an agent to quickly perceive changes in the environment and respond in a timely manner. This characteristic enables agents to react quickly and effectively in the face of unexpected events or emergencies. Reactivity is crucial for agents in real-time systems and dynamic environments, such as in robot control, where an agent needs to perceive the presence of obstacles immediately and adjust its path to avoid collisions. While reactivity typically implies an immediate response to the current state, advanced agents can also incorporate historical data and predictive information, making their reactions more intelligent and flexible.
[0153] Proactiveness is the ability of an agent to set goals, plan actions, and take measures to achieve these goals proactively, rather than just reacting to changes in the environment. Proactiveness enables agents to go beyond passive responses to external stimuli and actively pursue their internal goals and motivations. For example, a smart home system can proactively learn a user's daily habits and adjust indoor temperature or lighting in advance to improve user comfort. Agents with proactiveness can autonomously explore the environment, discover problems, and propose solutions, demonstrating greater flexibility and creativity in achieving long-term goals.
[0154] Sociality refers to the ability of an agent to interact, collaborate, and communicate with other agents or humans. Agents with sociality can understand and follow social norms, coordinate actions with other individuals to accomplish complex tasks. For example, in a multi-agent system, individual agents need to share information, allocate tasks, and achieve team goals through collaboration by using communication protocols. Sociality is also reflected in human-machine interactions, such as intelligent voice assistants that can understand user instructions and provide feedback and suggestions through dialogue. By enhancing sociality, agents can demonstrate higher efficiency and effectiveness in team work, group decision-making, and collaborative environments.
[0155] Evolutionary is the ability of an agent to improve its capabilities over time through learning and adaptation. An agent with evolutionary characteristics can gradually improve its performance by adjusting and optimizing itself when facing new environments or tasks. This characteristic is often combined with machine learning, evolutionary algorithms or reinforcement learning, enabling the agent to remain competitive in a changing environment. For example, a reinforcement learning agent continuously adjusts its strategy to maximize long-term rewards through continuous interaction with the environment. Evolutionary enables the agent to cope with uncertainty and complexity, making it perform well in long-term tasks or unknown environments, and becoming more intelligent and efficient over time.
[0156] Step d2, all agents share the same reward function configuration;
[0157] Step d3, based on the panic emotion value caused by the emergency source, the four types of emotion states of the agent are discretely generated;
[0158] Step d4, the relative importance of different behavior reward signals of the agent is dynamically adjusted;
[0159] The runtime simulation stage includes:
[0160] According to the calculated current emotion classification of the agent, the behavior corresponding to the emotion mapping is configured for the evacuation agent in real time, so that the agent can dynamically respond to the changing emotional state during the simulation process.
[0161] As for the case of the training stage, in this embodiment, the training strategy adopted by the training stage is set as follows:
[0162] Step e1, obtain the observable state of each agent observed by itself around other agents; wherein, the observable state includes the position of the agent, the speed of the agent and the radius of the agent;
[0163] Step e2, each agent evaluates the state gap between the current state and the desired state, and adjusts its own movement according to its real-time evaluation of the environment and the behavior of other agents; wherein, the state of the agent at the current time t is marked as The observed state of the neighbor agent j at the current time t is marked as The combined input state of the agent is represented as
[0164] Step e3, the preset strategy control mode maps the current state of the agent to its action behavior to obtain the maximum expected return; wherein, the preset strategy control mode is as follows:
[0165]
[0166] Wherein, R(st a t ) is a reward obtained by the agent at time t, a t is an action taken by the agent at time t, p is a discount factor, p △t is a discount factor in a time period At, v pref is a normalization term in the discount factor p, P(s t a t s t +△t ) is a transition probability from time t to time t+At, V * is an optimal value function, which is well known to those skilled in the art.
[0167] In addition, in the training strategy used in the above training phase, the agent is in an emergency scenario and the total reward return received by the agent at time t is calculated as follows:
[0168]
[0169] wherein R t is the total reward return received by the agent at time t when the agent is in an emergency scenario;
[0170] is a behavior weight configuration value of the agent in an emergency scenario and at time t for executing a behavior of moving towards a target point; is a behavior weight configuration value of the agent in an emergency scenario and at time t for executing a behavior of avoiding collision with other agents;
[0171] is a behavior weight configuration value of the agent in an emergency scenario and at time t for executing a behavior of following or staying in a group;
[0172] R g represents that when the current position of the agent coincides with the target point or is within a minimum threshold distance of the target point, it is confirmed that this condition is met;
[0173] R gt represents that the distance between the agent and the target point in the current frame is less than that in the previous frame, indicating that the agent moves towards the target;
[0174] R gm represents that the distance between the agent and the target point in the current frame is greater than that in the previous frame, indicating that the agent moves away from the target;
[0175] R ft represents that when the orientation of the agent points to the target point and the angle between the moving direction of the agent and the straight line pointing to the target is less than 45 degrees, this condition is met;
[0176] R farepresents that this condition is met when the agent's orientation is not pointing towards the goal point, and the angle between the agent's moving direction and the straight line pointing to the goal is greater than 45 degrees;
[0177] R ca represents that this condition occurs when the agent is in contact with another agent in the simulated surrounding environment;
[0178] R cw represents that this condition occurs when the agent is in contact with the boundary wall of the simulated environment;
[0179] R co represents that the agent collides with an object in the simulated surrounding environment that should not be crossed;
[0180] R rn represents the proximity between the agent and its nearest neighbor individual;
[0181] R ao represents that the agent performs the behavior of moving forward and backward in an emergency scenario;
[0182] R po represents that the agent is considered to perform the behavior of continuous oscillation when it turns left and right more than 70 times in succession.
[0183] In this embodiment, the emotion state and behavior weight mapping relationship is shown in Table 1:
[0184] Table 1
[0185]
[0186] Although the preferred embodiments of the present application have been described in detail above, it should be apparent that various modifications and changes can be made by those skilled in the art without departing from the spirit and principles of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the scope of the present application.
Claims
1. An emotion contagion method based on parameterized reinforcement learning, characterized in that, The method comprises the following steps: Step 1, obtaining a real crowd video with social identity attributes; Step 2, performing identification processing on the obtained real crowd video to obtain all groups constituting the real crowd and the social identity of each individual; Step 3, calculating the emotional intensity of each individual based on the influence of the emergency event source in the surrounding environment of each individual and the emotional influence of the surrounding neighbor individuals on each individual; Step 4, calculating the behavior weight of each individual corresponding to the behavior under different emotional intensities based on the obtained emotional intensity of each individual; wherein the emotional intensity of the individual corresponds to the behavior of the individual one by one; Step 5, obtaining the evacuation behavior of each individual under different emotional intensities and environmental conditions based on the same reward function and the behavior weight of all individuals.
2. The emotion contagion method based on parameterized reinforcement learning according to claim 1, characterized in that, In step 2, the real crowd video is identified by using a Gaussian mixture model to obtain all groups of the real crowd and the social identity of each individual; wherein the identification processing process comprises the following steps: Step a1, collecting the historical trajectory information of each individual in the real crowd video; wherein the historical trajectory information of the individual is the movement information of the individual in the real crowd video within the time period corresponding to the real crowd video, and the movement information includes the motion path, position coordinates and movement direction; Step a2, performing normalization processing on the collected historical trajectory information of each individual to obtain the normalized historical trajectory information of each individual; Step a3, pre-calculating the weight of each group, and evaluating the optimal group number of the real crowd in theory and the number of individuals in each group based on the contour coefficient; wherein the weight value of the group is the mixing coefficient of each group in the Gaussian mixture model; Step a4, performing spatiotemporal consistency clustering processing on the crowd in the real crowd video by using the Gaussian mixture model and the optimal group number obtained by evaluation to obtain a plurality of groups composed of different individuals; wherein the group and the cluster center are in one-to-one correspondence; Step a5, calculating the attention weight of each individual to each group by using the Mahalanobis distance between each individual and each cluster center and the direction similarity between the movement direction of each individual and the direction from the individual position to the cluster center; χ ij = a exp(MDIS ij ) + (1 - a) Sim ij ; a = 0.5; Sim ij ∈[0,1] where χ ij is the attention weight of individual i to the jth group, a is the relative importance index for controlling the Mahalanobis distance and the direction similarity, MDIS ij is the Mahalanobis distance of individual i to the cluster center j of the jth group, Sim ij represents the direction similarity between the moving direction of individual i and the direction from the position of individual i to the cluster center j; x i represents the position of individual i, σ j represents the position of the cluster center j, σ j -x i represents the vector from the position of individual x i to the position of the cluster center σ j ; v i represents the moving velocity vector of individual i; Sim ij = 1 indicates that the moving direction of individual i is completely directed to the cluster center j, Sim ij = 0 indicates that the moving direction of individual i is completely deviated from the cluster center j. Step a6, calculating the social identity strength of each individual to each group; wherein: wherein, is the strength of social identification of individual i to the jth group, κ ij is the degree of belongingness individual i feels to the jth group.
3. The emotion contagion method based on parameterized reinforcement learning according to claim 2, characterized in that, The influence of the individual on the surrounding environment of the emergency event source is the panic emotional change caused by the individual; wherein: wherein, represents an emergency event source H a an emotional infection intensity of a panic emotion change caused by an emotion of an individual i, P represents an emergency event source H a an arbitrary position in a danger field formed by the emergency event source H, represents an emergency event source H a a position in a danger field, dis sv represents a distance perceived by an individual i from the individual i to an emergency event source H a , U represents a duration of the emergency event source.
4. The emotion contagion method based on parameterized reinforcement learning according to claim 3, characterized in that, Further comprising: Constructing an individual emotional infection model affected by socialization; wherein the construction process of the individual emotional infection model comprises the following steps: Step b1, taking the OCEAN model as a measuring tool for measuring individual personality; wherein, in the OCEAN model, Φ=(ψ 0 ,ψ C ,ψ E ,ψ A ,ψ N ); Φ is individual personality, and the personality includes five mutually orthogonal factors of openness ψ 0 , conscientiousness ψ C , extraversion ψ E , agreeableness ψ A and neuroticism ψ N ; Step b2, constructing an individual initial empathy model affected by the individual personality; wherein: ε i = (0.35ψ 0 + 0.177ψ C + 0.135ψ E + 0.312ψ A + 0.021ψ N + 1) / 2; where ε i is the individual initial level of empathy for individual i; Step b3, calculating the social identity degree of each individual according to the obtained social identity strength of each individual to each group; wherein: where g i,social_identity represents the social identity of individual i, ε i,C represents the level of empathy of individual i to the members of extreme social identity; ε i,U represents the level of transcendental empathy of individual i to the members without obvious relationship or social identity relationship; represents the exponential adjustment value of the difference between individual identity and collective identity of individual i; Step b4, constructing an empathy breakdown model of the individual empathy affected by the emotion in the emergency scene; wherein: wherein ε I represents the critical threshold of empathy of the individual in the extreme irrational case, ε I is taken from a Gaussian distribution with a mean of 0.7-0.4ε and a standard deviation of (0.7-0.4ε) / 10; panT represents the critical threshold of the influence of the individual's own panic emotion on empathy, and the critical threshold panT is negatively correlated with the neuroticism factor of the individual, and the critical threshold panT is taken from a Gaussian distribution with a mean of 0.5-0.5ψ N and a standard deviation of (0.5-0.5ψ N ) / 10; the variable E represents the emotional intensity of the individual, the variable ε represents the empathy level of the individual, and the variable E and the variable ε are positively correlated; Step b5, obtaining the corrected empathy level of the individual according to the individual personality, the social identity degree and the empathy breakdown model of the individual; wherein: ε i,modify = f(OCEAN, E) * g i,social_identity ; where ε i,modify denotes the revised level of empathy of individual i.
5. The emotion contagion method based on parameterized reinforcement learning according to claim 4, characterized in that, Further comprising: A parameterized reinforcement learning method is used to drive the behavior of a simulated heterogeneous crowd with diverse behaviors; wherein the parameterized reinforcement learning method comprises a training phase and a runtime simulation phase; wherein: The training phase comprises: configuring a plurality of agents in an emergency scenario with the same emergency source and different environmental conditions; making all agents share the same reward function configuration; discretely generating four types of emotional states of each agent based on the panic emotion value caused by the emergency source; dynamically adjusting the relative importance of different behavior reward signals of the agent; The runtime simulation phase comprises: According to the calculated current emotional classification of the agent, real-time configuration of the evacuation agent with the corresponding behavior of the emotional mapping enables the agent to dynamically respond to the changing emotional state during the simulation process.
6. The emotion contagion method based on parameterized reinforcement learning according to claim 5, characterized in that, The training strategy used in the training phase is set as follows: Obtain the observable state of each agent observed by the agent around the agent; wherein the observable state includes the position of the agent, the speed of the agent and the radius of the agent; Let each agent evaluate the state gap between the current state and the desired state, and adjust its own movement according to its real-time evaluation of the environment and the behaviors of other agents; wherein the state of the agent at the current time t is denoted as The observed state of the neighbor agent j at the current time t is denoted as The combined input state of the agent is denoted as Map the current state of the agent to its action behavior in a preset strategy control manner to obtain the maximum expected return; wherein the preset strategy control manner is as follows: where R(s t ,a t ) is the reward obtained by the agent at time t, a t is the action taken by the agent at time t, p is the discount factor, p △t is the discount factor over a time period At, v pref is a normalization term in the discount factor p, P(s t ,a t ,s t+△t ) is the transition probability from time t to time t + At, and V * is the optimal value function.
Citation Information
Cited By
Digital customer service information generation method based on multi-modal information, medium and equipment
CN121561823A
Digital customer service information generation method, medium and device based on multi-modal information
CN121561823B