A sequential decision strategy collection method and system for high neurotic people

CN121117495BActive Publication Date: 2026-09-08AIR FORCE MEDICAL CENT PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511301836.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-09-08
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

[0004]为了克服现有技术的不足,本发明的目的是提供一种面向高神经质人群的序列决策策略采集方法及系统,通过在马尔可夫决策环境中引入情绪刺激与奖励耦合机制,结合探索与执行双阶段的高精度行为采集、拉丁方平衡顺序控制以及二进制/文本双格式归档,不仅克服了传统方法生态效度不足、数据采集粗糙和顺序效应偏差的缺陷,而且实现了对高神经质人群序列决策特征的真实复现与可重复研究

Benefits of technology

第一,本发明通过在网格化决策地图(矩阵型/中央辐射型)上耦合情绪刺激与奖励来构建决策环境,克服了传统Mouselab-MDP仅呈现金钱数额与空白卡片、生态效度不足的缺陷;系统支持悲伤/愤怒/高兴/中性四类情绪及可调情绪效价比率与概率设置,并以随机种子与加权抽样确保刺激呈现的可控性与一致性,从而更贴近高神经质人群在负性刺激下的真实决策情境。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121117495B_ABST
    Figure CN121117495B_ABST
Patent Text Reader

Abstract

The application provides a sequence decision strategy collection method and system for high neurotic people. The application establishes a full-screen drawing canvas and loads a task set containing a decision map type, an emotional valence ratio, a reward parameter, a random seed and a sequence control parameter; collects user gender selection and binds a corresponding emotional picture library; generates a decision map based on the Markov decision principle under the control of the parameter set, defines nodes and edges, randomly schedules and lays out emotional pictures according to the valence ratio and assigns rewards, and forms an environment coupling emotions and rewards; based on the environment, record the mouse position and path in the exploration stage, record the key time, action direction and reward in the execution stage, and establish a step sequence association with a timestamp to form sequence decision data; the task sequence is balanced using a Latin square, and the process is promoted under user triggering; finally, the data is archived in binary and text files for subsequent analysis. The application improves the data authenticity and repeatability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of psychological and behavioral science, and in particular to a method and system for collecting sequence decision-making strategies for highly neurotic individuals. Background Technology

[0002] Systematic bias, or systematic error, refers to a directional and repeatable deviation caused by non-random factors during measurement or experimentation. Unlike random error, systematic bias continuously affects the accuracy of results and cannot be eliminated by increasing sample size. Efforts to reduce the impact of systematic error in psychology began at the inception of experimental psychology. With the development of technology, the focus has gradually expanded from behavioral factors such as experimental design, experimental environment, and experimenter operation to emphasizing the control of systemic error in advanced research methods such as brain imaging and computational modeling. It can be said that minimizing the impact of systematic error on experimental results in the process of revealing scientific principles has always been the most fundamental, and even the core, pursuit of psychology researchers.

[0003] Human decision-making is not purely rational or perfect. The existence of cognitive biases in decision-making is widely acknowledged (Gilovich et al., 2002; Tversky & Kahneman, 1974). Kahneman's core insight is that "humans are boundedly rational, but not randomly prone to error; their biases are patterned and predictable." Neuroticism, a core dimension in the Five Factor Personality Model, is essentially a high sensitivity to negative emotions and potential threats, exhibiting stability across cultures, time periods, and research areas. Individuals with high neuroticism demonstrate significant systematic biases in risky situations, primarily exhibiting a harm-avoidance tendency. Their decision-making patterns are often accompanied by excessive risk aversion and emotion-driven cognitive biases. While neuroticism is not a disease, individuals with high neuroticism often experience more stress and negative emotions in decision-making scenarios that others consider normal, leading to worse choices and increased inconvenience in their lives. Therefore, elucidating the mechanisms underlying the biases in the decision-making process of neurotic individuals can help them adjust their decision-making behavior more effectively. However, current research largely focuses on the static representation of behavioral outcomes, lacking a systematic exploration of the dynamic characteristics of the decision-making process. Specifically, existing paradigm designs generally overlook a key dimension: the dynamics of cognitive processing over time, such as the contribution mechanisms of attention, value assessment, and action execution to the overall risk decision-making process. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for collecting sequence decision-making strategies for highly neurotic individuals. By introducing an emotional stimulus and reward coupling mechanism into a Markov decision environment, combined with high-precision behavioral collection in two stages of exploration and execution, Latin square balanced sequence control, and binary / text dual-format archiving, this invention not only overcomes the shortcomings of traditional methods such as insufficient ecological validity, coarse data collection, and sequence effect bias, but also achieves realistic reproduction and reproducible research on the sequence decision-making characteristics of highly neurotic individuals.

[0005] To achieve the above objectives, the present invention provides the following solution: A method for collecting sequence decision-making strategies for highly neurotic individuals, comprising: Create a full-screen drawing canvas and load a task parameter set; the task parameter set includes at least the decision map type and depth, emotion valence ratio, reward parameter and random seed, and presentation order control parameter. Collect user gender selection data and determine the binding relationship of the emotional image library for subsequent presentation, so that the stimulus set corresponds to user attributes; Under the control of the task parameter set, a gridded decision map is generated based on the Markov decision principle, nodes and actionable edges are defined, and emotional images are weighted and randomly scheduled according to the emotional valence ratio. At the same time, rewards are assigned at node or path positions according to the reward parameters and random seeds, forming a decision environment that couples stimuli and rewards. Based on the aforementioned decision-making environment, an exploration phase is first conducted to record mouse position, exploration time, and movement path. Then, an execution phase is conducted to record key press time, possible action directions, and rewards obtained. The exploration phase and the execution phase are linked in sequence to form sequential decision-making behavior data. The task presentation order is controlled based on Latin square balance, and the process is advanced upon user triggering. The sequence decision-making behavior data is archived in the form of binary files and text files for subsequent strategy reconstruction and statistical analysis.

[0006] Preferably, a full-screen drawing canvas is created and a task parameter set is loaded, including: Call the Screen function in the Psychtoolbox package to set the computer screen as a full-screen drawing canvas and initialize the display resolution parameters to achieve multi-resolution compatibility; The task prompt text is drawn in the center on the full-screen drawing canvas, and users can customize the font, font size, color and background highlight parameters to ensure the correct presentation of the task prompt information; When loading the task parameter set, the decision map type and depth, emotional valence ratio, reward parameters, random seed and presentation order control parameters are uniformly preset and bound to the experimental process after the canvas is initialized. A missing value detection and display environment recovery mechanism is implemented during the canvas drawing process to ensure that the system default display settings can be restored in the event of an abnormal exit.

[0007] Preferably, the user's gender selection is collected, and the binding relationship of the emotional image library used for subsequent presentation is determined accordingly, so that the stimulus set corresponds to the user attributes, including: Display the optional gender text entry in the center of the full-screen drawing canvas; The system receives the user's gender selection via mouse clicks and collects the gender text from the computer. The system uses a conditional statement to access an emotion image library corresponding to the collected gender. The emotional image library is linked to subsequent task flows to ensure that the stimulus set corresponds to user attributes.

[0008] Preferably, under the control of the task parameter set, a gridded decision map is generated based on the Markov decision principle, nodes and actionable edges are defined, and emotional images are weighted and randomly scheduled according to the emotional valence ratio. Simultaneously, rewards are assigned at node or path positions based on the reward parameters and random seeds, forming a decision environment that couples stimuli and rewards, including: Call the map type and depth parameters in the task parameter set, generate a matrix map or a central radial four-arm map based on the Markov decision principle, and define the coordinates, actionable edges and target nodes of each node. Based on the emotional valence ratio in the task parameter set, a cumulative probability distribution vector is constructed, and weighted random sampling is performed on four types of images: sad, angry, happy, and neutral. The sampling results are then distributed to the node positions. Based on the reward parameters and random seed in the task parameter set, a reward is assigned to the node or path position. Specifically, the expected reward is calculated using the following value function and probability weight function:

[0009] in, The expected reward for a node or path location; This refers to the actual reward value displayed at the node or path location; To obtain reward values The objective probability; The value function is constructed based on prospect theory; This is a probability weighting function based on nonlinear probability sensing; The concavity parameter of the revenue curve, The convexity parameter of the loss curve, The loss aversion coefficient, The curvature parameter of the probability weighting function; By combining the results of emotion setting and reward distribution, a complete decision-making environment that couples emotional stimuli and reward information is generated and linked to the exploration and execution phases.

[0010] Preferably, the value function and probability weight function in the formula are defined as follows: ; in, This represents the actual reward value for a single node or path location. To obtain reward values The objective probability; the function is used to calculate the expected reward. This completes the assignment of reward values ​​to node or path locations.

[0011] Preferably, based on the decision-making environment, an exploration phase is first performed to record mouse position, exploration time, and movement path; then an execution phase is performed to record key press times, possible action directions, and rewards obtained; and the exploration and execution phases are linked in sequence to form sequential decision-making behavior data, including: During the exploration phase, the mouse position is read cyclically at a sampling frequency of no less than 60Hz, and a timestamp and event marker are recorded for each frame. During the execution phase, a non-blocking key detection and debouncing mechanism is employed to record the start / end time, possible action direction, and corresponding reward value for each key press, while simultaneously maintaining a progressively updated cumulative result. Based on timestamps, the exploration samples are compared with the first... The sequence of key presses is associated with the first key press, defining the relationship with the first key press. The set of exploration samples adjacent to the time of the next key press is used to calculate the exploration intensity feature value corresponding to the current step. The formula is: ;in, For the first Average displacement intensity within the window before step execution; In order to be with the first Key press time The corresponding set of exploration sample indices; For the first The timestamp of the next key press; This is a constant representing the duration of the time window used for pre-step association; For the first The timestamp of the frame exploration sample, and Sampling frequency and For the first Frame mouse position vector; , The first The horizontal and vertical coordinates of the mouse cursor in the screen coordinate system; Represents the Euclidean norm; The exploration intensity characteristic value of each step, the associated key press time, action direction and reward value are organized into a recording unit according to the step sequence and written into a machine-readable file, which includes at least two formats: binary data file and tab-delimited text file, for subsequent strategy reconstruction and statistical analysis.

[0012] Preferably, the task presentation order is controlled based on Latin square balance, and the process is advanced upon user triggering, including: Obtain the number of experimental blocks and load or generate a Latin square matrix of the same order, and determine the presentation order framework across subjects based on the Latin square balance principle. Based on the random seed in the task parameter set, the Latin square is subjected to row and column permutations to obtain the presentation sequence corresponding to the current subject, and the presentation sequence is written into the sequence control parameter port for process scheduling; After each block is completed, the user input listening state is entered, the next step prompt message is displayed and the preset start button is listened for; Upon receiving user-triggered input, the timestamp is recorded and the process proceeds to the next block according to the presented sequence until the sequence traversal is complete, thus ending the task flow.

[0013] Preferably, the sequence decision-making behavior data is archived in the form of binary files and text files for subsequent policy reconstruction and statistical analysis, including: Based on a unified timestamp, the data of the exploration phase and the execution phase are matched and sorted step by step to generate a record unit containing timestamp, mouse position, event marker, key press time, action direction, reward value and cumulative reward; Complex data structures are managed using cell arrays, and the I / O write interface is called to save the record units as .mat binary data files; The record unit is expanded into tab-delimited text in a preset column order and written to a .txt file to achieve double archiving corresponding to the binary data; After archiving is completed, the data writing process for this block ends, and we are ready to proceed to the subsequent data processing and analysis stage.

[0014] A sequence decision-making strategy acquisition system for highly neurotic individuals, comprising: The canvas initialization and parameter loading module is used to create a full-screen drawing canvas and load the task parameter set; the task parameter set includes at least the decision map type and depth, emotion valence ratio, reward parameter and random seed, and presentation order control parameter. The gender collection and image library binding module is used to collect the user's gender selection and determine the binding relationship of the emotion image library for subsequent presentation, so that the stimulus set corresponds to the user's attributes. The decision environment construction module is used to generate a gridded decision map based on the Markov decision principle under the control of the task parameter set, define nodes and actionable edges, and perform weighted random scheduling and deployment of emotion images according to the emotion valence ratio. At the same time, it assigns rewards at node or path positions according to the reward parameters and random seeds to form a decision environment that couples stimuli and rewards. The behavior collection and step sequence association module is used to first conduct an exploration phase based on the decision environment to record the mouse position, exploration time and movement path, and then conduct an execution phase to record the key press time, possible action direction and reward obtained, and associate the exploration phase and execution phase with the step sequence to form sequential decision behavior data. The sequence control and process advancement module is used to control the task presentation order based on Latin square balance and advance the process when triggered by the user. The data archiving and export module is used to archive the sequence decision behavior data in the form of binary files and text files for subsequent strategy reconstruction and statistical analysis.

[0015] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: First, this invention constructs a decision-making environment by coupling emotional stimuli and rewards on a gridded decision map (matrix type / central radial type), overcoming the shortcomings of traditional Mouselab-MDP, which only presents monetary amounts and blank cards and lacks ecological validity. The system supports four types of emotions: sadness, anger, happiness, and neutrality, as well as adjustable emotional valence ratios and probability settings. Random seeds and weighted sampling are used to ensure the controllability and consistency of stimulus presentation, thus more closely resembling the real decision-making situation of highly neurotic individuals under negative stimuli.

[0016] Second, this invention performs unified data collection and step sequence association for the exploration-execution dual stages, providing millisecond-level time accuracy and frame-level recording of no less than 60 Hz. It not only saves process features such as mouse position, exploration time, and movement path, but also records key press time, action direction, and corresponding reward simultaneously. It is archived in both binary and text formats, which is significantly better than the coarse-grained recording method that only retains "selection-result", facilitating subsequent strategy reconstruction and statistical analysis.

[0017] Third, this invention introduces Latin square balance to eliminate chunk order bias in terms of presentation order and repeatability, and drives stimulus and reward distribution with system time or user-defined random seeds, which not only ensures comparability across subjects, but also improves the reproducibility of repeated experiments under different devices and scenarios, overcoming the shortcomings of insufficient randomization in existing experimental procedures and the interference of results by order effects.

[0018] Fourth, this invention provides an engineering and robust implementation: the full-screen drawing canvas supports multi-resolution compatibility, parameterized presentation of font / font size / color / background highlighting, and has a missing value detection and display environment recovery mechanism; at the same time, it uses cell arrays to manage complex data structures and archives them in both .mat and .txt formats, improving the stability and data interoperability of the system on different terminals, and solving the problems of poor resolution adaptation, anomaly recovery and weak data organization in existing implementations. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart of the method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the system structure provided in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] The purpose of this invention is to provide a method and system for collecting sequence decision-making strategies for highly neurotic individuals, which significantly improves the authenticity and reproducibility of decision-making data.

[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] Figure 1 A flowchart of the method provided in the embodiments of the present invention, such as Figure 1 As shown, this invention provides a method for collecting sequence decision-making strategies for highly neurotic individuals, including: Step 100: Create a full-screen drawing canvas and load the task parameter set; the task parameter set includes at least the decision map type and depth, emotion valence ratio, reward parameters and random seeds, and presentation order control parameters; Step 200: Collect user gender selection and determine the binding relationship of the emotion image library for subsequent presentation, so that the stimulus set corresponds to the user attribute; Step 300: Under the control of the task parameter set, a gridded decision map is generated based on the Markov decision principle, nodes and actionable edges are defined, and the emotion pictures are weighted and randomly scheduled according to the emotion valence ratio. At the same time, rewards are assigned at the node or path position according to the reward parameters and random seeds to form a decision environment that couples stimuli and rewards. Step 400: Based on the decision-making environment, first conduct an exploration phase to record the mouse position, exploration time, and movement path, then conduct an execution phase to record the key press time, possible action directions, and rewards obtained, and link the exploration phase and execution phase in sequence to form sequential decision-making behavior data; Step 500: Control the task presentation order based on Latin square balance and advance the process upon user triggering; Step 600: Archive the sequence decision behavior data in the form of binary files and text files for subsequent strategy reconstruction and statistical analysis.

[0025] Specifically, step 100 in this embodiment includes: This embodiment calls the Screen function in the Psychtoolbox package at the start of the experiment to create a drawing canvas in full-screen mode, reads the current display resolution and sets the canvas size accordingly, thereby achieving compatibility with displays of different resolutions and sizes. Then, task prompt text is drawn in the center of the canvas, and parameterized configurations for font, font size, color, and background highlighting are enabled to ensure correct presentation and consistent readability of the task prompt information. During the same stage of canvas initialization, this embodiment uniformly loads a "task parameter set," which includes: the decision map type and depth pre-set by the administrator, the emotional valence ratio input by the subject, reward parameters and random seeds generated by the system based on user definition or system time, and presentation order control parameters pre-generated based on the Latin square balance principle. After loading, this parameter set is bound to the entry point of subsequent processes (map generation, stimulus scheduling, reward assignment, and order control). To ensure robustness, this embodiment executes a missing value detection and display environment recovery mechanism during canvas creation and parameter loading. When a key parameter is detected to be empty or an initialization anomaly is detected, an anomaly recovery process is triggered and the system's default display settings are restored to avoid affecting the operating system's display environment.

[0026] The full-screen drawing canvas refers to the drawing window that covers the entire screen and is opened on the current main display via the Screen function. It serves as the sole interface for task presentation and interaction and adapts to the actual resolution returned by the operating system to ensure consistent presentation of graphics and text across different display devices. The task parameter set refers to the set of parameters loaded once during the task initialization phase and used throughout the entire process. It includes at least five types of parameters: decision map type and depth (set by the administrator), emotional valence ratio (input by the subject), reward parameters and random seed (generated by the system based on user definition or system time), and presentation order control parameters (pre-generated by Latin square balancing). This set is bound to modules such as map generation, weighted scheduling of emotional images, reward assignment, and order control to ensure consistent parameter sources and unified scheduling.

[0027] In this embodiment, the following key quantities are determined and recorded during the task initialization phase: there are two map types (matrix type and central radial four-arm type), four emotion categories (sad, angry, happy, neutral), the random seed source is millisecond-level system time or user-specified constant, and the minimum sampling frequency in the subsequent behavior collection phase is not less than 60 Hz; the above values ​​are written into the parameter set as running constraints after initialization and used throughout the task.

[0028] Optionally, step 200 in this embodiment includes: In this embodiment, after task initialization, optional gender text entries such as "Male" and "Female" are displayed in the center of the full-screen drawing canvas. Subjects select their gender by clicking the mouse. The computer collects this selection in the background and uses the collected gender text as a judgment condition for branching operations. Specifically, the computer executes conditional judgment statements to identify the user's selected gender and calls the corresponding emotional image library to ensure that the emotional stimuli presented in subsequent task flows are consistent with the user's attributes. This binding relationship is established at the beginning of the task flow and is maintained throughout the exploration and execution phases to ensure the immersion and consistency of the experiment.

[0029] The emotion image library refers to a pre-built image set based on different genders, containing filtered and labeled facial images categorized into four emotions: sadness, anger, happiness, and neutral. During task execution, the system will call the corresponding image set based on the user's gender to ensure that the presented stimulus matches the user's attributes. This embodiment provides at least two types of entry text ("male" and "female") during the gender selection phase and completes condition judgment and library call within 50 milliseconds after collecting the click event. The emotion image library has no fewer than 50 images in each category to ensure the controllability of random scheduling and repeated experiments.

[0030] Further, step 300 of this embodiment includes: This embodiment first reads the map type and depth parameters from the task parameter set, and generates a gridded decision map on the full-screen drawing canvas based on the Markov decision principle. The map can be a matrix type or a central radial four-arm structure. During the generation process, screen coordinates, actionable edges and corresponding target nodes are determined for each grid node, and the node images are ensured to be distributed at the same scale across the full screen and adapted to different resolutions, thereby forming an executable decision space composed of "state, action and transition" elements.

[0031] Subsequently, this embodiment constructs a cumulative probability distribution vector based on the emotional valence ratio in the task parameter set, performs weighted random sampling on four types of images: sad, angry, happy, and neutral, and distributes the sampling results on the aforementioned grid nodes according to the node coordinates. To ensure the temporal stability and reproducibility of the presentation, this embodiment uses a random seed generated by the system time or a preset constant for initialization during the scheduling phase, and caches the texture of the images to be presented in the background, thereby reducing switching latency and maintaining presentation accuracy.

[0032] After completing the emotion setting, this embodiment assigns reward values ​​to node or path locations based on reward parameters and random seeds in the task parameter set. Specifically, for each possible reward value, its psychological value is first obtained through a value function, and then its probability of occurrence is non-linearly weighted through a probability weight function. The weighted results are then summed to obtain the expected reward for that node or path location. The value function adopts a segmented form with concave gain segments, convex loss segments, and loss aversion characteristics. The probability weight function is controlled by curvature parameters to control the degree of subjective distortion of low and high probabilities. After the calculation is completed, this embodiment jointly writes the emotion setting and reward distribution into the decision-making environment and binds them to the subsequent exploration and execution stages, so that the subjects' navigation, dwelling, and key-pressing behaviors are triggered and recorded in a unified environment.

[0033] Among them, the cumulative probability distribution vector refers to the monotonically non-decreasing probability threshold table generated according to the target emotion valence ratio, which is used to map uniform random numbers in the interval to specific emotion categories, and realize weighted sampling and node deployment according to the preset ratio; its function is to ensure that the presentation frequency of the four types of emotions in the whole graph is consistent with the ratio set in the task parameter set, and to ensure cross-experiment repeatability when the random seed is fixed.

[0034] The parameters of the aforementioned "value function" and "probability weighting function" are sourced and set as follows: the concavity parameter of the gain segment and the convexity parameter of the loss segment are both real numbers in the open interval from zero to one, used to control the curvature of the gain and loss curves respectively; the loss aversion coefficient is a real number greater than zero, used to adjust the subjective asymmetry of loss relative to gain of the same magnitude; the probability curvature parameter is a real number in the open interval from zero to one, used to control the subjective weighting of objective probability (e.g., amplifying small probability and compressing larger probability); the actual reward value comes from the high, medium and low level reward settings in the task parameter set, and the objective probability comes from the distribution and sampling mechanism of node or path positions. The above parameters are set when the task starts and remain unchanged throughout the experiment to ensure repeatability. In this embodiment, there are two map types (matrix type and central radial four-arm type), and four emotion categories (sad, angry, happy, and neutral). The three interval parameters of the value function and probability weight function are all in the range of zero to one, and the loss aversion coefficient is greater than zero. The random seed is generated and fixed by the system time or a preset constant when the task starts. In order to cooperate with the unified time base for subsequent data collection, the sampling frequency of the exploration and execution phase is set to no less than 60.

[0035] Specifically, step 400 in this embodiment includes: In this embodiment, after entering the data acquisition phase, the input acquisition and timing module of the exploration phase is first activated. The mouse position in the screen coordinate system is read frame by frame at a frequency of no less than 60Hz, and a timestamp and event marker under a unified time base are written for each frame, thereby forming a continuous exploration sample sequence. Then, the execution phase is entered. A non-blocking key detection and debouncing mechanism is used to synchronously record the start time, end time, action direction and corresponding reward of each key press, and the cumulative reward is gradually updated during the recording process. In order to integrate the data of the two phases, this embodiment pairs each key press event with a small segment of exploration samples before it according to the timeline. The coordinate change of adjacent frames is averaged within the pairing window to obtain the exploration intensity feature value of the current step. This feature, together with the key press time, action direction and reward value of the step, constitutes a recording unit, and is finally written to a binary data file and a tab-delimited text file respectively, so as to facilitate subsequent strategy reconstruction and statistical analysis.

[0036] The exploration intensity feature value refers to the quantified result obtained by averaging the coordinate changes of adjacent frames in continuous exploration samples within a preset time window before a button press, used to characterize the exploration activity level before the execution of that step. Its function is to compress trajectory-level fine-grained information into statistical features that correspond one-to-one with the step sequence without changing the original timeline, thereby achieving unified alignment and analysis between the exploration and execution phases. In this embodiment, the sampling frequency in the exploration phase is no less than 60 frames / second, and the timestamp recording accuracy reaches the order of 1 millisecond. Each "step" archives at least 4 types of core elements (exploration intensity, button press time, action direction, and reward value). The archived files adopt two formats (binary and tab-separated text), corresponding one-to-one with the same "step," facilitating review and reconstruction.

[0037] Optionally, step 500 in this embodiment includes: In this embodiment, the number of blocks required for the experiment is first determined during the initialization of the experimental task, and a Latin square matrix of the same order as the number of blocks is loaded or generated accordingly. This matrix satisfies the characteristic that each row and column contains all blocks without repetition, which is used to balance the presentation order of blocks among different subjects, thereby avoiding the bias of results caused by order effects or learning effects.

[0038] After obtaining the basic matrix, this embodiment uses a random seed from the task parameter set to permutate the rows and columns of the matrix, resulting in a unique presentation sequence for the current subject. This sequence is then written into the sequence control parameter port for the process scheduling module to use. In this way, the balance of the sequence across subjects is ensured, while maintaining the randomization and repeatability of the task process within the same subject.

[0039] After each experimental block is completed, this embodiment enters a user input listening state. The interface displays a prompt message and waits for the subject to trigger the preset start button. When user input is detected, the timestamp of the trigger is immediately recorded, and the presentation sequence stored in the sequential control parameter port is followed to advance to the next block until all blocks have been traversed, thereby realizing full sequential control based on Latin square balance and user-driven process advancement.

[0040] Among them, the Latin square matrix refers to the arrangement structure in which the order of the experimental blocks is the same as the number of blocks, and each symbol appears only once in each row and column. It is used to balance the presentation order of blocks among subjects and reduce the order effect and learning effect. The sequence control parameter port refers to the parameterized interface (key value) used in the program to access the current subject presentation sequence. The process scheduler reads the next block identifier based on this to drive the task forward.

[0041] In this embodiment, the order of the Latin square is consistent with the number of experimental blocks: when the number of blocks is 2 / 3 / 4, a Latin square of order 2 / 3 / 4 is used respectively; each subject corresponds to 1 row or 1 column in the Latin square, and the length of the presented sequence is equal to the number of blocks; the row and column permutations are determined by a single random seed, which is set at the start of the task and remains unchanged; each time the user triggers the advancement, a timestamp is recorded, with a time resolution of 1 millisecond.

[0042] Further, step 600 of this embodiment includes: Before data archiving, the data from the exploration and execution phases are first matched and sorted according to a unified timestamp: a one-to-one correspondence is established between the samples generated frame by frame in the exploration phase (including mouse position and event markers) and each key press event in the execution phase (including start / end time, action direction and corresponding reward) according to the timeline, and a record unit is generated accordingly. The record unit contains at least the fields of timestamp, mouse position, event marker, key press time, action direction, reward value and cumulative reward, which are used as the basic entries for subsequent storage and retrieval.

[0043] This embodiment uses a cell array to organize and manage the above record units to adapt to complex data structures containing multiple types of fields. After organization, the I / O write interface is called to write the record units in batches to a binary data file in .mat format to fully preserve the internal representation of timestamps and structured fields, facilitating efficient loading and calculation processing.

[0044] To ensure data interoperability and verifiability, this embodiment expands the same batch of record units into tab-delimited plain text in a preset column order and writes them to a .txt file while completing the binary write, thus achieving double archiving that corresponds one-to-one with the binary data; once the archiving of this block is completed, the current write transaction is terminated and the subsequent data processing and statistical analysis stage begins.

[0045] In this embodiment, a recording unit refers to the smallest data entry that integrates information from the exploration and execution phases on a unified timeline, using "single step" or "single key press" as the granularity. It includes at least the following fields: timestamp, mouse position, event marker, key press time, possible action direction, reward value, and cumulative reward, used to support subsequent retrieval, reconstruction, and statistical analysis. The sampling frequency in this embodiment is no less than 60 frames per second; the timestamp recording accuracy is on the order of 1 millisecond; each recording unit contains at least 7 core fields (timestamp, mouse position, event marker, key press time, possible action direction, reward value, and cumulative reward); the archived files are fixed in two formats (.mat and .txt), maintaining consistency in the number and order of recording units between the two formats.

[0046] As an optional implementation, this embodiment first calls the Screen function in the Psychtoolbox package to draw a text entry for gender selection in the center of the screen. It supports custom fonts, sizes, colors, and background highlights to improve interactivity. A missing value detection and display recovery mechanism is also included to ensure robustness. Next, two types of topological models are generated based on the map depth parameters set by the administrator: a matrix map and a centrally radiating four-arm map. In the map, each node is defined with an emotion image identifier, node coordinates, upstream and downstream action directions, target node coordinates, and a reaction time recording entry. Path edges are formed using connecting lines and directional arrows, thus creating a gridded decision-making environment adapted to different screen resolutions.

[0047] After the user selects the gender text by clicking with the mouse, this embodiment collects the selection and uses conditional statements to call the corresponding gender's emotional image library, ensuring that the presented face image matches the user's gender and enhancing immersion. Subsequently, based on the emotional valence-arousal model, sad, angry, happy, and neutral images from the image library are stored in independent arrays. A cumulative probability distribution vector is constructed by combining the valence ratio input by the user, and image retrieval is achieved through weighted random sampling, ensuring that the emotional stimulation conforms to the preset ratio. At the same time, a background caching mechanism is used to preload textures to reduce switching latency and ensure the accuracy of presentation time.

[0048] Regarding reward settings, this embodiment uses expected utility theory to parametrically model reward values: defining the concavity / convexity parameters of the gain and loss curves, the loss aversion coefficient, and the probability curvature parameter; processing the actual reward value and objective probability using a value function and a probability weight function respectively, thereby obtaining the expected utility value of the node or path location; in specific operation, calling the user-defined or system-generated random seed distribution reward levels, and binding the utility value with emotional stimuli, so that different nodes and paths in the map carry both emotional images and reward parameters, thereby achieving the coupling of stimulus and reward.

[0049] During the exploration phase, this embodiment initializes the input detection and timing module, cyclically collecting mouse position and hover time at a frequency of no less than 60Hz, recording event markers and millisecond-level timestamps. During the execution phase, a non-blocking arrow key detection and debouncing mechanism is employed to collect the time, direction, and corresponding reward value of each key press, dynamically updating the cumulative reward results. The collected data is simultaneously written to both binary and text files during storage to ensure data repeatability and multi-format compatibility.

[0050] Regarding task sequence control, this embodiment employs a modular parameter passing method to reserve sequence ports. During each experiment, task blocks are loaded based on user-defined or Latin square-balanced presentation sequences. After each block is completed, the system enters a user input listening state. When the system detects the user pressing the start button, it immediately records the timestamp and calls the next block's program, continuing until all blocks have been traversed. Through this implementation process, this embodiment can accurately collect complete behavioral data on exploration and execution under realistic emotional stimuli and provides a highly ecologically valid sequence decision-making experimental environment for highly neurotic individuals.

[0051] Corresponding to the above methods, such as Figure 2 As shown, this embodiment also provides a sequence decision-making strategy acquisition system for highly neurotic individuals, including: The canvas initialization and parameter loading module is used to create a full-screen drawing canvas and load the task parameter set; the task parameter set includes at least the decision map type and depth, emotion valence ratio, reward parameter and random seed, and presentation order control parameter. The gender collection and image library binding module is used to collect the user's gender selection and determine the binding relationship of the emotion image library for subsequent presentation, so that the stimulus set corresponds to the user's attributes. The decision environment construction module is used to generate a gridded decision map based on the Markov decision principle under the control of the task parameter set, define nodes and actionable edges, and perform weighted random scheduling and deployment of emotion images according to the emotion valence ratio. At the same time, it assigns rewards at node or path positions according to the reward parameters and random seeds to form a decision environment that couples stimuli and rewards. The behavior collection and step sequence association module is used to first conduct an exploration phase based on the decision environment to record the mouse position, exploration time and movement path, and then conduct an execution phase to record the key press time, possible action direction and reward obtained, and associate the exploration phase and execution phase with the step sequence to form sequential decision behavior data. The sequence control and process advancement module is used to control the task presentation order based on Latin square balance and advance the process when triggered by the user. The data archiving and export module is used to archive the sequence decision behavior data in the form of binary files and text files for subsequent strategy reconstruction and statistical analysis.

[0052] The original Mouselab-MDP program only displays monetary amounts and blank card backgrounds, making it suitable only for studying strategy formation and selection in sequential decision-making processes. This method for explicit sequential decision-making studies in highly neurotic individuals improves upon the image presentation of the original Mouselab-MDP by introducing four emotional scenarios: sadness, anger, happiness, and neutrality, along with two gain probabilities: high gain and high loss. Furthermore, the ratio of background images for each emotion and the gain probabilities are adjustable, allowing for convenient manipulation of the intensity of negative stimuli. This facilitates research into the mechanisms by which negative stimuli influence the decision-making strategy formation process in highly neurotic individuals.

[0053] Under the condition of a matrix map displaying money directly, within 20 seconds of map presentation, users only need to remember the monetary gains of each path on the map to plan their chosen path. However, during the path execution phase, blank images are replaced with facial expressions with high emotional valence. Theoretically, facial expressions are expected to evoke strong emotional arousal in individuals, with negative expressions evoking moderate levels of negative stimulation. Based on the hypothesis that highly neurotic individuals are more sensitive to negative stimuli, their decision-making strategies will be influenced by negative stimuli in this scenario. This manifests as poorer strategies and shorter reaction times when negative faces appear, as they are eager to escape the negative stimulus. This experimental procedure can track reaction times at each step and communicate with eye-tracking, EEG, and MRI equipment to track brain activity during corresponding behaviors, facilitating research into the decision-making characteristics and brain mechanisms of highly neurotic individuals in this scenario.

[0054] Under the condition of displaying a matrix map by hovering the mouse, participants initially exhibit different emotional faces. Highly neurotic individuals, due to their sensitivity to and avoidance of negative stimuli, will engage in less and less effective exploration when negative faces appear, thus gaining less benefit. The program records the mouse hover path and dwell time, facilitating the reconstruction of the decision-making strategy formation process of highly neurotic individuals.

[0055] In a center-radial map scenario, the four arms are configured with varying ratios of emotional faces and rewards. Specifically, one arm has the most positive emotional faces and a moderate reward, another arm has the highest reward and the most negative face combination, and the remaining two arms have moderate rewards. Highly neurotic individuals, due to their sensitivity to and avoidance of negative stimuli, tend to avoid the arm with the highest reward but the most negative reward, opting instead for the path with the most positive emotion but a moderate reward. Their hovering exploration time, path, and step-by-step reaction time are also recorded to facilitate the reconstruction of the strategy formation process.

[0056] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0057] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for collecting sequence decision-making strategies for highly neurotic individuals, characterized in that, include: Create a full-screen drawing canvas and load the task parameter set; The task parameter set includes at least the decision map type and depth, emotional valence ratio, reward parameters and random seeds, and presentation order control parameters. Collect user gender selection data and determine the binding relationship of the emotional image library for subsequent presentation, so that the stimulus set corresponds to user attributes; Under the control of the task parameter set, a gridded decision map is generated based on the Markov decision principle, nodes and actionable edges are defined, and emotional images are weighted and randomly scheduled according to the emotional valence ratio. At the same time, rewards are assigned at node or path positions according to the reward parameters and random seeds, forming a decision environment that couples stimuli and rewards. Based on the aforementioned decision-making environment, an exploration phase is first conducted to record mouse position, exploration time, and movement path. Then, an execution phase is conducted to record key press time, possible action directions, and rewards obtained. The exploration phase and the execution phase are linked in sequence to form sequential decision-making behavior data. The task presentation order is controlled based on Latin square balance, and the process is advanced upon user triggering. The sequence decision-making behavior data is archived in both binary and text file formats for subsequent strategy reconstruction and statistical analysis. Under the control of the task parameter set, a gridded decision map is generated based on the Markov decision principle, nodes and actionable edges are defined, and emotion images are weighted and randomly scheduled according to the emotion valence ratio. Simultaneously, rewards are assigned at node or path positions based on the reward parameters and random seeds, forming a decision environment that couples stimuli and rewards, including: Call the map type and depth parameters in the task parameter set, generate a matrix map or a central radial four-arm map based on the Markov decision principle, and define the coordinates, actionable edges and target nodes of each node. Based on the emotional valence ratio in the task parameter set, a cumulative probability distribution vector is constructed, and weighted random sampling is performed on four types of images: sad, angry, happy, and neutral. The sampling results are then distributed to the node positions. Based on the reward parameters and random seed in the task parameter set, a reward is assigned to the node or path position. Specifically, the expected reward is calculated using the following value function and probability weight function: in, The expected reward for a node or path location; This refers to the actual reward value displayed at the node or path location; To obtain reward values The objective probability; The value function is constructed based on prospect theory; This is a probability weighting function based on nonlinear probability sensing; By combining the results of emotion setting and reward distribution, a complete decision-making environment that couples emotional stimuli and reward information is generated and linked to the exploration and execution phases; The value function and probability weight function in the formula are defined as follows: ; in, The concavity parameter of the revenue curve, The convexity parameter of the loss curve, The loss aversion coefficient, The curvature parameter of the probability weighting function; This represents the actual reward value for a single node or path location. To obtain reward values The objective probability; the function is used to calculate the expected reward. This completes the assignment of reward values ​​to node or path locations.

2. The sequence decision-making strategy acquisition method for highly neurotic individuals according to claim 1, characterized in that, Create a full-screen drawing canvas and load the task parameter set, including: Call the Screen function in the Psychtoolbox package to set the computer screen as a full-screen drawing canvas and initialize the display resolution parameters to achieve multi-resolution compatibility; The task prompt text is drawn in the center on the full-screen drawing canvas, and users can customize the font, font size, color and background highlight parameters to ensure the correct presentation of the task prompt information; When loading the task parameter set, the decision map type and depth, emotional valence ratio, reward parameters, random seed and presentation order control parameters are uniformly preset and bound to the experimental process after the canvas is initialized. A missing value detection and display environment recovery mechanism is implemented during the canvas drawing process to ensure that the system default display settings can be restored in the event of an abnormal exit.

3. The sequence decision-making strategy acquisition method for highly neurotic individuals according to claim 1, characterized in that, Collect user gender selection data and determine the binding relationship of the emotion image library for subsequent presentation, so that the stimulus set corresponds to user attributes, including: Display the optional gender text entry in the center of the full-screen drawing canvas; The system receives the user's gender selection via mouse clicks and collects the gender text from the computer. The system uses a conditional statement to access an emotion image library corresponding to the collected gender. The emotional image library is linked to subsequent task flows to ensure that the stimulus set corresponds to user attributes.

4. The sequence decision-making strategy acquisition method for highly neurotic individuals according to claim 1, characterized in that, Based on the aforementioned decision-making environment, an exploration phase is first conducted to record mouse position, exploration time, and movement path. Then, an execution phase is conducted to record key presses, possible action directions, and rewards. The exploration and execution phases are linked sequentially to form sequential decision-making behavior data, including: During the exploration phase, the mouse position is read cyclically at a sampling frequency of no less than 60Hz, and a timestamp and event marker are recorded for each frame. During the execution phase, a non-blocking key detection and debouncing mechanism is employed to record the start / end time, possible action direction, and corresponding reward value for each key press, while simultaneously maintaining a progressively updated cumulative result. Based on timestamps, the exploration samples are compared with the first... The sequence of key press events is associated with the first key press event, defining the relationship with the first key press event. The set of exploration samples adjacent to the time of the next key press is used to calculate the exploration intensity feature value corresponding to the current step. The formula is: ;in, For the first Average displacement intensity within the window before step execution; In order to be with the first Key press time The corresponding set of exploration sample indices; For the first The timestamp of the next keystroke; This is a constant representing the duration of the time window used for pre-step association; For the first The timestamp of the frame exploration sample, and Sampling frequency and For the first Frame mouse position vector; , The first The horizontal and vertical coordinates of the mouse cursor in the screen coordinate system; Represents the Euclidean norm; The exploration intensity characteristic value of each step, the associated key press time, action direction and reward value are organized into a recording unit according to the step sequence and written into a machine-readable file, which includes at least two formats: binary data file and tab-delimited text file, for subsequent strategy reconstruction and statistical analysis.

5. The sequence decision-making strategy acquisition method for highly neurotic individuals according to claim 1, characterized in that, The task presentation order is controlled based on Latin square balance, and the process is advanced upon user triggering, including: Obtain the number of experimental blocks and load or generate a Latin square matrix of the same order, and determine the presentation order framework across subjects based on the Latin square balance principle. The Latin square is subjected to row and column permutations based on a random seed in the task parameter set to obtain the presentation sequence corresponding to the current subject, and the presentation sequence is written into the sequence control parameter port for process scheduling. After each block is completed, the user input listening state is entered, the next step prompt message is displayed and the preset start button is listened for; Upon receiving user-triggered input, the timestamp is recorded and the process proceeds to the next block according to the presented sequence until the sequence traversal is complete, thus ending the task flow.

6. The sequence decision-making strategy acquisition method for highly neurotic individuals according to claim 1, characterized in that, The sequence decision-making behavior data is archived in both binary and text file formats for subsequent policy reconstruction and statistical analysis, including: Based on a unified timestamp, the data of the exploration phase and the execution phase are matched and sorted step by step to generate a record unit containing timestamp, mouse position, event marker, key press time, action direction, reward value and cumulative reward; Complex data structures are managed using cell arrays, and the I / O write interface is called to save the record units as .mat binary data files; The record unit is expanded into tab-delimited text in a preset column order and written to a .txt file to achieve double archiving corresponding to the binary data; After archiving is completed, the data writing process for this block ends, and we are ready to proceed to the subsequent data processing and analysis stage.

7. A sequence decision-making strategy acquisition system for highly neurotic individuals, characterized in that, For implementing the sequence decision strategy acquisition method as described in any one of claims 1 to 6, the sequence decision strategy acquisition system comprises: The canvas initialization and parameter loading module is used to create a full-screen drawing canvas and load the task parameter set; the task parameter set includes at least the decision map type and depth, emotion valence ratio, reward parameter and random seed, and presentation order control parameter. The gender collection and image library binding module is used to collect the user's gender selection and determine the binding relationship of the emotion image library for subsequent presentation, so that the stimulus set corresponds to the user's attributes. The decision environment construction module is used to generate a gridded decision map based on the Markov decision principle under the control of the task parameter set, define nodes and actionable edges, and perform weighted random scheduling and deployment of emotion images according to the emotion valence ratio. At the same time, it assigns rewards at node or path positions according to the reward parameters and random seeds to form a decision environment that couples stimuli and rewards. The behavior collection and step sequence association module is used to first conduct an exploration phase based on the decision environment to record the mouse position, exploration time and movement path, and then conduct an execution phase to record the key press time, possible action direction and reward obtained, and associate the exploration phase and execution phase with the step sequence to form sequential decision behavior data. The sequence control and process advancement module is used to control the task presentation order based on Latin square balance and advance the process when triggered by the user. The data archiving and export module is used to archive the sequence decision behavior data in the form of binary files and text files for subsequent strategy reconstruction and statistical analysis.

Citation Information

Patent Citations

  • Modular neurocognitive function test method for drug addiction evaluation

    CN113633256A

  • Unsupervised continuous emotion electroencephalogram analysis method and device based on deep reinforcement learning

    CN119474948A