A human-computer interaction method based on decision maker's preference
Through neural network training methods, precisely capturing decision makers' preferences and selecting the best decision results, solving the problem of difficult to accurately capture decision makers' preferences and strategies in the existing technology, and achieving efficient strategy selection and improvement in decision quality.
Patent Information
- Application Number
- CN202510038611.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-10
AI Technical Summary
When existing human-computer interaction methods deal with complex multi-agent decision-making problems, it is difficult to accurately capture decision makers' preferences, resulting in inconsistent optimal strategies, poor generalization performance, and unable to adapt well to new problem scenarios.
The neural network training method is adopted to generate sample data and train the neural network, output the preference scores and weights of each time step, perform combination preference predictions, and calculate the loss function update parameters to accurately capture decision makers' preferences and select the best decision result.
It realizes precise capture and quantification of decision makers' preferences, and can choose the most consistent with decision makers' expectations among a variety of strategies, improves the degree of personalization of the strategy and the quality of decision making, and is suitable for a variety of practical problems.
Smart Images

Figure CN119440264B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human-computer interaction, and in particular relates to a human-computer interaction method based on decision maker preference. Background Art
[0002] Human-computer interaction methods play an important role and significance in the field of artificial intelligence. They help computer systems understand, store and process human knowledge, and enable computers to use human preferences to assist in decision-making based on pure machine decision-making. In addition, human-computer interaction methods can take human knowledge and experience into account in the problem-solving process, thereby providing strategies that are more in line with human ideas, and help computers extract meaningful key information from complex relationships and network data, and realize intelligent data analysis and decision support tasks. Human-computer interaction methods are widely used in many practical problems, including hotel robot cooperative delivery and water resource intelligent configuration and scheduling.
[0003] For some complex multi-agent decision-making problems, agents need to constantly adjust their strategies to cope with changes in their opponents, forming a dynamic game process. This mutual suppression mechanism makes the final solution not only the optimal result for the individual, but also the equilibrium result of group interaction. Therefore, the game cycle suppression generated by pure machine decision-making leads to the situation where the optimal strategy given by the agent is not unique. When solving such complex decision-making problems, it is very important to select the strategy that best meets the decision maker's preferences from a series of equilibrium strategies.
[0004] There are many academic studies on the solution methods of human-computer interaction, such as taking human thinking into account in the decision-making process through reference point selection, pre-set rules and real-time intervention. Therefore, this technology covers the development of multiple fields and provides theoretical and methodological support for the new paradigm of collaborative solution between humans and machines.
[0005] However, the above methods are too limited in their representation and consideration of human thoughts, often using artificially formulated rules to characterize the preferences of decision makers, while ignoring the intrinsic relationship between decision makers' preferences and strategies. At the same time, the above methods are too dependent on prior expert knowledge and have poor generalization performance. When applied to new problem scenarios, they often cannot well represent the intentions of decision makers, thus affecting the effect of human-computer interaction solutions. Summary of the invention
[0006] In view of this, the present invention provides a human-computer interaction method based on decision maker preference, which can accurately capture and quantify the decision maker's preference for different strategies and select the best decision result.
[0007] In order to achieve the above object, the technical solution of the present invention is:
[0008] A human-computer interaction method based on decision maker preferences, the specific process is:
[0009] Sample data generation: Interact with the simulation environment to generate trajectory data; randomly intercept trajectory data of a set length, and combine them in pairs to annotate the decision maker's preference labels to form sample data for training;
[0010] Neural network training: Use the sample data to train the neural network, and the neural network outputs each time step corresponding preference scores and corresponding weights; based on the preference scores and corresponding weights, respectively performing combined preference predictions; based on the preference prediction results, calculating the loss function and updating the neural network parameters;
[0011] Selection of the best strategy for human-computer interaction: Input several given equilibrium strategies into the simulation environment to obtain trajectory data, use the trained neural network to obtain the preference score and corresponding weight of the trajectory data, and further calculate the optimal strategy.
[0012] Furthermore, the combined preference prediction of the present invention includes:
[0013] Combined trajectory data Better than trajectory data Probability
[0014]
[0015] Combined trajectory data Better than trajectory data Probability
[0016]
[0017] in, and Represents trajectory data and trajectory data Time step The corresponding preference score, and Represents trajectory data and trajectory data Time step The corresponding preference weight.
[0018] Furthermore, the present invention calculates the loss function based on the preference prediction result as follows:
[0019] For each combination in the training data set, calculate its loss function ,
[0020]
[0021] in, for , Represents the expectation of all combinations in the dataset.
[0022] The total loss function is:
[0023]
[0024] in, Represents the total amount of training data.
[0025] Furthermore, the neural network parameters are updated as follows:
[0026]
[0027] in, Represents the parameters of the network model. Represents the learning rate.
[0028] Furthermore, the preference label of the present invention is expressed as , preference tags Expressing preference trajectory More than tracks ; Preference tags Expressing preference trajectory More than tracks ; Preference tags Representation trajectory With trajectory Same preferences.
[0029] Furthermore, the trajectory data of the present invention is a state-action pair tuple form, where Indicates Status information of the step, Indicates Step action information.
[0030] Furthermore, the neural network of the present invention comprises a linear mapping layer and a self-attention layer, wherein the linear mapping layer converts the state-action pair of each time step into Encoded as , the self-attention layer is used to encode , calculate the corresponding weight of preference and preference scores .
[0031] Furthermore, the linear mapping layer encoding of the present invention obtains for:
[0032]
[0033] in, represents the concatenation between vectors, and Represents a linear transformation matrix.
[0034] Furthermore, the self-attention layer of the present invention consists of three parts: query, key and value, and the corresponding linear transformation matrix is ;
[0035]
[0036] in, Represent the query, key, and value at each time step respectively;
[0037] Calculating preference weights for:
[0038]
[0039] in, is the normalized exponential function, is the hidden layer dimension;
[0040] Calculating preference scores for:
[0041]
[0042] Furthermore, the optimal strategy of the present invention is:
[0043] Calculate the score corresponding to each equilibrium strategy , select the strategy with the highest score as the optimal strategy
[0044]
[0045] in, is the length of the trajectory.
[0046] Beneficial effects:
[0047] First, in the training process of the neural network, the method of the present invention adopts a combination of two trajectory data for training, performs preference prediction based on the training results, and then realizes the calculation of the loss function. This method can make it necessary to only label the preference comparisons of the two trajectories in the process of generating sample data, thereby greatly reducing the workload of labeling the sample data preference labels.
[0048] Second, the present invention uses linear mapping layer and self-attention layer model design, and the neural network outputs the preference weight and preference score corresponding to each time step (i.e., key time frame segment), which can realize importance analysis of discretized trajectory sequences, provide decision makers with a more intuitive explanation, and provide auxiliary information for subsequent strategy adjustment and selection, thereby improving the interpretability and transparency of the method.
[0049] Third, by introducing a quantitative model of decision maker preferences, the decision maker's preferences for different strategies can be accurately captured and quantified. The present invention can select the solution that best meets the decision maker's expectations from a variety of possible strategies, which not only improves the personalization of the strategy, but also makes the decision-making process more in line with actual application needs.
[0050] Fourth, the present invention introduces the decision maker's preference into the decision-making process, combines neural networks with decision maker's preference, and effectively improves the quality of decision-making; the method proposed by the present invention can be widely applied to a variety of practical problems, such as intelligent services, intelligent manufacturing, intelligent logistics and other fields. Taking the application of the hotel service robot path planning scenario as an example, the present invention can improve decision-making efficiency and enhance user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A flow chart of the human-machine collaborative solution method based on decision maker preference provided by the present invention;
[0052] Figure 2 A structural diagram of the decision maker preference quantification model provided by the present invention;
[0053] Figure 3 This is a graph of the strategy importance analysis results provided by the present invention. DETAILED DESCRIPTION
[0054] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0055] It should be noted that the following embodiments and features in the embodiments may be combined with each other in the absence of conflict; and, based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making any creative work are within the scope of protection of the present disclosure.
[0056] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0057] The embodiment of the present application provides a human-computer interaction method based on decision maker preferences, establishes a human-computer collaborative solution mode with a human in the loop, can effectively capture the preference relationship of decision makers for different strategies, and can also score given strategies based on quantified human preferences, and finally provide a strategy that best meets the decision maker's preferences. The process of the method is as follows:
[0058] Sample data generation: Interact with the simulation environment to generate trajectory data; randomly intercept trajectory data of a set length, and combine them in pairs to annotate the decision maker's preference labels to form sample data for training;
[0059] Neural network training: Use the sample data to train the neural network, and the neural network outputs each time step corresponding preference scores and corresponding weights; based on the preference scores and corresponding weights, respectively performing combined preference predictions; based on the preference prediction results, calculating the loss function and updating the neural network parameters;
[0060] Selection of the best strategy for human-computer interaction: Input several given equilibrium strategies into the simulation environment to obtain trajectory data, use the trained neural network to obtain the preference score and corresponding weight of the trajectory data, and further calculate the optimal strategy.
[0061] This embodiment takes the hotel service robot path planning scenario as an implementation example. The goal of this scenario is to formulate a reasonable delivery route for the hotel robot, taking into account the personalized needs of customers while minimizing costs, and improving customer satisfaction. Figure 1 As shown, the method is implemented by the following steps:
[0062] Step 1: Interact with the simulation environment to realize data collection.
[0063] For the simulation environment corresponding to the specified problem, the agent interacts with the simulation environment with any strategy time steps, record the state and action information corresponding to each time step, and form a length of Each time step in the time series is The data is represented as state-action pairs tuple form, where Indicates Status information of the step, Indicates Step action information, .
[0064] For example, in the hotel service robot path planning scenario, the agent represents the hotel service robot, and the state at each time step is Including the robot's absolute position coordinates and surrounding obstacle information, action The simulation environment is a simulation environment loaded with a hotel map, and the hotel robot can simulate its own strategy in the simulation environment.
[0065] Step 2: Data preprocessing and training data set generation.
[0066] After acquiring the time series, randomly intercept a section of length The fragment is taken as trajectory one and is expressed as ; Randomly intercept a length of The fragment is taken as trajectory 2 and is expressed as Repeat the above steps. The training data set consists of There are two pairs of trajectories, and the length of each trajectory is .
[0067] For example, in the hotel service robot path planning scenario, the time series represents the movement trajectory of the hotel service robot in the hotel, which includes the robot's position, speed, and environmental information within the hotel.
[0068] Step 3: Label the training set and initialize the network parameters.
[0069] After generating the training data set, the decision maker performs Mark the preferred labels. , the above steps are repeated times, to achieve Preferences are given to all pairs of trajectories. Then the neural network to be trained is set and all its parameters are randomly initialized.
[0070] For example, in the path planning scenario of a hotel service robot, the trajectory of the hotel service robot can be visualized through simulation software. Decision makers can compare the performance of different trajectories and make decisions on the trajectories in the training data. Make preference judgments and annotations.
[0071] The specific implementation of the preference label in this step is: the preference label is expressed as , preference tags Expressing preference trajectory More than tracks ,Right now ; Preference tags Expressing preference trajectory More than tracks ,Right now ; and preference tags Representation trajectory With trajectory Same preferences.
[0072] The following steps 4 to 6 are an iterative process:
[0073] Step 4: Neural network training.
[0074] After getting a batch of examples, each trajectory in it is composed of state and actions composition, In order to transform the trajectory data into a high-dimensional space and facilitate the representation of features, a linear mapping is used to transform the state-action pair of each time step Encode, the corresponding output represents the hidden layer encoding is .
[0075] For example, in the hotel service robot path planning scenario, the hidden layer encoding A high-dimensional representation of the state and action information of the hotel robot at each time step.
[0076] like Figure 2 As shown, based on the hidden layer encoding at each time step , the self-attention layer is used to further explore the correlation between variables and predict the preference score and corresponding weight at each time step. Among them, the input of the self-attention layer is the hidden layer encoding , and the corresponding outputs are the preference scores of the current time step and the corresponding weights .
[0077] For example, in the hotel service robot path planning scenario, the preference score represents the degree of conformity between the performance of the hotel service robot at the current time step and the decision maker's preference, and the weight Represents the temporal importance of the current time step to the entire trajectory.
[0078] Step 5: Prediction of preference relations.
[0079] After step 4, the preference score corresponding to each trajectory can be obtained. and the corresponding weights Based on this, the preference relationship between trajectory pairs is predicted and the trajectory is calculated. Better than track Probability and trajectory Better than track Probability .
[0080] In this embodiment, non-Markov probability calculation is used, where the preferred trajectory Better than track The non-Markov probability expression of probability is:
[0081]
[0082] in, and The trajectory At time step The preference scores and corresponding preference weights under and The trajectory At time step The preference scores and corresponding preference weights under .
[0083] Step 6: Calculate the loss function and update the parameters.
[0084] Using cross entropy as the loss function, for each pair of trajectories in the dataset And corresponding labels , the mathematical form of its loss function is
[0085]
[0086] in, ,and Represents the batch size, so the mathematical form of the overall loss function is
[0087]
[0088] Calculating Losses The gradient of each parameter in the network is calculated and the network parameters are updated using the gradient back propagation mechanism. The mathematical form of gradient back propagation is
[0089]
[0090] in, Represents the parameters of the network model. Represents the learning rate.
[0091] Step 7: Determine whether the termination condition is met. If so, save all parameters of the current network as the final model parameters. Otherwise, obtain a new batch of examples and return to step 4 for the next iteration.
[0092] Step 8: Obtain the trajectory of the interaction between the equilibrium strategy and the simulation environment and input it into the neural network to obtain the preference score corresponding to each time step and the corresponding weights .
[0093] Given several equilibrium strategies, the strategies are interacted with the simulation environment respectively, and the corresponding lengths are obtained by sampling Load the saved network model parameters, input the sampled trajectory into the loaded preference network model, and the network model can give preference scores to the trajectory. The network model outputs each time step Corresponding preference scores and the corresponding weights , where the weight That is, it reflects the importance of the current time step, provides auxiliary decision-making information for decision makers, improves the reliability of decisions, and facilitates the selection and adjustment of subsequent strategies.
[0094] For example, in the hotel service robot path planning scenario, this step can provide decision makers with the most critical time frame in the trajectory segment, improving the interpretability and transparency of the method.
[0095] Step 9: Preference score for the given strategy. and the corresponding weights Perform weighted averaging:
[0096]
[0097] in, This indicates the degree to which the strategy is consistent with human preferences. is the length of the trajectory. Therefore, according to the preference score corresponding to each strategy, a final strategy that is more in line with human preferences can be selected from a series of equilibrium strategies.
[0098] For example, in the hotel service robot path planning scenario, the trajectory preference represents the degree of preference of human experts for the corresponding strategy. The higher the value, the more the corresponding strategy conforms to human preferences.
[0099] Furthermore, the specific implementation of the linear mapping of the input data in step 4 of the embodiment of the present application is: assuming that the input action and actions The corresponding dimensions are and , the dimension of the hidden code is , then we have linear transformation matrices and Mapping states and actions to dimensional feature space. For each time step state-action pair , its encoding The specific mathematical form is
[0100]
[0101] in, Indicates The encoding of the state-action pairs for time steps, Represents the concatenation between vectors.
[0102] Furthermore, the specific implementation of the self-attention layer in step 5 of the embodiment of the present application is:
[0103] (1) Construction of self-attention layer input.
[0104] The input of the self-attention layer is divided into three parts: query, key, and value, which are obtained by linear mapping the input encoding of each time step. The linear transformation matrix used in the above linear mapping is , the specific mathematical form of linear mapping is
[0105]
[0106] in, Represents the query, key, and value at each time step, respectively, for preference scoring and the corresponding weights Calculation.
[0107] (2) Calculation of preference weights. Perform a dot product attention operation on the query and the key to calculate the similarity between the query and the key. The specific mathematical form of the dot product attention operation is:
[0108]
[0109] in That is the The weight corresponding to the time step preference is is the normalized exponential function, and are the query and key corresponding to the time step, respectively. is the hidden layer dimension.
[0110] (3) Calculation of preference scores. time step, directly set the value of the current time step As the corresponding preference score Combined with the preference weights calculated in step 2, the final time step The corresponding weighted preference score can be calculated as .
[0111] The embodiment of the present application establishes a human-machine collaborative solution mode with people in the loop, which can effectively capture the preference relationship between decision makers for different strategies, and can also score given strategies based on quantified human preferences, and finally provide the strategy that best meets the decision maker's preferences. At the same time, the present invention can output the strategy that best meets the decision maker's preferences while giving the key time frame fragments of the strategy, providing a more intuitive explanation for the decision maker. This method can effectively introduce the decision maker's preferences into the decision-making process, help improve the efficiency and quality of decision-making, and can improve the overall benefits of the application scenario.
[0112] In order to further illustrate the effectiveness of the provided method, the human-machine collaborative solution method based on decision maker preferences provided by the present invention is simulated in the environment of a hotel map. Three different equilibrium strategies are provided in the simulation verification: Strategy 1, Strategy 2 and Strategy 3, and the above three strategies have different characteristics and advantages. Considering the three standards of timeliness, security and stability, the above strategies can be ranked under different indicators, as shown in Table 1.
[0113] Table 1
[0114]
[0115] According to the above steps 1 to 8, we can obtain the pre-trained preference network models corresponding to the three preference standards according to the preference standards in Table 1. Using the above pre-trained model to conduct a verification experiment on the validation set strategy, we can find that the preference ranking results of the validation set strategy are consistent with the preference standards given in advance. At the same time, the importance distribution given by the method is as follows Figure 3 As shown in the figure, the importance of a single trajectory in different time frames is different. The method assigns higher importance values to the key frames of decision-making, which improves the interpretability and transparency of the method.
[0116] Based on the above experiments, the human-computer collaborative solution method based on decision maker preferences provided by the present invention can effectively capture the decision maker's preference relationship for different strategies, and can also score given strategies based on quantified human preferences, and finally provide a strategy that best meets the decision maker's preferences, thereby improving decision-making efficiency and enhancing user satisfaction.
[0117] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A human-computer interaction method based on decision maker preference, characterized in that: The specific process is: Sample data generation: The agent interacts with the simulation environment time steps to interact and generate trajectory data, which is the state and action information of the agent; randomly intercept the set length The trajectory data of the decision makers are combined in pairs to label the decision makers’ preference labels, forming sample data for training; Neural network training: Use the sample data to train the neural network, and the neural network outputs each time step The corresponding preference scores and corresponding weights; Based on the preference scores and corresponding weights, respectively, combined preference predictions are performed; Based on the preference prediction results, calculating the loss function and updating the neural network parameters; Selection of the best strategy for human-computer interaction: Input several given equilibrium strategies into the simulation environment to obtain trajectory data, use the trained neural network to obtain the preference score and corresponding weight of the trajectory data, and further calculate the optimal strategy; The combined preference prediction includes: Combined trajectory data Better than trajectory data Probability Combined trajectory data Better than trajectory data Probability in, and Represents trajectory data and trajectory data Time step The corresponding preference score is the degree to which the agent's performance at the current time step conforms to the decision maker's preference. and Represents trajectory data and trajectory data Time step The corresponding preference weights; The trajectory data is a state-action pair tuple form, where Indicates Status information of the step, Indicates Step action information; The neural network includes a linear mapping layer and a self-attention layer. The linear mapping layer converts the state-action pair of each time step into Encoded as , the self-attention layer is used to encode , calculate the corresponding weight of preference and preference scores ; The linear mapping layer encoding obtains for: in, represents the concatenation between vectors, and represents a linear transformation matrix; The self-attention layer consists of three parts: query, key and value. The corresponding linear transformation matrix is ; in, Represent the query, key, and value at each time step respectively; Calculating preference weights for: in, is the normalized exponential function, is the hidden layer dimension; Calculating preference scores for: ; The optimal strategy is calculated as: Calculate the score corresponding to each equilibrium strategy , select the strategy with the highest score as the optimal strategy in, Indicates the degree to which the strategy is consistent with human preferences, is the length of the intercepted trajectory.
2. The human-computer interaction method based on decision maker preference according to claim 1, characterized in that: Based on the preference prediction result, the loss function is calculated as: For each combination in the training data set, calculate its loss function , in, for , Represents the expectation of all combinations in the data set; The total loss function is: in, Represents the total amount of training data.
3. The human-computer interaction method based on decision maker preference according to claim 2, characterized in that: The updated neural network parameters are: in, Represents the parameters of the network model. Represents the learning rate.
4. The human-computer interaction method based on decision maker preference according to claim 1, characterized in that: The preference label is represented as , preference tags Expressing preference trajectory More than tracks ; Preference tags Express preference trajectory More than tracks ; Preference tags Representation trajectory With trajectory Same preferences.
Citation Information
Patent Citations
Man-machine hybrid formation intelligent decision generation method based on diffusion model and feedback learning
CN118838164A